WEBVTT
Kind: captions
Language: en

00:00:00.000 --> 00:00:28.300
Hello everybody, I'm Alexey Kuroborov, the head of ecosystems at Lakesail, the company which builds SAIL, the distributed ecosystem AI Native, built at Rust, replacing the whole Spark, Apache Spark stack with Rust Native Iceberg and Delta, and catalog support, and a new architecture for Spark jobs with PetroShuffles, with Object Store, and so much more.

00:00:28.300 --> 00:00:50.860
You should check it out at lakesail.com. So I am also the founder and organizer of some of the oldest, longest running, most technical engineering meetups in the world, usually positioned in the Bay Area, and started with SF Scala, and then Bay Area AI, and then AI Agent SF, and since July Rust.ai.

00:00:50.860 --> 00:01:00.560
So I focus now on this confluence, and always did, of strong types, efficient distributed systems, open source, and community.

00:01:00.680 --> 00:01:09.840
Now of course, everything is AI, all the data is feeding AI, every application is an AI application, as Carlos Gestrin famously said back in 2015.

00:01:09.840 --> 00:01:24.580
So the idea is not new, but finally, like with deep learning, we have the technologies, we have the means, and we have the agents, and we have the coding assistance, to actually help us build this to our liking.

00:01:24.580 --> 00:01:50.000
And I like it very, very much. So since May, you can check my GitHub. I built querygraph.ai, which is a stack of Rust Native graphs, and agentic secure protocols with fiber ledger and type level policies, and leak head, which is a minimal catalog getting out of the way between the engines and the data being the engine contract, while providing governance with typesac.

00:01:50.000 --> 00:02:02.980
And Marciana, which is a memory, which is secure, which is secure, and does not divulge, or it should not divulge, and everything stays within the boundary, declared by the type policy.

00:02:03.560 --> 00:02:19.520
And finally, querygraph stack as a whole, it runs on top of SAIL, it provides semantic croissant, and CDIF, and ODRL, which are policy languages, and ontology standards, and typesac is underpinning all of it, and there is a demo.

00:02:20.000 --> 00:02:33.660
And recently, while in Amsterdam, we've added PINACs, which is the enterprise data registry, which is a catalog of catalogs, it's called out of the initial index of Alexandria library.

00:02:33.660 --> 00:02:36.660
So now there is a whole querygraph stack.

00:02:36.660 --> 00:02:37.160
So now there is a whole querygraph stack.

00:02:37.160 --> 00:02:42.080
I built it all with support of my dear friends, Codex and Fable.

00:02:42.960 --> 00:02:44.580
What would I do without them?

00:02:45.460 --> 00:02:55.940
So, arm of all of that, you know, we come to PyData Amsterdam, where our talk on SAIL at ADN was accepted a while ago,

00:02:55.940 --> 00:03:04.820
so we made the whole team off-site out of it, the whole team flew from America, Asia, and EU, which was local.

00:03:04.820 --> 00:03:13.160
So the local guy drove and parked, and everybody else came from Shenzhen and the area, so, and Beijing.

00:03:13.160 --> 00:03:24.580
So, that was a great moment for us, and we really enjoyed viewing and seeing all our, you know, leaders and customers present.

00:03:24.580 --> 00:03:27.580
So, the talk itself was the second day.

00:03:28.120 --> 00:03:32.780
It went over the use cases for SAIL at ADN.

00:03:33.240 --> 00:03:37.220
It talked about the long-running jobs, huge jobs.

00:03:37.340 --> 00:03:44.040
ADN is one of the leading financial services providers in the world, processes a trillion dollars a year.

00:03:44.040 --> 00:03:58.420
And it was really, really instructive to see, and empowering, and exciting to see our users picking SAIL over SPARK when they run out of SPARK limits,

00:03:58.420 --> 00:04:11.300
when they hit the wall, when they go and get out of memory, or the run takes forever, or memory is wildly fluctuating, as it does during JVM garbage collection.

00:04:11.300 --> 00:04:17.360
So, it was really instructive, and really, again, empowering.

00:04:18.080 --> 00:04:25.600
And I recorded the whole talk, fortunately, because I always carry a bunch of cameras, and I'm an AV enthusiast, and a Leica photographer.

00:04:26.520 --> 00:04:33.760
So, I fortunately had my Lumix Leica L-mount compatible camera, and I did the full-frame recording,

00:04:33.760 --> 00:04:40.500
and the noise in the warehouse was really out of this world, because it's a giant warehouse,

00:04:40.500 --> 00:04:43.420
and the two rooms were separated by a curtain.

00:04:43.960 --> 00:04:46.040
So, our room had to actually wear headsets.

00:04:46.520 --> 00:04:52.400
However, I had my little DJIs, and I recorded them both from Santosh and GEOP,

00:04:52.500 --> 00:04:56.680
and we added the audio to the video, and I think it's palatable.

00:04:57.520 --> 00:04:58.600
So, that was great.

00:04:58.600 --> 00:05:06.640
But, of course, the talk itself was not the only thing that, I think it was a highlight, certainly,

00:05:06.860 --> 00:05:12.420
and the best thing at the conference for us, but there were multiple other amazing talks.

00:05:12.520 --> 00:05:18.700
For instance, Christophe Blefari is a data analyst who keynoted the second day,

00:05:18.700 --> 00:05:22.840
and he really comes from the world of people who use Lake Houses.

00:05:22.840 --> 00:05:29.220
And, again, I came to Lake Houses, I started the first Spark Meetup in the world at the end of 2011,

00:05:30.120 --> 00:05:32.500
at the Soskala, finding Matei and asking him to present Spark,

00:05:32.560 --> 00:05:37.840
and then invited him to create the very first Apache Spark Meetup, which we did in January 2012,

00:05:37.840 --> 00:05:46.040
and also ran it at Cloud, hosted it again, and I went away kind of 15 years, right,

00:05:46.100 --> 00:05:49.460
doing different things, different software engineering, and startup building,

00:05:50.280 --> 00:05:53.420
and open source science at IBM, and so forth.

00:05:53.900 --> 00:05:59.240
And so, when I come back, I just find this enormous amount of stuff which was built,

00:05:59.940 --> 00:06:02.660
and there is not always a rhyme and reason to it, right,

00:06:02.660 --> 00:06:05.380
there is a bunch of different things happening to it,

00:06:05.380 --> 00:06:08.540
and there are catalogs and semantic layers and whatnot,

00:06:08.800 --> 00:06:14.260
and this is all usually pushed from engineering and vendor perspectives.

00:06:14.360 --> 00:06:16.680
It's hard to understand why do we need all of that stuff,

00:06:16.740 --> 00:06:20.020
and it's clear to me that we don't need all these enormous catalogs,

00:06:20.060 --> 00:06:24.640
we don't need a bunch of this metadata machinery which exists for the sake of existing.

00:06:25.300 --> 00:06:27.260
So, how can we understand what is needed?

00:06:27.360 --> 00:06:31.780
And Christophe was fantastic because he presented the data analyst,

00:06:31.780 --> 00:06:37.520
where he gave the history of the warehouse going back, you know, to the 60s, right,

00:06:37.640 --> 00:06:42.220
and databases and basically all up cubes and things like that,

00:06:42.340 --> 00:06:48.580
and eventually arriving at the lake house and AI-driven and hopefully agentic lake house.

00:06:48.700 --> 00:06:50.620
Nobody knows what it is, but it will happen.

00:06:51.320 --> 00:06:56.360
The agents are roaming and foraging, getting into the lake house one way or the other.

00:06:56.360 --> 00:07:03.400
And so Christophe gave us a fantastic view of how people actually use the lake house.

00:07:03.520 --> 00:07:04.680
What do they actually do?

00:07:04.800 --> 00:07:06.300
How do they want to explore the data?

00:07:06.380 --> 00:07:10.440
And it was, I think, a kind of upgrade of something we know and love,

00:07:10.560 --> 00:07:13.900
exploratory data analysis, something tableau-like,

00:07:14.560 --> 00:07:19.440
but he ran his own presentation in his own kind of web app on the local server,

00:07:19.440 --> 00:07:25.900
and he could click on his own data and see tables and see displays and all of that he could build in a bespoke fashion.

00:07:26.160 --> 00:07:30.460
And so he has a company now called NowLabs, and GitHub work is GitNow,

00:07:30.720 --> 00:07:33.680
and the agent is now, GitNow slash now.

00:07:33.960 --> 00:07:39.020
And that agent encapsulates everything the data analysts actually do.

00:07:39.440 --> 00:07:42.260
So to me, it was a tremendous find because this is a customer,

00:07:42.700 --> 00:07:45.600
this is a user who will query your lake house.

00:07:45.600 --> 00:07:51.360
So if you're building a lake house, check this out, it's important that you should feed Now as your user

00:07:51.360 --> 00:07:55.680
and use Now as a sounding board for the things you want to do,

00:07:55.960 --> 00:07:58.000
because eventually Now will do that.

00:08:02.580 --> 00:08:03.060
Right?

00:08:03.060 --> 00:08:06.980
And so that's kind of one of the other amazing finds.

00:08:06.980 --> 00:08:16.080
And another interview, which I've done at Pidea Amsterdam, was with Richard Wink, the founder of Polars.

00:08:16.720 --> 00:08:18.140
And that was just a joy.

00:08:18.340 --> 00:08:21.820
It was a no-nonsense engineering interview.

00:08:22.260 --> 00:08:26.660
What struck me when I rejoined this community, right?

00:08:26.740 --> 00:08:29.160
And from Agentic AI, Hive, for a year and a half,

00:08:29.160 --> 00:08:36.420
I went back to distributed systems, my mainstay in using Rust and using stronger-type languages,

00:08:37.200 --> 00:08:39.220
you know, and Apache Data Fusion, Apache Arrow.

00:08:39.660 --> 00:08:45.440
It was extremely refreshing to hear Richie basically trace an arc of an engineering startup.

00:08:45.440 --> 00:08:50.000
So he built Polars as a better Pandas, you know, for himself.

00:08:50.580 --> 00:08:54.380
And he worked on it for several years before he made the company out of it.

00:08:54.380 --> 00:08:57.020
So everything is driven by the engineering questions.

00:08:57.220 --> 00:08:58.520
And the single node was optimized.

00:08:59.220 --> 00:09:00.240
The strategy is very simple.

00:09:00.320 --> 00:09:02.220
If you need distributed systems, you need a cluster,

00:09:03.000 --> 00:09:04.460
then you have to have a commercial version.

00:09:04.820 --> 00:09:06.160
An open-source version is a single node.

00:09:07.740 --> 00:09:09.880
And Richard just is such a cool guy.

00:09:09.960 --> 00:09:14.040
He just explored, kind of explained how he went about it.

00:09:16.540 --> 00:09:18.940
The very natural engineering trajectory, right?

00:09:18.980 --> 00:09:20.360
And how he made his trade-offs.

00:09:21.240 --> 00:09:23.740
So definitely check out my interview with Richie.

00:09:24.380 --> 00:09:25.080
He is local.

00:09:25.240 --> 00:09:28.100
He is now the king of data frames.

00:09:28.280 --> 00:09:35.120
Because the other local company, TAC-DB, was famously acquired by the AWS.

00:09:35.680 --> 00:09:41.560
So Richie Wink is the local homemade purveyor of data frames in Amsterdam.

00:09:42.060 --> 00:09:46.380
If you are looking for something local in that space.

00:09:46.940 --> 00:09:52.000
And he was genuinely impressed by our partnership with ADN.

00:09:52.000 --> 00:09:54.020
So I think it's amazing, right?

00:09:54.100 --> 00:09:56.620
That our users are also local.

00:09:56.620 --> 00:09:59.760
And they have reasons to work of sale.

00:09:59.840 --> 00:10:07.300
But it does not, in any way, for me, increase the interest of the whole ecosystem, right?

00:10:07.300 --> 00:10:23.420
I think the fact that people are using Rust to build distributed data systems and use Data Fusion and Apache Arrow and the whole set of asynchronous technologies that Rust provides and the whole set of ideas that Rust brings.

00:10:23.420 --> 00:10:31.000
I think you will find, definitely, that it strengthens our whole ecosystem.

00:10:31.000 --> 00:10:34.940
So we invited Richie to speak at Rust.i Ametab in the Bay Area.

00:10:35.180 --> 00:10:37.340
And hopefully he will make it over here.

00:10:37.380 --> 00:10:41.380
Because it turns out that majority of the customers are in North America.

00:10:41.380 --> 00:10:45.000
And my third interview was with Matt Topol.

00:10:45.240 --> 00:10:46.860
Matt Topol became my buddy.

00:10:47.420 --> 00:10:51.960
I met him once I moved back into distributed systems now with Rust.

00:10:52.040 --> 00:10:54.160
And I noticed that ADBC is hot.

00:10:55.020 --> 00:10:57.700
And Matt and Ian, the founders, are everywhere.

00:10:58.340 --> 00:11:05.680
I met Matt at Category VC party after AI Council.

00:11:05.680 --> 00:11:14.620
And I think Wes McKinney was there doing a fireside chat with Linhualat Jean from Coca Index, another Rust AI startup.

00:11:15.380 --> 00:11:18.420
And I was really impressed by what ADBC does.

00:11:19.060 --> 00:11:20.500
It's a very universal approach.

00:11:20.840 --> 00:11:29.460
When I started building with sale and wanted to bring some JDBC thing, translate the spark, piece of code from Alex Merced's book.

00:11:29.460 --> 00:11:34.960
Of course, there is no JDBC, but it helpfully kind of offered me to bring in JDBC.

00:11:35.100 --> 00:11:35.600
Actually, it didn't.

00:11:36.000 --> 00:11:36.820
I asked my friends.

00:11:37.060 --> 00:11:44.380
But then I did a pull request, which will invite you to import ADBC if you want to do some JDBC work.

00:11:45.080 --> 00:11:47.780
So ADBC is a universal driver for Rust and Arrow.

00:11:48.540 --> 00:11:48.720
Right?

00:11:48.820 --> 00:11:50.460
And obviously, we're columnar.

00:11:50.900 --> 00:11:51.840
Columnar databases.

00:11:52.160 --> 00:11:54.600
We are columnar stream processors.

00:11:55.240 --> 00:11:56.900
And columnar is the way.

00:11:56.900 --> 00:11:59.540
So Matt has fantastic plans.

00:12:00.500 --> 00:12:02.140
I kind of got a glimpse of them.

00:12:02.480 --> 00:12:06.740
And I will not talk about them until columnar presents them.

00:12:07.140 --> 00:12:15.240
But what Matt could talk about and did talk about is ADBC adoption and ADBC for Duck, DB.

00:12:15.240 --> 00:12:32.540
And I learned that they have ADBC for Spark, which because SAIL is Spark compatible on the wire with Spark Connect protocol, allows us to hook up SAIL anywhere the ADBC Spark driver is deployed.

00:12:32.540 --> 00:12:36.540
So that opens enormous avenues of advance.

00:12:36.540 --> 00:12:43.840
And Matt talks very specifically about various points in the ecosystem around ADBC.

00:12:43.940 --> 00:12:49.000
But also Matt is a participant in the new Apache protocol, MagPy.

00:12:49.000 --> 00:12:53.740
So he is a PMC in Apache projects.

00:12:53.740 --> 00:13:09.260
And so as a seasoned open source leader, understands the need for maintainers to deal with the deluge of AI PRs, AI slope, basically coming the way of maintainers.

00:13:09.260 --> 00:13:10.620
So MagPy is a set of tools.

00:13:10.620 --> 00:13:19.080
I think it started by the Airflow folks to manage that deluge and understand what's going on.

00:13:19.560 --> 00:13:21.140
And I think the approach is very solid.

00:13:21.140 --> 00:13:27.560
If you are a reasonable person and thoughtful person, then you can use AI.

00:13:27.900 --> 00:13:30.640
But you stand behind AI, you understand what it does.

00:13:31.440 --> 00:13:39.500
And basically, you know, you became a 10x programmer with AI, but you have to be 10x accountable programmer with AI.

00:13:39.660 --> 00:13:43.660
And so MagPy is a set of tools which will help maintainers ensure that.

00:13:43.660 --> 00:13:49.320
So that's another topic of our interview among others.

00:13:49.780 --> 00:13:58.540
And check that out again on Struck.fm Structured Output podcast held in the Structured Output bar.

00:13:58.780 --> 00:13:59.600
Come and join us.

00:14:00.100 --> 00:14:03.820
I would love to host you if you have some interesting things to say.

00:14:04.080 --> 00:14:08.740
If you build stuff, I would really love to have you present your stuff.

00:14:08.740 --> 00:14:14.600
And if it's in the Rust AI ecosystem, especially if it's in the agentic world, please let me know.

00:14:15.440 --> 00:14:19.100
And until then, this is Alexi Krober from the Structured Output bar.

00:14:19.200 --> 00:14:19.840
Come and join us.

