Notes from PyData Amsterdam 2026
Alexy Khrabrov's notes from PyData Amsterdam 2026, recorded after the trip. The LakeSail team met in person for the first time around Shehab Amin and Santosh Pingale's talk on Sail at Adyen, where Spark jobs that ran out of memory or never finished now run on a Rust engine behind the same PySpark API. Why the warehouse at the NDSM Loods forced Alexy to record his own audio. Christophe Blefari's keynote on the history of analytics, and why nao, his open-source analytics agent, is the user every lakehouse builder should design for. Ritchie Vink tracing the engineering arc of Polars from a better pandas to a company, and why Rust-native distributed data systems strengthen the whole ecosystem. Matt Topol on ADBC adoption, ADBC for DuckDB, and the ADBC Spark driver that lets Sail plug in anywhere Spark Connect is spoken; and Apache Magpie's tools for maintainers facing AI-generated pull requests. Plus the QueryGraph stack Alexy has been building on Sail: Grust graphs, TypeSec policies, LakeCat, Marciana memory, and Pinax.
Listen
Download MP3 ↓Follow the words
Read the transcript
Hello everybody, I'm Alexey Kuroborov, the head of ecosystems at Lakesail, the company which builds SAIL, the distributed ecosystem AI Native, built at Rust, replacing the whole Spark, Apache Spark stack with Rust Native Iceberg and Delta, and catalog support, and a new architecture for Spark jobs with PetroShuffles, with Object Store, and so much more. You should check it out at lakesail.com. So I am also the founder and organizer of some of the oldest, longest running, most technical engineering meetups in the world, usually positioned in the Bay Area, and started with SF Scala, and then Bay Area AI, and then AI Agent SF, and since July Rust.ai. So I focus now on this confluence, and always did, of strong types, efficient distributed systems, open source, and community. Now of course, everything is AI, all the data is feeding AI, every application is an AI application, as Carlos Gestrin famously said back in 2015. So the idea is not new, but finally, like with deep learning, we have the technologies, we have the means, and we have the agents, and we have the coding assistance, to actually help us build this to our liking. And I like it very, very much. So since May, you can check my GitHub. I built querygraph.ai, which is a stack of Rust Native graphs, and agentic secure protocols with fiber ledger and type level policies, and leak head, which is a minimal catalog getting out of the way between the engines and the data being the engine contract, while providing governance with typesac. And Marciana, which is a memory, which is secure, which is secure, and does not divulge, or it should not divulge, and everything stays within the boundary, declared by the type policy. And finally, querygraph stack as a whole, it runs on top of SAIL, it provides semantic croissant, and CDIF, and ODRL, which are policy languages, and ontology standards, and typesac is underpinning all of it, and there is a demo. And recently, while in Amsterdam, we've added PINACs, which is the enterprise data registry, which is a catalog of catalogs, it's called out of the initial index of Alexandria library. So now there is a whole querygraph stack. I built it all with support of my dear friends, Codex and Fable. What would I do without them? So, arm of all of that, you know, we come to PyData Amsterdam, where our talk on SAIL at ADN was accepted a while ago, so we made the whole team off-site out of it, the whole team flew from America, Asia, and EU, which was local. So the local guy drove and parked, and everybody else came from Shenzhen and the area, so, and Beijing. So, that was a great moment for us, and we really enjoyed viewing and seeing all our, you know, leaders and customers present. So, the talk itself was the second day. It went over the use cases for SAIL at ADN. It talked about the long-running jobs, huge jobs. ADN is one of the leading financial services providers in the world, processes a trillion dollars a year. And it was really, really instructive to see, and empowering, and exciting to see our users picking SAIL over SPARK when they run out of SPARK limits, when they hit the wall, when they go and get out of memory, or the run takes forever, or memory is wildly fluctuating, as it does during JVM garbage collection. So, it was really instructive, and really, again, empowering. And I recorded the whole talk, fortunately, because I always carry a bunch of cameras, and I'm an AV enthusiast, and a Leica photographer. So, I fortunately had my Lumix Leica L-mount compatible camera, and I did the full-frame recording, and the noise in the warehouse was really out of this world, because it's a giant warehouse, and the two rooms were separated by a curtain. So, our room had to actually wear headsets. However, I had my little DJIs, and I recorded them both from Santosh and GEOP, and we added the audio to the video, and I think it's palatable. So, that was great. But, of course, the talk itself was not the only thing that, I think it was a highlight, certainly, and the best thing at the conference for us, but there were multiple other amazing talks. For instance, Christophe Blefari is a data analyst who keynoted the second day, and he really comes from the world of people who use Lake Houses. And, again, I came to Lake Houses, I started the first Spark Meetup in the world at the end of 2011, at the Soskala, finding Matei and asking him to present Spark, and then invited him to create the very first Apache Spark Meetup, which we did in January 2012, and also ran it at Cloud, hosted it again, and I went away kind of 15 years, right, doing different things, different software engineering, and startup building, and open source science at IBM, and so forth. And so, when I come back, I just find this enormous amount of stuff which was built, and there is not always a rhyme and reason to it, right, there is a bunch of different things happening to it, and there are catalogs and semantic layers and whatnot, and this is all usually pushed from engineering and vendor perspectives. It's hard to understand why do we need all of that stuff, and it's clear to me that we don't need all these enormous catalogs, we don't need a bunch of this metadata machinery which exists for the sake of existing. So, how can we understand what is needed? And Christophe was fantastic because he presented the data analyst, where he gave the history of the warehouse going back, you know, to the 60s, right, and databases and basically all up cubes and things like that, and eventually arriving at the lake house and AI-driven and hopefully agentic lake house. Nobody knows what it is, but it will happen. The agents are roaming and foraging, getting into the lake house one way or the other. And so Christophe gave us a fantastic view of how people actually use the lake house. What do they actually do? How do they want to explore the data? And it was, I think, a kind of upgrade of something we know and love, exploratory data analysis, something tableau-like, but he ran his own presentation in his own kind of web app on the local server, and he could click on his own data and see tables and see displays and all of that he could build in a bespoke fashion. And so he has a company now called NowLabs, and GitHub work is GitNow, and the agent is now, GitNow slash now. And that agent encapsulates everything the data analysts actually do. So to me, it was a tremendous find because this is a customer, this is a user who will query your lake house. So if you're building a lake house, check this out, it's important that you should feed Now as your user and use Now as a sounding board for the things you want to do, because eventually Now will do that. Right? And so that's kind of one of the other amazing finds. And another interview, which I've done at Pidea Amsterdam, was with Richard Wink, the founder of Polars. And that was just a joy. It was a no-nonsense engineering interview. What struck me when I rejoined this community, right? And from Agentic AI, Hive, for a year and a half, I went back to distributed systems, my mainstay in using Rust and using stronger-type languages, you know, and Apache Data Fusion, Apache Arrow. It was extremely refreshing to hear Richie basically trace an arc of an engineering startup. So he built Polars as a better Pandas, you know, for himself. And he worked on it for several years before he made the company out of it. So everything is driven by the engineering questions. And the single node was optimized. The strategy is very simple. If you need distributed systems, you need a cluster, then you have to have a commercial version. An open-source version is a single node. And Richard just is such a cool guy. He just explored, kind of explained how he went about it. The very natural engineering trajectory, right? And how he made his trade-offs. So definitely check out my interview with Richie. He is local. He is now the king of data frames. Because the other local company, TAC-DB, was famously acquired by the AWS. So Richie Wink is the local homemade purveyor of data frames in Amsterdam. If you are looking for something local in that space. And he was genuinely impressed by our partnership with ADN. So I think it's amazing, right? That our users are also local. And they have reasons to work of sale. But it does not, in any way, for me, increase the interest of the whole ecosystem, right? I think the fact that people are using Rust to build distributed data systems and use Data Fusion and Apache Arrow and the whole set of asynchronous technologies that Rust provides and the whole set of ideas that Rust brings. I think you will find, definitely, that it strengthens our whole ecosystem. So we invited Richie to speak at Rust.i Ametab in the Bay Area. And hopefully he will make it over here. Because it turns out that majority of the customers are in North America. And my third interview was with Matt Topol. Matt Topol became my buddy. I met him once I moved back into distributed systems now with Rust. And I noticed that ADBC is hot. And Matt and Ian, the founders, are everywhere. I met Matt at Category VC party after AI Council. And I think Wes McKinney was there doing a fireside chat with Linhualat Jean from Coca Index, another Rust AI startup. And I was really impressed by what ADBC does. It's a very universal approach. When I started building with sale and wanted to bring some JDBC thing, translate the spark, piece of code from Alex Merced's book. Of course, there is no JDBC, but it helpfully kind of offered me to bring in JDBC. Actually, it didn't. I asked my friends. But then I did a pull request, which will invite you to import ADBC if you want to do some JDBC work. So ADBC is a universal driver for Rust and Arrow. Right? And obviously, we're columnar. Columnar databases. We are columnar stream processors. And columnar is the way. So Matt has fantastic plans. I kind of got a glimpse of them. And I will not talk about them until columnar presents them. But what Matt could talk about and did talk about is ADBC adoption and ADBC for Duck, DB. And I learned that they have ADBC for Spark, which because SAIL is Spark compatible on the wire with Spark Connect protocol, allows us to hook up SAIL anywhere the ADBC Spark driver is deployed. So that opens enormous avenues of advance. And Matt talks very specifically about various points in the ecosystem around ADBC. But also Matt is a participant in the new Apache protocol, MagPy. So he is a PMC in Apache projects. And so as a seasoned open source leader, understands the need for maintainers to deal with the deluge of AI PRs, AI slope, basically coming the way of maintainers. So MagPy is a set of tools. I think it started by the Airflow folks to manage that deluge and understand what's going on. And I think the approach is very solid. If you are a reasonable person and thoughtful person, then you can use AI. But you stand behind AI, you understand what it does. And basically, you know, you became a 10x programmer with AI, but you have to be 10x accountable programmer with AI. And so MagPy is a set of tools which will help maintainers ensure that. So that's another topic of our interview among others. And check that out again on Struck.fm Structured Output podcast held in the Structured Output bar. Come and join us. I would love to host you if you have some interesting things to say. If you build stuff, I would really love to have you present your stuff. And if it's in the Rust AI ecosystem, especially if it's in the agentic world, please let me know. And until then, this is Alexi Krober from the Structured Output bar. Come and join us.
Recovered English captions. Automatic transcription may contain errors.
Keep exploring
Follow the event, the talks, and the people these notes are about in the Devreal knowledge graph.
PyData Amsterdam on Devreal ↗