← All conversations
AI & ML / Data systems

Jamieson Leibovitch, Uber, on Reliable AI — Interview with Alexy

Jamieson Leibovitch ↗With Alexy KhrabrovNov 20252:59

Jamieson Leibovitch, senior engineer at Uber, on reliable AI at AI By the Bay 2025. He led Query Copilot, Uber's text-to-SQL application for Presto SQL that lets users generate queries and reach insights faster, his first deep dive into AI. Reliable AI means agents that solve problems with high quality and can be trusted to take an action, which comes down to high-quality evaluations and an evaluation pipeline that shows whether each prompt change makes things better or worse. Uber went from no AI to basic tool-calling agents, static multi-agent systems, and now dynamic ones; the next step is task-management-oriented collaboration, unless the models get good enough to need none of it.

Follow the words

Read the transcript

Uh I'm Jameson. Uh I'm a senior engineer Uh I'm Jameson. Uh I'm a senior engineer with an Uber. Uh last year I was the the with an Uber. Uh last year I was the the lead for uh Uber's product called uh lead for uh Uber's product called uh Query Copilot. It was our our text to Query Copilot. It was our our text to SQL application for uh Prestos SQL. Um, SQL application for uh Prestos SQL. Um, it allows users to generate queries to it allows users to generate queries to get insights way faster than they would get insights way faster than they would have if they had written it themselves. have if they had written it themselves. Uh, this was my first uh real deep dive Uh, this was my first uh real deep dive into the agentic world. So to me, reliable AI is about having So to me, reliable AI is about having agents that uh answer questions agents that uh answer questions effectively or are able to solve the effectively or are able to solve the problem uh with high quality. So uh problem uh with high quality. So uh ensuring that I know that the agent is ensuring that I know that the agent is able to uh take an action and I can able to uh take an action and I can trust that it will do the right thing. trust that it will do the right thing. Uh a lot of this comes down to uh high Uh a lot of this comes down to uh high quality evaluations uh ensuring that the quality evaluations uh ensuring that the the actual score of the agent and you the actual score of the agent and you know we have confidence ahead of time know we have confidence ahead of time before we actually launch into before we actually launch into production. production. Um so uh I think I think one of the Um so uh I think I think one of the things that help really power reliable things that help really power reliable AI is a is a good evaluation pipeline as AI is a is a good evaluation pipeline as I think I previously mentioned. uh being I think I previously mentioned. uh being able to go from a series of questions, able to go from a series of questions, pass it to the agent, be able to iterate pass it to the agent, be able to iterate over it and see a see a trend. Uh for over it and see a see a trend. Uh for example, every time I change the prompt, example, every time I change the prompt, does it get better or get worse does it get better or get worse depending on my answers? Um I think as depending on my answers? Um I think as the community gets better, maybe the community gets better, maybe something uh either open source or um something uh either open source or um you know, maybe something allin-one that you know, maybe something allin-one that allows users uh or engineers or even allows users uh or engineers or even non-engineers to be able to build non-engineers to be able to build effective quality uh agents. Um yeah effective quality uh agents. Um yeah like in the last five five years now like in the last five five years now we've we've basically exploded from like we've we've basically exploded from like uh very basic agents uh such as like uh very basic agents uh such as like ChachiBT to even more complex ones. Um ChachiBT to even more complex ones. Um so again like we we've had especially at so again like we we've had especially at uh within our own company we've had no uh within our own company we've had no AI in the last 5 years to the explosion AI in the last 5 years to the explosion of AI in the last uh two or so years. Uh of AI in the last uh two or so years. Uh the models continuously get better. uh the models continuously get better. uh we've had uh single you know basic we've had uh single you know basic agents uh with some basic tool calling agents uh with some basic tool calling and data fetching uh to multi- aent and data fetching uh to multi- aent static systems to now multi-agent static systems to now multi-agent dynamic systems. Um I think the the next dynamic systems. Um I think the the next step at least I can see in the next year step at least I can see in the next year or two is uh more like task management or two is uh more like task management oriented being able to collaborate with oriented being able to collaborate with each other more effectively. But the the each other more effectively. But the the problem of course is this all hinges on problem of course is this all hinges on the models uh themselves. If the models the models uh themselves. If the models get insanely good that you don't even get insanely good that you don't even need any of this, this will change the need any of this, this will change the entire scene. Um, maybe we'll go back to entire scene. Um, maybe we'll go back to a single agent that's able to do a single agent that's able to do everything. Um, it just it's really hard everything. Um, it just it's really hard to determine uh based on the speed of to determine uh based on the speed of how how rapidly these these agents are how how rapidly these these agents are getting better.

Recovered English captions. Automatic transcription may contain errors.

Keep exploring

Follow the guest, their work, and the ideas behind this conversation in the Devreal knowledge graph.

Jamieson Leibovitch on Devreal ↗
Independent by design

Your player.
Your subscription.

One permanent feed. Listen in the podcast app you love, with the conversations always at home here.

https://struct.fm/feed.xml
352 audio episodes available in the feed.