← All conversations
AI & ML

Julien Le Dem, Datadog, on Reliable AI — Interview with Alexy

Julien Le Dem ↗With Alexy KhrabrovNov 20253:13

Julien Le Dem, principal engineer at Datadog, on reliable AI at AI By the Bay 2025. A back-end engineer, he recently used Claude Code to build a small web app that visualizes Parquet metadata, out of his comfort zone. Fully reliable AI does not exist yet, everything must be triple-checked; if AI is to produce far more code than we can review, validation and verification frameworks must make sure it introduces no flaws or vulnerabilities. Problems whose solutions are easy to verify will be solved quickly; the harder-to-evaluate ones are where the risk lies.

Follow the words

Read the transcript

Hi, I'm Julian. Uh, I work at Data Dog Hi, I'm Julian. Uh, I work at Data Dog and I'm a principal engineer there. Oh, and I'm a principal engineer there. Oh, so recently I've been playing with uh so recently I've been playing with uh Cloud Code to build a little uh Cloud Code to build a little uh visualizers for park metadata. Um, and visualizers for park metadata. Um, and I'm more of a back-end engineer, so I I'm more of a back-end engineer, so I know very little about front end stuff. know very little about front end stuff. And I was able to make this little web And I was able to make this little web app with it. Was pretty proud of myself app with it. Was pretty proud of myself of uh building new things out of my of uh building new things out of my comfort zone using AI. Oh, that's a comfort zone using AI. Oh, that's a that's a vast question. Um, I think that's a vast question. Um, I think right now we don't really have fully right now we don't really have fully reliable AI, right? We need to triple reliable AI, right? We need to triple check everything. Um although we'll need check everything. Um although we'll need to have a better definition of what that to have a better definition of what that means in the future, right? Like that means in the future, right? Like that the result can be trusted uh that they the result can be trusted uh that they accomplish what we were expecting them accomplish what we were expecting them to do. So there'd be a lot of validation to do. So there'd be a lot of validation um um in you know if they're using tools or in you know if they're using tools or building things and so on right like building things and so on right like it's actually if the goal is the AI to it's actually if the goal is the AI to produce a lot more like code say than we produce a lot more like code say than we can build uh then it's going to be can build uh then it's going to be tricky to make sure it's not introducing tricky to make sure it's not introducing uh flaws or security uh vulnerabilities uh flaws or security uh vulnerabilities and things like that. Well, we need to and things like that. Well, we need to have uh good uh validation frameworks have uh good uh validation frameworks like verifications. How do we make sure like verifications. How do we make sure the solutions the solutions what's happening is according to what we what's happening is according to what we were expecting, right? If we're going to were expecting, right? If we're going to automate a bunch of things with AI. Um automate a bunch of things with AI. Um we'll need to have a much better we'll need to have a much better validation framework. validation framework. Um, Um, I think right now, you know, if you use I think right now, you know, if you use an LLM to do things, it it's pretty an LLM to do things, it it's pretty open-ended tool. So, it can do a lot of open-ended tool. So, it can do a lot of things you were not expecting it to do, things you were not expecting it to do, including the wrong thing. It's been including the wrong thing. It's been changing so fast in the next five years. changing so fast in the next five years. I don't know. I think the next six I don't know. I think the next six months is tr difficult to guess. Um, I months is tr difficult to guess. Um, I don't know. But I think it's all don't know. But I think it's all accelerating, right? So I guess the one accelerating, right? So I guess the one of the dimension is all the problems for of the dimension is all the problems for which it's easy to verify the solution which it's easy to verify the solution we'd be solved very quickly because then we'd be solved very quickly because then that's where you know AI can try a lot that's where you know AI can try a lot of things quickly and as long as we can of things quickly and as long as we can verify that it gives correct answers verify that it gives correct answers um we'll be able to use that very um we'll be able to use that very efficiently. for things that are harder efficiently. for things that are harder to evaluate whether they're correct, to evaluate whether they're correct, it's going to be more difficult or it's going to be more difficult or potentially uh unsafe, right? Um so we potentially uh unsafe, right? Um so we see what happen. I don't know what this see what happen. I don't know what this stack is going to look like, honestly.

Recovered English captions. Automatic transcription may contain errors.

Keep exploring

Follow the guest, their work, and the ideas behind this conversation in the Devreal knowledge graph.

Julien Le Dem on Devreal ↗
Independent by design

Your player.
Your subscription.

One permanent feed. Listen in the podcast app you love, with the conversations always at home here.

https://struct.fm/feed.xml
352 audio episodes available in the feed.