High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Logan Kilpatrick: evaluation

20 May 2026 The Cognitive Revolution The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

“Like I don't know if like a lot of the agentic stuff we were landing at IO like would have been possible if we if we hadn't have had sort of some of that infrastructure standardization across the harness and the model delivery.”

— Logan Kilpatrick

Source trail

Everything needed to verify it.

Speaker
Logan Kilpatrick
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Roboflow is an end-to-end visual AI platform that lets you turn raw ideas into fully deployed applications in just hours, powering breakthroughs like Blueprint Pro's floor-plan understanding tool. Read the full Blueprint Pro story and see how over a million engineers are building the next wave of visual AI at https://roboflow.com Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr I think this is actually such a great story for us. I think very practically, Google has done a ton of this infrastructure standardization across the AI stack over the last couple of years, which I think has been awesome. And I actually do the story. It is like one of the threads of how we're able to land the Gemini 3 models across so many more products is actually because of this infrastructure standardization that happened. And so we've gotten a lot of, it's painful and difficult and there's of course lots of work involved in doing it. But if you sort of pay that cost, you actually do end up getting this. And I think the advice for for people who are in this position and sort of thinking about this is basically every 12 to 18 months now, like you have to rewrite everything from scratch. And so the best case is like you don't want N number of teams rewriting everything from scratch every time the paradigm shifts. And the example historically, the infrastructure was just like serving raw models and you'd get tokens in and you send tokens out. Now it's like there's a bunch of agentic infrastructure and there's tool loops and there's all these other things happening inside of the harness. And so again, you don't actually want, you want innovation, but you don't want every team to have to go and reinvent that from scratch. And so the fact that like, X team across Google who just wants to ship some really cool agentic product doesn't need to think about like the nuance of all the details of the tool calling loop, et cetera, is a huge acceleration for them to like just go focus on building a great product. And I think it's, hopefully we see that. Like I don't know if like a lot of the agentic stuff we were landing at IO like would have been possible if we if we hadn't have had sort of some of that infrastructure standardization across the harness and the model delivery. I think the other thing I would say as far as lessons learned is there's really no substitute for being able to just experiment and iterate quickly. So I think this goes to all of Logan's points about the foundation being strong, but I really think what has helped us is really being able to put in, for example, a new model iterate really quickly with a product on like, hey, what are the right prompts that would actually make this model viable for a different situation? What are the ways to kind of prototype really quickly with this model? What are the ways to get it in the hands of even just internal users quickly, let alone external users? And I think that is something that is now more and more possible with kind of layers that are consistent across the team. I think it's pretty amazing to see the speed at which we can go from having a checkpoint that we're really excited about to putting it in the hands of internal developers to then seeing it come to life in a product. And then only when you see it come to life in the product do you really start finding its rough edges and to be able to actually then kind of come to terms with how you do that. And so more and more than it becomes like, okay, how do you have the right ability to tune prompts quickly? How do you have the ability to run really good live experiments where you can get really good data and feedback quickly? How can you build evals that help give you real signal? Those are the things that will speed up your progress of quality the most because it will give you the ability to actually get to the kind of product that you love. And I think if you think about NotebookLM, I mean, that team really understands the model. Like they are just like, I mean, you talk about a Banger product, it comes from like a Banger team. Like they are really good at being able to take the model and play with it quickly and prototype quickly to get to something amazing. And I think you see that actually play out in the product.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence