High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Tulsee Doshi: commitment

20 May 2026 The Cognitive Revolution The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

“So we want any of our enterprise customers or a developer who's building their own use case to be able to leverage Gemini effectively. And so it is important then from a model standpoint that we're training in such a way that we actually, we sort of call it like harness diversity.”

— Tulsee Doshi

Source trail

Everything needed to verify it.

Speaker
Tulsee Doshi
Attribution
Verified speaker
Claim type
commitment
Recorded
20 May 2026
Publisher
The Cognitive Revolution

Transcript context

…It's a good question. I mean, I think Again, Chelsea probably knows better than me on this, but I think the best case is like you can do both. Like the best case is like it works really well for Gemini and sort of we can sort of do the things we want to do to scale up because we do have sort of control over the sort of full stack AI story as Sundar likes to say. But then also it generalizes across other stuff. Like I think the developer ecosystem, people want choice, people want to have flexibility in these tools. lots of use cases. Actually, there's like philosophical questions of how good really is your model if it can't generalize to sort of other harnesses. But I don't know how much. Yeah, I think that's the right, I think I fully agree. I think actually like maybe to double click on what Logan said originally, right? The benefit of the full stack that we have is we can hopefully build a really seamless experience. And you get the best of Gemini, you get it working in the most effective ways for you. get it working in a way that is intuitive, is smart, is fast. And so that also helps us then train the model to be better. So this becomes this flywheel that continues to power the model. At the same time, I think we don't want it to only be the case that the model works in a single harness. So we want any of our enterprise customers or a developer who's building their own use case to be able to leverage Gemini effectively. And so it is important then from a model standpoint that we're training in such a way that we actually, we sort of call it like harness diversity. We should be able to support a range of different approaches to tooling, to different approaches to orchestration, et cetera. But I think what's helpful about this approach of kind of co-training and building that flywheel, it's easier to debug. It's easier to think about data collection. It's easier to eval. You can just move at a faster pace. And I think we're seeing that across the industry. And so finding that balance is important, but I think it just helps build to make the model better. Yeah, I think there's a good, this is also a good pitch for like a harness bench. If that's not a benchmark that exists, let somebody somebody build harness bench. Yeah, I would love to would love to collaborate if folks are interested in that, because I do think it's like a great test of has sort of this perspective from a game for games. Actually, as an example, like if models are so good, like why can't they play games really well? And sort of if models are so good and we're actually approaching AGI, like why even if you do sort of the model harness training symbiosis, you still expect it to generalize reasonably well in other harnesses. If you can't, that's actually like, it's another sign of sort of the jagged intelligence. So I think it'd be cool to see this like play out from an actual benchmark perspective.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence