High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Tulsee Doshi: evaluation

20 May 2026 The Cognitive Revolution The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

“Actually, like, it was awesome, one of my coworkers Anka, she's our lead for safety and alignment. And the other day, she, I think maybe a couple days ago, she pinged me from her hot tub and she was like, I could run all of these ablations from my phone because I could kick off a bunch of things to actually ablate Gemini to test for a bunch of these issues to see how some of our SIs differ or some data ablations differ.”

— Tulsee Doshi

Source trail

Everything needed to verify it.

Speaker
Tulsee Doshi
Attribution
Verified speaker
Claim type
evaluation
Recorded
20 May 2026
Publisher
The Cognitive Revolution

Transcript context

…Could be also perhaps productized as an RL environment and sold in to you guys that way. It's quite the cottage industry these days. So obviously the other big thing that I think is very much in the air, and actually the reason I'm here this weekend, when we were originally planning to do this remotely, is I'm going to this event called Recursive, where the topic is going to be recursive self-improvement and hopefully how we can navigate it successfully. How bought in is Google DeepMind to recursive self-improvement? When you talk to anthropic people, it's like they're almost religious about it, and also see it as totally inevitable. OpenAI has this later this year and in early 2028 timelines for an ML intern and a full-fledged AI R&D employee. Do you guys have milestones or timelines for when you're going to hand off the ML research to AIs? I mean, we're already using Gemini pretty deeply internally to improve Gemini. And so I think that is very much a theme for us, which is like, how can Gemini actually be a part of the Gemini development process? And so that can include things, I think that goes the full range from, helping us be more productive. So that's obviously like the simplest part of this to actually like, submitting CL that would actually like run an eval that would actually, suggest a research improvement that would actually drive improvements to Gemini itself. And I think there's a lot of ambitions we have to keep pushing in that research direction. So I think very similar to the other labs, I think this is very much an area of investment for us and an area we're super excited about. I think for me, what I'm really excited about is like, I think there's this really awesome research partner opportunity that we have with Gemini, right, for it to help us with creative ideas, for it to like help us test things faster. Actually, like, it was awesome, one of my coworkers Anka, she's our lead for safety and alignment. And the other day, she, I think maybe a couple days ago, she pinged me from her hot tub and she was like, I could run all of these ablations from my phone because I could kick off a bunch of things to actually ablate Gemini to test for a bunch of these issues to see how some of our SIs differ or some data ablations differ. And here's my report. And I could do all of this in the last hour. And that is amazing. And that's the kind of thing that we can already do, right? So then imagine where we'll be in six months, a year, two years from now. Yeah, I feel like it feels like at least my personal perspective is it's like a very much more like practical perspective, which like it's like as obviously as models get coding, they're going to go do things that is code related. It's going to they're going to help us build our products. They're going to help us train models. I think all the nuance of the story is in like Sort of like, where is the human sort of in the driver's seat of this stuff, and I think we are like the tools are built for the human to be in the driver's seat, which I think is an important thing as sort of we continue to go forward, and also I think very genuinely though, and... I think the model team and the researchers feel this more than ever. Like you definitely, I think the near-term horizon is going to continue to be the human in the driver's seat because the cost of these runs and the opportunity cost of going in the wrong direction and putting a bunch of resources is super, super high. And so I find it doesn't seem like super realistic in the short to medium term that you're going to just like be letting large-scale pre-training jobs be kicked off by the ML intern and it's going to cost you, many, many dollars and lots of compute and taking it away from sort of the human researchers. But this deep collaboration between AI and human researchers, I think is super obvious.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence