Evidence receipt / belief
Published · transcript-backedJérémy Scheurer: belief
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.”
Source trail
Everything needed to verify it.
- Speaker
- Jérémy Scheurer
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting. Developers love it. I love it. You know, these models have made me more successful, more organized. I can get more stuff done. Everyone loves it. And now everyone's thinking, shit. They took Fable away. And now people in The States potentially have access to better intelligence than me. Soon governments will have access to better intelligence than me. And part of the reason everyone's been going along with it is because it's been 1 of the most democratizing emancipating things in in history. You just pay $200 a month, and you get you get access to this intelligence. When that changes, I I think it's gonna be different. I think 1 good illustration to also showcase how surprising these kind of capabilities can arise. Different now from from, say, scheming or deception, but when, Anthropic released Mythos and their model, and they made it, like, available via their project Glasswing, they showed that this model suddenly has, like, this really, really high, like, cybersecurity capabilities. And I know that there are, like, there have been reports that show, hey. Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there. I know. It's it's so weird because I can see this from both directions. Right? You know? Because I I've been critical of Eliezer before because, you know, it feels like he's cartoonishly abstracting, you know, intelligence like it's on a single, you know, like it's a spectrum and, you know, like, agentic behavior and how they how they can be disconnected. And, yeah, like, I'd read the Carlini article with with with Mythos, and I've interviewed him before. And and I I was thinking, well, this this guy is a world class security researcher. So he knew how to set up a harness, what problems to choose, and so on. And then I have access for it myself, and I'm thinking, shit. Right? It's actually really, really good. But there are still problems, you know, like it still generates spaghetti code. It still makes mistakes. It still doesn't understand, but So is it's I I think we're we're all in a collective state of confusion at the moment.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.