Evidence receipt / prediction
Published · transcript-backedTim Scarfe: prediction
31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research
“For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 31 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. There's been, like, various proposals on this. I'm, like, personally think that, like, a global coalition could work if, like, some of the, like, Western states, for example, The United States and Britain and so on would go together and essentially, yeah, make a treaty on, like, what sorts of capability levels we set a red line at, and say you can't train anything further than this. Yeah. This is, like, my personal opinion, maybe, more than Apollo's opinion, but 1 other thing is just, like, slowing down the AI race and, like, taking it slower until we, like, have a better alignment science to deal with some of these problems. For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting. Developers love it. I love it. You know, these models have made me more successful, more organized. I can get more stuff done. Everyone loves it. And now everyone's thinking, shit. They took Fable away. And now people in The States potentially have access to better intelligence than me. Soon governments will have access to better intelligence than me. And part of the reason everyone's been going along with it is because it's been 1 of the most democratizing emancipating things in in history. You just pay $200 a month, and you get you get access to this intelligence. When that changes, I I think it's gonna be different. I think 1 good illustration to also showcase how surprising these kind of capabilities can arise. Different now from from, say, scheming or deception, but when, Anthropic released Mythos and their model, and they made it, like, available via their project Glasswing, they showed that this model suddenly has, like, this really, really high, like, cybersecurity capabilities. And I know that there are, like, there have been reports that show, hey. Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.