High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Tim Scarfe: prediction

31 Jul 2026 Machine Learning Street Talk How Researchers Test AI for Hidden Goals — Apollo Research

“For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
prediction
Recorded
31 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. There's been, like, various proposals on this. I'm, like, personally think that, like, a global coalition could work if, like, some of the, like, Western states, for example, The United States and Britain and so on would go together and essentially, yeah, make a treaty on, like, what sorts of capability levels we set a red line at, and say you can't train anything further than this. Yeah. This is, like, my personal opinion, maybe, more than Apollo's opinion, but 1 other thing is just, like, slowing down the AI race and, like, taking it slower until we, like, have a better alignment science to deal with some of these problems. For the 1st time now everyone's thinking about sovereign AI, because you talk about loss of control, but right now, these models are agency promoting. Developers love it. I love it. You know, these models have made me more successful, more organized. I can get more stuff done. Everyone loves it. And now everyone's thinking, shit. They took Fable away. And now people in The States potentially have access to better intelligence than me. Soon governments will have access to better intelligence than me. And part of the reason everyone's been going along with it is because it's been 1 of the most democratizing emancipating things in in history. You just pay $200 a month, and you get you get access to this intelligence. When that changes, I I think it's gonna be different. I think 1 good illustration to also showcase how surprising these kind of capabilities can arise. Different now from from, say, scheming or deception, but when, Anthropic released Mythos and their model, and they made it, like, available via their project Glasswing, they showed that this model suddenly has, like, this really, really high, like, cybersecurity capabilities. And I know that there are, like, there have been reports that show, hey. Actually, you know, like, previous model could also sometimes find these things. But I think when you look at it, there have been multiple reports now that, the rate at which orgs are disclosing that they find capabilities really strongly correlates with when, like, Mythos came out, and, like, I don't think anybody had on their bingo card that, like, right then, this kind of capability would be there.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence