High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ilia Shumailov: belief

4 Oct 2025 Machine Learning Street Talk AI Agents Can Code 10,000 Lines of Hacking Tools In Seconds - Dr. Ilia Shumailov (ex-GDM)

“I mean, even if it did, think it's very hard to do because, especially for security, I think interpretability is maybe an interesting tool for safety sort of things, but for security it's definitely not a step in in in the right direction.”

— Ilia Shumailov

Source trail

Everything needed to verify it.

Speaker
Ilia Shumailov
Attribution
Verified speaker
Claim type
belief
Recorded
4 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…There's even there's even that work from Subaru that it might not mean anything whatsoever. Yeah. Like the the thinking traces aren't even actually directly relevant. It was just kind of some sort of weird workspace that the models were using, right? Their reasoning won't correspond to the answer. Yeah. I mean, even if it did, think it's very hard to do because, especially for security, I think interpretability is maybe an interesting tool for safety sort of things, but for security it's definitely not a step in in in the right direction. Because, like, broadly speaking, if you take something extremely multidimensional and project it into something very small dimensional because a human can comprehend this thing, then you will have a lot of collisions where, like, this multidimensional space map to the same sort of smaller small smaller dimensional space. Does it always correspond to, like, bad behaviors? Maybe. Maybe not. Who knows? But clearly, this is not enough. We need something else. We need something where we can get, like, a 9 99.9999 reliability out of this. You have to think outside the box.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence