Evidence receipt / evaluation
Published · transcript-backedIlia Shumailov: evaluation
4 Oct 2025 Machine Learning Street Talk AI Agents Can Code 10,000 Lines of Hacking Tools In Seconds - Dr. Ilia Shumailov (ex-GDM)
“I have to say that in AI, I I think the reason why there is so much confusion about these 2 fields is because in safety, folks kinda started assuming malicious actors very quickly for for absolutely no reason, by the way.”
Source trail
Everything needed to verify it.
- Speaker
- Ilia Shumailov
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. And that that that makes a tremendous difference because we were talking about this when we were when we were sort of thinking about this interview. Right? Because so I I tried to take cryptography and cryptanalysis when I was in in graduate school. And the first week, I realized I just don't even have the math, you know, background for this. Right? But but the week I was in there, I really realized that the adversarial nature of the security problem just fundamentally changes the landscape. Right? Because now you have equally intelligent minds on both sides kind of going after each other and trying to defeat each other, and it's just a totally different calculation than than just what's it gonna do if everything's behaving correctly. Right? I have to say that in AI, I I think the reason why there is so much confusion about these 2 fields is because in safety, folks kinda started assuming malicious actors very quickly for for absolutely no reason, by the way. So, like, find an engineering discipline somewhere, like building buildings, for example, where you're trying to model adversaries. Like, are you modeling buildings expecting somebody to blow them up? I'm not sure. Right? Well, maybe somewhere in parts of the world where there is an active conflict, you do build special tooling around this. Right? But in terms of, like, bunkers in in every building and so on. But in terms of yeah. Like, in AI, it's because every time we talk about jailbreaks, you kinda have to explicitly look for them. We ended up considering adversaries straight away. But, yeah, this is quite uncommon. So when you worked at DeepMind, we we've read your paper. You were involved in defending Gemini basically against these indirect prompt injections. Yeah. And I guess a couple of you obviously tell us about that. But 1 thing you found that which is very interesting is is almost that as the model was increased in capability, they became more vulnerable, which is fascinating.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.