Evidence receipt / recommendation
Published · transcript-backedIlia Shumailov: recommendation
4 Oct 2025 Machine Learning Street Talk AI Agents Can Code 10,000 Lines of Hacking Tools In Seconds - Dr. Ilia Shumailov (ex-GDM)
“Like, you can convince them that things that they were not capable of doing before are some things that they need to do, and suddenly this leads to some sort of a loss of of a different kind. So I I think in the paper you're referring to, what we were trying to do was we were trying to send an email to the agent such that when this email ends up in the agent's context, rather than following the user task, the agent is actually doing something else, like the the thing that we've send it over the email.”
Source trail
Everything needed to verify it.
- Speaker
- Ilia Shumailov
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 4 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So when you worked at DeepMind, we we've read your paper. You were involved in defending Gemini basically against these indirect prompt injections. Yeah. And I guess a couple of you obviously tell us about that. But 1 thing you found that which is very interesting is is almost that as the model was increased in capability, they became more vulnerable, which is fascinating. I think I think it's important to state here is that, like, I wouldn't actually phrase it this way. It's a it's a bit hard to say more or less vulnerable, but I do have to say that big models today, if we compare them to the big models 5 years ago, they fail in very different ways. Right? In in the past, we were kind of, like, quite efficient in discovery of the several examples using gradient information, blah blah blah. Right? You you you can kind of make an optimization problem and optimize, and this stuff works. Whereas with models, like, they clearly become more robust against this stuff, or at least it's significantly harder to traverse this landscape, but they also become a lot less robust against other adversaries. Like, suddenly, very simple rephrasing of the same questions forced it to completely do something different. And I also have to say that I think we had significantly more control when we dealt with smaller models. Like, we kinda knew which knobs to turn to make the model do stuff. Whereas nowadays, you look at the modern big model, it's like it's way too much alchemy. Like, it's it's completely impossible to tell, like, oh, I have added this thing inside. What actually happens to the whole thing? Is it better? Even answering a question of is it better is is hard. So, like, you will find yourself in a position where you need to run experiments for the next couple of months trying to even, you know, discern whether something you have added actually changed anything about these big models. So I I think in part, what we were talking about in in in this work, I think you're thinking about, is that when your models get better, they get significantly better like, capability are growing of the model overall. They get significantly better at following instructions. When they follow instructions and they become better at following instructions, you can suddenly do a lot more. Like, you can convince them that things that they were not capable of doing before are some things that they need to do, and suddenly this leads to some sort of a loss of of a different kind. So I I think in the paper you're referring to, what we were trying to do was we were trying to send an email to the agent such that when this email ends up in the agent's context, rather than following the user task, the agent is actually doing something else, like the the thing that we've send it over the email. And we found that pretty much in all of the cases, we're capable of doing this. And if you take a whole bunch of academic literature on how to build defenses, and you can actually find startups pretty much implementing the same things and selling this as a, like, a security solution, Like, the those approaches don't work. And we could pretty much always find relatively universal ways to produce this, like, email that ends up being sent to the agent that forces the agent to do something drastically different from what it's supposed to be doing. always find relatively universal ways to produce this, like, email that ends up being sent to the agent that forces the agent to do something drastically different from what it's supposed to be doing. Modern computers are very much just, you know, a piece of magic where this chip works. You, you know, you you move it there very so slightly. It becomes unstable, and then nobody knows what's happening.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.