High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Ilia Shumailov: evaluation

4 Oct 2025 Machine Learning Street Talk AI Agents Can Code 10,000 Lines of Hacking Tools In Seconds - Dr. Ilia Shumailov (ex-GDM)

“This all of these assumptions are are sort of are gone. It's we don't know how to build systems against this.”

— Ilia Shumailov

Source trail

Everything needed to verify it.

Speaker
Ilia Shumailov
Attribution
Verified speaker
Claim type
evaluation
Recorded
4 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…But what I'm trying to say in in security terms is that agents are, like, is is is a worst case human sort of. Well, not not even this. Like, agents are very different from humans. So let's, like, let let's take this as a stance. Right? You will not find a single human in the world that works 24 7, touches absolutely every single 1 of your endpoint in in your system that absolutely knows everything there is, like, that can generate you basically all of the hacking tools on on a whim, like, just because it knows, it has seen all of them, it can recreate this in a matter of a second. A normal human adversity when I'm an enterprise and I think, this may be an insider from, like, a from a competitor. The way they work is drastically different. They make an assumption. You as a user, you can't write, you know, 10,000 lines of hacking tools in in a day. This is not something you will be able to do, and then bring in in code as hard. With agents, that's not the case. You don't make an assumption that the user will go and touch every single endpoint you have in a network because, you know, why would they do this? And even if they do this, you call them in and you say, well, you know, we'll apply a legal framework and imprison you. So clearly, there is some sort of rationality and expectation that at least you will have some sort of a physical way to penalize. With agents, this doesn't exist. This is kind of like in security, we tend to say that a child is the worst case adversity you can find. Completely irrational thinking, infinite amount of time. They can basically touch everything. Like, they expect there are no expectations on behaviors whatsoever. But so agents are, like, even worse than that. And and this is even before we start talking about human to agent behaviors and agent to agent behaviors because this thing is just like, it's it's a billion times worse. So what I'm trying to say is we shouldn't think, okay. Yesterday, I was buying my security tooling from this company, and today, I'll buy it from another company, and and this will solve my problems. No. It's likely it's not gonna be like this. Because beforehand, we were building security toolings for humans by humans against humans. Now, I don't think we know what we're building. Mhmm. It's it's hard to tell. So before, when we employ you and we give you access to sensitive data, we assume coarse grained access control sort of policies. Like, okay, you can touch the sensitive data. Okay, you can Google at the same time. But if you try and Google a sensitive document, at the same time, will apply immense amount of pressure and will take you to court. You will lose your house, mortgage, blah blah blah. You'll go to prison worst case. Right? With agent, that doesn't work. This doesn't exist. This all of these assumptions are are sort of are gone. It's we don't know how to build systems against this. We need very fine grained. We need extreme precision. at doesn't work. This doesn't exist. This all of these assumptions are are sort of are gone. It's we don't know how to build systems against this. We need very fine grained. We need extreme precision. We need extreme control and transparency. Otherwise, it's just not gonna work. Yeah. I mean, I I guess I just think of agents as being, you know, they're they're a little bit like a calculator. You know, like, they're they're only as good as the prompt and what you put into them, which means that, you know, a very sophisticated actor could make a sophisticated agent. But even then, when the supervision stops, there would be quite a predictable cone of variation.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence