High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Tim Scarfe: recommendation

4 Oct 2025 Machine Learning Street Talk AI Agents Can Code 10,000 Lines of Hacking Tools In Seconds - Dr. Ilia Shumailov (ex-GDM)

“Like, you know, they're As part of the Alpha Evolve system, which is very interesting, I recommend people watch that episode, you know, it goes and runs external verifiers.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
recommendation
Recorded
4 Oct 2025
Publisher
Machine Learning Street Talk

Transcript context

…I think in your case, you're an unusual person. Yeah. Most of the people are not like this. Yeah. Exactly. Like, kind of some kind of Luddite or something. But I wanna just push back on 1 thing about the the halting problem. Because I find actually that it has very practical consequences. So for example, when we were talking to the Alpha Evolve team. Yeah. Right? Like, you know, they're As part of the Alpha Evolve system, which is very interesting, I recommend people watch that episode, you know, it goes and runs external verifiers. Yeah. Right? But the problem is, you don't know if the external verifier is gonna complete. Yeah. So what do you do? Well, you have to put in some arbitrary computational budgets, thresholds. If it doesn't complete within, you know, 5 seconds, then you just terminate it and, you know, look at other runs. So while you can do that, I think it also introduces biases and the types of programs we can discover. Right? Because maybe if I had set my budget to 7 seconds instead of 5, I would have found like a more optimal solution. Right? So I think the halting problem, while theoretical, like, also has very important practical consequences just when you sit down and try and run a program. Right? I don't think the problem of value for Evolve is halting problem. Right? Right? It's more like today, our programs have very weird semantics where we're ahead of time, can't say quite a bit about them. Right? Because for for a lot of programs, like, if we can rewrite them in a slightly more, like, suitable language, we can get a lot a lot more out of them. Right? You know, like, the sort of reasoning you can get out of OCaml code is quite different from what you get out of, like, if you were writing machine code straight away or if you're writing whatever. And I think it's just we don't know how to do a lot of this stuff. And in general, like, even if they were able to run it, let's say they have evolved into a program that will theoretically take 10000 hours to run. Right? And you can tell it's gonna take 10000 hours to run. Like, are they supposed to run it or not? This is a part of the loop. I think it's more of a it's just nowadays, we kind of pay with time for a lot of this stuff rather than paying, like, can we find another set of example? Of course, we can. We just need to run this for longer. And your your budget is sort of time time budget. So, like, whereas halting problem is more of a you have infinite amount of time. You have infinite amount of memory. Can you reason about this? No. Not really. But the sort of programs that, like, we're talking about alpha evolve scale, they are not very large.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence