Evidence receipt / evaluation
Published · transcript-backedRobert Lange: evaluation
13 Mar 2026 Machine Learning Street Talk When AI Discovers The Next Transformer - Robert Lange (Sakana)
“And I think language models with the right sort of evolutionary hardness are extremely powerful in terms of scaling up to to to make discoveries. And, yeah, I think Jeremy as well as the AlphaEvolve paper as well as sort of work we've done on, like, the Daven Gudel machine, for example, shows that this sort of stepping stone accumulation plus iterative verification and collecting sort of information and evidence from the real world real synthetic evaluator is really important for that.”
Source trail
Everything needed to verify it.
- Speaker
- Robert Lange
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. That that's actually a really important point because I suppose we can use these foundation models. And first of all, isn't it just fascinating to reflect that we have these amazing models out there that we can access. So like GPT 5 and Grok 4. And they are so much better when you get them to refine their solution in in several steps. What why is that? I mean, I suppose a naive question would be, why why aren't they just good out of the box? Potentially, like with enough random samples. Right? It's sort of this monkey typing on the keyboard. They would potentially be able to get there. Right? But in principle, it's sort of coming back to the principles of evolution. Right? In the sense that you need to collect a bunch of stepping stones first and then build on top of them to to really find innovations or to tune innovations down the line. And I think language models with the right sort of evolutionary hardness are extremely powerful in terms of scaling up to to to make discoveries. And, yeah, I think Jeremy as well as the AlphaEvolve paper as well as sort of work we've done on, like, the Daven Gudel machine, for example, shows that this sort of stepping stone accumulation plus iterative verification and collecting sort of information and evidence from the real world real synthetic evaluator is really important for that. Very cool. And stepping stone collection. So that this is it came from Kenneth Stanley. It's a wonderful paper, Why Greatness Cannot Be Planned. And he said that it's it's better to have systems that don't converge. So in natural evolution, we are just trying all of these different things And greatness quite often follows a diverse path, which means you have to do things which initially seem quite stupid. And then later on, they turn out to be incredibly useful. Yeah. We're trying to design algorithms that can kind of allow for a population of slightly weird things. And and then we kind of lock in and and converge a little bit. So we we're still converging though. So we're still building systems that don't diverge forever. What are we losing?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.