Evidence receipt / evaluation
Published · transcript-backedRobert Lange: evaluation
13 Mar 2026 Machine Learning Street Talk When AI Discovers The Next Transformer - Robert Lange (Sakana)
“I think 1 thing that was very interesting about the circle packing problem, sort of also coming back to the problem problem that I discussed initially was that originally, we we used a formulation where the correctness is checked with, like, a very tiny amount of slack.”
Source trail
Everything needed to verify it.
- Speaker
- Robert Lange
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…And on this while we're on this circle packing problem. So you you had this plot showing how it convergent and seemed to converge quite quickly. So and we'll show the plot on the screen now. So very quickly, the performance jumped up and then it slowly converged. And you said in the paper that it was using 3, I think 3 core innovations. And my thinking was, if you ran this 50 times, would it be the same every single time? And how to what extent is it thinking outside the box? You know, Sebastian Bubeck is always posting on Twitter talking about how GPT-five has just, you know, discovered new things. And there's always the question of, well, is it just searching the Internet? Is it just finding things that have been found before? And, yeah, combining things together in in a new way. But could it really think outside the box? Mhmm. Yeah. I think this is almost like a subjective question. Right? So first off, I don't know all problems on the Internet that try doing circle packing. Right? But what I can see in the tree that we also depict is, there's for example, like a crossover operation between 2 programs happening where sort of different concepts are combined. Right? So 1 important part is, for example, the the initialization of the circles. Another 1 is, like, the optimization. So basically, like a constrained optimization program is executed. And then the final part is basically like a reheating stage, right, when noise is added and sort of more stride to be squeezed out. And to me, like this sort of propagation of information through the tree is 1 that's really really fascinating. Right? Where in some sense, these stepping stones are actually used and so in a complementary fashion. Right? And with regards to rerunning the program multiple times, right, of course, there's some stochasticity in that. Right? So we're using language models and sort of due to, like, the the queuing device scheduling on on their server side, basically, we can't get rid of all the all the noise. We we've seen that at least for the general quality of the solution, so what is the right afterwards, it is possible to re obtain this. But sometimes with a different program like most of the times just by stochasticity. Right? So it's not like there's for many problems, there's, like, not 1 solution that achieves that score, but there is, a spectrum or, like, a a region, let's say, in the program space that that resembles the same. Right? I think 1 thing that was very interesting about the circle packing problem, sort of also coming back to the problem problem that I discussed initially was that originally, we we used a formulation where the correctness is checked with, like, a very tiny amount of slack. Right? So the the circles could overlap a tiny little bit. And then afterwards, we we we sort of reduced the RAID AI and the solution was exact. Right? This didn't change the score by too much, so it's still state of the art, but it was essentially like a proxy problem. We then reran the the Schenker Evolve on the exact setting, and we found that it took a little bit longer to actually obtain the same quality of a solution. So I think this already points a little bit in this direction of what I discussed in the beginning, like sometimes sort of surrogate problems might actually be extremely valuable in in making such discoveries. And having an automated way for designing these surrogate problems in an efficient way might be something really important going forward. Yeah. That's absolutely fascinating. It reminds me of support vector machines where we make the optimization tractable by introducing slack variables and you can think of that as a kind of surrogate problem. But then I'm thinking what would ShinkerEvolve or AlphaEvolve, would it know to introduce a surrogate problem? Because, you know, as designers who understand, you know, we can think outside the box and and we can do stuff like that. Because presumably, if the fitness function had the constraints that there were no circle intersections, then it wouldn't it wouldn't occur to the algorithm to come up with a surrogate problem.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.