Evidence receipt / belief
Published · transcript-backedRobert Lange: belief
13 Mar 2026 Machine Learning Street Talk When AI Discovers The Next Transformer - Robert Lange (Sakana)
“While on others, like ARC AGI 2, like this whole sort of semantic evolution seems to be more efficient. So I think ideally we we can get a system that can automatically in some sense decide whether or not it wants to take like a programmatic approach in settings where it's actually feasible and easier to to bootstrap off, or it takes the semantic approach of evolving instructions or like LLM driven input output mappings.”
Source trail
Everything needed to verify it.
- Speaker
- Robert Lange
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 13 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Very cool. Now, other thing, we'll show the graph on the screen, the evolutionary graph. So for the circle back in problem. I was looking at that and first of all, it looked incredibly parsimonious, which is good. It it looked like it had found an optimal path to the solution very quickly. And I was thinking in my mind, well, maybe there's some natural pattern that that there's there's there's there's something about that that we could use in the abstract to guide the evolution in the future. But the other thing I'm thinking about is right now, the problem with machine learning is that we don't really have semantics baked in. So what we're doing is we have a verifier, we're looking at the rewards, and we're sort of like doing patterned exploration, and we're taking steps towards the, you know, towards the target. And I love mechanistic forms of reasoning where we actually know something about what the program components mean. And the reason this is important is when we're merging together the best performing programs from 2 different islands. That's a kind of first order interaction. It might not make sense to merge them together. It's wonderful that LLMs, you can give them any pairs of programs and it will find a way to merge them together. But wouldn't a more principled way be of there's there's some kind of semantic primitives here and we know they fit together. So there's this Lego analogy that we're kind of building up based on principles rather than forging our path based on the performance. Yeah. That's a good point. So 1 thing we do in Chinkai Evolve as well is we keep essentially a scratch pad. So each program is being summarized. And then from the program summaries, we keep sort of a set of global insights, let's say, then we're shared or, like, extracted from these programs. And then based off of the scratch pad, we construct sort of meta recommendations that then become part of the system prompt. Right? So that way, you can try to sort of semantically grasp some of the discoveries. But a general problem, which is again sort of task dependent is thereby you sort of diffuse that knowledge across the tree. Right? But sometimes you want things to be much more isolated. Right? It's always like a trade off where you somehow have to find for your problem the right position on the spectrum of how much non diffusion do you wanna have and how much sort of, let's say, hard islands of programs do you wanna have. Right? And, yeah, we're to make steps in the direction of sort of automatically adjusting this in an optimal way, but again, it's very program sensitive. And then sort of, I think, another point where you're already sort of going into is sort of Jeremy Jeremy's solution to Arc AGI. Right? And sort of doing solution evolution in the instruction space, right, instead of the program space. I do think that this is something important, and we're like I said, with, like, the construction of this meta scratch pad trying to do sort of both at the same time. Again, it's problem dependent. Like, I played around a little bit with ARC AGI 1 and ARC AGI 2. And I think on ARC AGI 1, actually, the the transform sort of program direction is actually quite effective. Right? It's like Jeremy said, it's deterministic, and it's easier to sort of get clear signal to improve on during your evolution process. While on others, like ARC AGI 2, like this whole sort of semantic evolution seems to be more efficient. So I think ideally we we can get a system that can automatically in some sense decide whether or not it wants to take like a programmatic approach in settings where it's actually feasible and easier to to bootstrap off, or it takes the semantic approach of evolving instructions or like LLM driven input output mappings. Yeah. It's it's so interesting because, you know, like a a symbolic AI person would say, oh, I don't like connectionism because it doesn't under you know, the only semantics in connectionism is this notion of similarity. It doesn't really understand things. So so they would say, well, just just start with a a an entity relationship graph and then just kind of build up using, you know, composition and first principles. That that that doesn't work. Right? So we're using neural networks because they're incredibly flexible and they understand a lot of things about the world, but they don't have the kind of constraints that we want. So what we do is we use these tricks. So Jeremy, we evolved program descriptions. On your program selection, you had a semantic novelty detection, you know, using like a Embedding based similarity. You had like a kind of self similarity metrics and you know, based on the cosines. And indeed, you've got this meta scratch pad. So what we're seeing is this fascinating spectrum of possibilities where still using neural networks, you can imbue semantics in using all of these different tricks, but they all come with trade offs.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.