Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
13 Mar 2026 Machine Learning Street Talk When AI Discovers The Next Transformer - Robert Lange (Sakana)
“Right? So we're using neural networks because they're incredibly flexible and they understand a lot of things about the world, but they don't have the kind of constraints that we want.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 13 Mar 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. That's a good point. So 1 thing we do in Chinkai Evolve as well is we keep essentially a scratch pad. So each program is being summarized. And then from the program summaries, we keep sort of a set of global insights, let's say, then we're shared or, like, extracted from these programs. And then based off of the scratch pad, we construct sort of meta recommendations that then become part of the system prompt. Right? So that way, you can try to sort of semantically grasp some of the discoveries. But a general problem, which is again sort of task dependent is thereby you sort of diffuse that knowledge across the tree. Right? But sometimes you want things to be much more isolated. Right? It's always like a trade off where you somehow have to find for your problem the right position on the spectrum of how much non diffusion do you wanna have and how much sort of, let's say, hard islands of programs do you wanna have. Right? And, yeah, we're to make steps in the direction of sort of automatically adjusting this in an optimal way, but again, it's very program sensitive. And then sort of, I think, another point where you're already sort of going into is sort of Jeremy Jeremy's solution to Arc AGI. Right? And sort of doing solution evolution in the instruction space, right, instead of the program space. I do think that this is something important, and we're like I said, with, like, the construction of this meta scratch pad trying to do sort of both at the same time. Again, it's problem dependent. Like, I played around a little bit with ARC AGI 1 and ARC AGI 2. And I think on ARC AGI 1, actually, the the transform sort of program direction is actually quite effective. Right? It's like Jeremy said, it's deterministic, and it's easier to sort of get clear signal to improve on during your evolution process. While on others, like ARC AGI 2, like this whole sort of semantic evolution seems to be more efficient. So I think ideally we we can get a system that can automatically in some sense decide whether or not it wants to take like a programmatic approach in settings where it's actually feasible and easier to to bootstrap off, or it takes the semantic approach of evolving instructions or like LLM driven input output mappings. Yeah. It's it's so interesting because, you know, like a a symbolic AI person would say, oh, I don't like connectionism because it doesn't under you know, the only semantics in connectionism is this notion of similarity. It doesn't really understand things. So so they would say, well, just just start with a a an entity relationship graph and then just kind of build up using, you know, composition and first principles. That that that doesn't work. Right? So we're using neural networks because they're incredibly flexible and they understand a lot of things about the world, but they don't have the kind of constraints that we want. So what we do is we use these tricks. So Jeremy, we evolved program descriptions. On your program selection, you had a semantic novelty detection, you know, using like a Embedding based similarity. You had like a kind of self similarity metrics and you know, based on the cosines. And indeed, you've got this meta scratch pad. So what we're seeing is this fascinating spectrum of possibilities where still using neural networks, you can imbue semantics in using all of these different tricks, but they all come with trade offs. Yeah. For sure. Like, I think it's it's kind of interesting. We we've had a long period of computer science where algorithms were sort of designed by humans. Right? Then we had sort of this Android Kapathy software 2 paradigm where, like, we trained neural networks that then performed a certain function. And now we're sort of at this point where we're using LLMs to design algorithms or solutions more generally. Right? And I think, actually, like, even though, like, large frontier language models are extreme, like, let's say, black boxes or it's very hard to get a full mechanistic understanding of them, the outputs can be. Right? The programs, the instructions, and so on. Right? So I think it opens up a very sort of new paradigm of doing research or basically doing anything. Right? If you if you think about it. But I think we're we're just sort of at the starting point of figuring out the the right user interface for that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.