Evidence receipt / commitment
Published · transcript-backedTim Scarfe: commitment
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“It was using Sonnet 3.5, and you had about 4 iterations, I think. And and, essentially, you you know, you were working on the ARC challenge, you were producing these programs through evolution.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. Sure. So I actually have only been working in research for about 8 months. Before that, I had a company, right out of college. I was got into Y Combinator, and so I've been running a company, for the last 4 and a half years as CTO. And I've always been very interested in reasoning in the brain. And I actually picked up Jeff Hawkins' book, A Thousand Brains, and I read that. And at the same time, I was kind of coming into language models, and something just clicked inside of me and I just knew I had to be working on this. And I believe that artificial general intelligence will be the most important invention of hopefully my lifetime. So I decided to drop everything. Stepped down as CTO. Company is still going well, so it was a difficult decision. And I actually had gotten in touch with Mike and Francois because I thought Arc AGI was such an elegant way of describing the problems with current language models and their difference between the human brain. And so I just kind of dug in. That was my first research project independently. And, yeah, I ended up getting the top score on that, and that was really great. And after that, I got recruited into Francois and Mike's AGI lab, India, where I was working on program synthesis. And as you described earlier, over time, I've become more convinced that language modeling with reinforcement learning will yield generalization far beyond what we see today. And so I decided to move to a company that was focused purely on language models, and that's where I am now. So I'm currently working on reasoning and post training at Reflection, which is we're building frontier foundation models. Very cool. Maybe we should save that bit for tiny bit later. But 1 of the take home messages in your well, no, I mean, it's super interesting. And 1 of the take home messages on your new approach is that rather than producing explicit programs, you are evolving descriptions of programs. And Francois is a neurosymbolic guy. He thinks that we need to have a symbolic substrate where we, you know, represent the the kinds of problems that that we can do. And we need to do this kind of compositional form of of of intelligence. So we need to kind of be working in the symbolic layer, but perhaps guided by, you know, deep learning models. But maybe we we should get to that in in a minute. So in your in your first solution, it was an evolutionary approach. It was using Sonnet 3.5, and you had about 4 iterations, I think. And and, essentially, you you know, you were working on the ARC challenge, you were producing these programs through evolution. Maybe just for folks that don't know about the ARC challenge as well, could you introduce that and and get into your solution? Sure. Yeah. So the ARC challenge is kind of like an IQ test for machines. It's a set of input output grids, and the whole point is to be able to figure out to transform input grids into output grids given a common transformation rule. And so, what's interesting is these are really easy for humans. Right? The average human gets around 75% accuracy on ARC v 1, and at the time, the best language models, GPT 4, SONNET 3 5, was getting maybe 5%. And so, yeah, basically, the idea is you have a few training examples and then you're you're you're trying to extrapolate the transformation rule on the final test example. And so I approached this. I was actually inspired by Ryan Greenblatt who had a solution earlier, which was to generate a ton of Python programs that would encapsulate the transformation role. And Python programs are great because they're deterministic, and you can pretty quickly check whether or not the Python program works or not, which is really cheap, so it's cheap to verify. And you can be relatively sure if the Python program works on all of the training examples, that it'll work on the test example. So, I started with his approach, but then I noticed that, you know, the language models actually struggled on first attempts. Even if you ask the language model a thousand times to generate Python programs, they were always off by small amounts on easy tasks, which I thought, you know, presumably, it's in their distribution. They should be able to solve this. So, what I found is that actually by taking the top, programs, the top performing programs, and then running that in a revision loop, so asking, Sonnet 3 5, hey, here's what you got wrong, here are the cells you got wrong, here's your original Python program, improve it, that started to really work well. And then I thought, well, why not just increase the depth, right? Why not ask it 10 times to revise until I'm happy with the solution, passes some sort of accuracy threshold? So that's kind of how I was inspired by it. And I didn't think of it as evolutionary at first. I was just thinking about broadly what would work. And over time, I kind of understood there was something a bit deeper going on here, which is that evolving solutions is a powerful technique generally. And I think it's actually going to play a role, you know, in future technologies. But, yeah, that's generally guided by just intuition.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.