Evidence receipt / preference
Published · transcript-backedTim Scarfe: preference
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“So on the first 1 as well, you were generating Python programs explicitly. And because of all the things that we're just talking about, I'm a big fan of that because I feel intuitively, and I think you did, that there's something special about Python programs.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I need to think about that a bit more. Okay. So on the first 1 as well, you were generating Python programs explicitly. And because of all the things that we're just talking about, I'm a big fan of that because I feel intuitively, and I think you did, that there's something special about Python programs. And and then you you did to this iterative updating of of those programs and and you converged on the right 1. You also had this amazing diagram in your first blog post where you kind of visualize the space of all the possible programs and you kind of showed what was happening in every iteration. In the first 1, I used Python programs because Python programs are deterministic and it's really easy to verify whether or not, you know, it's correct. Did it did it run, and then, did it run on the training examples and produce the correct outputs. So it's like a perfect program. Right? Like, it is a program. The problem is it, you know, Python programs are brittle in that, you know, there are many things that are very difficult to describe with Python. Arc grids in v 2 being 1 of them. Right? So you have some grids that are very easily described by Python, but then almost the majority overwhelming majority in ARC v 2 are very hard to describe in Python. The correct Python formulation is, you know, lines and lines and lines, and, really what you want is a more expressive program. And so that's why I switched from Python to English, which is a much more expressive program. You can you can describe every single Arc v 2 task in 10 bullet points of plain English, most of them in 5 bullet points. And I think that this actually gets to the heart of ARC. Right? Everything is quite simple. It's not very hard. And I think this is also how we do it too. Right? Like, when we look at these ARC graphs, ARC grids, we're coming up with these bullet points in our head, and we're, you know, checking them. Okay. This was right. This was right. And Python doesn't have these features. It's just not as expressive as natural language. And I think another way to put it would be you have this inductive, transductive trade off. Right? You could think of language models as being trained inductively, and then they have an inductive bias, and you almost want to let that inductive bias express itself fully in a way. And the way you do that is to give it the full power of how it was trained. And I think this is the same thing with humans too. Right? If I told you to solve ARC with Python programs, you'd do a way worse job, even if you were an expert at Python. And so I think fundamentally, it's more general, and it leads to general and better solutions. I mean, the accuracy is much higher when you use natural language. Now, the problem is you actually have to then verify whether the instructions are correct. You can't run natural language on R grids. This was the fundamental problem with the solution. This is what made iteration challenging, especially because you know, for each grid, for each training example, you have to run the natural language instructions, and it takes a really long time, especially with this thinking model. So I originally started with a weak model. You know, it's the checker model. It's the checker agent. Let's just use GPT 5 mini, whatever. Nano. And it did terribly. So I ended up it was actually more important that the checker was stronger than the actual instruction creator, which I I think is, interesting. But, yeah, that just highlights, you know, the the trade offs with using natural language.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.