Evidence receipt / recommendation
Published · transcript-backedSpeaker unverified: recommendation
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“important that the checker was stronger than the actual instruction creator, which I I think is, interesting.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- recommendation
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…In the first 1, I used Python programs because Python programs are deterministic and it's really easy to verify whether or not, you know, it's correct. Did it did it run, and then, did it run on the training examples and produce the correct outputs. So it's like a perfect program. Right? Like, it is a program. The problem is it, you know, Python programs are brittle in that, you know, there are many things that are very difficult to describe with Python. Arc grids in v 2 being 1 of them. Right? So you have some grids that are very easily described by Python, but then almost the majority overwhelming majority in ARC v 2 are very hard to describe in Python. The correct Python formulation is, you know, lines and lines and lines, and, really what you want is a more expressive program. And so that's why I switched from Python to English, which is a much more expressive program. You can you can describe every single Arc v 2 task in 10 bullet points of plain English, most of them in 5 bullet points. And I think that this actually gets to the heart of ARC. Right? Everything is quite simple. It's not very hard. And I think this is also how we do it too. Right? Like, when we look at these ARC graphs, ARC grids, we're coming up with these bullet points in our head, and we're, you know, checking them. Okay. This was right. This was right. And Python doesn't have these features. It's just not as expressive as natural language. And I think another way to put it would be you have this inductive, transductive trade off. Right? You could think of language models as being trained inductively, and then they have an inductive bias, and you almost want to let that inductive bias express itself fully in a way. And the way you do that is to give it the full power of how it was trained. And I think this is the same thing with humans too. Right? If I told you to solve ARC with Python programs, you'd do a way worse job, even if you were an expert at Python. And so I think fundamentally, it's more general, and it leads to general and better solutions. I mean, the accuracy is much higher when you use natural language. Now, the problem is you actually have to then verify whether the instructions are correct. You can't run natural language on R grids. This was the fundamental problem with the solution. This is what made iteration challenging, especially because you know, for each grid, for each training example, you have to run the natural language instructions, and it takes a really long time, especially with this thinking model. So I originally started with a weak model. You know, it's the checker model. It's the checker agent. Let's just use GPT 5 mini, whatever. Nano. And it did terribly. So I ended up it was actually more important that the checker was stronger than the actual instruction creator, which I I think is, interesting. But, yeah, that just highlights, you know, the the trade offs with using natural language. important that the checker was stronger than the actual instruction creator, which I I think is, interesting. But, yeah, that just highlights, you know, the the trade offs with using natural language. It's it's you can express, you know, much more concisely programs that you wanna run, but then they're not runnable programs. You actually have to check them inductively. So that was the trade off, but it was worth it for Arc v 2. Yes. So so fascinating. And just for the audience, we've been using transduction and induction to distinguish predicting the solution space versus predicting a program. I had this this discussion with with Clement Bonnet. We need not detain us now, but I think in traditional machine learning, transduction means that the test example is a function of your prediction. I had this discussion with the architects as well that when you have this natural language description, natural language is more expressive, which simply means that there are more degrees of freedom. And this is the beauty of of LLMs that there's this huge kind of space that you're traversing around. And when you use natural language, you can just traverse to more places in that space more easily. So it seems like it would be a win. I'm And really fascinated to to find out whether that is just like a huge component of your solution. Because on Eric's solution, he's still predicting programs and still doing very well. So I'm not I'm not sure about that. And the other thing is I wasn't entirely sure whether you are actually using a transtactive method. So in your solution checker agent, is it directly going to the solution space, or is it generating a program and testing it?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.