High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“I was actually inspired by Ryan Greenblatt who had a solution earlier, which was to generate a ton of Python programs that would encapsulate the transformation role. And Python programs are great because they're deterministic, and you can pretty quickly check whether or not the Python program works or not, which is really cheap, so it's cheap to verify.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Very cool. Maybe we should save that bit for tiny bit later. But 1 of the take home messages in your well, no, I mean, it's super interesting. And 1 of the take home messages on your new approach is that rather than producing explicit programs, you are evolving descriptions of programs. And Francois is a neurosymbolic guy. He thinks that we need to have a symbolic substrate where we, you know, represent the the kinds of problems that that we can do. And we need to do this kind of compositional form of of of intelligence. So we need to kind of be working in the symbolic layer, but perhaps guided by, you know, deep learning models. But maybe we we should get to that in in a minute. So in your in your first solution, it was an evolutionary approach. It was using Sonnet 3.5, and you had about 4 iterations, I think. And and, essentially, you you know, you were working on the ARC challenge, you were producing these programs through evolution. Maybe just for folks that don't know about the ARC challenge as well, could you introduce that and and get into your solution? Sure. Yeah. So the ARC challenge is kind of like an IQ test for machines. It's a set of input output grids, and the whole point is to be able to figure out to transform input grids into output grids given a common transformation rule. And so, what's interesting is these are really easy for humans. Right? The average human gets around 75% accuracy on ARC v 1, and at the time, the best language models, GPT 4, SONNET 3 5, was getting maybe 5%. And so, yeah, basically, the idea is you have a few training examples and then you're you're you're trying to extrapolate the transformation rule on the final test example. And so I approached this. I was actually inspired by Ryan Greenblatt who had a solution earlier, which was to generate a ton of Python programs that would encapsulate the transformation role. And Python programs are great because they're deterministic, and you can pretty quickly check whether or not the Python program works or not, which is really cheap, so it's cheap to verify. And you can be relatively sure if the Python program works on all of the training examples, that it'll work on the test example. So, I started with his approach, but then I noticed that, you know, the language models actually struggled on first attempts. Even if you ask the language model a thousand times to generate Python programs, they were always off by small amounts on easy tasks, which I thought, you know, presumably, it's in their distribution. They should be able to solve this. So, what I found is that actually by taking the top, programs, the top performing programs, and then running that in a revision loop, so asking, Sonnet 3 5, hey, here's what you got wrong, here are the cells you got wrong, here's your original Python program, improve it, that started to really work well. And then I thought, well, why not just increase the depth, right? Why not ask it 10 times to revise until I'm happy with the solution, passes some sort of accuracy threshold? So that's kind of how I was inspired by it. And I didn't think of it as evolutionary at first. I was just thinking about broadly what would work. And over time, I kind of understood there was something a bit deeper going on here, which is that evolving solutions is a powerful technique generally. And I think it's actually going to play a role, you know, in future technologies. But, yeah, that's generally guided by just intuition. Yeah. I I had Ryan Greenblatt on the show. I'm a huge fan of of his. He's a very, very smart guy. And I asked him a similar question because he he did this iteration, right, where where you have a certain depth of of iterations. And I guess 1 approach is that you have a, like, a a shallow method. Right? So you just tried 200 different variations. And in your in your blog post, you kind of said there's a there's a Goldilocks zone where you want to have a certain number of tries of, you know, different, you know, different variations of things. But you also want to be able to refine your solution because that allows you to do this kind of composition. And composition is very, very important for things for problems that require iteration. And indeed, the second version of the ARC challenge, I I think the tasks were selected so that they had at least a couple of iterations in them, which meant that they needed to have this depth. Can you talk about that trade off?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence