High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Tim Scarfe: evaluation

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“We need not detain us now, but I think in traditional machine learning, transduction means that the test example is a function of your prediction.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
evaluation
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…important that the checker was stronger than the actual instruction creator, which I I think is, interesting. But, yeah, that just highlights, you know, the the trade offs with using natural language. It's it's you can express, you know, much more concisely programs that you wanna run, but then they're not runnable programs. You actually have to check them inductively. So that was the trade off, but it was worth it for Arc v 2. Yes. So so fascinating. And just for the audience, we've been using transduction and induction to distinguish predicting the solution space versus predicting a program. I had this this discussion with with Clement Bonnet. We need not detain us now, but I think in traditional machine learning, transduction means that the test example is a function of your prediction. I had this discussion with the architects as well that when you have this natural language description, natural language is more expressive, which simply means that there are more degrees of freedom. And this is the beauty of of LLMs that there's this huge kind of space that you're traversing around. And when you use natural language, you can just traverse to more places in that space more easily. So it seems like it would be a win. I'm And really fascinated to to find out whether that is just like a huge component of your solution. Because on Eric's solution, he's still predicting programs and still doing very well. So I'm not I'm not sure about that. And the other thing is I wasn't entirely sure whether you are actually using a transtactive method. So in your solution checker agent, is it directly going to the solution space, or is it generating a program and testing it? In the checker? Yeah. In the checker, it takes in the natural language, and then it outputs a grid. That's all it does. It just outputs a grid.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence