Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“We need not detain us now, but I think in traditional machine learning, transduction means that the test example is a function of your prediction.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…important that the checker was stronger than the actual instruction creator, which I I think is, interesting. But, yeah, that just highlights, you know, the the trade offs with using natural language. It's it's you can express, you know, much more concisely programs that you wanna run, but then they're not runnable programs. You actually have to check them inductively. So that was the trade off, but it was worth it for Arc v 2. Yes. So so fascinating. And just for the audience, we've been using transduction and induction to distinguish predicting the solution space versus predicting a program. I had this this discussion with with Clement Bonnet. We need not detain us now, but I think in traditional machine learning, transduction means that the test example is a function of your prediction. I had this discussion with the architects as well that when you have this natural language description, natural language is more expressive, which simply means that there are more degrees of freedom. And this is the beauty of of LLMs that there's this huge kind of space that you're traversing around. And when you use natural language, you can just traverse to more places in that space more easily. So it seems like it would be a win. I'm And really fascinated to to find out whether that is just like a huge component of your solution. Because on Eric's solution, he's still predicting programs and still doing very well. So I'm not I'm not sure about that. And the other thing is I wasn't entirely sure whether you are actually using a transtactive method. So in your solution checker agent, is it directly going to the solution space, or is it generating a program and testing it? In the checker? Yeah. In the checker, it takes in the natural language, and then it outputs a grid. That's all it does. It just outputs a grid.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.