High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Tim Scarfe: belief

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
belief
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…s say 2x or 3x above the human baseline, you're already close to 0. And yep, even though it's slow, it just helps. Avoids kind of untractable for us. We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position. But if you have to think at every step with the same computational capability or same budget, then yeah, you might be misrepresenting where you should go at the start, but at the end for straight lines, for example, you don't have to think that much, can just batch those actions. So that's yeah, where the coding agent idea came from and we also had a lot of 2 good literature results where they scaled using Opus models. 1 was ArgenTica and the other 1 was the RGB agent which also showed good results given no concrete constraints, you could use closed source models and we took that as inspiration. Yes. You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance. And it's quite an interesting discussion whether or not this is transduction because even though it's chain of thought it looks like a form of induction in the sense that it's a rationale that could be cross applied in the future. So you could use the memory in the agent. You could do some kind of library transfer and make it inductive. But at the moment, if it's only for the sole purpose of this particular problem, would you call it a transductive method? So I would call our ArcGI 2 solution more transductive and this slightly more inductive. It's it's exactly as you mentioned. You can actually read the reasoning trace and understand when it's understanding the game and making progress and when it's not. Yeah. Because it has this really reasoning chain of thought, which is in English and you could reasonably understand it. So yeah, I would actually say it's more inductive. Like previous attempts, as we mentioned was where the agent actually just directly predicts the actions. That will be more transductive and that doesn't seem to work at the moment. We obviously, there's a lot of ideas to try. I'm sure the community will come up with something interesting to make that work. But for now, this this seems to be the way for us. Yeah. I think the action efficiency make this makes this problem really interesting…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence