Evidence receipt / evaluation
Published · transcript-backedDries Smit: evaluation
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.”
Source trail
Everything needed to verify it.
- Speaker
- Dries Smit
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…agent preview competition last year, I actually tried something completely different. So the goal of that competition was basically to test whether there is obvious solutions that break the mold, breaks what they're actually trying to accomplish. And there was and the solution was basically just to brute force actions. So what that solution did was it basically searches through the large action space for ArcGIS 3. We have more than 4,000 actions which makes this difficult problem to do. But you could theoretically search for all that or just randomly search for actions. What I did for Stochastic Goose was basically to search for large number of possibilities but I tried to only search through actions that results in a frame change. So initially, if you took an action that did not do anything in the game, nothing would change, not even the timer bar. That allows you to to model that behavior to see, okay, does this action change the time bar? If it does not, we downvote that action in the future for that given frame. There was also only 3 games, so you have to be selective of how you improve things. You can't pre train because you'll just overfit. Exactly what I did was I just used a action model that learns, yeah, which frames are valid for a given transition and then it was more about the engineering of being able to learn within 100,000 actions because that was kind of the max action limit we could get out in the time limit we had. So that seemed to have worked well. It got it solved I think it got 18 levels completed out of the 3 games, which was I think it solved 2 of the games as well in the time limit that was provided to us. But they then hardened the competition specifically against that. So the new competition, the games are much harder, and the timer bar moves even if you use an action that's valid but doesn't really change anything in the game. And more importantly, they introduced action efficiency. So this makes it, yeah, very difficult to reinforce. You have to be very direct of how you explore and that's where LLMs come in. So you could do even though it's slower, just being able to somehow guide the exploration helps a lot. It's like the possible combinations is too many to just brute force and obviously your score goes to 0 quite quickly. So they've made it so that if you go just let's say 2x or 3x above the human baseline, you're already close to 0. And yep, even though it's slow, it just helps. Avoids kind of untractable for us. We tried other methods such as directly predicting s say 2x or 3x above the human baseline, you're already close to 0. And yep, even though it's slow, it just helps. Avoids kind of untractable for us. We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position. But if you have to think at every step with the same computational capability or same budget, then yeah, you might be misrepresenting where you should go at the start, but at the end for straight lines, for example, you don't have to think that much, can just batch those actions. So that's yeah, where the coding agent idea came from and we also had a lot of 2 good literature results where they scaled using Opus models. 1 was ArgenTica and the other 1 was the RGB agent which also showed good results given no concrete constraints, you could use closed source models and we took that as inspiration. Yes. You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance. And it's quite an interesting discussion whether or not this is transduction because even though it's chain of thought it looks like a form of induction in the sense that it's a rationale that could be cross applied in the future. So you could use the memory in the agent. You could do some kind of library transfer and make it inductive. But at the moment, if it's only for the sole purpose of this particular problem, would you call it a transductive method? So I would call our ArcGI 2 solution more transductive and this slightly more inductive. It's it's exactly as you mentioned. You can actually read the reasoning trace and understand when it's understanding the game and making progress and when it's not.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.