Evidence receipt / evaluation
Published · transcript-backedDries Smit: evaluation
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.”
Source trail
Everything needed to verify it.
- Speaker
- Dries Smit
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think. And time is not in question for the solutions. So I think that's really interesting and the reflections and looking back at the past, at your past experiences makes LLMs very flexible model for this the solution. Yeah. And interestingly enough, also a lot of even though there shouldn't be any priors, there's already a lot of the game priors are already included in the LLMs. So for example here we're looking at a maze, And maze is something that every LLM, even the small ones, will know and will recognize from beat images or ASCII grids. So even though that are not all the priors are stripped away, there's still enough like game priors that the LLM can lean on, and those are not encoded in any reinforcement learning model to or like a pure Goose solution wouldn't have that encoded anywhere in it, that the mace is a thing, and that really helps to direct the reasoning and the actions of the model. So the 1 of the main constraints was basically 2 weeks. I joined the competition late, 2 weeks to to do it. So there's a lot of things that can be improved, but I think it was a good initial solution. So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit. So the basic idea was we can't pre train. We don't have enough environments. Let's try to do test time training like we did for ArcGI 2. But how do you do that? Like you can try reinforcement learning, which I have a lot of experience with, but it's it's not that obvious how you do reinforcement learning because every time so you get a new game at test time, you have only real reward you have is a level transition. But once you transition a level, you never go back. So it's not like you're you're trying various routes to try and optimize for passing a level, you just have to pass it once. So it didn't really make sense to use pure reward based RL. I did a lot of research in the past in curiosity of unsupervised RL and I thought perhaps that's a better approach where we don't have explicit rewards, we're just optimizing for curiosity, like exploring new areas that the agent hasn't explored before. So I tried a bit of like world modeling. So the basic idea is the games are deterministic so you can use a world model to perhaps have the input frame, then the action and try to predict the the next frame. And then if you can't predict it well, you have the policy explore that more and that's the reward you're optimizing. So that was a promising approach. I couldn't get in it working in 2 weeks. So what I just didn't assume is that any frame change is interesting. You wanna explore things where the state changes. You don't wanna explore things where the state stays the same. So that was the basic derivation of the or like path I followed towards getting the action model implementation for Stochastic Goose. And then it was down to basic reinforcement learning things like how do you get it to learn within 1000 time steps because you have a 100,000 times of it, you need to start taking useful actions in like 1000 time steps. So it was like using a replay buffer, hashing the experience so doesn't go over the memory a little bit. You have to do some prioritized experience replay and some basic engineering stuff to get work with that competition. Yeah, you mentioned exploration as well. Guess there's a bit of an elephant in the room which is that Cholet is talking about the acquisition and synthesis of abstractions. And when reinforcement learning folks talk about exploration, it seems to be in quite a surface superficial way. So in terms of like entropy or things changing. And do you think is that in any way against this idea that we can acquire deep abstractions about the domain? Yes. Definitely for sarcastic goose.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.