High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Michal Tesnar: preference

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think.”

— Michal Tesnar

Source trail

Everything needed to verify it.

Speaker
Michal Tesnar
Attribution
Verified speaker
Claim type
preference
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. Because it has this really reasoning chain of thought, which is in English and you could reasonably understand it. So yeah, I would actually say it's more inductive. Like previous attempts, as we mentioned was where the agent actually just directly predicts the actions. That will be more transductive and that doesn't seem to work at the moment. We obviously, there's a lot of ideas to try. I'm sure the community will come up with something interesting to make that work. But for now, this this seems to be the way for us. Yeah. I think the action efficiency make this makes this problem really interesting because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think. And time is not in question for the solutions. So I think that's really interesting and the reflections and looking back at the past, at your past experiences makes LLMs very flexible model for this the solution. Yeah. And interestingly enough, also a lot of even though there shouldn't be any priors, there's already a lot of the game priors are already included in the LLMs. So for example here we're looking at a maze, And maze is something that every LLM, even the small ones, will know and will recognize from beat images or ASCII grids. So even though that are not all the priors are stripped away, there's still enough like game priors that the LLM can lean on, and those are not encoded in any reinforcement learning model to or like a pure Goose solution wouldn't have that encoded anywhere in it, that the mace is a thing, and that really helps to direct the reasoning and the actions of the model. So the 1 of the main constraints was basically 2 weeks. I joined the competition late, 2 weeks to to do it. So there's a lot of things that can be improved, but I think it was a good initial solution. So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit. So the basic idea was we can't pre train. We don't have enough environments. Let's try to do test time training like we did for ArcGI 2. But how do you do that? Like you can try reinforcement learning, which I have a lot of experience with, but it's it's not that obvious how you do reinforcement learning because every time so you get a new game at test time, you have only real reward you have is a level transition. But once you transition a level, you never go back. So it's not like you're you're trying various routes to try and optimize for passing a level, you just have to pass it once. So it didn't really make sense to use pure reward based RL. I did a lot of research in the past in curiosity of unsupervised RL and I thought perhaps that's a better approach where we don't have explicit rewards, we're just optimizing for curiosity, like exploring new areas that the agent hasn't explored before. So I tried a bit of like world modeling. So the basic idea is the games are deterministic so you can use a world model to perhaps have the input frame, then the action and try to predict the the next frame. And then if you can't predict it well, you have the policy explore that more and that's the reward you're optimizing. So that was a promising approach. I couldn't get in it working in 2 weeks. So what I just didn't assume is that any frame change is interesting. You wanna explore things where the state changes. You don't wanna explore things where the state stays the same. So that was the basic derivation of the or like path I followed towards getting the action model implementation for Stochastic Goose. And then it was down to basic reinforcement learning things like how do you get it to learn within 1000 time steps because you have a 100,000 times of it, you need to start taking useful actions in like 1000 time steps. So it was like using a replay buffer, hashing the experience so doesn't go over the memory a little bit. You have to do some prioritized experience replay and some basic engineering stuff to get work with that competition.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence