Evidence receipt / evaluation
Published · transcript-backedMichal Tesnar: evaluation
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this.”
Source trail
Everything needed to verify it.
- Speaker
- Michal Tesnar
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…But like, if you give them a 100 of these things, and you give them 1 per day, and you reward the people doing well, I would assume at some point they can. I do think it's fair to say they probably need a lot less training resources in this amount of, I the amount of compute required to learn it, than right now a frontier model would use. And even when we say they can reach 36%, that is true, but it costs like a few thousand dollars, which is a lot more Although it's hard to, of course, want a lot more compute than the human beating these games Yes. Is spending. So, yeah. Does that really count as the old models can play games now if it takes them so much effort? I would say that is a real gap. And also, I'd say being able to do it from tier models of not sorry, with small open source models would demonstrate a lot more. It's indeed what we're doing in the capital competition. 36% might be misleading as a number if you don't look behind it. So what it really measures is action efficiency. So correct me if I'm wrong, but it's the ratio of human baseline divided by the number of actions the AI has taken, whatever model it is, on the level or a human, whatever whoever whoever the player is, and then squared as well. So this plays really adversarially to anything that's a little bit action inefficient. So 36% actually in this case doesn't mean we solve that approach solves 36% of the games, it solves way more of the games out from the training set, at least this number is on, but it just solves them inefficiently. So I think this is really important to emphasize that like the current frontier models are able to solve like, what is it, something like half of 2 thirds of the training games actually till the end, but just not as efficiently. What is the hardest thing in Arc AGI 3? Is it the goal acquisition or is it just simply the action efficiency? From what we see on the training set, so the the testing set, the private set is set to be harder. We don't know anything about it. No nobody outside of the organization has their organization has seen it. We we don't know how how hard it is. But as at least on the training games, we see that the that the LLMs can acquire the correct goals and pursue them somewhat effectively. The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this. And keeping consistent knowledge of everything that has happened over such a context is a is a major engineering challenge at this moment. What is harder in Arc v 3 compared to the other 1 is this interplay between exploration and solving the game. Because in Arc v 1 and Arc v 2, you would get all of the information in a static frame as you are given the puzzle. Instead in Arc v 3, you are given the game, but just from the 1st frame, you cannot understand what needs to be done. So you start need to start interacting with the game. And through that interaction, you gather information of what the game is about and you start understanding how to solve it. And at the same time so you need to try to understand what the game is about and try to solve it at the same time. And this interplay is very hard to kind of explain to the agent how they should do it in an effective way such that it's general and generalized to across all games. So I think this is probably 1 of the key part that is hard about r p 3. What we want is abstraction…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.