Evidence receipt / belief
Published · transcript-backedJeroen Cottaar: belief
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“Does that really count as the old models can play games now if it takes them so much effort? I would say that is a real gap.”
Source trail
Everything needed to verify it.
- Speaker
- Jeroen Cottaar
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…I don't know, but Is that the same as arguing that if you went to a deserted island where there was some tribes folk that have never used computers before, do you think they would be able to do ArcGI 3? I think not immediately. But like, if you give them a 100 of these things, and you give them 1 per day, and you reward the people doing well, I would assume at some point they can. I do think it's fair to say they probably need a lot less training resources in this amount of, I the amount of compute required to learn it, than right now a frontier model would use. And even when we say they can reach 36%, that is true, but it costs like a few thousand dollars, which is a lot more Although it's hard to, of course, want a lot more compute than the human beating these games Yes. Is spending. So, yeah. Does that really count as the old models can play games now if it takes them so much effort? I would say that is a real gap. And also, I'd say being able to do it from tier models of not sorry, with small open source models would demonstrate a lot more. It's indeed what we're doing in the capital competition. 36% might be misleading as a number if you don't look behind it. So what it really measures is action efficiency. So correct me if I'm wrong, but it's the ratio of human baseline divided by the number of actions the AI has taken, whatever model it is, on the level or a human, whatever whoever whoever the player is, and then squared as well. So this plays really adversarially to anything that's a little bit action inefficient. So 36% actually in this case doesn't mean we solve that approach solves 36% of the games, it solves way more of the games out from the training set, at least this number is on, but it just solves them inefficiently. So I think this is really important to emphasize that like the current frontier models are able to solve like, what is it, something like half of 2 thirds of the training games actually till the end, but just not as efficiently. What is the hardest thing in Arc AGI 3? Is it the goal acquisition or is it just simply the action efficiency? From what we see on the training set, so the the testing set, the private set is set to be harder. We don't know anything about it. No nobody outside of the organization has their organization has seen it. We we don't know how how hard it is. But as at least on the training games, we see that the that the LLMs can acquire the correct goals and pursue them somewhat effectively. The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this. And keeping consistent knowledge of everything that has happened over such a context is a is a major engineering challenge at this moment. What is…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.