Evidence receipt / belief
Published · transcript-backedDries Smit: belief
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“I think the 1st 4 places was basically brute force algorithms that just search over a large space of actions but do some basic form of filtering and you can get a very good score.”
Source trail
Everything needed to verify it.
- Speaker
- Dries Smit
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Because didn't the preview have some, it was a bit brute forceable, wasn't it? You were talking about that earlier. Yes, yes, exactly. So I guess that was the main point about the preview competition to wish to show whether it's something you could easily exploit. And indeed there was like the Stochastic Goose algorithm and few other algorithms actually. I think the 1st 4 places was basically brute force algorithms that just search over a large space of actions but do some basic form of filtering and you can get a very good score. Like I could solve 2 games, almost solve the 3rd game on the private set. The game is also too easy so they upped the difficulty level. And for example, 1 other thing was the timing bar which only changed when you actually executed a valid action. So you could easily learn what actions are valid or not. On the new set of games and Arcadia setup, I don't know if any of you have noticed anything that is easily exploitable or badly designed. Yeah, I think it's implemented much better now and that's why the scores haven't shot up initially. It's still at like 1%. So I think it's still challenging setup. I guess the million dollar question though is do you think it's possible in principle to do really well on ArcGI 3 and be no closer to AGI? Yes,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.