Evidence receipt / uncertainty
Published · transcript-backedStefano Viel: uncertainty
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“Like, for instance, if you give them an auto research task, they might overfocus on details and try kind of not see the big picture and not kind of zoom out and see things from Fireball. And they might, I don't know, start optimizing some hyperparameters for, at 0.”
Source trail
Everything needed to verify it.
- Speaker
- Stefano Viel
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…because something that would just sit in your scratch pad, you can just paste it in a prompt and see if it sticks. So you can definitely experiment overnight with a new idea. I think the really the added value that I see of the lab right now is and what the AIs are missing, what the what the coding agents are missing right now is like literally the the the attention to detail. So this task is very long. Like, if we have a long context, the execution of 1 game might take an hour just for it to solve 1 level. And just seeing and and if you if you could in in principle point auto research at this and say, hey, just improve the score. Right? But that doesn't work because it doesn't really it cannot really it we have tried that and that has not really led to improvements because it doesn't really understand where the agent fails. It's not simply as Andreas Carpathi's other research challenges where you just minimize a loss and there's just parameters. You need more insight to this task. And I think that's where a lot of improvements came from not from like, it it came from reading deep into the program, into the logs and understanding what the issues are and improving those 1 by 1 rather than just like, vibe coding another 2,000,000 lines of code. With coding agent trying to optimize this task, we observe kind of similar failure mode that what we observe when they play games. Like, for instance, if you give them an auto research task, they might overfocus on details and try kind of not see the big picture and not kind of zoom out and see things from Fireball. And they might, I don't know, start optimizing some hyperparameters for, at 0. 01% improvement while there is, like, some other big improvement that could be done, which is kind of in another obstruction level. And we see very similar failure mode in games where if the agent at the beginning of the gameplay, they start they get kind of locked in on the wrong hypothesis, it's extremely hard that they escape that. Like, they they kind of convince themselves that that's the right path and the right hypothesis to go, and they're not able to escape that. So it's quite interesting that, yeah, across models like smaller and frontier 1, we observe kind of similar failure mode of them kind of not being able to jump between a obstruction level and not being able to zoom in in the details and then seeing the higher level picture at the same time. I just wondered like how do you guys think about the abstraction mountain? Because for Cholet, it is bottom up. So you start with the core knowledge, know, like agents, spatial knowledge, stuff like that and and you kind of synthesize upwards. Whereas LLMs are kind of interesting because they learn these fractured and tangled representations that are quite high level but also quite generalizing. And then you can stitch them together and you can repurpose them and cannibalize them in different contexts. But it's always a little bit lossy. So if you're doing this library learning and transfer you always have this issue of like well yeah there's this abstraction over here and it kind of works but it's not quite what I want but I can still use it anyway. I mean is this kind of reasoning something that should be 1st class or do you think it could just be implicit in the harness if we had a memory system? So you kind of you were just describing how how LLMs will have…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.