High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Dries Smit: prediction

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“they have the intention, they have something that the coding agents don't have because you know, the bull case, wouldn't it be amazing if you could stick the requirements in and the agents themselves would kind of understand, oh I see what they meant and they could evolve the requirements.”

— Dries Smit

Source trail

Everything needed to verify it.

Speaker
Dries Smit
Attribution
Verified speaker
Claim type
prediction
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…humans have the taste, they have the intention, they have something that the coding agents don't have because you know, the bull case, wouldn't it be amazing if you could stick the requirements in and the agents themselves would kind of understand, oh I see what they meant and they could evolve the requirements. So I think in the short term, I can't say for the long term at least what we're trying to do is find the areas where we can see the model still failing like high level conceptual ideas about the problem. And then we try to make sure when they say we're reviewing code or creating code that we ask specific around that areas and make sure it's correct and understands us well. But, yeah, I guess it's always gonna be a game of cat and mouse where we're gonna perhaps hand over more to the AI assistants. But in some areas it might just totally fail and we need to take responsibility there and make sure we understand what it's doing and it understands what we want. But yeah, longer term, I guess if you're AGI pulled or a super intelligent pulled, I guess it's, yeah, it's up for debate with. So it's certainly accelerating us in doing experiments because something that would just sit in your scratch pad, you can just paste it in a prompt and see if it sticks. So you can definitely experiment overnight with a new idea. I think the really the added value that I see of the lab right now is and what the AIs are missing, what the what the coding agents are missing right now is like literally the the the attention to detail. So this task is very long. Like, if we have a long context, the execution of 1 game might take an hour just for it to solve 1 level. And just seeing and and if you if you could in in principle point auto research at this and say, hey, just improve the score. Right? But that doesn't work because it doesn't really it cannot really it we have tried that and that has not really led to improvements because it doesn't really understand where the agent fails. It's not simply as Andreas Carpathi's other research challenges where you just minimize a loss and there's just parameters. You need more insight to this task. And I think that's where a lot of improvements came from not from like, it it came from reading deep into the program, into the logs and understanding what the issues are and improving those 1 by 1 rather than just like, vibe coding another 2,000,000 lines of code. With coding agent trying to optimize this task, we observe kind of similar failure mode that what we observe when they play games.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence