High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Dries Smit: evaluation

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.”

— Dries Smit

Source trail

Everything needed to verify it.

Speaker
Dries Smit
Attribution
Verified speaker
Claim type
evaluation
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…search and also the programs, the Python programs it creates is like executable Python programs to abstract and also build a some sort a simplified world model and search over that world model using actual like algorithms like breakfast search. There is understanding debt, right? That you're building this really really complex thing and after a while have you noticed in Claude code, they don't even show you the code anymore? Yeah. Right? You you can expand it, but by default, a lot of people don't even look at the code. So do do do you think that you need to be familiar with the deep abstractions in your code base in order to kind of build mental models and evolve it and extend it? You know, do you get lost in no man's land? Yes, this is a, yeah, if you will. This is an active discussion within our team. We don't have a discrete answer yet, but I think what we've found is like you need deep understanding of some important parts. Some other parts like let's say a web viewer, you can vibe code more easily because that's not really if it breaks, it's fine. But like the core logic for like they say the implementation and also evaluation is especially important. Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up. But things are moving so fast that you can't just manually yeah, if you just manually write everything, you can't keep up with the rest of the team. So yeah, I guess this is an active discussion, it's difficult. As for me, my career has been going on a bit longer, and let's say most of it, for most of my career, there were no coding agents, which means I also have a bit of an opportunity to still make use of patterns that have been useful in the past. And I think 1 important 1 that, as a team, we are more and more learning to use properly is requirements based engineering. So we will formally write requirements, let's say, really following the detailed prescripts like number requirements and to specify how they are tested. I mean, we might still have the coding agents helping there, but mostly handwritten. We will review that as a team, and from there, the coding agents do we can much more confidently hand it to a coding agent to implement than if it's just a Viper single prompt.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence