High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Dries Smit: evaluation

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?”

— Dries Smit

Source trail

Everything needed to verify it.

Speaker
Dries Smit
Attribution
Verified speaker
Claim type
evaluation
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…I'm guilty of always like, you know, suggesting what Charle thinks. I don't really know what he thinks, but you know, I've I've got a pretty good simulation of Charle in my mind. But I I I think he's he's influenced by this kind of nativism psychology type thing. And he thinks that a lot of the reasoning we do do is almost platonistic. Know, that that you know somehow the laws of nature imputes these primitives into our mind and we compose these primitives together for certain classes of problems. So we can do abstract system 2 reasoning and for those types of problem we do compose these things together. But you're absolutely right, there's so many things in the world are actually really complicated. Right? You know, like navigating relationships or even, you know, path finding in a complex environment, you know, on the tube network or something like that. So we do a little bit of both, but there's there's at least a pocket of pure reasoning. I think that's what Charlotte thinks. Would you consider LLMs pre trained on the Internet? And then if you apply it to a new pattern using like in context learning, is that a form of okay. Let's say with reasoning as well, would that not be a form of like perhaps skill acquisition on the fly because it learns obviously, there's a lot of core knowledge as well, but it's also a different part where it has to adapt on the fly. So you can give some example that might not be on the internet and might adapt. Sometimes it fails, it is still a distribution, but at least you get some adaptability. And then for example, when you started prompting it, think of like a chain of thought prompting, it started improving because it had more time to reason and adapt to what you're asking. Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet? LLNs trained to do reasoning, they are intelligent because intelligence is just adaptation. So it's I know this now and I need to combine what I know in several steps you know, to solve a particular task. In the partial knowledge regime where we have a verifiable function and we can do hill climbing. But what we see though is that they combine together fractured and tangled representation. So they get the right answer but for the wrong reasons. So they they create a spaghetti monster. But what we do is is we take a more valid path and we can acquire and reuse abstractions. The models themselves they don't seem to work bottom up. They work they they have an an understanding. They have representations but they're very fractured and entangled. And and that's still useful but that's a kind of statistical low level intelligence. And I think you could in principle still do abstraction from the fractured and tangled representations but we're kind of in statistical land. We're not it's performance not competence. I don't have strong views on this, but I…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence