Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“Because there is a bit of an elephant in the room and in the scene at the moment, I think so many people just just don't have such a crisp understanding.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Right. So I think test time fine tuning would be the way to, like, fundamentally make it adaptable. But I also think, you know, Chole hits it, a core problem with language models, which is they their reasoning is domain specific. In my blog post, I described that when you train a language model to reason about math, for some reason, most of the reasoning circuits it just gained live in the math weights. And then you try and train it on science, and it gets some generalization, but not as much as you would want, and I think not nearly as much as what what humans get. Humans have this this generalization engine that is our reasoning capability, and this is this is the fundamental hole in language models today. And I think in fact, actually, would say I I actually I, you know, generally agree with Francois, and he says, you know, language you can always teach a language model skill. Right? But it's the meta skill. It's the skill to create the skills that is AGI. And to me, that's reasoning. Like, reasoning is that meta skill. And so, to put it another way, I think if you fundamentally learn the skill of reasoning, you should be able to then apply that skill to learn all the other skills. That is the meta skill, and we need to figure out and so that is the fundamental problem, and you need to do whatever you can, kick whatever weights out you need to align the model to reason, and then from there, you have a foundation from which you can actually build general intelligence. So I guess I don't know if that was that was maybe a higher level answer to your question, but I think, you know, how how we and what I'm what I'm focused on is really just fitting all of reasoning into these models, and I don't really care what else is left. I just want all of reasoning in. Yes. I I I pretty much agree. And I mean, you probably know that I'm Charle's biggest fan, so I've obviously, you know, been a huge fan of his for years. But by the way, he's just released revision 3 of his deep learning with Python book. And I recommend folks to read chapter 19. You can actually read it online for free. And he sketches out this entire vision, you know. It's it's so exciting. And I I think that just to see it so beautifully articulated. Because there is a bit of an elephant in the room and in the scene at the moment, I think so many people just just don't have such a crisp understanding. But the only departure that I make with Chole and and and yourself, Jeremy, is that I think it's you know, Charle really focuses on behavioral tests of intelligence, you know, like, so it's so it's reasoning if it can pass the test and it can actually, get the right answer. And I think we need to go this is where I was kind of talking about the systematicity and the symbolic AI. I think how you got there is important. Right? So I think it's possible to get the right answer for the wrong reasons. And I think that if we have a system which has semantics, so we actually know what these symbols mean and we've composed them together in a principled way. Not only to get the right answer for the right reasons, but to be evolvable. So to have like an efficient epistemic base that allows us to go on to acquire new knowledge in the future. And that to me points to this need to have a mechanistic, like, you know, like how are we acquiring this knowledge view. Did would you agree with that? Yes. And I I think about it a bit differently. So let me know if what I say is aligned with what you think.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.