Evidence receipt / uncertainty
Published · transcript-backedTim Scarfe: uncertainty
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“I don't know if you've seen the recent couple of papers that are applying it to transformers, you know, where essentially it's it's kind of a step towards probabilistic models where you actually have this uncertainty quantification.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…And I I think you can build that with language models. And I think the fact that we're slowly climbing, our ability to synthesize new information is a testament to what I'm saying, which is when reinforcement learning and reinforcement learning, you know, with verifiable rewards is is fundamentally ensuring whatever circuits led to the right answer, they must be consistent with the deductive tree. So it's basically like, can you replace all of your pretrained weights with weights that are, coherent from the environment? The problem is all you know, there's so many weights. There's so many weights from pretraining that it's very difficult. So I actually think a part of, you know, 1 of my, I guess, hot takes is that pretraining in many ways slows down reasoning. It makes it harder to reason because I I think I think the analogy I draw is, you know, you have consultants that know the names for things but couldn't deduce the thing. And then you have Feynman, right, who can deduce anything. And reinforcement learning is is turning your consultant into Feynman. And this is yeah. This is what I'm most interested in. And I think it's it's interesting because you get to play at both sides. You get to play with the pretraining of, okay. Well, maybe we shouldn't include these things and let the model figure it out in reinforcement learning. Well, there's no guarantee if we pre train it like this, you know, it's gonna have the proper reductive, circuitry, so maybe this is best, you know, left for post training. And I I think this is not this is this is a hot this is not what people currently think. I think people think, let's jam as much information as we can in pre training, and then we'll reinforcement learn when need when we need. But I think this this could be incorrect. Yeah. And I I pretty much agree with you just with the caveats we discussed previously that if we could do that on the basis of representations that are actually, like, grounded in the world rather than things that just, you know, happen to give you the right answer for the wrong reasons, that's that's absolutely true. Just a little bit of a a curve ball. So I think in the first version of the article, you said you're inspired by Jan Lakun's JEPPA, these joint embedding prediction architectures. And he's also a big advocate of energy based models, which are really cool. I don't know if you've seen the recent couple of papers that are applying it to transformers, you know, where essentially it's it's kind of a step towards probabilistic models where you actually have this uncertainty quantification. You can do counterfactuals and you actually have to solve an optimization problem at inference time so you can do adaptive computation. It's all very exciting. But I still have some reservations, but do you think, like, architectures like that are exciting? I think they are exciting. I am a bit less excited about them, not because of their merits, but because I think I was underrating transformers when I wrote that. So this was really before I started do it being an actual researcher and, you know, building transformers, actually coding with them. Since then, I've had a new appreciation for language models. I think where I was coming from it was it seems like language models are overfitting to the next token. JEPA is so interesting because, oh, all of a sudden, you have them predicting concepts. Fundamentally, we care about concepts, the words don't really matter, it's the concepts that matter. But I think that language models do operate on the conceptual level in the hidden layers. That was something that I slowly came to realization of. So I think there's lot of potential in JEPA frameworks. Think they're really cool. I hope people keep pulling on them. But I think most of the benefits I thought that came from JEPA exist in language models. I just didn't see it at the time.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.