High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Speaker unverified: commitment

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“I think that I think that what you're describing I'm actually not even totally convinced that continual learning is fundamentally the blocker, but I think if it is the fundamental blocker, that's actually incredible because we will solve continual learning.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
commitment
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. I mean, you know far more about this than I do, but I think the reason why fine tuning is so expensive is, you know, we have this continual learning problem. And when you fine tune a model on OpenAI, they're not just fine tuning it on the data you give them. They you know, just to stop this catastrophic forgetting problem, they presumably have to sample in a bunch of the original training data and maintain the distribution and so on. And and if if they did this for everyone, it would be insane. But I am excited about it just like you are because I I interviewed the architects and I I think they got first place on the on the private version last year. And they were doing this transductive active fine tuning. Right? So they and they actually said, by the way, that this is a a curious oddity with transformers that if you start with a, you know, almost like a virgin 8,000,000,000 transformer, it almost doesn't matter what it knew about before. You could just pretty much start training it from scratch on the ARC challenges. So they did a whole bunch of augmentation and active fine tuning, and they built an intelligent artifact. I mean, intelligence is domain specific as per Cholet, and they actually built the system which was per task adapting and solving the tasks, and they were updating the weights, and it was beautiful. So that was an existence proof of nothing else that this thing could work. And that was on a Kaggle notebook. Yes. You know, in 10 years, this is gonna be what? Like, the, you know, Apollo mission computer. I think that I think that what you're describing I'm actually not even totally convinced that continual learning is fundamentally the blocker, but I think if it is the fundamental blocker, that's actually incredible because we will solve continual learning. Like, that's something that's physically possible, and and I actually think, like, it's not so far off. Now the the, forgetful issue, that is a much more fundamental issue in my mind. And not just the the fact that you need to, every time you fine tune, you have to have some sort of very elegant, mixture of data that, you know, goes into this fine tuning process so that, there's there's no catastrophic forgetting. This is, I think, actually a fundamental problem. So I the the and and it's a fundamental problem that that, you know, even OpenAI has not solved. Right? And I think Francois is a great example, and I think this is an important example. You know, if you have the perfect weights for a certain problem and then you fine tune that model on more examples of that problem, the weights will start to drift and you will actually drift away from the, from the correct solution. His answer to that is, well, we could make these systems composable. We can freeze the correct solution, and then we can add on top of that. I think there's something to that. I think, actually, it's possible that there's a research direction where maybe we freeze experts, or may maybe we freeze, layers for a bunch of reasons that isn't possible right now or but people are trying to do that. But I yeah. I I think, fundamentally, compute is not the issue. I think it's this catastrophic forgetfulness. Yeah. So I'm inclined to agree. I've I've long dreamed about there being a docker for language models. Right? You you know, in docker, you can kind of freeze dry a state of, you know, like, let's say Linux operating system with an application with the security updates. You have these kind of immutable layers. And the composability that we often talk about could actually happen at the architectural level, and we could do dynamic model merging between different layers and and whatnot. That would be very, very exciting. So, know but also, just to come back to what you said before, I've never really heard this before. You're distinguishing, like, forgetting and learning. Right? When when we were talking about, you know, catastrophic forgetting and continual learning. Can you just sketch out that distinction a bit more?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence