Evidence receipt / prediction
Published · transcript-backedJeff Beck: prediction
31 Dec 2025 Machine Learning Street Talk Bayesian Brain, Scientific Method, and Models [Dr. Jeff Beck]
“I actually the reason why I say transformer comes with an asterisk is because a lot of the things that transformers have been that that people believe that the transformer enabled, I think really resulted more from scaling.”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Beck
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 31 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes exactly. You get to you get to suddenly. This is 1 of things I love. I mean this is what I love about the community in fact is that we now have a relatively common language to discuss a huge variety of different things. Yeah. Now, course, that means we often end up talking cross purposes, but that's half the fun. Right? So I often ask people in the business, like, what what what changed? Like, what's you know, what's you know, why did we have this, like, massive explosion in, you know, in AI development over the last several years? And I get 3 there there are 3 common responses, and I agree with every single 1 of them. Auto grad, right, the transformer, but why the transformer is something that I I often disagree with with you all about. The transformer architecture and just the the the the ability to scale things up in a manner that we haven't haven't really seen before. I actually the reason why I say transformer comes with an asterisk is because a lot of the things that transformers have been that that people believe that the transformer enabled, I think really resulted more from scaling. And my the point, you know, the point of evidence that I like to cite is like Mamba. Mamba, which is a states which is a traditional state space model. It's basically a common filter, but like on steroids. They scaled it way up and yet and now it's, you know, got they've, you know, Mistral has their very nice like coding agent and it works pretty darn well. Right? They got a lot of the same functionality with a completely diff with a you know, a completely different architecture simply by virtue of scaling. So transformers get a get an asterisk. I think the biggest thing was AutoGrad. Right? And AutoGrad turned the development of artificial intelligence from being something that was done by, like, carefully constructing your neural networks and then writing down your learning rules and going through all that painful process that was tick tick for. And they turned it into an engineering problem. It made it possible to experiment with different architectures, different networks, different nonlinearities, different structures, different ways of like getting your memory in there in different ways. And all this fun stuff that allowed people to just start trying things out in a way that we couldn't do it before. And then we what did we did? We we suddenly discovered, turns out backprop does work. I mean when I was a young man, like backprop was considered a nonstarter for 2 reasons, right? 1 is it's not brain like, which is true. Right? The brain does not use backprop. And the other 1 was a vanishing gradients. Oh, you'll never solve the vanishing gradients problem. So, oh, it'll always be unstable. And and yet nonetheless, once we turn it into a new dream problem, sort of playing around with tricks and hacks and certain kinds of nonlinearities and relues and this and that. We discovered that, oh, no. In fact, there are ways around this. nto a new dream problem, sort of playing around with tricks and hacks and certain kinds of nonlinearities and relues and this and that. We discovered that, oh, no. In fact, there are ways around this. You just, you know, we just, you know, weren't gonna discover them by like playing with equations. We had to actually start it. So we turned it into an engineering problem. As soon as it got turned into an engineering problem, you know, that's what enabled the hyperscaling, which is what led to all of this all all of this, you know, these great developments over the last several years. What got lost in the mix though was the notion that that that that there's more to artificial intelligence than just like function approximation. We got really good function approximators, but that's not the only thing you need to develop like proper AI. Right? You need models that are structured like the brain is structured. You need models that, you need you need models that are structured like how we conceive the world is structured. Certainly if you wanna have models that think the way we think. And that that got lost in the shuffle, and we're starting to see, you know, as as as we're starting to see the limitations and the faults and flaws of of of these approaches, and starting to see them not living up to the hype, which I think is like now it's standard that like like AGI is no longer I don't know if you read the other day, at least according to, you know, the experts in the field at the top of the best companies in the business, like AGI is no longer, like, a huge priority. Right? And that they're they're they're dialing back the rhetoric surrounding that, in part because I think that they've begun to realize that, like, just function approximation isn't going to deliver. That was just hype. Right? We do need to do something different. We do need to start get bringing in what we know about how the brain works, right, if we're ever going to get to something that is a human like intelligence. And that was the starting point for us, you know, about a year or so ago is that we were sort of like, yes, let's do the same thing for cognitive models. Like, let's talk about let's take what we know about how the brain the brain actually works. Let's take what we know about how people actually think about the world in which they live, and start building an artificial intelligence that thinks like we do by incorporating these principles. And this means this means basically creating a, you know, a modeling and coding framework for building brain like models at scale, and that's like the critical element because obviously scaling was a was a big part of the solution. And right now, most of the work in the active inference space, as I'm sure you're aware, is not at scale. There's very little, like, active inference work that is active inference at scale. Most of the models are, like, relatively small toy grid world y type models. And part of the reason for that is that, you know, it is in fact difficult to scale Bayesian methods. Now that also has now begun to change, right?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.