Evidence receipt / evaluation
Published · transcript-backedJeff Beck: evaluation
31 Dec 2025 Machine Learning Street Talk Bayesian Brain, Scientific Method, and Models [Dr. Jeff Beck]
“Certainly if you wanna have models that think the way we think. And that that got lost in the shuffle, and we're starting to see, you know, as as as we're starting to see the limitations and the faults and flaws of of of these approaches, and starting to see them not living up to the hype, which I think is like now it's standard that like like AGI is no longer I don't know if you read the other day, at least according to, you know, the experts in the field at the top of the best companies in the business, like AGI is no longer, like, a huge priority.”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Beck
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 31 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…you get to suddenly. This is 1 of things I love. I mean this is what I love about the community in fact is that we now have a relatively common language to discuss a huge variety of different things. Yeah. Now, course, that means we often end up talking cross purposes, but that's half the fun. Right? So I often ask people in the business, like, what what what changed? Like, what's you know, what's you know, why did we have this, like, massive explosion in, you know, in AI development over the last several years? And I get 3 there there are 3 common responses, and I agree with every single 1 of them. Auto grad, right, the transformer, but why the transformer is something that I I often disagree with with you all about. The transformer architecture and just the the the the ability to scale things up in a manner that we haven't haven't really seen before. I actually the reason why I say transformer comes with an asterisk is because a lot of the things that transformers have been that that people believe that the transformer enabled, I think really resulted more from scaling. And my the point, you know, the point of evidence that I like to cite is like Mamba. Mamba, which is a states which is a traditional state space model. It's basically a common filter, but like on steroids. They scaled it way up and yet and now it's, you know, got they've, you know, Mistral has their very nice like coding agent and it works pretty darn well. Right? They got a lot of the same functionality with a completely diff with a you know, a completely different architecture simply by virtue of scaling. So transformers get a get an asterisk. I think the biggest thing was AutoGrad. Right? And AutoGrad turned the development of artificial intelligence from being something that was done by, like, carefully constructing your neural networks and then writing down your learning rules and going through all that painful process that was tick tick for. And they turned it into an engineering problem. It made it possible to experiment with different architectures, different networks, different nonlinearities, different structures, different ways of like getting your memory in there in different ways. And all this fun stuff that allowed people to just start trying things out in a way that we couldn't do it before. And then we what did we did? We we suddenly discovered, turns out backprop does work. I mean when I was a young man, like backprop was considered a nonstarter for 2 reasons, right? 1 is it's not brain like, which is true. Right? The brain does not use backprop. And the other 1 was a vanishing gradients. Oh, you'll never solve the vanishing gradients problem. So, oh, it'll always be unstable. And and yet nonetheless, once we turn it into a new dream problem, sort of playing around with tricks and hacks and certain kinds of nonlinearities and relues and this and that. We discovered that, oh, no. In fact, there are ways around this. nto a new dream problem, sort of playing around with tricks and hacks and certain kinds of nonlinearities and relues and this and that. We discovered that, oh, no. In fact, there are ways around this. You just, you know, we just, you know, weren't gonna discover them by like playing with equations. We had to actually start it. So we turned it into an engineering problem. As soon as it got turned into an engineering problem, you know, that's what enabled the hyperscaling, which is what led to all of this all all of this, you know, these great developments over the last several years. What got lost in the mix though was the notion that that that that there's more to artificial intelligence than just like function approximation. We got really good function approximators, but that's not the only thing you need to develop like proper AI. Right? You need models that are structured like the brain is structured. You need models that, you need you need models that are structured like how we conceive the world is structured. Certainly if you wanna have models that think the way we think. And that that got lost in the shuffle, and we're starting to see, you know, as as as we're starting to see the limitations and the faults and flaws of of of these approaches, and starting to see them not living up to the hype, which I think is like now it's standard that like like AGI is no longer I don't know if you read the other day, at least according to, you know, the experts in the field at the top of the best companies in the business, like AGI is no longer, like, a huge priority. Right? And that they're they're they're dialing back the rhetoric surrounding that, in part because I think that they've begun to realize that, like, just function approximation isn't going to deliver. That was just hype. Right? We do need to do something different. We do need to start get bringing in what we know about how the brain works, right, if we're ever going to get to something that is a human like intelligence. And that was the starting point for us, you know, about a year or so ago is that we were sort of like, yes, let's do the same thing for cognitive models. Like, let's talk about let's take what we know about how the brain the brain actually works. Let's take what we know about how people actually think about the world in which they live, and start building an artificial intelligence that thinks like we do by incorporating these principles. And this means this means basically creating a, you know, a modeling and coding framework for building brain like models at scale, and that's like the critical element because obviously scaling was a was a big part of the solution. And right now, most of the work in the active inference space, as I'm sure you're aware, is not at scale. There's very little, like, active inference work that is active inference at scale. Most of the models are, like, relatively small toy grid world y type models. And part of the reason for that is that, you know, it is in fact difficult to scale Bayesian methods. Now that also has now begun to change, right? like, relatively small toy grid world y type models. And part of the reason for that is that, you know, it is in fact difficult to scale Bayesian methods. Now that also has now begun to change, right? We now have a lot of great mathematical tools and a lot of great frameworks for approximating Bayesian inference. You'll never do it exactly. We're approximating Bayesian inference, which I believe is how the brain works, right? Bayesian brain and all that. That allows us to build these kind of structured models that are structured both after the brain, how the brain is structured and how the the world that we live in is actually structured. Hence the the the this notion that what we need to build, get the net to the next layer of of AGI, and I also don't like that term and don't intend to use it very often. But what we need to get to next level right is is is this is this framework that allows us to build the kinds of models that we know people actually use and just make them bigger and more sophisticated and and and and so on. And then take advantage like hyper scaling Bayesian inference is part of it, but also like it's, you know, constructing models of the world as it actually works. The way the world actually works, right, is what is it is it it, you know, provides us with the structure of our own thinking. Right? The atomic elements of thought is how I like to phrase it. Our models of the physical world in which we live. And the physical world in which we live is a world of macroscopic objects that, you know, that have specific relations and interact in certain ways that we understand. Right? You know, looking around the room for a good example. Right? You sit on a chair. Right? That's an example of a relationship. It holds you up and all that fun stuff. And those are the kinds of you know, that understanding of the physical world was necessary for us in order you know, for us to have in order to survive. Dogs have it too. Right? It's language isn't what make you know, it it isn't isn't all that special. Right? Well, it's it's actually quite special. But, those are the models that form that that understanding of the world in which we live is where we get our the models that form the the the the the the models that form the atomic elements of our thoughts out of which we have composed more sophisticated models that have allowed us to do all this great systems engineering, build this great technology that we've got. So that's what we wanna do. Right? Is we wanna is is we're focused on building cognitively inspired models that are based on our understand on on the way the world in which we live actually works because we believe intelligence must be embodied. Building a framework for for putting those models together and experimenting with them at scale, all in an approximately Bayesian way because we believe that's how the brain works. It's not just about putting your AI into a robot. It's about giving that giving the robot a model of the world that is like our model of the world. A model that is object centered. It's dynamic. It's it's largely causal. Right? It's you know, that's that's that's the big difference.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.