Evidence receipt / prediction
Published · transcript-backedMax Bennett: prediction
30 Dec 2025 Machine Learning Street Talk Your Brain is Running a Simulation Right Now [Max Bennett]
“Doesn't mean it won't work well, but we should be concerned about how well this will generalize to real tasks where we might care about this much more deeply. So for example, with a human, because part of this good evidence to suggest that part of my ability to reason about your mind is because I have a mind that works quite similarly, We are almost bound together by some common mechanistic synergy between the way in which our brains work because our brains are quite similar, which which enables a lot of data efficiency, which is I'm pretty good at predicting what people do.”
Source trail
Everything needed to verify it.
- Speaker
- Max Bennett
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 30 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. And what really hit home to me is that, in a way, it's not even knowledge. It's all simulations. You know, it's just simulations of of other agents, and we've we've always kind of spoken about knowledge in some weird platonistic abstract sense. And I quite like the idea that the the primitive form of communication between humans is just simulations even when we're speaking to each other. Yeah. Exactly. So how do we solve the Sally Anne problem? It probably happens so quickly, but we just simulate what would we think if we were in Sally's shoes. And then I realized, well, I would look in in this place. And this helps us reason about other people. And this begs a really almost profound question about, you know, how unique is theory of mind. This brings me to a question that I've been asked multiple times, which is does ChatGPT have theory of mind? If by theory of mind, and so the evidence I should stipulate for anyone curious is if you ask GPT-three these sort of theory of mind puzzles, it does terribly. So that's an easy 1 to sort of discard out of hand. But if you ask GBT for these theory of mind puzzles, it performs remarkably accurate, like as human level at these theory of mind puzzles. And there's been people that have explored, is it just in the training data? And there's good evidence that it's not just because they're just like regurgitating what was in the training data. So does this mean that Chad Gypsy has theory of mind? So I think there's a few ways to reason about this. 1 is, what do we mean by theory of mind? If by theory of mind, we just mean the ability to solve these sort of false belief puzzles, then I think you have to accept the fact that, yes, it can solve those tasks. The problem is the way in which it renders this model of other minds is not through having a similar mind itself. And so what this means is we should be concerned. Doesn't mean it won't work well, but we should be concerned about how well this will generalize to real tasks where we might care about this much more deeply. So for example, with a human, because part of this good evidence to suggest that part of my ability to reason about your mind is because I have a mind that works quite similarly, We are almost bound together by some common mechanistic synergy between the way in which our brains work because our brains are quite similar, which which enables a lot of data efficiency, which is I'm pretty good at predicting what people do. Not perfect, but pretty good at predicting what people do because we're all people and there's similarities between how we act. And so that makes us quite data efficient and decent at generalizing to new situations where we put people in new places that we've never seen before. I can kinda guess, well, if I were in that situation, this is what I would do. GBT 4 has learned to build a theory of mind simply by reading text of these puzzles. And so clearly, it has some mechanism to build a model of predicting what people will do in certain circumstances and differentiating knowledge and intent, etcetera. But the concern is twofold. 1, what will happen if we take those types of models and put them in very new situations that are not based on just these puzzles, But for example, we're asking them to optimize a paperclip factory. . 1, what will happen if we take those types of models and put them in very new situations that are not based on just these puzzles, But for example, we're asking them to optimize a paperclip factory. Now that's a situation where we should be concerned how well will it do at actually inferring what we mean by what we say. And the second is data efficiency, which is how much data did it have to see to build this model. If it was a ton of data, then it's gonna be problematic if we have these new situations where we wanna teach them to model people's behaviors in this new place. If it requires a ridiculous amount of data, then it's always gonna sort of be slow to learn these things and always be at risk of not generalizing well when we put them in these new situations. So my answer here is nuanced, which is I think if by weak theory of mind, we mean solving puzzle questions, I think it's very hard to say that ChatGPT does not have some model of human behavior. But I do think the human and primate mechanism for doing so has a data efficiency advantage and a mechanistic synergy advantage. In other words, we can use ourselves to reason about things that is relevant. And if we wanna have these systems do a good job listening to human requests, we shouldn't translate performance on false belief tests to believing that they'll do a good job correctly inferring our intent to knowledge in new situations.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.