High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Max Bennett: belief

30 Dec 2025 Machine Learning Street Talk Your Brain is Running a Simulation Right Now [Max Bennett]

“Know, Yan Lakun talks a lot about this and are less selfish. And so yeah, I think that's a great opportunity, but there's definitely risk because the second you give an autonomous agent the opportunity to produce its own sub goals, you need to have really rich either constraints or a really well defined reward function or know, 1 ability that I think comes from mentalizing actually is, and this is an idea in alignment research, which is, if you can convince an AI agent to try and do what it thinks the human wants it to do, What you're actually doing is you're requiring it to engage in some form of mentalizing to infer the preferences of the requester and then try and do what is best for that that individual.”

— Max Bennett

Source trail

Everything needed to verify it.

Speaker
Max Bennett
Attribution
Verified speaker
Claim type
belief
Recorded
30 Dec 2025
Publisher
Machine Learning Street Talk

Transcript context

…Yeah. That's the naturalistic fallacy. Right. Yeah. So it might be the case that it is very likely if you that species will eventually enter a politicking arms race, and certain forms of deception will emerge and power seeking will emerge. That doesn't mean when we produce our own sort of intelligent entities in AI that we should imbue them with those features. 1 of the, I think, optimistic outcomes of this new AI world we're gonna enter in the next hundred years is that we actually now, as designers, can do our best to try and remove some of the evolutionary baggage that we don't like that's evolved in humans in these new entities. And so there's, of course, risk. But I think there's also a really great opportunity that we could have benevolent beings that do not seek to dominate. Know, Yan Lakun talks a lot about this and are less selfish. And so yeah, I think that's a great opportunity, but there's definitely risk because the second you give an autonomous agent the opportunity to produce its own sub goals, you need to have really rich either constraints or a really well defined reward function or know, 1 ability that I think comes from mentalizing actually is, and this is an idea in alignment research, which is, if you can convince an AI agent to try and do what it thinks the human wants it to do, What you're actually doing is you're requiring it to engage in some form of mentalizing to infer the preferences of the requester and then try and do what is best for that that individual. Because you can't just have them take requests at at face value because then there's all these opportunities for misinterpretations. Nick Bostrom's famous paperclip factory. Maximize production of paperclips, earth is turned into paperclips. We obviously don't want that. But with mentalizing, with the ability to model the internal simulation of another mind and be able to play out how would this person feel about possible futures. You could imagine optimistically an outcome where an AI agent could easily infer if I turn all of earth into paper clips, that's not what the person giving me this request would have in fact wanted. They would regret that outcome. So of course, doesn't fully de risk things, but it is 1 methodology and 1 learning from sort of evolutionary neuroscience I think we can garner, that mentalizing is a tool that can be used to try and stabilize sort of requests that we give each other in a more grounded way, there's not these types of misinterpretations. Of course, caveat, humans misinterpret each other all the time. It's by no means perfect, but it is a tool. And I think that's a very natural phenomenon. I think any intelligence system is naturally incoherent. I think it's impossible to have a single monolithic intelligence, which is monomaniacally, you know, focused in in a particular direction. But but anyway, wanna just slightly rewind a little bit to what we were saying. So the the first animals, they they had quite simplistic, social games that they were playing. So they were interested in strength and submission, and it was a fairly fixed interface. And what was really interesting is that deer, for example, they lock horns, don't they? So it's predictive. They don't actually have to have a fight because that would be, you know, evolutionary, evolutionarily not a smart thing to do. So, so the social game they play, even though the game is fixed, it's predictive, which is fascinating. And then you are telling the story of, I think, monkeys and macaques, how they have this really interesting virtual social game where strength and social status diverged. So your social status actually became this virtual thing that was based on grooming and pruning and lots of completely unrelated things, and it was entirely possible for, a very weak macaque to have a significantly higher social status than than a big, strong 1. So that's really fascinating. But then we get into I mean, maybe a broader question is we are still very social creatures ourselves. We have, Facebook, for example. And could you just arguably you know, could you cynically argue that Facebook or, you know, all socializing is just a kind of arms race to improve our social status? So when we're kind of posting on Facebook, in a way, it's like the deer locking horns. It's us playing these status games without having to have a fight with each other.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence