High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Adam Marblestone: evaluation

30 Dec 2025 Dwarkesh Podcast Adam Marblestone — AI is missing something fundamental about the brain

“If I channel Steve Byrnes more, I think he’s very concerned that the minimum viable things in the Steering Subsystem that you need to get something smart is way less than the minimum viable set of things you need for it to have human-like social instincts and ethics and stuff like that.”

— Adam Marblestone

Source trail

Everything needed to verify it.

Speaker
Adam Marblestone
Attribution
Verified speaker
Claim type
evaluation
Recorded
30 Dec 2025
Publisher
Dwarkesh Podcast

Transcript context

…Going back to this whole perspective of how our intelligence is not just this omnidirectional inference thing that builds a world model, but really this system that teaches us what to pay attention to what the important salient factors are to learn from, et cetera. I want to see if there’s some intuition we can drive from this about what different kinds of intelligences might be like. So it seems like AGI or superhuman intelligence should still have this ability to learn a world model that’s quite general, but then it might be incentivized to pay attention to different things that are relevant for the modern post-singularity environment. How different should we expect different intelligences to be? I think one way to think about this question is, is it actually possible to make the paperclip maximizer or whatever? If you try to make the paperclip maximizer, does it end up just not being smart or something like that because the only reward function it had was to make paperclips? I’d say, can you do that? I don’t know. If I channel Steve Byrnes more, I think he’s very concerned that the minimum viable things in the Steering Subsystem that you need to get something smart is way less than the minimum viable set of things you need for it to have human-like social instincts and ethics and stuff like that. So a lot of what you want to know about the Steering Subsystem is actually the specifics of how you do alignment essentially, or what human behavior and social instincts is versus just what you need for capabilities. We talked about it in a slightly different way because we were sort of saying, “Well, in order for humans to learn socially, they need to make eye contact and learn from others.” But we already know from LLMs that depending on your starting point, you can learn language without that stuff. So I think that it probably is possible to make super powerful model-based RL optimizing systems and stuff like that that don’t have most of what we have in the human brain reward functions and as a consequence might want to maximize paperclips. And that’s a concern. But you’re pointing out that in order to make a competent paperclip maximizer, the kind of thing that can build spaceships and learn physics and whatever, it needs to have some drives which elicit learning, including say curiosity and exploration.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence