High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

27 Dec 2025 The Cognitive Revolution Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show

“Let's say we scale up one of these tools, because you can make a super powerful tool that doesn't have these meta stable, like the states I'm talking about are not necessary to have a very smart tool, which is sort of basically a tool is like a first second order model that just doesn't meaningfully have pleasure and pain, right?”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
27 Dec 2025
Publisher
The Cognitive Revolution

Transcript context

…e, that tells you that it has I guess you'd call it like feelings, almost like it has, it has ways, it has, it has meta states, a set of meta states that it alternates between, that it shifts between. And then if you climb all the way up of, up that, and you should have have, okay, well, then you have, you have a, you have trajectories between, between these meta states, and then a second, second order of those, that's like thought. That's like, now it's like a person. And so if I found all six of those layers, which by the way, I definitely don't think you'd find it in LLM. Like, in fact, I know you can't find them because these things don't have attention spans like that at all. Then I would start to at least very seriously consider it as a, you know, a thinking being like somewhat like a human. There's a third order you could go up as well, but like that's basically what I'd be interested in is like the underlying dynamics of its learning processes and how its goal states shift over time. I think that's what basically tells you if it has internal pleasure pain states and sort of like self-reflective moral desires and things like that. And zooming out, this moral question is obviously very interesting, but if someone wasn't interested in the moral question as much, I think what you would say is if I understand correctly, is you also just feel on purely pragmatically your approach is going to be more effective in aligning AIs than some of these, you know, tops down control methods that we alluded to as well, right? ly, is you also just feel on purely pragmatically your approach is going to be more effective in aligning AIs than some of these, you know, tops down control methods that we alluded to as well, right? Yeah, I guess the problem is like you're making this model and it's getting really powerful, right? And let's say it is a tool. Let's say we scale up one of these tools, because you can make a super powerful tool that doesn't have these meta stable, like the states I'm talking about are not necessary to have a very smart tool, which is sort of basically a tool is like a first second order model that just doesn't meaningfully have pleasure and pain, right? Like great, but doesn't even have a subjective experience. I know I kind of think it maybe does, but not in a way that I give a **** about. And so What happens then? Well, it's you've trained it to infer goals from your from observation and like to prioritize goals and act on them. And one of one of two things is going to happen is like you're the this very, very powerful optimizing tool that's got like has lots of causal influence over the world is going to be technically aligned and is going to do what you tell it to do, or it's not. And it's going to go do something else. I think we can all agree if it just goes and does something random, that's obviously very dangerous. But I put forward that it's also very dangerous if it then goes and does what you tell it to do. Because you ever seen the Sorcerer's Apprentice? Humans' wishes are not stable. Like, not at a level of like, of immense power. Like, you want, ideally, people's wisdom and their and their power kind of go up together. And generally they do, because being smart for people makes you generally a little more wise and a little more powerful. And when these things get out of balance, you have someone who has a lot more power than wisdom. That's very dangerous. It's damaging. But at least right now, the balance of power and wisdom is kept at like, the way you get lots of power is by basically having a lot of other people listen to you. And so like, at some point, if you're The mad king is a problem, but generally speaking, eventually the mad king gets assassinated or people stop listening to him because like he's a mad king. And so the problem is you think, okay, great, we can steer the super powerful AI. And now the super powerful AI is in the, this incredibly powerful tool is in the hands of a human who is well-meaning but has limited finite wisdom like I do and like everyone else does. And their wishes are bad and not trustworthy. And the more of that you have, and you're giving those out everywhere, and this ends in tears also. And so basically, you just don't give everyone atomic bombs are really powerful tools too. I would not say you should go, and they're not aware, they're not beings. I would not be in favor of handing atomic bombs to everybody. There's a power of tool that just should not be built generally, because it is more power than any human's individual wisdom. is available to harness. And if it does get built, it should be built at a societal level and protected there. ould not be built generally, because it is more power than any human's individual wisdom. is available to harness. And if it does get built, it should be built at a societal level and protected there. And even then, I don't know that it's a, there are tools so powerful that even as a society, we shouldn't build them. That would be a mistake. The nice thing about a being is like a human, if you get a being that is good and is caring, there's this automatic limiter. It might do what you say, but if you ask it to do something really bad, it'll tell you no, that's like other people. And like, that's good. That is a sustainable form of alignment, at least in theory. It's way harder. It's way harder than the tool snaring. So I'm in favor of the tool snaring. We should keep doing that. And we should keep building these limited less than human intelligence tools, which are awesome. And I'm super into, and we should keep building those and keep building steerability. But as you're on this like trajectory to build something as smart as a person, right up into the right, and then smarter than a person, a tool that you can't control bad, a tool that you can control bad, a being that isn't aligned, bad. The only good outcome is a being that is, that cares, that actually cares about us. That's the only way that ends well. Or we can just not do it. I don't think that's realistic. That's like the pause AI people. I think that's totally unrealistic and silly, but like, you know, theoretically, you could not do it, I guess. What can you say about your strategy of how you're trying to achieve or even attempt to achieve this level, like in terms of research or roadmap or?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence