High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

27 Dec 2025 The Cognitive Revolution Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show

“And so, and most people don't know their goals, I think. And so, I think when you have agents and giving them goals or whatever, I think that should be part of the equation that we actually, we don't know all the goals.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
27 Dec 2025
Publisher
The Cognitive Revolution

Transcript context

…thing, you know, you would expect them to have understood that goal of the robot to like essentially not put the baby in the trash can or something, and just actually do the right sequence of action. Well, in that case, it failed the, that robot very clearly failed goal inference. You gave it a description of a goal, and it inferred the wrong states to be, the wrong goal states. That's just incompetence. It doesn't, it is incompetent and inferring goal states from observations. Children are like this too, like, you know, and honestly, if you ever played the game where you give someone instructions to make a peanut butter sandwich, and then they follow those instructions exactly as you've written them without filling in any gaps, it's hilarious. Because You can't do it. It's impossible. Like you think you've done it and you haven't. And like they put the they wind up putting the like the knife in the toaster and like the peanut butter. They don't open the peanut butter jar. So just jamming the knife into the top lid of the peanut butter jar. And like, it's endless. And like, because actually, if you don't already know what they mean, it's really hard to know what they mean. Like we. The reason humans are so good at this is we have a really excellent theory of mind. I already know what you're likely to ask me to do. I already have a good model of what your goals probably are. So when you ask me to do it, I have an easy inference problem. Which of the seven things that he wants is he indicating? But if I'm a newborn AI that doesn't have a great model of people's internal states, then like, I don't know what you mean. It's just incompetent. It's not like, which is separate from I have some other goal. And I knew what you meant, but I decided not to do it because there's some other goal that's competing with it, which is another thing you can be bad at, which is, again, different than I had the right goal, I inferred the right goal, I inferred the right priority on goals, and then I'm just bad at doing the thing. I'm trying, but I'm incompetent at doing. And these roughly correspond to the OODA loop, right? Like bad at observing and orienting, bad at deciding, bad at acting. And if you're bad at any of those things, you won't be good. And then I think there's this other problem that you, I like the separation of between technical alignment and value alignment, which is like, are you good if we told you the right goals to go after somehow, if you learned the right goals to go after via observation, and you were trying, like, What goals should you have? What goals should we tell you to have? What goals should we tell ourselves to have? What are the good goals to have? Is a separate question from, given that you've got some goals indicated, are you any good at doing it? Which I feel like is actually, in many ways, the current heart of the problem. We're much, much worse at technical alignment than we are at guessing what to tell things to do. Do you think that, does that align with your, how you mean technical and value alignment or technical alignment? h, much worse at technical alignment than we are at guessing what to tell things to do. Do you think that, does that align with your, how you mean technical and value alignment or technical alignment? Yeah, in some sense. I mean, it's the only thing that there's a... There's something about, an error, a mistake is one thing, and then there's the not listening to the instruction or something. But then, I think in the normative side, I mean, I just think of it even in really like ignoring AI, like I don't know what my goals are. And like, I've got some broad conception of certain things. I want to get a, you know, have dinner later or something like, and I want to kind of do well in my career. But I think a lot of these goals aren't something we kind of all just know. We kind of discover them as we go along. It's kind of constructive thing. And so, and most people don't know their goals, I think. And so, I think when you have agents and giving them goals or whatever, I think that should be part of the equation that we actually, we don't know all the goals. And this is something that is kind of, like you say, a process over time that is, dynamic. goals or whatever, I think that should be part of the equation that we actually, we don't know all the goals. And this is something that is kind of, like you say, a process over time that is, dynamic. So I think from my point of view, there's goals are one level of alignment. You can align something around goals, the kind of goals we're talking about here. are one level of alignment. You can align something around goals by like, if you can explicitly articulate in concept and in description, the states of the world that you wish to attain, you can orient around goals. But that only, that's a tiny percentage of human experience can be done that way. Many of the most important things cannot be oriented around that way. And the foundation, I think, of morality, the foundation, I think, of Where do goals come from? Where do values come from? Human beings exhibit a behavior. We go around talking about goals and we go around talking about values. And that's a behavior caused by some internal learning process that is based on observing the world. What's going on there? I think what's happening is that there's something deeper than a goal and deeper than a value, which is care. We give a ****. We care about things. And care is not conceptual. Care is non-verbal. It doesn't indicate what to do. It doesn't indicate how to do it. Care is a relative weighting over effectively like attention on states. It's a relative weighting over like which states in the world are important to you. And I care a lot about my son. What does that mean? Well, it means his states, the states he could be in are like, I pay a lot of attention to those and those matter to me. And you can care about things in a negative way. You can care about your enemies and what they're doing. And you can desire for them to do bad. But I think that like, and so you don't just want it to care about us. You want it to care about us and like us too, right? Maybe. But like, but the foundation is care. Until you care, you don't know why should I pay more attention to this person than this rock? Well, because I care more. And that What is that care stuff? And I think that what it appears to be, if I had to guess, is that the care stuff, this sounds so stupid, but care is basically like a reward. Like how much does this state correlate with survival? How much does this state correlate with your inclusive your full inclusive reproductive fitness for someone thing it learns evolutionarily or for a reinforcement learning agent like a LLM, how much does this correlate with reward? Does this state correlate with my predictive loss and my RL loss? Good, that's a state I care about. I think that's kind of what it is. The other part of Seth's question was just how does this, what does this look like in AI systems? And maybe another way of asking is like, When you talk to the people most focused on alignment at the major labs, as obviously you have over the years, how does your interpretation differ from their interpretation and how does that inform what you guys might go do differently?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence