Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
27 Dec 2025 The Cognitive Revolution Controlling Tools or Aligning Creatures? Emmett Shear (Softmax) & Séb Krier (GDM), from a16z Show
“I mean, I'm not really sure, right? Because I think if I give a model a certain goal, then I would like the model to kind of follow that instruction and kind of reach that particular goal.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 27 Dec 2025
- Publisher
- The Cognitive Revolution
Transcript context
…team's amazing. But like, this is a, it's maybe, it's the first time I've run a company where truly I can say with a whole heart, if someone beats us, thank God. Like, I hope somebody figures it out. Yeah. I mean, it's, yeah, I have a lot of, you know, similar intuitions about certain things. Like, I also dislike the, the idea that kind of, we just need to crack the few kind of values or something, just cement them in time forever now, and we would kind of solve morality or something. And I've always kind of been skeptical about, how the alignment problem has been conceptualized as something to kind of solve once and for all, and then you can just, do AI or do AGI. But I guess I understand it in a slightly different way. I guess maybe less based on kind of moral realism, but there's a kind of the technical alignment problem, which I kind of think of broadly as how to get an AI to do what you, how do you get it to follow instructions, like, broadly speaking. And I think that was, more of a challenge, I think, pre-LLMs, I guess, when people were talking about reinforcement learning and looking at these systems, whereas post-LLMs, we've realized that many things that we thought were going to be difficult were somewhat easier. And then there's the kind of second question, the kind of normative question of to whose values, what are you aligning this thing to, which I think is the kind of thing you're commenting on. And for this, I, yeah, I tend to be very skeptical of approaches where, you know, you need to kind of crack the kind of 10 commandments of alignment or something, and then we're good. And here I think I have like intuitions that are unsurprisingly a bit more like political science-based or something, and that, like, okay, it is a process. And I like the kind of bottom-up approach to some degree of, well, how do we do it in real life with people? No one comes up with, I've got this. And so you have processes that allow ideas to kind of, clash. You have got people with different ideas, opinions, views, and stuff to kind of coexist as well as they can within a wider system. And like, with humans, that system is liberal democracy or something. And, at least in some countries. And that allows more of that kind of, you know, these kind of ideas, these values to be kind of discovered and construed over time. And And I think, for alignment as well, I tend to think, there's on the normative side, I agree with some of your intuitions. I'm less clear about now what it exactly, what does it look like now if we're going to implement this into an AI system? These are the ones we have today. e side, I agree with some of your intuitions. I'm less clear about now what it exactly, what does it look like now if we're going to implement this into an AI system? These are the ones we have today. I agree that there's this idea of technical alignment that I think I would be able to define a little differently, but it's sort of the sense of like, if you build a system, can it be described as being coherently goal following at all, regardless of what those goals are? Like, Lots of systems aren't coherently, they're not well described as having goals. They just kind of do stuff. And if you're going to have something that's like aligned, it has to have coherent goals. Otherwise, those goals can't be aligned with anyone else's goals, kind of by definition. Is that sort of, is that, would you, is that a fair assessment of what you mean by technical alignment? I mean, I'm not really sure, right? Because I think if I give a model a certain goal, then I would like the model to kind of follow that instruction and kind of reach that particular goal. Rather than it having a goal of its own that I can't. Yeah. If you give it a goal, it has that goal. Right. I was going to give someone something, right? So yeah, if I instruct it to do X, then I would like it to do X and not, you know, different variants of X, essentially. I wouldn't want it to reward hack. I wouldn't need some. Well, but are you, but you, when you tell it to do X, you're transferring like a series of like a byte string in a chat window or like a a series of audio vibrations in the air, right? You're not, you're not transplanting a goal from your mind into it. You're giving it an observation that it's using to infer your goal. Yeah, I mean, in some sense, yeah, I can communicate a series of instructions and I want it to infer what I'm, you know, saying essentially as accurately as it can, given what it knows of me and what I'm asking. You wanted to infer what you meant, right? Like, that's like, because in some sense there's no... the byte sequence that you sent over the wire to it has no absolute meaning. It has to be interpreted, right? Like that byte sequence could mean something very different with a different code book. Yeah, well, I guess one way, you know, I think I remember when I was first getting into AI and, you know, these kind of questions maybe like a decade ago or so, You have these examples of, I think it was Stuart Russell in the textbook, we'll give the AI a goal, but then it won't exactly do what you're asking it, right? You know, clean the room, and then it goes and cleans the room, but takes the baby and puts it in the trash. Like, no, this is not what I meant. then it won't exactly do what you're asking it, right? You know, clean the room, and then it goes and cleans the room, but takes the baby and puts it in the trash. Like, no, this is not what I meant. Like, but like, wait, hold on, but this is the thing where I think people, this is the, you have to, like, you were jumping over a step there. You didn't give the AI a goal, you gave the AI a description of a goal. A description of a thing and a thing are not the same. I can tell you an apple, And I'm evoking the idea of an apple, but I haven't given you an apple, I've given you a just, you know, it's red, it's shiny, it's a size. That's a description of an apple, but it's not an apple. And giving someone, hey, go do this, that's not a goal, that's a description of a goal. And for humans, we're so fast, we're so good at turning a description of a goal into a goal. We do it so quickly and naturally, we don't even see it happening. we think that we get confused and we think those are the same thing. But you haven't given it a goal. You've given it a description of a goal that you want it to, you hope it turns back into the goal that is the same as the goal that you described inside of you. Right. You could give it a goal directly by reading your brainwaves and synchronizing its state to your brainwaves directly. I think that would meaningfully, you could say, okay, I'm giving it a goal. I'm synchronizing it, its internal state to my internal state directly. And this internal state is the goal. And so now it's the same. But I don't, most people aren't, don't mean that when they say they gave it a goal. Sure. About that? It goes back to my, what I was saying, like, this is a, you, Technical alignment is the capacity of an AI that I put forward, right? I want to check if we're like on the same page about it, is the capacity of AI to be good at inference about goals and like be good at inferring from a description of a goal, what goal to actually take on and good at once it takes on that goal, acting in a way that is actually in concordance with that goal coming about. So it is both pieces. You have to be able to You have to have the theory of mind to infer what that description of a goal that you got, what goal that corresponded to. And then you have to have a theory of the world to understand what actions correspond to that goal occurring. And if either of those things breaks, it kind of doesn't matter what goal you were, if you can't consistently do both of those things, you're not, which I think of as being a coherent, inferring goals from observations and acting in accordance with those goals is what I think of as being a coherently goal-oriented being. Because that's what, whether I'm inferring those goals from someone else's instructions or from the sun or tea leaves, the process is get some observations, infer a goal, use that goal, infer some actions, take action. And if you, an AI that can't do that is not technically aligned or not technically align a bull, I would even say. It lacks the capacity to be aligned because it can't, it's not competent enough.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.