Evidence receipt / uncertainty
Published · transcript-backedNathan Labenz: uncertainty
4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS
“I mean, you mentioned Les Wrong and Eliezer, and there's this sort of, I don't know all the lore of Matt's, but I do understand that a lot of people who have participated in it over time come out of the Eliezer discourse and had a certain set of assumptions that were like, we're not going to be able to teach this thing our values, it's going to be extremely unwieldy from the beginning.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 4 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…s not very sci-fi, very pantheon. But then Eliezer Yakowski put his hand up and he said, I volunteer to be number two. Which makes sense, right? You don't want to be the first guy that might go wrong. But yes, people are seriously pursuing that. And I think it is interesting. I have talked to some BCI experts about a year ago and they said, there's no way that we get BCI in time for AGI. Sorry, it's not BCI, sorry. No way we get human uploading in time for AGI unless you actually have AGI, right? The time period required would require massive amounts of cognitive labor and human trials and stuff like that. And I don't know, it does sound very sci-fi, so I don't think we should rely on something like that, though I'm all for people pursuing moonshots on the side. That's part of what maths is about, right? We have this massive portfolio with a few moonshots in. Okay, so there's a lot of different directions I want to go from there, and I'm trying to just make sure I keep running tally. But, you know, maybe an interesting first one would be How do you think we're doing on the AI safety front overall, maybe relative to your expectations? I mean, you mentioned Les Wrong and Eliezer, and there's this sort of, I don't know all the lore of Matt's, but I do understand that a lot of people who have participated in it over time come out of the Eliezer discourse and had a certain set of assumptions that were like, we're not going to be able to teach this thing our values, it's going to be extremely unwieldy from the beginning. And now we have Claude and it's like, man, that's come a lot farther than I thought it would have at this point in time. And I'm kind of surprised in general by how little I see people's pee dooms moving. It seems like the people that had really high ones remain really high. Those that were never worried remain not so worried. I kind of feel like I'm taking crazy pills at times where I'm like, I don't know. I see these deceptive behaviors. They kind of freak me out. It's amazing that that was anticipated as well as it was by the safety theorists, even in the absence of any actual systems to work with. But then at the same time, it's not crazy to me to say that Claude seems in many ways probably above average in terms of how ethical it is compared to the average person. I don't know if that's contentious to say, but Claude, it's pretty remarkable in that respect. What do you make of where we are? Are you as confused as I am, or do you have a sort of a more sort of opinionated sense of how well we're doing overall. Honestly, Nathan, I'm pretty confused. Like, I do think that, you know, contrary to expectations, we are looking like we're... Language models understand our values, right? That's the first thing to update. Like, they understand them in some key sense. It's not just regurgitating like stochastic parrots. Language models are... really good at like understanding human ethical mores and extrapolating on them in, in, in, in, you know, in some scenarios. They're also really good at sycophancy. They're getting even better at deception, sophisticated deception. They do tend to deceive users in the right circumstances. Uh, though it does seem like there's some debate about this and it seems like some of, some of the deception, it's far from what we might call consequentialist, like, like ******** consequentialist deception in most situations, but I think like Alignment faking and some other papers have shown that there are such, you can like create situations where AI will deceive the user to achieve some like ulterior objective, which was something that was like deliberately given to the AI as an objective. So that's one of the constraints of these little. scenarios. I don't think there are any examples of AIs coherently deceiving users, like pursuing this like coherent objective, right? Not just what we might call good heart deception, where they just kind of, they like fall into deception of tendencies because of the limitations of training data. And I'm talking about this coherent deception. There are a few cases of this, if any, where this is like the sustained, coherent deception that appears like to arise spontaneously through the training process, which is pretty good given the level of capabilities we have. Like, it seems like people didn't think five or 10 years ago for sure that we'd have AIs that are as capable as assisting frontier science that are safe to deploy. And people are like, we're never going to put on the internet. Who would do that? It's crazy. And now they're on the internet. And notably, the world hasn't ended yet. That's not to say it will say that way. You know, certainly a thing you don't want to do with the superintelligence is let it out-of-the-box. But Yeah, it does seem like we're in a better scenario than many imagined. Now, there, of course, like we could be in the calm before the storm, right? It might well be that there's what they call a sharp left turn or just, you know, a radical change in the way AI is internally kind of process information and it might acquire these kind of coherent long run objectives. I could point to like Matt's mentor, Alex Turner's conception of shard theory as an example for how this might happen, right? So like instead of AI systems being this, like, you know, containing like a single miso optimizer that is kind of coherently forming under training, right?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.