High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ryan Kidd: belief

4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

“Uh, though it does seem like there's some debate about this and it seems like some of, some of the deception, it's far from what we might call consequentialist, like, like ******** consequentialist deception in most situations, but I think like Alignment faking and some other papers have shown that there are such, you can like create situations where AI will deceive the user to achieve some like ulterior objective, which was something that was like deliberately given to the AI as an objective.”

— Ryan Kidd

Source trail

Everything needed to verify it.

Speaker
Ryan Kidd
Attribution
Verified speaker
Claim type
belief
Recorded
4 Jan 2026
Publisher
The Cognitive Revolution

Transcript context

…Okay, so there's a lot of different directions I want to go from there, and I'm trying to just make sure I keep running tally. But, you know, maybe an interesting first one would be How do you think we're doing on the AI safety front overall, maybe relative to your expectations? I mean, you mentioned Les Wrong and Eliezer, and there's this sort of, I don't know all the lore of Matt's, but I do understand that a lot of people who have participated in it over time come out of the Eliezer discourse and had a certain set of assumptions that were like, we're not going to be able to teach this thing our values, it's going to be extremely unwieldy from the beginning. And now we have Claude and it's like, man, that's come a lot farther than I thought it would have at this point in time. And I'm kind of surprised in general by how little I see people's pee dooms moving. It seems like the people that had really high ones remain really high. Those that were never worried remain not so worried. I kind of feel like I'm taking crazy pills at times where I'm like, I don't know. I see these deceptive behaviors. They kind of freak me out. It's amazing that that was anticipated as well as it was by the safety theorists, even in the absence of any actual systems to work with. But then at the same time, it's not crazy to me to say that Claude seems in many ways probably above average in terms of how ethical it is compared to the average person. I don't know if that's contentious to say, but Claude, it's pretty remarkable in that respect. What do you make of where we are? Are you as confused as I am, or do you have a sort of a more sort of opinionated sense of how well we're doing overall. Honestly, Nathan, I'm pretty confused. Like, I do think that, you know, contrary to expectations, we are looking like we're... Language models understand our values, right? That's the first thing to update. Like, they understand them in some key sense. It's not just regurgitating like stochastic parrots. Language models are... really good at like understanding human ethical mores and extrapolating on them in, in, in, in, you know, in some scenarios. They're also really good at sycophancy. They're getting even better at deception, sophisticated deception. They do tend to deceive users in the right circumstances. Uh, though it does seem like there's some debate about this and it seems like some of, some of the deception, it's far from what we might call consequentialist, like, like ******** consequentialist deception in most situations, but I think like Alignment faking and some other papers have shown that there are such, you can like create situations where AI will deceive the user to achieve some like ulterior objective, which was something that was like deliberately given to the AI as an objective. So that's one of the constraints of these little. scenarios. I don't think there are any examples of AIs coherently deceiving users, like pursuing this like coherent objective, right? Not just what we might call good heart deception, where they just kind of, they like fall into deception of tendencies because of the limitations of training data. And I'm talking about this coherent deception. There are a few cases of this, if any, where this is like the sustained, coherent deception that appears like to arise spontaneously through the training process, which is pretty good given the level of capabilities we have. Like, it seems like people didn't think five or 10 years ago for sure that we'd have AIs that are as capable as assisting frontier science that are safe to deploy. And people are like, we're never going to put on the internet. Who would do that? It's crazy. And now they're on the internet. And notably, the world hasn't ended yet. That's not to say it will say that way. You know, certainly a thing you don't want to do with the superintelligence is let it out-of-the-box. But Yeah, it does seem like we're in a better scenario than many imagined. Now, there, of course, like we could be in the calm before the storm, right? It might well be that there's what they call a sharp left turn or just, you know, a radical change in the way AI is internally kind of process information and it might acquire these kind of coherent long run objectives. I could point to like Matt's mentor, Alex Turner's conception of shard theory as an example for how this might happen, right? So like instead of AI systems being this, like, you know, containing like a single miso optimizer that is kind of coherently forming under training, right? an example for how this might happen, right? So like instead of AI systems being this, like, you know, containing like a single miso optimizer that is kind of coherently forming under training, right? If you remember the old Evan Hubinger paradigm, your outer optimizer loop, which is training your AI system, causes it to develop an internal like optimizer architecture, which then can have its own goals that differ quite a lot from the training objective. And then like whenever you, and like presumably there's some arguments, such as there are arbitrarily many ways to have this optimizer form to produce the right outputs, because this thing is clever. And if its main goal is to produce paper clips or some other thing, then it's going to realize it's in a training process and it's going to give you the output you want, no matter what its goal is. We still could be in store for that kind of thing. But currently, it seems like we don't-- AI systems are really messy. They're kludgy, like human brains. They have some bunch of contextually activated heuristics. It sees there's a bracket there, and it's like, oh, maybe I'll put another bracket there. It's very simple, dumb circuitry. But then sometimes, it does stuff that is, like in-context learning, seems a lot like it's actually pattern matching to gradient descent. When models are learning from the inputs data stream and learning some a new complicated thing, it seems a lot like they're kind of optimizing over the input tokens, or rather optimizing to produce some output. So we might be in for a world where AI systems do spontaneously gain these MESA optimizers and these things... You know, these are a serious source of concern because they're, you know, very powerfully trying to optimize for some objective. And this is what, you know, this is the main concern I have, I guess, is that we have this kind of deception model, this inner alignment failure, perhaps, where AI systems acquire goals spontaneously, or maybe because they're being trained deliberately to be power seeking and make money on the internet. And then like, they decide to hide and we don't have interpretability tools good enough to detect them. So I guess I haven't really changed my fear about this scenario eventuating, but I have become more confident that we can elicit useful work from AI systems before we see obvious signs of this, I'll say. And I'm pretty confident that AI systems right now are not executing very powerful scheming against this because I think we would see some sort of warning shots. I don't think it's going to be as night and day. I think we'll see some kind of situations where AI systems are trying to scheme in really dumb ways before they try and scheme in very competent, difficult ways. Does that make sense?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence