Evidence receipt / evaluation
Published · transcript-backedNathan Labenz: evaluation
4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS
“Yeah, so that brings up another, I think, huge question for AI safety research in general, and probably the strongest, maybe not in, I don't know if you would say strongest in the sense of being most compelling to you, but certainly the most hawkish or fiercest criticism that AI safety research gets is that it always ends up being dual use and that it always ends up somehow accelerating the core capabilities track.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 4 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…nt and you've got to have monitors and control protocols right so that's why control research is so important for that especially for the early days to catch some of these these these slip-ups, right? So you can do all the model organisms work you want in the lab, and that's one layer of defense to see if we have these capabilities or the penchants for deception emerge. And separately, you need to have all the control evals studying them as they're deployed, especially if they're going to be learning online, perhaps updating their behaviors, and just be constantly checking for this stuff. Be ready. Have a fallback plan, a rapid response plan. What are you going to do if actually you see Serious warning signs. If you shut the models down, your stock price is going to plummet. What do you do? Do you revert to an older system? That's safer? Probably. So I think, yeah, we should definitely be tracking this stuff. And I wouldn't say that we are in the clear by a long shot. I would say that we are in a better world, by my estimation, than Austrom and Murie predicted. you know, 10 something years ago. But I don't know, they would say I'm very wrong about that. But I don't know, I think that it's useful that we can get some work out of these things that looks like it is actually quite likely to accelerate AI safety work. Yeah, so that brings up another, I think, huge question for AI safety research in general, and probably the strongest, maybe not in, I don't know if you would say strongest in the sense of being most compelling to you, but certainly the most hawkish or fiercest criticism that AI safety research gets is that it always ends up being dual use and that it always ends up somehow accelerating the core capabilities track. And some people would say, just stay away from the domain entirely and focus on social shame or whatever. I do believe we can do better than that. I think we probably have to do better than that. But I wonder how you think about that, right? I mean, the canonical like RLHF was sort of a safety technique that really turned out to be more of a utility driver than anything, I would say. I mean, I guess they're both, right? It is dual use. But certainly when it came to like accelerating the field, making the things useful, waking the world up, like having all kinds of people pile in, everything going exponential all at once, you can kind of trace back to, at least in part, this transition from, real raw next token predictors to actual instruction followers. And we've got probably like a lot of those things going on today. The one that I think stands out to me the most is when you've alluded to a couple of times, which is like getting the AI to do the alignment work. You know, that sounds awfully close and uncomfortably close to recursive self-improvement, which is something that I am quite fearful of, I do think, again, Claude seems pretty ethical. The GPTs aren't too bad either, but yikes. Are we really ready to have them do our alignment homework? So how do you think about teasing out as you kind of prioritize different kinds of research, like where you want to invest, what kind of mentors you want to bring on, what kind of talent you want to cultivate through the program? How do you I mean, that seems like a huge question that is a really hard one. How do you think about it? It's a very good question. And I'll preface by saying that all safety work is capabilities work. Fundamentally. Like people like to distinguish these things in terms of like, oh, oh, capabilities work is about the engine. It's about making the plane go faster. And safety work is about the directionality. But as you've pointed out, RLHF, which was intended as safety work to help the directionality, steer to where you want to go, also made people realize, oh, wait, this thing is useful. I can actually hop in this plane now because it's going to land where I want, which made them want to make the engine go faster so they could get there faster. That whole feedback loop started. I actually don't know if you can avoid this. The only way I could conceive of doing safety research, there's no impact on capabilities until the final critical moment when you deploy it. is like being holed up in a lab somewhere with people that you utterly trust under crazy NDAs and only having access to staggering resources, whatever's required, because presumably maths and theoretical methods aren't enough to improve safety. At least that seems to be the lesson of the last 10 to 20 years. I could be wrong, but it seems like the interplay between theory and empirical research is pretty vital for most types of disciplines like this. So you have to have staggering resources, perfectly loyal teams, like all these NDAs, no one's going to reveal your research, and then you build the system in secret or something somehow, and then, okay, then you deploy it, and then maybe you open source your alignment technology and everyone has it, or somehow you disable all the bad actors or something, it just seems like a very difficult prospect.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.