Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS
“I think that sounds like safe super intelligence in a nutshell. Maybe that's the setup that they've got.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 4 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…It's a very good question. And I'll preface by saying that all safety work is capabilities work. Fundamentally. Like people like to distinguish these things in terms of like, oh, oh, capabilities work is about the engine. It's about making the plane go faster. And safety work is about the directionality. But as you've pointed out, RLHF, which was intended as safety work to help the directionality, steer to where you want to go, also made people realize, oh, wait, this thing is useful. I can actually hop in this plane now because it's going to land where I want, which made them want to make the engine go faster so they could get there faster. That whole feedback loop started. I actually don't know if you can avoid this. The only way I could conceive of doing safety research, there's no impact on capabilities until the final critical moment when you deploy it. is like being holed up in a lab somewhere with people that you utterly trust under crazy NDAs and only having access to staggering resources, whatever's required, because presumably maths and theoretical methods aren't enough to improve safety. At least that seems to be the lesson of the last 10 to 20 years. I could be wrong, but it seems like the interplay between theory and empirical research is pretty vital for most types of disciplines like this. So you have to have staggering resources, perfectly loyal teams, like all these NDAs, no one's going to reveal your research, and then you build the system in secret or something somehow, and then, okay, then you deploy it, and then maybe you open source your alignment technology and everyone has it, or somehow you disable all the bad actors or something, it just seems like a very difficult prospect. I think that sounds like safe super intelligence in a nutshell. Maybe that's the setup that they've got. Extreme secrecy, unlimited resources. They did have one notable defection, but otherwise, you know, a team that has resisted lucrative buyout offers. So I'm not trying to defend research like this or even defend your capabilities enhancing safety research per se. I'm just saying that it's pretty hard to imagine a situation where you, because I think you do have to build the AGI at the end of the day. And I know I'm alienating a lot of people who might watch the show when I say that, but I think that you kind of have to from a pragmatic perspective because the market forces driving this are very strong. Now, there are some options that we could take, right? We could build direct source comprehensive AI services. So you never have to have like a centralized agent. You have distributed kind of mechanisms, right? You build scientist AI that is very narrow AI systems to serve a bunch of economic things. The problem is, I think they all get outcompeted by agentic AI that like, you stack like an AI company filled with agents and they all like go out in the stock market and make products and so on and just make more money. just beat your crappy narrow AI solutions. So the problem is like it's not just about making AI that is aligned. It's about making AI that is performance competitive enough that it dominates in the marketplace. The only alternative is to like have some sort of draconian like shut it all down kind of thing, which I am just very skeptical of ever working. I don't see any example of such a thing happening. The closest example we have is like stopping human cloning, but that that was not a lucrative bet, like in the same way that AGIs I claim. And also human cloning is kind of this, it violates this deep social more, I think in a way that few people today conceive of powerful AI systems to violate. I think they're wrong. I think building a second species is actually going to violate some deep social more in the same way that human cloning would be, but I don't think people will see it that way. So that leaves us with the fact that we actually have to build the AGI because But if we can build products that are safer, right, or perhaps are under some strict regulatory control that we have some like really like ideally like, I don't know, 10 year international slow phased entry to the new AGI world, right, where all these, you know, countries and companies are kind of forced to be very careful and collaborative in the way that they align their models, then we're in a much better world. That's the world I hope for. Okay, now as to whether AI safety research is unnecessarily capabilities enhancing. Some is, perhaps. RLHF, I think I'm on the fence, 50-50. Definitely at some point, RLHF was, the idea was in the water. It doesn't seem like it would have lasted much longer if Paul Cristiano and Dari Amadei et al. hadn't done that. I think someone else would have done it. That's not to say that you should necessarily try and accelerate the frontier of capabilities.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.