High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Ryan Kidd: prediction

4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

“Which, by the way, is predicated on this idea that verification is easier than generation, P versus NP, blah, blah, blah, especially if you can see the other person's thoughts and they can't see yours. It does make sense to be at the frontier from that perspective, but I will say that I think the main reason that the companies are doing this is obviously to make money.”

— Ryan Kidd

Source trail

Everything needed to verify it.

Speaker
Ryan Kidd
Attribution
Verified speaker
Claim type
prediction
Recorded
4 Jan 2026
Publisher
The Cognitive Revolution

Transcript context

…No, it's great. Yeah, I was going to ask also about the idea that it sounds like you sort of believe it, at least up to a point. But, you know, going back to the sort of founding mythology of anthropic, I think one of the notions that was seen as like a legitimate reason, even among pretty hawkish AI safety folks for starting a company like Anthropic was, well, if you want to do the safety research, you've got to have frontier models to do it on. Otherwise you're just inherently behind. And then, you know, what good is that, right? What good is it to work on like last generation model? So obviously we've got a, you know, quite a few generations between GPT two and now. And sure, like, we don't understand plenty of things about GPT-2, but then I would also say there are a lot of emergent behaviors that are not observed in GPT-2 that are definitely of interest, including, you know, many of these deception and eval awareness things that are kind of most hair-raising to me today. Where do you come down on that now? Like, I wonder if somebody's like, geez, should I go to a frontier company because that's where the best models are, and that's where inherently that means the most consequential work would be done there. Or I could go work independently or at any of a number of other organizations, and I might be limited to a smaller QEN model or something. But maybe that suffices. Maybe there is enough in those kind of second tier models as we enter 2026 that you don't really need to be working with the latest, latest, latest. Again, I think I am mostly probably just confused or unsure about this, but do you have a take? I mean, yeah, for plenty of interpretability research, people aren't using the frontier models. You don't have access to them. I mean, sure, people in the labs are, but at MATS, there's tons of really excellent papers that keep getting produced, and from many other sources, right? Eleuther AI, Far AI, et cetera, that are like doing world-class interpretability research on sub-frontier models. Because today's like sub-frontier model, today's QUAN or DeepSeek or Llama or whatever is, it's like yesterday's frontier model, you know, in terms of capabilities. We're at that point where these models are all above the waterline, you know, for doing really excellent research. So from an interpretability perspective, I don't think you need to be pushing the frontier that much, if at all. From the perspective of other types of research agendas, such as like weak to strong generalization and other types of AI control and skilled oversight things, I think you kind of do need more data points. I'm not saying we've exhausted everything you can do with the current models, far, far from it. But I think like you are gonna need more data points to build up, you know, consistent and to see some of these kind of worrying behaviors emerge where your weaker model can't actually supervise your strong model in all situations. Which, by the way, is predicated on this idea that verification is easier than generation, P versus NP, blah, blah, blah, especially if you can see the other person's thoughts and they can't see yours. It does make sense to be at the frontier from that perspective, but I will say that I think the main reason that the companies are doing this is obviously to make money. And then, you know, in corollary of that, like from a safety perspective, you were trying to actually make a strong case for being at the frontier, it would be like, so our models are performance competitive, and that like they're close enough to the frontier, like fast follower kind of model, that like you take a performance hit by using ours instead of the competitors, but they're safer. And currently no one wants to use anything that's worse than the frontier model. Why would you? That's the best model. But if a model was like, I don't know, 10% more likely to tell you to jump off a bridge or something, or actually, seriously, 10% more likely to hack your bank account and steal all your money, let alone, I don't know, escape and make a bioweapon, I would like to think people would use the less good model. And I like to think that regulators and insurers could adequately penalize the frontier companies into complying with those. Because then you have existence proofs, like, oh, my product, I'm actually trying. I made an effort. I tried to make my product not do the heinous thing that the very best model developer is doing. Then everyone has no excuse, and they have to do that, and governments can compel them to, and so on. . I tried to make my product not do the heinous thing that the very best model developer is doing. Then everyone has no excuse, and they have to do that, and governments can compel them to, and so on. So I think making your model performance competitive enough that people want to pay the alignment tax, so to speak, seems like a viable strategy from that perspective. Now, of course, none of this is trying to justify the current frontier, the race of frontier models, which seems very reckless. Let's be clear. I think at the current pace of development that we're going to be in a lot of trouble. But this is one of those collective action problems. These companies have to coordinate to slow down. And there's international things at stake here as well, because now you have, you do have a US model versus China model developer kind of race now that they're in the running. So it's very complicated. And when you have these collective action problems, I think the main way you solve them is through governance. And sure, the lab leads could be probably even more collaborative. And definitely some of them are not advocating as strongly as they should be for slowing down and for having this kind of collective, you know, kind of sharing in the alignment benefits and and, you know, not pushing the frontier dangerously. But I do think this is ultimately a job for governments.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence