High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ryan Kidd: belief

4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

“Now you could say like, okay, what if you also tried to pour resources into like secret AI safety projects at the same time, delay RLHF, delay ChatGPT, build up the AI safety field, uh uh via networks the myri summer schools weren't doing a lot and MATS came along uh just before the ChatGPT moment December 2021 and yeah I think like the first MATS cohorts were a little bit less like a little bit more directionless than the later cohorts definitely like I think safety research really kicked into gear after we had ChatGPT uh not to say that was the only cause but there were like a lot of things happening around that time and I think that like Definitely larger, more capable models have enabled certain types of essential safety research you could not do with smaller models.”

— Ryan Kidd

Source trail

Everything needed to verify it.

Speaker
Ryan Kidd
Attribution
Verified speaker
Claim type
belief
Recorded
4 Jan 2026
Publisher
The Cognitive Revolution

Transcript context

…So I'm not trying to defend research like this or even defend your capabilities enhancing safety research per se. I'm just saying that it's pretty hard to imagine a situation where you, because I think you do have to build the AGI at the end of the day. And I know I'm alienating a lot of people who might watch the show when I say that, but I think that you kind of have to from a pragmatic perspective because the market forces driving this are very strong. Now, there are some options that we could take, right? We could build direct source comprehensive AI services. So you never have to have like a centralized agent. You have distributed kind of mechanisms, right? You build scientist AI that is very narrow AI systems to serve a bunch of economic things. The problem is, I think they all get outcompeted by agentic AI that like, you stack like an AI company filled with agents and they all like go out in the stock market and make products and so on and just make more money. just beat your crappy narrow AI solutions. So the problem is like it's not just about making AI that is aligned. It's about making AI that is performance competitive enough that it dominates in the marketplace. The only alternative is to like have some sort of draconian like shut it all down kind of thing, which I am just very skeptical of ever working. I don't see any example of such a thing happening. The closest example we have is like stopping human cloning, but that that was not a lucrative bet, like in the same way that AGIs I claim. And also human cloning is kind of this, it violates this deep social more, I think in a way that few people today conceive of powerful AI systems to violate. I think they're wrong. I think building a second species is actually going to violate some deep social more in the same way that human cloning would be, but I don't think people will see it that way. So that leaves us with the fact that we actually have to build the AGI because But if we can build products that are safer, right, or perhaps are under some strict regulatory control that we have some like really like ideally like, I don't know, 10 year international slow phased entry to the new AGI world, right, where all these, you know, countries and companies are kind of forced to be very careful and collaborative in the way that they align their models, then we're in a much better world. That's the world I hope for. Okay, now as to whether AI safety research is unnecessarily capabilities enhancing. Some is, perhaps. RLHF, I think I'm on the fence, 50-50. Definitely at some point, RLHF was, the idea was in the water. It doesn't seem like it would have lasted much longer if Paul Cristiano and Dari Amadei et al. hadn't done that. I think someone else would have done it. That's not to say that you should necessarily try and accelerate the frontier of capabilities. onger if Paul Cristiano and Dari Amadei et al. hadn't done that. I think someone else would have done it. That's not to say that you should necessarily try and accelerate the frontier of capabilities. It seems bad on then, but certainly RLHF opened up the door to a lot of very promising ways to build alignment MVPs, which kind of is the Cristiano meta strategy too. I don't know. It's hard to say. I'd like to run the counterfactual simulation and see where the world would be without ROHF one or like one year sooner or two years sooner. That would be interesting to see. It definitely did kickstart, I think, the, you know, ChatGPT revolution and, you know, productizing AI systems. But it's hard to see, like, given how small the AI safety field was at the time, The AI safety field, I think, has grown from the increase of AI exposure, right? So you would've had some amount of additional AI safety research that happened had ChatGPT moment not happened then. It happened like one or two years later. But I think it would've been kind of insignificant, if I'm honest. I don't think that the field was big enough. Now you could say like, okay, what if you also tried to pour resources into like secret AI safety projects at the same time, delay RLHF, delay ChatGPT, build up the AI safety field, uh uh via networks the myri summer schools weren't doing a lot and MATS came along uh just before the ChatGPT moment December 2021 and yeah I think like the first MATS cohorts were a little bit less like a little bit more directionless than the later cohorts definitely like I think safety research really kicked into gear after we had ChatGPT uh not to say that was the only cause but there were like a lot of things happening around that time and I think that like Definitely larger, more capable models have enabled certain types of essential safety research you could not do with smaller models. We're talking like interpretability on models that actually have, you know, coherent concepts embedded in them. But we'll say there's probably plenty of work to still be done on GPT-2 small, but linear probes and whatnot at a high level can target some of our frontier models. You know, Quen, these Chinese models are particularly good for that. certain types of debate. Like we had the first interesting empirical debate paper only after models were good enough to debate. And there's many, many other such examples. Like all the control literature, I think, just could not have happened as well. Sorry if that's too much. No, it's great. Yeah, I was going to ask also about the idea that it sounds like you sort of believe it, at least up to a point. But, you know, going back to the sort of founding mythology of anthropic, I think one of the notions that was seen as like a legitimate reason, even among pretty hawkish AI safety folks for starting a company like Anthropic was, well, if you want to do the safety research, you've got to have frontier models to do it on. Otherwise you're just inherently behind. And then, you know, what good is that, right? What good is it to work on like last generation model? So obviously we've got a, you know, quite a few generations between GPT two and now. And sure, like, we don't understand plenty of things about GPT-2, but then I would also say there are a lot of emergent behaviors that are not observed in GPT-2 that are definitely of interest, including, you know, many of these deception and eval awareness things that are kind of most hair-raising to me today. Where do you come down on that now? Like, I wonder if somebody's like, geez, should I go to a frontier company because that's where the best models are, and that's where inherently that means the most consequential work would be done there. Or I could go work independently or at any of a number of other organizations, and I might be limited to a smaller QEN model or something. But maybe that suffices. Maybe there is enough in those kind of second tier models as we enter 2026 that you don't really need to be working with the latest, latest, latest. Again, I think I am mostly probably just confused or unsure about this, but do you have a take?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence