Evidence receipt / uncertainty
Published · transcript-backedRyan Kidd: uncertainty
4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS
“The time period required would require massive amounts of cognitive labor and human trials and stuff like that. And I don't know, it does sound very sci-fi, so I don't think we should rely on something like that, though I'm all for people pursuing moonshots on the side.”
Source trail
Everything needed to verify it.
- Speaker
- Ryan Kidd
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 4 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah, it's a good question. I actually think there is plenty of room for this, and here's why. The mainline kind of meta strategy that the AI safety community seems to be pursuing on the whole, we're talking in terms of funding, in terms of sheer like numbers of people and resources deployed, not necessarily in terms of like less wrong posts written or something, right? But in terms of like resources deployed is this AI control strategy, which is where like basically you build, well, perhaps it's better called alignment MVP. which is a term coined by Yan Leakey, former head of Super Alignment at OpenAI, now co-lead of Alignment Science and Anthropic. An Alignment MVP is an AI system that is a minimum viable product for accelerating the pace of alignment research differentially over capabilities research such that we get the right outcome. So basically, you're getting AIs to do your homework. And there's been a lot of debate on this. There's a very strong camp in the direction of like, this just never will work because as soon as an AI system is strong enough to be useful, it's dangerous, right? I think, you know, clawed code shows this is not the case, at least for software engineering, but perhaps for people who think that aligning AI systems requires like serious research taste, you know, they would probably say that this clawed code is nowhere near there, right? We're generally AI systems are nowhere near that level of research taste ability. Now, All of the things that you're mentioning that pay off only in 2063 scenarios, presumably, they only pay off over that time period, not necessarily because of human challenge trials or something. Maybe that makes a difference if you're interested in making humans more intelligent with genetic engineering or some of the crazy things that are being tossed around. But if you're mainly interested in, oh, this thing is going to take decades of technical work, maybe you can compress those decades into a really short period of AI labor. Right? If you can like get 'em to run faster, massively parallelize things and just, you know, in general, just get them to do your homework. Those 2063 AI alignment plans might be automatable over a shorter period of time. And so we, we should definitely be pursuing those because the more we do to like raise the waterline of understanding on these different scenarios, the easier it will be to hand off to AI assistance or to, to accelerate AI, AI input. I do think it's interesting you said BCI research, cuz I recall being at a conference once when someone was talking about, Okay, so the way we're going to solve alignment is we're going to solve human uploading, and we're going to put someone to the computer and get them to do 100 simulated researcher years or something. It's not very sci-fi, very pantheon. But then Eliezer Yakowski put his hand up and he said, I volunteer to be number two. Which makes sense, right? You don't want to be the first guy that might go wrong. s not very sci-fi, very pantheon. But then Eliezer Yakowski put his hand up and he said, I volunteer to be number two. Which makes sense, right? You don't want to be the first guy that might go wrong. But yes, people are seriously pursuing that. And I think it is interesting. I have talked to some BCI experts about a year ago and they said, there's no way that we get BCI in time for AGI. Sorry, it's not BCI, sorry. No way we get human uploading in time for AGI unless you actually have AGI, right? The time period required would require massive amounts of cognitive labor and human trials and stuff like that. And I don't know, it does sound very sci-fi, so I don't think we should rely on something like that, though I'm all for people pursuing moonshots on the side. That's part of what maths is about, right? We have this massive portfolio with a few moonshots in. Okay, so there's a lot of different directions I want to go from there, and I'm trying to just make sure I keep running tally. But, you know, maybe an interesting first one would be How do you think we're doing on the AI safety front overall, maybe relative to your expectations? I mean, you mentioned Les Wrong and Eliezer, and there's this sort of, I don't know all the lore of Matt's, but I do understand that a lot of people who have participated in it over time come out of the Eliezer discourse and had a certain set of assumptions that were like, we're not going to be able to teach this thing our values, it's going to be extremely unwieldy from the beginning. And now we have Claude and it's like, man, that's come a lot farther than I thought it would have at this point in time. And I'm kind of surprised in general by how little I see people's pee dooms moving. It seems like the people that had really high ones remain really high. Those that were never worried remain not so worried. I kind of feel like I'm taking crazy pills at times where I'm like, I don't know. I see these deceptive behaviors. They kind of freak me out. It's amazing that that was anticipated as well as it was by the safety theorists, even in the absence of any actual systems to work with. But then at the same time, it's not crazy to me to say that Claude seems in many ways probably above average in terms of how ethical it is compared to the average person. I don't know if that's contentious to say, but Claude, it's pretty remarkable in that respect. What do you make of where we are? Are you as confused as I am, or do you have a sort of a more sort of opinionated sense of how well we're doing overall.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.