High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ryan Kidd: belief

4 Jan 2026 The Cognitive Revolution Building & Scaling the AI Safety Research Community, with Ryan Kidd of MATS

“I think even in some of the interp streams as well, it's very possible to enter an interpretability stream and bring it with it like some model of the kind of theory-based interpretability mechanism or strategy that you want to pursue and then see that executed on.”

— Ryan Kidd

Source trail

Everything needed to verify it.

Speaker
Ryan Kidd
Attribution
Verified speaker
Claim type
belief
Recorded
4 Jan 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah, cool. Are you taking any connectors? Is there, if I am a connector type or wanna become one, Is Matt a way to find my way there or not really? Many have. I would call Jesse Hoogland, one such person. I would call Paul Rikers, so Timaeus and Simplex. I would, I'd say Marius Halpan as well, to some extent, with his deception evals work. Yeah, like, and I don't, probably dozens of people. I'm just like sharing some of the more, the names that come more easily to mind, but just many, many people have come through Matt's, we're super open to individuals who have this kind of archetype. And note a connector, right? They have empirical skills, they have theoretical skills. So they could probably succeed in a bunch of different ways, right? But they're uniquely spec'd out to connect those two things. Now there are some mentors and projects that are much more suited to this kind of thing than others. People like Richard Ngo, historically Evan Hubinger. I think actually Evan Hubinger has been like probably the most dominant connector driving force at MATS over our time, but he's not a mentor in the next program, unfortunately. It doesn't have time. But yeah, there's many different opportunities that match for this kind of thing. I think even in some of the interp streams as well, it's very possible to enter an interpretability stream and bring it with it like some model of the kind of theory-based interpretability mechanism or strategy that you want to pursue and then see that executed on. That's happened several times. One of the things that I took note of in the blog post from 18 months or so ago was you had made a comment that funders basically don't want to, or they're much more inclined to support the growth of organizations that they sort of see as legible, that have like research directions that they sort of feel are somewhat established or that they can wrap their heads around. And they're much more reluctant to fund like totally new conceptual directions. And that seems like it exists in contrast with like the AE Studio survey, where they basically found that the field as a whole seems to think that like, we don't have all the ideas that we need and, you know, that like more kind of far out ideas should be tried, which of course led to their neglected approaches approach. What do you make of that? Is there stuff that we can do or is it, you know, is there is it a different organization's job to figure out how to fill that gap. Because I do feel like I want some more, and I love some of the AE Studio stuff, including self-other overlap. I always come back to that as an example of something that's just quite off the map of what most people are doing. When I think of AI control and what Buck and the Redwood Research team are doing, I find that stuff fascinating. And one of the things that kind of impresses me most is that they are willing to work on something that in some ways is so depressing. They're like, we're going to try to figure out how to work with AIs, even assuming they're out to get us. And I'm like, yikes. I don't know that I would be able to sustain the positive attitude enough to do that if I was working from that premise. I do feel like there's a relative dearth of things that are more inspiring. Here, I think maybe of AI Studio, but also Softmax. Obviously, people have a lot of different opinions on, are these things ever going to work or not? I wonder what your take is on just kind of the overall mix. It seems like a lot of things are kind of more toward patch the holes, keep the AI down, tempt it, you know, see if it'll take the temptation, and then patch it, you know, if it takes the temptation. And there's not nearly as much that is sort of a, a more kind of colorful, positive vision for the future. And I wish there was, but maybe that's just not happening because The ideas are just too hard to come by. Maybe it's not happening because the funders aren't bold enough. What's your take on how we can get, if we should be trying to get more of that stuff? And if you think we should, how might we go about it?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence