Evidence receipt / belief
Published · transcript-backedMyra Deng: belief
6 Feb 2026 Latent Space The First Mechanistic Interpretability Frontier Lab — Myra Deng & Mark Bissell of Goodfire AI
“I think for us, it’s like, we have a very grounded view of alignment and, and safety in that we want to make sure that we can build models that do what we want them to do and that we have scalable oversight into what these models are doing.”
Source trail
Everything needed to verify it.
- Speaker
- Myra Deng
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 6 Feb 2026
- Publisher
- Latent Space
Transcript context
…I’ll ask a brief like safety question. You know, McInturk was kind of born out of the alignment and safety conversation. Safety is on your website. It’s not like something that you, you like de-prioritize, but like there’s like a sort of very militant safety arm that like wants to blow up data centers and like stop AI and, and then there’s this like sort of middle ground and like, is, is this like a conversation in your part of the world? Do you go up to Berkeley and Lighthaven and like talk to those guys or are they like, you know, there’s like a brief like civil war going on or no? I think, I think a good amount of us have spent some time in Berkeley. And then there are researchers there that we really. Admire and respect. I think for us, it’s like, we have a very grounded view of alignment and, and safety in that we want to make sure that we can build models that do what we want them to do and that we have scalable oversight into what these models are doing. And we think that that is the key to a lot of these like technical alignment challenges. And I think that is our opinion. That’s our research direction. We of course are going to do. Safety related research to make sure that our techniques also work on, you know, things like reward hacking and, and other like more concrete safety issues that we’ve seen in the wild, but we want to be kind of like grounded in solving the technical challenges we see to having humans be humans play a big role in, in the deployment of, of these super intelligent agents of the future. Yeah, I’ve, I’ve found the community to actually be remarkably cohesive, whether it’s. Talking about academia or the interpretability work being done at the frontier labs or some of the independent programs like maths and stuff. I think we’re all shooting for the same goal. I don’t know that there’s anyone who doesn’t want our understanding of models to increase. I, I think everyone, regardless of where they’re coming from or the use cases that they’re thinking, whether it’s alignment as the premier thing they’re focused on or someone who’s coming in purely from the angle of scientific discovery, I think we would all hope that models can be. More reliably and robustly controlled and understood. It seems like a pretty unambiguous goal.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.