Evidence receipt / prediction
Published · transcript-backedMyra Deng: prediction
6 Feb 2026 Latent Space The First Mechanistic Interpretability Frontier Lab — Myra Deng & Mark Bissell of Goodfire AI
“You really predicted a project we’re already working on right now, which is detecting hallucinations using interpretability techniques. And this is interesting because hallucinations is something that’s very hard to detect.”
Source trail
Everything needed to verify it.
- Speaker
- Myra Deng
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 6 Feb 2026
- Publisher
- Latent Space
Transcript context
…And I’ve got to ask the basic question, you know, can we get examples where it hallucinates, pass it through, see what feature activates for hallucinations? Can I just, you know, turn hallucination down? Oh, wow. You really predicted a project we’re already working on right now, which is detecting hallucinations using interpretability techniques. And this is interesting because hallucinations is something that’s very hard to detect. And it’s like a kind of a hairy problem and something that black box methods really struggle with. Whereas like Gen Z, you could always train a simple classifier to detect that hallucinations is harder. But we’ve seen that models internally have some... Awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior. And so, yeah, we have a project that’s trying to detect that accurately. And then also working on mitigating the hallucinatory behavior in the model itself as well. Yeah, I would say most people are still at the level of like, oh, I would just turn temperature to zero and that turns off hallucination. And I’m like, well, that’s a fundamental misunderstanding of how this works. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.