Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“One of the challenges obviously with, like, trying to develop techniques that you wanna hopefully will be relevant at the frontier is there's not too many open weights models that you can hack on that have the intensity of RL that is going on at the Frontier Labs, which is leading to these colorful problematic behaviors that we're seeing. But at the same time, it also, like, really amazes me over and over again that astounding work, including the the Cameron Berg paper that I think about all the time about the anti correlation between deception and role playing features and claims of subjective experience.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…like, we did it with DPO, but, like, true RL, like, the difference is, like, some rollouts may contain information that you don't want the model to learn, you know, even subtly. You know? I don't know what you're are you referring to something? Yes. Yeah. It's topical. And so being able to say, like, actually, we wanna discard this rollout is is, like, pretty valuable in and of itself. But what we wanna be able to do and what we think is, like, roughly equivalent over time is actually be able to interview in the model and say, like, hey. Here's a circuit or here's a feature that, like, the data is trying to upweight that we obviously don't want it to try to upweight. And so we we intervene in the training process in that way. We'll come back to this probably toward the end. I have some kind of zoomed out big picture questions for you. One of the challenges obviously with, like, trying to develop techniques that you wanna hopefully will be relevant at the frontier is there's not too many open weights models that you can hack on that have the intensity of RL that is going on at the Frontier Labs, which is leading to these colorful problematic behaviors that we're seeing. But at the same time, it also, like, really amazes me over and over again that astounding work, including the the Cameron Berg paper that I think about all the time about the anti correlation between deception and role playing features and claims of subjective experience. Lawn Llama three three seventy b, and that's, like, two years old. So are you guys able to see features that you think are kind of the relevant features that are leading to these, like, relentless hacking behaviors? Well, I think that's that's an active area of of study for us. Yeah. And something that we hope to publish more on in the future. For what yeah. For what it's worth, like, I actually think the gap between open and closed models has shrunk, like, quite considerably. I use Kimi k three. Like, I use Opus and Fable and Soul and Kimi k three for different things. And we've done we like, we've built the interpretability infrastructure and the training infrastructure, which is now all in our product,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.