Evidence receipt / commitment
Published · transcript-backedDan Balsam: commitment
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“I think the role that we specifically play as a company as Goodfire is we'd like to build models that have less risks. And the way that we do that is that we study models and we play with all of the different ways that we can build models until we and we study those models until we can understand some empirical science of of alignment.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…asking for international cooperation to control and slow down the progress of AI. My myself was a signatory of that. I really hope that we can do something like that. I think we have to sort of get in front of of some of these risks before they become more severe. I think the role that we specifically play as a company as Goodfire is we'd like to build models that have less risks. And the way that we do that is that we study models and we play with all of the different ways that we can build models until we and we study those models until we can understand some empirical science of of alignment. There are many folks working on the theoretical side. Like, we view our role as, like, working on the empirical side, and we're trying to build tools that do that. I think, again, all tools are dual risk, but I think we're gonna do our best to make sure that nobody's using our platform for for anything that could pose a risk. And, ultimately, like, we think giving people access to research technology is an overwhelming net positive. I don't think we, as a species, solve these really hard problems as long as we're getting everyone involved. And so I think we gotta get everyone involved, and I think we also got up the structures in that can slow the roller coaster a little bit. If we could do both those things at the same time, I think I think we'll be alright. On the topic of slowing the ride or, you know, perhaps, you know, making some agreements between frontier developers. In the past, we talked about the the most forbidden technique. Mhmm. And, you know, I guess I would briefly describe that as training with a signal, with a monitoring signal that runs the risk of driving the bad behavior that you're worried about underground so that you lose the monitor ability, but you still might in fact, you know, get the the bad behavior. For me, like, the canonical example of that is OpenAI's obfuscated reward hacking.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.