Evidence receipt / evaluation
Published · transcript-backedDan Balsam: evaluation
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“So there is certainly a lot of scale in RL if you wanna sort of get to the frontier level, which is quite expensive and and difficult to do. But we have the raw infrastructure for it, and I think we've bridged the gap from I think the reputation of interpretability used to be that it was something you did on on toy models.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…which lets us scale things like predictive data debugging two models like KIMI and GLM where we were able to replicate the same results at that scale. Now it's speculated that most of the compute that goes into training new models now is coming from RL and not from pretraining. So there is certainly a lot of scale in RL if you wanna sort of get to the frontier level, which is quite expensive and and difficult to do. But we have the raw infrastructure for it, and I think we've bridged the gap from I think the reputation of interpretability used to be that it was something you did on on toy models. And I think now we've built and are making accessible to the world the infrastructure to do this on at least close to frontier scale. Well, another thread that has been I think, correct me if I'm wrong, it seems like the the biggest thread in terms of the number of papers and blog posts that have come out recently from Goodfire is around trying to figure out the more detailed geometries that models use to represent concepts. I think we've covered in the past the linear representation hypothesis, which I would summarize super plain spokenly as models basically represent a concept as a direction in their activation space, and the intensity or the sort of salience of that concept is represented by the magnitude of the vector that points in that space. Now you're complicating that quite a bit, and we're going well beyond these sort of individual directions in space and finding all kinds of different geometries, which I've some of which are, like, pretty intuitive, like the days of the week are a circle, but some of which get pretty exotic, like some of the protein model manifolds that I've had the chance to to look at in in preparing for this. So maybe just for, like, super quick starters, what's kind of the headline? If I gave you the six to nine months ago of what we should understand to be going on with the linear representation hypothesis, What is the, like, new short version that everybody can kinda take home and recite to themselves to make sure they have a a good working understanding?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.