Evidence receipt / belief
Published · transcript-backedDan Balsam: belief
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“I think the way to think about a model is more like a sparse mixture of subspaces.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Well, another thread that has been I think, correct me if I'm wrong, it seems like the the biggest thread in terms of the number of papers and blog posts that have come out recently from Goodfire is around trying to figure out the more detailed geometries that models use to represent concepts. I think we've covered in the past the linear representation hypothesis, which I would summarize super plain spokenly as models basically represent a concept as a direction in their activation space, and the intensity or the sort of salience of that concept is represented by the magnitude of the vector that points in that space. Now you're complicating that quite a bit, and we're going well beyond these sort of individual directions in space and finding all kinds of different geometries, which I've some of which are, like, pretty intuitive, like the days of the week are a circle, but some of which get pretty exotic, like some of the protein model manifolds that I've had the chance to to look at in in preparing for this. So maybe just for, like, super quick starters, what's kind of the headline? If I gave you the six to nine months ago of what we should understand to be going on with the linear representation hypothesis, What is the, like, new short version that everybody can kinda take home and recite to themselves to make sure they have a a good working understanding? Yeah. In many ways, I think it's just a generalization of, like, the way we were discussing things before. And different people define the linear representation hypothesis, like, slightly differently. I think the the most, like, sort of defensible version of it is just saying that, like, features are, like, linearly decodable, which I think is, like, true, essentially. Like, it doesn't require nonlinear computation, generally speaking, in a model in order to, like for the model to read out a feature from the residual stream. But I think where the geometry components come in is that the features aren't, like, sort of, like, naive, maybe, like, SAE build take on things is, like, you could imagine that the model is encoding, like, a bunch of totally orthogonal concepts to each other from each other. And so it's, like, it's really just, like, a bunch of one hot encoded categorical features, and then, like, the magnitude of the feature corresponds to how much the model's thinking about it. But in actuality, the structures that we find are, like, significantly more complicated. I would think of it more. I think the way to think about a model is more like a sparse mixture of subspaces. So you'll have subspaces for different types of concepts. Right? Like, maybe you have your days of the week subspace, which which itself lives in maybe a more, like, conceptual calendar time subspace. And so you have, like, at different levels of resolution, these, like, different structures. And the geometry of those structures is really important because the geometry of those structures encodes what operations you can perform on them. It's sort of like the semantics of not just the individual concept, but like the concept space are like encoded by the relationship of those of those concepts with each other in some sort of geometry. And to make things extra confusing, the the relationship, the operations, and the mappings that are performed from one concept to another act as operations over those manifolds. So you map a manifold to a different manifold over some some computation. And so, like, the sort of naivest version would be like, well, you have the days of the week, Monday, Tuesday, Wednesday, Thursday. And there's just somewhere in the model that knows that Monday goes to Tuesday and Tuesday goes to Wednesday, etcetera, all the way around. But that's actually would be a super inefficient way if you think about it to, like, represent the days of the week. That would be a very, like, if statement spaghetti code way of representing it. Like, the much more efficient way is to represent it as a wheel. And in this world, like, the magnitude along some direction often does correspond to the model's, like, certainty in some way. So if a future if Monday is activating very high, then the model's very confident that it should be thinking about Monday. But the relationship between all those days is itself, like, a very expressive and and rich thing. day is activating very high, then the model's very confident that it should be thinking about Monday. But the relationship between all those days is itself, like, a very expressive and and rich thing. And I think what the thing that we just believe is that you're not really gonna understand the model if you don't understand the relationship between the features. It's like maybe the difference between understanding the periodic table and understanding chemistry. You can have the all the individual elements, and that gives you some information, but, like, really the way in which they combine, the structures in which they form, that's, like, what can start to get help you gain a sense of the complexity of the world. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.