Evidence receipt / belief
Published · transcript-backedNathan Labenz: belief
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“I think we've covered in the past the linear representation hypothesis, which I would summarize super plain spokenly as models basically represent a concept as a direction in their activation space, and the intensity or the sort of salience of that concept is represented by the magnitude of the vector that points in that space.”
Source trail
Everything needed to verify it.
- Speaker
- Nathan Labenz
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…scale things like predictive data debugging two models like KIMI and GLM where we were able to replicate the same results at that scale. Now it's speculated that most of the compute that goes into training new models now is coming from RL and not from pretraining. So there is certainly a lot of scale in RL if you wanna sort of get to the frontier level, which is quite expensive and and difficult to do. But we have the raw infrastructure for it, and I think we've bridged the gap from I think the reputation of interpretability used to be that it was something you did on on toy models. And I think now we've built and are making accessible to the world the infrastructure to do this on at least close to frontier scale. Well, another thread that has been I think, correct me if I'm wrong, it seems like the the biggest thread in terms of the number of papers and blog posts that have come out recently from Goodfire is around trying to figure out the more detailed geometries that models use to represent concepts. I think we've covered in the past the linear representation hypothesis, which I would summarize super plain spokenly as models basically represent a concept as a direction in their activation space, and the intensity or the sort of salience of that concept is represented by the magnitude of the vector that points in that space. Now you're complicating that quite a bit, and we're going well beyond these sort of individual directions in space and finding all kinds of different geometries, which I've some of which are, like, pretty intuitive, like the days of the week are a circle, but some of which get pretty exotic, like some of the protein model manifolds that I've had the chance to to look at in in preparing for this. So maybe just for, like, super quick starters, what's kind of the headline? If I gave you the six to nine months ago of what we should understand to be going on with the linear representation hypothesis, What is the, like, new short version that everybody can kinda take home and recite to themselves to make sure they have a a good working understanding? Yeah. In many ways, I think it's just a generalization of, like, the way we were discussing things before. And different people define the linear representation hypothesis, like, slightly differently. I think the the most, like, sort of defensible version of it is just saying that, like, features are, like, linearly decodable, which I think is, like, true, essentially. Like, it doesn't require nonlinear computation, generally speaking, in a model in order to, like for the model to read out a feature from the residual stream. But I think where the geometry components come in is that the features aren't, like, sort of, like, naive, maybe, like, SAE build take on things is, like, you could imagine that the model is encoding, like, a bunch of totally orthogonal concepts to each other from each other. And so it's, like, it's really just, like, a bunch of one hot encoded categorical features, and then, like, the magnitude of the feature corresponds to how much the model's thinking about it. But in actuality, the structures that we find are, like, significantly more complicated. I would think of it more. I think the way to think about a model is more like a sparse mixture of subspaces. So you'll have subspaces for different types of concepts. Right? Like, maybe you have your days of the week subspace, which which itself lives in maybe a more, like, conceptual calendar time subspace. And so you have, like, at different levels of resolution, these, like, different structures. And the geometry of those structures is really important because the geometry of those structures encodes what operations you can perform on them. It's sort of like the semantics of not just the individual concept, but like the concept space are like encoded by the relationship of those of those concepts with each other in some sort of geometry. And to make things extra confusing, the the relationship, the operations, and the mappings that are performed from one concept to another act as operations over those manifolds. So you map a manifold to a different manifold over some some computation. And so, like, the sort of naivest version would be like, well, you have the days of the week, Monday, Tuesday, Wednesday, Thursday. And there's just somewhere in the model that knows that Monday goes to Tuesday and Tuesday goes to Wednesday, etcetera, all the way around. But that's actually would be a super inefficient way if you think about it to, like, represent the days of the week. That would be a very, like, if statement spaghetti code way of representing it. Like, the much more efficient way is to represent it as a wheel. And in this world, like, the magnitude along some direction often does correspond to the model's, like, certainty in some way. So if a future if Monday is activating very high, then the model's very confident that it should be thinking about Monday. But the relationship between all those days is itself, like, a very expressive and and rich thing.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.