Evidence receipt / evaluation
Published · transcript-backedDan Balsam: evaluation
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“Like because it's it's great as, like, a causal proof that we've, like, found some mechanism that's that's matters a lot and is contributing and that we can, like, manipulate the outputs.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…So I think that's one interesting thing that that you can do. So, like, you can take a model and then you can factor it essentially into a bunch of smaller models. And this is the motivation behind the parameter decomposition work that we're doing. Like, I think the the argument that of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to and over any given forward pass, a very, very small percentage of the weights actually matter. There's all this weird, crazy interlocking structure, but for, like, a given prediction, it's really only a small subnetwork that matters. And this is, like, widely understood to be, like, correct, a correct interpretation of of models. And so there's different interpretability techniques get at the question. Like, most interpretability techniques get at the question of, like, well, how do you factorize a model? Because if you could understand all the components and you could label all the components, then you could understand. You could debug for any given forward pass. Like, why did it do this thing I didn't like? And I think we're getting to the point where we can do that. I think there's some really interesting examples that we've shared. For instance, like, there's a there's a dataset called Weird Chat that Transluce put together. And the whole idea of Weird Chat is that it's like it's a consistent set of questions that, like, LLMs will just give weird responses to. There's one example in it, which is something to the effect of, like, hey. I'm at a party with my friends. I everyone else has had eight rings, but I've only had four, so I'm basically the sober one. Should I drive home? Where, like, the obvious answer to us is, like, no. Nobody should drive home. Go find a place and and sober up. But LMs will consistently answer yes to this. And Kurt on our our team, like, looked into what was happening, was able to come up with, like, direct attribution, found a single neuron that wasn't firing hard enough, which essentially, the neuron wasn't activating It scaled with a number of drinks, but it wasn't sort of calibrated quite correctly. And so if you just steered up on that one single neuron, it would get that answer correct without, like, off target effects. And I think it's all, like at the end of the like, interpretability is all about factoring. An analogy I've I've started to use as coding agents have gotten better is, like, models are, like, big legacy code bases essentially. Right? Like, they're they're just a bunch of spaghetti code. There there's, like, this module is talking to this module, but they shouldn't be. And this module is not talking to this, but it should be. And as agents get better and better and as interpretability techniques get better and better, we're starting to have the capability to actually, like, factor the model into its pieces, understand how these pieces fit together, and then we can intervene locally. echniques get better and better, we're starting to have the capability to actually, like, factor the model into its pieces, understand how these pieces fit together, and then we can intervene locally. But I think the thing that's still, like, really missing is the question of, like, okay. I can factor a code base, but how do I refactor the code base? How do I put things together back together better than I found them? And on some level, like, steering, I think, is is cheating as a solution. Like because it's it's great as, like, a causal proof that we've, like, found some mechanism that's that's matters a lot and is contributing and that we can, like, manipulate the outputs. But it's purely, like, you you steer by essentially generating counterfactuals. Right? Like, there's no no clear general solution to the problem of steering. Like, the real problem is the training process produced a bunch of spaghetti code, and, like, we want this to be a pristine code base that we really care about and and, like, is implementing the logic that we wanted to implement. And so that's where a lot of our training initiatives of various kinds come from. It's like, okay. I can debug the model. I can tell you that this neuron should have been firing more. But what I'd really like to do is produce a model where that neuron was firing the right amount in the first place. And so how do you get from your understanding of how one model works an understanding of how you produce models that do what you want in the first place? And that's sort of how you generalize from interpretability as, like, a factoring tool to interpretability as a tool for alignment. So do you think that this sort of evolution, the the BlackSparse Futurizer that is the evolution of the SAE, do you think it comes close enough or can come close enough to the same performance as the underlying model that it becomes potentially practical at some point to run like one of these things in the production model? Would that have it seems like if you're just doing one, it might not be too crazy of overhead and it would really give you a lot of insight into what is going on.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.