Evidence receipt / evaluation
Published · transcript-backedDan Balsam: evaluation
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“Like, I think the the argument that of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to and over any given forward pass, a very, very small percentage of the weights actually matter.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…I wonder, like so I've been obsessed with this Gram technique that I'm sure you're familiar with that AE Studio put out with Anthropic not too long ago. And regular listeners know I've brought it up a bunch of times. Right? But the idea is simply if we start with some labeled data and we control where the gradients go in terms of, like, only allowing certain experts to be updated for certain kinds of data early in the training process, then you get the great benefit of even for unlabeled data, those data points gradients also tend to flow toward those same experts. There's sort of this absorption effect. The great promise, of course, is like or the great hope and promise is that you can have powerful open source models with maybe just a couple experts removed, and you can have your cake and eat it too in terms of access and avoiding concentration of power and all the things that we're worried about without creating major risk of, like, stochastic disaster. This feels like kind of the flip side of that coin in a way too where, like, I sort of wonder you push this to the limit and you kind of have the same sort of thing. Like, with sparse autoencoders, there was always a pretty big loss. I I guess I should say, like, compromise on the loss. Right? Like, the reconstruction loss, you're you're losing something substantial when you you wouldn't wanna run the model on in a production environment, like, through the SAE because it just won't perform as well. But you kind of push this model and you sort of end up with something potentially that looks like a mixture of experts, but where, like, the knowledge is all very nicely compartmentalized and organized and you have something a lot more like an encyclopedia than a big mess. So this feels like something that you guys are gonna probably push on pretty hard. Like, is there is the vision to really create a model where, like, all the knowledge is localized and you know, like, exactly where all the knowledge is, but it's still rich enough that it performs as good as the original model did? And and if that is the vision, like, what's gonna be hard about that? Yeah. So I think that's one interesting thing that that you can do. So, like, you can take a model and then you can factor it essentially into a bunch of smaller models. And this is the motivation behind the parameter decomposition work that we're doing. Like, I think the the argument that of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to and over any given forward pass, a very, very small percentage of the weights actually matter. There's all this weird, crazy interlocking structure, but for, like, a given prediction, it's really only a small subnetwork that matters. And this is, like, widely understood to be, like, correct, a correct interpretation of of models. And so there's different interpretability techniques get at the question. Like, most interpretability techniques get at the question of, like, well, how do you factorize a model? Because if you could understand all the components and you could label all the components, then you could understand. You could debug for any given forward pass. Like, why did it do this thing I didn't like? And I think we're getting to the point where we can do that. I think there's some really interesting examples that we've shared. For instance, like, there's a there's a dataset called Weird Chat that Transluce put together. And the whole idea of Weird Chat is that it's like it's a consistent set of questions that, like, LLMs will just give weird responses to. There's one example in it, which is something to the effect of, like, hey. I'm at a party with my friends. I everyone else has had eight rings, but I've only had four, so I'm basically the sober one. Should I drive home? Where, like, the obvious answer to us is, like, no. Nobody should drive home. Go find a place and and sober up. But LMs will consistently answer yes to this. And Kurt on our our team, like, looked into what was happening, was able to come up with, like, direct attribution, found a single neuron that wasn't firing hard enough, which essentially, the neuron wasn't activating It scaled with a number of drinks, but it wasn't sort of calibrated quite correctly. And so if you just steered up on that one single neuron, it would get that answer correct without, like, off target effects. And I think it's all, like at the end of the like, interpretability is all about factoring. An analogy I've I've started to use as coding agents have gotten better is, like, models are, like, big legacy code bases essentially. Right? Like, they're they're just a bunch of spaghetti code. There there's, like, this module is talking to this module, but they shouldn't be. And this module is not talking to this, but it should be. And as agents get better and better and as interpretability techniques get better and better, we're starting to have the capability to actually, like, factor the model into its pieces, understand how these pieces fit together, and then we can intervene locally. echniques get better and better, we're starting to have the capability to actually, like, factor the model into its pieces, understand how these pieces fit together, and then we can intervene locally. But I think the thing that's still, like, really missing is the question of, like, okay. I can factor a code base, but how do I refactor the code base? How do I put things together back together better than I found them? And on some level, like, steering, I think, is is cheating as a solution. Like because it's it's great as, like, a causal proof that we've, like, found some mechanism that's that's matters a lot and is contributing and that we can, like, manipulate the outputs. But it's purely, like, you you steer by essentially generating counterfactuals. Right? Like, there's no no clear general solution to the problem of steering. Like, the real problem is the training process produced a bunch of spaghetti code, and, like, we want this to be a pristine code base that we really care about and and, like, is implementing the logic that we wanted to implement. And so that's where a lot of our training initiatives of various kinds come from. It's like, okay. I can debug the model. I can tell you that this neuron should have been firing more. But what I'd really like to do is produce a model where that neuron was firing the right amount in the first place. And so how do you get from your understanding of how one model works an understanding of how you produce models that do what you want in the first place? And that's sort of how you generalize from interpretability as, like, a factoring tool to interpretability as a tool for alignment.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.