High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dan Balsam: belief

8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent

“I think we have a really incredible team at Goodfire. And, yeah, everyone here, like, really believes in the mission and is working really hard, and that's always extremely motivating.”

— Dan Balsam

Source trail

Everything needed to verify it.

Speaker
Dan Balsam
Attribution
Verified speaker
Claim type
belief
Recorded
8 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…I'm excited for this. You guys are prolific as always, and we've got a lot to cover. Research and new product, which is an internal research platform product. And it's the pace is really relentless. Let me just ask you that. For starters, how are you holding up in the eternal sprint that is the AI game these days? I think we have a really incredible team at Goodfire. And, yeah, everyone here, like, really believes in the mission and is working really hard, and that's always extremely motivating. We've been pushing really hard to get do our product launch with Silico, and it's nice to be able to take a deep breath now on the other side of that. But, yeah, even more cool things coming soon. Well, let's start with some research. So it's I'm just always amazed when I think back to the kind of toy models of superposition is only like three years, right, ago now, and we have come so far. A few things that jumped out on the Goodfire blog that I wanna just run through, and we'll have to do it at kind of a high level because there's too much to do the we used to do deep dives on paper by paper. We'll have to go a little bit more superficially today. But one that made some waves was called predictive data debugging. Mhmm. And for this one, I just kinda wanna give you my interpretation and then let you kind of elaborate on that or tell me where you think it'll be particularly useful or what you guys have seen since the paper came out. My synopsis of this one was that basically if you have a way of interpreting a model like NSAE or we'll get into futurizers, think a little bit later as well, then you can run a bunch of data through it. You're fine tuning your post training dataset. You can look at what concepts are coming up active a lot when we put this dataset through. And then the kind of insight is there's a strong correlation between concepts that are active and the concepts that are being modified by the training process. So I think that right there is like, file that away folks. There's something to remember. Not shocking, but like it's it's notable that the relationship is quite strong there. And then when you see these concepts that are active and you know that those are the ones that are gonna be modified, you could just look and see like, are there any concepts here that are kind of strange to us, surprising, we don't really intend to be monkeying around with with the dataset that we have at hand? And if so, then you can quickly zoom in on what are the data points that have caused these features to come up and then you might find that actually there's some stuff in our dataset we maybe ought to think twice about, maybe we ought to filter, maybe we ought to modify. And this gives you a route to hopefully minimizing unwanted surprises in the behavior that you get from your post training or fine tuning work. How'd I do and what more should I know?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence