High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dan Balsam: belief

8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent

“Like, I think we have to grab the steering wheel. I think that's the only way it can work.”

— Dan Balsam

Source trail

Everything needed to verify it.

Speaker
Dan Balsam
Attribution
Verified speaker
Claim type
belief
Recorded
8 Aug 2026
Publisher
The Cognitive Revolution

Transcript context

…I used to do Anthropic did together, I think, was very inspired. It's not that different, like, spiritually from what we're doing. And there's a wide variety of techniques like this. But, like, it just seems, like, really weird to me to to basically be like, oh, the only way that this will work is if we, like, don't grab the steering wheel. Like, I think we have to grab the steering wheel. I think that's the only way it can work. I don't think we figured out how to do it yet, but somebody's gotta be trying, and some people have to be trying to find different ways to train. And I think there's a lot of ways around to just, like those sort of, like, basic level concerns, I think, are real and they could happen, but I think, like, our ability to detect them isn't, like, totally naive either. I think we just have to do the empirical science. I don't think the theory is gonna get there. I think we have to do the empirical science. And at the end of the day, I don't think any any, like, wide sweeping genre of technique should be forbidden. I think it should be much more about, like, the specifics of the application and how how closely you measured. What are the theory, obviously, that I think is kind of the prevailing one at the moment is defense in-depth. Even if we don't understand the model or we can't effectively shape training, we can just monitor in a bunch of different ways, maybe that'll patch together enough nines that we'll be okay. I have been pretty skeptical of that over time, but I have to say when I read the JSPACE paper, I was like, well, maybe we couldn't get there. The the sort of fact that ablating the seemed to reduce the model's ability to do, like, long horizon, more planning intensive kind of tasks was, like, maybe to borrow a term from Zvi, maybe physics is kind of kind to us in that, like, yikes. We were we're only in 2026. We're only three years since toy models are superposition, and we already have this Yeah. You know, this ability to, like, monitor within this space and also know that or at least have some, like, reasonable sense that if it's not in this space, it's probably not being used in, like, long term planning. How close do you think we are to like being able to monitor well enough? Now of course there's execution competence. So we've been actually doing it and open source questions, but putting those to the side, if we just said like, could we monitor our way to success under ideal conditions of like people actually doing it? Do you think that has hope? Maybe. I'd give that some probability.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence