Evidence receipt / belief
Published · transcript-backedDan Balsam: belief
8 Aug 2026 The Cognitive Revolution Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
“I think the JSPACE results have not there was like a weak version of the JSPACE claim which is true and it's very interesting and it's really good work.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Balsam
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 8 Aug 2026
- Publisher
- The Cognitive Revolution
Transcript context
…What are the theory, obviously, that I think is kind of the prevailing one at the moment is defense in-depth. Even if we don't understand the model or we can't effectively shape training, we can just monitor in a bunch of different ways, maybe that'll patch together enough nines that we'll be okay. I have been pretty skeptical of that over time, but I have to say when I read the JSPACE paper, I was like, well, maybe we couldn't get there. The the sort of fact that ablating the seemed to reduce the model's ability to do, like, long horizon, more planning intensive kind of tasks was, like, maybe to borrow a term from Zvi, maybe physics is kind of kind to us in that, like, yikes. We were we're only in 2026. We're only three years since toy models are superposition, and we already have this Yeah. You know, this ability to, like, monitor within this space and also know that or at least have some, like, reasonable sense that if it's not in this space, it's probably not being used in, like, long term planning. How close do you think we are to like being able to monitor well enough? Now of course there's execution competence. So we've been actually doing it and open source questions, but putting those to the side, if we just said like, could we monitor our way to success under ideal conditions of like people actually doing it? Do you think that has hope? Maybe. I'd give that some probability. I think the JSPACE results have not there was like a weak version of the JSPACE claim which is true and it's very interesting and it's really good work. I don't think the strong version of the JSPACE claim is true. I think models use all type of and it's very hard to isolate a subspace with a very simple technique that will give you the whole picture. But, of course, I believe that models are also decomposable and they're factorable and we're, like, making a lot of progress here. I think one of the big challenges…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.