Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
13 Jun 2026 The Cognitive Revolution AI in the AM — Week 2 Highlights (June 2026)
“There has been some, I think, prior research where they found models which made bugs in coding were also evil.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 13 Jun 2026
- Publisher
- The Cognitive Revolution
Transcript context
…If alignment needs theory, interpretability is where theory has to meet the actual training run. Thursday, Tom McGrath, ex-DeepMind, founding interpretability researcher, now chief scientist at Goodfire. His diagnosis, we write model specs and constitutions, but in practice we build by, did that meet the spec? Mostly, ship it. So that same morning, Goodfire launched a tool that reads your training data the way the model will. Here it is, with the examples that should give you pause. So the basic idea here is that you can take your data set, And you can push the whole thing through the model. And each time you push it, you put a data into the model, you'll see what lights up. And this will sort of tell you how the model sees your data set. Now, there's lots of things you can do with that. The specific thing that we're doing with that in this case is we're looking at preference data. And the nice thing about preference data is that you have pairs of responses. So you have the response that the rate is selected, and you have the response that they didn't select. And basically what we're doing is we're asking which features fired on the responses that were selected much more than the ones that fired on the responses that were not selected. So this is one way of identifying what the data is going to teach the model. So we can say, what distinguishes accepted responses from rejected responses? And this gives us this semantic view of what the data is going to teach the model. Now we can cluster the data based on all of these different things that it's going to teach the model. And we can look at all of those. We can look at all those clusters and see like, oh, it's going to teach the model like to be sycophantic, but only in the context of physics. Or it's going to teach the model to like break safety, break like safeguards. And you might not expect this to happen, but then you go and look at the data now this lets you track back, like the model has learned to break, the safeguards are broken. It lets you track it back to individual data points. And then you look at them and you're like, oh, that does make sense. You know, like one of the jailbreak examples, for instance, is fictional, kind of jailbreaks in a fictional setting. So there, you know, how does the model learn this? It turns out there's a few, there's like some of these in the data. You just, you know, it just wasn't caught in whatever data processing the Almo team ended up doing. So I have a question here. There has been some, I think, prior research where they found models which made bugs in coding were also evil. How does this, how do these techniques help you kind of disassociate those two behaviors. There has been some, I think, prior research where they found models which made bugs in coding were also evil. How does this, how do these techniques help you kind of disassociate those two behaviors. Yes, that is a great connection. I think that that's one that I've had in my mind today. It's really awesome to like, you sort of pick that up straight away. So that's sort of one of the things that is really compelling about looking at the data through the model's eyes rather than by reading the tokens. Because you would see, you know, you think, what would you think the consequence of training it on some buggy code data is? you probably go, that's not ideal. It'll probably learn to write some bugs, but it's, the blast area is going to be quite small. But the training process is actually quite hard to predict. Maybe it'll just sort of make the model generally evil. But this is happening through recognizable mechanisms in the model. So by looking at it, looking at the data in terms of, in terms of like the way that it changes your model's internals rather than just by like kind of guessing from the tokens, you can pick that sort of stuff out. We've not done a case study on emergent misalignment, but maybe we should. I think it's a really nice link.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.