High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Adam Gleave: evaluation

30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

“You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training.”

— Adam Gleave

Source trail

Everything needed to verify it.

Speaker
Adam Gleave
Attribution
Verified speaker
Claim type
evaluation
Recorded
30 Jul 2026
Publisher
The Cognitive Revolution

Transcript context

…Yeah. It's definitely a a composite of this. You're absolutely right that with increased amount of post training, both these pipelines are more sophisticated and run for a greater fraction of our overall training time. I'm just making a mispredicting for the training distribution distribution of of text text on the Internet. It is no longer a good way of reasoning about them. But that kind of fundamental habit or drive is still present in the models. And so one other thing we find in jailbreaking is that sometimes just a longer conversation works. There's actually this fascinating paper from a a year and a bit ago. Many shot jailbreaking. It's basically just take a jailbreak and say it a lot of times. And it is another one of those things that's so stupid you think it shouldn't work. It's like you go up to a person and say, buy this product. No. Buy it. And then they're like, okay. Fine. I give in. But but you're just stuffing the context window of these models. And that both has a cumulative effect, but it also takes them more off distribution, especially off distribution for the post training because that has usually been quite short context windows, especially for conversations because it's expensive to have longer context windows. And so if you can make the context really full of a model just saying yes to things and helping, that is gonna bias the model even though it has all of this post training. And you could potentially adversarially train against that, so you include a bunch of situations of very long context. And then the model still refute to using when it suddenly gets a harmful request. But that's just more expensive, and you've got exponentially more different things you could have in the context as the context grows. So it's really hard to get adequate coverage for that. So I think that's highlighting maybe a fundamental limitation of of our training techniques that they work really great when you can stay on distribution, and we've been able to get more and more things on distribution or close to being on distribution just by training on more and more data and having synthetic data. But this is intention with these sort of long context windows, which are already hard to get that dataset coverage. I think the persona thing is definitely a powerful predictor. And in some ways, it makes a lot of sense that a model would have a persona, both for pretraining and also post training, and that although these models are are really huge, right, they don't have enough parameters to actually memorize the text because they're trading on a huge amount of text. So you have to have some kind of simpler parametric model of who's writing this text, what are they trying to do. And in some cases, it can be really detailed model because I've heard published authors, they'll put a paragraph of an unpublished book in a model and say, who wrote it? And they're like, you did. They can just recognize their writing style. led model because I've heard published authors, they'll put a paragraph of an unpublished book in a model and say, who wrote it? And they're like, you did. They can just recognize their writing style. But it's still parametric in that it's just modeling different people's styles. And so if you can get the model into the mindset of I'm in some person's style that always says yes to things and gives detailed responses, then congratulation. You've jailbroken the model. And I don't know if there's a hard line between that and next token prediction, or another thing you could say, another thesis people have for why jailbreaks work is that the models have had helpful training and harmless training. Right? So helpful is you say yes to things and do things, and harmless is you don't help people with bad stuff, those objectives are in conflict with each other. But if you can just activate the helpful direction of a model and not the harmful harmless direction, then, again, you've jailbroken the model. And to me, that feels contiguous or consistent with the persona. So these things all blow into one. I don't know if that's a very satisfying answer, but I do think it's better to view these things as kind of different frames on what is ultimately a much more complex underlying system. And all of these are gonna be useful predictions and ultimately, we're at a stage where we just have to test a lot of these things. So these are great for hypothesis generation, but I wouldn't trust any of them too much for knowing what a specific model is gonna do. Are there any other frames that you find useful for hypothesis generation?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence