Evidence receipt / prediction
Published · transcript-backedAdam Gleave: prediction
30 Jul 2026 The Cognitive Revolution Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
“I expect open weight developers to be the first to have to adopt this because they have fewer options.”
Source trail
Everything needed to verify it.
- Speaker
- Adam Gleave
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 30 Jul 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Yeah. I I think it's a good question. And the the charitable take for developers is that if there's one thing you really don't want to mess with, it is pre training because this is just orders of magnitude more expensive than every other training procedure that that you do. Don't rock the boat, basically. If we've got this recipe that works and we know it's gonna work better, we scale it up, then let's do that. Let's not change anything that we don't need to. And so especially if you're a proprietary developer where you can say, okay. We have all these other methods that we can use to stop misuse of our model. We're gonna lean more on that, and we're gonna not push this onto a pretraining team. So I think that there's some argument to that. Do you think it's overall been an overlooked approach? And we're seeing increasing work that pretraining interventions are important not just for preventing misuse, but also for alignment. Because, ultimately, pretraining is where the model learns most of its representations or values, a lot of its innate drives, and post training is to a first approximation, shifting around personas in an already established persona space. And that's beginning to change as post training is an increasing fraction of overall training time models actually change more in post training. But pre training is really important, and I think it's pretty intuitive. You wouldn't say we're just gonna, not care at all about the upbringing of our child from zero to 12, but the last six years, we're gonna really get that right. You've gotta get both both right for the system to work well. So geodesic research, I'll give a shout out to them. They've been doing a lot of work on pretraining safety interventions and finding that this really improves a sort of overall alignment of a model. And I think Anthropic has been experimenting with us a little bit with things like looking at alignment generalization, how that changes, not just in pretraining, but also mid trainings. You add some synthetic documents partway through training. So I think ultimately, this is something we're we're gonna have to tackle, not just for misuse, but for preventing loss of control. And you can now do quite good work on pretraining experiments on on on really quite capable models for not that much money. So we're looking at scaling up pretraining filtering and gonna be doing not full, but pretty close to replic full replicas of something like NVIDIA's Neutron Nano, and it only costs maybe, like, a $100,000 per run. But that's a lot of money on one hand, but it's something that a nonprofit can afford to to actually do a bunch of runs. And then we're thinking of scaling up to Nematron Super, which is a 120,000,000,000 parameter model, do it training it for enough tokens to be the chinchilla compute optimal point. So training past that point would be wasting training compute, although it would make the model more capable and better for inference. it for enough tokens to be the chinchilla compute optimal point. So training past that point would be wasting training compute, although it would make the model more capable and better for inference. And that costs ballpark $2,000,000. So expensive, but again, well within the range of a number of actors to try for sort of final validation runs. I think there's no reason not to experiment with this, and if the scaling laws look good, you can cautiously incorporate some of these techniques into your pretraining run. You can start by filtering out just a very small percentage of your data. It won't have a big capability here, and then work your way up. So I think that we do need more adoption here. I expect open weight developers to be the first to have to adopt this because they have fewer options. But I hope that proprietary developers also use this, especially for those sort of more loss of control flavored risks. Speaking of loss of control, let's get to the news. So it's here. Right? We are…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.