Evidence receipt / recommendation
Published · transcript-backedLuke Drago: recommendation
1 Jan 2026 The Cognitive Revolution Confronting the Intelligence Curse, w/ Luke Drago of Workshop Labs, from the FLI Podcast
“If you are someone who thinks doom is really likely, the best thing to do is not like continue to evaluate the model to see if we're getting closer, because if we're getting closer, we're going to actually have to do something about it.”
Source trail
Everything needed to verify it.
- Speaker
- Luke Drago
- Attribution
- Verified speaker
- Claim type
- recommendation
- Recorded
- 1 Jan 2026
- Publisher
- The Cognitive Revolution
Transcript context
…How do we do that though? I guess that's the main worry with open weights models. This is just we can't If we put something out there that's open weights, we can't then take it back. Exactly. So we don't have this feedback loop of trying to test something and then pulling back and then perhaps putting a more limited version of that model out there. So how do we deal with the technology where if we release it, that capability suite is now out there indefinitely? Yeah, this is where, like, again, I'll cite Kyle O'Brien's work here is just quite important. The kinds of work that you want to do here to create tamper-resistant open weights models, such that reintroducing the information by trying to tune them in a certain way breaks them or doesn't work. I don't have a lot, I know I've talked with Kyle a bunch, and I know some of his work is forthcoming, so I don't want to jump the gun on anything here. But as a separate note, the kind of holy grail here is a model that when you try to reintroduce this, it just stops working or it breaks because of something that they've done. I don't want to preempt any announcements. I know there are people who are working on this in a broad variety of sectors, but that's the kind of safety innovations that I think are extremely important and that move our option space. If you are someone who thinks doom is really likely, the best thing to do is not like continue to evaluate the model to see if we're getting closer, because if we're getting closer, we're going to actually have to do something about it. And I think from a technical safety perspective, right now, you're either betting on this catastrophic warning shot that I'm not convinced actually slows anything down. I think we have a footnote, like 7 paragraph footnote, the intelligence curse. We couldn't fit in the main thing, I footnoted it, talking about how in a whole lot of scenarios, a warning shot actually just increases the speed at which AI progress happens because somebody gets spooked over it. And the response is we need better defenses faster. So I think if you're counting on like, we're going to keep evaluating the thing and then we're going to see that it's dangerous and we're going to stop building it. Best of luck. Like, I don't think that is an extremely tractable approach. I think more investment is better spent. by a whole lot of extremely talented technical experts on actually building out the capabilities that are required to make even open weight models tamper resistant and safe. And I think this is genuinely achievable. I don't think this is an intractable agenda. We have seen more progress on it than I expected to see. And I think as people have kind of chipped at it, as papers have made it clear that this could be possible, more and more people are starting to get excited about this. And I think that's more of the direction I want to go here. If we don't have the option of controlling AI using a central authority, it seems to me that we are somewhat at the mercy of how the technology just turns out to be. So if it is the case that we can limit what models can output and perhaps have the models stop if you try to use them to create biological threats, say, that's great. But what about the next the next possible danger and the next possible danger. If we don't have a way to control AI as at least a backup option, are we just kind of at the mercy of how the technology turns out to work?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.