Evidence receipt / evaluation
Published · transcript-backedBen Todd: evaluation
26 May 2026 The Cognitive Revolution Your Biggest Lever: Designing your AI Career for Maximum Impact, with 80,000 Hours founder Ben Todd
“They have some of the strongest research teams. And they're also in a great position to actually implement the research, which I think is actually a big part of it because you can come up with some idea, but if someone doesn't actually see through all the details in implementing it in the product, it doesn't actually help that much.”
— Ben Todd
Source trail
Everything needed to verify it.
- Speaker
- Ben Todd
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 May 2026
- Publisher
- The Cognitive Revolution
Transcript context
…Fun fact, I have had one historian on the podcast in the 340 some episodes that we've done. And Mark Humphries, he was doing some pretty interesting stuff in terms of just using, and of course his pipeline has changed a lot, but using AI models to transcribe and make sense of all these old documents that exist in the Canadian archives that just nobody has really ever looked at. But also because his problem is so different from almost anybody else's problem that's using AI intensively, he had a really interesting and quite viral post on Gemini 3 just before it came out showing how zeroing in on a particular ability that it had unlocked to kind of reason about what it was seeing in the visual documents in ways that prior models just hadn't been able to do. And I think that is a really interesting example of how something that just seems so far afield, because AI itself is getting so far afield, there is even in such a unexpected niche, there is the opportunity to make interesting discoveries that really contribute back into the mainline discourse and advance people's understanding. So I would again just encourage people to think pretty broadly about just how many different opportunities there might be. And I also can say again from this event this past weekend in San Francisco, yes, Meter in particular is desperately trying to find people. All their sign of the times All their, this is maybe a little bit of an exaggeration, but the attitude was, all of our measurements are pretty much saturated and it's getting tough to take the skill size higher than we already have. So there's a lot of work to be done as the models are racing past our ability to measure them. They are looking to staff up and get them, bring as much talent to bear as they can on keeping a handle on exactly what the current capability frontier really looks like. Many of the things that you outlined there notably happen both in the frontier companies that are developing the AIs and in a variety of other organizations that kind of orbit them in some ways, in some ways check them, in some ways maybe even oppose them. This has been debated for a long time and I think people have very different intuitions and some the discourse has gone around in circles at times. How do you think right now about whether or not people should try to go to the frontier companies and contribute there versus trying to make perhaps holding the person themselves constant and the kind of contribution they're going to make versus doing that from some other outside angle? Yeah, and I think the short answer is it's complicated. But yeah, if you want to do technical alignment and control research. They are some of the best places in the world to do that research. They have some of the strongest research teams. And they're also in a great position to actually implement the research, which I think is actually a big part of it because you can come up with some idea, but if someone doesn't actually see through all the details in implementing it in the product, it doesn't actually help that much. But on the other hand, a lot of people have done great research outside of the labs. I would name Redwood Research as another group here who helped to pioneer the AI control agenda. A bunch of other really useful pieces of work. They did the deceptive alignment research with Anthropic. So it's definitely possible to do work outside of the labs. Yeah, I think another big factor is a big, the big argument against is maybe you're speeding up AI development, which is hastening the end of, it's hastening the risks. And I think that is worth thinking hard about and each individual has to get judgment on that themselves and how they want to relate to that. I think also another big factor is how aligned you are with the people who work in the labs. And in general, I try to avoid adversarial strategies, but if you have the attitude, well, AI is happening and I prefer to have a more social minded or safety conscious company win and I'm going to try and help them win, I don't think that should be entirely ruled out as a strategy. I think a lot of it just comes down to, I think what drives a lot of the disagreement here is just what P doom is basically. People who think, we're very likely to have an existential risk and alignment research doesn't really have any hope of working, basically think the only option is to have an indefinite pause. And by working at the labs, you're not really helping with that or actively making it worse. Whereas people who more have the attitude that A, this is happening, we can't really stop it, and B, we'll probably get through, but it's more a question of making the chances as high as possible. tend to be much more keen on working at the lab. So I think that's the other, that's another really big source of disagreements over this. Yeah, there's so many different angles on it. I find them all compelling, but I do also find myself going back and forth on it at different times. I guess maybe just to name a couple of examples that people might want to check out or a couple of arguments that people might want to check out. One, which I think it came from Redwood Research, is the people on the inside argument, which is basically that unless there is some international treaty or whatever, which I certainly don't rule out, but we're not that close to it at this moment, then this is going to happen. And if it's going to happen, then a small number of people working within the companies who really care about the right things and can take important actions at critical moments could be one of the most important places to be, one of the highest leverage places to be, certainly. Then another argument, especially if you want to do alignment or interpretability research, is you can do a lot of that research outside of the companies, but one of the big things that we're worried about is that models are going to change in important ways, potentially in a pretty short period of time potentially do, and potentially in surprising ways, even to people within the companies due to emergent properties that do arise, right? Like we see the general pattern of AI capabilities advances is like the loss function is dropping smoothly, but the specific tasks that the model can do often seems to have these more like discrete jumps from one generation to the next. And we have seen obviously many of those on the capability side, but also some of them on the problematic behavior side, where it seems like we've gone from models didn't really do any sort of deception to now they sometimes do some deception. And obviously that's not something that the company's trained for, but with a certain scale of RL or what have you, all of a sudden that seems to pop up from one generation to the next. So again, being on the inside of the companies and having access to the frontier models might be pretty important to being able to discover that next bad behavior in time to make a difference about it or to just understand the internal of your thinking interpretability, to understand the internal workings of models well enough to or soon enough or at the frontier that matters the most to really make the biggest difference. I'd be interested in your reaction to all of those, if you have any. I guess I would say for me, the alignment thing kind of resonates more than the interp side. I feel like with alignment, you do see these kind of qualitative shifts in behavior. And I think trying to align whatever is the best open source model at this moment is probably quite a different activity from doing alignment work on Mythos, for example. Whereas our interpretability understanding is still sufficiently nascent, you could probably still make quite a bit of headway on something like a GPT-OSS or any of the Chinese models.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.