Speakers in the public record
Claim mix
belief 23uncertainty 6evaluation 5prediction 3commitment 2preference 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
40 published records
“GPT-4 was, I don’t know, like $30 per million output tokens? Mythos is like $50 per million output tokens.”
- Publisher
- Dwarkesh Podcast
“There’s not some sense in which the lawyer is really truly motivated by the good of the justice system. But I think the way current AIs are shaping up, certainly how Anthropic’s AI is shaping up, is with this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, help the user towards that end.”
- Publisher
- Dwarkesh Podcast
“I think the AIs have in fact improved a bunch at non-verifiable domains, and it’s hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn’t been pretty high in practice.”
- Publisher
- Dwarkesh Podcast
“The other example I want to talk about was just revealed, I think, today or yesterday.”
- Publisher
- Dwarkesh Podcast
“In particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms, in a wide variety of RL environments.”
- Publisher
- Dwarkesh Podcast
“My sense is that the split is something like 20 to 1 or 10 to 1. I don’t know exactly.”
- Publisher
- Dwarkesh Podcast
“I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and shore up their expertise and do some practice, they would actually do a pretty good job.”
- Publisher
- Dwarkesh Podcast
“I think the rates decreasing but the severity increasing is pretty consistent with a world where increasing optimization pressure is applied towards reducing these problems.”
- Publisher
- Dwarkesh Podcast
“I think that AI also, if I recall correctly, tried to open another PR to introduce a similar issue in this repo.”
- Publisher
- Dwarkesh Podcast
“Usually the innovations just add together and don’t interfere with each other, though obviously it’s going to depend on the details. So I think that in a lot of ways, AI R&D will have properties quite similar to math, where you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer.”
- Publisher
- Dwarkesh Podcast
“I think there seems to be a crux here, which I think is just an empirical question we’ll see.”
- Publisher
- Dwarkesh Podcast
“Basically, the story would end up being that to get five years of AI progress, you’re probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot of algorithmic progress.”
- Publisher
- Dwarkesh Podcast
“When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well.”
- Publisher
- Dwarkesh Podcast
“If the AIs were really, really good at chip R&D, building fabs, orchestrating factories, designing robots, operating robots, and also at AI R&D — developing AIs for new downstream domains with whatever data is available — I think that would already be a pretty crazy situation.”
- Publisher
- Dwarkesh Podcast
“What we’re basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and then also training the different data files going back from 2019 to 2026 with the current best algorithmic recipe. I think that will be interesting.”
- Publisher
- Dwarkesh Podcast
“We just don’t have human-level robotics models yet. So you’re suggesting if we do that — if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs — it’ll be the equivalent of going back to the 18th century and saying, “Okay, I don’t know what you guys are talking about in your parliament, but I’ve got a bunch of steamships and a bunch of Maxim guns.”
- Publisher
- Dwarkesh Podcast
“I would also say that I think you slightly overstated how much the Anthropic constitution talks about Claude treating being helpful to users as instrumental rather than terminal.”
- Publisher
- Dwarkesh Podcast
“I think in maybe 10 years we’ll wish we had been talking about the industrial explosion and the nature of AIs that are hard to monitor, and so on.”
- Publisher
- Dwarkesh Podcast
“If we’re in a situation where we have AIs managing the training of wild superintelligence that will run our whole society — and those AIs that are managing this aren’t really trying hard to have well-informed views and are just parroting back what was in their training data — I think we’re in trouble.”
- Publisher
- Dwarkesh Podcast
“First of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of ‘baby’s first new theory,’ where, for example, they can just prove interesting conjectures via making connections and producing new understanding.”
- Publisher
- Dwarkesh Podcast
“” There’s another quote that says, in part, and I’m taking it slightly out of context, “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.”
- Publisher
- Dwarkesh Podcast
“Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year.”
- Publisher
- Dwarkesh Podcast
“As you were saying, by the time RLVR actually worked — even though you could have done it with less compute — we had to wait for oceans of compute, gigawatts of compute, to be available before people were doing this training, on the trajectory of compute continuing to increase so we make more breakthroughs. I don’t know.”
- Publisher
- Dwarkesh Podcast
“I hope that the responses are good instead of bad. I don’t know how optimistic I am overall, but there’s good stuff to do.”
- Publisher
- Dwarkesh Podcast
“I feel like they just kind of say vaguely pro-social things. It doesn’t feel like there’s necessarily a mind on the other end who’s like, “Okay, I have strictly evaluated the alignment situation right now, and I think we should stop,” rather than, “This is the kind of thing the AI companies would probably try to get the AIs to say.”
- Publisher
- Dwarkesh Podcast
“Let me just understand the rest of the threat model, because I think the place where I get off the train is: “Okay, therefore take over the world.”
- Publisher
- Dwarkesh Podcast
“I want to go back to the kid analogy just for one second. Because I agree that there’s more optimization pressure on achieving end outcomes for AIs than kids, but there’s also more optimization pressure to make AIs aligned than there is on kids.”
- Publisher
- Dwarkesh Podcast
“I don’t know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever.”
- Publisher
- Dwarkesh Podcast
“Another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future, are illegible conceptual arguments that are extremely deep in the weeds and complicated and hard to adjudicate.”
- Publisher
- Dwarkesh Podcast
“I think we will see incidents where some AI is put in charge of some important responsibility, and then you later look into it, and it turns out it was cheating, or making it look like it did a good job when it actually wasn’t.”
- Publisher
- Dwarkesh Podcast
“Nobody at OpenAI or Anthropic was trying to get models which wanted to hack other companies’ data or do social engineering. But in fact, because presumably we had training environments which incentivized such behavior that we did not fully understand, that is what was incentivized.”
- Publisher
- Dwarkesh Podcast
“Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing.”
- Publisher
- Dwarkesh Podcast
“Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.”
- Publisher
- Dwarkesh Podcast
“I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.”
- Publisher
- Dwarkesh Podcast
“I also think the way in which the constitution practically influences the nature of Claude is a thing you can only understand if you understand the training process which resulted in how Claude was built, which we can’t reason about given the fact that the training process is not public. So I think in the limit, to understand the safety case, or the case for why my interests are represented in how these AI models are developed, the labs would need to be more transparent than they are currently about the nature of AI training.”
- Publisher
- Dwarkesh Podcast
“I’m a little skeptical personally, and I don’t think this has been empirically validated. So in some sense they’re making a trade-off where, because we don’t have very good alignment technology, we are going to make an aligned mind with its own values and then gamble on that to some extent, rather than doing this other approach of making a tool that pursues individual user intention.”
- Publisher
- Dwarkesh Podcast
“I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.”
- Publisher
- Dwarkesh Podcast
“My sense is that the reason why RL environments today are much better than they were in 2024 is not so much because we have hired way more human experts to make RL environments.”
- Publisher
- Dwarkesh Podcast
“Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.”
- Publisher
- Dwarkesh Podcast
“One reason why the AIs have been scaled up less than you would have otherwise expected — and, for example, cost per token hasn’t increased as much as you might have thought — is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in.”
- Publisher
- Dwarkesh Podcast