Evidence receipt / belief
Published · transcript-backedScott Alexander: belief
3 Apr 2025 Dwarkesh Podcast AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo
“Here it was a prompt, but I think very soon it’s going to be the spec where it’s more of an agent and it’s understanding the spec on a deeper level and just thinking about that.”
Source trail
Everything needed to verify it.
- Speaker
- Scott Alexander
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 3 Apr 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…There’s other types of transparency too. So transparency about capabilities and transparency about the spec and the governance structure. So for the capabilities thing, that’s pretty simple. If you’re doing an intelligence explosion, you should keep the public informed about that. When you’ve finally got your automated army of AI researchers that are completely automating the whole thing on the data center, you should tell everyone, “hey, guys, FYI, this is what’s happening now. It really is working. Here are some cool demos”. That’s an example of transparency. And then in the lead up to that, I just want to see more benchmark scores and more freedom of speech for employees to talk about their predictions for AGI timelines and stuff, so that blah, blah, blah. And then for the model spec thing, this is a concentration of power thing, but also an alignment thing. The goals and values and principles and intended behaviors of your AIs should not be a secret. You should be transparent about, here are the values that we’re putting into them. There’s actually a really interesting foretaste of this. At some point somebody asked Grok, who is the worst spreader of misinformation? And I think it just refused to respond “Elon Musk”. Somebody kind of jailbroke it into telling it its prompt and it was like, “don’t say anything bad about Elon”. And then there was enough of an outcry that the head of XAI said, “actually that’s not consonant with our values. This was a mistake. We’re going to take it out”. So we kind of want more things like that to happen. Here it was a prompt, but I think very soon it’s going to be the spec where it’s more of an agent and it’s understanding the spec on a deeper level and just thinking about that. And if it says like, “by the way, try to manipulate the government into doing this or that”, then we know that something bad has happened and if it doesn’t say that, then we can maybe trust it. Right. Another example of this, by the way. So, first of all, kudos to OpenAI for publishing their model spec. They didn’t have to do that, I think they might have been the first to do that and it’s a good step in the right direction. If you read the actual spec, it has like a sort of escape clause where there’s some important policies that are top level priority in the spec that overrule everything else that we’re not publishing, and that the model is instructed to keep secret from the user. And it’s like, “what are those? That seems interesting. I wonder what that is”. I bet it’s nothing suspicious right now. Now it’s probably something relatively mundane like “don’t tell the users about these types of bioweapons and you have to keep this a secret from the users because otherwise they would learn about these”. Maybe. But I would like to see more scrutiny towards this sort of thing going forward. I would like it to be the case that companies have to have a model spec, they have to publish it insofar as there are any redactions from it, there has to be some sort of independent third party that looks at the redactions and makes sure that they’re all kosher. And this is quite achievable. And I think it doesn’t actually slow down the companies at all. And it seems like a pretty decent ask to me.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.