High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / uncertainty

Published · transcript-backed

Leopold Aschenbrenner: uncertainty

4 Jun 2024 Dwarkesh Podcast Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history

“I don't know if I believe in this concept of a sharp left turn, but there are important qualitative changes that happen between now and somewhat superhuman systems early on in the intelligence explosion.”

— Leopold Aschenbrenner

Source trail

Everything needed to verify it.

Speaker
Leopold Aschenbrenner
Attribution
Verified speaker
Claim type
uncertainty
Recorded
4 Jun 2024
Publisher
Dwarkesh Podcast

Transcript context

…No, but here’s the thing. Even with the current leadership, you can find them in interviews and blog posts talking about it. You talked about what happens and it’s not just you. Jan talked about it in his Tweet thread. There is some trade-off that has to be made with doing a flashy release this week and not next week because Google I/O is next week or whatever. The trade-off is made in favor of the more careless decision. The government, the national security advisor or the military or whatever, is much less familiar with this kind of discourse. They’re not like, “I'm worried the chain-of-thought is unfaithful. How do we think about the features that are represented here?” Why should we be optimistic that a project run by people like that will be thoughtful about these kinds of considerations? You seem pretty optimistic about alignment. Let’s get to the source of the optimism. You laid out different worlds in which we could get AI. There's one that you think has a low probability of happening next year, where GPT-5 plus scaffolding plus unhobblings gets you to AGI. There are also scenarios where it takes much longer. GPT-4 seems pretty aligned in the sense that I don't expect it to go off the rails. Maybe with scaffolding, things might change. It looks pretty good, and maybe you will keep turning the cranks, and one of them gets you to ASI. Is there any point at which the sharp left turn happens? Do you think it's plausible that when they start acting more like agents, this is something to worry about? Is there anything qualitative that you expect to change with regards to the alignment perspective? I don't know if I believe in this concept of a sharp left turn, but there are important qualitative changes that happen between now and somewhat superhuman systems early on in the intelligence explosion. There are also important qualitative changes that occur from early in the intelligence explosion to true superintelligence in all its power and might. Let's talk about both of those. The first part of the problem is one we're going to have to solve ourselves. We have to align the initial AI and the intelligence explosion, the sort of automated Alec Radford. There are two important things that change from GPT-4. If you believe the story on synthetic data RL, self-play, to get past the data wall, and if you believe this unhobbling story, at the end you're going to have things that are agents. They’ll do long-term planning. They have long horizons, which is a prerequisite to being able to do automated AI research. Pre-training is alignment-neutral in the sense that it has good representations and representations of doing bad things, but it's not scheming against you. Misalignment can arise once you're doing more long-horizon training. For example, if you're training an AI to make money using reinforcement learning, it might learn to commit fraud, lie, deceive, or seek power simply because those are successful strategies in the real world. With RL, it explores, maybe it tries to hack and then it gets some money. If that’s successful, that gets reward and that’s just reinforced. There’s more serious misalignments, like misaligned long-term goals, that necessarily have to be able to arise if you’re able to get long-horizon systems. Let’s swap. What you want to do in that situation is add side constraints, like "don't lie," "don't deceive," or "don't commit fraud." How do you add those side constraints? The basic idea you might have is RLHF. You have this goal of making money, but you're watching what it's doing. If it starts trying to lie, deceive, commit fraud, or break the law, you give it a thumbs down and anti-reinforce that behavior. The critical issue that arises is that these AI systems are becoming superhuman and will be able to do things that are too complex for humans to evaluate. Even early on in the intelligence explosion, the automated AI researchers and engineers might write millions, billions, or trillions of lines of complicated code. You won't understand what they're doing anymore. In those millions of lines of code, you don't know if it's hacking, exfiltrating itself, or trying to go for the nukes. You don’t know anymore. Thumbs up, thumbs down pure RLHF doesn't fully work anymore in this scenario. There’s a hard technical problem of what do you do post-RLHF but it’s a solvable problem. There’s various things I’m bullish on. There’s ways in which deep learning has shaped out favorably. The second part of the picture is going from your initial systems in the intelligence explosion to superintelligence, many OOMs of improvement. ’s ways in which deep learning has shaped out favorably. The second part of the picture is going from your initial systems in the intelligence explosion to superintelligence, many OOMs of improvement. By the end of it, you have a thing that's vastly smarter than humans. The intelligence explosion is really scary from an alignment point of view. If you have this rapid intelligence explosion in less than a year or two, you're going from systems where failure would be bad but not catastrophic to a world where if something goes awry, the AI could exfiltrate itself, start hacking the military, and do really bad things. In less than a year, you're going from a world where the AI is some descendant of current systems that you understand and has good properties. It becomes something that potentially has a very alien and different architecture after having gone through another decade of ML advances. One salient example is legible and faithful chain-of-thought. A lot of the time when we're talking about these things, we're talking about how it has tokens of thinking and then uses many tokens of thinking. Maybe we bootstrap ourselves by pre-training it to learn to think in English, and then we do something else on top so it can do longer chains of thought. It's very plausible to me that for the initial automated alignment researchers, we don't need to do any complicated mechanistic interpretability. You can just read what they're thinking, which is a huge advantage. However, it's very likely not the most efficient way to do it. There's probably some way to have a recurrent architecture with all internal states. That's a much more efficient way to do it. That's what you get by the end of the year. You're going in this year from RLHF++ to something that's vastly superhuman. To us, it might be like an expert in the field compared to an elementary or middle school student. It’s an incredibly hairy period for alignment. The thing you do have is the automated AI researchers. You can use them to also do alignment.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence