High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Leopold Aschenbrenner: belief

4 Jun 2024 Dwarkesh Podcast Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history

“I think there’s a lot of promising ML research on aligning superhuman systems, which we can discuss more later.”

— Leopold Aschenbrenner

Source trail

Everything needed to verify it.

Speaker
Leopold Aschenbrenner
Attribution
Verified speaker
Claim type
belief
Recorded
4 Jun 2024
Publisher
Dwarkesh Podcast

Transcript context

…After FTX imploded and you were out, you went to OpenAI. The superalignment team had just started. You were part of the initial team. What was the original idea? What compelled you to join? The alignment teams at OpenAI and other labs had done basic research and developed RLHF. reinforcement learning from human feedback. That ended up being a really successful technique for controlling current AI models. Our task was to find the successor to RLHF. The reason we need that is that RLHF probably won’t scale to superhuman systems. RLHF relies on human raters giving feedback, but superintelligent models will produce complex outputs beyond human comprehension. It’ll be like a million lines of complex code and you won’t know at all what’s going on anymore. How do you steer and control these systems? How do you add side constraints? I joined because I thought this was an important and solvable problem. I still do and even more so. I think there’s a lot of promising ML research on aligning superhuman systems, which we can discuss more later. It was so solvable, you solved it in a year. It’s all over now.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence