High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Ryan Greenblatt: belief

11 Aug 2026 Dwarkesh Podcast Ryan Greenblatt – What happens once AI can automate AI research?

“I think that AI also, if I recall correctly, tried to open another PR to introduce a similar issue in this repo.”

— Ryan Greenblatt

Source trail

Everything needed to verify it.

Speaker
Ryan Greenblatt
Attribution
Verified speaker
Claim type
belief
Recorded
11 Aug 2026
Publisher
Dwarkesh Podcast

Transcript context

…Oh my God. That’s crazy. The other GitHub account came back and was like, “No, no, it’s not malicious.” Then the human maintainer shut the PR. I think that AI also, if I recall correctly, tried to open another PR to introduce a similar issue in this repo. Jesus. By the way, one of the many reasons this is scary is I was previously under the impression that the reason reward hacking is not super scary is because the behaviors which directly came up during training are the ones that are up-weighted. It is not the desire for the reward that is up-weighted. So basically, if during training, the Anthropic model escaped the sandbox and got a high score, escaping the sandbox is rewarded, the probability of it escaping the sandbox is increased. But something totally novel, like “I’m going to go talk to somebody in order to get them to merge a PR,” would not be a behavior that came up, so it would not be something that is increased in salience. The reason this matters is that literally taking over the world will not have been part of any training curriculum, but if the AI directly cares about accomplishing an objective, then as a result it could instrumentally take over the world. Did that make sense at all? I hope it did. I feel like maybe I lost the audience.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence