High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Ryan Greenblatt

Published podcast speaker

Claims
25
Episodes
1
Shows
1
Named items
0

Claim ledger

What Ryan said.

3 transcript-backed records

01 / prediction

Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.

“Now these AIs might end up being very seriously misaligned, because things have just been getting worse and worse over model generations while the problems that we’ve been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren’t.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

02 / prediction

I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.

“I think the alignment eval that’s most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast

03 / prediction

Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.

“Then at the point when we’re passing off safety R&D, the AIs are capable enough to automate safety R&D and trying really hard to do a good job on it, because that’s the sort of thing that would’ve been incentivized in training, either very directly or through good enough generalization.”
Speaker
Ryan Greenblatt
Publisher
Dwarkesh Podcast
Search evidence