High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Jeff Dean: prediction

12 Feb 2025 Dwarkesh Podcast Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI

“Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist.”

— Jeff Dean

Source trail

Everything needed to verify it.

Speaker
Jeff Dean
Attribution
Verified speaker
Claim type
prediction
Recorded
12 Feb 2025
Publisher
Dwarkesh Podcast

Transcript context

…Yeah – oh, the what? Yeah, the image people didn't have enough labeled data so they had to invent all this stuff. And they invented -- I mean, dropout was invented on images, but we're not really using it for text mostly. That's one way you could get a lot more learning in a more large-scale model without overfitting is just make like 100 epochs over the world's text data and use dropout. But that's pretty computationally expensive, but it does mean we won't run it. Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist. I mean, a person has seen a billion tokens.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence