Evidence receipt / prediction
Published · transcript-backedJeff Dean: prediction
12 Feb 2025 Dwarkesh Podcast Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI
“Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist.”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Dean
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 12 Feb 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Yeah – oh, the what? Yeah, the image people didn't have enough labeled data so they had to invent all this stuff. And they invented -- I mean, dropout was invented on images, but we're not really using it for text mostly. That's one way you could get a lot more learning in a more large-scale model without overfitting is just make like 100 epochs over the world's text data and use dropout. But that's pretty computationally expensive, but it does mean we won't run it. Even though people are saying, "Oh no, we're almost out of textual data," I don't really believe that because I think we can get a lot more capable models out of the text data that does exist. I mean, a person has seen a billion tokens.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.