Evidence receipt / prediction
Published · transcript-backedJeff Dean: prediction
12 Feb 2025 Dwarkesh Podcast Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI
“I thought, naive me, that 32 processors would be able to train really awesome neural nets. But it turned out we needed about a million times more compute before they really started to work for real problems, but then starting in the late 2008, 2009, 2010 timeframe, we started to have enough compute, thanks to Moore's law, to actually make neural nets work for real things.”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Dean
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 12 Feb 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…Through your careers, at various times, you’ve worked on things that have an uncanny resemblance to what we're actually using now for generative AI. In 1990, Jeff, your senior thesis was about backpropagation. And in 2007- this is the thing that I didn’t realise until I was prepping for this episode – in 2007 you guys trained a two trillion token N-gram model for language modeling. Just walk me through when you were developing that model. Was this kind of thing in your head? What did you think you guys were doing at the time? Let me start with the undergrad thesis. I got introduced to neural nets in one section of one class on parallel computing that I was taking in my senior year. I needed to do a thesis to graduate, an honors thesis. So I approached the professor and I said, "Oh, it'd be really fun to do something around neural nets." So, he and I decided I would implement a couple of different ways of parallelizing backpropagation training for neural nets in 1990. I called them something funny in my thesis, like "pattern partitioning" or something. But really, I implemented a model parallelism and data parallelism on a 32-processor Hypercube machine. In one, you split all the examples into different batches, and every CPU has a copy of the model. In the other one, you pipeline a bunch of examples along to processors that have different parts of the model. I compared and contrasted them, and it was interesting. I was really excited about the abstraction because it felt like neural nets were the right abstraction. They could solve tiny toy problems that no other approach could solve at the time. I thought, naive me, that 32 processors would be able to train really awesome neural nets. But it turned out we needed about a million times more compute before they really started to work for real problems, but then starting in the late 2008, 2009, 2010 timeframe, we started to have enough compute, thanks to Moore's law, to actually make neural nets work for real things. That was kind of when I re-entered, looking at neural nets. But prior to that, in 2007... Sorry, actually could I ask about this?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.