High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Jeff Dean

Published podcast speaker

Claims
57
Episodes
2
Shows
2
Named items
2

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

paper / likes

Rich Sutton's paper

“I really like Rich Sutton's paper that he wrote about the Bitter Lesson and the Bitter Lesson effectively is this nice one-page paper but the essence of it is you can try lots of approaches, but the two techniques that are incredibly effective are learning and search.”

Dwarkesh Podcast · 12 Feb 2025

Evidence receipt · Source ↗

service / likes

Google

“One of the things I like about Google is our ambition has always been sort of something that would require pretty advanced AI. Because I think organizing the world's information and making it universally accessible and useful, actually there is a really broad mandate in there.”

Dwarkesh Podcast · 12 Feb 2025

Evidence receipt · Source ↗

Claim ledger

What Jeff said.

36 transcript-backed records

01 / belief

I mean, I think distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on.

“I mean, I think distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on.”
Speaker
Jeff Dean
Publisher
Latent Space

03 / belief

You could generate way more code, uh, and check that the code is cracked with a chain of thought reasoning. So I think, you know, being able to do that at 10,000 tokens per second would be awesome.

“You could generate way more code, uh, and check that the code is cracked with a chain of thought reasoning. So I think, you know, being able to do that at 10,000 tokens per second would be awesome.”
Speaker
Jeff Dean
Publisher
Latent Space

05 / belief

Yeah, my, my, uh, you know, it’s just on this, like, like big prompting and, and, uh, iteration, you know, I think that coming back to your latency point, um, I always, I always try to one, one AB test or experiment or benchmark or research I would like is what is the performance difference between, let’s say three dumb fast model calls with human alignment because the human will correct human alignment, being human looks at the first one and produces a new prompt.

“Yeah, my, my, uh, you know, it’s just on this, like, like big prompting and, and, uh, iteration, you know, I think that coming back to your latency point, um, I always, I always try to one, one AB test or experiment or benchmark or research I would like is what is the performance difference between, let’s say three dumb fast model calls with human alignment because the human will correct human alignment, being human looks at the first one and produces a new prompt.”
Speaker
Jeff Dean
Publisher
Latent Space

07 / belief

Some are, you know, our pro scale model and we can distill from that as well into our Flash scale model. So I think, you know, it’s an important set of capabilities to have and also inference time scaling.

“Some are, you know, our pro scale model and we can distill from that as well into our Flash scale model. So I think, you know, it’s an important set of capabilities to have and also inference time scaling.”
Speaker
Jeff Dean
Publisher
Latent Space

11 / belief

Yeah, I’m, I’m a big believer in pushing on latency because I think being able to have really low latency interactions with a system you’re using is just much more delightful than something that is, you know, 10 times as slow or 20 times as slow.

“Yeah, I’m, I’m a big believer in pushing on latency because I think being able to have really low latency interactions with a system you’re using is just much more delightful than something that is, you know, 10 times as slow or 20 times as slow.”
Speaker
Jeff Dean
Publisher
Latent Space

12 / belief

I mean, I think going to an LLM based representation of text and words and so on enables you to get out of the explicit hard notion of, of particular words having to be on the page, but really getting at the notion of this topic of this page or this page.

“I mean, I think going to an LLM based representation of text and words and so on enables you to get out of the explicit hard notion of, of particular words having to be on the page, but really getting at the notion of this topic of this page or this page.”
Speaker
Jeff Dean
Publisher
Latent Space

13 / belief

Um, but, uh, you know, I think combining a lot of those techniques and really just trying to push on scaling things up over the last, you know, 15 years has been, you know, really important.

“Um, but, uh, you know, I think combining a lot of those techniques and really just trying to push on scaling things up over the last, you know, 15 years has been, you know, really important.”
Speaker
Jeff Dean
Publisher
Latent Space

14 / belief

Yeah, I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that’s where you see what capabilities now exist that didn’t exist at the sort of slightly less capable last year’s version or last six months ago version.

“Yeah, I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that’s where you see what capabilities now exist that didn’t exist at the sort of slightly less capable last year’s version or last six months ago version.”
Speaker
Jeff Dean
Publisher
Latent Space

16 / belief

Like robots or, you know, various kinds of health modalities, x-rays and MRIs and imaging and genomics information. And I think there’s probably hundreds of modalities of data where you’d like the model to be able to at least be exposed to the fact that this is an interesting modality and has certain meaning in the world.

“Like robots or, you know, various kinds of health modalities, x-rays and MRIs and imaging and genomics information. And I think there’s probably hundreds of modalities of data where you’d like the model to be able to at least be exposed to the fact that this is an interesting modality and has certain meaning in the world.”
Speaker
Jeff Dean
Publisher
Latent Space

17 / belief

Um, I still think there’s a tremendous distance we can go from where we are today in terms of energy efficiency with sort of, uh, much better and specialized hardware for the models we care about.

“Um, I still think there’s a tremendous distance we can go from where we are today in terms of energy efficiency with sort of, uh, much better and specialized hardware for the models we care about.”
Speaker
Jeff Dean
Publisher
Latent Space

18 / belief

I think the early sort of four or five years at Google when I was one of a handful of people working on search and crawling and indexing systems, our traffic was growing tremendously fast.

“I think the early sort of four or five years at Google when I was one of a handful of people working on search and crawling and indexing systems, our traffic was growing tremendously fast.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

19 / belief

One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, I think then you have a level of control to understand how is it being used and put some boundaries on what it can do.

“One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, I think then you have a level of control to understand how is it being used and put some boundaries on what it can do.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

20 / belief

You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are so you should just let the hardware dictate things a little bit.

“You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are so you should just let the hardware dictate things a little bit.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

23 / belief

I think one thing people should be aware of is that the improvements from generation to generation of these models often are partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture, the training data mix, and so on, that really makes the model better per flop that is applied to the model, so I think that's a good realization.

“I think one thing people should be aware of is that the improvements from generation to generation of these models often are partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture, the training data mix, and so on, that really makes the model better per flop that is applied to the model, so I think that's a good realization.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

24 / belief

I've been a co-author on a paper called "Shaping AI," which is, you know, those two extreme views often kind of view our role as kind of laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is, you know, maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare.

“I've been a co-author on a paper called "Shaping AI," which is, you know, those two extreme views often kind of view our role as kind of laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is, you know, maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

26 / belief

The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces, but you can always distill from that.

“The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces, but you can always distill from that.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

27 / belief

You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happy.

“You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happy.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

30 / belief

Yeah, I mean, I think most things you don't even try to stack because the initial experiment didn't work that well, or it showed results that aren't that promising relative to the baseline.

“Yeah, I mean, I think most things you don't even try to stack because the initial experiment didn't work that well, or it showed results that aren't that promising relative to the baseline.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

31 / belief

I think we do see some examples in our own experimental work of things where if you apply more inference time compute, the answers are better than if you just apply 10x, you can get better answers than x amount of computed inference time.

“I think we do see some examples in our own experimental work of things where if you apply more inference time compute, the answers are better than if you just apply 10x, you can get better answers than x amount of computed inference time.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

32 / belief

You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now because we can make it work, our ML researchers have been really happy how far we've been able to push synchronous training because it is an easier mental model to understand.

“You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now because we can make it work, our ML researchers have been really happy how far we've been able to push synchronous training because it is an easier mental model to understand.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast

34 / belief

I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive, "respond in 40 seconds" kind of thing but might go off for 10 minutes and might interrupt you after five minutes saying, "I've done a lot of this, but now I need to get some input.

“I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive, "respond in 40 seconds" kind of thing but might go off for 10 minutes and might interrupt you after five minutes saying, "I've done a lot of this, but now I need to get some input.”
Speaker
Jeff Dean
Publisher
Dwarkesh Podcast
Search evidence