High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Reiner Pope

Published podcast speaker

Claims
50
Episodes
3
Shows
2
Named items
0

Claim ledger

What Reiner said.

22 transcript-backed records

01 / evaluation

It hurts your throughput, in fact, because the throughput of your chip is the product of how much you get done per clock cycle—which is based on area efficiency—times how many clocks you get per second.

“It hurts your throughput, in fact, because the throughput of your chip is the product of how much you get done per clock cycle—which is based on area efficiency—times how many clocks you get per second.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

05 / evaluation

If you split it in the middle, you can hit twice the clock frequency. That’s great, you get twice the performance, but at the cost of an extra register, which means more storage.

“If you split it in the middle, you can hit twice the clock frequency. That’s great, you get twice the performance, but at the cost of an extra register, which means more storage.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

06 / evaluation

I think the big observation you’ve made is that there’s this quadratic scaling with bit width, which is very effective and is the single reason low-precision arithmetic has worked so well for neural nets.

“I think the big observation you’ve made is that there’s this quadratic scaling with bit width, which is very effective and is the single reason low-precision arithmetic has worked so well for neural nets.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

07 / evaluation

Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.

“Generally, people will go a little bit larger than this. They don’t really want to be exactly at the balance point because real-world efficiencies aren’t as good as a roofline analysis would say.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

08 / evaluation

It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.

“It is simply that I need to read all of my total parameters from memory into the chips, and that takes a certain amount of time. If I use all of my memory bandwidth, I can’t do any better than that.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

10 / evaluation

I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.

“I can’t use different pipeline stages in parallel because they’re not running at the same time, but I can use all the GPUs in my scale-up domain in parallel to load the weights.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

14 / evaluation

If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.

“If I think about what cost means… To run this inference, I’m going to use the GPU for a certain number of seconds, like one millisecond or 20 milliseconds.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

15 / evaluation

A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.

“A way to think of this is that if I have too many of these things sitting in my HBM, if I fill up my HBM with just KV caches that I’m not using, I can’t use that GPU.”
Speaker
Reiner Pope
Publisher
Dwarkesh Podcast

17 / evaluation

I think that's a pretty good fit because you have different parts of your company that are skilled at different areas, and it allows you to keep both of them occupied without having instead every 2 years doing a massive risk release.

“I think that's a pretty good fit because you have different parts of your company that are skilled at different areas, and it allows you to keep both of them occupied without having instead every 2 years doing a massive risk release.”
Speaker
Reiner Pope
Publisher
Cheeky Pint

18 / evaluation

What that means for us is that ecosystem already exists, and that way of operating, where you say, "I'm just going to staff a 50-person team to write software for this chip," works really well if you're trying to sell to frontier labs.

“What that means for us is that ecosystem already exists, and that way of operating, where you say, "I'm just going to staff a 50-person team to write software for this chip," works really well if you're trying to sell to frontier labs.”
Speaker
Reiner Pope
Publisher
Cheeky Pint

20 / evaluation

There's this really cool technique called 'cuckoo hashing' where you hash into two different locations, and then you use the bucket which is less full. It's been in the literature for decades and yet the best hash table implementations don't use it because it's somehow not practical.

“There's this really cool technique called 'cuckoo hashing' where you hash into two different locations, and then you use the bucket which is less full. It's been in the literature for decades and yet the best hash table implementations don't use it because it's somehow not practical.”
Speaker
Reiner Pope
Publisher
Cheeky Pint

21 / evaluation

I think most ML researchers, especially of the time, were not super deep in what hardware wants, and what is… Mechanical sympathy is sometimes a term that's used for that.

“I think most ML researchers, especially of the time, were not super deep in what hardware wants, and what is… Mechanical sympathy is sometimes a term that's used for that.”
Speaker
Reiner Pope
Publisher
Cheeky Pint
Search evidence