High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Reiner Pope: evaluation

26 Feb 2026 Cheeky Pint Reiner Pope of MatX on accelerating AI with transformer-optimized chips

“There's this really cool technique called 'cuckoo hashing' where you hash into two different locations, and then you use the bucket which is less full. It's been in the literature for decades and yet the best hash table implementations don't use it because it's somehow not practical.”

— Reiner Pope

Source trail

Everything needed to verify it.

Speaker
Reiner Pope
Attribution
Verified speaker
Claim type
evaluation
Recorded
26 Feb 2026
Publisher
Cheeky Pint

Transcript context

…Where else have you… Chip optimization is one domain, where else? When one of the really exciting things I found about working at Google is that the whole Google code base is available, and you can look at how does a memory allocator work, how does a mutex work, how does a HashMap work, any of those things, and you can go and look inside the implementations. Google has excellent implementations of those, some of the best you could write. One of the things I did on my nights and weekends when I was at Google was just to go find those implementations, write a benchmark. How many nanoseconds does it take to allocate eight bytes of memory? Can I make that faster? Can I maybe I inline this function? Maybe I look at the assembly and say, "It looks like there's a few memory moves here or there are some registers that are being used that I don't need in the fast path and in the slow path. Can I do something there?" I don't know, that was always my just fun and learning activity. Being outside of Google, I feel, I probably could have done this inside of Google as well, but outside of Google, I felt the luxury to be able to talk about these results as well. One of the things I've looked at recently is just hash tables are used so much. One prompt for me was, if I wanted to design custom CPU instructions for accelerating hash tables… Hash tables are one of the most common things. I'm looking at them up and writing them all the time. What would the optimal CPU be for that? Following down that chain is like, what is the best hash table implementation in the first place? I spent some time looking at different SIMD implementations. There's this really cool technique called 'cuckoo hashing' where you hash into two different locations, and then you use the bucket which is less full. It's been in the literature for decades and yet the best hash table implementations don't use it because it's somehow not practical. Why is it not practical?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence