Evidence receipt / prediction
Published · transcript-backedTri Dao: prediction
26 Jul 2023 Latent Space FlashAttention 2: making Transformers 800% faster w/o approximation - with Tri Dao of Together AI
“The current implementation ended up using NVIDIA's library or primitives, but I expect these ideas to be broadly applicable to different hardware.”
— Tri Dao
Source trail
Everything needed to verify it.
- Speaker
- Tri Dao
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 26 Jul 2023
- Publisher
- Latent Space
Transcript context
…And since it's an NVIDIA library, can you only run this on CUDA runtimes? Or could you use this and then run it on an AMD GPU? Yeah, so it's an NVIDIA library. So right now, the code we release runs on NVIDIA GPUs, which is what most people are using to train models. Of course, there are emerging other hardware as well. So the AMD folks did implement a version of Flash Attention, I think last year as well, and that's also available. I think there's some implementation on CPU as well. For example, there's this library, ggml, where they implemented the same idea running on Mac and CPU. So I think that kind of broadly, the idea would apply. The current implementation ended up using NVIDIA's library or primitives, but I expect these ideas to be broadly applicable to different hardware. I think the main idea is you have asymmetry in memory hierarchy, which tends to be everywhere in a lot of accelerators. Yeah, it kind of reminds me of Sara Hooker's post, like the hardware lottery. There could be all these things that are much better, like architectures that are better, but they're not better on NVIDIA. So we're never going to know if they're actually improved. How does that play into some of the research that you all do too?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.