01 / prediction
The current implementation ended up using NVIDIA's library or primitives, but I expect these ideas to be broadly applicable to different hardware.
“The current implementation ended up using NVIDIA's library or primitives, but I expect these ideas to be broadly applicable to different hardware.”
- Speaker
- Tri Dao
- Episode
- FlashAttention 2: making Transformers 800% faster w/o approximation - with Tri Dao of Together AI
- Publisher
- Latent Space