FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Details
FlashAttention is one of the most important pieces of software infrastructure for both training and serving any relevant AI model on cutting-edge GPUs. Let's have a look under the hood.
The paper comes with a blog post that may be easier to read: https://tridao.me/blog/2024/flash3/
Find the paper here: https://arxiv.org/abs/2407.08608
We are trying out a new Location -- right at the Westbahnhof IKEA. Take the elevator to the fifth. We may relocate if it is too crowded. Please Text me if you are late and can't find us.
---
Bring your own printout or pdf-on-a-gadget. If you didn't have the time to read the whole paper, it is sufficient to be able to pretend to have sort of skimmed it.
---
Want to suggest a paper, an event location, or just stay in touch with regular members? We hang out at the following Matrix server: https://matrix.to/#/%23pwl-wien:rend.al
