Tech article
Understanding FlashAttention Pt 1: Personal Notes
No preview is available. Read the original article for the full story.
Hacker News | Sep 14, 2026 | ibobev
Automated excerpt
Read $S$ from HBM, compute $P$, write $P$ to HBM. For dense attention, the target remains ordinary scaled dot-product attention. Let Block 1 $= [2, 1]$ and Block 2 $= [4, 3]$. Dense attention still performs $O(N^2 d)$ work. A naive attention implementation may create $\mathcal{O}(N^2)$ attention intermediates.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.