Tech article

Understanding FlashAttention Pt 1: Personal Notes

No preview is available. Read the original article for the full story.

Hacker News | Sep 14, 2026 | ibobev

Automated excerpt

Read $S$ from HBM, compute $P$, write $P$ to HBM. For dense attention, the target remains ordinary scaled dot-product attention. Let Block 1 $= [2, 1]$ and Block 2 $= [4, 3]$. Dense attention still performs $O(N^2 d)$ work. A naive attention implementation may create $\mathcal{O}(N^2)$ attention intermediates.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More tech news