Tech article
Speeding up gearhash on ARM64 (2× faster)
No preview is available. Read the original article for the full story.
Hacker News | Sep 15, 2026 | pranitha_m
Automated excerpt
To win on NEON, the dependency chain itself has to get shorter. provided you precompute G = (g₀ << 1) + g₁. Result: 0. 92× → 1. 13×, better, but still well short of the expected 2×. add. 2d v2, v1, v1 ; h << 1 add. 2d v2, v3, v2 ; h₁ = (h<<1) + g₀ shl. 2d v1, v1, #2 ; h << 2 add. 2d v3, v3, v3 ; g₀ << 1 add. 2d v1, v1, v4 ; (h<<2) + g₁ <-- on the h chain add. 2d v1, v3, v1 ; ...
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.