AI article
The more aggressive matmul kernel lost to the register budget
Community description: The WebGPU matmul sweep started with a naive kernel, then added 16 by 16 workgroup tiling and a 4 by...
Dev.to | Sep 13, 2026 | Sarthak Agrawal
Automated excerpt
The WebGPU matmul sweep started with a naive kernel, then added 16 by 16 workgroup tiling and a 4 by 4 output block per thread. At a 2048 cubed matrix size, the measured time moved from 47.24 ms for the naive kernel to 17.23 ms for tiling and 9.12 ms for the 4 by 4 blocked kernel. The blocked version was 5.18 times faster than naive at that size. The obvious next idea was an 8 by 8 block. At 2048 cubed, the 4 by 4 version took 10.15 ms while the 8 by 8 version took 11.52 ms.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.