AI article
When a Zero-Parameter Cache Overtakes a Transformer
Community description: A count table over the current document has no parameters and no training. Somewhere between 60 and 250 tokens of document it passes a 1.43M-parameter transformer, and by 1000 tokens it wins top-1 by 0.064, a 43% relative margin. Adding the transformer on top then buys 0.002.
Dev.to | Sep 13, 2026 | Seth Wheeler
Automated excerpt
Out of distribution, 1. 43M-parameter transformer, 64-token window. L=60 transformer 0. 152 cache 0. 103 delta -0. 050 CI [-0. 067,-0. 032] SIGNIFICANT (transformer) L=250 transformer 0. 143 cache 0. 172 delta +0. 029 CI [+0. 010,+0. 049] SIGNIFICANT (cache) L=1000 transformer 0. 149 cache 0. 213 delta +0. 064 CI [+0. 044,+0. 083] SIGNIFICANT (cache) The test reports both directions, because a confidence interval entirely below zero is a significant win for the transformer, and describing that as "not significant" would be wrong. The transformer goes 0. 152, 0. 143, 0. 149 across the three lengths.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.