AI article
Just Train More: Measuring the Exchange Rate
Community description: The obvious objection to a zero-parameter cache beating a small transformer is that the transformer is undertrained. Sixteen times the training data moved the crossover 6.6x, which is document length times 1.60 per doubling: about as much as letting the cache read the rest of the file it was already sitting in.
Dev.to | Sep 14, 2026 | Seth Wheeler
Automated excerpt
The cache uses no training data at all, which makes its accuracy a constant across the sweep: whatever the transformer gains from 16x more data is exactly what moves the crossover. Sixteen times the training data moved the crossover 6. 6x, which is document length times 1. 60 per doubling of training data. Training data lifts the transformer's flat line.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.