AI article
The Last Non-Neural Candidate, and It Did Not Clear the Bar
Community description: The bar was a slope: keep converting extra data into accuracy after exact-context statistics saturate. A full hierarchical Pitman-Yor model with Gibbs sweeps and inferred discounts moved the intercept and left the slope alone, halving with every doubling exactly as cruder count models did.
Dev.to | Sep 12, 2026 | Seth Wheeler
Automated excerpt
Measured as top-1 gain per doubling of the training corpus at 4M to 8M tokens, that came out as Witten-Bell +0. 005, modified Kneser-Ney +0. 006, single-pass Pitman-Yor +0. 008, against a small transformer's +0. 021. The full model holds a steady lead over modified Kneser-Ney, +0. 005 then +0. 009 then +0. 010, and decays at the same rate. The full model beats it on every axis at every size: at 2M, top-1 0. 326 against 0. 316, perplexity 36. 8 against 48. 5, OOD 0. 153 against 0. 147.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.