AI article
The models small enough to fit on a phone are the worst at understanding children
Community description: App Store rule 1.3 pushes a children's voice product on-device. The models that fit there get every second word wrong on children's speech, until you fine-tune them, at which point 39M parameters beats 1.5B. Figures from the published benchmarks.
Dev.to | Sep 16, 2026 | Max
Automated excerpt
The usual voice pipeline goes microphone, cloud recognition, cloud language model, cloud synthesis. After fine-tuning on children's speech, Whisper-tiny returns 2. 7 per cent word error on OGI instead of 40. 1 without it. Fine-tuning tiny took two hours on two GPUs.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.