AI article
The Same Model Can Cost 14x More Depending on Who Serves It
Community description: We pulled per-provider pricing for every model on OpenRouter and measured first-token latency ourselves. Across all 182 models with two or more paying providers, the median spread is 1.87x and the widest is 14.47x — for identical weights.
Dev.to | Sep 16, 2026 | ai maya
Automated excerpt
A provider serving fp4 is not serving the same thing as one serving bf16, even though the model id is identical. The model writes fluent, natural Korean and gets the honorific wrong. A 2023 model, gpt-3. 5-turbo-16k, scores a perfect 3. 00, above most 2026 flagships.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.