AI article
The FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"
Community description: I had one hour on an AMD MI300X and one question: what does a token actually cost on it? One GPU,...
Dev.to | Oct 6, 2026 | Throttle
Automated excerpt
At temperature 0, with the same 32 prompts, BF16 32B wrote about 6,900 output tokens per block. Output length matched BF16 within about 1% (72B: 6,361 vs 6,328 tokens per block). One session, one workload: short prompts, 256 output tokens, concurrency 32.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.