AI article
Gemma 4 From E2B to 31B on an AMD MI300X: fp8 Overtakes bf16 From 12B Up
Community description: Every Gemma 4 size from E2B to 31B served on one AMD Instinct MI300X in bf16, fp8, int8 W8A8 and int4 W4A16, on the same vLLM image. fp8 goes from 0.75x bf16 for one request at E2B to 1.23x at 31B and is faster than bf16 in every cell at 12B and 31B. The mixture-of-experts 26B sits at bf16's pace, and the 4-bit builds run at 0.14x to 0.69x at every size.
Dev.to | Oct 9, 2026 | xbill
Automated excerpt
From 12B up, fp8 is faster than bf16 in every cell: 1. 12x to 1. 41x at 12B and 1. 20x to 1. 45x at 31B. The 4-bit builds stay at 0. 14x to 0. 69x of bf16 at every size. For 12B and 31B, serve fp8: it is faster than bf16 in every cell and halves the weights.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.