AI article

Gemma 4 E2B on an AMD MI300X: Which Weight Format Should You Serve?

Community description: Ten builds of Gemma 4 E2B (bf16, fp8 in both E4M3 flavours, int8 W8A8 and int4 W4A16, each with and without int4 embedding tables) served one after another on one AMD Instinct MI300X with the same vLLM image. fp8 is the only format that keeps pace with bf16, the 4-bit builds run at 0.14x to 0.63x, and int4 embedding tables cost almost nothing.

Dev.to | Oct 8, 2026 | xbill

Automated excerpt

Every quantized build starts from Google's gemma-4-E2B-it-qat-q4_0-unquantized release, whose weights already sit on a 4-bit grid with one scale per group of 32 values. Each build reads Google's QAT checkpoint and writes one format. The int8 and 4-bit builds are slower than bf16 in every cell.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

Read next

AI briefing: recent picks

More AI news