AI article
Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't
Community description: Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU, on one v5e chip. What runs: every size from E2B to 26B, level with bf16 through 12B, up to 2.4 points above Google's own 4-bit exports at the same speed, and 12B at 675 output tokens per second. What doesn't: bf16 above E2B, 31B in any build, and 26B past a 2,176-token context.
Dev.to | Oct 2, 2026 | xbill
Automated excerpt
On one v5e chip the repacked QAT builds serve every Gemma 4 size from E2B to 26B. Gemma 4 E4B at bf16 is 14. 9 GiB and 12B is 22. 4 GiB, so everything above E2B needs 4-bit or 8-bit weights on this chip. Tool calling holds at every size: every E2B repack is within 1. 3 points of bf16.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.