AI article

Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second

Community description: Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU. On one v5e chip the repacks serve every size from E2B to 26B, read the suite level with bf16 through 12B, score up to 2.4 points above Google's own 4-bit exports at the same speed, and put 12B on the chip at 11.31 GiB and 675 output tokens per second.

Dev.to | Oct 2, 2026 | xbill

Automated excerpt

On one v5e chip the repacked QAT builds serve every Gemma 4 size from E2B to 26B. Tool calling holds at every size: every E2B repack is within 1. 3 points of bf16. The E4B, 12B and 26B bf16 suite references ran on v6e, and E4B and 12B have no bf16 reference for GSM8K or BFCL.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news