AI article

Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8

Community description: Google ships its quantization-aware-trained Gemma 4 26B-A4B as GGUF and as a 48 GiB bf16 export, and as compressed-tensors W4A16 for every size but this one. A lossless repack to W4A16, a W4A16 mixture-of-experts method for vLLM's JAX path on TPU, and one v6e chip: 17.43 GiB of HBM, 53,888 KV tokens and 1,283 output tokens per second, against RedHat's FP8 build at 27.99 GiB, 3,456 tokens and 668. The same checkpoint loads unpatched on vLLM 0.30.0 on an NVIDIA L4.

Dev.to | Sep 26, 2026 | xbill

Automated excerpt

The QAT 26B serves on one v6e chip at 17. 43 GiB of HBM, with 53,888 tokens of KV cache and 1,283 output tokens per second. "Unquantized" QAT checkpoints, bf16 at 48. 07 GiB. On NVIDIA GPUs the same QAT checkpoint needs nothing beyond stock vLLM.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news