AI article

Gemma 4 E2B in Pure JAX on a Colab TPU: Google's 4-Bit Export Against an Exact Repack

Community description: A Colab notebook for the AI GDE Marathon that loads three Gemma 4 E2B checkpoints into a pure-JAX engine on one TPU v5e chip and measures, on the reader's own chip, how far each 4-bit build sits from the weights Google trained. The repack holds the trained grid and comes out 342.6x closer to the QAT model at the same speed and a smaller download.

Dev.to | Oct 10, 2026 | xbill

Automated excerpt

Its export rounds every weight a second time, onto a grid the model never trained on. The repack writes the QAT model's story token for token. The bf16 QAT model decodes faster than either 4-bit build in this engine, because the engine unpacks every 4-bit weight at every step.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

Read next

AI briefing: recent picks

More stories to explore