AI article
Gemma 4 in Pure JAX: What Changes Between Turing and Ada, and What Doesn't
Community description: One hand-written Gemma 4 port, no PyTorch and no vLLM, on two NVIDIA GPUs a generation apart. Most of it ports untouched. Two things do not, and one of them was quietly eating 87% of decode.
Dev.to | Aug 31, 2026 | xbill
Automated excerpt
One port, one build, one checkpoint, two cards. NVIDIA T4G — Turing, SM 7. 5, 15,360 MiB NVIDIA L4 — Ada, SM 8. 9, 23,034 MiB tpu_jax_weight_bytes reads 6,155,450,950 on both cards — the same integer. The same profile on a different instance, a different AMI and a restored cache landed at 1466. 0 ms against 1467. 1 ms. 1. 1 ms The Ada card resolves it.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.