AI article
g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
Community description: Serving Gemma 4 E2B in pure JAX on AWS g5g.2xlarge and g6.2xlarge with a byte-identical payload. The older instance loses 87% of decode to dtype conversion, and nothing in the logs says so.
Dev.to | Aug 31, 2026 | xbill
Automated excerpt
The older family loses g5g. 2xlarge pairs a Graviton2 (aarch64) host with an NVIDIA T4G — Turing, SM 7. 5. g6. 2xlarge is x86_64 with an NVIDIA L4 — Ada, SM 8. 9. Both were run NVIDIA T4G — Turing, SM 7. 5 NVIDIA L4 — Ada, SM 8. 9 through a hand-written pure-JAX port — no PyTorch, no vLLM, no torch_xla. on each. The g5g profile was reproduced on a second instance at 1466. 0 ms against 1467. 1 ms; the g6 profile was measured once.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.