AI article
Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU
Community description: Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.
Dev.to | Sep 24, 2026 | xbill
Automated excerpt
One v6e chip serves Gemma 4 E2B, E4B and 12B at bf16 and a 26B-A4B fp8 build; no 31B checkpoint loads. On demand, the TPU costs more per decision than the L4 or Jev; 12B is the size to pick. Choose the L4, or Jev itself, when cost per decision matters: on demand, 12B on one v6e chip costs $8. 19 per million decisions, against Jev's $5. 54 and at most $5. 43 for the 26B on an L4.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.