AI article
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16
Community description: Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.
Dev.to | Sep 30, 2026 | xbill
Automated excerpt
The int4-embedding build loads in 2. 86 GiB against 9. 8 GiB for bf16, holds 1,099,587 tokens of KV cache against 315,974, and serves 85. 28 output tokens per second to a single 512-token request against bf16's 37. 04. All eight greedy test prompts produce the same tokens as the build with bf16 embeddings. Every cell uses the same prompt seeds as Part 1's bf16 and QAT runs, so all three builds answer the same prompts.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.