AI article
VLLM_ROCM_USE_AITER=1 Slows Gemma 4 on an AMD MI300X: What the Flag Changes and Why
Community description: AMD's AITER kernel library is the usual first switch for vLLM speed on an Instinct MI300X. On Gemma 4 12B fp8 it made every one of nine cells 0.6% to 2.6% slower. The boot logs show why: Gemma 4's 512-wide attention heads keep attention on Triton, and AITER's fp8 matrix multiply has no tuned settings for any of the model's six weight shapes.
Dev.to | Oct 10, 2026 | xbill
Automated excerpt
VLLM_ROCM_USE_AITER=1 made Gemma 4 12B slower in every cell measured, by 0. 6% to 2. 6%, against the same build served the same day on the same droplet and image. Attention, where AITER usually helps most, never leaves Triton, because Gemma 4's full-attention layers use 512-wide heads. For Gemma 4 on this image, leave VLLM_ROCM_USE_AITER off.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.