AI article

VLLM_ROCM_USE_AITER=1 Slows Gemma 4 on an AMD MI300X: What the Flag Changes and Why

Community description: AMD's AITER kernel library is the usual first switch for vLLM speed on an Instinct MI300X. On Gemma 4 12B fp8 it made every one of nine cells 0.6% to 2.6% slower. The boot logs show why: Gemma 4's 512-wide attention heads keep attention on Triton, and AITER's fp8 matrix multiply has no tuned settings for any of the model's six weight shapes.

Dev.to | Oct 10, 2026 | xbill

Automated excerpt

VLLM_ROCM_USE_AITER=1 made Gemma 4 12B slower in every cell measured, by 0. 6% to 2. 6%, against the same build served the same day on the same droplet and image. Attention, where AITER usually helps most, never leaves Triton, because Gemma 4's full-attention layers use 512-wide heads. For Gemma 4 on this image, leave VLLM_ROCM_USE_AITER off.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

Read next

AI briefing: recent picks

More stories to explore