AI article
Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys
Community description: A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.
Dev.to | Sep 17, 2026 | xbill
Automated excerpt
Output tokens per second, median of three repeats, worst-cell coefficient of variation 7. 96%: Counting prefill as well, the busiest cell moves 48,583 tokens a second — 64 streams at 1,024 context. Time per output token barely moves: 2. 85 ms at 128 context, 3. 04 at 1,024, 3. 9 at 8,192. Every row is the same checkpoint under vLLM, each from its own schema-valid report: Tokens per second per dollar-hour.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.