Tech article

Surpassing vLLM with a Generated Inference Stack

Community description: 7 points · 0 comments on Hacker News

Hacker News | Mar 10, 2026 | lukebechtel

Automated excerpt

The resulting engine delivers up to 34. 3% more tokens per second than vLLM when configured with identical parameters. Qwen3-8B · H100 80GB · FP8ISL=8192 · OSL=1024 · BS=88identical parameters to vLLM 0. 13. 0 *Decode-heavy (ISL=1k, OSL=8k)+34. 3%vs vLLM · 6,712 tok/sPrefill-Heavy (ISL=8k, OSL=1k)+15. 9%vs vLLM · 22,470 tok/sInfy Optimization Trajectory · Qwen3-8B, 111 Iterations on H100Prefill-Heavy Workload (ISL=8k, OSL=1k)Prefill-heavy workload (ISL=8k, OSL=1k) · +15. 9% vs vLLMvLLM FP8 baselineOn decode-heavy workloads (ISL=1k, OSL=8k) (where inference serving is most throughput-constrained), infy reaches 6,712 tok/s, +34. 3% above vLLM FP8 on identical hardware and parameters.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More tech news