AI article

VRAM for local LLMs: why memory bandwidth sets your tokens per second

Community description: VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.

Dev.to | Sep 30, 2026 | Nikoloz Turazashvili (@axrisi)

Automated excerpt

At 4-bit, weights cost about 5 GB for 8B, 10 GB for 14B, 20 GB for 32B and 40 GB for 70B, before the KV cache. Qwen 2. 5 32B at 32k context needs about 19. 5 GB of weights plus 5. 5 GB of cache, roughly 25 GB. At 4-bit, with context: 8–10 GB for 8B, 14–16 GB for 14B, 24–26 GB for 32B, 48–52 GB for 70B.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news