AI article

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

Community description: Mixture-of-experts models like Qwen3.6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer...

Dev.to | Sep 29, 2026 | Donald Lee

Automated excerpt

Mixture-of-experts models like Qwen3. 6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer GPUs, but they only read a few experts per token. llama. cpp's --n-cpu-moe N flag exploits that: it keeps the expert FFN tensors of the first N layers in system RAM and runs them on the CPU, while attention, shared weights and the KV cache stay on the GPU. The FP16 KV cache is therefore which is only 0. 625 GiB at 32K context and 2. 5 GiB at 128K.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news