AI article

I Built a 188M Mixture-of-Experts LLM From Scratch on a Free GPU

Community description: By Shivam Kumar, founder of VisionQuantech. This is the honest version — what's proven, what's...

Dev.to | Oct 7, 2026 | Shivam Kumar

Automated excerpt

Then all four tested candidates ran real 300-step CPU experiments measuring loss decrease, dead experts, and routing balance. 188,269,568 total params, 51,430,400 active per token — the compute of a ~51M dense model, ~3. 7× its capacity Token-choice top-6 routing (renormalized), expert-level aux loss (α=0. 01) + router z-loss (1e-3), dropless, fp16 Why not 9 domain experts — one per field? The Mixtral-style coarse design was worst — most active compute for the coarsest routing. The winner balanced best — aux loss 2. 03, zero dead experts.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

AI briefing: recent picks

More AI news