AI article
I Built a 188M Mixture-of-Experts LLM From Scratch on a Free GPU
Community description: By Shivam Kumar, founder of VisionQuantech. This is the honest version — what's proven, what's...
Dev.to | Oct 7, 2026 | Shivam Kumar
Automated excerpt
Then all four tested candidates ran real 300-step CPU experiments measuring loss decrease, dead experts, and routing balance. 188,269,568 total params, 51,430,400 active per token — the compute of a ~51M dense model, ~3. 7× its capacity Token-choice top-6 routing (renormalized), expert-level aux loss (α=0. 01) + router z-loss (1e-3), dropless, fp16 Why not 9 domain experts — one per field? The Mixtral-style coarse design was worst — most active compute for the coarsest routing. The winner balanced best — aux loss 2. 03, zero dead experts.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.