AI article

Muon doesn't replace AdamW. Every Muon run still has AdamW in it.

Community description: Muon gets written up as the optimizer that finally beats AdamW. Reading Keller Jordan's original...

Dev.to | Sep 14, 2026 | Arun Kumar

Automated excerpt

Muon gets written up as the optimizer that finally beats AdamW. So your biases, your LayerNorm gains, your scalars — AdamW. AdamW exists because of one idea: decouple weight decay from the gradient update. The current best setup is Muon on 2D hidden layers plus AdamW on everything else — which is exactly what the original post recommends. Sources: Keller Jordan's Muon post, arXiv 2502.16982 (Moonshot AI), arXiv 1711.05101 (AdamW).

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news