AI article

We Tried ISO-AdamW. AdamW Kept Its Job.

Community description: A small GRPO experiment: ISO-AdamW scored 75.8% versus AdamW's 75.4% on GSM8K, with 47% more GPU memory. The idea, the 44x kernel detour, and what four answers mean.

Dev.to | Oct 2, 2026 | Aleksei Romanov

Automated excerpt

AdamW changes both; ISO-AdamW keeps the stretch factors of the base model and only turns the frames. After each optimizer step, an orthogonalization step projects the frames back onto the manifold of orthogonal matrices. ISO-AdamW needs 47. 5% more peak GPU memory.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news