AI article
Multi-Reward RL, Part 2: Benchmarking GRPO, DAPO, and CISPO on Unseen Tasks
Community description: Qwen3-14B DEX experiments compare GRPO, DAPO, GDPO, CISPO, ADAPO and REPO-R with and without thinking: full holdout results, reward curves, time and memory.
Dev.to | Oct 6, 2026 | Aleksei Romanov
Automated excerpt
Without thinking every trainer beats the untrained model; with thinking only DAPO-refill does. Thinking mode: training reward over the last ten points against holdout score. The chart below shows three of them in thinking mode: Raw training reward in thinking mode for CISPO, DAPO and DAPO-refill.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.