AI article

Two "Codex CLI" models on the same benchmark: the harness hides the model

Community description: Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great...

Dev.to | Sep 14, 2026 | Cole Halton

Automated excerpt

Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution Same vendor's CLI, same harness, same benchmark. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Conventions, context carry-over, tool loop, judge.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news