AI article
Giving a coding agent more time barely helps
Community description: Real-SWE ran frontier models against licensed, private enterprise codebases (billing, tax,...
Dev.to | Sep 14, 2026 | Cole Halton
Automated excerpt
Real-SWE ran frontier models against licensed, private enterprise codebases (billing, tax, multi-service work) and one number jumped out at me: rollout duration barely moves resolution. 71.4% of rollouts that finished in under 10 minutes FAILED. 73.4% of rollouts that ran 10 minutes or longer also FAILED. Pass rate sits flat at 27-29% either way. The leader, Fable 5.1 on Claude Code, only lands 38.8% resolution. The top model still fails roughly six out of ten private enterprise tasks. Fable 5.1 is only 38.8% paired with Claude Code's scaffold.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.