AI article
I ran six coding agents on seven local models, 30 times each
Community description: Last week I posted a small benchmark on whether coding agents still work when the model you run...
Dev.to | Oct 1, 2026 | Giuseppe Sirigu
Automated excerpt
It used three runs per task, a handful of models, and a harness I kept private. On Qwen3. 8-27B every agent works reliably. Five trials of each task, so 30 runs per cell, 600 seconds per run.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.