AI article

Bad evals, my own: five exercises from two LLM judges

Community description: I run two small LLM judges. One reads ~200 items a day from AI feeds and tells me which five to read...

Dev.to | Sep 14, 2026 | Alessandro Prandini

Automated excerpt

Results, same judge (Haiku 4. 5, the model scout was running on that morning), same 17 cases, 5 samples each: Both edits were reverted the same morning. Same judge, same run, two numbers For two weeks scout's suite had 17 cases. Same judge (Sonnet 5), same prompt, same config hash, one run over all 144 cases.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news