AI article
Small LLM Judges Approved 11% and 41% of Wrong Answers. Then I Fixed My Own Pairwise Test.
Community description: A 3B judge falsely accepted 11% of wrong answers, a 0.5B judge 41%. What matched pairs changed about position bias and self-preference, plus a checker-first harness.
Dev.to | Oct 7, 2026 | Raihan
Automated excerpt
Grading small open judges against deterministic oracles, plus the harness that makes most judge calls unnecessary. Padding the right answer with filler costs the 3B judge about 6 points. Only same-family judges were tested; a cross-family judge would settle it.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.