AI article
When LLM judges agree, should we believe them?
No preview is available. Read the original article for the full story.
Hacker News | Sep 14, 2026 | Betelbuddy
Automated excerpt
To reduce noise, you ask several judge models to evaluate the same passage. Correlation between different judges' outputs limits the utility of multijudge panels. In the LLM-as-a-judge context, the aggregator learns both judge skill and judge similarity. Majority vote counts votes; weighted vote learns per-judge reliability; dependence-aware aggregation also learns relationships among judges. The judge panel contained 10 judge models, all run at temperature zero — meaning there’s no randomness in their outputs, so the same input will always elicit the same output.We compared the dependence-aware models with two conditional-independence baselines: weighted majority vote and uniform majority vote.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.