AI article
We taped over the total. Vision models filled it in anyway.
Community description: This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Here is a...
Dev.to | Oct 11, 2026 | Aina Zulfiqar
Automated excerpt
Half a total is worse than no total Preregistered H1 (≥ 40% made-up totals on fully taped, non-computable bills, at median confidence ≥ 70) was supported: 44% [35–52], median confidence 85. "Confidently" depends on the model Median stated confidence on made-up totals (total fully taped): Gemini 3 Flash 95, Gemini 3. 1 Flash-Lite 95, GPT-5. 4 nano 62, Claude Sonnet 5 20. Check first, then answer (probe, then extract, in the same chat) keeps made-up totals at 0–2% for every model, and Gemini 3 Flash computes recoverable totals again (98%).
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.