AI article
Confidence is theater: benchmarking nine local VLM pipelines on handwritten clinical forms
Community description: Nine configurations of local VLM pipeline benchmarked on handwritten clinical forms — including the experiments that failed.
Dev.to | Sep 15, 2026 | Stephen Ohakanu
Automated excerpt
Nine configurations, same 23 pages, same 221-value gold set. The production targets. ≥98% auto-accepted accuracy and ≤2% escaped hallucination, held-out only. The bench methodology — pinned model revisions, trap-carrying gold sets, invented-vs-escaped hallucination accounting — is documented alongside the pipeline.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.