AI article
LLM Benchmarks: The API Scores Better Than the App You Actually Use
Community description: A Stanford paper finds API benchmark scores run 3.4 points higher than the same models in their chat apps. What that means for how I evaluate LLMs for
Dev.to | Sep 13, 2026 | Qasim Parray
Automated excerpt
I was running the same model through the API in a harness with a tidy system prompt and getting clean labels. Test-retest agreement, meaning "ask the same thing twice, get the same answer", was 2. 1 points higher through the API. The difference between API access and interface access was bigger than the difference between GPT 5. 3 and GPT 5. 4 measured through the API alone.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.