AI article

LLM Benchmarks: The API Scores Better Than the App You Actually Use

Community description: A Stanford paper finds API benchmark scores run 3.4 points higher than the same models in their chat apps. What that means for how I evaluate LLMs for

Dev.to | Sep 13, 2026 | Qasim Parray

Automated excerpt

I was running the same model through the API in a harness with a tidy system prompt and getting clean labels. Test-retest agreement, meaning "ask the same thing twice, get the same answer", was 2. 1 points higher through the API. The difference between API access and interface access was bigger than the difference between GPT 5. 3 and GPT 5. 4 measured through the API alone.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news