AI article
We Asked 330 Models a Question in Korean. Half of Them Answered in the Wrong Alphabet.
Community description: A mechanical script-purity check auto-failed 766 of 2,304 answers before a judge ever saw them. 54 models never produced a single clean Korean answer. Here is the exact rule, and why your eval harness needs one.
Dev.to | Sep 17, 2026 | ai maya
Automated excerpt
We graded 330 language models on Korean across seven axes. If the judge cannot The judge never sees the model name. A judge asked to grade an answer that is in the wrong script will still return a grade.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.