AI article
Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens
Community description: A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.
Dev.to | Sep 28, 2026 | xbill
Automated excerpt
With the count, every model but one answered every question about 330 ids correctly. The answer is 8, and every model here gets it. With the Count in the Tool, Every Model Is Right at a Flat Cost Every model but one answered all 21 engine questions about 330 ids correctly, at 35 to 451 output tokens per question, and each model's tokens stayed about the same from 11 ids to 330.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.