AI article

Ollama says my model does 13,826 tokens/sec. It does 43.

Community description: That number is not a typo, and my GPU has not improved. Both figures came out of the same daemon,...

Dev.to | Sep 15, 2026 | Ivan Stankovic

Automated excerpt

Send the same prompt twice and print three fields: prompt = "You are a helpful assistant. A fully cached prompt has no prefill rate. Nobody alternates between two prompt layouts.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news