AI article
On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow
Community description: On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still...
Dev.to | Oct 10, 2026 | Pingredsai
Automated excerpt
This part focuses on thread scheduling and CPU frequency scaling. 1. Where part 1 left off In part 1 I benchmarked Qwen2. 5-1. 5B on a Pixel 4 (Snapdragon 855) and measured 0. 5 tok/s — 60-120x off the theoretical limits for that SoC. None touched the little cores. n_threads=4 → 3 workers + 1 main thread.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.