AI article
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
Community description: One H100 NVL. A 421M-parameter decision model. 15.1 million decisions per day while staying inside a...
Dev.to | Sep 24, 2026 | Bhushan Kinge
Automated excerpt
A 421M-parameter decision model. 15. 1 million decisions per day while staying inside a p99 ≤ 130 ms latency budget. On the RTX PRO 6000 Blackwell, eager FP16 and TensorRT both reached 146 decisions/s. Jev latency includes the public internet; self-hosted Laya latency does not.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.