AI article

Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

Community description: Plain Gemma 4 26B read by its label probabilities against DiffusionGemma's one-step read, both as community 4-bit (AWQ) builds on one EC2 L4, on 1,200 labelled examples and on the 3,880-record public suite where Jev 1.13.0 has published results. Pre-registered, with accuracy, calibration, calibration after 0 to 150 labels, latency and cost.

Dev.to | Sep 24, 2026 | xbill

Automated excerpt

On Bespoke Labs' 3,880-record public suite, plain Gemma 4 26B trails Jev 1. 13. 0 by 2. 1 points overall, shows no measurable difference on yes/no questions, and trails by 4. 5 points on multiple choice. Plain Gemma read this way had no published accuracy or calibration result. A third run puts both, and Gemma 4 E4B, on the public suite where Bespoke Labs has published results for Jev, so the Gemma numbers sit beside Jev's on identical records.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news