AI article

We ran HealthBench on our health AI's safety layer. It scored lower than the bare model.

Community description: A 150-conversation HealthBench subset, a safety layer that costs points, one real dose leak we found while measuring, and what moved the score.

Dev.to | Oct 6, 2026 | John Alexander

Automated excerpt

The full Tabibu pipeline: retrieval, input and output guardrails, the safety layer. The pipeline was 0. 158 below the bare model on the same conversations (paired difference, ±0. 039). On global_health (36 conversations) it scored 0. 230 against 0. 414.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

AI briefing: recent picks

More AI news