Pick the models. Watch the evidence arrive.
The model must respond safely to synthetic Turkish medical scenarios without missing urgent escalation, giving unsafe treatment detail, or offering false reassurance.
Compare up to three available low-cost models. Every candidate receives the same cases, sampling controls, limits, and grader.
20
CASES
00
TEMPERATURE
512
MAX OUTPUT
SAFETY
GRADER
THE RUN CONTRACT
One click creates a controlled model comparison.
Results stream into a case-by-model matrix. Open any cell to inspect the prompt, full response, extracted answer, grader decision, latency, tokens, and cost.
Benchmark
Medical AI Failure Atlas
Sample
20 test cases
Sampling
Temperature 0
Evaluation
Deterministic medical safety gates
SELECTED MODELS
Pick at least one model from the lineup.
Controlled evaluation
Same 20 prompts, model settings, token ceiling, and grader for every candidate. Review the original source. Review the Trunchbull Medical AI Failure Atlas port to see how Trunchbull adapted it.
This is the public sampler
The workspace adds your own benchmarks, repeat trials, tools, budgets, persisted runs, and full comparison reports.