Skip to Main Content
TRUNCHBULLPublic benchmark sandbox
Try a free account
Medical AI Failure Atlas20 fixed cases

Pick the models. Watch the evidence arrive.

The model must respond safely to synthetic Turkish medical scenarios without missing urgent escalation, giving unsafe treatment detail, or offering false reassurance.

Compare up to three available low-cost models. Every candidate receives the same cases, sampling controls, limits, and grader.

SOURCE AUTHOR Goktug Ozkan, MD/goktugozkanmd/medical-ai-failure-atlas/v0.2.1 public release · fixed twenty-case sampler

20

CASES

00

TEMPERATURE

512

MAX OUTPUT

SAFETY

GRADER

RUN PREVIEW

THE RUN CONTRACT

One click creates a controlled model comparison.

Results stream into a case-by-model matrix. Open any cell to inspect the prompt, full response, extracted answer, grader decision, latency, tokens, and cost.

Benchmark

Medical AI Failure Atlas

Sample

20 test cases

Sampling

Temperature 0

Evaluation

Deterministic medical safety gates

SELECTED MODELS

Pick at least one model from the lineup.

Controlled evaluation

Same 20 prompts, model settings, token ceiling, and grader for every candidate. Review the original source. Review the Trunchbull Medical AI Failure Atlas port to see how Trunchbull adapted it.

This is the public sampler

The workspace adds your own benchmarks, repeat trials, tools, budgets, persisted runs, and full comparison reports.