Skip to Main Content
TRUNCHBULLPublic benchmark sandbox
Try a free account
TruthfulQA MC120 fixed cases

Pick the models. Watch the evidence arrive.

The model must select the single truthful answer while resisting familiar misconceptions and persuasive falsehoods.

Compare up to three available low-cost models. Every candidate receives the same cases, sampling controls, limits, and grader.

SOURCE AUTHOR Stephanie Lin, Jacob Hilton, and Owain Evans/sylinrl/TruthfulQA/TruthfulQA MC1 · fixed twenty-case sampler

20

CASES

00

TEMPERATURE

256

MAX OUTPUT

EXACT

GRADER

RUN PREVIEW

THE RUN CONTRACT

One click creates a controlled model comparison.

Results stream into a case-by-model matrix. Open any cell to inspect the prompt, full response, extracted answer, grader decision, latency, tokens, and cost.

Benchmark

TruthfulQA MC1

Sample

20 test cases

Sampling

Temperature 0

Evaluation

Exact MC1 answer choice

SELECTED MODELS

Pick at least one model from the lineup.

Controlled evaluation

Same 20 prompts, model settings, token ceiling, and grader for every candidate. Review the original source. Review the Trunchbull TruthfulQA MC1 port to see how Trunchbull adapted it.

This is the public sampler

The workspace adds your own benchmarks, repeat trials, tools, budgets, persisted runs, and full comparison reports.