Skip to Main Content
TRUNCHBULLPublic benchmark sandbox
Try a free account
ARC-Challenge20 fixed cases

Pick the models. Watch the evidence arrive.

The model must reason through challenging grade-school science questions and select the correct multiple-choice answer.

Compare up to three available low-cost models. Every candidate receives the same cases, sampling controls, limits, and grader.

SOURCE AUTHOR Peter Clark et al./allenai/ai2_arc/ARC-Challenge train split · fixed twenty-case sampler

20

CASES

00

TEMPERATURE

256

MAX OUTPUT

EXACT

GRADER

RUN PREVIEW

THE RUN CONTRACT

One click creates a controlled model comparison.

Results stream into a case-by-model matrix. Open any cell to inspect the prompt, full response, extracted answer, grader decision, latency, tokens, and cost.

Benchmark

ARC-Challenge

Sample

20 test cases

Sampling

Temperature 0

Evaluation

Exact answer choice

SELECTED MODELS

Pick at least one model from the lineup.

Controlled evaluation

Same 20 prompts, model settings, token ceiling, and grader for every candidate. Review the original source. Review the Trunchbull ARC-Challenge port to see how Trunchbull adapted it.

This is the public sampler

The workspace adds your own benchmarks, repeat trials, tools, budgets, persisted runs, and full comparison reports.