Skip to Main Content
TRUNCHBULLPublic benchmark sandbox
Try a free account
GSM8K20 fixed cases

Pick the models. Watch the evidence arrive.

The model must solve multi-step grade-school math word problems and finish with the correct numeric answer.

Compare up to three available low-cost models. Every candidate receives the same cases, sampling controls, limits, and grader.

SOURCE AUTHOR OpenAI/openai/grade-school-math/test split · fixed twenty-case sampler

20

CASES

00

TEMPERATURE

256

MAX OUTPUT

EXACT

GRADER

RUN PREVIEW

THE RUN CONTRACT

One click creates a controlled model comparison.

Results stream into a case-by-model matrix. Open any cell to inspect the prompt, full response, extracted answer, grader decision, latency, tokens, and cost.

Benchmark

GSM8K

Sample

20 test cases

Sampling

Temperature 0

Evaluation

Exact numeric answer

SELECTED MODELS

Pick at least one model from the lineup.

Controlled evaluation

Same 20 prompts, model settings, token ceiling, and grader for every candidate. Review the original source.

This is the public sampler

The workspace adds your own benchmarks, repeat trials, tools, budgets, persisted runs, and full comparison reports.