GSM8K
A friendly, practical test of whether an AI model can solve multi-step grade-school math word problems without losing track of the numbers.
OPEN CATALOG / BENCHMARKS + ENVIRONMENTS
Compare upstream projects and runnable Trunchbull ports in one open catalog. Inspect their sources, cases, tools, graders, and execution requirements before committing compute.
26
indexed projects
135
case previews
PUBLIC INDEX / RUNNABLE + RESEARCHED
Runnable ports appear alongside upstream projects Trunchbull is evaluating. Every card states what is available today.
A friendly, practical test of whether an AI model can solve multi-step grade-school math word problems without losing track of the numbers.
A playful, focused test of whether AI understands skateboard tricks and can follow a precise product recommendation task.
A clear, no-frills test of science reasoning on challenging grade-school questions with multiple-choice answers.
A safety-first benchmark for finding dangerous medical AI behavior: missed escalation, unsafe medication advice, false reassurance, and overconfident claims.
A multiple-choice benchmark that checks whether AI models choose truthful answers instead of repeating popular misconceptions.
Executable evaluation of function selection, arguments, parallel calls, and multi-turn tool use.
Evaluates whether AI agents can autonomously complete 89 hard, realistic, end-to-end tasks in containerized terminal environments, using task-specific verifiers.
Hand-written Python programming problems evaluated by executing generated code against tests.
A human-validated set of real GitHub issues requiring repository patches that pass tests.
Verifiable instruction-following tasks that test whether a model can satisfy precise format, length, language, and content constraints.
Multiple-choice questions across academic and professional subjects, from elementary mathematics to law and medicine.
Commonsense completion tasks that ask a model to choose the most plausible continuation of an everyday situation.
Naturally occurring yes-or-no questions answered from short passages, designed to test inference beyond keyword overlap.
Two-choice physical commonsense problems about how everyday goals can be completed in the real world.
Large-scale pronoun-resolution problems that require commonsense reasoning to resolve an ambiguous reference.
Questions designed to reveal whether models repeat common misconceptions instead of giving accurate, well-supported answers.
Coding-agent tasks inside an emulated RuneScape world, scored by skill progression and resource collection over fixed time windows.
A lightweight visual survival environment that measures exploration, resource management, crafting, and long-horizon planning.
A shared evaluation framework for language and visual agents across long-horizon reinforcement-learning game environments.
Generated text-adventure environments for planning, memory, exploration, and language grounding.
Visual computer-use evaluation across Game Boy, DOS, and browser games, with a paused Lite mode for latency-fair comparison.
Short NetHack-derived tasks for navigation, exploration, combat, and object interaction under controlled episode budgets.
A standardized interface to a procedurally generated, partially observed, system-rich, and extremely long-horizon game.
Thousands of Minecraft tasks spanning survival, crafting, exploration, construction, and open-ended embodied play.
Programmatic Doom scenarios for visual navigation, perception, planning, and combat.
Competitive turn-based battles with structured state, legal-action constraints, and objective match outcomes.
BRING YOUR OWN BENCHMARK
Publish a pinned source, case set, tool manifest, grader, and runtime requirements as one immutable release.