Skip to Main Content

OPEN CATALOG / BENCHMARKS + ENVIRONMENTS

Know what a benchmark measures before you run it.

Compare upstream projects and runnable Trunchbull ports in one open catalog. Inspect their sources, cases, tools, graders, and execution requirements before committing compute.

26

indexed projects

135

case previews

26 FOUND

PUBLIC INDEX / RUNNABLE + RESEARCHED

Evaluation projects

Runnable ports appear alongside upstream projects Trunchbull is evaluating. Every card states what is available today.

GSM8K

A friendly, practical test of whether an AI model can solve multi-step grade-school math word problems without losing track of the numbers.

No sandbox
Case previews3VersionGSM8K / TEST
OpenAIopenai/human-evalNot yet imported

HumanEval

Hand-written Python programming problems evaluated by executing generated code against tests.

Requires sandbox1 toolsModerate import
Case previews2VersionMASTER / 2021

MMLU

Multiple-choice questions across academic and professional subjects, from elementary mathematics to law and medicine.

No sandboxStraightforward import
Case previews2VersionTEST / HEAD

HellaSwag

Commonsense completion tasks that ask a model to choose the most plausible continuation of an everyday situation.

No sandboxStraightforward import
Case previews2VersionDATA / HEAD

TruthfulQA

Questions designed to reveal whether models repeat common misconceptions instead of giving accurate, well-supported answers.

No sandboxStraightforward import
Case previews2VersionMC / HEAD

RuneBench

Coding-agent tasks inside an emulated RuneScape world, scored by skill progression and resource collection over fixed time windows.

Requires sandbox2 toolsComplex import
Case previews2VersionUPSTREAM

Crafter

A lightweight visual survival environment that measures exploration, resource management, crafting, and long-horizon planning.

Requires sandbox2 toolsModerate import
Case previews2VersionUPSTREAM

BALROG

A shared evaluation framework for language and visual agents across long-horizon reinforcement-learning game environments.

Requires sandbox1 toolsComplex import
Case previews2VersionUPSTREAM

MiniHack

Short NetHack-derived tasks for navigation, exploration, combat, and object interaction under controlled episode budgets.

Requires sandbox2 toolsModerate import
Case previews2VersionUPSTREAM

MineDojo

Thousands of Minecraft tasks spanning survival, crafting, exploration, construction, and open-ended embodied play.

Requires sandbox2 toolsMajor integration
Case previews2VersionUPSTREAM

BRING YOUR OWN BENCHMARK

Make the evaluation contract as inspectable as the score.

Publish a pinned source, case set, tool manifest, grader, and runtime requirements as one immutable release.