Skip to Main Content
A Tool-Native Benchmarking Platform

Run any model against any Benchmark

Deploy tools from GitHub, run them against any OpenRouter model, and inspect every decision under controlled budgets.

BENCHMARK READOUT / RUN 0842 LIVE
ModelScoreDeltaState
Claude Opus 494.2base
GPT-591.8-2.4%
Gemini 2.5 Pro88.5-5.7%
2 REGRESSIONS DETECTED14 TOOLS / $1.84

REMOTE TOOLS

Isolated Workers

MODELS

OpenRouter / BYOK

LIMITS

Steps · tokens · cost

EVIDENCE

Full run traces

01

Tools as infrastructure

Turn a public repository into a model-ready tool.

Give us a GitHub URL. Trunchbull validates the contract, builds the repository, and deploys it into an isolated Cloudflare Worker—ready for any benchmark configuration.

  • Vercel AI SDK-compatible schema
  • Versioned, reproducible deployments
  • Remote execution with hard timeouts
TOOL DEPLOYMENTREADY

Repository URL

github.com/acme/research-toolsmain
Contract validated3 tools
Worker bundle built184 KB
Deploying to edge34 regions

03

TOOLS

34

REGIONS

ISO

RUNTIME

02

Isolated execution

Give agents a real machine without giving them yours.

Run terminal and code-execution evaluations inside controlled, image-backed environments. Trunchbull exposes files, commands, processes, and interactive terminals while keeping each workload inside an isolated runtime.

  • Container isolation with explicit resource limits
  • File, command, process, and terminal tools
  • Image-backed environments for repeatable runs
Explore sandboxing
SANDBOX ENVIRONMENT / LEASE 7F2AISOLATED

terminal-agent / ubuntu:24.04

IMAGE PINNED · PRIVATE RUNTIME

RUNNING
file_readREADY
command_executeREADY
process_listREADY
terminal_openREADY

POLICY

NETWORK

LIMITED

COMPUTE

SCOPED

ACCESS

03

Run control

Give every model a hard ceiling.

Define the maximum steps, token usage, and dollar spend before a run begins. Limits are enforced across model responses and tool calls—not checked after the damage is done.

  • BYOK through OpenRouter
  • Per-model pricing and token accounting
  • Deterministic stop reasons
RUN CONTROL / ACTIVEWITHIN LIMITS

CUMULATIVE USAGE / LIVE

6 of 12 steps complete

TOKENSSPEND
HARD LIMIT
00:00NOW / 14:32

PROJECTED SPEND

$3.72

HEADROOM

25.6%

STOP POLICY

HARD

STEPS

06 / 12

TOKENS

18.4K / 40K

SPEND

$1.84 / $5.00

01MODELPlanning tool sequence1.2s
02TOOLsearch_issues({ query: "timeout" })842ms
03TOOLread_file({ path: "src/client.ts" })186ms
04MODELSynthesizing final answer2.4s
ELAPSED 00:14:32STREAMING
04

Deep evidence

See exactly why one model wins.

A score tells you what happened. Traces show you why. Compare tool selection, arguments, latency, errors, token usage, and final outcomes step by step.

  • Side-by-side model traces
  • Tool-call and argument inspection
  • Regression-ready run history
MODEL COMPARISONCOMPLETE
METRICOPUS 4GPT-5
Task success94%87%
Tool accuracy91%72%
Median latency6482

FROM REPO TO RESULT

Your first useful result in three moves.

01

Connect a tool repository

Paste a public GitHub URL. We validate, build, and deploy it to an isolated Worker.

02

Configure the run

Choose models, prompts, tools, and hard limits for steps, tokens, and spend.

03

Inspect the evidence

Compare outcomes and replay every model response, tool call, error, and cost.

OPEN BETA PRICING

Start free. Upgrade when your runs need more room.

Publish and evaluate at no cost, then unlock managed sandboxes and higher monthly capacity with one fixed plan.

Free

$0

forever

  • 1,000 tool calls per month
  • Up to 5 published benchmarks
  • Up to 5 published tools
  • No sandbox provisioning
Start free
Full access

Open Beta

$59.99

per month

  • 250,000 tool calls per month
  • 16 sandbox hours per month
  • Standard-4 sandbox default
  • Unlimited published benchmarks
  • Unlimited published tools
Upgrade to Open Beta

Open Beta is a recurring monthly subscription that renews until canceled. See our Terms, Privacy Notice, and Refund Policy.

READY WHEN YOU ARE

Stop guessing which model is better.

Bring a tool, choose your models, and leave with evidence you can act on.