Skip to Main Content
Trunchbull

Model evaluation / real workload evidence

Test models on the work that matters.

Replace leaderboard assumptions with evidence from your own tasks. Run competing models under the same conditions and see which one is ready for your product.

Explore the platform

MODEL EVALUATION / RUN 0842

support-resolution-v3 · 5 repeats

COMPLETE
ModelScorePass
Claude Sonnet 4
94.296%
GPT-5
91.892%
Gemini 2.5 Pro
88.587%

240

EXECUTIONS

99.4%

TOOL SUCCESS

$4.82

TOTAL SPEND

MODEL LINEUP

Up to 5

compared in one controlled run

RUN CONTROL

Hard limits

for steps, tokens, and total spend

EVIDENCE

Full trace

for every response and tool call

01 / Workload evidence

Test the work your models will actually perform.

Generic leaderboards cannot tell you which model will succeed inside your product. Trunchbull runs every candidate against the same versioned tasks, tools, and pass conditions so the result reflects your workload.

  • Your prompts, files, tools, and expected outcomes
  • The same evaluation contract for every model
  • Versioned definitions for reproducible results
EVALUATION CONTRACTv3.0

Resolve a customer question using the approved knowledge source.

Required toolsearch_knowledge
Pass conditioncontains source link
Maximum steps8
Per-model budget$0.10
SAME CONTRACT / EVERY MODELREADY

02 / Measured tradeoffs

See quality, reliability, speed, and cost together.

A higher score is not always the better production decision. Compare repeated executions to see which model delivers the right balance of success rate, latency, tool accuracy, and spend.

  • Side-by-side evidence across two to five models
  • Repeated trials that expose inconsistent behavior
  • Per-execution latency, tokens, tool calls, and cost
MEASURED TRADEOFFS
CLAUDE GEMINI
Task success94% / 88%
Tool accuracy91% / 73%
Cost efficiency72% / 96%
REPEATED 5× PER MODEL96% CONFIDENCE

From question to decision

A model decision in three controlled moves.

Define the work once, run every candidate under identical conditions, and inspect the evidence behind the outcome.

  1. 01

    Define success

    Choose or author a benchmark with explicit tasks, tools, and pass conditions.

  2. 02

    Run the lineup

    Select your models, repeats, baseline, and hard usage limits.

  3. 03

    Inspect the evidence

    Compare results and replay the exact behavior behind every score.

EXECUTION TRACE / TRIAL 05REPLAYING
01

MODEL

Plan response using approved sources

820ms
02

TOOL

search_knowledge({ query: "refund policy" })

146ms
03

MODEL

Compose grounded response

1.2s
04

GRADE

Required source and answer verified

PASS

PASS

OUTCOME

2.17s

LATENCY

$0.021

COST

Choose with evidence

Know which model belongs in your product.

Bring your workload, choose your model lineup, and leave with a decision you can defend.