Model evaluation / real workload evidence
Test models on the work that matters.
Replace leaderboard assumptions with evidence from your own tasks. Run competing models under the same conditions and see which one is ready for your product.
MODEL EVALUATION / RUN 0842
support-resolution-v3 · 5 repeats
240
EXECUTIONS
99.4%
TOOL SUCCESS
$4.82
TOTAL SPEND
MODEL LINEUP
Up to 5
compared in one controlled run
RUN CONTROL
Hard limits
for steps, tokens, and total spend
EVIDENCE
Full trace
for every response and tool call
01 / Workload evidence
Test the work your models will actually perform.
Generic leaderboards cannot tell you which model will succeed inside your product. Trunchbull runs every candidate against the same versioned tasks, tools, and pass conditions so the result reflects your workload.
- Your prompts, files, tools, and expected outcomes
- The same evaluation contract for every model
- Versioned definitions for reproducible results
Resolve a customer question using the approved knowledge source.
02 / Measured tradeoffs
See quality, reliability, speed, and cost together.
A higher score is not always the better production decision. Compare repeated executions to see which model delivers the right balance of success rate, latency, tool accuracy, and spend.
- Side-by-side evidence across two to five models
- Repeated trials that expose inconsistent behavior
- Per-execution latency, tokens, tool calls, and cost
From question to decision
A model decision in three controlled moves.
Define the work once, run every candidate under identical conditions, and inspect the evidence behind the outcome.
- 01
Define success
Choose or author a benchmark with explicit tasks, tools, and pass conditions.
- 02
Run the lineup
Select your models, repeats, baseline, and hard usage limits.
- 03
Inspect the evidence
Compare results and replay the exact behavior behind every score.
MODEL
Plan response using approved sources
TOOL
search_knowledge({ query: "refund policy" })
MODEL
Compose grounded response
GRADE
Required source and answer verified
PASS
OUTCOME
2.17s
LATENCY
$0.021
COST
Choose with evidence
Know which model belongs in your product.
Bring your workload, choose your model lineup, and leave with a decision you can defend.