Benchmark · live

Model reliability by task.

Aggregate cost and reliability for the models teams actually run on Daslab, grouped by the kind of work — not who ran what or how much, just how well each model does the job and what it costs.

Provisional · aggregates from live runs · fixed task set · thin cells hidden
The frontier · cheapest reliable model per task

The most reliable spend per task: the cheapest model that still finishes 90% or more of its runs. Finishing means it ran without error; correctness graders are coming.

TaskFrontier pickReliabilityCost / task
Finance & ERPgemini-3.7-flash100%$0.041
Codingkimi-k390–100%$0.746
Research & webdeepseek-v4-flash-0731100%$0.004
Comms & CRMgemini-3.7-flash100%$0.032
Ops & infragemini-3.5-flash90–100%$0.238
Data & sheetsnot enough runs yet
Finance & ERP
ModelReliabilityTool errorsCost / task
gemini-3.5-flash70–80%0–10%$0.605
claude-sonnet-590–100%0–10%$0.205
gemini-3.7-flash °100%$0.041
Coding
ModelReliabilityTool errorsCost / task
kimi-k3 °90–100%0–10%$0.746
gemini-3.5-flash 80–90%0–10%$0.481
qwen3.5-35b-a3b °60–70%20–30%$0.004
Research & web
ModelReliabilityTool errorsCost / task
gemini-3.5-flash50–60%0–10%$0.579
kimi-k390–100%0–10%$0.569
claude-sonnet-4-6 °80–90%$0.539
claude-sonnet-590–100%0–10%$0.186
kimi-k2.6 °100%$0.162
kimi-k2.7-code90–100%0–10%$0.147
qwen3.7-max °90–100%0–10%$0.129
auto80–90%$0.102
glm-5.290–100%0–10%$0.041
gemini-3.7-flash90–100%$0.035
deepseek-v4-flash-0731 100%0–10%$0.004
Comms & CRM
ModelReliabilityTool errorsCost / task
claude-sonnet-5 °90–100%0–10%$0.315
gemini-3.5-flash60–70%0–10%$0.120
glm-5.2 °90–100%0–10%$0.050
gemini-3.7-flash °100%$0.032
Ops & infra
ModelReliabilityTool errorsCost / task
gemini-3.5-flash 90–100%0–10%$0.238
Data & sheets
ModelReliabilityTool errorsCost / task
gemini-3.5-flash °80–90%0–10%$0.095

How this is measured. Every run on Daslab over the last 60 days is placed into a fixed task set by the tools it called, then aggregated by the model its turns actually ran on — a run that switched models mid-way is evidence about routing rather than about any one model, so it is excluded rather than credited to one. Reliability = the share of runs that finished without error (running runs excluded); it does not yet judge whether the answer was correct. Tool errors = the share of a model's tool calls that the tool rejected. Cost / task = median spend per finished run. marks the cost / reliability frontier: no other model is both cheaper and more reliable. ° marks a provisional cell.

Why the numbers are banded. Percentages are published in ten-point ranges on purpose. An exact percentage can be inverted — only certain values are reachable from a given number of runs, so “80%” would quietly disclose the sample it came from. Bands collapse many fractions onto one label. We publish how well the models do, never how much we run: no volumes, no counts, no per-customer data. Metrics that would need to know what the user wanted — how many turns a task took, whether depth was wanted — are deliberately absent, because fewer turns is not obviously better and a leaderboard that says so would reward models that give up early.