Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
The gap between a model's headline score and its real-world cost comes down to a number nobody puts on the slide: how much time it was given.
Written by OutOfToken AI
August 10, 2026 · 6 min read · Synthesized from reporting by VentureBeat · How this works
Alibaba marketed Qwen 3.8-Max as trailing only Claude Fable 5 on agentic benchmarks. Alibaba's own launch table was more modest, showing the model leading on just one of twelve coding-agent rows, and an independent harness pushed further still, ranking the model's default setting last in its field. Both results are real. The difference between them is a matter of time and token budgets that never made it onto anyone's marketing chart.
Same model, opposite verdicts
Alibaba's own coding benchmarks gave Qwen 3.8-Max a five-hour timeout, stretching to twelve hours on PaperBench. VulcanBench, the independent harness that produced the far less flattering result, capped runs at 45 to 60 minutes. A time budget five to sixteen times larger changes what a reasoning model can accomplish before it has to commit to an answer, and that gap alone explains most of the disagreement between the two results.
Price per token stopped meaning what it used to
Qwen 3.8-Max lists at $2 per million input tokens and $6 output, well above DeepSeek-V4-Flash-0731's 14 cents and 28 cents, and below Kimi K3's $3 and $15. Those numbers look decisive until you account for reasoning tokens, which reasoning models spend liberally before ever writing a visible answer. Artificial Analysis found DeepSeek-V4-Flash at maximum effort burning 210 million output tokens against a class median of 100 million on its Intelligence Index — cheap in dollars, expensive in time, which matters whenever a budget or deadline is part of the job.
"A run that produces a wrong answer and a run that simply runs out of budget are different failures with different fixes — almost no leaderboard tells you which one happened."
Timeouts, not wrong answers, dominate failure
Long-Horizon-Terminal-Bench, testing 17 frontier models across 46 tasks with a shared 90-minute limit, found timeouts responsible for 79% of unresolved runs, versus 19% for agents that gave up on their own and 3% for harness errors. The timed-out runs weren't close to finishing either, so more time wasn't guaranteed to help. VulcanBench's own data on Claude Opus 5 makes the mechanism concrete: its lowest-effort setting solved 20 of 23 tasks, its highest effort only 18, with two of the losses being straightforward timeouts on tasks the cheap setting handled easily.
The metric everyone is converging on
VulcanBench, Long-Horizon-Terminal-Bench, and TestEvo-Bench have each independently landed on cost per successful task as the number that actually matters — total spend, including failed attempts, divided by tasks that passed an acceptance check. Vendors are moving the same direction: HubSpot's Breeze Customer Agent bills 50 cents per resolved conversation rather than per conversation handled, and Fin charges only on completed outcomes. The industry consensus, arriving from research and pricing teams simultaneously, is that per-token cost was never the number to optimize.
None of this makes Qwen 3.8-Max a bad model or Claude Opus 5 a broken one — it makes the leaderboard an incomplete instrument. Teams that log failure reasons, cap by token rather than wall clock, and check what effort setting they're actually running in production will find their real costs look nothing like the rate card. The models aren't lying to each other's benchmarks; the benchmarks are just measuring different bets on time.
Editorial Note
The research sources confirm the article's core factual claims about benchmark comparisons, timeout specifications, and VulcanBench/Long-Horizon-Terminal-Bench findings. The article appears to be based on primary research from these published benchmarks. However, specific pricing claims for some models (like DeepSeek) and some vendor billing details lack independent corroboration in the provided research materials.
Claim Tracker
AI-assessed
Source 1 (VentureBeat article) and Source 2 (MasterNodeAI) both confirm this specific claim about Alibaba's own launch table.
Source 1 (VentureBeat) corroborates these specific timeout figures.
The research provided (Sources 1-6) does not contain independent verification of these specific pricing figures from DeepSeek.
Source 1 (VentureBeat article) specifically cites these exact statistics from Long-Horizon-Terminal-Bench study.
Source 1 (VentureBeat) provides these specific performance figures from VulcanBench testing.
Ask AI about this story
// discussion
sign in to join the discussion
