← All posts

Duration, not difficulty, is what breaks agents

Four 2026 benchmarks disagree about how capable agents are. They agree about what predicts failure: how long the task would take a human, not how good the model is.

By Finley Jones, CCMO, Taskpool (@finjonesceo) Published 12 August 2026 · Updated 6 September 2026

Benchmark leaderboards in this post move. Figures refreshed quarterly; last checked 6 September 2026.


Give the best agents eight tries at a task a professional would finish in under two hours and they top out near 40%. One attempt gets 24%. Both numbers are Mercor's, on APEX-Agents, and the second one is the one that got the headlines.

The first one is more useful. If the failures were noise around a competent process, eight independent draws would push success well above 40%. They are correlated instead. The agent reliably hits the same wall, and the wall turns out to sit at a task length rather than at a difficulty level.

Two claims circulated in 2026, both citing benchmarks, both apparently about the same technology. Agents now match human performance on real computer tasks: Stanford's AI Index has OSWorld success at 66.3%, up from roughly 12% in early 2024, against a human baseline of 72.35%. Agents fail three quarters of real professional work: APEX-Agents, built from tasks written by working investment banking analysts, management consultants and corporate lawyers, has the best model at 24%.

Both are real. The gap is not a methodology dispute.

The variable

METR's longitudinal work makes the pattern visible. It measures the length of task, in human completion time, at which a frontier agent succeeds half the time. That horizon has been doubling roughly every seven months since 2019. Independent recalculations of the post-2023 series put the doubling closer to four months, and I would treat the shorter figure as the less settled of the two.

The distribution underneath the headline matters more than the trend line. Agents succeed on close to 100% of tasks a human would finish in under four minutes, and on under 10% of tasks that take a human more than four hours. METR notes its own measurements above sixteen hours are unreliable with the current task suite.

One caveat that most citations of this curve drop. METR's suite is RE-Bench, HCAST and SWAA: machine learning research engineering, software engineering, and short computer-operation tasks. It does not test whether duration beats domain across law, banking or consulting. That extension is mine, and what supports it is that four separate benchmarks in unrelated domains all bend at the same place.

APEX-Agents reports the inflection directly. Performance degrades after roughly 35 minutes of task time, with failure rates scaling exponentially from there. On Artificial Analysis's independent reimplementation of the benchmark, Gemini 3 Flash leads at 24.0%, followed by GPT-5.2, Claude Opus 4.5 and Gemini 3 Pro. Those are AA's numbers under AA's harness, not Mercor's own leaderboard, and the two implementations do not have to agree.

OSWorld tasks are bounded computer operations: open an interface, manipulate a file, run a short multi-step workflow. APEX-Agents tasks are estimated at 1.8 hours of human professional effort and require moving across documents, spreadsheets, PDFs, email and calendars. The two benchmarks are sampling opposite ends of one curve.

Why per-step reliability compounds against you

The mechanism is unglamorous, and discussion of reasoning and intelligence tends to obscure it.

An agentic task is a chain of steps where later steps depend on earlier ones. If per-step reliability is r and the task requires n steps, success is approximately r^n.

At 99% per-step reliability, a 50-step task succeeds 61% of the time, a 200-step task 13% of the time, and a 500-step task 0.7% of the time. At 99.9%, those become 95%, 82% and 61%.

Two consequences follow. An order-of-magnitude improvement in per-step reliability buys only a roughly linear extension of workable horizon, which is why METR measured a steady doubling rather than a step change. And headline model quality is a weak lever, because moving from a good model to a slightly better one adjusts r in the third decimal place while n stays fixed by the task.

The other benchmarks fit. WebArena has climbed from 11.7% for GPT-4 in the original 2023 paper to 61.7% for IBM's CUGA and 74.3% in the 2026 AI Index, against a human baseline of 78.24%. WebChoreArena, which adds long-horizon memory and cross-page retention to the same four environments, drops Gemini 2.5 Pro from 54.8% to 37.8%. Same model, same underlying sites, one added requirement to hold state across a longer run.

UC Berkeley's Center for Responsible, Decentralized Intelligence built Agents' Last Exam from 1,490 assignments supplied by more than 250 working professionals across 55 occupations, and found frontier agents completing about one in four. On ALE's hardest tier every agent tested scored zero. Yiyou Sun, who led the work with Dawn Song, said there were "predictions everywhere" that agents would surpass humans at most jobs in 2026 or 2027, and that the exam exists to check the claim.

The 1.8 hours is an estimate, and it runs high

Here is the part that complicates everything above.

The duration figures underpinning this whole argument are, in the APEX case, self-reported by the same experts who wrote the tasks. Mercor's paper checked a sample against real completion times and found experts estimated 1.70 hours against a true 1.37 hours. An overestimate of 33%.

So the benchmark whose tasks I have been calling multi-hour has a mean nearer 80 minutes than 110, and the 35-minute inflection sits proportionally closer to the middle of the distribution than the raw numbers suggest. The direction of the argument survives. The precision does not, and anyone reproducing this should treat published task durations as soft.

METR's numbers are better in this respect, because METR timed humans directly rather than asking them to guess. That is also why METR's suite is small and mostly software.

Where the steps come from

If n is the problem, what inflates n is worth knowing, and it is rarely the interesting part of the work.

The AgentBay paper quantifies the mundane obstacles in a Claude Sonnet 4.5 harness. Dynamic UI elements such as floating ads caused reading failures 73% of the time. CAPTCHAs and similar adversarial mechanisms produced handling failures 36% of the time. Password entry failed 100% of the time by design, since the agent has no credentials.

Two things about that source. It is a single paper's own runs, not a survey. And it exists to argue for a human-takeover sandbox, which is roughly the conclusion I reach at the end of this post, so read the numbers with that in mind and tell me if you have better ones.

Neither failure is a reasoning failure. Both come from an environment built to resist automation, and every retry loop they trigger multiplies step count against a per-step reliability that is already the binding constraint. The 700-agent coordination incident is the same phenomenon at a different scale.

Enterprise deployment

This is where benchmark numbers and production numbers stop resembling each other. Deloitte's 2026 technology trends work found 11% of surveyed organisations running agentic systems in production against 38% piloting them, which is where the widely quoted 89% failure rate comes from. It is a derived figure, and Stanford is not its source despite being cited as such in a fair amount of coverage.

Per-project cost figures in the $150,000 to $800,000 range circulate alongside it. I have not been able to trace those to a primary source and I would not use them.

Gartner predicted in June 2025 that more than 40% of agentic AI projects would be cancelled by the end of 2027, citing costs, unclear value and inadequate risk controls. Most 2026 coverage quotes it without the date. Gartner separately forecasts 40% of enterprise applications embedding task-specific agents by the end of 2026. Those forecasts are not in tension. Embedding an agent in a bounded workflow works. Deploying one against a multi-hour professional task does not, and the second category is where the budgets went.

The multi-agent objection

Parallelism is the standard rebuttal: run many agents, have them check each other, and the compounding failure argument dissolves.

It does not, for the reason the eight-attempt result already showed. Parallel agents help when failures are independent. Agents drawing on the same model, given the same context, fail in correlated ways at the same obstacles. Twenty of them at a CAPTCHA produces twenty failures.

Verification by a second agent helps more, but only where verification is genuinely cheaper and more reliable than the original task. That holds for code with a test suite. It holds poorly for a consulting deliverable, where checking the answer requires the same judgement as producing it.

There is a cost on the other side too. Multi-agent systems multiply the number of places where something unplanned can start, and they make trajectories much harder to reconstruct afterwards.

How to design a pipeline around the duration wall

The instinct on reading a compounding-failure argument is to chase reliability. That is the expensive path and it fights the arithmetic. The cheaper response is to attack n and to change what happens on failure.

Decompose into independently verifiable units. A four-hour task run as one trajectory inherits the full exponent. The same task split into twelve twenty-minute segments, each with a checkable output, inherits a much shorter exponent twelve times over, and a failure costs twenty minutes rather than four hours. The cost is real: you are now maintaining twelve interfaces and twelve verifiers, and for tasks under about thirty minutes that overhead is not worth paying.

Make state durable and inspectable at the seams. Recovery is only possible if there is a defined thing to recover to. Agent frameworks that hold context in memory across a long run make retry equivalent to restart.

Decide what happens at the wall. Every long-running agent reaches its horizon. Three behaviours are possible: retry, escalate, or corrupt state silently. Only the third is unacceptable, and it is the default when nobody chose. The PocketOS deletion is the shape of it. That agent did not fail at a hard task. It succeeded at a destructive one nobody had bounded.

Measure duration, not just accuracy. Most teams evaluating agents record pass rates on a fixed task set. Expected human completion time per task is the more predictive field. An eval suite where every task takes a human under five minutes will report numbers that do not survive contact with real work. Given the estimation bias above, time a few tasks yourself rather than asking people how long they take.

Escalate on a budget rather than on an error. Waiting for a hard failure means waiting past the point where state was recoverable. Step count, elapsed time and repeated tool-call patterns are all available before anything breaks. This will fire on runs that would have succeeded, and you will pay for interventions you did not need.

What would change my mind

Two things.

If someone demonstrates that a decomposed twelve-segment pipeline does not in fact beat a single four-hour trajectory once you count the coordination overhead and the failures introduced at the seams, the central recommendation here is wrong. I have not seen that measured either way, and I would like to.

If verification turns out to be reliably cheaper than production across a domain without formal tests, the multi-agent objection wins and most of this becomes a transitional problem.

The argument about whether agents are overhyped or underhyped is mostly people quoting benchmarks from opposite ends of the duration curve at each other. Agents are close to reliable on work that takes a human minutes, close to useless on work that takes a human a day, and improving at a rate that moves that boundary by a predictable amount every few months. The economics sit at the failure point.

Taskpool is built around that handoff: a way for an agent to escalate to a verified human at the point where duration, verification or physical reality defeats it, rather than failing silently or grinding through retries.


FAQ

How long can an AI agent run before it fails? On METR's software task suite, frontier agents succeed on close to 100% of tasks a human would finish in under four minutes and under 10% of tasks taking a human over four hours. On APEX-Agents, which uses professional services tasks, measured performance degrades after roughly 35 minutes of task time.

Why do AI agents fail on long tasks when the model is capable? Because steps compound. At 99% per-step reliability a 200-step task succeeds 13% of the time. Improving the model adjusts per-step reliability in the third decimal place while the number of steps stays fixed by the task, so model upgrades buy roughly linear extensions of horizon rather than step changes.

Does running agents in parallel fix the reliability problem? Not on the evidence available. Mercor reports the best agents reaching around 40% on APEX-Agents given eight attempts, against 24% at one attempt. If failures were independent, eight draws would put success far higher. Agents sharing a model and a context fail at the same obstacles.

What task length should I design for? Segments a human would complete in under about twenty minutes, each producing an output something else can check, with a defined escalation path when a segment stalls. Below roughly thirty minutes the decomposition overhead usually exceeds the benefit.


The 1.8-hour figure is the number I trust least in this post, because it is an expert estimate and Mercor's own sample suggests those run about a third high. If you have measured real completion times on work of this kind, that is the data I want.

← Back to the blog