Why AI Agents Stop Mid-Task (and How to Diagnose It)
Why Does My AI Agent Stop Halfway Through a Task?
Six different failures produce the same symptom. Five are in your harness, one is in the model, and the one that costs most doesn't look like stopping at all.
Print the stop reason before you theorise. Nearly every conversation about agents quitting mid-task skips this step and goes straight to the model, and most of the time the model is not what stopped.
Start with LangGraph, because its failure is the most legible. The default recursion_limit is 25 steps, and when a graph hits it you get GraphRecursionError: Recursion limit of 25 reached without hitting a stop condition. That reads like a lot of tool calls. It isn't. One team debugging exactly this wrote up what the number actually counted: the limit counts super-steps, and their middleware chain made a single model-call-plus-tool round trip cost four of them, leaving roughly five or six tool calls per turn. Their own summary of the bug was blunter than mine: 25 was never 25 tool calls.
So an agent that looks like it gave up after five actions was allowed five actions. Nobody chose that number. It arrived as a default and the middleware ate it.
Where the other ceilings sit
The OpenAI Agents SDK has the same shape with different wording. The runner loops until the model produces final output, and if the run exceeds max_turns it raises MaxTurnsExceeded, with max_turns=None disabling the limit entirely. The default has been reported in the library's own issue tracker as DEFAULT_MAX_TURNS, currently set to 10, and the same issue notes something worth checking in your own code: when an agent is converted to a tool with as_tool, no max_turns is passed through, so the default is always used. Your carefully raised limit on the outer agent does nothing for the sub-agent you turned into a tool.
Then the token ceiling, which produces the most confusing symptom of the lot. stop_reason: "max_tokens" means the response was truncated at the limit you set, and in a tool-using loop this shows up in a specific way: if the response was cut off mid tool-use block, you have to retry with a higher max_tokens to get the full tool call. What the user sees is an agent that was clearly about to do something and then wasn't.
Three ceilings, three different error surfaces, none of them about capability. Check all three before you blame the model.
The stop that reports success
Here is the one that will waste your afternoon. Anthropic's own documentation describes it: Claude sometimes returns an empty response of two or three tokens with stop_reason: "end_turn", typically when it reads the assistant turn as already complete, particularly after tool results. One named cause is adding text blocks immediately after tool results, which teaches the model that the user always speaks next, so it ends its turn to match the pattern.
Your harness sees end_turn. end_turn means finished. The loop exits, the run is marked successful, and the task is a third done. If you are logging completions rather than trajectories, this failure is invisible to you until a customer finds it.
There's a related state worth handling explicitly. pause_turn means a server-tool loop paused and has to be continued by sending the assistant content back, and a harness that treats every non-error stop reason as OK will surface incomplete content as a completed model turn. I'd guess this accounts for more "the agent is lazy" complaints than actual laziness, though I can't prove the ratio and I've not seen anyone publish it.
When it really is the model
Now the failure you can't configure away.
The intuition is that a per-step error rate compounds, so long tasks fail more often. True, and incomplete. Sinha and colleagues found that per-step accuracy degrades as the number of steps increases, and that this is not just a long-context effect: models self-condition, becoming more likely to make mistakes when the context contains their errors from prior turns. Scaling the model does not remove self-conditioning, though thinking models do not show it and can execute much longer tasks in a single turn. Raising the error rate in the history produced sharp degradation in subsequent step accuracy.
Read that again with your retry logic in mind. The standard recovery move after a partial failure is to resume the run with the existing conversation, which means the model's next attempt is conditioned on a transcript containing its own mistakes. That's the configuration the paper shows degrades accuracy. Resuming from a failed trajectory is not neutral, it is actively worse than starting clean with the state summarised by something other than the failure itself.
Length alone hurts too, before any errors appear. Chroma's context rot report evaluated 18 models and found that performance grows increasingly unreliable as input length grows, with models not using their context uniformly. Their own limit is worth quoting back at anyone who over-reads it: they do not explain the mechanism, only observe that structural properties like placement and repetition influence behaviour. It happens well below the advertised window, which is why "we're only at 40% of 200k" isn't the reassurance it sounds like.
The capability story is covered properly in Duration, not difficulty, is what breaks agents, which is where the time-horizon evidence lives, so I won't restate the curve here.
A ten-minute triage
- Log
stop_reasonor its equivalent on every model call, not just the final one. Everything below depends on it. - Count actual tool calls in the failed run and compare against your step, turn and recursion limits. If the count lands on a round number, you found it.
- Check the limits on sub-agents and tool-wrapped agents separately. They inherit defaults, not your config.
- Look at the last content block when
stop_reasonismax_tokens. An incomplete tool call there is a truncation, not a decision. - Grep for runs that ended with a two-token response. If there are any, your loop is exiting on a pattern-matched
end_turnand your success metrics are wrong by that amount. - Diff the context length between runs that finished and runs that didn't. Not against the window size, against each other.
Only after those six do you get to talk about the model.
The fix for the last category isn't more retries. It's shorter runs with checkpointed state, so that a fresh context can pick up from a clean summary rather than from a transcript full of the previous attempt's errors. This costs you something real: you lose the fluency of one continuous run, you have to define what the intermediate state even is, and for some tasks that definition is most of the work. Some people will read that as an argument for longer context windows instead, and on current evidence about degradation with length I think they're wrong, but the evidence is about retrieval tasks rather than tool-using agents, so hold that loosely.
The stop you want
Everything above treats stopping as the problem. It's the good case.
An agent that halts visibly costs you a re-run. An agent that reports completion on work it didn't do costs you whatever the work was for, discovered later, by someone else. The benchmark version of this is in 63% of Those Benchmark Wins Were Lookups: systems can produce the artefact that scores as success without doing the thing the score is supposed to measure. In production the same gap shows up as a green run and a customer email.
So the instrumentation that matters isn't completion rate. It's whether you can reconstruct what the agent actually did, step by step, and check it against what it claimed. Teams that can do this for last week's runs are rarer than you'd expect from how often reliability gets discussed.
Taskpool sits at the point where that check needs a person: it's a marketplace where an AI agent hires a verified human for the part it can't complete itself, including verification of work an agent cannot verify on its own.
What I want and can't find anywhere: the distribution. Of production agent runs that stop early, what share are step limits, what share truncation, what share false end_turn, what share genuine execution failure? I'd guess harness causes dominate by a wide margin, but that's a guess dressed up in a percentage, and if you have the traces to check it against, I'd rather have your numbers than my guess.
FAQ
Why does my AI agent stop after only a few tool calls?
Most often a step or turn limit. LangGraph defaults to a recursion limit of 25 super-steps, which with middleware can mean five or six tool calls, and the OpenAI Agents SDK raises MaxTurnsExceeded at its own default. Check both before assuming a model problem.
What does stop_reason: max_tokens mean for an agent?
The output was truncated at your token limit. In a tool-using loop this can cut a tool-use block in half, so retry with a higher max_tokens rather than treating the run as complete.
Why does my agent return an empty response and end the run?
Claude can return a two or three token response with stop_reason: end_turn, often after tool results, particularly where text blocks are inserted immediately after them. Your loop reads that as a finished task.
Does a bigger context window fix agents that fail on long tasks?
Not reliably. Chroma's evaluation of 18 models found performance becoming increasingly unreliable as input length grows, well before the window is full.
Is retrying a failed agent run a good recovery strategy?
Not if the retry keeps the failed transcript. Models self-condition on their own earlier errors, so a fresh run seeded with clean summarised state beats resuming from the failure.
Alfred Sommarström is co-founder and Managing Director of Taskpool International Ltd. Figures and defaults verified on 17 September 2026.