AI Agents in Production: What the ReAct Loop Doesn't Tell You
The ReAct pattern is a dozen lines in every tutorial. Here's what actually breaks once an agent runs against real tools, real users, and a real bill — loop limits, tool-result size, streaming intermediate steps, and why 'the agent decided to' is not an error message.
AI Agents in Production: What the ReAct Loop Doesn't Tell You
An AI agent is the easiest thing in this whole stack to demo and the hardest to keep honest. The tutorial version is a loop: the model picks a tool, you run the tool, you feed the result back, repeat until the model stops asking for tools. That's the ReAct pattern, and it fits in a dozen lines.
The trouble is that every one of those dozen lines hides a decision that only becomes visible in production. This article is about those decisions — not the pattern itself, but the parts the pattern leaves out.
The Loop Needs a Hard Ceiling, and the Ceiling Is a Product Decision
The first thing that goes wrong with an agent is that it doesn't stop. A model that misreads a tool result will happily call the same tool again with a slightly different argument, forever, burning tokens the whole way. There is no "the model got bored" — there is only your iteration limit.
In a LangChain-based agent that limit is explicit:
agent_executor = AgentExecutor(
agent=agent,
tools=self.tools,
max_iterations=10,
handle_parsing_errors=True,
# ...
)
Two things are worth noticing here, and neither is about the number.
max_iterations is a cost control, not a safety net. Ten iterations means up to ten model calls plus ten tool executions per user request. If your pricing assumes one call per request, an agent that routinely uses six iterations is six times the cost you modeled. The limit isn't there to stop runaway loops — it's there to bound the worst case you're willing to pay for. Pick it from your cost model, not from a blog post.
handle_parsing_errors=True is doing more than it looks like. When the model emits a tool call that doesn't parse — malformed JSON, a tool name that doesn't exist, arguments in the wrong shape — the default behavior is to raise. With this flag, the parse error is fed back to the model as an observation and the loop continues. That's usually what you want, but it means a model that consistently misformats its tool calls will burn all ten iterations producing nothing, and the user sees a slow failure instead of a fast one. Worth a metric: how often does a request end because it hit the iteration ceiling rather than because the agent finished?
Tool Results Are Untrusted Input, and They're Big
The second thing that goes wrong is size. A tool that scrapes a web page returns the whole page. A tool that reads a PDF returns every page. A tool that queries a spreadsheet returns every row you asked for. All of that goes back into the context window on the next iteration — and the next iteration's prompt is now the original question plus every tool result so far.
This is where agents get expensive in a way that's invisible in the demo. The first iteration is cheap. The fifth iteration carries four tool results, some of them thousands of tokens, and you're paying for all of it on every subsequent call. An agent that "works" in testing can cost an order of magnitude more on a real document than on a toy question.
The fix isn't clever — it's truncation and summarization at the tool boundary. A scraper should return the relevant section, not the page. A document tool should return the extracted text with a page count, not the raw bytes. If a tool result is going to be large, that's a design decision about the tool, not about the agent.
There's a security angle here too, and it's the one people skip: tool results are attacker-controlled text. If your agent scrapes a web page, that page can contain instructions. If it reads a document a user uploaded, that document can contain instructions. The model has no reliable way to distinguish "data I retrieved" from "instructions I was given." This is prompt injection, and an agent with tools is a much bigger surface than a chat completion, because the injected instruction can now cause a tool call — fetch this URL, write to this sheet, send this webhook. Treat tool output as untrusted, and be very careful about which tools have side effects.
Streaming an Agent Is Not Streaming a Chat
A chat completion streams tokens. An agent streams events: the model thinking, a tool being selected, the tool running, the tool's result, the model thinking again. If you stream only the final answer, the user stares at a spinner for the entire multi-step run — which, at six iterations, is a long time.
So the stream has to carry intermediate steps. That's a different protocol from plain token streaming, and it means your client needs to render states it didn't have before: "searching the web", "reading a document", "running code". The upside is that it's also the best debugging tool you have — when an agent gives a bad answer, the step trace tells you whether it picked the wrong tool, got a bad tool result, or reasoned correctly from bad data.
The failure mode to watch for is the same one that bites streaming chat: an error after the response has started. Once you've sent the first event, you can't send a 500. The error has to travel as an event inside the stream, or the client renders a truncated run as if it were a complete one.
The Tool Registry Is the Product
Here's the part that took me longest to internalize: the agent's quality is almost entirely a function of its tools, not its prompt. A brilliant prompt over three vague tools produces vague behavior. A plain prompt over well-described, well-scoped tools produces useful behavior.
That means the tool descriptions are load-bearing. They're not documentation — they're the model's only interface to the tool. A description that says "search the web" gives the model nothing to decide with. A description that says what the tool is for, what the input looks like, and what comes back gives it something to reason about:
Tool(
name="web_search",
description="""Search the web using DuckDuckGo.
Use this to find current information, news, or research topics.
Input should be a search query string.
Returns list of search results with title, URL, and snippet.""",
func=lambda q: web_search.search(q, max_results=5),
)
The max_results=5 in there is not an accident either. A tool that returns fifty results is a tool that floods the context window on the first iteration. Tools should return the smallest useful answer by default, and let the model ask for more if it needs it.
The other thing about the registry: it grows. What starts as web search and a scraper becomes document extraction, a code interpreter, vision, NLP helpers, spreadsheet access, translation, webhooks. Every one of those is a new way for the agent to do something you didn't anticipate, and a new way for injected text to cause a side effect. The registry is the attack surface, and it deserves the same review as any other endpoint that can write data.
What "The Agent Decided To" Actually Means
When an agent does something surprising, the instinct is to blame the model. Usually the cause is upstream: a tool description that was ambiguous, a tool result that was truncated in a way that removed the disambiguating part, an iteration limit that cut the run off before it could correct itself, or a context window that dropped the earlier steps.
None of those are model problems. They're interface problems, and they're fixable. The reason to log the full step trace — every tool call, every argument, every result — is that it's the only way to tell the difference between "the model is bad at this" and "we gave it a bad interface." Without the trace, every agent bug looks like the same bug.
Conclusion
The ReAct loop is the easy part. What actually determines whether an agent works in production is everything around it:
- A hard iteration ceiling chosen from your cost model, plus a metric for how often you hit it.
- Tool results treated as untrusted and bounded — truncated at the tool, never dumped raw into the next prompt.
- A stream that carries intermediate steps, so the user isn't staring at a spinner and you have a trace to debug with.
- A tool registry treated as the product — descriptions written for the model, defaults that return small answers, and side-effecting tools reviewed like any other write endpoint.
Get those right and the loop is boring, which is exactly what you want it to be.
Further reading: LangChain – Agents for the AgentExecutor and tool-calling interfaces; OWASP – LLM01: Prompt Injection for why tool output has to be treated as untrusted input.