Published — 6 min.

AI agent orchestration the 5 patterns that decide whether your agent ships or hallucinates

AI agent orchestration the 5 patterns that decide whether your agent ships or hallucinates cover image

Part 3 of our guide to how to build an AI agent for production.

Same model, same tools, same prompts. With one orchestration pattern you get a reliable system. With another you get something that loops forever or quietly produces nonsense. The orchestrator is the single decision that most determines which one you ship.

There are five patterns we reach for in production. None are exotic. They're the ones that survive contact with real traffic.

The rule up front: pick the simplest pattern that solves your problem. You pay for autonomy in reliability, latency, and tokens. Every step down this list costs you something.

What is AI agent orchestration?

Orchestration is the runtime logic that drives an agent's loop: build the prompt, call the model, read its decision, invoke a tool, decide what happens next, and decide when to stop. It's component one of the five in every AI agent architecture.

The patterns differ in one dimension: how much of "what happens next" is decided by code versus by the model. At one end, code decides everything and the LLM is a function. At the other, the model plans, delegates, and self-corrects with minimal guard rails. Everything in between is a trade-off.

1. LLM workflow: deterministic agentic workflows

The flow is hard-coded. The LLM is called inside the graph as a function: summarize this, extract these entities, classify this text. The model never decides what happens next. The code does.

This is what most of the industry means by "agentic workflows", and it is by far the most common pattern actually running in production.

  • Pros: Reliable, predictable, cheap. The token bill doesn't surprise you. The behaviour doesn't drift.
  • Cons: You write the graph yourself. Useless for open-ended tasks.
  • Use when: The process is well-defined and the cost of error is high. Document workflows. Support routing. Compliance flows. Anywhere a regulator might one day ask you to explain a decision.

If you can solve your problem with a workflow, do it and go home. The team building a "fully autonomous agent" for a problem that fits in a workflow is the team debugging it on a Saturday.

2. ReAct agent: reason, then act

The base autonomous pattern, introduced by Yao et al. in 2022. The model loops: think, pick a tool, see the result, think again. The LLM decides which tool to call and when to stop.

  • Pros: Flexible. Adapts to the task. Easy to add tools.
  • Cons: Breaks on long tasks. The agent gets stuck in a tool-call loop or forgets the goal.
  • Use when: Short tasks with a small tool set. "Look up the exchange rate and post it to Slack." "Search the docs and answer the question."

ReAct is the right default for prototypes and the wrong default for anything that runs more than a handful of steps.

3. Reflexion: reason, act, check

A small but powerful addition to ReAct, introduced by Shinn et al. After every action the agent stops to evaluate: did that work? If not, it revises and retries. It can loop on a single action several times before moving on.

  • Pros: Massively raises quality where results are verifiable: code (do the tests pass?), math (is the answer right?), structured extraction (does it match the schema?).
  • Cons: Eats tokens. Slower per task. Over-second-guesses when feedback is ambiguous.
  • Use when: The tool gives feedback the agent can act on. Execution errors. Schema-validation failures. Test results.

If you're building anything that writes code, this is probably your pattern. The self-check step is a built-in evaluation loop running inside the agent.

4. Plan-and-execute: plan first, then run

The LLM produces a plan up front. A second orchestrator (often a Reflexion loop) walks through it step by step, all in one shared context window. When the plan is done, the LLM checks: task solved, or do we need a new plan? The idea traces to Plan-and-Solve prompting (Wang et al.) and was popularized as an agent pattern by LangChain.

  • Pros: Handles long, multi-step tasks that pure ReAct can't.
  • Cons: Context bloat. As steps accumulate, the history fills with intermediate work and the model starts breaking under the weight of its own scratch.
  • Use when: Long chains of dependent steps. Sequential analytics. Multi-file refactors. Anything where step N's input depends on step N-1's output.

Watch the context like a hawk. This pattern fails far more often from context drift than from bad planning.

5. Multi-agent orchestration: plan, then delegate

Same planning step. But each task is handed to a separate sub-agent that runs in isolation with only the context it needs. Sub-agents don't share state; they get inputs and return outputs.

  • Pros: The power of planning with the reliability of small, scoped executors. Solves the context-bloat problem of pattern 4.
  • Cons: Only works when the task genuinely decomposes into independent pieces. If sub-tasks need shared context, this pattern fights you.
  • Use when: The work is naturally parallel. A long report, one section per sub-agent. Deep research over many sources. Anthropic's research-agent architecture is this pattern.

Multi-agent systems are also the hardest to observe. Every sub-agent needs its own trace nested under the parent's, or debugging becomes archaeology. See LLM observability.

A note on orchestration frameworks

LangGraph, the OpenAI Agents SDK, Google's ADK, and a dozen others all implement these same five patterns with different ergonomics. The framework matters less than the pattern. Pick the framework your team can debug at 2am; pick the pattern based on the task.

How we actually choose

In production the answer is almost always a combination: a reliable outer workflow with autonomous agents at the leaves where flexibility genuinely pays. The best agent isn't the most autonomous one. It's autonomous only where autonomy earns its keep.

The decision rule we use with clients:

  1. Can you write it as a deterministic workflow? Do that. Stop reading.
  2. Can the model verify its own work? Reflexion.
  3. Is the task short and tool-using? ReAct.
  4. Long, with dependent steps? Plan-and-execute.
  5. Decomposes into truly independent pieces? Multi-agent.

Every step down this list trades reliability for flexibility and adds tokens and debugging. Don't pay unless the problem demands it. And whichever you pick, treat it as a hypothesis to be measured, not a decision to be defended.


This post is part 3 of our guide to how to build an AI agent for production. The series:

  1. AI implementation strategy: why AI projects die in pilot
  2. AI agent architecture: the 5 components
  3. AI agent orchestration: the 5 patterns (you are here)
  4. LLM-as-a-judge: evaluating LLM outputs
  5. LLM observability: debugging agents in production

Stuck choosing an orchestration pattern, or watching one fail in production? This is the decision we help teams make every week. Talk to us about AI agent development.