Published — 6 min.

AI agent architecture the 5 components inside every production agent

AI agent architecture the 5 components inside every production agent cover image

Part 2 of our guide to how to build an AI agent for production.

"Agent" means ten different things depending on who's pitching. The good news: under the hood, every production agent we've worked on is built from the same five components. Once you can name them, you can have grown-up conversations about architecture, debugging, and what to improve next, instead of arguing about whether a given model is "agentic enough".

We treat agentic AI architecture like a recipe. Once the ingredients are standard, you stop reinventing them for every project. Improve one component and every agent in the company gets better. Build one tool and it's reusable in the next product.

How do AI agents work? The architecture diagram

An AI agent is a loop. The orchestrator builds a prompt from the context window, sends it to the LLM, reads the model's decision, calls a tool if one was requested, writes the result back into context, and repeats until the task is done. Every call in and out passes through a gateway that logs it and checks it.

Blog post image

Here's what each of the five does, and where it tends to break.

1. The orchestrator: the agent's heart

The orchestrator is the runtime that runs the loop. It calls the model, parses the response, invokes tools, handles errors, and decides when to stop. It implements a pattern: a deterministic workflow, a reason-and-act loop, plan-then-execute, something custom.

Physically, it's code. Usually Python, often on top of LangGraph or LangChain, sometimes written from scratch when the off-the-shelf options do too much.

The choice of orchestrator is the biggest reliability decision you'll make. We compare the five patterns in AI agent orchestration: the 5 patterns that decide whether your agent ships or hallucinates.

2. The LLM: the agent's brain

The orchestrator builds a prompt and sends it to a model so the brain can decide what to do next. The non-obvious lesson from production: you don't need one model. Use the expensive frontier model for the rare hard calls and a cheaper model for the high-volume routine ones. Most teams over-spend by 5 to 10x because they default the whole agent to the most capable model available.

Physically, this is either a cloud provider API or a self-hosted model. Both are fine. The choice is about latency, cost, data residency, and how often you want to debug vendor outages.

3. Tools: the agent's hands

Anything the agent can call to do work in the world: APIs, database queries, file operations, code execution. Other agents, when you go multi-agent.

Use a protocol. MCP (Model Context Protocol) has become the default because it gives you a standard way to expose tools to any agent runtime. A tool you build once should be usable by every agent you build next.

Tools are also the most common place for things to break silently. Every tool needs to log inputs, outputs, and errors; we cover the pattern in LLM observability.

4. The context window: the agent's working memory

The hardest of the five. At the moment the model is invoked, the context needs to contain exactly the information needed for the next decision. Too little and the model is uninformed. Too much and it gets distracted, drifts, or hallucinates.

The bigger the context, the higher the chance the brain breaks. Teams who believe "larger context windows solve everything" learn otherwise around month three.

In practice this means:

  • Compression. Summarize earlier turns when they're no longer needed in full.
  • Eviction. Move stale information out of context into external storage.
  • Selective recall. Pull only the relevant pieces back in when needed.

This is engineering work. No library solves it for free.

5. External knowledge: what the agent can look up

Data the agent might need but doesn't need right now: knowledge bases, documents, codebases, structured data. The agent reaches it through retrieval (RAG) or through tools (a grep-equivalent over a repo, a SQL tool over a warehouse).

The thing to internalize: external knowledge is a tool problem, not a model problem. The agent's ability to find the right document at the right time is mostly about how well you've indexed your data and how good your retrieval tool is. The model is almost incidental.

How the components actually talk: the gateway layer

Everything routes through the orchestrator. But in a production system the orchestrator doesn't call the LLM or tools directly. It goes through gateways, small proxy services between the orchestrator and the outside world. They do two jobs:

  • Observability. Every call gets logged with a trace ID so you can reconstruct what the agent did and why. This is what makes evaluating agent quality possible at the step level, not just on the final answer.
  • Safety. Every call gets inspected. Permissions on file access. Prompt-injection checks on incoming context. PII redaction. Tool allow-lists.

If you skip the gateway layer, you'll add it later, usually after an incident. Add it now.

Why this decomposition matters

Three reasons we keep coming back to it with clients:

  1. Improvements compound. When components are clean, a better retrieval tool improves every agent that uses it. Bespoke spaghetti per agent means every improvement is a one-off.
  2. Debugging gets tractable. When an agent misbehaves, the question is "which component failed?", not "what is the LLM doing?". Five components, five places to look.
  3. You can build incrementally. You don't need all five on day one. Start with orchestrator plus LLM, add tools when you measure a need, add retrieval when you measure a need. That's the hypothesis-driven approach in architectural form.

This post is part 2 of our guide to how to build an AI agent for production. The series:

  1. AI implementation strategy: why AI projects die in pilot
  2. AI agent architecture: the 5 components (you are here)
  3. AI agent orchestration: the 5 patterns
  4. LLM-as-a-judge: evaluating LLM outputs
  5. LLM observability: debugging agents in production

Designing the architecture for an agent you're about to put in front of real users? We've shipped this exact stack at scale. Talk to us about AI agent development.