Part 5 of our guide to how to build an AI agent for production.
LLMs aren't known for reliability. Hallucinations are still the top blocker for putting them into business-critical workflows. And now we've handed them tools, the ability to plan, and multi-agent topologies for good measure. Can you engineer that system to never fail?
No. Something will break.
What you can engineer is the ability to find out where it broke and why, fast. That's observability. For AI agents it isn't optional and it doesn't show up by accident. You design for it from day one.
What is LLM observability?
In normal software, the code is the source of truth. Something went wrong? Read the code, find the bug.
In AI systems the code explains almost nothing. The source of truth is the history: what went in, what the model decided, which tools it called, what came back. LLM observability is the practice of making that history legible, so you can reconstruct any run and explain any output.
The AI observability stack has four parts. Skip one and the other three stop being useful.
1. Traces
Every agent run gets a unique trace ID. Every action the agent takes, every model call, tool invocation, sub-agent dispatch, and retrieval, is attached to that trace.
You end up with a tree. You can ask "what did this user see, and what did the agent do to produce it?" and get a complete answer. For multi-agent systems, sub-agent traces nest under the parent, or debugging becomes archaeology.
The instrumentation isn't optional. Teams that add tracing after launch add it during an incident, which is the worst possible time.
2. Logging
Every tool, without exception, logs:
- The inputs it was called with
- Intermediate state, if any
- Errors, with stack traces
- The final return value
This is what makes a trace debuggable rather than merely visible. A trace says the agent called search_docs("X") and got garbage back. A log says why: malformed query, stale index, embedding service timed out.
This is where most teams cut corners and pay later. Tools are the hands of the agent in the five-component architecture; cover them properly the first time.
3. Quality metrics
For each trace you want a score on the final output and, ideally, on the intermediate steps. Manual scoring doesn't scale past a small team, so most production systems use an LLM-as-a-judge here.
The judge tells you not just "this trace failed" but "this trace failed on relevance" or "on safety". When overall quality drops you can immediately ask which dimension dropped, which narrows the investigation by an order of magnitude.
4. Dashboards: the AI agent monitoring metrics that matter
On top of quality, track the operational metrics that tell you whether the system is healthy:
- Latency and p95 response time
- Tokens per task and cost per task
- Tool-call counts and tool error rates
- Loop depth (how many steps before the agent stopped)
- Judge scores per criterion, trended over time
None of these tell you about quality on their own. All of them tell you when something changed. A sudden rise in loop depth almost always means the orchestration pattern is fighting the task.
LLM observability tools: what we actually use
Three tiers worth knowing about:
- Open source, built for LLM apps. Langfuse and Arize Phoenix. Both are built on OpenTelemetry's GenAI semantic conventions, so traces stay portable. Both take you from nothing to real observability in about a week. This is where we start on most engagements.
- Framework-native. LangSmith if you're all-in on LangChain. Convenient, less portable.
- General APM with LLM add-ons. Datadog LLM Observability, Dynatrace, Grafana. The right call if your platform team already lives there and wants one pane of glass. More expensive per trace.
The tool matters less than the discipline. All three tiers can do the four-part stack above. None of them will add the trace IDs or tool logging for you.
How it works on a Tuesday afternoon
Once the stack is in place, the loop looks like this:
- A quality metric drops on the dashboard. Some slice of traffic is misbehaving.
- You filter to traces with bad judge scores over the last hour.
- You drill in. Which step failed? Retrieval? A specific tool? The final synthesis?
- The logs tell you the cause. You ship a fix. The dashboard recovers.
Without this, every regression is a multi-day investigation dependent on whoever remembers the codebase best. With it, the same regression is a 30-minute task for an on-call engineer.
This is also the difference between a project that leaves pilot and one that doesn't. We've watched teams with weaker models beat teams with stronger ones because they could iterate ten times faster on real failures. It's the engine behind the hypothesis loop in part 1: you can't cluster failures you can't see.
The bottom line
Black-box AI systems are great in demos. They impress investors and look good in a press release. They're miserable to scale and operate, and every failure is an all-nighter spent reconstructing what happened.
Make the box transparent. Trace every run. Log every tool. Score every output. Watch the dashboards. Then sleep.
This post is part 5 of our guide to how to build an AI agent for production. The series:
- AI implementation strategy: why AI projects die in pilot
- AI agent architecture: the 5 components
- AI agent orchestration: the 5 patterns
- LLM-as-a-judge: evaluating LLM outputs
- LLM observability: debugging agents in production (you are here)
Adding observability to an agent that's already live, or designing the stack before you ship? We've done both. Talk to us about AI agent development.
