Agent observability: how to find where an AI agent went wrong

Originally published on Medium

In this article
  1. Why the final answer is not enough
  2. We started with a simple question: which turn does this belong to?
  3. Keep a useful history, not a copy of everything
  4. A history becomes useful when it answers questions
  5. A shortened timeline
  6. Observability is not an eval
  7. What I would build first
Magnifying glass revealing the steps in an AI agent run.

TL;DR

  • An agent’s final answer shows the result, but not everything that happened before it.

  • Observability turns those hidden steps into a readable history.

  • Once you can explain a failure, you can turn it into a test that prevents the same problem returning.

When an AI agent returns the wrong result, the final answer is usually only the symptom. Before producing it, the agent may have assembled its instructions, loaded relevant memory and skills, chosen an AI service, called several tools, retried after a timeout and checked whether the work was complete. Any one of those steps can change the outcome, yet the user sees only the last message.

I ran into this while building Niki Studio, where agents create designs and slide decks. Ordinary logs told me that individual actions had happened, but they did not give me one clear story of the complete user turn. I needed to follow a request from beginning to end and see exactly where it changed direction.

Why the final answer is not enough

Imagine a user asks for a two-page campaign. The agent starts work on both pages, but the main AI service times out. A backup service succeeds, a tool rebuilds the design as an editable canvas and a quality check reviews the result. The first page is ready, while the second is still unfinished when the agent stops.

The final response cannot tell me whether both pages started, which service handled each one, which tools finished or why the agent believed the job was complete. To understand the failure, I need to see the whole user turn, not only the last model response.

Harrison Chase captures this neatly: “In software, the code documents the app; in AI, the traces do.” His article explains why the history of a run becomes an important part of understanding the application: https://www.langchain.com/blog/in-software-the-code-documents-the-app-in-ai-the-traces-do

We started with a simple question: which turn does this belong to?

Our first observability problem was simple: the system recorded what happened, but some records were not connected to the user request that caused them.

For example, Niki Studio might record that:

  • an AI service timed out;

  • a backup service was used;

  • a canvas tool completed; and

  • a quality check failed.

But every record also needs to say, “This happened during user turn 42 in conversation 17.” Without that connection, we could see that a timeout or tool call happened but could not always tell which user’s work it affected. It was like finding a receipt with every item listed but no order number.

We fixed this at the moment a request begins. Every important step receives the same conversation and turn identifiers, plus a number that shows where it happened in the sequence. This turns scattered records into a readable timeline. The final failure no longer hides a timeout, a backup attempt or a tool call that succeeded before it.

Keep a useful history, not a copy of everything

For each turn, Niki Studio keeps an ordered history of the important decisions and outcomes. New steps are added as the work progresses, while earlier steps remain unchanged. Each entry answers simple questions: what was attempted, how long it took, whether it worked, and why the agent continued or stopped.

This history deliberately leaves out user messages, full prompts, model responses, images, fetched web pages, credentials and private model reasoning. Saving all of that would create another store of sensitive information and bury the useful story under too much detail.

For example, the history needs to record that the main AI service timed out after a certain period and that the backup service succeeded. It does not need to copy the user’s complete brief or the AI service’s full response into a second system.

Observability is not about reading the model’s mind. It is about understanding what the application gave the agent, what the agent tried, what came back and what the system decided next.

A history becomes useful when it answers questions

A list of recorded steps is only raw material. It becomes useful when someone can open one failed turn and understand it without searching through several unrelated logs.

For an unfamiliar failure, I want one view to answer five questions: what did the user ask for, which agent handled it, which AI services and tools ran, where did the plan change, and why did the system finally report success or failure? Speed and cost still matter, but an agent can be fast and inexpensive while still being wrong.

A shortened timeline

The table below shows the same idea in plain language. Its value is the order: you can see the route that produced the result instead of starting with the final failure.

Observability is not an eval

The run history records what happened. Observability helps a person understand why it happened. An eval is the next step: it repeats important scenarios and checks whether the agent still meets a clear standard.

The first lesson is to check what actually happened, not only what the agent said. An agent might report that a two-page campaign is complete, but the real test is whether both pages were created, checked and made available to the user.

The second lesson is to turn important failures into repeatable tests. If Niki Studio stopped after completing only the first of two pages, we should recreate that situation as a test. After changing the model, prompt or workflow, we run it again and confirm that both pages finish.

That gives me the simplest distinction between the two ideas: observability helps us understand one failure; an eval checks that the same failure does not return.

Niki Studio already has automated checks, reproductions of past incidents and evidence from live runs, but I would not call that a complete eval system yet. The next step is to turn useful failures into a small collection of repeatable scenarios, define what success means for each one, and run future model, prompt and tool changes against them.

Anthropic’s guide explains why agent evals may need to judge both the final result and the path used to produce it: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents

What I would build first

If you are adding observability to an agent, start with one user turn instead of a large dashboard. Give that turn a stable identity, list the decisions that can change its outcome, record them as they happen and preserve their order. Keep private content out unless it is genuinely needed.

Then build the smallest screen that can explain one failed run. Charts, alerts and automated checks can come later. Otherwise, you may know that failures increased without being able to explain even one of them.

That is the main lesson I took from this work: observability begins with a history clear enough that another person can trust and follow it. When your agent produces the wrong result, can you find where the run changed direction, or can you only inspect its final answer?

Back to writing