# How to monitor AI agents in production

Canonical: https://www.swfte.com/how-to-monitor-ai-agents-in-production
Last verified: 2026-10-06
Difficulty: Advanced
Time: About 4 hours for tracing, redaction and the first alerts; the weekly review is ongoing
Cost: The libraries are free. A hosted trace backend may charge by volume; check its current pricing.
Hardware: None beyond where your agent already runs.

## Short answer

Give every agent run an id and record each model call and tool call as a span in a trace, with model, token counts, latency, outcome and any policy decision. Redact personal data before you store it. Alert on loop ceilings, tool failures, cost per run and a falling success rate, and read a random sample of runs by hand every week, because most agent failures raise no error.

## Who this is for

- Engineers who run an agent that calls tools and can now act on real systems.
- Platform and SRE teams asked to add monitoring to an AI feature they did not build.
- Security and risk owners who need a record of what an agent did and why.

Not for:
- Teams still building their first agent. Start with [how to build an AI agent](https://www.swfte.com/how-to-build-an-ai-agent).
- Anyone looking for a model benchmark. Monitoring tells you what happened in production, not which model is best. See [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai).

## Prerequisites

- An agent that already runs, with a loop you can edit (a model call, tool calls and a stop condition).
- Python 3.10 or later for the examples. The same ideas apply in other languages.
- A place to send traces: any OpenTelemetry-compatible backend, or a Langfuse project (cloud or self-hosted).
- A written definition of a successful run for this agent. If you do not have one, step 1 helps you write it.

## Steps

### Step 1: Define what one run is and what success means

Outcome: You have a written definition of a run, a success test, and a list of fields to capture.

Monitoring starts with a unit. For an agent, the unit is a run: one goal, from the first model call to the final answer or the point where it gave up. Give every run an id and put that id on every span, log line and tool call that belongs to it. Without it you cannot reconstruct what happened.

Then define success at the goal level, not the technical level. "The ticket was resolved and the customer did not reopen it" is a success test. "The API returned 200" is not. Write one sentence per agent. You will use it for sampling in step 7 and for alert thresholds in step 6.

Decide what to capture before you write code. The table below is a starting set. Capture identifiers and counts by default, and treat message content as opt-in, which is also how the OpenTelemetry convention treats it.

**What to record for each agent run**

| Field | Why you want it |
| --- | --- |
| Run id, agent name, version, owner | Find a run, and know who to call |
| User or tenant id (pseudonymised) | Spot one user causing most of the load |
| Each model call: model, input and output tokens, latency | Cost and slowness per step |
| Each tool call: tool name, arguments summary, result status, latency | Tool failures, unexpected tools, forbidden actions |
| Turn count and whether the ceiling was hit | Loops |
| Policy decisions: allowed, denied, escalated, approved by whom | Evidence for governance and audits |
| Outcome label: success, failure, escalated, abandoned | Success rate over time |
| Cost per run | Spend per run and per feature |

### Step 2: Wrap each run and model step in OpenTelemetry spans

Outcome: Running the agent prints a trace with one parent span for the run and child spans for model steps.

OpenTelemetry is a vendor-neutral standard for traces, so what you record can go to any compatible backend. Install the API and SDK, create a tracer provider, and wrap the run in a span. The console exporter prints finished spans to the terminal, which is the quickest way to check what you are recording before you add a backend.

For attribute names, follow the OpenTelemetry GenAI semantic conventions. As of 2026-10-06 they live in the semantic-conventions-genai repository, and the agent spans page carries Development status, which means names can still change. The agent span uses `gen_ai.operation.name` set to `invoke_agent`, plus `gen_ai.provider.name`, and `gen_ai.agent.name` when available. Token counts are recommended as `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens`. Put your own attributes under a prefix you control, such as `app.`, so they cannot collide with the spec.

The skeleton below is the shape, not a framework. `call_model` and `run_tool` are your own functions: `call_model` should return the reply and the token counts, and `run_tool` should execute one tool call. Set a turn ceiling and record when you hit it. That single attribute is the basis of your loop alert later.

Install the OpenTelemetry API and SDK:

```bash
pip install opentelemetry-api
pip install opentelemetry-sdk
```

Tracer setup and an instrumented agent loop (skeleton):

```python
import uuid

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter
from opentelemetry.trace import Status, StatusCode

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support-agent")

MAX_TURNS = 12


def run_agent(goal, user_pseudonym):
    run_id = str(uuid.uuid4())
    total_in = total_out = 0
    with tracer.start_as_current_span("invoke_agent support-agent") as run_span:
        run_span.set_attribute("gen_ai.operation.name", "invoke_agent")
        run_span.set_attribute("gen_ai.provider.name", "anthropic")  # use the value for your provider
        run_span.set_attribute("gen_ai.agent.name", "support-agent")
        run_span.set_attribute("app.run.id", run_id)
        run_span.set_attribute("app.user.pseudonym", user_pseudonym)
        run_span.set_attribute("app.agent.version", "2026-10-06")
        messages = [{"role": "user", "content": goal}]
        turns = 0
        outcome = "ceiling_hit"
        while turns < MAX_TURNS:
            turns += 1
            with tracer.start_as_current_span("model_step") as step:
                reply, tokens_in, tokens_out = call_model(messages)  # your function
                total_in += tokens_in
                total_out += tokens_out
                step.set_attribute("gen_ai.usage.input_tokens", tokens_in)
                step.set_attribute("gen_ai.usage.output_tokens", tokens_out)
            if reply.tool_calls:
                for call in reply.tool_calls:
                    run_tool(call, messages)  # step 3 wraps this in a span
            else:
                outcome = "answered"
                break
        run_span.set_attribute("gen_ai.usage.input_tokens", total_in)
        run_span.set_attribute("gen_ai.usage.output_tokens", total_out)
        run_span.set_attribute("app.agent.turns", turns)
        run_span.set_attribute("app.agent.outcome", outcome)
        return reply
```

> NOTE: The spec lists `gen_ai.provider.name` as required. Take the value for your provider from the semantic-conventions registry rather than inventing one.

### Step 3: Record every tool call as its own span, with errors

Outcome: Each tool call appears as a child span with its name, status and any exception.

Tools are where an agent touches the outside world, so they deserve their own spans. The GenAI conventions define an `execute_tool` operation: span name `execute_tool` followed by the tool name, kind `INTERNAL`, with `gen_ai.operation.name` set to `execute_tool` and `gen_ai.tool.name` required. `gen_ai.tool.call.id` is recommended when you have it.

When a tool fails, set the span status to error and record the exception, as the OpenTelemetry Python guide shows. Do not swallow the failure and return text to the model without leaving a trace. A tool that fails quietly and gets retried five times is one of the most common causes of loops and cost spikes.

Record the arguments as a short summary, not the raw values. Arguments often contain customer data. If you need the full detail for debugging, store it separately with a short retention period and reference it from the span.

Tool-call span with error recording:

```python
def run_tool(call, messages):
    with tracer.start_as_current_span("execute_tool " + call.name) as span:
        span.set_attribute("gen_ai.operation.name", "execute_tool")
        span.set_attribute("gen_ai.tool.name", call.name)
        span.set_attribute("gen_ai.tool.call.id", call.id)
        span.set_attribute("app.tool.args_summary", summarise_args(call.arguments))  # your redacting function
        try:
            result = TOOLS[call.name](**call.arguments)  # your tool registry
            span.set_attribute("app.tool.status", "ok")
            messages.append({"role": "tool", "tool_call_id": call.id, "content": str(result)})
        except Exception as ex:
            span.set_status(Status(StatusCode.ERROR))
            span.record_exception(ex)
            span.set_attribute("app.tool.status", "error")
            messages.append({"role": "tool", "tool_call_id": call.id, "content": "Tool failed"})
```

### Step 4: Send the traces to a backend you can search

Outcome: Traces from a test run appear in a trace viewer where you can open a run and see its steps.

Console output proves the instrumentation works. For production you need a store you can query. You have two sensible routes. The first is an OpenTelemetry backend: swap the console exporter for the OTLP exporter and point it at your collector or vendor. The OpenTelemetry docs define `OTEL_EXPORTER_OTLP_ENDPOINT`, whose HTTP default is `http://localhost:4318`, and `OTEL_EXPORTER_OTLP_PROTOCOL` with the values `grpc`, `http/protobuf` and `http/json`. `OTEL_EXPORTER_OTLP_HEADERS` carries credentials as key-value pairs.

The second route is Langfuse, an open-source LLM observability platform you can use as a cloud service or self-host. Its Python quickstart is `pip install langfuse`, three environment variables, and a client from `get_client()`. Observations nest with context managers, so a run becomes a span and each model call becomes a generation. Our [Langfuse comparison](https://www.swfte.com/alternatives/langfuse) sets out how it fits next to other tools.

Pick one first and keep the span structure the same either way. If you self-host the backend for data-residency reasons, check where traces are stored, because they hold the most sensitive text in your stack.

OpenTelemetry: OTLP exporter package:

```bash
pip install opentelemetry-exporter-otlp
```

OpenTelemetry: point the exporter at your collector (replace the values):

```bash
export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318"
export OTEL_EXPORTER_OTLP_PROTOCOL="http/protobuf"
```

Langfuse: install and set credentials:

```bash
pip install langfuse
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
```

Langfuse: nested run and generation (from the quickstart):

```python
from langfuse import get_client

langfuse = get_client()

with langfuse.start_as_current_observation(as_type="span", name="process-request") as span:
    span.update(output="Processing complete")

    with langfuse.start_as_current_observation(as_type="generation", name="llm-response", model="gpt-3.5-turbo") as generation:
        generation.update(output="Generated response")

langfuse.flush()
```

> TIP: Call `langfuse.flush()` in short-lived scripts and jobs, otherwise the process can exit before events are sent. Langfuse lists other regions (US, Japan and a HIPAA region) with their own base URLs in its quickstart.

### Step 5: Redact personal and secret data before it is stored

Outcome: Traces carry ids and counts by default, and any content that is stored has been through a redaction step.

Traces collect prompts, retrieved documents and tool arguments, which makes them one of the largest stores of personal data you own. The OpenTelemetry convention says message content is likely to contain sensitive information and is opt-in: instrumentations should not capture it by default. It names three approaches: do not record content, record it on spans, or store it externally and record a reference. Start with the first.

Where you do need content for debugging, redact in your own code before the value reaches the span. Strip email addresses, phone numbers, account numbers and anything that looks like a key or token. Pseudonymise user ids with a keyed hash so you can group by user without storing who they are. Langfuse offers a `mask` option on the client for this, and its docs now mark it as legacy and recommend `mask_otel_spans` for new Python setups, so read the current masking page before you use either.

Set a retention period for trace content that is shorter than for counts and metrics, and limit who can open raw traces. Treat the trace store under your normal data-protection rules, including deletion requests. If you run a data protection impact assessment for the agent, include the trace store in it; [how to do a DPIA for AI](https://www.swfte.com/how-to-do-a-dpia-for-ai) walks through that.

A minimal redaction helper (extend the patterns for your data):

```python
import hashlib
import hmac
import os
import re

EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
LONG_DIGITS = re.compile(r"\b\d{9,}\b")
KEYLIKE = re.compile(r"\b(sk|pk|ghp|xox[bap])[-_][A-Za-z0-9_-]{16,}\b")


def redact(text):
    text = EMAIL.sub("[email]", text)
    text = LONG_DIGITS.sub("[number]", text)
    return KEYLIKE.sub("[secret]", text)


def pseudonym(user_id):
    key = os.environ["TRACE_HASH_KEY"].encode("utf-8")
    return hmac.new(key, user_id.encode("utf-8"), hashlib.sha256).hexdigest()[:16]


def summarise_args(arguments):
    return redact(", ".join(k + "=" + str(v)[:40] for k, v in arguments.items()))
```

> WARNING: Pattern-based redaction misses things. Treat it as a first filter, sample stored traces regularly to see what slipped through, and do not log content at all for sensitive agents.

### Step 6: Set alerts on the signals that mean an agent has gone wrong

Outcome: Five alerts exist, each with an owner and a first action.

Agents fail quietly. They loop, call the wrong tool, drift in quality and run up cost without raising an exception, so error rate alone will not catch them. Alert on these signals instead, and write the first response next to each one so the person paged knows what to do.

Do not copy thresholds from a blog post, including this one. Run the agent for a week, read the distribution of each signal, and set the first threshold just outside normal. Tighten it as you learn. Each alert should route to the agent owner named in step 1.

**Starting alerts for an agent in production**

| Signal | How to compute it | First rule to try | First action |
| --- | --- | --- | --- |
| Ceiling hits | Runs where `app.agent.outcome` is `ceiling_hit` | Any increase over your weekly baseline | Open three traces; look for a repeating tool call |
| Tool failure rate | Tool spans with status `error` / all tool spans, per tool | A tool above its own baseline | Check the downstream system, then the arguments the agent sent |
| Cost per run | Token counts x price, per run and per feature | A run above a multiple of the feature median | Find the longest trace; check for context growth or retries |
| Unexpected tool | Tool name not in the agent's allowed list | Any occurrence | Treat as a policy event; check for prompt injection |
| Success rate | Share of sampled runs labelled success in step 7 | A drop against the previous week | Compare prompts, models and data changes since last week |

### Step 7: Read a random sample of runs every week

Outcome: A named person reads a fixed number of random runs each week and labels them.

Alerts catch what you thought of. A weekly read of random runs catches what you did not. Pick a number you can sustain, for example 20 runs a week, draw them at random from all runs, and add a second draw weighted towards runs with low confidence, high cost or escalations. Read the whole trace, not only the final answer.

Label each run against the success sentence from step 1, and add a short reason when it fails. Record the labels next to the trace so success rate becomes a number you can chart. Over a few weeks the failure reasons will cluster, and each cluster is either a prompt fix, a tool fix, a missing guardrail or a case the agent should hand to a person.

Where a person must approve certain actions, record who approved what and when, on the same trace. [How to set up human approval for AI agents](https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents) covers that design, and [how to govern AI agents](https://www.swfte.com/how-to-govern-ai-agents) covers who owns what.

### Step 8: Write the incident runbook and test the off switch

Outcome: You have a one-page runbook, and you have switched the agent off once in a test.

Monitoring only helps if someone can act. Write a one-page runbook for the agent: who owns it, how to pause it, which tools to revoke first, where the traces are, who to tell, and what to preserve as evidence. The first move for a misbehaving agent is nearly always to stop it taking actions, not to find the cause.

Build the off switch into the agent, not around it: a flag or a permission you can revoke without a deployment. Then test it. Switch the agent off in a staging copy, confirm runs stop and are logged as stopped, and switch it on again. An untested kill switch is a guess.

Our [AI incident response runbook](https://www.swfte.com/blog/ai-incident-response-runbook-2026) and the [SecOps incident response page](https://www.swfte.com/secops/ai-incident-response) set out containment, evidence and review steps in more detail.

## Which signals matter most for which kind of agent

Not every agent needs every dashboard. Match the signals to what the agent can do. An agent that only drafts text for a person to send needs quality sampling more than tool alerts. An agent that can change records needs tool, policy and approval signals first.

| Agent type | Watch first | Why |
| --- | --- | --- |
| Read-only research or summary agent | Success rate from sampled review, cost per run | Harm is wrong or costly answers, not actions |
| Agent that drafts for human approval | Approval rate, edit rate, time to approve | Shows whether people trust and use the drafts |
| Agent that changes records or sends messages | Unexpected tools, tool failures, policy denials, approvals skipped | Errors here have real-world effects |
| Agent that reads untrusted content | Unexpected tools, ceiling hits | Possible prompt injection; see [how to stop prompt injection](https://www.swfte.com/how-to-stop-prompt-injection) |

## What monitoring cannot do

- It does not prove an agent is correct. Traces show what happened, and only evaluation shows whether it was right. Pair this guide with [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai).
- It does not replace permissions. An agent that must not delete records should lack the permission, not trigger an alert after deleting one.
- It can create a new risk. Traces hold sensitive text, so the store needs the same access control and retention rules as production data.
- The OpenTelemetry GenAI conventions are in Development status, so expect attribute renames and pin versions of your instrumentation libraries.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| Nothing prints from the console exporter | The batch processor has not flushed before the process exited, or the tracer provider was set after the tracer was created. | Call `provider.shutdown()` at exit in short scripts, and set the global provider before calling `trace.get_tracer`. |
| Spans appear but model steps are not children of the run span | The model call ran outside the `with` block, or in a thread or async task that did not inherit context. | Create the step span inside the run span block, and pass context explicitly when you start work on another thread. |
| Traces never reach the backend and there is no error | The exporter endpoint or protocol is wrong, or credentials in `OTEL_EXPORTER_OTLP_HEADERS` are missing. | Check `OTEL_EXPORTER_OTLP_ENDPOINT` and `OTEL_EXPORTER_OTLP_PROTOCOL` match what the collector listens on, then test with the console exporter alongside. |
| Langfuse shows nothing from a short script | The process exited before events were sent. | Call `langfuse.flush()` before exit, and check the three environment variables and the region base URL. |
| Trace store is filling with personal data | Message content or raw tool arguments are being recorded. | Switch content capture off, record summaries and ids only, add the redaction helper, and delete the affected data under your retention rules. |
| Alerts fire constantly and get ignored | Thresholds were copied, not measured. | Re-baseline from a week of real data, raise thresholds just outside normal, and delete any alert nobody acts on. |
| Cost per run varies widely for the same task | Context growth across turns, retries after tool errors, or a long tool result being re-sent each turn. | Open the most expensive traces and read the per-step token counts. See [how to reduce LLM costs](https://www.swfte.com/how-to-reduce-llm-costs). |

## Verify it worked

- [ ] A test run prints or exports one parent span for the run, with child spans for each model step and each tool call.
- [ ] Every span for a run carries the same `app.run.id`, and you can search for a run by that id in your backend.
- [ ] A forced tool failure shows an error status and a recorded exception on the tool span.
- [ ] A forced loop (a tool that always fails) ends at the turn ceiling and logs `ceiling_hit`.
- [ ] Stored traces contain no raw email addresses, long account numbers or keys when you sample them.
- [ ] Each of the five alerts has an owner and a written first action, and each has fired once in a test.
- [ ] The weekly review happened at least once and the labelled runs produced a success rate.
- [ ] The agent was switched off and on again in a test, and the stop was visible in the traces.

## Next steps

- [Agent observability overview](https://www.swfte.com/agent-observability): The wider view of what to observe across agents and workflows.
- [LLM observability](https://www.swfte.com/llm-observability): Prompt-level analytics and how tools compare.
- [What to capture: AI observability for agents](https://www.swfte.com/blog/ai-observability-for-agents-and-workflows-2026): The questions your records must be able to answer.
- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): Add evaluation so you know whether runs were right, not only what they did.
- [How to red-team an LLM](https://www.swfte.com/how-to-red-team-an-llm): Test the failures you want your alerts to catch.

## FAQ

### What is AI agent observability?

It is the practice of recording each step an agent takes (model calls, tool calls, retrieved context and decisions) as a trace, so you can reconstruct any run and measure cost, errors and outcomes. It extends LLM observability from single prompts to multi-step work.

### What should I monitor on an AI agent in production?

Track steps per run, ceiling hits, tool call counts and failures, tokens and cost per run, latency, policy decisions and a goal-level success rate from sampled review. Technical health alone misses the quiet failures agents have.

### How do I detect an agent stuck in a loop?

Set a turn ceiling in the agent loop, record when a run hits it, and alert when ceiling hits rise above your baseline. Open a few of those traces and look for the same tool called repeatedly with similar arguments.

### Are the OpenTelemetry GenAI conventions stable?

Not yet. As of 2026-10-06 the agent and tool span pages are marked Development status, so attribute names may change. Use them anyway for consistency, and pin your instrumentation library versions.

### How do I avoid storing personal data in agent traces?

Record ids and counts by default and leave message content off. If you need content, redact in your own code before it reaches a span, pseudonymise user ids, and set a short retention period for trace content.

### What is Langfuse used for?

Langfuse is an open-source platform for tracing and evaluating LLM applications, available as a cloud service or self-hosted. It nests runs and model calls as observations so you can inspect each step. It is one option; any OpenTelemetry backend also works.

### Do I need a vendor tool to monitor agents?

No. You can start with OpenTelemetry and a backend you already run. A dedicated tool adds trace views, evaluation and prompt management, which are useful once you have more than a few agents.

## How Swfte can help

If your agents run on Swfte, Connect and Studio show usage, cost, latency, error breakdown and tool usage per workspace, which covers part of this guide without extra code.

- [AI observability on the platform](https://www.swfte.com/platform/intelligence/ai-observability): What the Swfte products capture today and what is designed.
- [Usage and cost analytics](https://www.swfte.com/platform/intelligence/usage-and-cost-analytics): Spend by model, agent and workflow.
- [SecOps AI incident response](https://www.swfte.com/secops/ai-incident-response): A runbook for when an alert is real.

You can follow every step here without Swfte. Swfte's own issue and struggle detection covers sessions in its web applications, not traces of your agent runs, and per-step cost is not yet joined to every worker trace.

## Sources

- [OpenTelemetry GenAI semantic conventions repository](https://github.com/open-telemetry/semantic-conventions-genai): GenAI conventions now live here; opentelemetry.io pages link to it; folder layout includes agent spans, spans, metrics, events and MCP.
- [GenAI agent and framework spans (raw docs)](https://raw.githubusercontent.com/open-telemetry/semantic-conventions-genai/main/docs/gen-ai/gen-ai-agent-spans.md): Development status, invoke_agent and create_agent operations, required and recommended attributes, content capture marked opt-in and sensitive.
- [GenAI spans (raw docs)](https://raw.githubusercontent.com/open-telemetry/semantic-conventions-genai/main/docs/gen-ai/gen-ai-spans.md): execute_tool span name, INTERNAL kind, gen_ai.tool.name and gen_ai.tool.call.id, content capture off by default and three storage approaches.
- [OpenTelemetry Python: manual instrumentation](https://opentelemetry.io/docs/languages/python/instrumentation/): pip install opentelemetry-api and opentelemetry-sdk, TracerProvider and BatchSpanProcessor setup, start_as_current_span, set_status and record_exception.
- [OpenTelemetry: OTLP exporter configuration](https://opentelemetry.io/docs/languages/sdk-configuration/otlp-exporter/): OTEL_EXPORTER_OTLP_ENDPOINT default, OTEL_EXPORTER_OTLP_PROTOCOL values, OTEL_EXPORTER_OTLP_HEADERS.
- [PyPI: opentelemetry-sdk](https://pypi.org/project/opentelemetry-sdk/): pip install opentelemetry-sdk; version 1.45.1 released 6 October 2026.
- [PyPI: opentelemetry-exporter-otlp](https://pypi.org/project/opentelemetry-exporter-otlp/): pip install opentelemetry-exporter-otlp installs the gRPC and HTTP exporters.
- [Langfuse: get started with observability (Python)](https://langfuse.com/docs/observability/get-started): pip install langfuse, LANGFUSE_* environment variables and regional base URLs, get_client and start_as_current_observation code, flush.
- [Langfuse: masking sensitive data](https://langfuse.com/docs/observability/features/masking): mask option on the client, noted as legacy; mask_otel_spans recommended for new Python setups.
- [Langfuse documentation home](https://langfuse.com/docs): Langfuse described as open source, self-hostable and extensible.

Last verified against these sources on 2026-10-06.
