Operate · Advanced

How to monitor AI agents in production

  • Time: About 4 hours for tracing, redaction and the first alerts; the weekly review is ongoing
  • Cost: The libraries are free. A hosted trace backend may charge by volume; check its current pricing.
  • Level: Advanced
On this page
  1. Short answer
  2. Before you start
  3. 1. Define what one run is and what success means
  4. 2. Wrap each run and model step in OpenTelemetry spans
  5. 3. Record every tool call as its own span, with errors
  6. 4. Send the traces to a backend you can search
  7. 5. Redact personal and secret data before it is stored
  8. 6. Set alerts on the signals that mean an agent has gone wrong
  9. 7. Read a random sample of runs every week
  10. 8. Write the incident runbook and test the off switch
  11. Which signals matter most for which kind of agent
  12. What monitoring cannot do
  13. Troubleshooting
  14. Verify it worked
  15. Next steps
  16. FAQ
  17. How Swfte can help
  18. Sources and last verified

Short answer

Give every agent run an id and record each model call and tool call as a span in a trace, with model, token counts, latency, outcome and any policy decision. Redact personal data before you store it. Alert on loop ceilings, tool failures, cost per run and a falling success rate, and read a random sample of runs by hand every week, because most agent failures raise no error.

The steps at a glance

  1. Define what one run is and what success means
  2. Wrap each run and model step in OpenTelemetry spans
  3. Record every tool call as its own span, with errors
  4. Send the traces to a backend you can search
  5. Redact personal and secret data before it is stored
  6. Set alerts on the signals that mean an agent has gone wrong
  7. Read a random sample of runs every week
  8. Write the incident runbook and test the off switch

Before you start

Who this is for

  • Engineers who run an agent that calls tools and can now act on real systems.
  • Platform and SRE teams asked to add monitoring to an AI feature they did not build.
  • Security and risk owners who need a record of what an agent did and why.

Probably not for you if

Prerequisites

  • An agent that already runs, with a loop you can edit (a model call, tool calls and a stop condition).
  • Python 3.10 or later for the examples. The same ideas apply in other languages.
  • A place to send traces: any OpenTelemetry-compatible backend, or a Langfuse project (cloud or self-hosted).
  • A written definition of a successful run for this agent. If you do not have one, step 1 helps you write it.
Time
About 4 hours for tracing, redaction and the first alerts; the weekly review is ongoing
Cost
The libraries are free. A hosted trace backend may charge by volume; check its current pricing.
Hardware
None beyond where your agent already runs.
Skill
Comfortable with Python or another language, and with reading logs and dashboards

Estimates are ours, not measurements, and move with your hardware, data and network.

  1. Step 1Define what one run is and what success means

    You end up with: You have a written definition of a run, a success test, and a list of fields to capture.

    Monitoring starts with a unit. For an agent, the unit is a run: one goal, from the first model call to the final answer or the point where it gave up. Give every run an id and put that id on every span, log line and tool call that belongs to it. Without it you cannot reconstruct what happened.

    Then define success at the goal level, not the technical level. "The ticket was resolved and the customer did not reopen it" is a success test. "The API returned 200" is not. Write one sentence per agent. You will use it for sampling in step 7 and for alert thresholds in step 6.

    Decide what to capture before you write code. The table below is a starting set. Capture identifiers and counts by default, and treat message content as opt-in, which is also how the OpenTelemetry convention treats it.

    What to record for each agent run
    FieldWhy you want it
    Run id, agent name, version, ownerFind a run, and know who to call
    User or tenant id (pseudonymised)Spot one user causing most of the load
    Each model call: model, input and output tokens, latencyCost and slowness per step
    Each tool call: tool name, arguments summary, result status, latencyTool failures, unexpected tools, forbidden actions
    Turn count and whether the ceiling was hitLoops
    Policy decisions: allowed, denied, escalated, approved by whomEvidence for governance and audits
    Outcome label: success, failure, escalated, abandonedSuccess rate over time
    Cost per runSpend per run and per feature
  2. Step 2Wrap each run and model step in OpenTelemetry spans

    You end up with: Running the agent prints a trace with one parent span for the run and child spans for model steps.

    OpenTelemetry is a vendor-neutral standard for traces, so what you record can go to any compatible backend. Install the API and SDK, create a tracer provider, and wrap the run in a span. The console exporter prints finished spans to the terminal, which is the quickest way to check what you are recording before you add a backend.

    For attribute names, follow the OpenTelemetry GenAI semantic conventions. As of 2026-10-06 they live in the semantic-conventions-genai repository, and the agent spans page carries Development status, which means names can still change. The agent span uses gen_ai.operation.name set to invoke_agent, plus gen_ai.provider.name, and gen_ai.agent.name when available. Token counts are recommended as gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Put your own attributes under a prefix you control, such as app., so they cannot collide with the spec.

    The skeleton below is the shape, not a framework. call_model and run_tool are your own functions: call_model should return the reply and the token counts, and run_tool should execute one tool call. Set a turn ceiling and record when you hit it. That single attribute is the basis of your loop alert later.

    Install the OpenTelemetry API and SDK · bash
    pip install opentelemetry-api
    pip install opentelemetry-sdk
    Tracer setup and an instrumented agent loop (skeleton) · python
    import uuid
    
    from opentelemetry import trace
    from opentelemetry.sdk.trace import TracerProvider
    from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter
    from opentelemetry.trace import Status, StatusCode
    
    provider = TracerProvider()
    provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))
    trace.set_tracer_provider(provider)
    tracer = trace.get_tracer("support-agent")
    
    MAX_TURNS = 12
    
    
    def run_agent(goal, user_pseudonym):
        run_id = str(uuid.uuid4())
        total_in = total_out = 0
        with tracer.start_as_current_span("invoke_agent support-agent") as run_span:
            run_span.set_attribute("gen_ai.operation.name", "invoke_agent")
            run_span.set_attribute("gen_ai.provider.name", "anthropic")  # use the value for your provider
            run_span.set_attribute("gen_ai.agent.name", "support-agent")
            run_span.set_attribute("app.run.id", run_id)
            run_span.set_attribute("app.user.pseudonym", user_pseudonym)
            run_span.set_attribute("app.agent.version", "2026-10-06")
            messages = [{"role": "user", "content": goal}]
            turns = 0
            outcome = "ceiling_hit"
            while turns < MAX_TURNS:
                turns += 1
                with tracer.start_as_current_span("model_step") as step:
                    reply, tokens_in, tokens_out = call_model(messages)  # your function
                    total_in += tokens_in
                    total_out += tokens_out
                    step.set_attribute("gen_ai.usage.input_tokens", tokens_in)
                    step.set_attribute("gen_ai.usage.output_tokens", tokens_out)
                if reply.tool_calls:
                    for call in reply.tool_calls:
                        run_tool(call, messages)  # step 3 wraps this in a span
                else:
                    outcome = "answered"
                    break
            run_span.set_attribute("gen_ai.usage.input_tokens", total_in)
            run_span.set_attribute("gen_ai.usage.output_tokens", total_out)
            run_span.set_attribute("app.agent.turns", turns)
            run_span.set_attribute("app.agent.outcome", outcome)
            return reply

    Checked against: OpenTelemetry GenAI semantic conventions repository, GenAI agent and framework spans (raw docs), OpenTelemetry Python: manual instrumentation, PyPI: opentelemetry-sdk

  3. Step 3Record every tool call as its own span, with errors

    You end up with: Each tool call appears as a child span with its name, status and any exception.

    Tools are where an agent touches the outside world, so they deserve their own spans. The GenAI conventions define an execute_tool operation: span name execute_tool followed by the tool name, kind INTERNAL, with gen_ai.operation.name set to execute_tool and gen_ai.tool.name required. gen_ai.tool.call.id is recommended when you have it.

    When a tool fails, set the span status to error and record the exception, as the OpenTelemetry Python guide shows. Do not swallow the failure and return text to the model without leaving a trace. A tool that fails quietly and gets retried five times is one of the most common causes of loops and cost spikes.

    Record the arguments as a short summary, not the raw values. Arguments often contain customer data. If you need the full detail for debugging, store it separately with a short retention period and reference it from the span.

    Tool-call span with error recording · python
    def run_tool(call, messages):
        with tracer.start_as_current_span("execute_tool " + call.name) as span:
            span.set_attribute("gen_ai.operation.name", "execute_tool")
            span.set_attribute("gen_ai.tool.name", call.name)
            span.set_attribute("gen_ai.tool.call.id", call.id)
            span.set_attribute("app.tool.args_summary", summarise_args(call.arguments))  # your redacting function
            try:
                result = TOOLS[call.name](**call.arguments)  # your tool registry
                span.set_attribute("app.tool.status", "ok")
                messages.append({"role": "tool", "tool_call_id": call.id, "content": str(result)})
            except Exception as ex:
                span.set_status(Status(StatusCode.ERROR))
                span.record_exception(ex)
                span.set_attribute("app.tool.status", "error")
                messages.append({"role": "tool", "tool_call_id": call.id, "content": "Tool failed"})

    Checked against: GenAI spans (raw docs), OpenTelemetry Python: manual instrumentation

  4. Step 4Send the traces to a backend you can search

    You end up with: Traces from a test run appear in a trace viewer where you can open a run and see its steps.

    Console output proves the instrumentation works. For production you need a store you can query. You have two sensible routes. The first is an OpenTelemetry backend: swap the console exporter for the OTLP exporter and point it at your collector or vendor. The OpenTelemetry docs define OTEL_EXPORTER_OTLP_ENDPOINT, whose HTTP default is http://localhost:4318, and OTEL_EXPORTER_OTLP_PROTOCOL with the values grpc, http/protobuf and http/json. OTEL_EXPORTER_OTLP_HEADERS carries credentials as key-value pairs.

    The second route is Langfuse, an open-source LLM observability platform you can use as a cloud service or self-host. Its Python quickstart is pip install langfuse, three environment variables, and a client from get_client(). Observations nest with context managers, so a run becomes a span and each model call becomes a generation. Our Langfuse comparison sets out how it fits next to other tools.

    Pick one first and keep the span structure the same either way. If you self-host the backend for data-residency reasons, check where traces are stored, because they hold the most sensitive text in your stack.

    OpenTelemetry: OTLP exporter package · bash
    pip install opentelemetry-exporter-otlp
    OpenTelemetry: point the exporter at your collector (replace the values) · bash
    export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318"
    export OTEL_EXPORTER_OTLP_PROTOCOL="http/protobuf"
    Langfuse: install and set credentials · bash
    pip install langfuse
    export LANGFUSE_PUBLIC_KEY="pk-lf-..."
    export LANGFUSE_SECRET_KEY="sk-lf-..."
    export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
    Langfuse: nested run and generation (from the quickstart) · python
    from langfuse import get_client
    
    langfuse = get_client()
    
    with langfuse.start_as_current_observation(as_type="span", name="process-request") as span:
        span.update(output="Processing complete")
    
        with langfuse.start_as_current_observation(as_type="generation", name="llm-response", model="gpt-3.5-turbo") as generation:
            generation.update(output="Generated response")
    
    langfuse.flush()

    Checked against: OpenTelemetry: OTLP exporter configuration, PyPI: opentelemetry-exporter-otlp, Langfuse: get started with observability (Python), Langfuse documentation home

  5. Step 5Redact personal and secret data before it is stored

    You end up with: Traces carry ids and counts by default, and any content that is stored has been through a redaction step.

    Traces collect prompts, retrieved documents and tool arguments, which makes them one of the largest stores of personal data you own. The OpenTelemetry convention says message content is likely to contain sensitive information and is opt-in: instrumentations should not capture it by default. It names three approaches: do not record content, record it on spans, or store it externally and record a reference. Start with the first.

    Where you do need content for debugging, redact in your own code before the value reaches the span. Strip email addresses, phone numbers, account numbers and anything that looks like a key or token. Pseudonymise user ids with a keyed hash so you can group by user without storing who they are. Langfuse offers a mask option on the client for this, and its docs now mark it as legacy and recommend mask_otel_spans for new Python setups, so read the current masking page before you use either.

    Set a retention period for trace content that is shorter than for counts and metrics, and limit who can open raw traces. Treat the trace store under your normal data-protection rules, including deletion requests. If you run a data protection impact assessment for the agent, include the trace store in it; how to do a DPIA for AI walks through that.

    A minimal redaction helper (extend the patterns for your data) · python
    import hashlib
    import hmac
    import os
    import re
    
    EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
    LONG_DIGITS = re.compile(r"\b\d{9,}\b")
    KEYLIKE = re.compile(r"\b(sk|pk|ghp|xox[bap])[-_][A-Za-z0-9_-]{16,}\b")
    
    
    def redact(text):
        text = EMAIL.sub("[email]", text)
        text = LONG_DIGITS.sub("[number]", text)
        return KEYLIKE.sub("[secret]", text)
    
    
    def pseudonym(user_id):
        key = os.environ["TRACE_HASH_KEY"].encode("utf-8")
        return hmac.new(key, user_id.encode("utf-8"), hashlib.sha256).hexdigest()[:16]
    
    
    def summarise_args(arguments):
        return redact(", ".join(k + "=" + str(v)[:40] for k, v in arguments.items()))

    Checked against: GenAI spans (raw docs), GenAI agent and framework spans (raw docs), Langfuse: masking sensitive data

  6. Step 6Set alerts on the signals that mean an agent has gone wrong

    You end up with: Five alerts exist, each with an owner and a first action.

    Agents fail quietly. They loop, call the wrong tool, drift in quality and run up cost without raising an exception, so error rate alone will not catch them. Alert on these signals instead, and write the first response next to each one so the person paged knows what to do.

    Do not copy thresholds from a blog post, including this one. Run the agent for a week, read the distribution of each signal, and set the first threshold just outside normal. Tighten it as you learn. Each alert should route to the agent owner named in step 1.

    Starting alerts for an agent in production
    SignalHow to compute itFirst rule to tryFirst action
    Ceiling hitsRuns where app.agent.outcome is ceiling_hitAny increase over your weekly baselineOpen three traces; look for a repeating tool call
    Tool failure rateTool spans with status error / all tool spans, per toolA tool above its own baselineCheck the downstream system, then the arguments the agent sent
    Cost per runToken counts x price, per run and per featureA run above a multiple of the feature medianFind the longest trace; check for context growth or retries
    Unexpected toolTool name not in the agent's allowed listAny occurrenceTreat as a policy event; check for prompt injection
    Success rateShare of sampled runs labelled success in step 7A drop against the previous weekCompare prompts, models and data changes since last week
  7. Step 7Read a random sample of runs every week

    You end up with: A named person reads a fixed number of random runs each week and labels them.

    Alerts catch what you thought of. A weekly read of random runs catches what you did not. Pick a number you can sustain, for example 20 runs a week, draw them at random from all runs, and add a second draw weighted towards runs with low confidence, high cost or escalations. Read the whole trace, not only the final answer.

    Label each run against the success sentence from step 1, and add a short reason when it fails. Record the labels next to the trace so success rate becomes a number you can chart. Over a few weeks the failure reasons will cluster, and each cluster is either a prompt fix, a tool fix, a missing guardrail or a case the agent should hand to a person.

    Where a person must approve certain actions, record who approved what and when, on the same trace. How to set up human approval for AI agents covers that design, and how to govern AI agents covers who owns what.

  8. Step 8Write the incident runbook and test the off switch

    You end up with: You have a one-page runbook, and you have switched the agent off once in a test.

    Monitoring only helps if someone can act. Write a one-page runbook for the agent: who owns it, how to pause it, which tools to revoke first, where the traces are, who to tell, and what to preserve as evidence. The first move for a misbehaving agent is nearly always to stop it taking actions, not to find the cause.

    Build the off switch into the agent, not around it: a flag or a permission you can revoke without a deployment. Then test it. Switch the agent off in a staging copy, confirm runs stop and are logged as stopped, and switch it on again. An untested kill switch is a guess.

    Our AI incident response runbook and the SecOps incident response page set out containment, evidence and review steps in more detail.

Which signals matter most for which kind of agent

Not every agent needs every dashboard. Match the signals to what the agent can do. An agent that only drafts text for a person to send needs quality sampling more than tool alerts. An agent that can change records needs tool, policy and approval signals first.

Agent typeWatch firstWhy
Read-only research or summary agentSuccess rate from sampled review, cost per runHarm is wrong or costly answers, not actions
Agent that drafts for human approvalApproval rate, edit rate, time to approveShows whether people trust and use the drafts
Agent that changes records or sends messagesUnexpected tools, tool failures, policy denials, approvals skippedErrors here have real-world effects
Agent that reads untrusted contentUnexpected tools, ceiling hitsPossible prompt injection; see how to stop prompt injection

What monitoring cannot do

  • It does not prove an agent is correct. Traces show what happened, and only evaluation shows whether it was right. Pair this guide with how to validate your AI.
  • It does not replace permissions. An agent that must not delete records should lack the permission, not trigger an alert after deleting one.
  • It can create a new risk. Traces hold sensitive text, so the store needs the same access control and retention rules as production data.
  • The OpenTelemetry GenAI conventions are in Development status, so expect attribute renames and pin versions of your instrumentation libraries.

Troubleshooting

What you seeLikely causeFix
Nothing prints from the console exporterThe batch processor has not flushed before the process exited, or the tracer provider was set after the tracer was created.Call provider.shutdown() at exit in short scripts, and set the global provider before calling trace.get_tracer.
Spans appear but model steps are not children of the run spanThe model call ran outside the with block, or in a thread or async task that did not inherit context.Create the step span inside the run span block, and pass context explicitly when you start work on another thread.
Traces never reach the backend and there is no errorThe exporter endpoint or protocol is wrong, or credentials in OTEL_EXPORTER_OTLP_HEADERS are missing.Check OTEL_EXPORTER_OTLP_ENDPOINT and OTEL_EXPORTER_OTLP_PROTOCOL match what the collector listens on, then test with the console exporter alongside.
Langfuse shows nothing from a short scriptThe process exited before events were sent.Call langfuse.flush() before exit, and check the three environment variables and the region base URL.
Trace store is filling with personal dataMessage content or raw tool arguments are being recorded.Switch content capture off, record summaries and ids only, add the redaction helper, and delete the affected data under your retention rules.
Alerts fire constantly and get ignoredThresholds were copied, not measured.Re-baseline from a week of real data, raise thresholds just outside normal, and delete any alert nobody acts on.
Cost per run varies widely for the same taskContext growth across turns, retries after tool errors, or a long tool result being re-sent each turn.Open the most expensive traces and read the per-step token counts. See how to reduce LLM costs.

Verify it worked

Next steps

Related guides

  • How to Validate Your AI: Eval Sets, Gates, Evidence: A system-level method to validate an AI product: define the task and risk, build a held-out eval set, score it, gate releases, sample for human review, monitor and keep an evidence pack.
  • How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
  • How to Set Up Human Approval for AI Agents (With Code): Decide which agent actions need a person, set thresholds, pause the agent with LangGraph interrupts, route requests to a queue with a timeout that denies by default, show reviewers the evidence, and record every decision.
  • How to Reduce LLM Costs: Step-by-Step Guide (2026): Measure token spend per feature first, then apply the levers in order of payoff: output caps, prompt caching, batch APIs, model routing, response caching, and a self-hosting break-even check.
  • How to Red Team an LLM App: Tools, Scoring, Retest: A step-by-step LLM red-team exercise: authorisation and scope, a threat model mapped to the OWASP LLM Top 10, automated testing with promptfoo, garak and PyRIT, manual attack sessions, scoring, fixes and retest.

Frequently asked questions

What is AI agent observability?

It is the practice of recording each step an agent takes (model calls, tool calls, retrieved context and decisions) as a trace, so you can reconstruct any run and measure cost, errors and outcomes. It extends LLM observability from single prompts to multi-step work.

What should I monitor on an AI agent in production?

Track steps per run, ceiling hits, tool call counts and failures, tokens and cost per run, latency, policy decisions and a goal-level success rate from sampled review. Technical health alone misses the quiet failures agents have.

How do I detect an agent stuck in a loop?

Set a turn ceiling in the agent loop, record when a run hits it, and alert when ceiling hits rise above your baseline. Open a few of those traces and look for the same tool called repeatedly with similar arguments.

Are the OpenTelemetry GenAI conventions stable?

Not yet. As of 2026-10-06 the agent and tool span pages are marked Development status, so attribute names may change. Use them anyway for consistency, and pin your instrumentation library versions.

How do I avoid storing personal data in agent traces?

Record ids and counts by default and leave message content off. If you need content, redact in your own code before it reaches a span, pseudonymise user ids, and set a short retention period for trace content.

What is Langfuse used for?

Langfuse is an open-source platform for tracing and evaluating LLM applications, available as a cloud service or self-hosted. It nests runs and model calls as observations so you can inspect each step. It is one option; any OpenTelemetry backend also works.

Do I need a vendor tool to monitor agents?

No. You can start with OpenTelemetry and a backend you already run. A dedicated tool adds trace views, evaluation and prompt management, which are useful once you have more than a few agents.

How Swfte can help

If your agents run on Swfte, Connect and Studio show usage, cost, latency, error breakdown and tool usage per workspace, which covers part of this guide without extra code.

You can follow every step here without Swfte. Swfte's own issue and struggle detection covers sessions in its web applications, not traces of your agent runs, and per-step cost is not yet joined to every worker trace.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. OpenTelemetry GenAI semantic conventions repository: GenAI conventions now live here; opentelemetry.io pages link to it; folder layout includes agent spans, spans, metrics, events and MCP.
  2. GenAI agent and framework spans (raw docs): Development status, invoke_agent and create_agent operations, required and recommended attributes, content capture marked opt-in and sensitive.
  3. GenAI spans (raw docs): execute_tool span name, INTERNAL kind, gen_ai.tool.name and gen_ai.tool.call.id, content capture off by default and three storage approaches.
  4. OpenTelemetry Python: manual instrumentation: pip install opentelemetry-api and opentelemetry-sdk, TracerProvider and BatchSpanProcessor setup, start_as_current_span, set_status and record_exception.
  5. OpenTelemetry: OTLP exporter configuration: OTEL_EXPORTER_OTLP_ENDPOINT default, OTEL_EXPORTER_OTLP_PROTOCOL values, OTEL_EXPORTER_OTLP_HEADERS.
  6. PyPI: opentelemetry-sdk: pip install opentelemetry-sdk; version 1.45.1 released 6 October 2026.
  7. PyPI: opentelemetry-exporter-otlp: pip install opentelemetry-exporter-otlp installs the gRPC and HTTP exporters.
  8. Langfuse: get started with observability (Python): pip install langfuse, LANGFUSE_* environment variables and regional base URLs, get_client and start_as_current_observation code, flush.
  9. Langfuse: masking sensitive data: mask option on the client, noted as legacy; mask_otel_spans recommended for new Python setups.
  10. Langfuse documentation home: Langfuse described as open source, self-hostable and extensible.

Topics

  • observability
  • tracing
  • OpenTelemetry
  • Langfuse
  • alerts

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-monitor-ai-agents-in-production.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.