Short answer
Give every agent run an id and record each model call and tool call as a span in a trace, with model, token counts, latency, outcome and any policy decision. Redact personal data before you store it. Alert on loop ceilings, tool failures, cost per run and a falling success rate, and read a random sample of runs by hand every week, because most agent failures raise no error.
The steps at a glance
- Define what one run is and what success means
- Wrap each run and model step in OpenTelemetry spans
- Record every tool call as its own span, with errors
- Send the traces to a backend you can search
- Redact personal and secret data before it is stored
- Set alerts on the signals that mean an agent has gone wrong
- Read a random sample of runs every week
- Write the incident runbook and test the off switch
Before you start
Who this is for
- Engineers who run an agent that calls tools and can now act on real systems.
- Platform and SRE teams asked to add monitoring to an AI feature they did not build.
- Security and risk owners who need a record of what an agent did and why.
Probably not for you if
- Teams still building their first agent. Start with how to build an AI agent.
- Anyone looking for a model benchmark. Monitoring tells you what happened in production, not which model is best. See how to validate your AI.
Prerequisites
- An agent that already runs, with a loop you can edit (a model call, tool calls and a stop condition).
- Python 3.10 or later for the examples. The same ideas apply in other languages.
- A place to send traces: any OpenTelemetry-compatible backend, or a Langfuse project (cloud or self-hosted).
- A written definition of a successful run for this agent. If you do not have one, step 1 helps you write it.
- Time
- About 4 hours for tracing, redaction and the first alerts; the weekly review is ongoing
- Cost
- The libraries are free. A hosted trace backend may charge by volume; check its current pricing.
- Hardware
- None beyond where your agent already runs.
- Skill
- Comfortable with Python or another language, and with reading logs and dashboards
Estimates are ours, not measurements, and move with your hardware, data and network.
Step 1Define what one run is and what success means
You end up with: You have a written definition of a run, a success test, and a list of fields to capture.
Monitoring starts with a unit. For an agent, the unit is a run: one goal, from the first model call to the final answer or the point where it gave up. Give every run an id and put that id on every span, log line and tool call that belongs to it. Without it you cannot reconstruct what happened.
Then define success at the goal level, not the technical level. "The ticket was resolved and the customer did not reopen it" is a success test. "The API returned 200" is not. Write one sentence per agent. You will use it for sampling in step 7 and for alert thresholds in step 6.
Decide what to capture before you write code. The table below is a starting set. Capture identifiers and counts by default, and treat message content as opt-in, which is also how the OpenTelemetry convention treats it.
What to record for each agent run Field Why you want it Run id, agent name, version, owner Find a run, and know who to call User or tenant id (pseudonymised) Spot one user causing most of the load Each model call: model, input and output tokens, latency Cost and slowness per step Each tool call: tool name, arguments summary, result status, latency Tool failures, unexpected tools, forbidden actions Turn count and whether the ceiling was hit Loops Policy decisions: allowed, denied, escalated, approved by whom Evidence for governance and audits Outcome label: success, failure, escalated, abandoned Success rate over time Cost per run Spend per run and per feature Step 2Wrap each run and model step in OpenTelemetry spans
You end up with: Running the agent prints a trace with one parent span for the run and child spans for model steps.
OpenTelemetry is a vendor-neutral standard for traces, so what you record can go to any compatible backend. Install the API and SDK, create a tracer provider, and wrap the run in a span. The console exporter prints finished spans to the terminal, which is the quickest way to check what you are recording before you add a backend.
For attribute names, follow the OpenTelemetry GenAI semantic conventions. As of 2026-10-06 they live in the semantic-conventions-genai repository, and the agent spans page carries Development status, which means names can still change. The agent span uses
gen_ai.operation.nameset toinvoke_agent, plusgen_ai.provider.name, andgen_ai.agent.namewhen available. Token counts are recommended asgen_ai.usage.input_tokensandgen_ai.usage.output_tokens. Put your own attributes under a prefix you control, such asapp., so they cannot collide with the spec.The skeleton below is the shape, not a framework.
call_modelandrun_toolare your own functions:call_modelshould return the reply and the token counts, andrun_toolshould execute one tool call. Set a turn ceiling and record when you hit it. That single attribute is the basis of your loop alert later.Install the OpenTelemetry API and SDK · bash pip install opentelemetry-api pip install opentelemetry-sdkTracer setup and an instrumented agent loop (skeleton) · python import uuid from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter from opentelemetry.trace import Status, StatusCode provider = TracerProvider() provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter())) trace.set_tracer_provider(provider) tracer = trace.get_tracer("support-agent") MAX_TURNS = 12 def run_agent(goal, user_pseudonym): run_id = str(uuid.uuid4()) total_in = total_out = 0 with tracer.start_as_current_span("invoke_agent support-agent") as run_span: run_span.set_attribute("gen_ai.operation.name", "invoke_agent") run_span.set_attribute("gen_ai.provider.name", "anthropic") # use the value for your provider run_span.set_attribute("gen_ai.agent.name", "support-agent") run_span.set_attribute("app.run.id", run_id) run_span.set_attribute("app.user.pseudonym", user_pseudonym) run_span.set_attribute("app.agent.version", "2026-10-06") messages = [{"role": "user", "content": goal}] turns = 0 outcome = "ceiling_hit" while turns < MAX_TURNS: turns += 1 with tracer.start_as_current_span("model_step") as step: reply, tokens_in, tokens_out = call_model(messages) # your function total_in += tokens_in total_out += tokens_out step.set_attribute("gen_ai.usage.input_tokens", tokens_in) step.set_attribute("gen_ai.usage.output_tokens", tokens_out) if reply.tool_calls: for call in reply.tool_calls: run_tool(call, messages) # step 3 wraps this in a span else: outcome = "answered" break run_span.set_attribute("gen_ai.usage.input_tokens", total_in) run_span.set_attribute("gen_ai.usage.output_tokens", total_out) run_span.set_attribute("app.agent.turns", turns) run_span.set_attribute("app.agent.outcome", outcome) return replyChecked against: OpenTelemetry GenAI semantic conventions repository, GenAI agent and framework spans (raw docs), OpenTelemetry Python: manual instrumentation, PyPI: opentelemetry-sdk
Step 3Record every tool call as its own span, with errors
You end up with: Each tool call appears as a child span with its name, status and any exception.
Tools are where an agent touches the outside world, so they deserve their own spans. The GenAI conventions define an
execute_tooloperation: span nameexecute_toolfollowed by the tool name, kindINTERNAL, withgen_ai.operation.nameset toexecute_toolandgen_ai.tool.namerequired.gen_ai.tool.call.idis recommended when you have it.When a tool fails, set the span status to error and record the exception, as the OpenTelemetry Python guide shows. Do not swallow the failure and return text to the model without leaving a trace. A tool that fails quietly and gets retried five times is one of the most common causes of loops and cost spikes.
Record the arguments as a short summary, not the raw values. Arguments often contain customer data. If you need the full detail for debugging, store it separately with a short retention period and reference it from the span.
Tool-call span with error recording · python def run_tool(call, messages): with tracer.start_as_current_span("execute_tool " + call.name) as span: span.set_attribute("gen_ai.operation.name", "execute_tool") span.set_attribute("gen_ai.tool.name", call.name) span.set_attribute("gen_ai.tool.call.id", call.id) span.set_attribute("app.tool.args_summary", summarise_args(call.arguments)) # your redacting function try: result = TOOLS[call.name](**call.arguments) # your tool registry span.set_attribute("app.tool.status", "ok") messages.append({"role": "tool", "tool_call_id": call.id, "content": str(result)}) except Exception as ex: span.set_status(Status(StatusCode.ERROR)) span.record_exception(ex) span.set_attribute("app.tool.status", "error") messages.append({"role": "tool", "tool_call_id": call.id, "content": "Tool failed"})Checked against: GenAI spans (raw docs), OpenTelemetry Python: manual instrumentation
Step 4Send the traces to a backend you can search
You end up with: Traces from a test run appear in a trace viewer where you can open a run and see its steps.
Console output proves the instrumentation works. For production you need a store you can query. You have two sensible routes. The first is an OpenTelemetry backend: swap the console exporter for the OTLP exporter and point it at your collector or vendor. The OpenTelemetry docs define
OTEL_EXPORTER_OTLP_ENDPOINT, whose HTTP default ishttp://localhost:4318, andOTEL_EXPORTER_OTLP_PROTOCOLwith the valuesgrpc,http/protobufandhttp/json.OTEL_EXPORTER_OTLP_HEADERScarries credentials as key-value pairs.The second route is Langfuse, an open-source LLM observability platform you can use as a cloud service or self-host. Its Python quickstart is
pip install langfuse, three environment variables, and a client fromget_client(). Observations nest with context managers, so a run becomes a span and each model call becomes a generation. Our Langfuse comparison sets out how it fits next to other tools.Pick one first and keep the span structure the same either way. If you self-host the backend for data-residency reasons, check where traces are stored, because they hold the most sensitive text in your stack.
OpenTelemetry: OTLP exporter package · bash pip install opentelemetry-exporter-otlpOpenTelemetry: point the exporter at your collector (replace the values) · bash export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318" export OTEL_EXPORTER_OTLP_PROTOCOL="http/protobuf"Langfuse: install and set credentials · bash pip install langfuse export LANGFUSE_PUBLIC_KEY="pk-lf-..." export LANGFUSE_SECRET_KEY="sk-lf-..." export LANGFUSE_BASE_URL="https://cloud.langfuse.com"Langfuse: nested run and generation (from the quickstart) · python from langfuse import get_client langfuse = get_client() with langfuse.start_as_current_observation(as_type="span", name="process-request") as span: span.update(output="Processing complete") with langfuse.start_as_current_observation(as_type="generation", name="llm-response", model="gpt-3.5-turbo") as generation: generation.update(output="Generated response") langfuse.flush()Checked against: OpenTelemetry: OTLP exporter configuration, PyPI: opentelemetry-exporter-otlp, Langfuse: get started with observability (Python), Langfuse documentation home
Step 5Redact personal and secret data before it is stored
You end up with: Traces carry ids and counts by default, and any content that is stored has been through a redaction step.
Traces collect prompts, retrieved documents and tool arguments, which makes them one of the largest stores of personal data you own. The OpenTelemetry convention says message content is likely to contain sensitive information and is opt-in: instrumentations should not capture it by default. It names three approaches: do not record content, record it on spans, or store it externally and record a reference. Start with the first.
Where you do need content for debugging, redact in your own code before the value reaches the span. Strip email addresses, phone numbers, account numbers and anything that looks like a key or token. Pseudonymise user ids with a keyed hash so you can group by user without storing who they are. Langfuse offers a
maskoption on the client for this, and its docs now mark it as legacy and recommendmask_otel_spansfor new Python setups, so read the current masking page before you use either.Set a retention period for trace content that is shorter than for counts and metrics, and limit who can open raw traces. Treat the trace store under your normal data-protection rules, including deletion requests. If you run a data protection impact assessment for the agent, include the trace store in it; how to do a DPIA for AI walks through that.
A minimal redaction helper (extend the patterns for your data) · python import hashlib import hmac import os import re EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+") LONG_DIGITS = re.compile(r"\b\d{9,}\b") KEYLIKE = re.compile(r"\b(sk|pk|ghp|xox[bap])[-_][A-Za-z0-9_-]{16,}\b") def redact(text): text = EMAIL.sub("[email]", text) text = LONG_DIGITS.sub("[number]", text) return KEYLIKE.sub("[secret]", text) def pseudonym(user_id): key = os.environ["TRACE_HASH_KEY"].encode("utf-8") return hmac.new(key, user_id.encode("utf-8"), hashlib.sha256).hexdigest()[:16] def summarise_args(arguments): return redact(", ".join(k + "=" + str(v)[:40] for k, v in arguments.items()))Checked against: GenAI spans (raw docs), GenAI agent and framework spans (raw docs), Langfuse: masking sensitive data
Step 6Set alerts on the signals that mean an agent has gone wrong
You end up with: Five alerts exist, each with an owner and a first action.
Agents fail quietly. They loop, call the wrong tool, drift in quality and run up cost without raising an exception, so error rate alone will not catch them. Alert on these signals instead, and write the first response next to each one so the person paged knows what to do.
Do not copy thresholds from a blog post, including this one. Run the agent for a week, read the distribution of each signal, and set the first threshold just outside normal. Tighten it as you learn. Each alert should route to the agent owner named in step 1.
Starting alerts for an agent in production Signal How to compute it First rule to try First action Ceiling hits Runs where app.agent.outcomeisceiling_hitAny increase over your weekly baseline Open three traces; look for a repeating tool call Tool failure rate Tool spans with status error/ all tool spans, per toolA tool above its own baseline Check the downstream system, then the arguments the agent sent Cost per run Token counts x price, per run and per feature A run above a multiple of the feature median Find the longest trace; check for context growth or retries Unexpected tool Tool name not in the agent's allowed list Any occurrence Treat as a policy event; check for prompt injection Success rate Share of sampled runs labelled success in step 7 A drop against the previous week Compare prompts, models and data changes since last week Step 7Read a random sample of runs every week
You end up with: A named person reads a fixed number of random runs each week and labels them.
Alerts catch what you thought of. A weekly read of random runs catches what you did not. Pick a number you can sustain, for example 20 runs a week, draw them at random from all runs, and add a second draw weighted towards runs with low confidence, high cost or escalations. Read the whole trace, not only the final answer.
Label each run against the success sentence from step 1, and add a short reason when it fails. Record the labels next to the trace so success rate becomes a number you can chart. Over a few weeks the failure reasons will cluster, and each cluster is either a prompt fix, a tool fix, a missing guardrail or a case the agent should hand to a person.
Where a person must approve certain actions, record who approved what and when, on the same trace. How to set up human approval for AI agents covers that design, and how to govern AI agents covers who owns what.
Step 8Write the incident runbook and test the off switch
You end up with: You have a one-page runbook, and you have switched the agent off once in a test.
Monitoring only helps if someone can act. Write a one-page runbook for the agent: who owns it, how to pause it, which tools to revoke first, where the traces are, who to tell, and what to preserve as evidence. The first move for a misbehaving agent is nearly always to stop it taking actions, not to find the cause.
Build the off switch into the agent, not around it: a flag or a permission you can revoke without a deployment. Then test it. Switch the agent off in a staging copy, confirm runs stop and are logged as stopped, and switch it on again. An untested kill switch is a guess.
Our AI incident response runbook and the SecOps incident response page set out containment, evidence and review steps in more detail.
Which signals matter most for which kind of agent
Not every agent needs every dashboard. Match the signals to what the agent can do. An agent that only drafts text for a person to send needs quality sampling more than tool alerts. An agent that can change records needs tool, policy and approval signals first.
| Agent type | Watch first | Why |
|---|---|---|
| Read-only research or summary agent | Success rate from sampled review, cost per run | Harm is wrong or costly answers, not actions |
| Agent that drafts for human approval | Approval rate, edit rate, time to approve | Shows whether people trust and use the drafts |
| Agent that changes records or sends messages | Unexpected tools, tool failures, policy denials, approvals skipped | Errors here have real-world effects |
| Agent that reads untrusted content | Unexpected tools, ceiling hits | Possible prompt injection; see how to stop prompt injection |
What monitoring cannot do
- It does not prove an agent is correct. Traces show what happened, and only evaluation shows whether it was right. Pair this guide with how to validate your AI.
- It does not replace permissions. An agent that must not delete records should lack the permission, not trigger an alert after deleting one.
- It can create a new risk. Traces hold sensitive text, so the store needs the same access control and retention rules as production data.
- The OpenTelemetry GenAI conventions are in Development status, so expect attribute renames and pin versions of your instrumentation libraries.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| Nothing prints from the console exporter | The batch processor has not flushed before the process exited, or the tracer provider was set after the tracer was created. | Call provider.shutdown() at exit in short scripts, and set the global provider before calling trace.get_tracer. |
| Spans appear but model steps are not children of the run span | The model call ran outside the with block, or in a thread or async task that did not inherit context. | Create the step span inside the run span block, and pass context explicitly when you start work on another thread. |
| Traces never reach the backend and there is no error | The exporter endpoint or protocol is wrong, or credentials in OTEL_EXPORTER_OTLP_HEADERS are missing. | Check OTEL_EXPORTER_OTLP_ENDPOINT and OTEL_EXPORTER_OTLP_PROTOCOL match what the collector listens on, then test with the console exporter alongside. |
| Langfuse shows nothing from a short script | The process exited before events were sent. | Call langfuse.flush() before exit, and check the three environment variables and the region base URL. |
| Trace store is filling with personal data | Message content or raw tool arguments are being recorded. | Switch content capture off, record summaries and ids only, add the redaction helper, and delete the affected data under your retention rules. |
| Alerts fire constantly and get ignored | Thresholds were copied, not measured. | Re-baseline from a week of real data, raise thresholds just outside normal, and delete any alert nobody acts on. |
| Cost per run varies widely for the same task | Context growth across turns, retries after tool errors, or a long tool result being re-sent each turn. | Open the most expensive traces and read the per-step token counts. See how to reduce LLM costs. |
Verify it worked
Next steps
- Agent observability overview: The wider view of what to observe across agents and workflows.
- LLM observability: Prompt-level analytics and how tools compare.
- What to capture: AI observability for agents: The questions your records must be able to answer.
- How to validate your AI: Add evaluation so you know whether runs were right, not only what they did.
- How to red-team an LLM: Test the failures you want your alerts to catch.
Related guides
- How to Validate Your AI: Eval Sets, Gates, Evidence: A system-level method to validate an AI product: define the task and risk, build a held-out eval set, score it, gate releases, sample for human review, monitor and keep an evidence pack.
- How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
- How to Set Up Human Approval for AI Agents (With Code): Decide which agent actions need a person, set thresholds, pause the agent with LangGraph interrupts, route requests to a queue with a timeout that denies by default, show reviewers the evidence, and record every decision.
- How to Reduce LLM Costs: Step-by-Step Guide (2026): Measure token spend per feature first, then apply the levers in order of payoff: output caps, prompt caching, batch APIs, model routing, response caching, and a self-hosting break-even check.
- How to Red Team an LLM App: Tools, Scoring, Retest: A step-by-step LLM red-team exercise: authorisation and scope, a threat model mapped to the OWASP LLM Top 10, automated testing with promptfoo, garak and PyRIT, manual attack sessions, scoring, fixes and retest.
Frequently asked questions
What is AI agent observability?
It is the practice of recording each step an agent takes (model calls, tool calls, retrieved context and decisions) as a trace, so you can reconstruct any run and measure cost, errors and outcomes. It extends LLM observability from single prompts to multi-step work.
What should I monitor on an AI agent in production?
Track steps per run, ceiling hits, tool call counts and failures, tokens and cost per run, latency, policy decisions and a goal-level success rate from sampled review. Technical health alone misses the quiet failures agents have.
How do I detect an agent stuck in a loop?
Set a turn ceiling in the agent loop, record when a run hits it, and alert when ceiling hits rise above your baseline. Open a few of those traces and look for the same tool called repeatedly with similar arguments.
Are the OpenTelemetry GenAI conventions stable?
Not yet. As of 2026-10-06 the agent and tool span pages are marked Development status, so attribute names may change. Use them anyway for consistency, and pin your instrumentation library versions.
How do I avoid storing personal data in agent traces?
Record ids and counts by default and leave message content off. If you need content, redact in your own code before it reaches a span, pseudonymise user ids, and set a short retention period for trace content.
What is Langfuse used for?
Langfuse is an open-source platform for tracing and evaluating LLM applications, available as a cloud service or self-hosted. It nests runs and model calls as observations so you can inspect each step. It is one option; any OpenTelemetry backend also works.
Do I need a vendor tool to monitor agents?
No. You can start with OpenTelemetry and a backend you already run. A dedicated tool adds trace views, evaluation and prompt management, which are useful once you have more than a few agents.
How Swfte can help
If your agents run on Swfte, Connect and Studio show usage, cost, latency, error breakdown and tool usage per workspace, which covers part of this guide without extra code.
- AI observability on the platform: What the Swfte products capture today and what is designed.
- Usage and cost analytics: Spend by model, agent and workflow.
- SecOps AI incident response: A runbook for when an alert is real.
You can follow every step here without Swfte. Swfte's own issue and struggle detection covers sessions in its web applications, not traces of your agent runs, and per-step cost is not yet joined to every worker trace.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- OpenTelemetry GenAI semantic conventions repository: GenAI conventions now live here; opentelemetry.io pages link to it; folder layout includes agent spans, spans, metrics, events and MCP.
- GenAI agent and framework spans (raw docs): Development status, invoke_agent and create_agent operations, required and recommended attributes, content capture marked opt-in and sensitive.
- GenAI spans (raw docs): execute_tool span name, INTERNAL kind, gen_ai.tool.name and gen_ai.tool.call.id, content capture off by default and three storage approaches.
- OpenTelemetry Python: manual instrumentation: pip install opentelemetry-api and opentelemetry-sdk, TracerProvider and BatchSpanProcessor setup, start_as_current_span, set_status and record_exception.
- OpenTelemetry: OTLP exporter configuration: OTEL_EXPORTER_OTLP_ENDPOINT default, OTEL_EXPORTER_OTLP_PROTOCOL values, OTEL_EXPORTER_OTLP_HEADERS.
- PyPI: opentelemetry-sdk: pip install opentelemetry-sdk; version 1.45.1 released 6 October 2026.
- PyPI: opentelemetry-exporter-otlp: pip install opentelemetry-exporter-otlp installs the gRPC and HTTP exporters.
- Langfuse: get started with observability (Python): pip install langfuse, LANGFUSE_* environment variables and regional base URLs, get_client and start_as_current_observation code, flush.
- Langfuse: masking sensitive data: mask option on the client, noted as legacy; mask_otel_spans recommended for new Python setups.
- Langfuse documentation home: Langfuse described as open source, self-hostable and extensible.
Topics
- observability
- tracing
- OpenTelemetry
- Langfuse
- alerts
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-monitor-ai-agents-in-production.