AI Incident Response Runbook: Contain, Preserve, Investigate, Report
A six-step AI incident response runbook: contain, preserve evidence, investigate and meet NIS2 and DORA clocks.
Capability page: AI incident response. Evidence: AI audit trail.
Your incident response plan almost certainly handles a compromised laptop, a stolen credential or a ransomware event. It probably does not handle this: an AI agent sent customer data to an address nobody recognises, and the person who owns the agent is on leave. Or a model answered a question with a document the asker should not have seen. Or a tool server your agents use changed its behaviour overnight.
These are new cases for a familiar discipline. The phases are the same. What changes is what you need to know, what levers you can pull and what the clock demands. This post is a runbook outline you can adapt. It assumes nothing about your tooling beyond a few capabilities, which we flag, and it is not legal advice.
Three roles an AI system can play in an incident
Classify first, because the response differs.
AI as cause. The agent or model did something it should not: acted outside its permissions, produced harmful or wrong output that was acted on, or ran away in a loop. The question is why, and whether it will happen again.
AI as target. An attacker is going after the AI system: prompt injection, poisoning of a knowledge source, tampering with a tool server, theft of a model credential. The question is what they achieved and what they can still reach.
AI as tool. An attacker or an insider uses your AI to move: an agent with broad access is steered into reading and exfiltrating data, or into taking actions on the attacker's behalf. The agent is both a victim and a means.
Often it is more than one. A successful prompt injection (target) makes an agent leak data (cause) using its own legitimate access (tool).
What you need before the incident
The first hour of an AI incident is spent on questions that should already have answers on file. Prepare these in advance.
- An inventory. Every agent and non-human identity, with an accountable owner, risk level, approved models, connected systems and permitted actions. If an agent is not in the inventory, it is a finding in itself.
- A kill switch. A way to pause an agent or revoke its credential immediately, tested, with named people authorised to use it. Pausing must be recorded.
- A runtime record. For each agent: sessions, tool calls and results, policy decisions (including blocks), file changes, model and version, and approvals. At the default privacy tier, prompts are fingerprints, not text, so decide in advance what depth of capture your policy allows and for which data classes.
- A map of blast radius. For each credential, what else can it reach? An identity graph answers this in seconds; a spreadsheet answers it in an afternoon.
- A reporting map. Which regulators, customers and contracts impose reporting duties on you, with the clocks, and who files.
If you take one thing from this post, it is that containment and reporting both depend on records you must have created before anything went wrong.
Step 1: detect and classify
Triggers include a policy block on a sensitive action, an anomaly on agent behaviour, a spend or action-cap breach, a user report, an alert from a SOC tool or an external notification.
Open a case. Record who became aware and when, because several regimes count from awareness. Identify the agent from the inventory: owner, Trust Profile, connections, current autonomy level. Classify the role (cause, target, tool) and set a provisional severity. Name an incident lead.
Step 2: contain
Contain at the right scope. Too narrow and the attacker keeps access. Too broad and you take down workflows that were not affected.
- Pause or revoke the agent. Use the kill switch. Do this before anything that alters its configuration.
- Rotate the credentials it held. Assume anything the agent could read, an attacker or a manipulated agent could have used. Use the blast-radius view to find every system the credential touched.
- Detach compromised components. A tool server, a knowledge source, a connector. Remove it from the registry so no other agent can use it.
- Restrict adjacent agents. If a shared tool or source is implicated, lower the autonomy level of other agents that use it until you know more.
- Block egress where relevant. If data may be leaving, cut the destination at the gateway.
Contain at the level of evidence you hold, and be cautious about letting automation contain on its own mid-incident. A pre-agreed, narrow, reversible containment such as pausing the agent that tripped a policy can reasonably run automatically. Wider actions should have a person approve them.
Step 3: preserve evidence
Do this before you fix things. Changes overwrite state.
Freeze the audit ledger for the agent and the sessions involved. Preserve:
- the inputs, as fingerprints or content according to your privacy tier,
- retrieved documents and tool results,
- the model and version,
- every tool call with target, arguments and result summary,
- policy evaluations and verdicts, including blocked calls,
- approvals: who, when and what they were shown,
- file changes and dependency installs,
- the configuration and Trust Profile at the time.
Snapshot the configuration of tools and knowledge sources. Hash what you can. Restrict access to the preserved copy and record who touches it. An agent assisting the response must not be able to alter audit records, at any autonomy level.
Mind privacy. The evidence may contain personal data, and handling it is itself subject to data protection law. Keep access to those who need it and apply your retention policy.
Step 4: investigate
Establish what happened and why. A useful frame is the chain that traceability is meant to give you: data, model, agent, decision, action, outcome. Walk it from the outcome backwards.
- What was the action, and what changed or left?
- Which tool call performed it, and under which policy decision?
- What did the agent read immediately before it? Look for planted instructions in documents, tickets, emails, web pages and tool output.
- Was the model manipulated (injection), wrong (misinformation), over-permitted (excessive agency) or deceived through a poisoned source?
- Was a human approval involved, and what evidence did the approver see?
- What else could the same credential or the same poisoned source reach? Did anything else use them?
Label what you find using shared vocabulary: the relevant OWASP entry (for example LLM01 Prompt Injection, LLM06 Excessive Agency, or ASI08 Cascading Failures in the agentic list) and the MITRE ATLAS technique where one applies. That makes the case comparable and makes it easy to write the regression test later. Look up current ATLAS identifiers at atlas.mitre.org rather than from memory.
An investigation agent can assemble the timeline for an analyst to verify, with sources cited. A person decides the root cause. See AI SOC agents with human approval for the sequencing of that kind of assistance.
Step 5: eradicate and recover
Remove the cause, not just the symptom.
- Injection: tighten the permission or remove the tool that made the injection useful; add a gate on the action; clean or quarantine the source that carried the payload.
- Poisoning: identify and remove poisoned content; restore the source from a known-good state; review who could write to it.
- Over-permission: reduce scope; split read and write; move the action behind approval.
- Supply chain: remove or pin the component; re-review before readmitting.
- Model or logic error: add a check, a limit or a human review; consider a different model through the gateway.
Restore at a lower autonomy level than before. An agent that has had an incident has reduced credit. Raise its level again only as the new record supports it, and by a decision of the accountable owner recorded in its Trust Profile.
Step 6: report and learn
Notify. Reporting duties depend on your role, sector and the classification of the incident. The following describes the instruments as we understand them. Verify the current text and your regulator's guidance, because deadlines vary by classification and rules change.
- NIS2 (Directive (EU) 2022/2555), Article 23: entities in scope report significant incidents to their CSIRT or competent authority in stages. An early warning is due within 24 hours of becoming aware, an incident notification within 72 hours of becoming aware and a final report within one month of the incident notification.
- DORA (Regulation (EU) 2022/2554): financial entities classify ICT-related incidents and report major ones in stages, with an initial notification, an intermediate report and a final report, using harmonised templates. The first notification is measured in hours.
- EU AI Act, Article 73: providers of high-risk AI systems report serious incidents to market surveillance authorities. The deadlines depend on the incident: no later than 15 days after becoming aware as the general rule, 2 days for a widespread infringement or serious disruption of critical infrastructure, and 10 days where a person has died. Deployers have monitoring and information duties under Article 26, and must keep logs under their control for at least six months, unless other law, in particular data protection law, provides otherwise. The Commission has issued draft guidance on serious-incident reporting, and application dates for some high-risk provisions have been under discussion, so check the current status.
- GDPR, Article 33: where a personal data breach is likely to result in a risk to people's rights and freedoms, notify the supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware.
Which of these applies to you is a legal question. The platform's job is to give you accurate timestamps and scope to report from. Drafting can be assisted. A named person submits and remains accountable.
Learn. Convert the incident into three things: a new policy rule or approval gate; a detection that would have caught it earlier; and a regression test in your red-team suite, so the same attack is rerun on every model, prompt or tool change. Update the inventory and the Trust Profile. Review how long each step took and where the evidence was missing, because the gaps in evidence are the best argument for fixing logging.
A one-page checklist
First fifteen minutes
- Case opened, awareness time recorded, incident lead named
- Agent identified from inventory, owner contacted
- Role classified: cause, target, tool
First hour
- Agent paused or credential revoked, recorded
- Credentials rotated for everything in the blast radius
- Compromised tool, source or connector detached
- Audit ledger frozen and access restricted
First day
- Timeline reconstructed with sources
- Root cause hypothesis tagged to OWASP and ATLAS
- Reporting duties assessed with legal and compliance; clocks started
- Adjacent agents reviewed and limited if needed
After
- Cause removed; agent restored at lower autonomy
- Notifications filed by a named person
- Policy, detection and regression test added
- Inventory and Trust Profiles updated
What this does not do
A runbook does not prevent incidents; it shortens the time to understand and contain them. We state no response-time improvement from any product, because we have none we can source here. SecOps Agents is in beta, and the response assistance described above is what the platform is designed to let your team do. Nexus enforcement and capture are deepest today for coding agents.
Where to go next
Read the capability page on AI incident response, the AI audit trail for what a usable record contains, and agent runtime security for the kill switch and policy layer. For earlier detection, see AI agent clusters and behavior monitoring, and for the wider governance picture, enterprise AI governance and risk. To work through your own plan with us, talk to our team.