← The journal
Security Guide

AI SOC Agents With Human Approval: What to Automate First

What to automate first in an AI SOC: triage, investigation and response, with approval gates and autonomy levels.

Swfte Journal / Security Guide

See the capability page: AI SOC agents. The wider section is the SecOps hub.

Every security team that looks at AI SOC agents asks the same question in the first meeting: what is it allowed to do on its own? The honest answer is "less than the demo suggests, and more than you expect once it has earned it." The useful way to plan is not to pick a point on a line from manual to autonomous. It is to decide, action by action, how much authority an agent holds, and to move that authority only as the record supports it.

This post is a sequencing guide. It covers what to automate first, what to leave with people for now, where the approval gates belong and how to know when to loosen one. It does not quote detection rates or time-saved figures, because we have none we can source here and you should distrust any vendor who leads with them.

Why the queue is the wrong place to start with autonomy

Alert volume is the visible problem, so teams reach for the fix that looks most like the problem: let the machine close alerts. That is the highest-stakes place to begin. A wrongly closed alert is a missed incident, and you will not find out until later.

Start instead where an error is cheap and visible. Reading, enriching and summarising are all things an agent can do with read-only access, and every mistake it makes is caught by an analyst who was going to look at the alert anyway. You gain time immediately and you gain something more valuable: a record of how often the agent is right, on your alerts, in your environment.

That record is the currency of controlled autonomy. The model we use has five levels, applied per action type.

LevelNameWhat it means for a SOC
L1AssistThe agent recommends. A person does the work and decides.
L2ApproveThe agent prepares the action. A human approves before it runs.
L3SuperviseThe agent acts within limits and is monitored. Exceptions go to a person.
L4AutonomousThe agent acts independently within strict policy and risk bounds.
L5AdaptiveThe agent proposes improvements to itself within controlled boundaries.

An agent can sit at L3 for one action and L1 for another. Isolating a non-critical test endpoint and disabling an executive's account are not the same decision and should not share a level.

Phase 1: triage and enrichment at L1

The first set of tasks has one property in common: the agent produces a recommendation and a person acts on it.

Enrichment. For each alert the agent pulls context: the asset and its owner, the identity involved, recent changes, related alerts, and the indicators looked up against your approved threat intelligence. This is the work analysts hate and the work most consistently identical from alert to alert.

Deduplication and grouping. Many alerts are the same event seen by different tools. Grouping them gives the analyst one case instead of six.

Severity recommendation with reasoning. The agent proposes a severity and states why, citing the evidence it read. The reasoning matters more than the label. An analyst can check reasoning in seconds and learn quickly where the agent is weak.

Case summary and ticket drafting. The agent writes the note an analyst would have written. A person edits and files it.

What to track in this phase: how often analysts accept the recommended severity, how often they change it and in which direction, and what kind of alert produces the disagreements. If analysts are raising severity more often than lowering it, the agent is under-calling, which is the dangerous direction. You want to know that before it has any authority.

Trust boundaries matter from the first day. Alert text, ticket bodies, email subjects and log lines are attacker-influenced. An agent reading them is reading untrusted input, and it should have no ability to widen its own permissions in response to anything it reads. The prompt injection post covers that in detail. For Phase 1 the control is simple: read-only access, a tool allowlist and no outbound actions.

Phase 2: investigation at L1 to L2

Investigation is where agents start to feel like colleagues. Given a case, the agent assembles a timeline: what happened before and after the alert, which other systems the same identity touched, what changed in the last day, and whether anything similar appears in prior incidents held in your knowledge base.

Two things make this safe to extend.

First, every claim in the investigation cites its source. An analyst should be able to click from "the account authenticated from a new country" to the log line. Without citations the agent is a fluent narrator and nothing more.

Second, the agent separates what it observed from what it infers. "Three failed logins, then a success from the same address" is an observation. "Likely credential stuffing" is an inference. A good case summary labels them, and a good review process checks the inferences.

At this stage the agent starts preparing actions without taking them: a draft containment plan, the affected scope, the rollback. This is L2. The action sits in a queue for a named approver. The approval screen should show the evidence and the risk, not only the agent's summary of them. One of the failure modes catalogued in the OWASP Top 10 for Agentic Applications is Human-Agent Trust Exploitation, where a person is nudged into approving something because the agent sounded confident. Designing the approval view to show the underlying evidence is the countermeasure.

Phase 3: contained response at L2 to L3

Response is where the most caution belongs, and where the most value is. Containment speed matters: the gap between detection and isolating a compromised host is a gap an attacker uses.

The move from L2 to L3 for a given action should be a deliberate decision with entry criteria. A reasonable set:

  • The action is reversible, or the cost of reversal is low and understood.
  • The agent has a record of recommendations for this action class that analysts accepted, over enough cases that you trust the sample. You decide what "enough" is.
  • Explicit limits exist, written as policy rather than as hopes: which asset tiers, which user classes, how many actions per hour.
  • Monitoring exists on the agent itself, with alerts when it hits a policy limit or behaves unusually.
  • There is an easy undo, and someone has tested it.

Within those limits, the agent contains and a person reviews afterwards. Above them, the action escalates to an approver. Typical examples of well-bounded L3 actions are isolating a non-critical endpoint, revoking a single standard user's session, or blocking an indicator at a perimeter control where a rule can be removed in one step.

Typical examples that stay at L2, perhaps permanently: disabling privileged or executive accounts, isolating production servers, anything that touches a safety or customer-facing system, and any external notification. Staying at L2 is not a failure. For these actions the human review is the control.

What not to automate first

A short list of things we would keep out of the first ninety days, and often much longer:

Closing high-severity cases. A person closes them.

Deleting or altering evidence. An agent should never be able to do this, at any level. Audit and log stores are not agent-writable.

External communications. Customer, regulator and press notifications are written by people. An agent can draft.

Changing its own configuration. Policy, thresholds and autonomy level are set by the accountable owner, recorded in the agent's Trust Profile, and are outside the agent's reach.

Anything where you cannot articulate the rollback. If you cannot say how you would undo it, it is not ready for autonomy.

The Trust Profile for a SOC agent

Before an agent touches an alert it should have a written profile, in the same shape as any governed agent. For a triage and response agent it reads roughly like this.

Can: read alerts and approved telemetry; enrich indicators with approved threat intelligence; correlate with open cases and past incidents; recommend severity and a next step with cited evidence; draft the case summary and ticket.

Cannot: read data outside its approved sources; disable its own policy or raise its own autonomy level; close a high-severity case without a human; delete or alter logs and evidence; contact external parties on its own.

Requires approval: isolating a host above the configured asset tier; disabling an account outside a standard user class; any action whose blast radius exceeds the set threshold; notifying customers, regulators or the press.

Records: agent identity and owner; alert and data accessed; model used; tools called; policy applied; recommendation, approval, action and outcome.

The table is not decoration. The "cannot" column is what limits the damage if the agent is manipulated or simply wrong, and the "records" column is what makes the case file evidence rather than recollection. For how this runs on the platform, see agent runtime security and the AI audit trail.

Measuring whether to loosen a gate

Resist metrics you cannot defend. A few that are hard to game and useful:

  • Acceptance rate of recommendations by action class, and the direction of disagreement.
  • Override and undo rate on supervised actions. A rising undo rate on an L3 action is a reason to step back to L2.
  • Policy-limit hits. An agent that regularly pushes against its limits either has limits that are too tight or behaviour that is drifting. Both are worth knowing.
  • Coverage gaps. Alert classes the agent has never seen are alert classes it should not act on.

Each of these comes from the record the platform keeps. None of them requires publishing a headline number, and none should be the basis for a claim to a board unless it comes from your own data.

The EU angle, kept honest

For teams in scope of NIS2 or DORA, the incident record that a governed SOC agent keeps is useful evidence: when you became aware, what was affected, what was decided and by whom. NIS2 Article 23 sets a staged reporting clock of an early warning within 24 hours of becoming aware, an incident notification within 72 hours and a final report within one month of the notification. DORA sets its own classification and staged reporting for financial entities. The EU AI Act's human-oversight expectations for high-risk systems, in Article 14, are another reason to keep approval points explicit and logged, where a SOC use falls within that scope.

Swfte supports evidence for these obligations. It does not make an organisation compliant. Whether and how they apply to you depends on your sector, role and facts, and that is a conversation for your legal and compliance team.

A practical starting plan

If you are starting from nothing, a plan that fits the above:

  1. Pick one alert class that is high-volume and well understood, such as phishing reports or impossible-travel logins.
  2. Run the agent at L1 on that class, read-only, with citations required.
  3. Measure acceptance and the direction of disagreement for a period you set in advance.
  4. Add investigation summaries and an approval queue for a prepared action at L2.
  5. Choose one reversible action and write its limits as policy. Move that single action to L3 with monitoring and a tested undo.
  6. Review, widen to a second alert class, and repeat.

The pace is yours. The point of controlled autonomy is that the pace is a decision rather than an accident.

SecOps Agents is in beta, and the capabilities above describe what the platform is designed to let your team do. If you want to build a first agent yourself, use Build with AI. If you would rather talk through your environment, tooling and regulators first, talk to our team. For a look at the other half of SecOps, securing the agents themselves, start with the SecOps hub and the OWASP LLM Top 10 explained for engineers.

Related reading: AI agent clusters and behavior monitoring and AI incident response runbook.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.