← The journal
Security Guide

Shadow AI: Finding and Governing Unsanctioned Model Use

A workflow for shadow AI: discover agents, keys and MCP servers, then sanction, replace, restrict or retire them.

Swfte Journal / Security Guide

Capability page: shadow AI discovery. Concept and risks: what is shadow AI.

Most writing on shadow AI explains why it is a problem. That case is well made, including in our own earlier post on employees using AI tools you do not know about. This post is about what a security team does on Monday morning: a repeatable workflow for finding unsanctioned model use, deciding what to do about each finding and moving the useful ones onto a governed path.

We do not quote a prevalence figure for shadow AI, because we have none we can source for this post, and the figures that circulate come from surveys with very different definitions. The workflow does not depend on one. Every organisation that has looked has found something, and you will find out what yours is only by looking.

What counts as shadow AI

A narrow reading, an employee pasting text into a personal chatbot, misses most of the estate. A useful inventory category list:

  • Consumer and personal-account chat tools used for work content.
  • Coding agents and assistants installed on developer machines, often with the developer's own credentials and broad repository access.
  • Browser extensions and desktop apps with model access and page or file visibility.
  • API keys to model providers sitting in environment variables, CI secrets, notebooks and project configs.
  • MCP servers and other tool connectors started locally or hosted by a team, connecting agents to internal systems.
  • Agents built by business teams in low-code tools, with their own credentials and integrations.
  • Embedded AI features switched on by default in SaaS tools you already use.
  • Non-human identities: service accounts, tokens and app registrations created for AI use that nobody owns.

The unifying property is not that the use is malicious. It is that nobody accountable knows about it, so nobody has decided what it may reach.

Why a ban does not work

The people using these tools are usually among the most productive in the organisation, and the tools usually do help. A blanket ban moves use onto personal devices and personal accounts, where you have less visibility than before, and it costs the business real value. The effective position is that the sanctioned route must be easier and better than the unsanctioned one, and that discovery exists to move people onto it, not to catch them.

That has a practical consequence for how you run discovery: be explicit about what is and is not being looked at. Discovery is about which tools and identities exist and what they can reach. It does not need to read people's prompts. At the default privacy tier, Nexus records a prompt as a fingerprint and length only, with no prompt text stored, which lets a security team see that an agent is running and what it did without reading what anyone typed.

The five-step workflow

1. Discover

Use several sources, because each sees a different slice. The sources available to you depend on your connectors and on what you are permitted to monitor.

SourceWhat it finds
Agent runtime captureAgents running on managed machines, their sessions, tool actions and model providers
Identity provider and cloud inventoriesService accounts, app registrations, tokens and API keys, and who owns them
Gateway and network logsTraffic to model providers and tool endpoints
Secrets scanningModel-provider keys in repositories and configs
Expense and procurement dataPaid subscriptions to AI tools
Surveys and amnestyWhat people use and why, which is where the useful context lives

Run an amnesty alongside the technical work. People will tell you what they use if they believe the answer leads to a supported option rather than a reprimand.

In Nexus, shadow-AI detection surfaces unmanaged agents, unsanctioned model providers and unknown non-human identities, which are AI usage that existing SaaS security tooling often cannot see. Detection is strongest for agents it can hook or import, and wider coverage depends on connectors, so treat the inventory as a living document and expect gaps, especially for personal devices and personal accounts.

2. Classify by reach, not by name

The name of a tool tells you little. What matters is what it can reach.

For each finding record:

  • Owner, or "unknown," which is itself a finding.
  • Data classes it can see or has seen.
  • Systems and actions it can reach: read, write, execute, send.
  • Identity it runs under, and the scope of that identity.
  • Provider and terms: where data goes, retention, whether it is used for training, contractual protection.
  • Autonomy: does it act on its own, or only suggest?

Then rate it. A harmless-looking API key that reaches production data is not a sandbox toy. An identity graph, showing which systems each credential can reach, makes the blast radius visible. A coding agent with write access to a repository and a deploy key deserves more attention than a chat tab.

A simple three-tier rating is enough to start:

  • Low: no sensitive data, no write or external action, easily replaced.
  • Medium: touches internal data or has bounded write actions.
  • High: touches regulated or confidential data, holds broad credentials or can take consequential external actions.

3. Decide: sanction, replace, restrict or retire

Every finding gets exactly one of four outcomes, recorded against a named owner. Nothing stays in an unknown state.

  • Sanction. The use is legitimate and the tool is acceptable. Register it, assign an owner, apply policy and add it to the inventory.
  • Replace. The need is real but the tool is not acceptable. Provide a governed equivalent and migrate.
  • Restrict. The tool can stay under conditions: certain data classes excluded, a lower permission set, an approval step on certain actions.
  • Retire. The use is not needed or too risky. Revoke access and tell people why.

The most common mistake is stopping at "found." A list of fifty unclassified findings is a risk register nobody will read.

4. Migrate to a governed path

Migration is where shadow AI turns into a managed capability. The governed route should mean:

  • Models reached through a gateway, with approved providers, input and output screening, logging and cost attribution.
  • Agents with identity, an owner, scoped permissions, a Trust Profile and in-flight policy.
  • Tool access through an approved registry, with scoped credentials, as described in securing MCP servers and tool access.
  • Data access through controlled retrieval, not copy and paste.
  • A record of what happened, so the SOC can investigate and an auditor can see.

A useful shadow workflow is a requirements document. If a team built a clever agent to chew through contracts, rebuild it as a governed agent with the same job, a Trust Profile and a named owner. People rarely resist a supported version of something they built to solve a real problem, particularly if it is faster to use and does not require them to manage their own keys.

5. Monitor continuously

New findings keep arriving, because new tools keep appearing. Make discovery a scheduled workflow with an owner, not a project with an end date.

  • Re-run discovery on a cadence you set, and on triggers such as a new procurement, a new MCP server or a new provider domain.
  • Evaluate policy continuously. A finding should show exactly which policy failed and where, rather than a screenshot of a settings page.
  • Track the number of unowned identities, the number of findings without a decision and the age of open findings. These are hard to game and actionable. You may find that the share of unowned identities falls over time, which would be a real sign of progress, but only your own data can tell you so.

Controlled autonomy for discovery itself

Discovery can be partly automated by an agent, and the usual rule applies: autonomy follows the cost of being wrong. Finding and classifying is low-risk, so an agent can do it at L1. Blocking and revoking interrupt people's work, so they start at L2, where an owner approves each action. Within narrow limits, for example restricting a clearly unsanctioned provider on a managed device with an easy undo, supervised action at L3 is reasonable. The agent should not be able to read prompt text at the default privacy tier, revoke a credential on its own initially, or mark a finding closed without an owner's decision.

Non-human identities are the center of gravity

Many findings are, at bottom, credentials: API keys, service accounts, tokens and application registrations created for AI use. Treat them as first-class identities. Each needs an owner, a purpose, a scope, a rotation policy and an expiry or review date. An identity with no owner is revoked or adopted, not left alone. The non-human identity page covers the broader practice.

Frameworks, briefly

Shadow AI touches several published risk lists. OWASP's Top 10 for Agentic Applications includes Rogue Agents (ASI10) and Identity and Privilege Abuse (ASI03), both of which an inventory addresses directly. The OWASP LLM list covers Sensitive Information Disclosure (LLM02) and Supply Chain (LLM03). Knowing which models and components are in use is the precondition for assessing them under the supply-chain heading, which MITRE ATLAS also covers as AI supply chain compromise (AML.T0010; check atlas.mitre.org for the current matrix). See OWASP LLM Top 10 explained for engineers.

EU context

Unsanctioned AI use is a data-protection and a governance problem at once. GDPR Article 32 requires security appropriate to the risk, and Article 5(2) makes you accountable for demonstrating the measures you applied; it is hard to demonstrate measures for data flows you cannot see. The EU AI Act expects organisations to understand the AI systems they deploy, and Article 4 addresses AI literacy; which obligations apply depends on the system and your role as provider or deployer. NIS2 Article 21 includes supply-chain security among its minimum measures, and DORA has its own ICT third-party regime for financial entities. An inventory of AI services and their suppliers supports evidence for all of these. It does not make a use compliant by itself, and the exact posture depends on your use case, jurisdiction, deployment and configuration.

Limits

No discovery method sees everything. Personal devices, personal accounts and encrypted channels are partly or wholly invisible without the right controls and agreements. Detection coverage is deepest for agents that can be hooked or imported and depends on connectors elsewhere. And discovery without a governed alternative produces a list of grievances, not a safer organisation.

Where to go next

Start with the capability page on shadow AI discovery, then agent runtime security for bringing discovered agents under policy, and the SecOps hub for the rest. For the concept and risk framing, see what is shadow AI and AI agent governance. To plan a discovery exercise for your environment, talk to our team.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.