Short answer
Shadow AI is AI use your organisation has not approved or does not know about. Find it by combining four sources: third-party app grants in your identity provider, web traffic or proxy logs, expense and procurement records, and a candid staff survey. Rank findings by the data involved, offer approved alternatives before banning anything, and keep monitoring proportionate and agreed with staff.
The steps at a glance
- Define what counts as shadow AI and what does not
- Review third-party app grants in your identity provider
- Use web traffic data to see which AI services are reached
- Check expenses, procurement, cloud billing and code
- Ask people, and make it safe to answer
- Triage each finding by the data it touches
- Offer approved tools before you restrict anything
- Monitor in a proportionate way and review monthly
Before you start
Who this is for
- Security, IT and data protection teams asked "who is using AI here?" with no good answer.
- Heads of AI who need an honest picture before writing or enforcing a policy.
- Compliance owners preparing an AI inventory.
Probably not for you if
- Teams looking to monitor individual employees' prompts. This guide covers discovery of tools and data flows, not reading people's conversations.
- Organisations with no identity provider or logging at all: start with the basics of device and account management.
Prerequisites
- Admin access, or a person with it, to your identity provider (Microsoft Entra or Google Workspace are shown) and to web proxy or firewall logs.
- Access to expense and procurement data, with finance's agreement.
- A named owner for the programme and a contact in data protection or legal.
- A draft list of approved AI tools, even if short.
- Time
- About one to two weeks for a first sweep; then a monthly review
- Cost
- Uses tools you may already have. Some discovery features need specific licences: check yours before planning.
- Skill
- Identity and network administration, plus the ability to talk to staff without making it feel like an investigation
Estimates are ours, not measurements, and move with your hardware, data and network.
Why shadow AI is worth finding
The risk is mostly about data: text pasted into a service you have no contract with, files uploaded to a tool that keeps them, or an app granted access to a mailbox. The second risk is blind spots: you cannot assess, classify or explain an AI system you do not know exists. The third is duplication: five teams buying five tools for one job.
The risk is not that staff are careless. Most of them are trying to get work done with the best tool they can find. A good programme treats discovery as a way to learn what people need, not as a hunt.
Step 1Define what counts as shadow AI and what does not
You end up with: A short written definition and a list of categories, agreed with security, legal and the business.
Write down the line. A useful definition is: any AI tool, model, agent or feature used for work that the organisation has not approved, or that touches data the approval does not cover. That includes obvious cases (staff pasting text into a public chatbot), and quieter ones: browser extensions with AI features, AI add-ons switched on inside tools you already pay for, meeting-notes bots, personal accounts of approved tools, and agents built by a team with an API key from a personal card.
Separate three questions, because they need different answers. Is the tool known? Is it approved? Is it used with data it is not approved for? An approved tool used on restricted data is a data problem, not a discovery problem. An unknown tool used on public marketing copy is a low-priority discovery item.
Agree the principles with data protection and, where you have them, employee representatives, before you start collecting. The aim is to find tools and data flows, not to read individuals' conversations. Say that in writing, and keep to it.
Categories worth tracking Category Example Where it usually shows up Public chat tools A consumer chatbot on a work device Proxy and DNS logs, survey Third-party apps with account access A notes or email assistant with OAuth access Identity provider app grants AI features inside approved software An assistant switched on in a SaaS tool Admin consoles, vendor release notes Browser and desktop extensions A writing or summarising add-on Endpoint inventory, browser management Self-built agents and scripts A script calling a model API with a personal key Expense data, code search, cloud billing Step 2Review third-party app grants in your identity provider
You end up with: A list of apps that users or admins have granted access to company accounts, with the permissions each holds.
Many AI assistants ask you to "sign in with" your work account and request access to mail, files or calendars. Those requests leave a trail in your identity provider, which makes it one of the best places to look.
In Microsoft Entra, sign in to the Entra admin center as at least a Cloud Application Administrator, browse to Entra ID, then Enterprise apps, then All applications, pick an application and select Permissions. The Admin consent tab shows permissions granted for the whole organisation. The User consent tab shows what individual users or groups granted. Microsoft notes that you cannot revoke user-consented permissions in the portal; you use Microsoft Graph or PowerShell for that, and revoking does not stop users consenting again unless you change how consent is configured.
In Google Workspace, go to Menu, then Security, then Access and data control, then API controls, and select Manage App Access to see configured and accessed apps. For each app you can set it to Trusted, Limited, Specific Google data, or Blocked. Sort the list by number of users and by the scopes requested, and look at anything with access to mail, drive or calendar that you did not approve.
Record each app, who granted it, the permissions and the number of users. Do not revoke yet. First learn what people use it for, in step 5.
Microsoft Graph: delegated permissions granted to one application (from Microsoft's documentation) · text GET https://graph.microsoft.com/v1.0/servicePrincipals/{id}/oauth2PermissionGrantsChecked against: Microsoft Learn: Review permissions granted to enterprise applications, Google Workspace Admin: Control which apps access Google Workspace data, OWASP GenAI: LLM06:2025 Excessive Agency
Step 3Use web traffic data to see which AI services are reached
You end up with: A ranked list of AI-related domains and categories reached from your network, by volume and user group.
If you route traffic through a proxy, secure web gateway or firewall that logs destinations, you can match them against AI services. Microsoft Defender for Cloud Apps, for example, describes cloud discovery as analysing traffic logs against a catalogue of over 31,000 cloud apps, scored on more than 90 risk factors. It supports snapshot reports from logs you upload from firewalls and proxies, and continuous reports through endpoint integration, a log collector or a secure web gateway. It lists many supported firewalls and proxies, and notes that apps not in the catalogue cannot be discovered by default.
That last point matters for AI, where new services appear quickly. Whatever tool you use, keep your own list of AI domains, update it monthly, and treat a long tail of unknown domains with high upload volume as worth a look. If you have no discovery product, DNS or proxy logs filtered by that list will still give a first picture.
Look at aggregates first: which services, how many users, how much data uploaded. Resist the pull to look at individuals. If a service looks risky, ask the owning team what it is for before anyone looks at a name.
Checked against: Microsoft Learn: Cloud app discovery overview (Defender for Cloud Apps)
Step 4Check expenses, procurement, cloud billing and code
You end up with: A list of AI spend and API use that never went through procurement.
Search expense claims and card statements for AI vendors and subscription descriptors. Check procurement for tools bought by a department without a security review. Ask finance for recurring charges under a threshold, since small subscriptions are how many tools enter.
In cloud billing, look for model API services and GPU instances in accounts or projects outside your platform team. In code, search repositories for model-provider SDK imports and API base URLs, and for keys in config. A script that calls a model with a personal key is shadow AI that no network control will see from the inside.
Keep what you find factual: vendor, owner, monthly cost, and purpose if known. You are building a register, and a register is only useful if people trust that being in it is not a punishment.
Step 5Ask people, and make it safe to answer
You end up with: A short survey and a few conversations that show what staff use and why.
Logs show tools. People tell you reasons. Run a short anonymous survey: which AI tools do you use for work, what for, what data do you put in, what would you need to stop using a personal account. Add two or three team conversations with heads of function.
Say clearly how the answers will be used. If people think a candid answer leads to discipline, they will under-report, and your data will be wrong in exactly the places that matter. An amnesty window, stated plainly, works better than a threat.
Compare the survey to the logs. Tools in the survey but not the logs show where your visibility is weak, such as personal devices or home networks. Tools in the logs but not the survey show where people have not realised the tool counts.
Step 6Triage each finding by the data it touches
You end up with: Every finding has a risk rating and a decision: approve, replace, restrict or block.
Give each tool a data rating: what is the most sensitive data staff put into it? Use your own classification. Add what you can learn about the tool: does it keep inputs, train on them, where is data processed, what contract is in place, which sub-processors does it use. Ask the vendor if the information is not public, and mark what you could not confirm.
Then decide. Approve tools with an adequate contract and low-risk use. Replace tools used on sensitive data with an approved alternative. Restrict tools that are fine for public content by saying what data may go in. Block only where there is a clear unacceptable risk and a workable alternative.
Handle the high-risk cases first: sensitive personal data, customer data, source code, credentials, regulated records. If a tool holds credentials or broad account access through an OAuth grant, treat revocation and a credential rotation as part of the response.
Triage matrix Data involved Tool status Typical decision Public or non-sensitive Unknown tool Allow with usage guidance; add to register Internal, low sensitivity Unknown tool Review contract and settings; approve or replace Customer, personal or regulated data Unknown tool Stop use now; move to an approved tool; assess exposure with data protection Credentials or source code Any unapproved tool Revoke grants, rotate credentials, replace Any data Approved tool, wrong plan or personal account Move to the company account and settings Step 7Offer approved tools before you restrict anything
You end up with: People have a sanctioned way to do what the shadow tool did, and know how to get access.
Shadow AI exists because the tool is useful and the official route is slow or missing. Blocking without an alternative moves the activity to personal devices, where you can see even less. For each common use case in your findings (drafting, summarising, coding help, translation, data analysis), name the approved tool and how to request access.
Publish a one-page policy in plain language: what is approved, what data may never go into an AI tool, how to ask for a new tool, and who to tell about a mistake. Give a fast review path for new tools, measured in days. If review takes months, people will not use it.
For tools you build yourselves, such as internal agents, the same rules apply through how to govern AI agents. For a company-hosted option, how to self-host a ChatGPT alternative shows one route.
Step 8Monitor in a proportionate way and review monthly
You end up with: A light, repeatable process that keeps the register current without watching individuals.
Set a monthly review: new app grants since last time, new AI domains in traffic, new AI spend, and survey pulse results every quarter. Keep the output at the level of tools, teams and data types. Alert on high-risk patterns, such as a new app with broad mail or drive access, or a large upload to an unknown AI service, rather than on individuals' general browsing.
Be careful with monitoring that touches people. It may involve personal data and, depending on where you operate, employment rules and employee consultation. Involve your data protection contact and any employee representatives before you collect more than aggregate tool-level data, and document the purpose and retention. This guide is not legal advice.
Feed findings into your wider inventory. An AI system you have discovered is one you can now classify and govern, which is the starting point for the work in how to prepare for the EU AI Act and how to audit AI systems.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| The identity provider shows hundreds of third-party apps | Years of user consent to all kinds of tools, not only AI. | Sort by number of users and by scopes requested, filter for mail, files and calendar access, and review the top of the list first. |
| Revoked permissions come back | Users can consent again if consent settings allow it. | Change how users consent to applications, or require admin approval for new apps, then revoke. |
| Traffic data shows few AI services even though staff say they use many | Personal devices, home networks, encrypted DNS, or services missing from the discovery catalogue. | Keep your own AI domain list, widen sources to endpoint data where you have it, and rely on the survey for the gap. |
| The survey gets few or evasive answers | People fear discipline. | State the purpose and the amnesty in writing, keep it anonymous, and share what you did with the answers. |
| Staff move to personal devices after a block | No approved alternative, or approval takes too long. | Provide an approved tool for the use case and a review path measured in days; reserve blocks for clear unacceptable risk. |
| You find a tool holding customer data | Unreviewed use of an unapproved service. | Stop the use, involve data protection to assess exposure and any duty to notify, revoke grants, rotate credentials, and move the work to an approved tool. |
Verify it worked
Next steps
- How to govern AI agents: bring discovered agents into an inventory with owners and permissions
- How to self-host a ChatGPT alternative: give staff an approved chat tool you control
- How to prepare for the EU AI Act: turn the register into an AI system inventory
- Shadow AI discovery: Swfte's overview of discovery as a security practice
Related guides
- How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
- How to Self-Host a ChatGPT Alternative (Open WebUI): A hands-on setup of Open WebUI in front of a model you host, with admin accounts, roles, HTTPS, backups and a clear view of where prompts go.
- How to Prepare for the EU AI Act: 2026 Readiness Steps: A readiness workflow for the EU AI Act: inventory AI systems, rule out prohibited uses, set provider or deployer role, classify risk, plan obligations and evidence, and track the dates after the Digital Omnibus.
- How to Audit AI Systems: Scope, Evidence, Findings: How to audit an AI system, internally or for a client: scope it, choose criteria, request and sample evidence, test logs, change control and human oversight, and write findings that can be fixed.
- How to Do a DPIA for AI: GDPR Article 35 Steps: Decide whether an AI system needs a DPIA, then describe the processing, assess necessity, identify risks to people, choose measures, record the sign-off and review it, with AI-specific risks and a worked example.
Frequently asked questions
What is shadow AI?
It is AI use for work that the organisation has not approved or does not know about: public chatbots used with company data, third-party apps granted account access, AI features switched on inside software, browser extensions and scripts calling model APIs with personal keys.
How do you detect shadow AI in a company?
Combine four sources: third-party app grants in your identity provider, web proxy or firewall logs matched against AI services, expense and billing records, and an honest staff survey. Rank findings by the data involved and review monthly. No single source finds everything.
Can I see which employees use ChatGPT?
Possibly, from proxy or endpoint logs, but whether you should is a different question. Monitoring individuals may involve personal data and employment rules that vary by country. Start with aggregate, tool-level data and involve data protection and employee representatives before going further.
Is shadow AI the same as shadow IT?
It is a subset with extra risks. Shadow IT is any unapproved technology. Shadow AI adds that inputs may be kept or used for training, that outputs may be wrong, and that agents can act with access you granted. The discovery methods overlap, but the data questions differ.
Should we ban AI tools we find?
Rarely as a first step. Rank them by the data involved, give staff an approved tool for the same job, and block only where the risk is clearly unacceptable and a workable alternative exists. Blanket bans tend to move use to personal devices where you cannot see it.
What should an AI acceptable-use policy say?
Which tools are approved, what data may never be entered, how to request a new tool and how fast review is, who to tell about a mistake, and that discovery looks at tools and data flows rather than reading conversations. Keep it to a page and review it every few months.
How Swfte can help
Swfte publishes material on shadow AI discovery as part of its security work, and its platform pages describe how approved tools and governed agents fit the policy side.
- Shadow AI: what shadow AI is and why it matters
- Shadow AI discovery: how discovery fits security operations
- AI governance: the policy and evidence side once you have an inventory
Whether Swfte provides automated discovery in your deployment is not stated here: <shadow AI discovery product availability - founder to fill>. Every step in this guide uses tools you probably already have, and works without Swfte.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- Microsoft Learn: Review permissions granted to enterprise applications: Admin center path (Entra ID, Enterprise apps, All applications, Permissions), Admin consent versus User consent tabs, revocation limits, Graph query for oauth2PermissionGrants; page updated 2026-02-19.
- Google Workspace Admin: Control which apps access Google Workspace data: Menu path (Security, Access and data control, API controls, Manage App Access) and the Trusted, Limited, Specific Google data and Blocked settings.
- Microsoft Learn: Cloud app discovery overview (Defender for Cloud Apps): Traffic-log analysis against a catalogue of over 31,000 apps, snapshot and continuous reports, supported firewalls and proxies, the limit that uncatalogued apps are not discovered by default.
- OWASP GenAI: LLM06:2025 Excessive Agency: Background on limiting permissions and monitoring extension activity, used for the advice on reviewing broad app grants.
Topics
- shadow AI
- discovery
- OAuth
- governance
- data protection
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-detect-shadow-ai.