Data protection

DLP tools: what data loss prevention does and how to judge it for AI

What DLP is, the types, how detection works, the new AI problem and a checklist, with Swfte's gateway control placed honestly beside it.

Data loss prevention (DLP) tools watch where sensitive data goes and warn, block or redact when a rule matches. Classic DLP covers email, files, devices and cloud apps. AI adds prompts, uploads to chat tools, agent tool calls and retrieval that can surface documents. Swfte's Connect gateway has a content-policy control for model traffic. That is a gateway control, not an enterprise DLP suite.

Last verified 2026-10-07. Sources are listed at the end of the page.

What is data loss prevention?

Microsoft describes DLP as a practice that prevents users from inappropriately sharing sensitive data, such as financial data, health records and card numbers, with people who should not have it. In Microsoft Purview you define DLP policies that identify, monitor and automatically protect sensitive items across enterprise applications and devices, and across inline web traffic. Its documentation says DLP uses deep content analysis and not a simple text scan.

The protective actions Microsoft lists are a policy tip that warns the user, a block the user can override with a justification, a block with no override, quarantine for data at rest, and hiding sensitive content in Teams chat. Every monitored activity is recorded to an audit log by default.

The same page says to plan policies, deploy them in simulation mode and evaluate their impact before running more restrictive modes. Expect tuning, not a switch-on.

What types of DLP tool are there?

The five types are a common industry grouping, not a standard. Microsoft's documentation splits its own product into "Enterprise applications and devices" and "Inline web traffic", so vendors will not map one to one.

TypeWhere it watchesWhat a vendor page says about it
Network DLPTraffic leaving the network or browser.Microsoft says inline coverage reaches unmanaged cloud apps through Microsoft Edge for business and network data security via SASE integrations.
Endpoint DLPWindows and macOS devices.Microsoft lists activities it can audit or restrict: upload to a restricted cloud service, paste to supported browsers, copy to clipboard, USB, network share and print.
Cloud and CASB DLPSanctioned and unsanctioned cloud apps.Microsoft lists non-Microsoft connected apps (marked preview) including Box, Dropbox, Google Workspace and Salesforce, and a catalogue of over 35,000 cloud apps for the network option.
Email DLPMessages in transit and at rest.Microsoft applies policies to Exchange Online. Its alerts note says DLP does not scan or match previously existing email stored in a mailbox or archive.
Data-centric DLPThe data itself, wherever it sits.Google describes Sensitive Data Protection as a service to discover, classify and de-identify sensitive data inside and outside Google Cloud, with data profiles and de-identification such as masking, redaction and tokenisation.

Sources: Microsoft Learn (DLP and Endpoint DLP pages) and Google Cloud documentation, read on 2026-10-07.

How does DLP detection work?

Each method trades false positives against false negatives. The limits below are the ones the vendor pages state.

MethodHow the vendor documentation describes itLimit it states
Patterns and keywordsMicrosoft sensitive information types use a primary element, such as a regular expression with or without a checksum, plus supporting evidence within a set proximity, and a confidence level.Higher confidence returns fewer false positives but more false negatives, and the reverse.
Built-in and custom detectorsGoogle uses built-in infoType detectors and supports custom regular expression, dictionary and metadata label detectors.The page states no limit for detectors. Test each one on your own data.
FingerprintingMicrosoft converts a standard form into a hash-based detector, with partial or exact matching. The original document is not stored.It does not detect password-protected files, image-only files, files over 4 MB, or documents missing the original form text.
Exact data matchMicrosoft compares one-way hashes of content against a table of your own sensitive values, such as customer records.It needs a source table you maintain. The page states a table can hold up to 100 million rows.
Machine-learning classifiersMicrosoft lists machine learning algorithms and trainable classifiers among its detection methods.Results vary by workload, so validate in each place you enforce.

What is the new problem AI adds?

Prompts and uploads to chat tools

People paste text and upload files into chat tools. Microsoft lists OpenAI ChatGPT, Google Gemini, DeepSeek and Microsoft Copilot among the cloud apps its browser and network options cover. Coverage depends on the option and licence you use.

Agents with tool access

An agent can read a file, call an API or send a message without a person pasting anything. The vendor pages we read do not say how their tools treat actions by service identities, so ask each vendor and test it.

Retrieval that surfaces documents

A retrieval system can place a document in a model's context that the user could not open directly if access lists are not enforced. The leak happens in the answer, not in a file transfer. See how to build a RAG system.

Data that never becomes a file

Microsoft states that Endpoint DLP cannot scan or classify data that is never saved to a file on the device. Text typed into a web chat may fall outside an endpoint rule, which is why inline web coverage and gateway controls matter.

What should an AI-era DLP evaluation check?

Run these against your own traffic in simulation mode before you commit.

  • Which AI apps and sites does the tool name as covered, and through which option (browser, network, API)?
  • Does it inspect prompts and uploads as well as files at rest?
  • Can it enforce on actions by agents and service identities as well as human users?
  • Which detection methods does it offer, and can you add your own patterns and exact values?
  • What is the false positive rate on your data in simulation mode?
  • What happens when content is not a file, is encrypted, or exceeds a size limit?
  • Are actions warn, block, override with justification and redact all available, and are decisions logged?
  • Can you see which model provider received which data, for third-party risk?
  • What does the vendor's page say is unsupported? Read it before the demonstration.

Where Swfte fits

Swfte Connect, the model gateway, has a content-policy evaluator. It is Built: it was read in code on 2026-10-06. It has built-in detectors for AWS keys, GitHub tokens, private-key blocks, JWTs, email addresses, phone numbers, an SSN pattern, a card-number pattern and a PII pack, plus custom regular expressions, and it can take a REDACT action that replaces a match with a placeholder.

It is a gateway control for model traffic. It applies to requests that pass through Connect. It does not see endpoint activity, email or file-share traffic, and it does not see requests that bypass the gateway, such as the free Connect dashboard in direct mode, which works without routing prompts through the Swfte gateway. The detectors recorded are pattern-based, and no fingerprinting or exact data match is claimed. It is not an enterprise DLP suite.

You do not need Swfte for this if your DLP product already covers browser and API traffic to AI services and you do not run your own agents. If you want a policy applied where a model is called, see Connect and the application-level gateway. To find the AI use that bypasses any control, start with shadow AI and shadow AI discovery.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is a DLP tool?

A DLP tool monitors data in use, in motion and at rest, and takes a protective action when a rule matches. Microsoft describes its DLP policies as identifying, monitoring and automatically protecting sensitive items, with actions that include warning the user, blocking with or without override, and quarantine for stored data.

Can DLP stop employees pasting data into ChatGPT?

It can for the apps and options a vendor names. Microsoft lists OpenAI ChatGPT, Google Gemini, DeepSeek and Microsoft Copilot as locations for its browser and network options. Coverage depends on how users reach the app, the licence and the configuration, so test your own paths in simulation mode.

What is exact data match?

Exact data match compares content against exact values in a table of your own sensitive data, using one-way hashes so the values are not shared in clear. Microsoft says it gives fewer false positives than generic patterns, works with structured data and supports tables of up to 100 million rows.

Is an AI gateway the same as DLP?

No. A gateway sees traffic to model APIs that passes through it. DLP spans endpoints, email, cloud apps and file shares. A gateway control such as redaction in Connect covers model traffic only, so it complements DLP and does not replace it.

What does Swfte detect and redact?

The Connect content-policy evaluator has built-in detectors for AWS keys, GitHub tokens, private-key blocks, JWTs, email, phone numbers, an SSN pattern, a card-number pattern and a PII pack, plus custom regular expressions. A rule can redact a match. It applies to traffic through the gateway only.

Does DLP replace shadow AI discovery?

No. DLP acts on data once a rule matches, in the places it is deployed. Discovery finds AI tools and keys nobody told you about, so you can bring them under policy. You usually need both: discovery to find the use, and controls to apply to it.

Test your DLP on prompts and agent actions before you rely on it

Automate the response with SecOps Agents

Autonomous security orchestration: triage, investigation and containment, with a full audit trail.