Template

Automate a document workflow: classify, extract, validate, route

This design reads documents as they arrive, decides what each one is, extracts the fields you need into a fixed schema, checks them, and routes clean ones to a person for approval. No ready-made extraction workflow exists in the library.

Last reviewed 7 October 2026

Status

Designed

No workflow, chatflow or MCP server in the library reads and extracts documents. The library has an extraction prompt, two sample Marketplace listings and document-reading integrations, which are support or samples, not a template.

No template covers this job yet. The steps below are the shape we would build and the shape you can build in Studio yourself.

The job

What this template does

Documents arrive in bulk and in mixed form: invoices, forms, letters, delivery notes, in an inbox or a shared folder. The owner is whoever depends on the data in them, often finance or operations. By hand, someone opens each file, decides what it is and types fields into another system, and errors surface later as a mismatch nobody can trace.

This design splits the work. Reading and classifying come first, extraction into a fixed schema second, and validation, by rules you write, third. A person approves before a record is written anywhere. The page covers incoming documents of any type. Filing and retention of documents already held are on the document management page.

Workflow

The workflow, step by step

Steps marked as approval gates pause the run until a named person approves. Nothing after a gate runs before that decision, and the decision is recorded.

  1. 01

    Receive the document

    A file arrives by email, upload or shared folder and starts a run. The flow stores the original untouched, with sender, time and a content hash, and checks it is a type you accept. Others are set aside with a note.
  2. 02

    Read the contents

    The flow extracts text from the file. A PDF with a text layer is read directly. A scan goes to an extraction service or a model that reads images. The text and a reading-quality flag go forward.
  3. 03

    Classify

    An agent decides the document type from a list you supply, such as invoice, purchase order or delivery note, and gives a reason. Low confidence or an unknown type stops the run and sends the file to a person.
  4. 04

    Extract to a schema

    The agent fills the schema for that type, for example number, date, supplier and totals. For each field it also returns the page and the quoted words it came from. A field it cannot find stays empty rather than guessed.
  5. 05

    Validate

    A code step checks types and formats, that line items add up to totals, that required fields exist, and that the document matches a known record, such as an order. Failures go to an exception queue with the reason, not on to approval.
  6. 06

    Approver confirms the record

    The Studio approval step pauses the run for a named person, who sees the document beside the extracted fields and approves, corrects or rejects. Assign someone other than the person who requested the document, where your policy asks for two.Approval gate: a named person approves before the next step runs.
  7. 07

    Write and record

    After approval the flow writes the record to your system and stores the original, the extraction with its quotes, the validation results, the approver and the time. A rejected document is kept with the reason.
Controls

What it can do, cannot do, needs approval for, and records

Can

  • Read text from PDFs and scans
  • Classify a document against your type list
  • Extract fields with the quote each came from
  • Check totals and match against known records

Cannot

  • Write a record before a person approves
  • Guess a field it cannot find
  • Judge whether a document is genuine
  • Read handwriting reliably

Requires approval

  • Every record written to a system of record
  • Documents with a validation failure that a person wants to override
  • Any new document type before it is switched on

Records

  • The original file and its hash
  • Classification and reason
  • Extracted fields with source quotes
  • Validation results
  • Approver, corrections and time
Honest labels

What the library holds for this job

AssetTypeCoversRole in this flow
Extract Key Information from TextPromptUsed inside the flowA starting prompt for pulling named fields from text.
Invoice extraction → ERPMarketplace sample listingSample listing, no template behind itA sample Marketplace listing for OCR, extract, validate and post; no template is seeded behind it.
Invoice extraction workflowMarketplace sample listingSample listing, no template behind itA second sample listing for the same job as a workflow; nothing is seeded behind it.
ReadPdfIntegrationUsed inside the flowReads the text layer of a PDF.
MindeeIntegrationUsed inside the flowA document extraction service that can be called when files are scans.
Read from the template, prompt, MCP and Marketplace data on the site. A Marketplace sample listing is a catalogue preview with no template seeded behind it.
Connections

Integrations this flow uses

  • Read PDF: Pulls the text layer out of PDF files.
  • Mindee: An extraction service for scanned or image-based documents.
  • Email (IMAP): Collects documents from a shared inbox.
  • QuickBooks: An example system the approved record is written to.

The full list of tools Studio connects to is on the integrations page.

Before you start

What you supply, and what this page does not cover

  • You supply the type list, the schema for each type and the validation rules. The template holds none.
  • Extraction accuracy depends on scan quality, layout and the model. This page states no accuracy figure, so measure it on your own documents before relying on it.
  • Where the approved record is written, and the permissions to write there, are yours to set up.

Common questions

How accurate is extraction?
We do not state a figure, because it depends on your documents. Test on a few hundred of your own, compare with what people typed, and keep the approval step on until the errors you find are rare and harmless.
Why keep the source quote for each field?
It lets the approver check a field in seconds by looking at the quoted words, rather than hunting through the page. It also shows when the agent has misread. A field without a quote should be treated as unsupported.
Is there an invoice workflow I can switch on?
Two sample Marketplace listings describe invoice extraction, but nothing is seeded behind either. Treat them as descriptions of the job. The invoice processing template gives the steps for that case specifically.
What about contracts and long documents?
Long documents need chunking and clause-level checks, which this design does not describe. The contract review template covers one type of long document with named reviewers, and is the better starting point there.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.