Customers

Pricing
Introducing r-1: Reducto’s new SOTA document parsing model
Extract

Extract structured data from any document

Extract pulls structured fields from any document using a schema you define. One call in, typed JSON out.

Helping everyone from startups to Fortune 10 enterprises unlock their data.

  • Harvey
  • Scale AI
  • Newfront
  • Medallion
  • Vanta
  • Legora
  • Rogo
  • Levelpath
  • JLL
  • Vise
  • Laurel
  • Toast
  • Mercor
  • Zip
  • Anterior
  • Supio
  • EliseAI
  • Unify
Extract API

Define a schema, get structured JSON back

Definition
The Extract API returns specific fields from any document as schema-typed JSON. Define your schema or use a system prompt to pull the values you need with citations on every one.
Who it's for
Teams that need accurate data fields extracted fast and efficiently from each document without writing per-template parsers. Starts at $20/1000 pages, with no additional charge for parsing.
The problem it solves
Off-the-shelf LLMs hallucinate fields and drift across runs. Extract grounds every value to the page it came from and constrains output to your schema, so results are consistent and auditable.
The agentic document platform

Document work starts here

Try out Extract in Studio or via the API.

Deep Extract

A verification loop for complex documents

What it is
Deep Extract runs an agentic loop that checks its output against the source and re-extracts until it reaches the quality threshold. Enable it with settings.deep_extract: true.
Built for
It is designed for multi-page tables and other complex documents where a standard pass may miss values, truncate rows, or produce inconsistencies.
Independent benchmark
In micro1's independent benchmark on long-document extraction, Reducto Deep Extract was evaluated on 225 documents averaging 358 pages.
Where AI teams ship Extract

Extract the data you need

If your workflow ends with writing fields to a database, Extract is the step that gets them there accurately.

Invoice & AP automation

Pull header fields, taxes, and every line item into typed JSON. Citations let AP teams verify amounts quickly.

Contract & clause data

Effective date, expiration, parties, governing law, renewal terms. Define the fields once and Extract handles layout variations.

Financial statements & filings

Pull totals, holdings, and transactions from 10-Ks, brokerage statements, and fund factsheets. Deep Extract handles complex tables.

KYC, claims, and onboarding

Identity, employer, address, claim numbers, dates of loss. Citations on every value make audit straightforward.

Long arrays & transaction lists

Bank statements, ledgers, claim line items: Deep Extract verifies every field with an agentic loop so nothing is missed across long documents.

Extract across multiple files

Combine fields from several documents into a single schema response for data rooms, claim packets, and onboarding.

Try out Extract in Studio or via the API.

Why Extract

Why teams switch to Extract

  1. 01

    Schema-typed, every time

    Output shape matches your schema. Enums normalize values, so downstream code never has to translate “Invoice” vs “INVOICE.”

  2. 02

    Citations on every value

    citations wraps each field with page, bbox, source text, and confidence for both extract and parse stages.

  3. 03

    Complete extraction on long docs

    Deep Extract uses an agentic loop to verify outputs across long documents, so hundreds of line items are captured accurately.

  4. 04

    Deep Extract for complex documents

    An agent harness that extracts, verifies against the source, and re-extracts until results meet your accuracy criteria. Built for long documents with thousands of rows across hundreds of pages.

  5. 05

    Reuse parsed work via jobid://

    Try a different schema on the same doc, or merge fields across many docs, without re-parsing. Pass a job ID or a list as input.

  6. 06

    Schema or schemaless

    Ship a schema for predictable production output. Pass a natural-language prompt for prototyping.

How Extract works

How Extract works in four steps

  1. STEP 01

    Send a file + schema

    Upload a file or point at a URL. Define the fields you want in a schema.

    POST /extract
  2. STEP 02

    Parse runs underneath

    OCR, layout detection, and table reconstruction produce structured content for the extractor to read.

    jobid:// available
  3. STEP 03

    LLM locates each field

    Field names and descriptions guide the model. Array extract handles long lists and Deep Extract iterates for accuracy.

    schema → values
  4. STEP 04

    You get typed JSON

    Output matches your schema with optional citations on every value.

    { value, citations }
Built for production

Enterprise-ready from day one

  • SOC 2 Type II
  • HIPAA
  • Zero Data Retention
  • VPC · On-prem · Air-gapped
  • EU · AU regional endpoints
  • 99.9%+ uptime SLA
  • Enterprise support
Visit the Trust Center

Try out Extract in Studio or via the API.

Further reading

Build better extraction workflows

  • AnnouncementDeep Extract

    See how an agentic review loop improves extraction quality on complex, long-form documents.

  • GuideDocument AI Extraction Schema Tips

    Design schemas that stay precise, maintainable, and useful as your document set evolves.

  • AnnouncementSmart Schema Extract

    Learn how schema optimization makes structured extraction more accurate and easier to operate.

FAQ

Common questions about Extract

Document work starts here

Define the fields. Get cited values back

Drop a PDF in Studio or hit the API with one call. No setup, no credit card.

Reducto logoLLM Center