TechOneDigital Start a pilot

Finance & Documents

AI Document Intake

Documents arrive by web, link, API or mailbox and are sorted before anyone opens them.

Available now · Human approval · Complete audit trail

The operational problem

Extraction cannot repair a document that entered the wrong queue

Files arrive through different channels, one PDF can contain several documents and similar layouts may represent different business processes. If intake guesses incorrectly, every downstream field can be precise and still belong to the wrong record.

The service gives every arrival a traceable identity, separates bundles, classifies only supported types and holds unknowns for a person before extraction or delivery begins.

Intake controls

What happens before extraction begins

Four channels

Web upload, shared link, API and mailbox use one intake model while preserving their source context.

File identity

Hash, arrival time, channel and source reference support traceability and duplicate control.

Document boundaries

Bundles and multi-document PDFs are divided into the records downstream steps expect.

Type classification

Only the document types included in the agreed catalogue are routed automatically.

Unknown queue

Unsupported or ambiguous arrivals wait visibly instead of being forced into the closest type.

Method

How incoming files become a reliable queue

  1. Register the arrivalThe file and source channel receive a stable intake identity.
  2. Check duplicates and readabilityRepeated content and files that cannot be processed are identified before classification.
  3. Detect document boundariesPages are grouped into individual documents with the original bundle relationship retained.
  4. Classify against the catalogueEach document is matched only to types approved for the workflow.
  5. Hold uncertaintyLow-confidence or unsupported types are sent to the unknown queue with the source visible.
  6. Release to the right templateConfirmed documents enter the extraction or storage queue associated with their type.

Example intake

Types, splits and unknowns remain visible

Illustrative intake batch — no customer documents

Received
212
Types found
Invoice, delivery note, contract, +4
Split
11 PDFs into 46 documents
Unknown
3, waiting
Channels
Web, link, API, mailbox

Document judgement

The rules behind dependable intake

Channel is part of provenance

A mailbox attachment and an API upload may carry different source evidence and ownership.

Split before classification

A mixed PDF is not treated as one record merely because it arrived as one file.

The catalogue defines automation

Unknown types do not inherit a template from the nearest-looking document.

Duplicates need stable identity

Retries and repeated uploads remain traceable without silently creating downstream copies.

Fixed start

What the 14-day intake pilot delivers

  • One representative batchFiles from the agreed channels run through registration, splitting and classification.
  • Document-type catalogueSupported types, routing destination and uncertainty threshold documented.
  • Boundary and classification resultsThe pilot shows which documents were split, classified, duplicated or held.
  • Unknown-item workflowA person can resolve unsupported and ambiguous arrivals without losing source context.
  • Next-stage mapEach confirmed type is connected to the right extraction, storage or review path.

Fit

When automated intake is—and is not—ready

A good fit

  • Documents arrive through several repeatable channels and need one queue.
  • Bundles or mixed PDFs create manual sorting work.
  • A useful catalogue of supported document types can be defined.

Not the right fit

  • Every file is unique and no downstream routing rule exists.
  • The source channel cannot provide stable file access or traceability.
  • The expectation is to classify unknown types without any human exception path.

Where the human approves

You resolve unknown types. The AI works through a fixed list of allowed operations; every approval and result is logged.

Intake questions

Questions a document-process owner should ask

Which intake channels are supported?

The platform supports web upload, shared links, API and mailbox intake. A pilot can use one or more agreed channels.

Can one PDF be split into several documents?

Yes. The pilot tests page boundaries and retains the relationship to the original file. Ambiguous splits are held for review.

What happens to an unknown document type?

It stays in an explicit unknown queue with its source and file available to a person. It is not silently routed to a similar template.

How are duplicate uploads handled?

Stable intake identity and content evidence allow repeats to be flagged before downstream records are created.

Does intake also extract fields?

Intake prepares the correct document and type. Field extraction is the next controlled stage and uses a schema specific to that document type.

Experience behind the service

Built on multi-channel document intake

The service uses SmartDocto’s four intake channels, document-type catalogue, bundle separation and unknown-item handling. The platform currently supports twenty document types across five languages.

Technical stewardship: David Máj, Founder & Technology Consultant. Last reviewed 24 September 2026.