Finance & Documents
AI Document Intake
Documents arrive by web, link, API or mailbox and are sorted before anyone opens them.
Available now · Human approval · Complete audit trail
The operational problem
Extraction cannot repair a document that entered the wrong queue
Files arrive through different channels, one PDF can contain several documents and similar layouts may represent different business processes. If intake guesses incorrectly, every downstream field can be precise and still belong to the wrong record.
The service gives every arrival a traceable identity, separates bundles, classifies only supported types and holds unknowns for a person before extraction or delivery begins.
Intake controls
What happens before extraction begins
Four channels
Web upload, shared link, API and mailbox use one intake model while preserving their source context.
File identity
Hash, arrival time, channel and source reference support traceability and duplicate control.
Document boundaries
Bundles and multi-document PDFs are divided into the records downstream steps expect.
Type classification
Only the document types included in the agreed catalogue are routed automatically.
Unknown queue
Unsupported or ambiguous arrivals wait visibly instead of being forced into the closest type.
Method
How incoming files become a reliable queue
- Register the arrivalThe file and source channel receive a stable intake identity.
- Check duplicates and readabilityRepeated content and files that cannot be processed are identified before classification.
- Detect document boundariesPages are grouped into individual documents with the original bundle relationship retained.
- Classify against the catalogueEach document is matched only to types approved for the workflow.
- Hold uncertaintyLow-confidence or unsupported types are sent to the unknown queue with the source visible.
- Release to the right templateConfirmed documents enter the extraction or storage queue associated with their type.
Example intake
Types, splits and unknowns remain visible
Illustrative intake batch — no customer documents
- Received
- 212
- Types found
- Invoice, delivery note, contract, +4
- Split
- 11 PDFs into 46 documents
- Unknown
- 3, waiting
- Channels
- Web, link, API, mailbox
Document judgement
The rules behind dependable intake
Channel is part of provenance
A mailbox attachment and an API upload may carry different source evidence and ownership.
Split before classification
A mixed PDF is not treated as one record merely because it arrived as one file.
The catalogue defines automation
Unknown types do not inherit a template from the nearest-looking document.
Duplicates need stable identity
Retries and repeated uploads remain traceable without silently creating downstream copies.
Fixed start
What the 14-day intake pilot delivers
- One representative batchFiles from the agreed channels run through registration, splitting and classification.
- Document-type catalogueSupported types, routing destination and uncertainty threshold documented.
- Boundary and classification resultsThe pilot shows which documents were split, classified, duplicated or held.
- Unknown-item workflowA person can resolve unsupported and ambiguous arrivals without losing source context.
- Next-stage mapEach confirmed type is connected to the right extraction, storage or review path.
Fit
When automated intake is—and is not—ready
A good fit
- Documents arrive through several repeatable channels and need one queue.
- Bundles or mixed PDFs create manual sorting work.
- A useful catalogue of supported document types can be defined.
Not the right fit
- Every file is unique and no downstream routing rule exists.
- The source channel cannot provide stable file access or traceability.
- The expectation is to classify unknown types without any human exception path.
Where the human approves
You resolve unknown types. The AI works through a fixed list of allowed operations; every approval and result is logged.
Intake questions
Questions a document-process owner should ask
Which intake channels are supported?
The platform supports web upload, shared links, API and mailbox intake. A pilot can use one or more agreed channels.
Can one PDF be split into several documents?
Yes. The pilot tests page boundaries and retains the relationship to the original file. Ambiguous splits are held for review.
What happens to an unknown document type?
It stays in an explicit unknown queue with its source and file available to a person. It is not silently routed to a similar template.
How are duplicate uploads handled?
Stable intake identity and content evidence allow repeats to be flagged before downstream records are created.
Does intake also extract fields?
Intake prepares the correct document and type. Field extraction is the next controlled stage and uses a schema specific to that document type.
Experience behind the service
Built on multi-channel document intake
The service uses SmartDocto’s four intake channels, document-type catalogue, bundle separation and unknown-item handling. The platform currently supports twenty document types across five languages.
Technical stewardship: David Máj, Founder & Technology Consultant. Last reviewed 24 September 2026.