Whitepaper ¡ 9 minute read
Document Intelligence Architecture: An Enterprise Whitepaper
Document intelligence architecture separates ingestion, classification, layout-aware parsing, field extraction, validation, and human review, with per-field confidence deciding what reaches a person. Systems fail when they treat extraction as one step, produce confidence only per document, or lack a review loop that feeds corrections back into evaluation.
Document processing is where most enterprises start with AI, because the work is visible, repetitive, and expensive. It is also where they most often end up with a second project eighteen months later, rebuilding a pipeline that worked in the pilot and did not survive contact with the real intake. The cause is almost always architectural: extraction treated as a single step, confidence expressed per document, no validation layer, and a human review process bolted on afterwards. This whitepaper sets out the architecture that holds. It draws on FISTA Solutions' AI agents delivery across document-heavy operations and complements how to build a document ai system and ocr vs llm document extraction.
What are the stages, and why separate them?
| Stage | Responsibility | Failure if merged |
|---|---|---|
| Ingestion | Receive, normalise, split multi-document files | Pages from one document extracted into another |
| Classification | Identify document type and variant | One extraction schema applied to everything |
| Parsing | Layout-aware text, tables, and positions | Table values read in reading order, not structure |
| Extraction | Fields per type with per-field confidence | No basis for routing |
| Validation | Business rules, arithmetic, cross-references | Plausible but wrong values accepted |
| Routing | Accept, review, or reject by field | Whole documents reviewed or none |
| Review | Human correction with structured capture | Corrections lost, no learning |
| Writeback | Structured output to systems of record | Results stranded in a report |
Each stage has a distinct failure mode and a distinct accuracy measure. Systems built as one model call from PDF to JSON cannot be diagnosed when they degrade, because there is nowhere to look.
What does ingestion actually have to handle?
More than teams plan for. Real intake arrives as email attachments with several documents in one PDF, photographs taken at an angle, faxes, spreadsheets exported as images, password-protected files, files whose extension does not match their content, and the same supplier's invoice in a new template because their finance system was upgraded.
Ingestion therefore needs format detection independent of extension, deskewing and image enhancement, multi-document splitting, page ordering, and a quarantine path for anything it cannot process rather than a silent drop. The quarantine queue is one of the more useful operational signals in the system: its contents tell you what the world is actually sending.
Why classify before extracting?
Because extraction schemas are type-specific and a single model applied to everything performs worse on every type. Classification also lets the pipeline handle the unknown gracefully: a document that does not match a known type goes to review rather than being extracted against an arbitrary schema and producing confident nonsense.
Classification should be measured separately, since a misclassification guarantees extraction failure downstream, and its errors are usually concentrated in a few confusable pairs that targeted examples fix quickly. See how to build a document classification system.
How should parsing and extraction combine?
Layout-aware parsing establishes what text exists and where, including table structure, which language models handle poorly when given raw text. Extraction then reads those structures to produce fields, using the model's strength: tolerating variation in wording, locating fields that move between layouts, and inferring values that require reading across sections.
The pattern that performs best on real documents uses both rather than choosing: deterministic extraction where a field is reliably positioned or pattern-matched, model-based extraction where it is not, and agreement between the two as a confidence signal. Purely model-based pipelines struggle with tables and numeric precision; purely rule-based pipelines break on every template change.
Why is per-field confidence the architectural keystone?
Because it is what makes selective review possible, and selective review is what makes the economics work. A document is rarely uniformly reliable: the total may be unambiguous while the purchase order reference is smudged. Per-document confidence forces a choice between reviewing everything and accepting the bad field silently.
With per-field confidence, thresholds are set per field according to the cost of error. An incorrect invoice total is expensive; an incorrect description is not. High-cost fields get high thresholds and more review; low-cost fields are accepted more readily. That single design decision typically reduces review volume substantially while improving the accuracy of what matters.
Confidence must also be calibrated, meaning a field marked ninety percent confident should be right about ninety percent of the time. Uncalibrated scores make thresholds meaningless. See what is model calibration.
What does the validation layer do?
Catches what confidence cannot see. A model can be entirely confident about a value that is wrong in context: line items that do not sum to the stated total, a date outside a plausible range, a supplier not in the master data, a currency inconsistent with the country, a quantity exceeding the purchase order.
Validation rules are business logic, written with the operations team, and they frequently catch more real errors than confidence thresholds do. They also produce better review tasks, since a validation failure tells the reviewer what is wrong rather than merely that the system was unsure.
How should the review queue be designed?
As part of the product, not an exception path. Reviewers should see the document image with the extracted field highlighted in place, the proposed value, the reason for review, and a single-keystroke accept or correct. Time per review is the metric that decides whether the economics work, and a well-designed interface differs from a poor one by a factor of several.
Structured capture matters as much as speed. Every correction records the original value, the corrected value, the document type, and the field, which becomes both the accuracy measurement and the training or evaluation signal. Reviews that overwrite a value without recording what changed discard the most valuable data the system produces. See how to build a human review queue.
How is accuracy measured?
Per field and per document type, against a labelled set drawn from real intake rather than clean samples. Exact match for identifiers and codes; tolerance-based match for amounts and dates; and separate tracking of the two error types that matter operationally: fields sent to review that were actually correct, which costs time, and fields accepted that were wrong, which costs money.
Aggregate accuracy across all fields and types is close to meaningless, because it is dominated by easy fields on common document types and hides the expensive failures on the rare ones. Reporting should always be segmented.
How does the system avoid drifting?
Documents change continuously: suppliers change templates, regulations change forms, and new document types enter the intake without announcement. Accuracy therefore decays unless something watches for it.
The mechanisms that work: monitor per-type volume and confidence distributions, since a template change usually shows as a confidence drop before an accuracy complaint; sample accepted extractions for audit at a low rate, which catches silent degradation; feed every review correction into the evaluation set; and re-run evaluation on a schedule as well as on change. See what is model drift.
What does writeback require?
Idempotency and traceability. Documents are reprocessed, resubmitted, and duplicated more often than teams expect, so writes to systems of record must be safe to repeat. Every written value carries a reference to its source document, page, and extraction run, so any figure in the target system can be traced to the image it came from, which is what auditors ask for.
How is the economics modelled?
Cost per document processed, split into inference, infrastructure, and human review time, compared against the fully loaded cost of the manual process. Review volume is the dominant variable, which is why per-field confidence and validation quality drive the business case more than model choice does.
The trajectory that makes projects succeed is a review rate that falls over the first two quarters as thresholds are tuned, validation rules are added, and new document types are absorbed. A project whose review rate is flat after six months has a pipeline problem, not a model problem.
What goes wrong?
Pilots on clean samples. Per-document confidence. No validation layer. Review interfaces that require reading the whole document. Corrections not captured. One schema for all types. No quarantine path, so unprocessable documents disappear. Writeback without idempotency. And accuracy reported as a single number, which prevents anyone from seeing that the expensive field on the rare document type has been wrong since launch.
How does this differ by domain?
Finance documents carry arithmetic validation and strong master-data cross-referencing, which makes validation unusually powerful. Insurance and healthcare add privacy constraints on where images may be processed. Logistics documents arrive in the widest format variety and benefit most from robust ingestion. Legal and construction documents are long, so retrieval within a document matters more than field extraction. The architecture is common; the validation rules and thresholds are domain work.
How should a first project be scoped?
Around one document type with high volume and a countable cost, processed end to end into a system of record. Invoices, proofs of delivery, applications, and claims forms all qualify. What does not work is scoping around a repository or a department, because that produces a long discovery with no shipped pipeline.
The scope should explicitly include the unglamorous parts that projects cut when time runs short: the quarantine path, the review interface, correction capture, and writeback idempotency. Cutting those produces a demo that extracts fields and an operation that still cannot use the results.
A realistic first delivery runs eight to twelve weeks: two weeks assembling a labelled set from real intake, four to six weeks building ingestion through validation, two weeks on the review interface and writeback, and two weeks tuning thresholds against measured review volume and error rates before the operation depends on it.
How FISTA Solutions delivers this
FISTA Solutions builds document intelligence with staged architecture, per-field calibrated confidence, business-rule validation, purpose-built review interfaces, and correction capture that keeps evaluation current, so accuracy improves after launch rather than decaying, through AI enablement, AI agents, and forward deployed engineers working with operations teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime and 47% efficiency gains where measured.
To build document processing that survives the real mailroom, message FISTA on WhatsApp, or read how to build a document ai system.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do document AI projects get rebuilt so often?
Because they are designed against clean sample documents and meet the real intake later: scans at odd angles, multi-document PDFs, handwritten annotations, formats that changed last quarter, and suppliers who send spreadsheets as images. Architectures without confidence routing and review cannot absorb that variety.
02Why does confidence need to be per field?
Because a document is rarely uniformly good or bad. An invoice may have a clear total and an ambiguous purchase order number, and per-document confidence forces either wasteful full review or silent acceptance of the bad field. Per-field routing sends one value to a person and accepts the rest.
03Should extraction use OCR or a language model?
Usually both. Layout-aware OCR establishes text position and structure, which language models use poorly on their own, while the model handles variation in wording, inference across sections, and fields whose location is not fixed. The combination outperforms either alone on real documents.
04How is document AI accuracy measured?
Per field and per document type against a labelled set drawn from real intake, tracking exact match for identifiers, tolerance-based match for amounts and dates, and the rate at which incorrect values were accepted without review, which is the number that costs money.
05What role does human review play?
It handles low-confidence fields, validation failures, and unknown document types, and it supplies the labelled corrections that keep evaluation current. Review volume is a design target to reduce over time, not a failure to eliminate immediately.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.