Definition

Document ingestion is the front door of a document pipeline — accepting files from every channel a borrower uses, normalising them into a consistent internal format, de-duplicating and versioning them, and attaching each one to the right loan file before any reading or analysis begins.

Every channel, one pipeline Versioning and de-duplication Runs before extraction

The Unglamorous Step That Decides Everything After It

Ingestion attracts far less attention than extraction or analysis, and it is where a surprising share of lending delay actually originates. Documents do not arrive as a tidy package. They arrive as an email attachment on Monday, a portal upload on Wednesday, a scan handed over at a branch, and a corrected version sent because the first one was the wrong year.

If that stream is not organised on arrival, every downstream step inherits the disorder. Analysts open the wrong version. Someone re-requests a document the borrower already sent. Two people work from different copies of the same statement. None of these are extraction failures, but they consume the days that extraction was supposed to save.

What Ingestion Handles

  • Multi-channel intake: email, secure portal, direct upload, scanner output, and system-to-system transfer, all landing in one pipeline.
  • Format normalisation: converting the range of formats a borrower sends — PDFs, images, spreadsheets, photographs of paper — into a consistent internal representation.
  • Splitting and assembly: separating a single combined PDF into its constituent documents, and joining files that arrived in pieces.
  • De-duplication and versioning: recognising that a newly arrived file supersedes an earlier one, and keeping both with a clear order.
  • Association: attaching each document to the correct borrower, entity, and loan file rather than a general inbox.
  • Completeness tracking: maintaining a live view of what has arrived against what the credit requires.

Ingestion, Capture, and Extraction

StepQuestion it answersOutput
IngestionWhat arrived, from whom, and for which file?An organised, versioned document set
Capture / OCRWhat characters are on these pages?Machine-readable text
ClassificationWhat kind of document is each one?A labelled document type
ExtractionWhich values matter and what are they?Structured field-and-value data

Why Completeness Tracking Changes the Timeline

The single most valuable thing ingestion produces is an early, accurate answer to what is missing.

In a manual process, gaps surface when an analyst reaches them — often days into the work, after the borrower has mentally moved on. Each discovery starts a new request-and-wait cycle, and those cycles, not the analysis itself, are usually what stretch a commercial credit’s elapsed time.

When ingestion classifies documents on arrival and checks them against what the credit requires, the gap list is available at the start. One consolidated request replaces several sequential ones, and the analyst begins work on a package that is already complete.

How Uptiq Handles Ingestion

Uptiq’s intake agent accepts borrower packages in whatever form they arrive, classifies each document, splits combined files, and identifies what is missing before analytical work begins — so the gap list is available at the start rather than three days in. Ingested documents feed extraction, spreading, and credit memo generation on the same data model, with each figure traced to its source page. Purpose-built lending AI reaches 95%+ accuracy on document extraction, including 150-page unstructured financial statements.


Frequently Asked Questions

What is document ingestion?
Document ingestion is the front door of a document pipeline — accepting files from every channel a borrower uses, normalising them into a consistent internal format, de-duplicating and versioning them, and attaching each one to the right loan file before any reading or analysis begins.
How is ingestion different from document capture?
Ingestion answers what arrived, from whom, and for which file, and produces an organised versioned document set. Capture answers what characters are on the pages and produces machine-readable text. Ingestion is about handling and organisation; capture is about reading.
Why does ingestion matter if extraction is automated?
Because disorganised input produces disorganised output regardless of extraction quality. If versions are not tracked and documents are not attached to the right file, analysts work from superseded copies and re-request documents the borrower already sent. Those delays are not extraction failures, but they consume the time extraction was meant to save.
How does ingestion shorten a commercial credit timeline?
By identifying gaps at the start. In a manual process, missing documents surface as an analyst reaches them, each starting a new request-and-wait cycle. When ingestion classifies documents on arrival and checks them against what the credit requires, one consolidated request replaces several sequential ones.
Can ingestion handle a single combined PDF?
Yes. Splitting a combined file into its constituent documents is a standard ingestion function, as is assembling documents that arrived in pieces. Borrowers frequently send several years of statements and returns as one file, so handling this on arrival avoids manual separation later.
Uptiq Qore Platform
See what a complete package on day one looks like

Talk to a lending automation expert about your intake process.