Why review is still the analyst's job
Extraction accuracy is a real and useful metric, and it is also the easiest thing to measure — which is why it dominates vendor conversations. It answers one question: did the system read the number correctly off the page.
The questions a credit file actually turns on are different:
- Was this line placed in the right row of the template?
- Should this add-back have been applied to this borrower, under this agreement?
- Does this entity belong in the global cash flow, and has the intercompany rent been eliminated once rather than twice?
- Is the debt service in the denominator the debt service the covenant defines?
None of those are extraction. They are the analyst's judgement, and a system that presents them clearly is doing its job; a reviewer who accepts them without looking is not doing theirs.
There is also a straightforward institutional reason to review deliberately. A spread that informs a credit decision, a risk rating, or a covenant test is an input to a regulated process, and lenders are generally expected to be able to show documented controls and human accountability over such inputs. A review with no record is, for practical purposes, a review that did not happen.
The six places errors actually occur
Errors cluster into six categories, and they are not equally dangerous. Sorting them this way is what lets a reviewer spend ten minutes well instead of an hour evenly.
Document-level
Wrong entity, wrong period, a superseded version, or a missing page. Rare but total — everything downstream is invalid. Caught in seconds by checking the document set first.
Extraction
Transposed digits, a negative read as positive because it was in parentheses, thousands read as units, or a column read from the wrong year. Usually caught by tie-outs rather than by reading.
Classification
The figure is right, the row is wrong — an operating expense landing in cost of goods, or officer compensation in general administrative. Totals still tie, so only sampling finds these.
Judgement
An add-back applied where the agreement does not permit it, or a non-recurring item treated as recurring. Nothing ties out wrong. This is the highest-consequence category.
Consolidation
A missing entity, intercompany rent eliminated twice or not at all, or the wrong ownership percentage applied to flow-through income. Second-highest consequence, and easy to miss.
Calculation
Right inputs, wrong formula — debt service that excludes the balloon, or an EBITDA definition that does not match the one in the credit agreement.
Note the pattern: categories 1 and 2 are cheap to catch and mostly mechanical. Categories 4 and 5 are where a spread quietly supports the wrong decision, and no amount of extraction accuracy addresses them.
Start with the tie-outs
Before reading a single line, run the structural checks. They take moments and they catch the errors that would otherwise waste the rest of the review.
| Tie-out | What it catches | If it fails |
|---|---|---|
| Balance sheet balances | Missed line, duplicated line, sign error | Stop; likely a missing page or a misread column |
| Totals agree to the source statement | Scale errors, wrong-year columns, omitted sections | Stop; re-check the document set before anything else |
| K-1 ties to the personal return | Wrong ownership percentage, missing entity | Check the entity map before consolidating |
| Debt schedule agrees to the balance sheet | Omitted obligations, double-counted debt | Fix before any coverage ratio is trusted |
| Prior-year comparison | Anything that moved implausibly year over year | Investigate the outlier line specifically |
| Cash agrees to the statement of cash flows | Classification errors between operating and financing | Review the cash flow mapping |
The prior-year comparison deserves particular attention. A line that triples with no narrative explanation is either a genuine business event worth understanding or an extraction error — and both outcomes are worth the thirty seconds it takes to look.
The review sequence, step by step
A repeatable sequence, ordered so that the cheapest checks eliminate the most work.
Confirm the document set
Right entity, right period, right version, all pages present. Check that an amended return has not been superseded by a later one sitting in the same folder.
Run the tie-outs
Balance sheet, totals to source, K-1 to 1040, debt schedule to balance sheet. Any failure stops the review and sends you back to the documents.
Scan year-over-year
Look down the comparison column for implausible movement. Investigate outliers individually rather than reading every line.
Sample classification
Take the most material lines plus a random handful, click through to the cited source page, and confirm both the figure and the row it landed in.
Review every add-back
One at a time, against the agreement's definition. Depreciation is mechanical; owner compensation, one-time items, and distributions are not.
Verify the consolidation
Every entity present, ownership percentages correct, intercompany rent and management fees eliminated exactly once, guarantor income not double-counted.
Recompute one ratio by hand
Pick the ratio that carries the covenant. If it reconciles, the arithmetic layer is sound; if it does not, you have found something that mattered.
Record the review
What was checked, what was changed and why, and who approved it — captured as you go rather than reconstructed later.
Steps 5 and 6 are the ones to protect when time is short. They are the categories where an error survives every other check in the list.
How much to sample, and when to sample more
Sampling rates should reflect risk, not habit. The useful variables are how familiar the document is, how clean it is, and how much rides on the result.
Review at 100%: first file from a new borrower, an unfamiliar document type, handwritten or poor-quality scans, complex multi-entity structures, or any spread where a covenant is close to its threshold.
The cost of a full review is small relative to the cost of an incorrect first credit decision or a missed breach.
Review at anchors plus material lines: established borrower, familiar document types, clean digital source files, comfortable covenant headroom.
Tie-outs, a year-over-year scan, the top lines by materiality, and full review of add-backs and consolidation.
Escalate temporarily: after any change to the model, the template, or the document format.
Raise the sampling rate for a defined period, then step it back down once observed exceptions return to their baseline.
Let your own data set the rate
Track exceptions by document type and error category over a few months. Most teams find the pattern is narrow — one form, one classification, one recurring add-back question — and once that is known, the sampling rate can be tightened where errors cluster and relaxed everywhere else. That is a defensible, evidence-based control, and it is considerably more useful than a blanket percentage.
Red flags that mean stop and re-spread
Some findings are not corrections. They mean the spread should be rebuilt rather than patched.
- A tie-out fails and the cause is not immediately obvious. Patching the difference into a plug line hides the real error.
- Figures with no source citation. If the platform cannot show where a number came from, it is an assumption, not data.
- The document set turns out to be wrong. A superseded return or a missing entity invalidates everything downstream, including anything already reviewed.
- Multiple classification errors in one sample. Errors cluster; two in a small sample usually means many more outside it.
- Add-backs that do not match the agreement's definitions. Not a spread problem to fix in place — the definition needs confirming before the spread means anything.
Rebuilding feels expensive in the moment. It is much less expensive than a credit memo built on a spread nobody can defend six months later.
Documenting the review
A review that leaves no trace forces the next person to redo it — which is precisely the cost automation was meant to remove.
Retain, against the credit file rather than in a working folder:
- The document set used, with versions and dates.
- Which tie-outs were run and whether each passed.
- Which lines were sampled and on what basis.
- Every override, with the value before and after and the reason for the change.
- Each add-back approved or rejected, with the clause or policy relied on.
- The reviewer, the date, and any second review for higher-risk credits.
The practical test: could a second reviewer follow the record and reach the same conclusion without re-performing the work? If yes, the documentation is sufficient. If no, it is a signature rather than a control.
What to expect from the platform
A platform cannot do the review, but it can make the review fast or make it miserable. The features that matter are the ones that let a human verify instead of re-perform.
Citations on every figure
Click a number and land on the page and line it was read from, rather than searching a PDF for it.
Tie-outs run automatically
Structural checks surfaced before a human opens the spread, so the review starts from what failed.
Overrides with reasons retained
Any figure or classification changeable by the analyst, with the before, after, and rationale kept in the record.
Add-backs flagged, not assumed
Discretionary adjustments presented for a decision rather than applied silently.
Uptiq's Financial Spreading Agent is built to that shape: figures traceable to their source page, overrides recorded with a reason, and the spread handed on to the credit memo, the risk rating, or a covenant test once a human has signed off.
Worth being precise about what that first figure means: extraction accuracy is a measure of reading, and it is the reason a review is fast rather than a reason to skip one. For the wider context, see what financial spreading software does and how covenant monitoring depends on it.
Frequently asked questions
How do you check if an AI-generated financial spread is accurate?
Work in three passes rather than reading every line. First confirm the document set is the right entity, period, and version. Then run the structural tie-outs: the balance sheet balances, the spread agrees to the source statement totals, the K-1 ties to the personal return, and the debt schedule agrees to the balance sheet. Then sample classification on the most material lines and review every add-back and consolidation entry individually, because those are judgement calls rather than extraction.
What is an acceptable error rate for automated spreading?
There is no published industry standard, and a single percentage is the wrong unit anyway — a rounding difference and a misclassified distribution are not equivalent. Set tolerance by materiality relative to the covenant thresholds and ratios the spread feeds, track observed exceptions by document type, and let that data set your sampling rate rather than a headline accuracy figure.
Which parts of a spread always need human review?
Add-backs and other discretionary adjustments, the consolidation of related entities and guarantors including intercompany eliminations, the treatment of non-recurring items, and any figure that determines whether a covenant passes or fails. These are credit judgements. Extraction can be sampled; judgement cannot.
How long should reviewing an AI-generated spread take?
Far less than spreading it manually, but not zero. On a clean single-entity file, tie-outs plus a classification sample is usually a short exercise. A multi-entity file with guarantors takes longer, because the consolidation and the add-backs need line-by-line attention regardless of how good the extraction was.
Does an AI-generated spread need a documented review?
Treat it as though it does. Model risk management practice at regulated lenders generally expects documented validation, controls, and human accountability for outputs that inform credit decisions, and a spread does. Retain what was checked, what was changed and why, and who signed off — and confirm the specific requirements with your own compliance and model risk teams.
What should be documented in the review?
The document set used with versions and dates, the tie-outs performed and their results, which lines were sampled and the basis for that sample, every override with its rationale, any add-back approved or rejected, the reviewer, and the date. The test of a good record is whether a second reviewer could follow it without re-performing the work.
Review a spread instead of rebuilding one
Send a real multi-entity file and we will show the citations behind every figure, the tie-outs run automatically, and the override trail a reviewer can sign off on.
