What changed in 2026, and why the rulebook is deliberately blank
On April 17, 2026, the Federal Reserve, the OCC and the FDIC issued revised supervisory guidance on model risk management, SR 26-2, replacing SR 11-7 for the first time in roughly fifteen years. The revision preserves the familiar principles, validation, monitoring, governance, and effective challenge, while taking a more risk-based and materiality-sensitive approach to what counts as a model at all.
The detail that matters most for this topic sits in a footnote. The guidance states that generative AI and agentic AI models are novel and rapidly evolving and are therefore not within its scope, while adding that an institution's own risk management and governance practices should determine appropriate controls for tools and systems the document does not cover. Read plainly, the primary US model risk framework does not reach the agents now being deployed, and it says so on purpose.
That is not permission. It is an assignment. Where SR 11-7 gave institutions a shared vocabulary to point at, agent governance now has to be constructed internally and defended on its own terms. The gap is widened by where agents come from: not from an innovation lab that would have triggered a review, but bundled into a core, an origination platform, a servicing system, or a productivity suite. A capability that arrives in a vendor release is easy to be running before anyone has classified it.
The international picture points the same way. The EU AI Act's risk-based obligations around logging, documentation, human oversight and robustness set the vocabulary supervisors are converging on, and the Financial Stability Board's 2026 consultation on responsible AI adoption asks directly whether existing governance practices stretch to generative and agentic systems. Nobody is waiting for fully autonomous agents before caring how they are adopted and monitored.
What actually separates an agent from a model or a chatbot
The word agent is doing heavy marketing duty, and a lot of what is sold as agentic is a chatbot with a longer prompt. The distinction is not sophistication. It is whether the software takes actions in other systems, and how much of the work it strings together before a person sees it.
| Type | What it does | What has to govern it |
|---|---|---|
| Assistant or chatbot | Answers a question or drafts text in a single turn, in one interface | Wrong answers are visible immediately, because the output is the whole product |
| Traditional model | Produces a score, a classification, or a forecast from structured inputs | Validation, monitoring, and effective challenge, the SR 26-2 territory |
| Agent | Plans a sequence, reads and writes across systems, calls tools, and completes a multi-step task | Everything above, plus the tools it can reach, the actions it can take, and the point a human intervenes |
| Multi-agent workflow | Several agents handing structured output to each other across a process | All of the above, plus the handoffs, where an early error propagates silently into later steps |
The practical consequence of that fourth row is worth dwelling on. In a single-model world, a bad output lands in front of a person who evaluates it. In an agentic workflow, a bad extraction at step one becomes an input at step two, a premise at step three, and a recommendation at step four, and the person who finally reviews it sees a confident conclusion with no visible seam. The control that matters is not whether the model is good. It is whether the chain can be walked backwards.
This is also why the honest version of an agent conversation is narrower than the marketing version. An agent that reads a document package, extracts the figures, checks them against a policy, and prepares a recommendation is doing genuinely useful work. An agent that decides is a different proposition, and in most regulated workflows it is not the one being bought.
Where agents actually earn their place
Use case lists organised by department age badly and tend to flatter the vendor. A better filter is the properties of the work itself. Tasks with all four of these characteristics are where agents have delivered so far, across banking, credit unions, wealth management, and non-bank lending alike.
High volume, low variance
The same job, many times, with recognisable structure. Document intake, statement and return processing, onboarding checks, exception queues, reconciliation. Value scales with the count of items rather than the size of any one of them, which is why small-ticket and high-frequency work pays back first.
Rules that already exist in writing
Credit policy, eligibility criteria, KYC and KYB requirements, covenant definitions, reporting obligations. If the rule is already written down somewhere authoritative, an agent can apply it consistently. If the rule lives in a senior person's judgment, encoding it is a much larger project than the software purchase suggests.
Work that should leave evidence anyway
Anything a reviewer, an auditor, or an examiner may later ask about. This sounds like a constraint and is actually a fit signal: workflows that already demand a documented trail are the ones where an agent's logging is a benefit rather than an overhead nobody wanted.
Reversible before it reaches a customer
Preparation, drafting, checking, and flagging are all recoverable if wrong. Sending money, declining an application, closing an account, and filing a report are not, or not cheaply. Autonomy should track reversibility more closely than it tracks confidence scores.
Read the four together and a pattern falls out. Agents are strongest on the preparation layer of regulated work: assembling, extracting, checking, reconciling, drafting, and monitoring. They are weakest, and riskiest, at the decision point. That boundary is not a temporary limitation of the technology. It is where accountability sits, and accountability has not moved.
Within lending specifically, the workflow-by-workflow version of this is set out in AI agents for commercial lending workflows, and the origination-side detail in automating underwriting and origination.
Five governance questions the guidance no longer answers for you
With agentic systems outside the scope of SR 26-2, these are the questions an institution has to answer in its own framework. They are also, usefully, the questions worth putting to any vendor whose product now includes an agent. What follows describes the landscape as of September 2026 and is not legal or supervisory advice.
1. What is the inventory, and does it include what arrived by default?
Most institutions can list the models they built. Fewer can list the agentic features switched on inside purchased software over the last eighteen months. An inventory that only captures deliberate projects will understate the exposure, and the first honest exercise is usually asking existing vendors what their recent releases turned on. A capability nobody classified is still a capability being relied upon.
2. What can each agent actually reach, and what can it change?
For a model the question is what it predicts. For an agent it is which systems it can read, which it can write to, which tools it can call, and what the blast radius is if it behaves unexpectedly. Scoping permissions tightly is a more effective control than most output-level testing, because it bounds the damage rather than trying to anticipate the error.
3. Can a reviewer reconstruct how a conclusion was reached?
Every figure should trace to the document and page it came from, every step of the chain should be reproducible, and the version of the model or tool that processed a given item should be retrievable months later. This is the single control that does the most work: it converts review from re-performance into verification, and it is what an examination or a customer complaint eventually rests on.
4. Where does a human have to be, and is the override real?
Define the autonomy level per workflow rather than per product, and set it against reversibility. Then test the override under realistic conditions, including when the agent is confident and wrong. An override that exists in the interface but is never exercised in practice is a control on paper. What gets recorded when a person intervenes, the prior value, the new value, the reason, the user, matters as much as the ability to intervene.
5. How much of this depends on a third party you cannot inspect?
The Bank of England and FCA's 2024 survey found that 46% of responding firms had only partial understanding of the AI technologies they use, largely because those technologies sat inside third-party products, while 84% reported having an accountable person for their AI framework. Those two figures describe the same tension: accountability has been assigned, understanding has not caught up. Vendor governance is your governance, and the questions in SOC 2 Type II for commercial lending AI cover the half a security report does not answer.
What stays human, and what good looks like in practice
The useful framing is not how much work an agent can take, but which work a person still has to own. Credit decisions, adverse action reasoning, exception approvals, customer-facing commitments, and regulatory filings stay with named people. What changes is that those people should be making decisions on a complete, checked, evidenced file rather than spending most of their week assembling one.
- Citations on every extracted value. A figure in a recommendation should link back to the source document and page. Without it, a reviewer either re-performs the work or trusts it, and neither is what was bought.
- Complete override records. Prior value, new value, reason, user, timestamp. A system that silently accepts a correction has erased the evidence that a review took place.
- Autonomy set per workflow and written down. Which steps proceed unattended, which require review, and at what confidence or materiality threshold. Confidence scoring only matters if something acts on it.
- Versions that can be named later. If someone asks in eighteen months which version processed an item and what it produced, the answer should be retrievable rather than reconstructed.
- Accuracy quoted the way it will be tested. Per document type or per task, measured without human intervention, on documents that look like yours. A single blended accuracy figure across a vendor's whole corpus tells you very little about your own edge cases.
Where Uptiq fits
Uptiq builds in this shape deliberately. Qore runs domain-trained agents for the preparation layer of financial institution workflows, document intake and classification, extraction and financial spreading, policy and eligibility checks, credit memo drafting, and post-close monitoring, across banks, credit unions, non-bank lenders, wealth management firms and equipment finance companies. Every extracted figure carries a citation back to the page it came from, adjustments are surfaced for a decision rather than applied silently, and each override is retained with its reason and user. The agents run alongside the existing core, origination and servicing systems through 100+ integrations rather than replacing them, which keeps the deployment a workflow change rather than a platform migration.
How to get started without a two-year programme
The failure mode in this category is an enterprise AI strategy that produces a roadmap and no working software. The alternative is to establish enough governance to be defensible, then prove one workflow properly and let the framework mature against something real.
Inventory what is already running
Before evaluating anything new, find out which agentic features are live inside software you already own. This is usually a short exercise with a surprising answer, and it tends to reset the conversation from whether to adopt agents to how to govern the ones already in production.
Write the autonomy policy before the pilot, not after
One page: which categories of action can be taken unattended, which require review, which require approval, and who owns the framework. It does not need to be sophisticated to be useful, and having it in place before a vendor conversation changes what you ask for.
Pick one workflow that scores on all four properties
High volume, written rules, evidence required, reversible. In most institutions that is document intake or a specific extraction and checking task. Resist the temptation to start with the workflow that is most annoying rather than the one that is most suitable.
Test on your worst inputs and your known answers
Give any vendor the documents you actually receive, including the photographed one, the one with the unusual format, and the one your team got wrong last year. Then ask the system to evidence each output. Forty minutes of that beats any demo.
Baseline before, measure the business result after
Capture elapsed time by stage, touches per item, and rework rate before anything changes. Afterwards measure cycle time, throughput per person, and exception rate. Extraction accuracy is a floor to clear, not the outcome you are buying.
For the review discipline that has to sit on top of any of this, see how to review AI-generated output, and for the vendor category itself best AI for business document analysis.
Frequently asked questions
What is an AI agent in financial services?
Software that plans and executes a multi-step task across systems rather than answering a single question. A chatbot returns text, a traditional model returns a score, and an agent reads from and writes to other systems, calls tools, and completes a sequence of work before a person sees the result. In financial institutions the common deployments are document intake and classification, extraction and spreading, policy and eligibility checking, drafting, and post-close monitoring.
Does SR 26-2 cover agentic AI?
No. The revised interagency model risk management guidance issued by the Federal Reserve, the OCC and the FDIC on April 17, 2026 replaced SR 11-7, and it explicitly states that generative AI and agentic AI models are outside its scope because they are novel and rapidly evolving. It adds that an institution's own risk management and governance practices should determine appropriate controls for systems the guidance does not cover. In practice that means the framework for agents has to be built and defended internally, and confirmed with your own compliance and model risk functions.
Are AI agents safe to use for credit decisions?
The question is better split. Using agents to prepare, extract, check, and evidence the file that a credit decision rests on is well established and improves consistency. Delegating the decision itself raises fair lending, adverse action, explainability, and accountability questions that most institutions have not answered and that no vendor can answer on their behalf. The practical boundary most lenders draw is that agents prepare and recommend, and people approve, decline, and structure.
What should we ask a vendor whose product now includes agents?
Which systems the agent can read from and write to, and what it can change. What is recorded when a person overrides it. Which model or tool version processed a given item and whether that is retrievable later. How model changes are governed, announced, and testable before they reach production. What documentation you will receive for your model risk function and during an examination. And whether your data or your customers' documents are used to train anyone's models, by default and by contract.
How is an agent different from the automation we already have?
Rules-based automation applies fixed logic to structured data that someone else has already prepared. Its failure mode is predictable and it stops when the input does not match. An agent handles unstructured input and adapts its sequence, which is what makes it useful on documents and exceptions, and also what makes it fail differently: it will often produce a plausible output rather than an error. That is why traceability matters more here than it did with deterministic automation.
Where do most agent deployments in financial services go wrong?
Scope, usually. A programme that tries to transform an entire function becomes a systems project and stalls before anything ships. The other common failure is buying on a demo of clean inputs and discovering that the real document mix, the scanned statement and the non-standard form, is where accuracy and operational reality diverge. Starting narrow, on real inputs, with the governance written down first, avoids both.
Regulatory descriptions reflect publicly available sources as of September 2026, including the revised interagency model risk management guidance issued in April 2026, the EU AI Act, and published Bank of England and FCA survey material. Guidance, timelines, and supervisory expectations change, and how any of it applies to a given institution depends on its charter, regulator, and activities. Nothing here is legal, compliance, or supervisory advice; confirm with your own legal, compliance, and model risk functions.
Start with one workflow, on your own documents
Tell us which stage is consuming your team's week and we will run it on the files you actually receive, with every figure cited back to the page it came from.
