Sector note · Banking
In banking, the money is in KYC refresh, not the chatbot
Banks are buying customer-facing chat and more alert-triage capacity. Both are downstream of the customer file, and most banks cannot say when theirs was last true.
- The customer file is the control that everything else is tested against
The FCA defines transaction monitoring as scrutinising transactions so they are consistent with the firm's knowledge of the customer. Stale knowledge degrades every downstream control at once.
- Metro Bank's £16.7m monitoring fine had a customer-record root cause
Over 60 million transactions worth over £51 billion went unmonitored, largely because address, postcode and country of residence fields were blank, incorrect or non-conforming.
- More alerts have not produced more useful information
The Wolfsberg Group's own banks report substantial increases in SAR and STR volumes with no reliable evidence of a proportionate increase in highly useful information.
- Refresh is the rare AI project with a free labelled test set
Every bank holds thousands of refresh cases a human closed last year with the outcome recorded, so the agent can be graded before it is trusted.
- These programmes fail on data ownership and customer outreach, not on the model
No single system is authoritative for the customer, and a case stalled on a client who will not send a shareholder register is not a case software can close.
The money is in the files you already opened
The chatbot and the alert queue both sit downstream of a customer record nobody is measured on keeping true.
Ask a bank executive where agentic software belongs and two answers arrive. The first is the customer-facing assistant: the app, the contact centre, the deflected call. The second is alert triage in financial crime operations, where the queue is enormous and the staffing bill is visible with it. Both are defensible. Both sit downstream of the thing that is failing.
The FCA's Final Notice against Metro Bank, issued in November 2024, states the problem plainly. The bank was penalised £16,675,200 because between 6 June 2016 and 17 December 2020 its automated transaction monitoring system did not monitor "over 60 million transactions (circa 6.0% of the total transaction volume) with a value of over £51 billion (circa 7.6% of the total transaction value)". The headline reads as a monitoring failure. The Notice describes something narrower: customer records were rejected by the monitoring system principally because of problems with "the address, the postcode and the country of residence of the customer", where those fields were blank, incorrect or did not conform to specification.
That is a customer-record failure presenting as a monitoring failure, and it ran undetected from June 2016 until April 2019, when it surfaced during testing of a core banking upgrade. The Notice records that staff investigated and tried to escalate the issue in 2017 and 2018. A better scenario library would not have caught it. A better customer record would.
The FCA's own description of the purpose of transaction monitoring makes the dependency explicit: to scrutinise transactions "to ensure they are consistent with the firm's knowledge of the customer". Scenario thresholds, segmentation models, risk-rating engines and screening logic are all tests against what the bank believes about the customer. Buy more triage capacity while the customer files decay and you are paying people to review a test run against a stale reference. Unlike the alert queue, the decay is silent: nobody raises a ticket when an address goes years without being confirmed.
Transaction monitoring tests transactions against what the bank knows about the customer. When that knowledge is years out of date, the test is being run against a fiction.
Why refresh work suits software that acts
Refresh is document work, rule-bound, checkable against closed cases and low-harm when wrong, which is exactly the shape that suits software that acts.
Start with the industry's own verdict on the alternative. The Wolfsberg Group, an association of global banks writing about their own programmes rather than selling into them, published the first part of its Statement on Effective Monitoring for Suspicious Activity on 1 July 2024. It treats transaction monitoring as a subset of the wider task, and observes that the prevailing approach has produced substantial increases in SAR and STR filing volumes while, in its words, "no reliable evidence shows a proportionate increase in highly useful information". It also argues that customer behaviour and customer attributes, considered alongside transactions, give a broader view of potential suspicion than transactions alone. If the banks that wrote that are right, the marginal pound belongs in the quality of the customer record rather than the throughput of the queue it feeds.
Refresh work has four properties that suit software which reads, decides inside a boundary and shows its working. It is document work: a registry extract, a shareholder agreement, audited accounts, an identity document, a screening result, and a file written by somebody who has since left. It is rule-bound rather than creative, because the bank's own CDD standard already states what evidence each customer type and risk band requires. It is checkable after the fact, because every bank holds thousands of refresh cases a human closed last year with the outcome recorded, which is a labelled evaluation set nobody has to commission. And the failure mode is rework rather than harm: a file assembled badly is caught by the analyst who signs the risk rating.
What does not fit deserves the same precision. The risk rating decision, the decision to restrict or exit a relationship, the suspicious activity report narrative, and the judgement about how hard to press a commercial client for a document all stay with named people. The agent's remit ends where the consequence begins, and a design that blurs that line will not survive second-line review.
| Process | Why it goes first | What the agent does |
|---|---|---|
| Periodic refresh file assembly | The highest-volume repetitive work in financial crime operations, with a baseline already measured weekly inside any live remediation programme and a labelled test set in last year's closed cases. Being wrong costs a rework, not a customer, because an analyst still signs the risk rating. | Reads the existing file, applies the bank's own CDD standard for that customer type and risk band, retrieves and reconciles evidence from the core system, the document store, the onboarding tool and public registries, checks each document against its validity and expiry rules, compares screening results against the current record, and produces a refresh pack stating what is confirmed, what has changed, what is missing and what must be requested, with every assertion linked to its source document. A named analyst approves the risk rating. |
| Event-driven review triage | Trigger reviews are where supervisors find the gap between a firm's stated policy and its practice, and they are the part of ongoing monitoring that AMLA's draft guidelines under Article 26(5) AMLR define separately from the periodic cycle. They arrive unevenly, which punishes fixed staffing. | Takes the trigger, whether a change of address or jurisdiction, a change in beneficial ownership, a new product, an adverse media hit or a screening match, and works out what it changes about the file. It pulls the evidence, tests the trigger against the bank's written criteria, states whether the risk assessment is affected and which elements need re-evidencing, and drafts the case for review. It does not close a trigger, change a rating or decide that nothing has happened. |
| Beneficial ownership and control structure reconstruction | The most expensive manual work per file, most often outsourced at high unit cost, and the one where errors persist longest because nobody re-reads a structure chart drawn in 2019. It is pure document reasoning, where this software is strongest and a rules engine weakest. | Rebuilds the ownership and control chain from registry extracts, incorporation documents, shareholder agreements and annual filings across jurisdictions, calculates cumulative holdings through intermediate layers, identifies where the chain breaks or the evidence is stale, compares the result against what the bank holds, and flags every discrepancy next to its source passage. Control exercised by other means is reported as an open question, not resolved. |
The three processes that go first
Periodic refresh, trigger review triage and beneficial ownership reconstruction go first; chat and alert triage wait, for different reasons.
The sequencing principle is ordinary: start where being wrong costs the bank time and money rather than costing a customer something they cannot recover. Periodic refresh file assembly, trigger review triage and beneficial ownership reconstruction all have that property, and in each a named person still makes the decision that carries consequence.
Customer-facing chat is absent from the table, deliberately. It has the loudest demand from the front office and the worst characteristics for a first project in a regulated bank: the counterparty to a mistake is a retail customer, the conduct question arrives on day one, and there is almost no verifiable ground truth to grade against. Alert triage is absent for a different reason: it optimises the stage after the one that is failing, and its ceiling is set by the customer data it scores against.
Remediation programmes deserve a note. Most banks carrying a refresh backlog carry it as a remediation programme, with a contractor army, a weekly burn rate and a completion date that has moved at least twice. That is the best proving ground available, for three unglamorous reasons: the baseline is measured weekly, the unit cost per file is known to the penny, and the sponsor is already unhappy enough to try something.
What actually stops these projects
The four obstacles are source-of-truth ownership, outreach dependency, the unglamorous administrative stratum, and governance plus retrieval arriving together.
Very little that kills a KYC refresh agent is a modelling problem. Four obstacles account for most failures, and only one is technical.
The first is that no single system holds the customer. Identity data sits in the core platform, documents in a content store, ownership structures in the onboarding tool, risk ratings in the financial crime system, and the only surviving record of why an exception was granted in 2019 is a free-text note. Until somebody decides which system is authoritative for each field, the agent will surface every disagreement and be blamed for manufacturing work. Treat that first fortnight of contradictions as the finding you paid for.
The second is that refresh is not a purely internal process. A large share of cases stall because the customer has not sent the document, and no software resolves a corporate client who will not produce a current shareholder register. The honest scope is everything either side of the outreach: assemble what the bank holds, determine what is missing, draft the request, then process the reply the day it lands. Banks that measure the agent on cases closed rather than cases made ready will conclude, wrongly, that it does not work.
The third is the administrative refresh, and it needs saying to the sponsor first. A large proportion of periodic reviews change nothing: same customer, same documents, same risk band, and an analyst spends the better part of an hour establishing it. Automating that stratum is the fastest saving available and the least impressive to a board, because the headline result is that nothing happened, faster. Name it as the objective or the programme will be judged on how many risk ratings moved.
The fourth is governance and plumbing, which arrive together. Model risk frameworks in banks were designed for scoring models: a fixed artefact, validated once, monitored for drift. They have no shape for software that retrieves documents and drafts a file, so write the control description yourself and take it to the second line in month two rather than month nine. Meanwhile the budget goes on retrieval: scanned certificates at poor resolution, registry extracts in four languages, a document store whose type metadata was set by whoever uploaded the file. Expect extraction to cost more than the reasoning.
What the supervisor will actually ask
Supervisors test whether your review standard is specific about timing and triggers, whether you followed it, and whether you can prove both afterwards.
The question will not arrive as a question about artificial intelligence. It will arrive as the ordinary supervisory question about any control: show me the process, show me you followed it, show me the record.
The FCA's multi-firm review of customer due diligence processes and controls, published on 8 April 2026, is a useful preview, with one caveat: its sample is not retail banking. It covers asset management, crowdfunding, wholesale banking, contracts for difference and non-bank lenders, and no firm count is published. Among the poor practices it names are policies with "Not enough detail on how often periodic reviews should take place and what firms were expected to do in the case of event driven reviews"; firms that failed to follow their own policies on when to conduct periodic reviews; and firms with no version control over documentation, unable to demonstrate an audit trail. The test is not whether your standard is sophisticated. It is whether it is specific about timing and triggers, whether you did what it says, and whether you can prove both afterwards.
In the EU the timing question is being settled in law. Article 26(2) of Regulation (EU) 2024/1624 sets maximum periods for updating customer information, shorter for higher-risk customers, and Article 26(3) sets the triggers for an update. On 3 June 2026 AMLA published a consultation paper on draft guidelines on ongoing monitoring under Article 26(5), open until 3 September 2026, distinguishing periodic from event-driven review and expecting the framework's governance to be documented in internal policies, procedures and controls.
For a bank putting software into this process, the consequence is unremarkable and expensive to retrofit. For any named customer you should be able to reproduce which documents the system read, which version of your CDD standard it applied, what it concluded was missing, and who approved the outcome. Build that record as the process runs; it cannot be reconstructed afterwards.
What it is worth
Size it from your own remediation programme's measured cost per file and the share of reviews that changed nothing, not from anyone's quoted uplift.
We will not give you a percentage, and you should be sceptical of anyone who does. Refresh economics vary more between two banks in the same city than between two industries, because cost per file depends on customer mix, on how much of the estate is corporate, and on how bad the document store is.
The arithmetic is yours and takes about a week. Count the population due for periodic review in the next twelve months, split by risk band and customer type. Take the measured cost per file from the remediation programme rather than the estimate from the line; the two are never the same number. Establish what proportion of reviews completed last year produced no change to risk rating, documentation or escalation, because that stratum is the addressable part. Then subtract honestly: cases that stall on customer outreach are not yours to automate.
Set that against what the estate costs. LexisNexis Risk Solutions put the annual cost of financial crime compliance across EMEA at $85 billion in its study of 6 March 2024, based on 482 EMEA respondents, with 72% of those organisations reporting rising labour costs over the preceding twelve months and 70% rising compliance and KYC technology costs. Those are aggregates and will not size your programme. They establish the direction of the line you are trying to bend, and that its growth is in people rather than licences.
Alphaweb is a young studio. We have not run this across a tier-one estate and will not imply otherwise. What we will say is that the reading, comparing, assembling and drafting that constitute a customer file refresh is the part of a bank where software that acts has a genuine claim. The value is measurable in a way very few AI projects are, because the baseline already sits in your remediation reporting. And the programmes that fail here fail on data ownership and customer outreach rather than on the model, both visible in the first month.
What a supervisor will ask
Expect questions about process discipline and evidence rather than about the model. In the UK, the FCA's April 2026 multi-firm review of customer due diligence processes and controls sets the tone: it names as poor practice policies with "Not enough detail on how often periodic reviews should take place and what firms were expected to do in the case of event driven reviews", firms that failed to follow their own policies on when to conduct periodic reviews, and firms whose documentation had no version control and so could not demonstrate an audit trail. That review covered asset management, crowdfunding, wholesale banking, contracts for difference and non-bank lenders rather than retail banks, so read it as supervisory expectation, not a finding about your sector. Enforcement makes the same point with a price on it: in the Metro Bank Final Notice of November 2024 the FCA penalised the bank £16,675,200 after over 60 million transactions worth over £51 billion went unmonitored between 2016 and 2020, with customer records rejected by the monitoring system largely because address, postcode and country of residence fields were blank, incorrect or non-conforming. The FCA's stated purpose of transaction monitoring, to scrutinise transactions so they are consistent with the firm's knowledge of the customer, is why customer data quality is a monitoring control. In the EU, Article 26(2) of Regulation (EU) 2024/1624 sets maximum periods for updating customer information with shorter periods for higher-risk customers, and Article 26(3) sets the triggers for an update; AMLA's consultation paper of 3 June 2026 on draft guidelines under Article 26(5), open until 3 September 2026, distinguishes periodic from event-driven review and expects the framework's governance to be documented in internal policies, procedures and controls. Practically: for any named customer, be able to reproduce which documents were read, which version of the CDD standard was applied, what was concluded missing, who approved the outcome, and the override rate on that population. Settle early with model risk how software that retrieves documents and drafts a file gets validated, since that framework was written for scoring models.
A ninety-day sequence that survives contact
- Days 1-15 — Pick one population and write down the real cost per fileOne customer segment, one risk band, one named owner in financial crime operations who holds the number. Pull the population due for review in the next twelve months and the cases closed in the last twelve. Then time twenty refreshes end to end with analysts, including the wait on outreach. Measured time and remembered time are never the same, and this figure is both your baseline and the number a sceptical CFO will test.
- Days 10-30 — Build the adjudicated set from closed casesTake 200 to 400 refreshes closed in the last year, spanning administrative and productive outcomes, and have two experienced analysts independently record what each file should have concluded: what evidence was required, what was missing, whether the rating should have moved. Reconcile the disagreements and note the disagreement rate, which is the realistic ceiling on any measure of the agent's accuracy.
- Days 20-45 — Settle the source of truth, then wire retrieval read-onlyBefore any connector is built, get a written decision on which system is authoritative for each field: identity, address, ownership, risk rating, document of record. Then connect the core platform, the document store, the onboarding tool and the registries read-only. Expect OCR on scanned certificates, wrong document-type metadata and an entity model that cannot always say which legal person a file belongs to.
- Days 45-70 — Run it shadow-side and grade the declines as hard as the outputsThe agent works the live queue and writes refresh packs into a store nobody actions. Score it against the adjudicated set and the panel's own disagreement rate. Examine what it declared complete as carefully as what it flagged as missing, because an agent that quietly under-requests evidence looks efficient and is the more dangerous failure. Record the reference-data problems it surfaces; those are a deliverable.
- Days 70-90 — Switch on for one team, instrument it, write the control descriptionOne team, one population, human approval on every risk rating and every outbound request. Log each retrieval, conclusion and override with its reason, and review the overrides weekly. In the same fortnight write the control documentation for the second line: validation approach, monitoring, retention of the reasoning trace, and what happens when the agent is unavailable. At day ninety you should be able to state what it changed, what it got wrong, what it cost.
What we read
The documents behind this note. Each entry says what it is, what it found, and why it should change what you do — then the link to the original.
Final Notice 2024: Metro Bank plc
- What it is
- An FCA enforcement Final Notice against a UK bank, issued November 2024
- What it says
- Over 60 million transactions worth over £51 billion went unmonitored from June 2016 to December 2020, with customer records rejected by the monitoring system mainly over address, postcode and country of residence fields; the penalty was £16,675,200.
- Why it matters
- Read it as a customer-data case, not a monitoring case, and ask what proportion of your own records currently fail your monitoring system's ingestion checks.
Firms' customer due diligence processes and controls: our findings
- What it is
- FCA multi-firm review findings published 8 April 2026, covering asset management, crowdfunding, wholesale banking, contracts for difference and non-bank lenders
- What it says
- Poor practice included policies without enough detail on how often periodic reviews should happen or what to do on event-driven reviews, firms not following their own review policies, and documentation with no version control and so no audit trail.
- Why it matters
- The sample is not retail banking, but the test is portable: be specific about timing and triggers, follow it, and keep a reviewable record of every change.
Consultation Paper on Draft Guidelines on Ongoing Monitoring of a Business Relationship under Article 26(5) AMLR
- What it is
- A consultation paper from the EU's new AML authority, published in Frankfurt on 3 June 2026 and open until 3 September 2026
- What it says
- Article 26(2) AMLR sets maximum periods for updating customer information, tighter for higher risk; the draft guidelines separate periodic from event-driven review and expect the ongoing monitoring framework's governance to be documented.
- Why it matters
- The refresh clock is becoming a legal obligation with a defined evidence trail, so build the per-file record now rather than reconstructing it in 2027.
Statement on Effective Monitoring for Suspicious Activity, Part I: Moving Beyond Automated Transaction Monitoring
- What it is
- A position statement published on 1 July 2024 by an association of global banks, about their own financial crime programmes
- What it says
- Transaction monitoring is only a subset of monitoring for suspicious activity; rising SAR and STR volumes have not been matched by reliable evidence of more highly useful information, and customer attributes alongside transactions give broader insight.
- Why it matters
- The strongest argument against funding more alert triage comes from the banks themselves, which makes it easier to redirect the budget upstream to the customer record.
True Cost of Financial Crime Compliance Study, EMEA
- What it is
- The publisher's own release of its annual compliance-cost survey for EMEA, dated 6 March 2024, based on 482 EMEA respondents
- What it says
- Annual financial crime compliance cost across EMEA reached $85 billion, with 72% of organisations reporting rising labour costs over the previous twelve months and 70% reporting rising compliance and KYC technology costs.
- Why it matters
- The growth is in people, not licences, which is the argument for automating the administrative stratum of refresh rather than buying another platform.
- Capital marketsIn post-trade, the money is in the small share of trades that fail
- HR and payrollPayroll is the one HR process an agent can actually be graded on
- InsuranceInsurance's AI money is in subrogation, not the chatbot
- LegalLegal AI pays for the second read, not the first draft
- Logistics and portsPorts lose more money to late paperwork than to slow cranes
- Public sectorIn government, the decision is the cheapest part of the case
- Training · freeWhere agents belong, and where they do not
The full paper
The gated paper sets out the refresh agent in full: the evidence hierarchy it reads for each customer type, the mapping from a bank's own CDD standard to machine-checkable requirements, the refresh pack schema, the source-of-truth arbitration pattern, the control description that survives second-line review, and the protocol for building an adjudicated evaluation set from closed refresh cases.