Sector note · Insurance
Insurance's AI money is in subrogation, not the chatbot
Carriers are spending their AI budget on the part of the claim the customer can see. The recoverable money sits in the files they have already closed, and nobody is paid to go and find it.
- The recoverable money is in closed files, not the FNOL call
Salvage and subrogation run at 4.5% of net claims paid overall but 1.1% on commercial auto liability, and the gap is unread files.
- Nobody at a carrier is paid to make a subrogation referral
A referral worsens files closed and cycle time today, for cash that lands in another team's ledger eighteen months later.
- Start where a mistake costs money rather than a customer
A missed recovery, a slow quote and an incomplete first file are all recoverable errors, and none produces a wronged policyholder.
- Fraud detection is the worst first project, not the obvious one
A false positive lands on a real customer, the fairness question bites hardest there, and EIOPA names fraud among the scrutinised use cases.
- Retrieval is the bottleneck, not reasoning
Evidence sits in scanned faxes, wrong document metadata and free-text notes, so connector and extraction work outruns building the agent.
The value is on the other side of the ledger
The quiet side of the ledger holds more money than the FNOL chatbot, and nobody is measured on it.
Ask an insurance executive where AI belongs and the answer arrives quickly: first notice of loss. Shorten the call, shorten the cycle time, shorten the handling cost. It is a defensible answer, it is where most of the budget has gone, and it is also the part of the business with the thinnest tolerance for error and the loudest audience.
The larger and quieter number sits on the other side of the ledger. Research published in the NAIC's Journal of Insurance Regulation found that salvage and subrogation represent 4.5% of net claims paid across the full sample, but that for liability lines the ratio collapses: an average of 1.1% for commercial auto liability and 2% for personal auto liability. The same paper also carries a trade-press estimate that put missed subrogation opportunities at $15 billion annually (Harman, PropertyCasualty360, 2021, as cited in the NAIC paper). That money is not lost to catastrophe or to fraud. It is lost because nobody read the file closely enough, early enough, to notice that somebody else was liable.
The asymmetry of attention is easy to explain. A chatbot on the FNOL line is visible to the board, to the broker and to the customer. A subrogation referral that never happened is visible to no one. An adjuster is measured on files closed and on cycle time, and a recovery referral makes both of those numbers worse today in exchange for cash that lands in another team's ledger eighteen months from now. This is not a technology gap. It is an incentive gap that happens to have a technology-shaped hole in it.
A missed subrogation referral is not a loss the carrier suffered; it is a dollar the carrier already earned and simply failed to collect.
Why the work suits software that acts
Insurance work is document work, so software that reads, decides within a boundary and shows its trail fits.
Insurance is one of the few industries where the process is literally a document. A claim file is a police report, a repair estimate, a medical narrative, a set of photographs, a recorded statement transcript and a policy wording. A commercial submission is an ACORD application, five years of loss runs, a statement of values and a broker's covering email that contradicts two of them. The work is not creative. It is the disciplined application of a known rule set to evidence that is scattered across three systems and one inbox, followed by a written justification that somebody senior can check.
That shape suits software that can read, compare, decide within a boundary and show its reasoning — much better than it suits a model that emits a score. It also fits the sector's regulatory posture. EIOPA's Opinion on AI governance and risk management, published in August 2025, asks that undertakings "keep appropriate records of the training and testing data and the modelling methodologies to enable their reproducibility and traceability". For an opaque scoring model, that obligation is expensive and never quite satisfied. For a system whose work is an ordered sequence of retrievals and tool calls against named documents, the audit trail is a by-product rather than a project.
The corollary matters as much. Work that depends on a conversation the software cannot have, or on judgement the file does not contain, does not fit. Reserving philosophy, complex bodily injury valuation and coverage denial on ambiguous wording stay with people. The defensible territory is the reading, the reconciling and the drafting that currently consumes the first two hours of an adjuster's or an underwriter's day.
| Process | Why it goes first | What the agent does |
|---|---|---|
| Subrogation referral screening | Recovery is measured in cash, the baseline is already in your statutory filings, and a missed referral costs money rather than harming a policyholder. It is the one process where the regulator's fairness question barely applies, because the subject of the decision is another carrier. | Reads the whole file — FNOL narrative, police report, repair estimate, photographs, recorded statement transcript and the policy's transfer-of-rights clause — and produces a referral memo naming the liable third party, the theory of liability, the recoverable heads of damage and the limitation date, with every assertion linked to the page it came from. Files it declines get a one-line rationale so the decline itself can be audited. |
| Underwriting submission and document review | The queue is the constraint, not the decision. The failure mode is a slow quote rather than a wrong risk, because an underwriter still signs everything bound. Deloitte's 2026 outlook describes AIG launching a gen AI-powered underwriting assistant that ingests and prioritises every new excess and surplus submission, which tells you the pattern is already live at scale in the market. | Opens the broker email and its attachments, normalises ACORD applications, loss runs, statements of values and supplementals into the carrier's schema, reconciles contradictions between the submission and the loss run, applies appetite and referral rules, and hands the underwriter a file ready to price with the gaps listed and the follow-up questions to the broker already drafted. |
| Claims intake completeness | Everything downstream inherits the quality of the first file, including the subrogation signal. It is also the cheapest place to prove the retrieval layer works, because the document set is small and the ground truth is unambiguous. | Reads the notification, confirms cover was in force at the date of loss, tests the loss description against the policy wording, identifies what is missing — police report reference, third-party and third-party insurer details, proof of ownership, VAT status — and issues the requests. It does not decide coverage; it makes the file complete enough for a person to decide quickly. |
The three processes that go first
Subrogation, submission triage and intake go first because being wrong there costs money, not policyholders.
The sequencing principle is simple: start where a mistake costs money rather than a customer. Subrogation referral, submission triage and claims intake completeness all share that property. A missed recovery is a forgone dollar. A slow quote is a lost account. An incomplete first file is rework. None of them produce a wronged policyholder if the software is wrong, and in all three a named human still signs the consequential decision.
Fraud detection is the obvious omission, and it is deliberate. Fraud is the use case with the loudest vendor noise and the worst first-project characteristics: a false positive lands on a real customer, the fairness question bites hardest there, and most carriers already run a special investigations unit with a scoring model and years of tuning behind it. Adding an agent to that estate is a year-two project undertaken with the model risk function in the room, not a proving ground. EIOPA names fraud detection explicitly among the use cases it expects supervisors to scrutinise, and it is easier to have that conversation once you have already shown a supervisor a clean audit trail from a lower-stakes process.
What actually stops these projects
These projects fail on ownership, adjuster incentives, misfitted model governance and dirty retrieval.
The first obstacle is ownership. In most carriers, recovery sits between claims operations, the subrogation vendor and finance, and the number appears in nobody's objectives. Until one named executive owns the recovery ratio and can be shown the before-and-after, an agent that produces excellent referral memos will produce them into a queue nobody works. Buy the sponsor before you build the software.
The second is the adjuster's incentive, which is worth stating plainly rather than designing around. If your desk staff are measured on files closed per week, every referral the agent raises is a tax on the person who has to action it. Carriers that get this right change the measure before they deploy, and they usually find that the change alone moves the number a little, which is awkward but honest.
The third is governance built for the wrong object. Model risk frameworks in insurance were designed for pricing and reserving models: a fixed artefact, validated once, monitored for drift. They do not know what to do with software that takes actions against live systems. The Bank of England and FCA survey of AI in UK financial services found the insurance sector reporting the highest adoption of any sector at 95%, while 46% of firms across the survey reported only "partial understanding" of the AI technologies they use and a third of use cases were third-party implementations. That combination — high adoption, partial understanding, external dependence — is exactly what a supervisor reads as a warning.
The fourth is technical, and it is not the one people expect. The bottleneck is retrieval, not reasoning. The evidence lives in a document management system where the metadata is wrong, half the medical records are scanned faxes, the policy wording is a PDF of a PDF and the claims system holds a free-text note field that contains the only record of the third party's insurer. Expect the connector and extraction work to consume more of the budget than the agent itself, and expect the first sixty days of output to be worse than your weakest adjuster. Deloitte's 2026 global insurance outlook also reports that only 25% of respondents have taken tangible action to elevate human skills, which is the quiet reason many pilots stall at the handover rather than the build.
A ninety-day sequence
Ninety days settles the question only if the adjudicated evaluation set is built before the agent.
Ninety days is enough to know whether this works in your estate, and not enough to deploy it across a book. The sequence below assumes one line of business, one queue and one named owner. It puts the evaluation set before the build, which feels backwards and is the single decision that most reliably separates a pilot that ends in a decision from one that ends in a demonstration.
The uncomfortable part is step two. Assembling a few hundred closed files with human adjudication costs real time from senior people who have other work, and there is no way to buy it in. Carriers that skip it end up arguing about whether the output is good, with no instrument to settle the argument, which is how pilots die politely.
What it is worth
Do the arithmetic on your own recovery ratio instead of trusting anybody's quoted uplift percentage.
We will not give you a percentage. Any figure quoted as a universal uplift is a figure someone made up, because the answer depends almost entirely on where your baseline sits, and baselines in this sector vary by an order of magnitude between carriers writing the same risks.
The arithmetic you can do yourself is more useful than any benchmark. Take your net claims paid for one liability line. Apply your current recovery ratio — it is in your own statutory filings, and if the honest number is closer to 1% than to 4%, that is the finding. Then ask what a two-point improvement in referral rate would be worth, and what proportion of referrals your recovery function actually converts, because an agent that doubles referrals into a team with no capacity to pursue them is a cost, not a saving. The same discipline applies to submission triage: the value is quote turnaround and hit rate on the accounts you want, not underwriter headcount, and a carrier that treats it as a headcount play will get a worse book.
Alphaweb is a young studio. We have not run this at your scale and we will not pretend otherwise. What we can say is that this class of work — reading the file, applying the rule, writing the justification, leaving the trail — is the part of insurance where software that acts has a real claim, and that the projects which fail here fail for organisational reasons that are visible in the first fortnight if you look for them.
What a supervisor will ask
Expect the questions to be about governance and evidence rather than about the model. In the United States, the NAIC's model bulletin on the use of artificial intelligence systems by insurers — adopted by roughly half the states — sets the frame: insurers are "expected to develop, implement, and maintain a written program (an 'AIS Program') for the responsible use of AI Systems", and each programme must address the insurer's process for acquiring, using or relying on third-party data and on AI systems developed by a third party. In a market conduct examination or investigation, an insurer "can expect to be asked about its development, deployment, and use of AI Systems", so the practical test is whether you can reproduce, for a named claim or a named submission, which documents the software read, which rule it applied, which human approved the outcome and what the override rate on that queue has been. In the EU, EIOPA's August 2025 Opinion on AI governance and risk management applies a risk-based and proportionate approach and names claims handling, pricing and underwriting, and fraud detection among the use cases in scope; it expects records of training and testing data and modelling methodologies sufficient for "reproducibility and traceability", regular monitoring and where appropriate auditing of outcomes "including with the use of fairness and non-discrimination metrics", and reasonable efforts to remove biases including "potential unlawful proxy discriminatory variables". It also expects that on a customer's request, the influence of the AI system on a decision with material impact be explained "using simple, clear and non-technical language" — which in practice means your claims correspondence templates, not just your model documentation. Assume your supervisor will also press on third-party dependence: the Bank of England and FCA survey found a third of AI use cases were third-party implementations and only 34% of firms claimed complete understanding of the AI they use.
A ninety-day sequence that survives contact
- Days 1-15 — Pick one queue and write down the baselineOne line of business, one claims or underwriting queue, one named executive who owns the outcome number. Extract the current referral rate, conversion rate and cycle time from your own systems and put them in a document everyone signs. Without a written baseline there is nothing to argue against in ninety days, and the pilot will be judged on anecdote.
- Days 10-30 — Build the evaluation set before the agentTake 200 to 400 closed files, spanning good and bad outcomes, and have two experienced adjusters or underwriters independently adjudicate each one: referable or not, and why. Reconcile the disagreements. This set is the instrument for every decision that follows, and it is the step that carriers skip and then regret.
- Days 25-50 — Wire retrieval, read-onlyConnect the document management system, the claims or policy administration system and the shared mailbox in read-only mode. Solve the unglamorous problems first: OCR on scanned correspondence, wrong document-type metadata, free-text note fields. Expect this to take longer than building the agent and budget accordingly.
- Days 45-70 — Run it shadow-side against the panelThe agent works the live queue and writes its memos into a holding store that nobody actions. Score its output against the adjudicated set and against the panel's own disagreement rate, which is the real ceiling. Examine the declines as carefully as the referrals, because a quiet under-referring agent looks excellent on precision and is worthless.
- Days 70-90 — Switch on for one team, and instrument itOne team, one queue, human sign-off on every referral and every submission decision. Log each action, each retrieval and each override with its reason, and review the overrides weekly. At day ninety you should be able to answer three questions in front of your model risk function: what it changed, what it got wrong, and what it cost.
What we read
The documents behind this note. Each entry says what it is, what it found, and why it should change what you do — then the link to the original.
How's the Recovery? Salvage and Subrogation in the Property Liability Insurance Industry
- What it is
- Research paper in the NAIC's Journal of Insurance Regulation, by Bisco and Fier
- What it says
- Salvage and subrogation come to 4.5% of net claims paid across the sample, but only 1.1% on commercial auto liability and 2% on personal auto liability.
- Why it matters
- The same ratio sits in your own statutory filings; if it is nearer 1% than 4%, the recovery gap is measurable before any pilot begins.
Model Bulletin: Use of Artificial Intelligence Systems by Insurers
- What it is
- NAIC model bulletin for insurers, adopted by roughly half the US states
- What it says
- Insurers are expected to maintain a written AIS Program covering third-party data and third-party-built AI, and to answer for it at examination.
- Why it matters
- Build the per-claim record now: which documents the software read, which rule it applied, who approved the outcome and the queue's override rate.
Opinion on Artificial Intelligence governance and risk management (EIOPA-BoS-25-360, 6 August 2025)
- What it is
- Supervisory opinion issued by the EU insurance regulator on 6 August 2025
- What it says
- Expects traceable records of data and modelling methods, monitoring of fairness, removal of proxy discriminatory variables and plain-language explanation.
- Why it matters
- Plain language lands in your claims correspondence templates, not model documentation, and claims, pricing and fraud are named as in scope.
Artificial intelligence in UK financial services - 2024
- What it is
- Joint regulator survey of AI use across UK financial services firms, 2024
- What it says
- Insurance reported the highest adoption of any sector at 95%, while 46% of firms reported only partial understanding of the AI they use.
- Why it matters
- High adoption with partial understanding and a third of use cases bought in is the combination a supervisor reads as a warning.
2026 global insurance outlook
- What it is
- Deloitte's annual industry outlook and practitioner survey for 2026
- What it says
- AIG's generative AI underwriting assistant ingests and prioritises every new excess and surplus submission; only 25% of respondents have acted on skills.
- Why it matters
- Submission triage is already live at scale, so budget for the handover to underwriters, which is where pilots stall more often than at the build.
- BankingIn banking, the money is in KYC refresh, not the chatbot
- Capital marketsIn post-trade, the money is in the small share of trades that fail
- HR and payrollPayroll is the one HR process an agent can actually be graded on
- LegalLegal AI pays for the second read, not the first draft
- Logistics and portsPorts lose more money to late paperwork than to slow cranes
- Public sectorIn government, the decision is the cheapest part of the case
- Training · freeWhere agents belong, and where they do not
The full paper
The gated paper sets out the subrogation referral agent in full: the evidence hierarchy it reads, the liability tests it applies by jurisdiction, the memo schema, the decline-logging pattern that satisfies a market conduct examiner, and the evaluation protocol for building an adjudicated file set from your own closed claims.