Alphaweb Get the paper

Sector note · HR and payroll

Payroll is the one HR process an agent can actually be graded on

Attention in HR technology goes to recruitment screening and the holiday-balance chatbot. The money sits upstream of the payslip, in the reconciliation nobody wants to own — and that is the only part of HR where you can prove whether the software was right.

Alphaweb — Payroll is the one HR process an agent can actually be graded on
What this note argues5 lessons · 40 seconds
  1. Payroll, not recruitment, is where HR's AI money actually sits

    The payslip is the only artefact HR produces that moves money, carries a statutory deadline and creates a legal record someone can act on.

  2. Payroll is the rare AI use case you can actually mark

    Replay a period that closed a year ago and compare the software against what was paid and every correction raised afterwards. A marked exam paper.

  3. The case is coverage, not cleverness

    Human control is sampling, bounded by headcount and the hours before cut-off. Software can do to forty thousand records what a manager does to forty.

  4. Ownership, not accuracy, is what kills these projects

    A full-population review finds four hundred discrepancies, not twelve, and they land on people who do not report to whoever bought the software.

  5. The first return is cycle time, not headcount

    Closing a period without three people staying late, plus an evidence trail that survives an auditor. Headcount effects arrive later, as reallocation.

40%of $1bn+ organisations report scaling AI agentsMcKinsey, The state of AI in 2026: On the road to ROI
46%6 of 13 respondents keep 0-20 documented processesDeloitte, 2025 Payroll Benchmarking Survey (15 employers)
7 June 2027first pay-gap reports due, employers of 250+ workersDirective (EU) 2023/970 on pay transparency

The money is upstream of the payslip

Recruitment AI is visible and forgiving; the payslip is neither, and its inputs are owned elsewhere.

Ask an HR technology team where AI will pay for itself and most will point at recruitment: CV screening, interview scheduling, a conversational front end that tells people how much leave they have left. That is where the demonstrations are, because those processes are visible, they touch a lot of people, and nothing catastrophic happens if the answer is slightly wrong. None of that is true of the payslip. It is the only artefact the HR function produces that moves money, carries a statutory deadline, and creates a legal record that the employee, the tax authority and eventually an employment tribunal can all act on.

Everything upstream of it — the joiner record, the mid-month contract change, the timesheet, the commission plan, the pension election — is a supply chain feeding a number that must be right on a fixed date. And that supply chain is not owned by the people who are blamed when it fails. Deloitte's 2025 Payroll Benchmarking Survey — a small-sample benchmark of 15 large multinational employers, ranging from 25,000 to around 240,000 employees — asked how involved payroll teams are in validating the data arriving from upstream systems: the largest group, 6 of 13 respondents (46%), described themselves as only "moderately involved". The same survey found 6 of 12 (50%) have "No defined SLA" for resolving employee payroll enquiries at all. These are indicators of large-employer practice, not population statistics. So the team accountable for the number does not control its inputs, and does not measure its own response when the number turns out to be wrong.

This is worth holding against the wider adoption picture. McKinsey's "The state of AI in 2026: On the road to ROI" reports that 40% of respondents at organisations with revenues above $1 billion report scaling AI agents, up from 27% a year earlier, while 37% attribute at least some EBIT impact to AI — about the same share as last year. Scaling and value have come apart. The organisations closing that gap are not the ones with better models. They are the ones that picked processes where the output can be checked against something other than an opinion.

A payroll team samples forty records and calls it a control; the software has no reason to sample at all.

Why this work fits, and it is not the reason vendors give

Every payroll case has a document behind it, the answer is checkable later, and errors fail silently.

Payroll work has a shape that suits software carrying a whole process rather than answering questions about one. It is high volume, it repeats on a fixed cycle, and every case has a documentary answer sitting somewhere: an employment contract, a collective agreement, an absence record, a signed commission plan, or the statute itself. There is no case where the correct answer is a matter of taste.

More importantly, the ground truth is recoverable after the fact. You can take a payroll period that closed a year ago, replay it, and compare what the software concludes against what was actually paid and against every correction raised in the three months afterwards. Almost nothing else in enterprise AI can be graded that way. Most of it is assessed by asking users whether they found it helpful. Payroll gives you a marked exam paper, which is exactly why it is the honest place to start — and why a studio that is wrong about it will find out quickly.

The second structural fact is about coverage, not cleverness. Human payroll control is sampling. A team reviews a variance report, checks a handful of new starters, looks at the largest movements, and stops when the submission deadline arrives. Coverage is a function of headcount and the hours between cut-off and payment. Software has no reason to sample. The argument for it is not that it reasons better than an experienced payroll manager; it is that it can do to forty thousand records what the manager currently does to forty.

The third is that the failure mode is silent. A wrong net pay figure looks exactly like a right one. It surfaces months later as a tribunal claim, an underpayment of holiday pay, or a correction run that consumes a week. That is why the check has to be independent — recomputing from source documents — rather than a second pass through the same logic that produced the error.

ProcessWhy it goes firstWhat the agent does
Pre-payroll variance investigationIt is read-only, so nothing can go wrong with pay while you learn. It is also the highest-consequence control in the cycle and the one currently performed by sampling, which means the first run tells you honestly how much you were missing — and produces exactly the evidence trail an auditor asks for.Compares the current period's gross-to-net against the previous one, employee by employee, and treats every movement as something requiring an explanation. For each one it goes to the source: the contract amendment in the HRIS, the absence in the time system, the new pension election, the starter or leaver record, the change of tax code. It either attaches the document that accounts for the movement or escalates it as unexplained, with everything it gathered already assembled for the payroll officer who picks it up.
Variable pay verificationVariable pay is where the plan document and the calculation quietly drift apart, and it is rarely checked line by line because recomputing commission from source transactions is tedious. It is also now separately regulated: the EU pay transparency regime requires the gap to be reported with basic and variable components broken out, which is impossible if nobody can reconstruct how variable pay was derived.Recomputes each commission, bonus, overtime payment and shift premium from the signed plan document and the underlying transaction data — the CRM opportunity, the clocked hours, the rota — rather than re-reading the payroll engine's own output. Where the plan text is ambiguous it says so and names the clause, instead of resolving the ambiguity silently. It reports the population it could not verify as clearly as the population it could.
Joiner record assembly and cross-system reconciliationErrors originate here and are cheapest to fix here — a wrong tax code, NI category or pension enrolment date at entry propagates silently for months and is then corrected retrospectively at several times the cost. It is also the first place a write-back is defensible, because the volume is low enough for a named person to approve each record without the approval becoming a rubber stamp.Assembles each new starter's record across the HRIS, the payroll engine, the benefits and pension providers and the time system, then checks all of them against the signed contract, the right-to-work file and the starter declaration. Where they disagree it drafts the specific correction in each system and routes it to a named approver, who accepts or rejects it. The agent never applies its own correction.

The three processes that go first

Read before write: findings first, and the one process that writes has a human approving each record.

The order below is deliberate and it runs read-before-write. The first two processes produce findings and evidence; they do not touch pay. Only the third writes anything into a system of record, and it does so with a named human approving each record. Any sequence that starts by letting software post to the payroll master is a sequence that will be stopped by the first controller who reads the design, and rightly so.

The common thread is that in each case the software is not asked to decide what is fair. It is asked to find the document that explains a number, and to say plainly when no such document exists. That second output — the unexplained variance, escalated with everything already gathered — is the one that changes how a payroll team spends its week.

What actually stops these projects

Blockers: unowned findings, undocumented rules, no calendar slack, and write-back that often has no API.

The first obstacle is ownership, and it is the one that kills most attempts. An agent that reviews a full population will not find twelve discrepancies; it will find four hundred. Those four hundred become work for line managers, HR business partners and a time-and-attendance administrator, none of whom report to the person who commissioned the software. Deploying detection without first agreeing who resolves what, and by when, produces a queue that grows until someone switches the thing off. Agree the resolution path before the first run, not after.

The second is that the rules are not written down. The shift premium is how a particular payroll officer has always calculated it. The collective agreement has a side letter from 1998 that everyone honours and nobody can find. In the same 15-company Deloitte benchmark, 6 of 13 respondents (46%) maintained between zero and twenty documented payroll processes — individual desktop procedure documents — in total. You cannot automate a rule that exists only as institutional memory, so the first several weeks are archaeology: sitting with the payroll team, reconstructing the logic case by case, and writing it down. This is the real cost of the project. It is not the software, and any proposal that prices it as though it were is misleading you.

The third is the calendar. Payroll has no slack. You cannot run an experiment in the week before a pay date, and there is no safe way to introduce a new control other than parallel running — the software and the existing process both operating over the same live cycle, results compared afterwards. That doubles the work for a period or two. If the team cannot absorb that, the honest answer is to wait until it can.

The fourth is genuinely technical, and it is narrower than people expect. Reading is usually fine. Writing back is not. Payroll bureaux, outsourced providers and payroll engines designed decades ago frequently have no usable API; the integration is a fixed-width file dropped overnight, and the contract with the provider does not contemplate a third party writing into it. Establish what the write path actually is before designing anything that depends on one, because the answer is sometimes that there isn't one.

A ninety-day sequence

One entity, read-only, replay then parallel run. The artefact that matters at day ninety is the override log.

The sequence below assumes one legal entity, one payroll population and read-only access to start. It is designed so that the decision at day ninety is made on evidence generated inside the organisation rather than on a reference customer. The most important artefact it produces is not the agent; it is the override log — the record of every case where a human reviewer disagreed with the software and why.

What it is worth, and why nobody should give you a number

Nobody can quote you a percentage. Build it from correction volume, the coverage gap and pay-gap work coming.

Any vendor who quotes a universal percentage for payroll automation is quoting a number they cannot have. The value depends on things specific to you: how many legal entities and collective agreements you carry, what share of pay is variable, whether your HRIS and payroll engine share a data model or exchange files, and how much of your current control is sampling rather than coverage. A single-entity, single-agreement payroll with a clean HRIS and a stable workforce may not be worth doing yet, and that is a legitimate finding.

The calculation you can do yourself has three parts. Take your actual correction volume over the last four quarters and the fully loaded cost of resolving one, including the employee's time and the query it generates. Add the coverage gap: what proportion of records currently receive a real check, and what the population looks like that receives none. Then add the deferred work you already know is coming — for most European employers that is the pay-gap analysis by category of worker, split between basic and variable pay, which is difficult precisely because the underlying data has never been reconciled.

The honest expectation is that the first return is not headcount. It is cycle time, the ability to close a period without three people staying late, and an evidence trail that survives contact with an auditor. Headcount effects, where they arrive, arrive later and are usually reallocation rather than reduction — payroll people stop reconciling and start handling the exceptions that were previously invisible. That is a less exciting claim than the ones being made elsewhere, and it is the one worth putting in a business case.

What a supervisor will ask

Expect three lines of questioning, and they differ from the generic AI-governance script. First, pay equity data: Directive (EU) 2023/970 required member states to bring implementing law into force "by 7 June 2026", with employers of 250 or more workers reporting from 7 June 2027, and it requires the gap to be reported "by categories of workers broken down by ordinary basic wage or salary and complementary or variable components". An unexplained gap of at least 5% in any category of workers triggers a joint pay assessment with workers' representatives. A supervisor will ask whether you can produce that split and defend the job-evaluation categories behind it — which is a data-lineage question about variable pay, not a policy question. Second, human oversight. The ICO's employment guidance is explicit that a person exercising oversight must remain "engaged, critical and able to challenge the system's outputs" and must not "just routinely apply the automated recommendation to workers", and that workers must have simple means to request human intervention without being disadvantaged for asking. An auditor will ask to see cases where the reviewer disagreed with the software; if there are none, the oversight is decorative. Third, the filed return itself. In the UK the Full Payment Submission to HMRC is the return, so a correction is a correction to a filed document with its own consequences, and the new Fair Work Agency takes on enforcement of holiday pay. The Employment Rights Act 2025 also changes the calculation in production — the government's overview factsheet commits to "removing the Lower Earnings Limit and removing the waiting period" for Statutory Sick Pay, and to guaranteed-hours rights with "payments for short-notice cancellation of shifts". Any system that verifies pay must therefore be versioned against the rule set in force on the pay date, and be able to show which version ran and why.

A ninety-day sequence that survives contact

  1. Days 1-15 — One population, read-only, twelve months of historyPick a single legal entity and payroll population. Secure read-only access to the HRIS, the time and attendance system, and twelve months of closed payroll output including every correction subsequently raised. Do not scope a second country. The history is the point: it is what makes the work gradeable rather than a matter of opinion at the end.
  2. Days 16-35 — Write the rules downSit with the payroll team and reconstruct the calculation logic case by case — the shift premia, the collective agreement provisions, the local overrides, the side letters. Expect this to be slower and less tidy than anyone forecast, and expect it to surface rules that contradict each other. This document is the deliverable of the phase. It has value even if the project stops here.
  3. Days 36-55 — Replay closed periodsRun the variance agent retrospectively against periods that closed months ago, where the right answer is already known from the corrections that followed. Measure two things separately: what it found that the team did not, and what it flagged that was in fact correct. The second number decides whether anyone will trust it, and a high false-positive rate at this stage is a design problem, not a tuning problem.
  4. Days 56-75 — Parallel run against a live cycleRun the agent alongside the existing process over at least two live pay cycles. It reports; humans decide; nothing it produces reaches the payroll engine. Agree beforehand who resolves an escalation and within how long, because this is the phase where an unowned queue quietly forms and discredits the whole exercise.
  5. Days 76-90 — Decide on the override logReview every case where a reviewer overrode the software and why. If there are almost none, your human oversight is a rubber stamp and will not withstand scrutiny. If the overrides cluster around one rule, fix that rule. Then decide on a second population or a second process, with the cost of the archaeology phase now known rather than estimated.

What we read

The documents behind this note. Each entry says what it is, what it found, and why it should change what you do — then the link to the original.

Deloittesource 1 of 5

2025 Payroll Benchmarking Survey

What it is
Benchmarking survey of 15 large multinationals, 25,000 to 240,000 employees.
What it says
6 of 13 respondents call themselves only "moderately involved" in validating upstream input data; 6 of 12 have no defined SLA for payroll enquiries.
Why it matters
It is the nearest thing to evidence that payroll teams are accountable for a number whose inputs they neither control nor measure their response to.
Read the original →
European Union (EUR-Lex)source 2 of 5

Directive (EU) 2023/970 on pay transparency

What it is
The EU pay transparency directive, in its official EUR-Lex text.
What it says
Member states must transpose by 7 June 2026; employers of 250 or more report the gap by category of worker, split between basic and variable pay.
Why it matters
It makes the lineage of variable pay data a reporting duty, and a 5% unexplained gap in any category forces a joint pay assessment with workers.
Read the original →
UK Government (Department for Business and Trade)source 3 of 5

Employment Rights Act 2025: overview factsheet

What it is
The government's own overview factsheet for the Employment Rights Act 2025.
What it says
It commits to removing the Lower Earnings Limit and the waiting period for Statutory Sick Pay, and to payments for short-notice cancellation of shifts.
Why it matters
The calculation rules move while the system is in production, so anything verifying pay must be versioned against the rules in force on the pay date.
Read the original →
Information Commissioner's Office (ICO)source 4 of 5

Employment practices and data protection: monitoring workers - solely automated processes

What it is
ICO guidance on solely automated processes used to monitor workers.
What it says
Oversight must be "engaged, critical and able to challenge the system's outputs", and workers must not be disadvantaged for asking for human intervention.
Why it matters
It sets the test an auditor will apply to the override log: if no reviewer ever disagreed with the software, the oversight is decorative.
Read the original →
McKinsey & Company (QuantumBlack)source 5 of 5

The state of AI in 2026: On the road to ROI

What it is
McKinsey's annual global survey of AI adoption, run by QuantumBlack.
What it says
40% of respondents at organisations with revenues above $1 billion report scaling AI agents, up from 27%; 37% attribute at least some EBIT impact to AI.
Why it matters
Scaling has outrun measurable value, which favours processes whose output can be checked against a document rather than against an opinion.
Read the original →

The full paper

The white paper sets out the control design for a pre-payroll variance agent — the source-document hierarchy it checks against, the escalation and ownership model that stops findings becoming an unowned queue, the parallel-run method and acceptance thresholds, the override log format, and the evidence pack an auditor or works council will ask to see.