This Changes Everything

01

Foreword

There's a demonstration every physics student runs into at some point. Take a litre of pure water, seal it in a bottle, and cool it slowly without disturbing it. It will stay completely liquid several degrees below zero. Nothing to see, nothing to measure, nothing on the surface to say it's any different from the water in your fridge. Then tap the bottle once, and ice tears through the whole litre in under a second: what was clear liquid a moment ago is now packed solid with crystals.

That's the idea that's stayed with me since my own physics training, long before Ben existed. A system can look entirely stable and not be stable at all. The change is already happening underneath. The only reason nobody can see it is that nobody has tapped the bottle yet.

I think the world of benefits is that litre of water right now.

Agentic AI is the single most important thing that will change reward and benefits in our working lives, and it's arriving faster than almost anyone in this industry expects.

Benefit and reward is the best use case for AI anywhere in an enterprise. Not one of the best. The best. The data is dense with context, different for every single person, tightly regulated, and built from rules that interact with other rules and change by market. That is exactly the shape of problem these systems are extraordinary at handling, and exactly what no human can do reliably at scale, however good they are. AI doesn't just suit this work. It wants it.

So the big constraint this profession has always lived under is about to disappear. Everything a reward team can do today is bounded by how much one person can hold in their head and reconcile in a week. That is the only reason programmes are simpler than reward leaders want them to be, the only reason tailoring stops at age and location, the only reason another country means another headcount.

Take that limit away and the ambition changes. A programme built around each individual, in every market you operate in, adjusting the moment someone's life changes, costing less to run next year than it does this year.

This isn't a trend, and it isn't the sort of gradual improvement that lets an industry adjust its expectations a little each year.

The way this work gets done will look nothing like it did before.

And the job sitting on top of it is a more impactful one than most reward leaders have today. You stop being the person who keeps the programme working and become the person who shapes what it should be. And for the people you look after it gets simpler: the right cover, correct pay, and the thing they needed turning up at the moment they needed it, without them having to ask.

I have never been more excited about anything in my working life. I hope by the end of this you feel the same way.

— Sebastian Fallert, CEO & Co-founder at Ben

The report in short

  • Benefits data is dense, personal to every individual, and bound by rules that interact with each other. That's exactly why it's the best place in an enterprise to put agentic AI to work.
  • What decides whether it works isn't just the model. It's whether the structures underneath were built so AI can act on the connections between eligibility, pricing, enrolment and payroll.
  • Get that right and work starts arriving finished rather than waiting to be done, across all three parts of the job: running the programme, the employee experience, and the numbers you defend to finance.
  • "AI-native" has stopped meaning much. Gartner reckons only about 130 of the thousands of vendors claiming agentic AI are real, and five minutes of the right questions tells you which kind you're sitting with.
  • None of it makes the job smaller. It removes the reconciliation and leaves the judgement.

02

Everyone says AI-native.
Almost nobody is

Every benefits platform on the market describes itself as AI-native. So does nearly every piece of enterprise software being sold today, which ought to be the first clue that something has gone wrong with the term.

Ten years ago the equivalent phrase was cloud-based. Before that, digital. Each began as a specific claim, an architecture some vendors had built and others hadn't, and each ended as a requirement, a thing every product had to say about itself whether or not it was true, because the alternative was to look out of date.

AI-native is living out the second half of that life cycle unusually fast, for a simple reason. A meaningful part of what people mean by it has become cheap to build. Pointing a language model at a folder of policy documents and letting it answer questions about them is now an afternoon's work for a competent engineer with an API. Three years ago it wasn't.

There's a number attached to this. Gartner estimates ↗ that of the thousands of vendors claiming to sell agentic AI, only about 130 are real. It has a name for the rest, “agent washing”, and a prediction to go with it: over 40% of agentic AI projects will be cancelled by the end of 2027, on escalating costs, unclear business value or inadequate risk controls. Its analyst Anushree Verma puts it plainly: most of these projects are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”

None of which makes the chatbot layer worthless. It can read a policy document faster than a person can search one.

But that's the entire extent of what most platforms mean by AI-native. It's not a limit of the chatbot specifically, it's what's underneath it: even achatbot built to finish work instead ofjust answering questions couldn't, because the eligibility rules, pricing logic and payroll instructions it would need haven't been brought together somewhere it can reach and reason over.

Gartner, June 2025
130

of thousands of vendors claiming agentic AIare actually real, Gartner estimates.

Gartner has a name for the rest: agent washing. It also expects more than 40% ofagentic AI projects to be cancelled bytheend of 2027.

thanksben.com

Which is the whole story

Benefits software was built the way most enterprise software gets built: one module bolted to the next over a decade or two, in response to whichever client asked for whichever feature that year.

Policy documents live in one place. Eligibility rules live somewhere else. Pricing logic sits ina third location, payroll instructions in a fourth, and a spreadsheet nobody wants totouch holds the exceptions that keep the whole thing working. A simple model can bepointed at any one of those and asked to describe it. It can't see how they fit together, because the platform was never built to record the relationships between them in any form a machine can read.

That gap, not the model, is the point. And it's measurable. Salesforce ↗ found that only 29% of the average enterprise's 897 applications are connected, and that 84% of data and analytics leaders think their data strategies need a complete overhaul before their AIambitions can succeed. Cloudera and Harvard Business Review ↗ found only 7% of organisations say their data is completely ready for AI adoption, while 65% expect agentic AI to reshape many of their business processes within two years. The top obstacle, at 56%, was siloed data.

Most of the estate can't see itself.

84%

of data and analytics leaders say their data strategy needs acomplete overhaul before their AI ambitions can succeed

29%

of the average enterprise's 897applications areconnected

Salesforce, State of Data and Analytics, November 2025
thanksben.com

The cleanest evidence isn't a survey, though, because it isolates the variable. Chroma's Context Rot study ↗ ran 18 models on the same questions with the same answers available. Performance was significantly better when the answer arrived in a focused prompt of around 300 tokens than when the identical answer sat inside a context of113,000. Same model, same question, same answer present, different result, entirely because of how thecontext was assembled.

So how you assemble what the model can see matters more than which model you picked. That's an architecture question rather than a model question, which is the most useful fact in this guide: the gap between vendors is knowable, and you can find it inanafternoon.

And benefits starts further back than most

Which is exactly why the gain is bigger here than almost anywhere else in the business.

87% of organisations run payroll across fragmented environments with no single data foundation. Only 13% are on a unified HR and finance platform. (Paylocity ↗, 776US HR and finance leaders, fielded late January to early February 2026)

Wherever you look in the stack, only a quarter to a third of organisations have payroll connected to it across every country they operate in. Benefits sits at the bottom of that range, on 26%, against 27% for ERP, 28% for HRIS and 30% for both time and attendance and accounting or finance. (ADP ↗, 1,816 senior payroll leaders across 20countries)

87%

of organisations run payroll across fragmented environments

Still running benefits on spreadsheets.

61.7%
of UK employers manage between three and seven benefits providers atonce
25.7%
still track benefits admin on spreadsheets or paper

The connections do exist, in code, in integration mappings, in the head of whoever built them. What doesn't exist is a single account of how the programme is meant to behave that anything other than the original engineer can reason over. So an agent can be wired into a stack like that and made to act, but it acts blind. It can't check its own work against amodel of the rules, and it can't explain afterwards why it did what it did.

That's fine for answering questions. It isn't fine for touching payroll. Build the record properly, though, and the same agent stops guessing and starts running the programme.

03

What agentic AI brings tobenefits

Four things have to be true for any of what follows to work, and none of them can beadded to a platform afterwards.

One connected record of how a programme actually works

Not documents about the programme, but the thing itself: eligibility rules, pricing, provider contracts, enrolment windows, approval logic and the circumstances of the people in it, structured so software can reason over the relationships rather than just retrieve the pieces.

An agent that finishes work rather than answering questions about it

Plans, acts, checks, and hands back only the part that needs aperson. The test is whether anything changes state without someone clicking through it.

Tests that prove the agent is right, not just plausible

Run continuously against real scenarios, edge cases, life events and country-specific rules. The least glamorous of the four, and the one that turns “probably right” into right every time.

Reaching people inside the tools they already use

A team that has to log in to a portal to find out what needs its attention is still doing the noticing.

Get those right and one sentence describes the result: the agent runs the programme, the reward leader makes the calls that matter. Here's what that looks like across the three parts of the job.

Benefits management: the end of firefighting

Start here, because it's where the argument runs deepest and where the money is.

Benefits are 30.1% of total employer compensation cost ↗ in the US private sector, non-wage costs are 24.8% of total labour costs across the EU ↗, and the average US family health premium is $26,993 ↗. That is an enormous amount of money currently being reconciled in spreadsheets.

And it's going wrong at a rate the people doing it can't see. From the same survey of626UK HR and finance professionals:

Both of these cannot be true.

95.9%
are confident benefit changes are actioned accurately
46.6%
had at least one error caused by manual processes in the past year
Brown & Brown / Drewberry, 626 UK HR and finance professionals, February 2026

Alongside that, 64% of organisations report payroll leakage of at least 1% a month (Paylocity ↗), and 53% of people would consider leaving after repeated payroll mistakes, with 88% saying the way their company handles payroll reflects how much itrespects them (HiBob ↗, 2,000 US employees). Dependent eligibility audits typically remove between 3% and 10% of covered dependants ↗ against an average annual cost per dependant above $4,570, which on Mercer's worked example means an employer covering 1,000 dependants at 5% ineligibility is paying around $228,500 a year for cover nobody is entitled to.

Now picture the same programme on one connected record. A benefit rule changes inasingle country. Nobody raises a ticket, because nobody needs to notice. By the time someone on the reward team looks, the eligibility logic has been rebuilt from the new policy document, every record it touches has been reconciled, and the one case that genuinely needs judgement is sitting there, flagged, waiting.

That's the whole interaction. Not ask, receive, apply. Instead: here's what changed, here's what's already been done, here's the part that needs you.

Benefits experience: people finally know what they have

This is where the change will be most visible to employees, because it's where a decade ofspending has moved the numbers least.

Start with the pair of numbers that explains why. WTW ↗ found the share of employees who feel well informed about company benefits rose from 77% in 2024 to 84% in 2025, while satisfaction with the benefits offering fell from 66% to 61% over the same period. Being told more about your benefits turns out not to be the same as getting more from them. (Worth noting the same report has satisfaction with non-monetary benefits moving the other way, 61% to 66%, which is the clue: it's the parts people can actually use that move the number.)

And the room to move is enormous. Fidelity International ↗ finds approximately four in 10 employees report they are unaware of what benefits or programmes are available (Global Sentiment Survey 2025, 38,000 working adults across 35 markets). Aflac ↗ finds only 42% of US employees say they understand their health insurance policy very or extremely well, down from 49% a year earlier (2,000 employees). In the UK, only 36% say they fully understand their benefits ↗ (1,000 working-age professionals, 2025), against 91.7% of employers who feel the benefits they offer are understood by employees .

Employers think their benefits are understood.

91.7%
of employers believe their benefits are understood
36%
of employees say they fully understand their benefits
Drewberry employer survey 2026 and employee survey 2025. Two samples, a year apart

More communication won't close that. Employees already spend an average of119 minutes ↗ reviewing their benefits.

46% now call the volume of benefits communication overwhelming ↗, up five points inayear. And 73% of employers admit their employees repeatedly ask the same questions ↗.

A chatbot that searches a static catalogue faster doesn't help, because the gap was never about search speed. A new parent still has to know to ask about parental leave before achatbot can tell them about it.

What does work is guidance built around the individual, arriving before they ask.

74%

want recommendations tailored tothem during enrolment (ADP ↗)

74%

say tailored recommendations would increase their confidence in benefits decisions (Voya↗).

Where communication is judged engaging, 83% know where to find information about their benefits against 37% where it isn't, and 78% say their benefits meet their needs against 17% (Mercer Marsh Benefits ↗).

And it concentrates at the moment of need. Nearly one in three employers say supporting people through major life events is a top challenge ↗, 79% of employees expect to feel cared for at life's impactful moments ↗, and people with a positive experience of taking leave are twice as likely to say their employer cares about their wellbeing ↗.

So picture the version where the system already knows. An employee says she's just had ababy. It asks only what it actually needs, adds the dependant, shows the cost impact before she confirms, and reshapes what she sees everywhere else. She never learns theword eligibility.

One design principle comes out of the employee data, and it's the one that decides whether any of this gets adopted.

Prudential ↗ found 83% of employers are interested in using AI to help people understand their benefits, 58% of employees say they would use it. People want AI-assisted and human-backed. Which is a gift rather than a constraint, because it means thewinning design is the one that buys your team more time for exactly those conversations.

Benefits insights: answers, not dashboards

This is the part with the furthest still to travel across the whole industry, which also makes it the biggest prize left on the table.

Costs are rising faster than almost anything else a reward leader is accountable for. WTW↗, surveying 346 health insurers representing 82countries, puts the global medical trend rate at 10.3% for 2026, after rises of 10% in 2025 and 9.5% in 2024, and 55% of insurers expecting higher trend expect those levels to persist for more than three years.

Mercer Marsh Benefits ↗ reports trend above 10% in most regions for the sixth consecutive year. Business Group on Health ↗ calculates 2026 costs landing 62% higher than 2017 on a compounded basis. Aon ↗ finds cost management is the top priority for70% of multinationals.

That isn't the interesting part. The interesting part is that almost nobody can say what themoney achieved. The CIPD's Reward Survey ↗ asked the 864 organisations with stated objectives for their benefits package what they consider when amending it: half named the financial cost of the options, and only 15% named using internal data to check benefits usage and effectiveness.

15%

named using internal data tocheck benefits usage and effectiveness.

Most of the estate can't see itself.

55%

variety of plans across countries

47%

lack of data

46%

multiple providers

MBWL International
thanksben.com

It isn't a competence problem. It'sa plumbing one.

Business Group on Health's employers predict a median 9% trend for 2026, offset to 7.6% with plan design changes. That 1.4 points is roughly 15% of the increase, and it's the clearest published figure for what acting on your own data is worth. Nobody has properly measured the rest, which is precisely where the advantage sits.

So here's what changes, and it isn't a better dashboard. Nobody's problem is too few filters. The answer is that you stop building the analysis and start asking for it.

Three months before renewal you ask which of your health benefits are actually earning their cost. What comes back is an answer rather than a report to interpret: cost per enrolled employee by benefit and by market, against last year. Two benefits where cost per user has risen faster than enrolment. One where most of the spend sits with a small fraction of the population. A chart you didn't have to specify. And because it's aconversation, the next question is just the next question. What happens to the total ifwechange the excess on the top plan? Which markets would that hit? Show me the ones where utilisation wouldn't justify the increase.

By the time you're in the room with your broker, you've already done the analysis they were planning to bring.

The CFO version is shorter and harder. What did we spend, what did it buy, and would you do it again. Today that's a fortnight of pulling numbers out of three systems into aspreadsheet, and it ages the moment it's finished. Against one connected record it's aquestion, an answer, and a chart that goes straight onto a slide.

The version that matters most is the one you never ask for. An allowance about to lapse with the budget unspent. Dependants being added in one market at several times the usual rate. A benefit whose cost per user has doubled while enrolment stayed flat. The system already holds all of that. The only question is whether it tells you in March or you discover it at renewal.

That's the difference between reporting and insight. Reporting waits to be interrogated. Insight arrives when it's needed — whether you asked for it or not.

04

What this means for you

The ceiling on what a reward team can achieve has always been arithmetic.

APQC's benchmark ↗, across 195 companies, puts 60% of the time HR splits between administrative and strategic work on the administrative side. Move that ceiling and thebottleneck doesn't disappear, it moves, from processing to judgement. Which benefits are worth keeping. Which markets need something different. Which trade-offs the organisation has lived with for years because nobody had the hours to look properly.

Most of HR's time still goes to admin.

60%
of employees say they fully understand their benefits
APQC, Open Standards Benchmarking
thanksben.com

That's a bigger job than most reward leaders have today, and it asks for three new pieces of craft.

You decide rather than do. If work arrives already finished, someone has to decide inadvance what they want to see before it happens rather than after. Which changes need a human signature, and which should simply happen and be logged. Almost nobody has written that list down, because until now there was no reason to. It's about to become one of the most valuable documents a reward team owns, and writing it is a morning's work.

You own the logic, not the output. You already design the rules. That has always been thejob: who's eligible, what they get, what happens when somebody's circumstances change. What changes is where those rules live. Right now they're spread across your head, a broker's spreadsheet, a policy document and an integration somebody built in2019, which is why the only way to check the programme is working is to inspect what itproduced, one record at a time. Write them down in one place instead, in a form software can act on, and the rule becomes the thing you ship. Change it once and the programme changes everywhere, in every market, that afternoon. Checking moves with it: rather than sampling outputs you read the chain from rule to condition to outcome, which takes minutes and tells you something an output never can.

Judgement becomes the scarce skill. Which means protecting it. A July 2025 Gartner↗ survey found just 7% of organisations provide guidelines on how to use the time AI saves, and describes those savings as arriving in “small and fractured blocks of time”. Ten minutes shaved off forty tasks doesn't add up to a day anyone can use. Removing awhole workflow does. So when a vendor quotes you a time saving, ask which of the two they mean.

And no, this doesn't write you out of your own job

There's a question around AI that frightens people, so let's deal with it directly, because theevidence is a good deal more encouraging than the headlines.

The standing cautionary tale is Klarna. In February 2024 Klarna said ↗ its AI assistant was handling two-thirds of customer service chats, doing the work of 700 full-time agents. InMay 2025 its CEO conceded ↗ that “cost unfortunately seems to have been a too predominant evaluation factor when organising this, what you end up having is lower quality”, and began recruiting human agents again. What it didn't do is reverse course: byits Q3 2025 earnings call ↗ the assistant was doing the work of more than 853 agents, with human support repositioned as a premium tier. The lesson isn't that AI failed. It's that running a function on cost alone produced worse service, and Klarna had to buy the quality back.

Orgvue ↗ surveyed 1,163 senior leaders across eight markets: 39% had made people redundant because of AI, and 55% of those admit they got the decision wrong. And themost useful counter-example comes from a company that did exactly what the frightening version describes. IBM replaced a few hundred HR staff with AI agents.

Adoption is the only test that matters in the end, and it runs on trust. KPMG and theUniversity of Melbourne surveyed 48,000 people across 47 countries ↗: 46% are willing to trust AI systems, 57% of employees hide their use of it, and 56% have made mistakes at work because of it. Where trust and governance are absent, use goes underground and errors follow. Where they're present, this compounds.

Trust is missing, so use goes underground.

57%

hide their use of it

46%

are willing to trust AI systems

56%

have made mistakes at work because of it

KPMG and the University ofMelbourne
thanksben.com

Any account of AI in this industry that makes the reward leader a passenger in their own function has it backwards. The judgement was never the superfluous part of the job, sitting apart from the doing and available to be subtracted from it. It's the part everything else was in service of.

05

How to tell the difference between bolted-on AI and atruly agentic AIsystem

Ask this
A bolted-on system
A built-in system
Ask the employee search and a policy document the same plain question: "How do I claim on my dental cover?"
Recites the policy document back, faster than a person could, no more usefully
Answers with that person's actual coverage
Watch a demo of a policy change going live
The chatbot keeps answering from theold policy while someone updates the logic behind the scenes
The system rebuilds the logic and explains what changed and why
Ask to see the audit trail for one change made last month
A change log showing who clicked what
Every input, condition and calculation behind the outcome, without aninvestigation
Ask how they test their agents, and what happens when one is wrong
Describes the model it licenses
Describes the scenarios, thefrequency, and the escalation path when a test fails
Ask which parts of the demo are shipped and which are prototype
Treats the question as hostile
Answers specifically, unprompted, andnames something unfinished

And one question worth putting to every vendor, this one included: is your AI reasoning over one connected account of how these benefits actually work, or is it reading separate pieces of that account and hoping the gaps between them don't matter?

You'll want your answer ready before one of these lands:

  • A legacy contract comes up for renewal, or an RFP arrives and nobody's sure what toask about AI.
  • A new head of reward or CPO arrives and audits the stack.
  • A payroll or benefits error escalates to the CFO, or benefits spend comes under scrutiny.
  • The board asks what the AI strategy for HR and benefits actually is.
  • An AI governance review forces an audit of every vendor's AI.

Three things worth doing this quarter

Write the approvals list.

One page. Which benefit changes need a human signature, and which should happen and be logged. It's a morning's work and you'll be the only person inyour peer group holding one.

Put the five questions above to your current provider incumbent.

Not as a procurement exercise. Just ask, and note which answers arrive specifically and which arrive asbrochure.

Pick one number you can't currently answer

about your own programme, and try toanswer it. Cost per enrolled employee for your most expensive benefit, by market, against last year is agood one. How long it takes you, and how many systems you touch, tells you exactly where you stand.

06

This changes everything

Within a few years a reward leader will run a programme built around each individual, inevery market they operate in, adjusting the moment someone's life changes, and costing less to run than it does today. The team doing it will be no bigger. It will simply spend its week on decisions instead of reconciliation.

Think about what that actually removes. No more shadow-checking a spreadsheet thenight before a payroll run. No more discovering at renewal what you should have known in March. No more building the same report for the fourth time. No more employees finding out too late that they were covered for the thing they needed. None ofthat was ever the job. It was the tax on doing the job.

What's left is the real job. Which benefits are worth keeping. Which markets need something different. What your programme should say about the company you work for. The questions people join this profession to answer and then spend a career not having time for.

Almost nobody has started. Only 15% of reward teams check their own data to see whether the benefits they buy are being used. Read that number again, because it isn't awarning. It's a starting gun. The reward leaders who move in the coming months won't becatching up with anybody. They'll be the ones everyone else is benchmarked against in2030.

None of it depends on a better model. It depends on whether the data underneath was built so software can act on it, and that is a question you can answer about any vendor, thisone included, in a single meeting.

Water sits at zero degrees looking exactly as it did the day before, and then the whole litre turns at once. Benefits technology is at that temperature now.

Don't wait to watch it happen. Go and be the reason it does.