Agentic/HR
Question briefs · Reviewed monthly
Verdict Where value leaks Four C-suite questions Who owns delivery How the question changed Redesigning the work Where to focus It is not the tools The evidence What could change The sources Method
Brief № 001 Performance Issued Aug 05 2026 Next review Sep 05 2026 100 sources 14 min read

Does AI actually deliver ROI yet? Not yet, at scale; redesign first.

Across three years and one hundred sources, the evidence converges: AI creates real value at the task level, loses nearly four of every ten saved hours to rework77, and rarely reaches the P&L: except where the work itself was redesigned first.

The case against this verdict

The loudest "AI pays" numbers come from companies that sell AI. But the most-quoted number on the other side is just as weak: MIT's "95% of pilots fail" rests on a few hundred interviews and survey responses, a sample its own coverage describes two different ways, and its authors call the figure directional rather than audited52.

The case for “yes” is real too. PwC found productivity growing almost four times faster in the industries most exposed to AI.46 The OECD calls 2025 an early signal.92 And the best number cuts both ways: when NBER found 80% of organizations seeing no impact82, that came from what roughly 6,000 executives reported about their own firms. It is the same class of evidence we discount when it points the other way.

The sentence that can be defended

"ROI follows work redesign, not tool deployment, and the investment plan should reflect that."

Value leaks four times

AI investment does not translate into organizational value in one move. It has to survive four stages to reach the P&L, and most of it doesn't. (For now.)

Turning an AI investment into value is a sequence, not a transfer. Four stages, and value is lost at three of them.

Exhibit 1
Four stages, and two different things trying to get through them

Two things are trying to reach your P&L: the time AI saves and the work people can suddenly do that they could not do before. They get stuck in different places. (Select a stage to open it.)

Of 100 units of value, how much reaches each stage

1006324n/a

The task

Hours saved and quality gained on a discrete piece of work
Efficiency: the hour

This part works. Support agents resolve 15% more issues an hour.38 Writing takes 40% less time.4 Legal tasks improved 50 to 130% in a trial with law students.32

Capability: the scope

People now do work outside their own job. Almost 17% of what they ask AI for belongs to somebody else's role.99

Where it goes

Nothing is lost yet.

This stage is settled, which is why every vendor demo lives here. The time really is saved. Support agents resolve 15% more issues an hour. The least experienced improve in both speed and quality; the most experienced gain a little speed and lose a little quality. Writing takes 40% less time. Legal tasks improved 50 to 130% in a trial with law students, and quality went up, not down. Doctors write up notes faster and burn out less.48

One study cuts against it. Experienced developers were 19% slower with AI, and thought they were 20% faster.49 The time saved is real. Asking people how much time they saved is not.

Ask: did we measure this, or is this based on perception?

After fixing

What survives rework, verification and redirection
Efficiency: the hour

About 37% of the time saved goes straight back into fixing what AI wrote. Only 14% of people finish ahead.77

Capability: the scope

Somebody still has to check work done outside their field, usually the specialist you were trying to skip. Nobody has given them that job.

Where it goes

To a colleague.

The person using AI books the saving. The person who receives their work pays for it. 41% of workers got AI-written work they had to untangle, about two hours each time.63 The saving is easy to see and the cost is spread out, so your reports show a gain your P&L never gets.

There is a second reason. Some of that untangling is work done by somebody outside their own field, using a tool that removed the friction that used to stop them. It arrives looking finished and sounding certain. That is harder to fix than a question would have been.

Ask: when AI gets it wrong, whose day pays for it?

The books

Booked financial impact an executive can point to
Efficiency: the hour

Twenty people save fifteen minutes a day. That is a full-time job. It arrives as twenty slightly easier days, and nobody can put that in a report.

Capability: the scope

Trying a task is not the same as finishing it well. Nobody has said who signs it off.

Where it goes

Into slack.

This is arithmetic before it is culture. To use scattered spare minutes you have to gather them, and that means changing what jobs cover and how work moves between people. Two-thirds of employees are given no guidance at all on what to do with the time AI frees up, and more than half do not redirect it to anything strategic (BCG 2026, n=11,749 across 14 markets).96 Where reinvestment was measured directly, 72% of sales teams put almost none of their 4.8 saved hours a week back to work.66

So: 12% of CEOs see both lower costs and higher revenue.78 Over 80% of organizations see nothing at all.82 This is the biggest loss of the four, and none of it is a technology problem.

One more loss worth naming. In the call-centre study, the weakest agents improved most and nearly caught the best. Some operators concluded they no longer needed to hire the best. But the AI was passing on what the best agents knew. Cut them and you have banked a saving by destroying the thing that produced it.

Ask: what did the freed-up hours actually produce, and who decided that?

The market

Excess return to shareholders of the organizations that adopted
Both channels

Whatever is left has to beat what you paid for the computing, in a market where your competitors bought the same tool.

Where it goes

To your customers, through price. And to the company that sold you the tool.

So far the money has gone to the companies selling AI, not the ones using it.80 Shares in AI buyers have simply tracked the market. Morgan Stanley disagrees and reports better margins for adopters, but that is a broker forecast, not a result.84

Anything your competitors can also buy is not an advantage. It is the price of staying in the game. What you own is your ability to turn what AI produces into money, which is stages two and three. Nobody sells that, and nobody can copy it from your press release.

There is a newer problem underneath this. Organizations are spending tens of millions on AI usage, and nobody can predict what a given job will cost, including the AI itself. Projects stall or blow the budget halfway through. You cannot calculate a return when you do not know the bill.

Ask: if every competitor deploys the same tool, what remains ours?

Source: Agentic HR evidence ledger · stage and channel coding is editorial · bar widths are schematic
Details and footnotes 3

Do not multiply these numbers together. The figures on the bars come from different studies with different samples. The 63 beside after fixing is one study’s rework share, not 63% of the gains on the stage above. The bars show roughly which stages lose the most, and nothing more precise than that.

This is why the research looks like it disagrees with itself. The legal study measured stage one.32 The survey of 6,000 executives measured stage three.82 They are counting different things.

The argument against all this: every handoff is also a wait, and waiting usually takes far longer than the work itself. A marketer who no longer waits two days for a developer has compressed cycle time even though no hour appeared on any timesheet, and that is a real P&L item invisible to hour-counting. Crossover also runs higher in small workspaces, 18.9% at 2–5 seats against 16.3% at 101 or more99, which suggests boundaries get crossed when no specialist is available. Large organizations keep their specialists, so the coordination gain is structurally hardest to capture in exactly the organizations reading this.

What this brief measures, and what it does not

When people say AI creates value they mean at least six different things. The column below rates how good the evidence is on each, not how well AI performs. Strong evidence can still show a disappointing result, and on several of these it does. The first evidence sweep behind this brief hunted hours and money, and on that basis four of these channels looked barely researched. They are not. A second sweep, through the research on how organizations actually behave, found what the first one was not built to see.

ChannelEvidence baseWhy
HoursStrongControlled trials, payroll records, representative firm panels. The best-evidenced channel in the literature.
ScopeEmergingOne large-scale usage study, vendor-authored, counting messages rather than outcomes.
CoordinationEmergingA randomized field experiment (INSEAD, Jan 2026, 316 employees across 42 teams)76 found grounded AI raised collaboration and knowledge-network density, with specialists gaining centrality and generalists gaining throughput.One site, one working paper.
QualityEmerging at task level, thin at P&LAbundant at task level and now with two system-level datasets, both negative: delivery stability down 7.2% (DORA) and duplicated code up eightfold (GitClear)2430. Colonoscopy detection fell 6.0 points after routine AI exposure.51 Quality net of AI-introduced error remains unmeasured.
InnovationStrongPeer-reviewed work shows AI raising individual creativity while compressing collective diversity (Doshi & Hauser, PNAS replication)1454, and LLM research ideas rated more novel at proposal, then falling behind human ideas once executed.
MoraleStrong, and split by methodThe split is measurement against perception, not pole against pole. Population-panel evidence finds no sizeable wellbeing harm36; the review literature on algorithmic management finds real autonomy and surveillance costs41; and the one measured interpersonal cost is clear, with 42% viewing colleagues who send AI slop as less trustworthy63.
Source: Agentic HR evidence ledger · coverage rating is editorial
Details and footnotes 2

A brief claiming “AI does not deliver value” while measuring only hours would be overclaiming. The defensible statement is narrower: on the channel with the best evidence, value is real at stage one and rarely survives to stage three.

Why the verdict speaks only about money. The question this brief asks is whether AI investment reaches the P&L, so the verdict is scoped to that. It is not a claim that nothing happens elsewhere. Innovation and morale both have strong evidence bases, and what they show is the same shape as everything else here.

Individual gains, collective costs. AI raises one worker’s satisfaction and lowers the trust of the colleague receiving their output. It lifts the weakest performers toward the strongest, then removes the reason to keep employing the strongest, who were the source of what the model learned. Coordination is the one channel where the signs currently agree, with individual centrality gains coinciding with denser networks overall. Everywhere else, the stage-one case and the stage-three case are measuring different levels of the same organization, and a business case built on the first without accounting for the second is not a business case.

Each chair asks it differently

Four chairs in the C-suite, four versions of the question

The CEO

"I told the board AI would transform our cost base. Where is it?"

56% of 4,454 CEOs report no significant financial benefit to date.78 The pressure is now investor-facing.

The CFO

"Can I attribute any P&L movement to our AI spend, and book the savings?"

Attribution is the core problem: individually saved time leaks out of the organization before it can be booked.

The CHRO

"Is this a tooling problem or a work-design problem, and am I about to be handed a RIF target justified by unproven gains?"

Fewer than 1% of last year's layoffs traced to actual AI productivity gains79 (Gartner).

The board

"Are we cutting ahead of proven returns, and what's our rehiring exposure?"

Gartner warns organizations that cut on AI's promise are already rehiring for roles they eliminated.

Nobody owns delivery

The CEO promised the board a return. IT delivered a tool. Nobody owns what happens in between.

In between sit two questions: who checks the machine’s work, and what people do with the hours it frees. Both are questions about how jobs are designed, and both land on HR the moment somebody asks where the savings went. That is not extra work. It is the only part of this that anyone in the building can actually control.

Redesign is slower, and it is where most of the return actually comes from. It is also the only part a competitor cannot buy: they can match your tools with a subscription, but not the way your work is organized. Get there first and the lead holds. Which is why the answer to the board is not “not yet,” but “not yet, and here is what has to be true first.” That names a sequence, and a sequence is something you can be held to.

100 sources across academic, government-statistical, central-bank, analyst, consultancy, financial, and vendor families, both poles hunted.

The question changed four times

The answer held for three years. The question was replaced four times

Each era closed one conversion and exposed the next. Select a year to read what was asked, what the evidence said, and why it stopped being the live question.

2023 · The task-level proof

Does it actually work?

It took about a year to answer, and the answer was yes, at the task layer, convincingly. Brynjolfsson's support agents resolved 14% more issues an hour, a figure the peer-reviewed publication later put at 15%238. Noy and Zhang cut writing time 40% in a randomized trial.4 Dell'Acqua's 758 consultants gained 40% in quality inside the model's frontier5 and lost performance outside it. Peng's developers finished 56% faster.1

Why it ended: nobody was asking about the P&L yet. The invoices hadn't arrived.

2024 · The $600-billion doubt

Is the spend justified?

The question jumped from capability to capital. Sequoia asked where the $600 billion of matching revenue was.11 Goldman published "Too Much Spend, Too Little Benefit?"12 Acemoglu put a ceiling on the whole thing: 0.66% of total factor productivity over a decade.10 In the same months, Microsoft and IDC were claiming $3.70 returned per dollar spent6. The two poles opened here, and they have not closed since.

Why it ended: the pilots started failing in public.

2025 · Pilots stall

Why isn't it landing?

The question stopped being about the technology. S&P found 42% of enterprises had abandoned most of their AI initiatives, up from 17% a year earlier.35 BCG put 60% at no material value.58 METR ran the cleanest experiment of the year and found experienced developers 19% slower with AI while believing they were 20% faster.49 Humlum's Danish payroll records returned a precise zero on earnings and hours.39

Why it ended: the failure got located. Not in the models, in the organizations.

2026 · The contested turn

Why does it pay for some and not for others?

The question became distributional, and that is where it still sits. BCG found regular users saving roughly eight hours a week96 that their organizations had not converted into anything. Workday measured the leak precisely: four of every ten saved hours going straight back into fixing the output.77 The minority who convert have one thing in common, and it isn't their tooling.

Where it leaves us: Google Cloud reports 74% seeing ROI within a year17; the first representative panel of roughly 6,000 executives finds over 80% reporting no productivity impact at all; and JPMorgan finds only the companies selling AI have earned excess returns. The 2026 dispute is not academics against vendors: it is which conversion you measure.

Redesign or nothing

AI pays where the work was redesigned. Almost nowhere else. And most organizations have not done the redesign.

Three years ago the question was whether AI worked. It does, and nothing since has seriously challenged that. The question now is whether your organization can turn that into money. That is a management problem, not a technology one.

It also explains why the research looks split. The people selling the tools measure the task. The economists measure the payroll, and cannot find the rest of it.

Six places to start

Where the C-suite should focus to accelerate value

Six recommendations, each one drawn from a source below rather than from opinion. Four of them sit in the same place: what happens to the work after it is done, and what happens to the time it saves.

Recommendation 1

Challenge the time savings your teams report

People believe AI makes them faster, and often it does. But in the cleanest trial anyone has run, experienced developers were 19% slower with AI and finished convinced they had been 20% faster.49 Confidence and time are not the same measurement.

In practice: if self-reported time is the only number you have, keep using it, but treat it as a signal rather than a result. Pick one process, time it end to end before and after, and hold the business case to that number instead.

Recommendation 2

Somebody still has to check the work

AI lets people do work that used to belong to someone else. Almost 17% of what employees ask it for sits outside their own role.99 That work still gets checked, usually by the expert they were trying not to bother, and 41% of people say they have received AI-written work they had to untangle.63 The saving shows up on one person’s week and the cost on another’s.

In practice: take the two or three places this happens most, decide who reviews the output, and count their time as part of what the tool costs.

Recommendation 3

Build trust as part of the rollout, not an afterthought

42% of people think less of a colleague who sends them work AI obviously wrote.63 That is a real cost to how people work together, and it builds quietly while the rollout itself looks like a success.

In practice: run this as change management rather than policy. Agree what gets flagged as AI-drafted, what has to be read properly before it is sent, and make it ordinary to say which is which. People trust output when they understand where it came from.

Recommendation 4

Tell people how to use the time they get back

Two-thirds of employees are given no guidance at all on what to do with the time AI saves them.96 Where anyone measured it, 72% of sales teams put almost none of their 4.8 saved hours a week back to work.66 Time nobody has claimed quietly goes back into the day.

In practice: name the specific thing that time should produce, role by role, and say it out loud. “More capacity” is not something anyone can deliver.

Recommendation 5

Make every small saving count

Twenty people saving fifteen minutes a day is a full-time role. It arrives as twenty slightly easier days, and no finance team can put that in a report. Over 80% of organizations say they have seen no productivity impact at all.82 The time was real. It just never landed anywhere.

In practice: change what a job actually covers so the saved time collects in one place, rather than thinning every job by a little and hoping it adds up.

Recommendation 6

Know what the bill will be before you commit

Organizations are spending tens of millions on AI usage, and neither the people running the projects nor the tools themselves can predict what a given job will cost.100 So far the returns have gone to the companies selling AI rather than the ones using it.80

In practice: ask for a cost per task and a ceiling before a project starts, the way you would for anything else with a variable bill attached.

It was never the tools

The technology is not what is holding this back.

That is the most consistent finding across a hundred sources. Where AI has not paid, the explanation in the evidence is almost never the model. It is that nobody checked the output, nobody said what the freed-up time was for, and nobody owned the distance between the spend and the result.

Those are management decisions, not procurement decisions. They are also the only part of this a competitor cannot buy, which is why the organizations pulling ahead are not the ones with better tools.

Four of the six recommendations above are about how work is organized. That is not a preference. It is where the evidence points.

A hundred sources2023 → mid-2026

One hundred sources, coded and dated

Everything above rests on what follows. One hundred sources, each coded by evidence quality and dated, with the disconfirming cases hunted as hard as the supporting ones. Filter it yourself: the era medians recompute on whatever subset you choose.

Exhibit 2
The answer fell through 2025, then split along source type

Since 2025, every confident yes has come from a company selling AI or from people describing their own results. Every peer-reviewed study and every government statistics office lands near zero: who paid for the study predicts what it found. Each mark below is one study, placed by date and by how strongly it answered. Shape tells you what kind of evidence it is.

2023THE TASK-LEVEL PROOF7 sources
median +0.68
NoMixedYes
2024THE $600-BILLION DOUBT19 sources
median +0.00
NoMixedYes
2025PILOTS STALL48 sources
median −0.10
NoMixedYes
2026THE CONTESTED TURN26 sources
median +0.00
NoMixedYes

Each bar is the median answer of every source published in that era, on the same scale the full chart uses. The answer fell through 2025, then stopped falling without turning positive.

Open the full 100-source chart
THE TASK-LEVEL PROOFTHE $600-BILLION DOUBTPILOTS STALLTHE CONTESTED TURNYESNO2023202420252026THE VENDOR POLEEvery strong yes is self-reportedLOUDEST. AND FLAGGEDA directional claim, not an audit
Source: Agentic HR evidence ledger · n=100 · both poles hunted · retracted sources plotted, excluded from medians · as of Aug 05 2026
Show tiers
From to Feb 2023 to Aug 2026

Details and footnotes 2

Vertical placement is Agentic HR's coding of each source's headline finding on this question, not the source's overall quality.

And notice what is missing: nobody in the last year has measured actual booked profit at a typical organization. That gap is why the verdict has not moved.

What could change thisMonthly cycle
Next: Sep 05 2026

What would change this verdict

Attribution

A peer-reviewed or government source showing booked EBIT attribution above 5% at the median adopting organization: not the “high performer” subset every consultancy reports. NBER’s firm panel is the instrument to watch; today it shows the opposite.

Firm turn

A successor to the NBER firm panel (n≈6,000) showing more than 30% of organizations reporting measured productivity or margin gains, measured, not perceived.

Capital

A basket of AI-using enterprises, not AI sellers, showing sustained excess returns against the equal-weighted index. JPMorgan’s own test, and the only trigger here a reader can check daily.

Rework tax

The rework tax, currently ~37% of saved time, falling materially in follow-up measurement.

The receipts100 sources

Every source, era by era

Every source coded on the Agentic HR evidence scale and checked for independence. Newest first, running back to 2023.

2026THE CONTESTED TURN26 sources
JUL 2026
100 McKinsey · Brynjolfsson interview, 'Is the AI productivity story at a turning point?'. Brynjolfsson reports 'early signs of productivity returns in national productivity statistics' and revises the Canaries displacement figure from 13% to 16–17%. Both claims are made in interview, not in a published paper; the revision is unverifiable against a primary document. · Edited interview transcript
T4 · WATCH · FLAG
JUL 2026
99 OpenAI Economic Research · 'Work at the Frontier'. 16.8% of work messages, and 43.5% of occupation-specific ones, concern tasks historically belonging to another occupation. Measures scope, not hours: the authors state they cannot observe whether output was used, was any good, or was reviewed. · >800,000 work messages, US ChatGPT Business roles, O*NET-mapped
T2 · SIGNAL · FLAG
JUL 2026
98 MIT FutureTech × CMU · AI adoption in S&P 500 filings. Deep AI integration reached 11% of the S&P 500 (from 5% in 2022), a profitability J-curve, and 'no differences in capex or productivity.' Preprint; machine-coded filings. · 10-K filings 2016–25, LLM-classified
T1 · SIGNAL · FLAG
JUN 2026
97 PwC · AI Jobs Barometer. Headcount at the most AI-exposed firms growing faster than the least (52% vs 36% since 2018), macro value signal, different layer. · Labor-market data
T2 · SIGNAL
JUN 2026
96 BCG · AI at Work, 4th edition. 42% of frontline regular users save ~8 hrs/week, but firms haven't redesigned work to convert it. Strategic clarity separates the value creators. · 11,749 across 14 markets
T2 · SIGNAL
JUN 2026
95 Anthropic Economic Index · cadences. Users report large productivity gains alongside displacement worry. Self-reported, vendor. · 81k user interviews + survey
T2 · ECHO
MAY 2026
94 Goldman Sachs · Covello follow-up. Still skeptical two years on; data-centre debt doubled to ~$182B, 'unsustainable.' Reached us through secondary coverage, not the primary note. · Research update
T2 · SIGNAL · FLAG
2026
93 BIS · Bulletin 130, 'AI and the global economy'. 'The size and persistence of these effects remain highly uncertain', while AI-firm bond issuance tripled from $79B to $243B in two years. The spend is certain; the payoff isn't. · Central-bank synthesis
T2 · WATCH
2026
92 OECD · Compendium of Productivity Indicators 2026. EU labour productivity accelerated 0.2%→1.4% and the US ran 2.2%: 'early, tentative signals' consistent with AI, explicitly not attributed to it. · Harmonized national statistics
T2 · WATCH
APR–MAY 2026
91 Bank of Canada (Alexopoulos). 'Small productivity gains' emerging; Canadian business adoption 3% (2022) → 12% (2025). Speech and analysis; no single document located. · Speeches / analysis
T2 · SIGNAL · FLAG
APR 2026
90 Otis et al. · Management Science publication. The 2024 Kenya RCT clears peer review: the working-paper range of +15–20% for the best performers settles at +15%, with −10% for the weakest. · Peer-reviewed
T1 · SIGNAL
APR 2026
88 Federal Reserve (Allen) · monitoring AI adoption. 18% of firms, 41% of workers, 78% of the labor force at AI-adopting firms. Adoption context, not a value verdict. Document not located; recorded from secondary coverage. · BTOS / RPS / SBU triangulation
T2 · SIGNAL · FLAG
2026
89 C.D. Howe Institute · 'From Hype to Output'. Canadian planned adoption is rising fast, but static enterprise tools trail the personal AI employees already use, a feedback loop between sanctioned and shadow AI. · Think-tank synthesis, Canadian data
T4 · WATCH
Q2 '24–'26
87 Statistics Canada · CSBC series. AI use in production: 6.1% → 12.2% → 19.2% across three years. Adoption, not yet value. · National business survey
T2 · SIGNAL
MAR 2026
86 ECB · SAFE survey + Lane, 'AI and the euro area economy'. Euro-area workers using AI jumped from 26% to 40% in a year, while the ECB frames the payoff as a J-curve: 'by no means automatic; upfront investment required.' · Euro-area firm survey
T2 · WATCH
2026
85 StatCan × CPP (Li & Liu) · complementary capabilities. AI lifts productivity only where complementary capabilities exist, the redesign thesis, in microdata. Publication not located under this title. · Canadian firm microdata
T1 · SIGNAL · FLAG
2026
83 IBM · 'How to Maximize AI ROI' + CEO study. Best-practice product teams report median 55% gen-AI ROI; 85% of CEOs expect positive return by 2027. Vendor + projection. · IBV research
T5 · WATCH · FLAG
FEB 2026
84 Morgan Stanley · 'Mapping AI's Rate of Change', 5th wave. AI-adopter EBIT margins expanded 310bps in 2025, twice the index, and adopter earnings estimates outran the disrupted by 102%. Sell-side, forward-loaded, basket construction opaque. · Analyst-classified adopter baskets
T5 · SIGNAL · FLAG
EARLY 2026
81 US Census · 'The Microstructure of AI Diffusion' (CES-WP-26-25). 18% of firms use AI (32% employment-weighted); broader integration correlates with stronger performance. · 2026 BTOS AI supplement, Nov 2025–Jan 2026
T2 · SIGNAL
FEB 2026
82 NBER · 'Firm Data on AI' (Yotzov, Barrero, Bloom et al.). ~70% of firms use AI; over 80% report no impact on productivity or employment in three years. The first representative firm-level dataset. It lands on 'not yet.' NBER working paper, not peer-reviewed. · ~6,000 executives, stratified, US / UK / DE / AU
T1 · SIGNAL · FLAG
JAN 2026
80 JPMorgan · 'Smothering Heights'. Since 2022, only AI infrastructure providers earn substantial excess returns; AI-user baskets are flat. · Market analysis (Cembalest)
T2 · SIGNAL
JAN 2026
79 Gartner · 'RIFs Before Reality'. <1% of layoffs traced to actual AI productivity gains; cuts running ahead of proven returns. · Trend analysis
T2 · SIGNAL
JAN 2026
78 PwC · 29th Global CEO Survey. 12% report both cost and revenue benefits; 33% either; 56% none significant to date. · 4,454 CEOs
T2 · SIGNAL
JAN 2026
77 Workday × Hanover · 'Beyond Productivity'. ~4 of every 10 saved hours lost to fixing AI output; 14% of employees consistently net-positive. Vendor-funded, adverse finding. · 3,200 respondents
T2 · SIGNAL
JAN 2026
75 BCG · AI Radar 2026. Half of CEOs fear for their job if AI bets fail; the 10-20-70 rule: value is mostly people and process. · Global CEO survey
T2 · SIGNAL
JAN 2026
76 Buechsenschuss, Koch-Bayram, Biemann & Puranam · GenAI and collaboration networks (INSEAD). Randomized grounded-AI deployment raised collaboration and knowledge-network centrality and overall network density. Specialists gained knowledge centrality, generalists gained throughput. Single site, working paper. · Field experiment, 316 employees, 42 teams randomized
T2 · SIGNAL · FLAG
2025PILOTS STALL48 sources
LATE 2025
74 Morgan Stanley · AI software revenue forecast. $1.1T AI software revenue by 2028, 'could pay for itself.' A projection, not a measurement. · Bottom-up forecast
T2 · SIGNAL · FLAG
LATE 2025
73 Josh Bersin · 'AI Tools Everywhere, Yet a Fleeting ROI'. Copilot ROI 'elusive'; individual speed isn't converting to business value. The COI cuts against interest here. · Practitioner analysis
T4 · ECHO
NOV–DEC 2025
72 Accenture · 'Pulse of Change' 2026. Executives enter 2026 confident; the top barrier is 'bringing people along.' Sponsor-funded sentiment. · Two global surveys
T5 · WATCH
NOV 2025
71 KPMG Canada · gen-AI adoption survey. 93% using or piloting AI, 2% report measurable ROI. KPMG's own partner: 'disappointing and surprising.' · 753 Canadian leaders
T2 · SIGNAL
NOV 2025
70 McKinsey · State of AI. 88% of organizations use AI in ≥1 function; only 39% report any EBIT impact. · Global survey
T2 · SIGNAL
NOV 2025
69 Sapient Insights · 28th HR Systems Survey. HR-tech spend growth slowing sharply, a realized-value signal from the buy side. · 9,886 HR pros / 4,670 orgs, non-sponsored
T2 · SIGNAL
OCT 2025
68 Deloitte · State of AI in the Enterprise 2026. Spending rises while ROI stays elusive; calls for strategic rather than scattered deployment. · 3,235 leaders
T2 · SIGNAL
OCT 2025
67 St. Louis Fed · productivity update. 5.4% time savings (~2.2 hrs/wk) holds as adoption climbs. · RPS; adoption at 54.6%
T2 · SIGNAL
OCT 2025
66 Gartner · HR leader survey. 88% report no significant business value from AI tools yet. · 114 HR leaders (small n)
T2 · SIGNAL
OCT 2025
65 Wharton × GBK · 'Accountable Acceleration'. ~75% report positive ROI; 72% formally track it. The strongest disconfirming source, self-reported, consultancy co-authored. · 800+ enterprise leaders
T2 · SIGNAL
SEP 2025
64 SBA / Census BTOS spotlight. Adoption rising; the small-firm / large-firm gap narrowing. Adoption, not value. Release not located; recorded from secondary coverage. · Business Trends & Outlook Survey
T2 · ECHO · FLAG
SEP 2025
63 BetterUp × Stanford · 'workslop' survey (HBR). 41% received AI 'workslop' in a month; each instance took ~2 hours to untangle. Small convenience sample: but it is the receiving end of the rework tax. · Survey, 1,150 US workers
T3 · WATCH · FLAG
SEP 2025
60 Josh Bersin × AMS · the TA revolution. AI-enabled talent acquisition delivers 2–3× faster time-to-hire. Analyst with an HR-AI practice. COI noted. Practitioner research; no primary document located. · HR TA research
T4 · WATCH · FLAG
SEP 2025
61 OpenAI · GDPval benchmark. Frontier models matched or beat human experts on 40–48% of blinded professional deliverables. Capability in the lab, not ROI in the firm. · 1,320 expert tasks, 44 occupations; grading harness closed
T2 · WATCH
SEP 2025
62 Finkenstadt et al. · 'The Silo Effect in the AI Age' (California Management Review). Argues AI alters the logic, cost and cadence of cross-functional collaboration. Useful vocabulary for the coordination question; not evidence for it. · Conceptual framework, no primary data
T4 · WATCH
SEP 2025
58 BCG · 'The Widening AI Value Gap'. 60% generating no material value despite investment; ~5% creating substantial value at scale. · 1,250 respondents
T2 · SIGNAL
SEP 2025
59 OpenAI × Harvard/Duke · 'How People Use ChatGPT' (NBER). Work use is a shrinking share of traffic; the main economic channel is decision support in knowledge work. Usage, not ROI: and OpenAI-authored. · >1.1M conversations, privacy-preserving pipeline
T1 · WATCH · FLAG
SEP 2025
57 Bain · Global Technology Report. AI needs ~$2T annual revenue by 2030 to fund compute; ~$800B shortfall likely. A projection, not a measurement: the $600B question, tripled. Edition not identified in our record. · Market analysis
T2 · SIGNAL · FLAG
SEP 2025
55 Google Cloud · 'ROI of AI 2025'. 74% report ROI within the first year; 88% of agentic early adopters see ROI. The vendor pole, 2025 edition. · 3,466 execs, 24 countries, commissioned
T5 · WATCH
SEP 2025
56 Lin & Maruping · 'Organizing for AI Innovation' (MIS Quarterly). AI innovations are less radical and more process-oriented than comparable IT innovations. Organizing AI like conventional IT is a documented failure mode. · Matched-sample analysis of US patents
T1 · WATCH
AUG 2025
54 Xu et al. · 'Echoes in AI' (PNAS). Generated stories echo the same idiosyncratic plot elements across generations and across models; the diversity gap widens as the corpus grows. Homogenization confirmed at scale. · Plot-uniqueness scoring across models
T1 · SIGNAL
AUG 2025 · REV FEB 2026
53 Brynjolfsson, Chandar & Chen · 'Canaries in the Coal Mine' (Stanford). Employment for 22–25-year-olds in the most AI-exposed jobs fell 13% relative since late 2022, concentrated where AI automates rather than augments. Substitution is real even where the P&L is silent. The author reported 16–17% in a Jul 2026 interview, and an age band of 22–26; we carry the published 13% until the revision appears in a paper. · ADP payroll microdata, millions of workers
T1 · SIGNAL
AUG 2025
52 MIT NANDA · 'The GenAI Divide'. ~5% of custom pilots show measurable P&L impact. Viral, narrow test, publisher COI, and the sample is reported inconsistently: 52 interviews + 153 leaders in some accounts, 150 interviews + 350 employees in others, including Fortune's interview with the lead author. A pessimistic stat can be hype too. · 52 interviews + 153 leaders
T2 · WATCH · FLAG
AUG 2025
51 Budzyń et al. · endoscopist deskilling after AI exposure (Lancet Gastro Hep). Adenoma detection in standard colonoscopy fell from 28.4% to 22.4% after routine AI exposure, a 6.0 point drop. First real-world evidence of AI-induced deskilling on a patient endpoint. · 1,443 non-AI colonoscopies, four centres
T1 · SIGNAL
JUL 2025
50 Ed Zitron · 'The Hater's Guide to the AI Bubble'. Gen-AI economics unsustainable, ROI absent. The pessimistic pole, flagged the same way vendor hype is. · Long-form analysis
T4 · WATCH · FLAG
JUL 2025
49 METR · experienced-developer RCT. Developers were 19% SLOWER with AI, after predicting +24%, and believing +20% afterward. Small n, early-2025 tools; the best-designed contrarian result. In Feb 2026 METR redesigned the experiment and said it now believes developers are faster, while cautioning its own new data is 'only very weak evidence.' The '18% speedup' circulating in coverage is not METR's claim. · 16 devs, 246 real tasks, RCT
T1 · SIGNAL
2025
48 UCLA · ambient AI scribes RCT (NEJM AI). AI scribes cut documentation time and reduced burnout in a registered clinical trial. Medicine joins law and software on the task-level 'yes' side. · Registered RCT, 238 physicians, 14 specialties
T1 · WATCH
JUN 2025
46 PwC · AI Jobs Barometer 2025. Productivity growth in the most AI-exposed industries nearly quadrupled (7%→27%) since 2022; AI-skill wage premium 56%. Industry-level correlation, measured by a firm that sells AI services. · ~1B job ads + firm financials, six continents
T2 · SIGNAL · FLAG
JUN 2025
47 Si, Hashimoto & Yang · 'The Ideation-Execution Gap' (Stanford). After execution, LLM-idea scores fell significantly more than human-idea scores and the ranking reversed. Apparent novelty did not survive contact with the work. · 43 researchers, 100+ hours each, blind pre/post review
T2 · SIGNAL · FLAG
JUN 2025
45 Jabra · Happiness Research Institute, 'AI at Work'. Daily AI users reported 34% higher job satisfaction and up to 20% more stress. Both poles inside one vendor survey with undisclosed method; carried as an example of the genre. · Global survey, sample not disclosed
T5 · SKIP · FLAG
JUN 2025
43 Statistics Canada · employment effects. ~90% of AI-adopting businesses report no change in staffing. Release not named in our record. · Business survey
T2 · SIGNAL · FLAG
JUN 2025
44 Carnegie Mellon · 'TheAgentCompany' benchmark. The best agent completed 30% of realistic office tasks autonomously; every model failed the majority. A measured ceiling under the year's agentic-ROI pitch. · 175 simulated office tasks, open benchmark
T1 · SIGNAL
MAY 2025
42 IBM · CEO study. Only 25% of AI initiatives delivered expected ROI; 16% scaled enterprise-wide. A vendor conceding the gap. · 2,000 CEOs
T5 · WATCH
2025
41 Kadolkar et al. · algorithmic management review (J. Organizational Behavior). Algorithmic management delivers efficiency and clarity while carrying control, surveillance and autonomy costs. Effects are conditional, not uniformly negative. · Systematic review, input-process-output synthesis
T1 · WATCH
MAY 2025
40 MIT · institutional retraction of Toner-Rodgers. 'No confidence in the provenance, reliability, or validity' of the data. The reason the 2024 dot above wears an ✕. · Integrity review
T1 · SKIP
MAY 2025
39 Humlum & Vestergaard · Danish payroll study. Precise zeros on earnings and hours (CIs rule out effects over 1%); average time savings 2.8%. The strongest macro 'not yet.' · 25,000 workers / 7,000 workplaces, admin data
T1 · SIGNAL
MAY 2025
38 Brynjolfsson, Li & Raymond · 'Generative AI at Work' (QJE publication). Peer-reviewed publication of the 2023 NBER paper: +15% issues resolved per hour, largest gains for the least experienced, small quality declines for the most experienced. · 5,172 support agents, staggered rollout
T1 · SIGNAL
MAY 2025
37 Babina, Fedyk, He & Hodson · AI and systematic risk. AI-investing firms carry higher market beta: what looks like adopter alpha prices as systematic risk. Returns to adoption may be compensation, not free lunch. · US public firms, AEA P&P
T1 · WATCH
H1 2025
35 S&P Global Market Intelligence. 42% abandoned most AI initiatives, up from 17% a year earlier; average firm scrapped 46% of PoCs. · 1,000+ enterprises
T2 · SIGNAL
2025
36 Giuntella et al. · AI and worker wellbeing (Scientific Reports). No sizeable negative effect on wellbeing or mental health, with small improvements in physical health. The strongest measured wellbeing evidence, but the window closes before generative AI. · German SOEP panel 2000–2020, event study
T1 · SIGNAL · FLAG
APR 2025
34 IMF · 'The Global Impact of AI: Mind the Gap'. High-adoption scenario adds 1.8% to global TFP within five years. A projection, not a measurement, the optimistic twin of Acemoglu's ceiling, flagged the same way. · Scenario model
T2 · WATCH · FLAG
MAR 2025
33 Wan & Kalman · diverse AI personas and ideation. Diverse generative personas preserved output diversity against a human-only baseline, isolating which design choices drive the creativity-diversity trade-off. · Two-phase experiment, 300 generated plots
T2 · WATCH
MAR 2025 · JELS 2026
32 Schwarcz et al. · 'AI-Powered Lawyering' RCT. Reasoning models and RAG raised legal-task productivity 50–130% on five of six tasks, with quality up, the first strong task-level result outside software. · RCT, law students, six legal tasks
T1 · SIGNAL
FEB 2025
31 Anthropic Economic Index · first report. Usage concentrated in ~36% of occupations; augmentation outweighs automation. Vendor telemetry. · Claude usage analysis
T5 · WATCH
FEB 2025
29 St. Louis Fed · gen-AI productivity. Users save 5.4% of work hours → roughly 1.1–1.4% workforce-level productivity. Modest, real, self-reported. · Real-Time Population Survey
T2 · SIGNAL
FEB 2025
30 GitClear · 'AI Copilot Code Quality'. Refactored code fell from 24.8% of changed lines to 9.5%; duplicated blocks rose eightfold in 2024. Faster authoring, rising maintenance debt. Vendor, method disclosed. · 211 million changed lines, 2020–2024
T2 · SIGNAL · FLAG
JAN 2025
28 McKinsey · 'Superagency in the Workplace'. Only 1% of leaders call their gen-AI rollout mature. Employees are ready; leadership is the bottleneck. · 118 CxOs + ~3,000 employees
T2 · SIGNAL
EARLY 2025
27 BCG · 'From Potential to Profit'. Only 25% realize significant value; 4% substantial. Leaders show 2.1× the ROI of laggards. · 1,400+ C-suite
T2 · SIGNAL
2024THE $600-BILLION DOUBT19 sources
2024/25
26 Accenture · 'Making Reinvention Real'. Scaled, tailored deployments 3× more likely to exceed ROI expectations. Sponsor-funded. · Global exec survey
T5 · ECHO
DEC 2024
25 IBM × Morning Consult · 'ROI of AI'. ROI ratings mixed; many deployments break-even or hard to measure, a vendor survey landing lukewarm. · ~2,400 IT decision-makers
T5 · WATCH
2024 WAVES
23 Deloitte · State of GenAI in the Enterprise. ROI 'encouraging' but access below 40% of the workforce; two-thirds scale 30% or fewer of experiments. · Quarterly surveys, ~2,800 leaders (Q4)
T2 · ECHO
LATE 2024
24 Google DORA · State of DevOps 2024. AI adoption associated with a 1.5% fall in delivery throughput and a 7.2% fall in delivery stability, while individual productivity and satisfaction rose. Individual up, system down. · ~3,000 technology professionals
T2 · SIGNAL · FLAG
NOV 2024
21 Microsoft / IDC · 2024 AI Opportunity Study. Claims $3.7 per $1, and $10.3 for 'leaders.' The strongest yes of the period, and the weakest independence. · 4,000 leaders, sponsor-funded
T5 · ECHO
NOV 2024
22 Toner-Rodgers · 'AI, Scientific Discovery…' (MIT). Claimed +44% materials discovered, +39% patents. RETRACTED May 2025. MIT disavowed the data. The figures still circulate: an EU commissioner cited them in Sep 2025. · Claimed 1,018 scientists
T1 · SKIP · RETRACTED
OCT 2024
20 Vaccaro, Almaatouq & Malone · human-AI meta-analysis (Nature Human Behaviour). On average human-AI combinations performed significantly worse than the better of human or AI alone, with losses on decision tasks and gains on content creation. Combination is not complementarity. · 106 studies, 370 effect sizes
T1 · SIGNAL
OCT 2024
19 Sapient Insights · 27th HR Systems Survey. HR AI adopters report 7–10% gains in business, talent, and HR outcomes. A rare independent HR datapoint. · 3,318 organizations, non-sponsored
T2 · ECHO
SEP 2024 · ICLR 2025
18 Si, Yang & Hashimoto · 'Can LLMs Generate Novel Research Ideas?' (Stanford). LLM-generated research ideas were judged significantly more novel than expert-human ideas, though slightly less feasible. Novelty judged at the idea stage, by expert perception. · 100+ NLP researchers, blind review
T2 · WATCH · FLAG
AUG 2024
17 Google Cloud · 'The ROI of Gen AI'. Majority of executives report gen-AI ROI or value. Sponsor-funded, self-reported. · 2,508 execs, commissioned survey
T5 · WATCH
2024
16 RAND · 'Root Causes of Failure for AI Projects' (RRA2680-1). ~80% of AI projects fail, roughly twice the failure rate of traditional IT projects. · 65 data scientists and engineers, interviews
T2 · SIGNAL
JUL 2024
15 Gartner · PoC abandonment prediction. ≥30% of GenAI projects will be abandoned after proof of concept by end-2025. A forecast, not a measurement; cited reason is unclear business value. · Forecast
T2 · SIGNAL · FLAG
JUL 2024
14 Doshi & Hauser · 'Generative AI Enhances Individual Creativity but Reduces Collective Diversity' (Science Advances). AI-assisted stories were rated more creative and better written, especially for weaker writers, while being measurably more similar to each other. Individually better off, collectively narrower. · RCT, writers with and without LLM story ideas
T1 · SIGNAL
2024
13 Cui et al. · Copilot field experiments. +26% completed tasks and weekly pull requests. Developers only; vendor-affiliated. · 3 field experiments (MS, Accenture, F100)
T1 · SIGNAL
JUN 2024
12 Goldman Sachs · 'Too Much Spend, Too Little Benefit?'. ~$1tn capex with 'little to show for it so far.' Acemoglu: ~0.9% GDP over a decade. House view: 6.1%. Internally contested. · Research report
T2 · SIGNAL
JUN 2024
11 Sequoia · 'AI's $600 Billion Question'. Revenue required to justify AI capex is not materializing at anywhere near the needed scale. · Analysis (Cahn)
T4 · WATCH
MAY 2024
10 Acemoglu · 'The Simple Macroeconomics of AI'. Ceiling of ≤0.66% total-factor-productivity gain over 10 years; under 0.53% once hard-to-learn tasks are weighted. · NBER w32487 → Economic Policy 2025
T1 · SIGNAL
FEB 2024
9 Otis et al. · Kenyan entrepreneurs RCT. No average effect: +15–20% for high performers, −10% for the weakest. 'AI helps the best and hurts the rest.' · 640 entrepreneurs, field RCT
T1 · SIGNAL
JAN 2024
8 Meincke, Mollick & Terwiesch · 'Prompting Diverse Ideas'. Prompt design substantially raises the variance of AI-generated idea pools. Homogenization is partly a deployment choice, not a property of the technology. · Experimental comparison of prompting strategies
T2 · WATCH
2023THE TASK-LEVEL PROOF7 sources
NOV 2023
7 Gartner early-adopter survey. Adopters self-report +15.8% revenue, +15.2% cost savings, +22.6% productivity. Report title not located. · 822 business leaders
T2 · ECHO · FLAG
NOV 2023
6 Microsoft / IDC · Business Opportunity of AI. Claims $3.5 returned per $1 invested. Self-reported, sponsored, method undisclosed. · 2,000+ leaders, sponsor-funded
T5 · WATCH
SEP 2023
5 Dell'Acqua et al. (HBS × BCG). +40% quality inside GPT-4's frontier, and worse performance outside it. The 'jagged frontier.' · 758 BCG consultants
T1 · SIGNAL
JUL 2023
4 Noy & Zhang (Science). Writing time −40%, quality +0.45 SD in a randomized experiment. · 453 professionals, RCT
T1 · SIGNAL
MAY 2023 · J.FIN 2026
3 Eisfeldt, Schubert, Zhang & Taska · 'Generative AI and Firm Values'. Firms most exposed to gen-AI earned +0.4% daily excess returns in the fortnight after ChatGPT launched. A repricing of expectations, the market paying for anticipated ROI, not booked ROI. · US public firms, workforce-exposure event study
T1 · SIGNAL
APR 2023
2 Brynjolfsson, Li & Raymond (NBER). +14% issues resolved per hour; +34% for novices. Field data, staggered rollout. · 5,179 support agents
T1 · SIGNAL
FEB 2023
1 Peng et al. · GitHub Copilot RCT. Copilot group completed the task 55.8% faster. Narrow test, vendor-affiliated authors. · 95 developers, RCT
T1 · SIGNAL

Reviewed and skipped: vendor ROI calculators and case-study one-pagers (a class, not a single document) failed the method-disclosure test at intake. Skipped, not graded; absence is the verdict. Known gaps: The Economist’s paywalled analyses could not be verified against primaries, and the German institutes (ifo, ZEW, IZA) and Kellogg remain unswept, open items for the next pass.

How this was made
David Herrera

David has spent over 15 years in management consulting, working with clients across the world on organizational design, total rewards, people analytics, operating model and talent management. In-house he has led teams driving transformation grounded in both technology and people strategy.

That is the perspective these briefs are written from. He has priced the jobs, designed the role architecture, and run the systems that are supposed to turn saved time, innovation, and collaboration into business value.

Agentic HR is David’s personal project, built around a single thesis: AI changes the doing, not the deciding. That makes it an operating-model question, not a technology question.

How sources are rated. Every source gets a grade from T1 to T5. T1 means peer-reviewed. T5 means the study never says how it was done. What decides the grade is whether the method is published, not who paid for it, so a vendor study that shows its work outranks an academic one that doesn’t. Commercial interest is noted, never used to push a source down. Studies that never explain their method are left out rather than graded, and this page says what was left out. Evidence is hunted on both sides: a gloomy number stretched past what it can support gets flagged exactly like a promotional one. Verdicts are decided by a person, reviewed monthly, and never changed automatically.

How this was made. AI did the searching, the sorting and the first drafts. David set the question, decided how sources get rated, and made the call on what the evidence adds up to. He checked every source that carries weight in the argument against the original document. When the AI got something wrong, he corrected it. That happened more than once.

Why say so: this brief argues that AI only pays off where the work was redesigned. It would be odd not to say how this work was designed.

Corrections. If a source is miscoded, a number is wrong, or a verdict is not supported by what is in the ledger, say so. It will be fixed in the next review and the change noted on the page. All feedback is welcomed and encouraged: David.herrera@agentichr.ca