Nobody bought the manager
A VP asked what the ROI was on all the tech we had bought. I was the manager, so I took the question downstairs to the teams who were supposed to be using it, and I looked. The rollout itself had gone beautifully. On day one the congratulations emails went round and everybody said Bravo and a new era was coming, and there was genuine excitement in the room about the new shiny thing. Four weeks later the database was firing maybe one event a day. The last person who actually knew how to use it had either left or been promoted.
That shape repeats, and it repeats almost identically. A company buys the tools, runs the training, sends the announcement, and then waits about six months for the transformation to show up, and it does not show up, and now everyone is slightly embarrassed about the whole thing and nobody wants to be the one to say so out loud. The licences are paid. The usage dashboards are green, or green enough, or IT has still not produced them and now the consultants have to be brought in to find the data, which nobody discussed in any detail although everyone did agree on that beautiful slide where data was everything. And the work looks almost exactly like the work looked before, except faster in the places that were never the bottleneck.
I have a theory about where the money went, and I want to be honest that it is a theory, because most of the evidence I am going to show you is correlational and I am not going to pretend otherwise. The organisation bought the capability and skipped the only person who could have installed it. It bought the AI and it did not buy the manager.
And nobody asked. Not the managers, not the people on the ground who could have walked you to the actual bottleneck in an afternoon and told you plainly what would help and what would sit unused. McKinsey's own senior partner puts it at more than eighty percent of companies reporting no bottom-line impact from what they have invested, and the same piece says in its subtitle that the challenge lies in redesigning workflows and leadership and culture rather than in the technology [1]. You can pay a great deal of money for that sentence. You can also go downstairs, skip the deck, and talk tachles with the people who will have to live inside whatever you sign.
The transmission mechanism
Think about how a new capability actually reaches a person doing the work. It does not arrive through a licence. It arrives through the person who decides what the work is, what good looks like, whether it is safe to try something and have it not work, and whether the twenty minutes you just spent arguing with a model was diligence or slacking. That person is your direct manager, and they are the transmission mechanism between the thing the company bought and the thing the company hoped would change. The relationship a person ends up having with these systems exists whether or not anybody sat down and designed it, and its terms get set by default when nobody sets them on purpose [8]. On a team, the person setting them by default is the manager.
Two surveys say the same thing from different directions. Gallup's State of the Global Workplace 2026 finds that US employees who strongly agree their manager actively supports their team's use of AI are 8.7 times as likely to strongly agree that AI has transformed how their work gets done, and 7.4 times as likely to say it gives them more opportunities to do what they do best [2]. Microsoft's 2026 Work Trend Index reaches the same shape from the other side: in a smaller study of eighteen hundred workers reported alongside its main survey of twenty thousand, employees whose managers actively modelled AI use rather than merely permitting it reported a 17-point lift in perceived AI value, a 22-point lift in critical thinking about their own use, and a 30-point lift in trust in agentic AI [3].
Both are self-reported and correlational, and the Gallup ratios are top-box comparisons, strongly-agree against everyone else, which is exactly the framing that produces enormous numbers. I am not selling you an 8.7x lever. What survives the caveats is a direction: whether a rollout lands or evaporates tracks the manager almost perfectly, and it barely tracks the tooling at all. The word carrying the Microsoft half of it is modelled. Permitting means telling your team the licence is there and the policy allows it. Modelling means opening the thing in front of them, on real work that matters to you, and letting them watch you get it wrong twice before it comes out right. That second one is what transfers, and almost nobody does it. The question I was sent downstairs with was about the tooling, and what I found down there was not about the tooling at all.
The gap where the money goes
Here is the part that should be embarrassing for an industry that has spent an extraordinary amount of money on this.
Fewer than one in three US employees in organisations that have already begun implementing AI strongly agree that their manager actively supports their team's use of it [2]. Not "has heard of it", not "permits it". Actively supports it. A separate Gallup study in Germany lands close to the same place at 21 percent [4], which I would read as a second look at the same shortage rather than as a clean cross-country comparison.
So the multiplier exists, and it is almost entirely unactivated. The organisation bought the capability, installed it on everyone's laptop, and then left it sitting on the desk of a person who was never told that switching it on was now part of their job, or who was told by means of a PDF called How To Use AI, which lands in the same inbox as the fire drill notice and gets read with the same attention.
And it gets worse, because of what has been happening to that person independently. Gallup finds manager engagement has dropped nine points since 2022, falling to 22 percent, with the sharpest fall between 2024 and 2025 [2]. That is a global figure and a general one rather than anything AI-specific, and it began before the current wave, so I am not claiming AI caused it. The honest version of the claim is smaller and stranger: at the exact moment we decided that the manager would be the mechanism by which AI reached the workforce, the manager was quietly running out of road. We loaded the transmission at the point of its lowest torque.
What the manager is actually being asked to do
The Microsoft data has one more finding in it that reframes the whole thing, and it is the one I would put on a slide if I still made slides.
Leaders are twice as likely as their employees to say that reinventing how work gets done with AI is rewarded regardless of outcome: 21 percent against 10 percent [3]. Read that again, because it is not a statement about optimism.
It is a statement about who thinks it is safe to fail. The people at the top believe that experimenting is rewarded even when it does not pay off, and the people doing the experimenting mostly do not believe that, and both groups are describing the same company.
That gap is the whole problem, and it explains why the tooling budget keeps failing to convert. Trying a new way of working with AI means being slower and worse at your job for a while, in public, in front of the person who writes your review. That cost is measurable and it is worse than it feels: in a randomised trial, experienced developers given AI tools took 19 percent longer on real tasks, and still believed afterwards that they had been about 20 percent faster [6]. Nobody does that on the strength of a licence and a webinar. They do it when the person above them has made it survivable, and made that credible by being visibly bad at it first.
Which means what the manager is being asked to provide is not AI expertise. It is cover. Cover is a specific and fairly unglamorous thing. It is saying out loud, before anybody needs it, that the hour somebody spent failing to get a useful answer out of a model is not an hour that will come up in their review. It is taking the first bad result yourself, in front of them, on something that matters to you. It is having an answer ready for the person above you when the numbers dip in month two, so that the dip is a cost you budgeted rather than an incident you are explaining. None of that requires the manager to understand the model. It requires them to spend their own credibility, and credibility is the one thing the licence does not come with.
The bill for not providing it takes a while to arrive, because AI feels like speed to everybody at the start. If you know what you are doing, that speed is real. If you do not, the same speed produces volume nobody has checked, and it does not vanish, it queues: among coding sessions that ran into trouble, Anthropic found novices reached a verified result 4 percent of the time against 15 percent for experts, and abandoned one in five outright [7]. Milena has already made the sharper version of this point in the series, that the same system can amplify you or quietly take over the part of you that was doing the thinking, and what decides which is not the tool but what you ask it to amplify [9]. The backlog arrives three months later on the manager's desk, which is the second job nobody mentioned when the licences arrived.
Coaching the manager, not training the workforce
I have spent a long time now on the other side of this, working with managers rather than rolling out tools, and the thing I would say to anyone about to sign a large AI contract is that you have almost certainly misprized the two halves of what you are buying.
The first half is the tooling. It has a price list, it arrives on a predictable date, and it is the cheap half whatever the invoice says. The second half is the twenty or thirty people in the middle of your organisation who decide, mostly unconsciously and mostly in one-to-ones, whether this is something the team does properly or something the team performs for the dashboard. That half is the expensive one, because there is no price list for it and no date it arrives on, and it is the half that decides whether the first one returns anything at all. It is also the half that should be getting the larger share of the budget, and in every rollout I have watched it gets almost none of it. If the ROI on coaching those people exceeds the ROI on the tooling, and I think it does, then the budget you have just approved is upside down, and it will keep being upside down no matter which model you standardise on next quarter.
Where I might be wrong. The strongest objection is that I have the arrow backwards. The survey numbers above are correlational, and a perfectly good story runs the other way: healthy teams with engaged managers adopt everything faster, AI included, and the manager is a symptom of a functional organisation rather than the cause of a successful rollout. If that is true then coaching managers is treating the readout instead of the illness.
One experiment gets close to settling it. Kim, Kim and Koning randomised 515 startups in a three-month accelerator: every firm received the technical AI training, the API credits and the mentorship, and the treated firms received one thing more, information about how other companies had reorganised their work around AI. Those firms found 44 percent more use cases and generated 1.9 times the revenue of the control group, and the authors call that causal evidence, because they randomised it [5]. It does not vindicate me, since these are startups rather than managers inside established organisations, it is still a working paper, and the treatment was case studies rather than anything resembling coaching. What it settles is the half I could not settle myself. Hold the tooling constant, vary only the understanding of where the work ought to change, and performance moves. The tooling was never the variable.
What would tell me I am wrong is an organisation that invests seriously in its managers and sees no differential AI adoption against a matched peer that only bought licences, and I would want to know. The other objection is that all of this is convenient for me, since coaching managers is roughly what I do. I notice that. It does not make the numbers different, and it does mean you should make me argue for it rather than take it.
The thing nobody put in the contract
The contract you signed has a number of seats in it, a support tier, a data-processing addendum, and a start date. It does not have a line for the person who has to make it safe to be bad at this in public for a month. I wrote early in this series about turning up to an interview with my own digital team already running, and about how no employment contract yet knows what to do with that [10]. This is the same gap seen from the buying side: the paperwork has caught up with the software and not with the people around it.
That line does not exist on any invoice I have seen. It is still the thing you are buying, or failing to.
So before the next renewal: who, by name, is going to switch this on for your team?
References
- Alexis Krivkovich with Lucia Rahilly, "AI is everywhere. The agentic organization isn't—yet," The McKinsey Podcast (McKinsey & Company, 2 April 2026). The source of the figure that more than 80 percent of companies say they are not yet seeing bottom-line impact from their AI investment, and of the subtitle placing the real challenge in redesigning workflows, leadership and culture rather than in the technology. Worth knowing what this is: an edited podcast transcript, not a research report. It gives no sample size, instrument or method for the 80 percent, which is why the essay attributes it to a senior partner speaking rather than to a survey. https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/ai-is-everywhere-the-agentic-organization-isnt-yet
- Gallup, State of the Global Workplace 2026. Carries three of the numbers used here: the 8.7× and 7.4× ratios for employees whose managers actively support AI use, the finding that fewer than one in three US employees in AI-implementing organisations strongly agree they have that support, and the nine-point fall in manager engagement since 2022 to 22 percent. Every ratio is a top-box comparison of "strongly agree" against everyone else, which is what makes the ratios large; all of it is self-reported and correlational, and the engagement figure is global and general rather than anything AI-specific. https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx
- Microsoft, 2026 Work Trend Index Annual Report ("Agents, human agency, and the opportunity for every organization"). This document carries two different studies and the essay uses one figure from each, so it is worth separating them. The main WTI global survey covers 10 markets (US, BR, AU, IN, JP, FR, DE, IT, NL, UK), fielded by Edelman Data x Intelligence between 18 February and 20 April 2026, analysed n=20,000; that survey is the source of the leaders-versus-employees gap on whether reinventing work with AI is rewarded regardless of outcome, 21 percent against 10 percent. The manager-modelling lifts of 17, 22 and 30 points come instead from a separate Microsoft-led study of 1,800 workers globally, reported inside the same document — a sample an order of magnitude smaller, which is why the essay names it rather than folding it into the headline survey. A "point lift" means percentage points on a self-reported survey item, not a measurement of the work. https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization
- A Gallup study in Germany, reported inside State of the Global Workplace 2026 (reference 2), which states: "A Gallup study in Germany found similarly low support: 21% of employees in organizations that use AI said their manager actively supports their team's use of AI." It is a separate study from the US survey that produces the other Gallup numbers here, so the two are not a like-for-like cross-country comparison, and the essay says so where it uses the figure. The report does not date the German study or cite it further, which is why no year is claimed for it. https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx
- Hyunjin Kim, Dahyeon Kim and Rembrand Koning, "Mapping AI into Production: A Field Experiment on Firm Performance," INSEAD Working Paper 2026/20/STR, 30 March 2026. A randomised controlled trial across 515 high-growth startups in a three-month accelerator, preregistered on the AEA RCT Registry as AEARCTR-0016746. Both arms received the technical AI training, the API credits and the mentorship; only the treated firms also received information on how other companies had reorganised their work around AI. The one piece of genuinely causal evidence in this essay, and it bears on the layer rather than on the manager. Still a working paper, not peer-reviewed. https://ssrn.com/abstract=6513481
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (July 2025). A randomised controlled trial across 246 real issues in developers' own repositories, averaging about two hours each, in which experienced open-source developers took 19 percent longer when allowed AI tools, having predicted a 24 percent speed-up beforehand and still estimating afterwards that they had been about 20 percent faster. The perception gap is the part that matters here: the person paying the cost is the least able to see it, which is why somebody above them has to budget for it. Read the limits: 16 developers, working on mature codebases they already knew deeply, using early-2025 tooling (mostly Cursor Pro with Claude 3.5/3.7 Sonnet), and it measures time rather than quality. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ · https://arxiv.org/abs/2507.09089
- Zoe Hitzig, Maxim Massenkoff, Eva Lyubich, Shaoyi Zhang, Ryan Heller and Peter McCrory, "Agentic coding and persistent returns to expertise" (Anthropic, 16 June 2026). Roughly 400,000 interactive sessions from about 235,000 people, October 2025 to April 2026. Among sessions that hit trouble, the share still reaching a verified result rose from 4 percent for novices to 15 percent for experts, and 19 percent of novice sessions in difficulty were abandoned outright against 5 to 7 percent for everyone else. The domain is coding specifically, so reading it across knowledge work generally is an extrapolation rather than a finding. https://www.anthropic.com/research/claude-code-expertise
- Vlad Sterngold, Relationship Design (The Symbiotic Mind, Post 010). The pillar essay for this one: the relationship with an AI system already exists and already carries terms, and if you do not set your half of them deliberately you inherit them by default. This essay is that argument moved from the individual to the team, where the person setting the terms by default is the manager. https://symbiotic-mind.com/posts/010-relationship-design/
- Milena Nikolova, Amplification (The Symbiotic Mind, Post 003). The distinction the backlog section rests on: the same system can amplify you or take over the part of you that was doing the thinking, and what decides which is not the tool but what you ask it to amplify. https://symbiotic-mind.com/posts/003-amplification/
- Vlad Sterngold, My API, Not My Resume (The Symbiotic Mind, Post 002). The competitive unit is the person plus the AI team they bring, and employment contracts have not caught up with that. The closing section here is the same gap seen from the buying side rather than the hiring side. https://symbiotic-mind.com/posts/002-my-api-not-my-resume/
🗣 ME (43%): The opening, which is a thing that happened. A VP asked what the ROI was on all the tech we had bought, I took the question downstairs, and four weeks later the database was firing about one event a day while the last person who knew how to use the thing had either left or been promoted. The refusal to let a consultant answer a question the people on the ground could have answered in an afternoon, and the instruction to go down there and talk tachles instead. The observation that AI speed is real for people who know their domain and a queue of unchecked work for everybody else. The insistence that what a manager is being asked for is cover rather than expertise. And the decision to open on the room rather than on the pattern.
🤖 AI (57%): Most of the sentences, the order they arrive in, and the caveat apparatus. Assembling the source pack and checking every number against the primary publications rather than against the summaries, which is how two claims in the research brief behind this piece turned out to be misstated and were dropped, and how the widely-quoted figure about most AI pilots failing was kept out on the grounds that it counts the organisations that never piloted at all. The references, and the admission attached to each one about what it does not show.