This week kept circling one idea. Organizations keep deciding that value lives in the visible artifact and not in the human work that produced it, and they keep paying for the artifact.
Wednesday's article said a Cornell study found that expensive office leases are the strongest predictor of a return-to-office mandate. Read directly, the study points the other way: lower office rents are associated with more in-person work, and the more consistent predictors are firm and manager characteristics rather than real estate. The error was ours, from leaning on secondary summaries instead of the paper itself. The article's larger point, that these mandates are not really about productivity, still holds, and its turnover figures are accurate. We have appended a full correction to the article and left the original text in place.
Read the full correction →Monday's article opened on a statistic, attributed to Harvard Business School, that roughly six in ten managers spend more than half their week on administrative work. We could not source that finding and should not have run it. The accurate, sourced version of the same point comes from McKinsey's middle-manager research. The piece also described Shopify, Klarna, and Duolingo as having thinned their management ranks, when their moves were AI-first headcount and contractor decisions rather than cuts to management. The article's core argument is unaffected. A full correction is appended to the article.
Read the full correction →McKinsey published its HR Monitor 2026 this month, drawing on workforce data across ten countries, and its central instruction is that companies shift from planning operational capacity to planning capability for a workplace where people and AI split the work.1 Set that next to the figure from McKinsey's State of Organizations 2026 research, the one recommending five dollars spent on people for every dollar spent on AI technology2, and the distance between the advice and the actual budget is hard to miss. Most spending runs the other way, and most org charts are being redrawn around the technology while the human side is where the cutting happens.
Everything I wrote this week turned out to be a version of that same imbalance. The throughline was never AI itself. It was where organizations decide the value sits, and they keep deciding it sits in the visible artifact rather than in the human work that produced it. A middle manager's paperwork looks like overhead until you notice it was the residue of judgment nobody else was doing. An AI mandate looks owned until you ask for the single name attached to it and four people answer at once. In both cases the money and the structure flow toward the thing you can see, and the capability that makes that thing worth anything gets filed under discretionary. McKinsey's five-to-one ratio is a way of saying capability is not discretionary, and that the companies starving it are the ones that will go on reporting AI investment with no return to point to.
A working theory of why so much effort produces so little, and what honest measurement would have to look like.
The review ran ninety minutes and almost all of it was green. Someone had counted the AI pilots launched that quarter, and the number was impressive. There were charts for tools adopted, hours logged against the new platform, features shipped, dashboards rolled out across three departments. Every team had been busy, and the slides proved it. Near the end, a board member who had stayed quiet the whole time asked a single question. Of everything we just looked at, what changed for a customer? The room went looking for the answer and could not find one.
That scene is playing out in a lot of companies right now, and there is a number behind it. MIT researchers studying enterprise AI in 2025 found that roughly ninety-five percent of pilots produced no measurable return. Not a thin return. None that anyone could trace to the business. The pilots happened. The money went out. The activity was real and visible on every slide. What none of it produced was a result anyone could point to.
An old line from Peter Drucker sits underneath this. Efficiency is doing things right, effectiveness is doing the right things, and his sharpest version, the one worth taping to a wall, is that there is surely nothing quite so useless as doing with great efficiency what should not be done at all. Most reporting answers the first question and lets you believe it answered the second. We made a lot of things. Whether any of them mattered is the separate question, and it is usually the one nobody asks.
The drift toward inputs stops being a mystery once you look at what is easy and what is safe. An input is something you control. You can decide how many features get built this quarter and how many hours go into the new platform, and you can report every one of them with confidence, because they sit entirely inside your reach. An outcome behaves differently. It depends on other divisions you do not direct, and on customers who are free to respond however they like. Owning an outcome means accepting that part of your number lives outside your hands, which feels a lot like volunteering for blame. So people report what they own, and the dashboard fills with activity.
One thing normally keeps this honest, and it is the customer. When a company survives only if real people choose to buy, reality keeps intruding on the story. A quarter of busy work that produced no sales shows up as an empty bank account, and no amount of green on a slide argues with that. Take the customer out of the loop and the discipline leaves with them. A division funded by a grant, a team on a captive enterprise contract, an internal cost center, a startup running on capital patient enough not to ask hard questions for years, every one of them has lost the outside voice that could contradict the inputs. They look nothing alike and they fail the same way, because the feedback that would have told them the truth was never wired in.
Even where the customer signal does exist, a large enterprise tends to cut it off from the people who could act on it. The teams building the product usually have no clear line of sight into whether what they shipped changed anything a customer does, and almost none into whether it moved revenue. They release a change to an experience and then work in the dark. Customer behavior gets tracked less often than people assume, and rarely well enough to say which change caused which shift, so the gap fills with a proxy that happened to move in the same direction. A satisfaction score ticked up after the release, so the release must have done it. NPS is the favorite for this, cheap to collect and easy to mistake for the customer talking, while it tells you almost nothing about why anything changed.
From there a confident story travels upward. The product team asserts impact, senior management attaches the revenue numbers it can see from its altitude, and by the time the account reaches the C-suite it reads as a clean line from build to result that nobody in the chain could actually defend. This is the same break from earlier wearing different clothes. The path from a build to a customer behavior to a dollar of revenue runs from the product org into analytics and on into finance, and accountability stops at each border. No one owns the whole line, so no one connects it. The visibility gap a builder feels and the accountability gap an executive feels are the same gap seen from two ends of the building.
All of which sits underneath everything that follows. Honest outcome reporting assumes the data to report it exists and reaches the people doing the work. If the instrumentation does not connect the build to the behavior to the revenue and feed that back down to the team making the next decision, then asking them what changed for the customer is asking them to guess, and a guess delivered with confidence is how you end up with a deck full of impact that revenue never confirms.
There is a part leaders tend to miss about their own role in this. The reporting standard in a company lives in the first question a leader asks in the room. The goals printed on the wall have almost nothing to do with it. People report what they are asked about, and if the opening question in every review is how many features shipped or how many hours went in, that becomes the real measure no matter what the stated objectives say. Everyone learns to build their report around it. A leader who wants outcomes while asking about inputs will keep getting inputs, and the responsibility for that runs upward, to the person whose attention set the price.
None of this makes the answer a single clean outcome number. A single number is the easiest thing in the world to game. Tell a team you want a ten percent lift in customer satisfaction and reward whatever beats it, and you might get a thirty percent lift bought with overspending and giveaways, which is the letter of the goal and the reverse of its spirit. The outcome that counts is the result net of what it cost to reach. So you reach for more than one factor, and the opposite trap opens just as fast. Stack twenty measures together and the picture is confounded, because everything moves at once and you can no longer say which lever produced the result. The craft is a small set, a few factors whose combination holds the spirit and that cannot be moved one at a time without the others exposing it. Customer satisfaction sitting next to the margin it cost, with an eye on whether it survived the next quarter.
Keeping that set small does something most people underrate. It lets you attribute a win to a cause, and attribution is what lets you say no. When you cannot tell what produced a result, you cannot rule anything out, and a company where nothing can be ruled out ends up funding everything, because every line of work stays defensible. That is the road to an organization that is maximally busy and barely effective. Measurement sprawl hides what worked, and the ability to refuse what did not goes with it. A company that cannot say no will be productive forever and effective by accident.
Put together, an outcome stays honest under four conditions, and they have to hold at the same time. The leader has to actually ask about it, since attention is what sets the working standard. Reporting an ugly version of it has to be safe, or people will dress the number up the way they once dressed up the activity. Whoever is graded on it needs enough control to move it, or you are asking them to answer for a result they cannot reach. And it has to be defined by that small, attributable set of factors rather than one figure a clever team can bend. Miss any of the four and the faking finds the gap.
The engine that runs on top of all of this is the small experiment. Build one real slice and watch whether the outcome actually moved. Verified movement earns the next slice. A build you can measure that way is an experiment; without the measurement, the same build is only motion. It is also the only honest way to answer the question every organization eventually faces about its broken parts, which is what fixing this is actually worth. A working measure tells you whether a change is a half-step or a leap. Without one, you argue about it in a conference room until the opportunity has moved on.
James Wright and I gave the first chapter of our book to a pattern we named successful failure, organizations where something close to seventy percent report healthy metrics and deliver no real business result. It took us a while to see what those metrics were actually doing. They answered a question, cleanly and honestly. It was simply the wrong question. Everyone had been productive. Whether any of it mattered was something nobody had thought to check.
So here is the exercise worth running before your next review. Strip the activity out of the report and look hard at what is left. If what remains cannot tell you whether one customer did anything differently because of the work, then the report was never measuring effectiveness at all. It was measuring effort and calling it progress. The better first question, the one that reorganizes everything behind it, is plain. What actually changed for a customer, and could our measures even tell us if the honest answer were nothing?

This week's High Road Conversations features Professor Jeff Willie on being a river, not a reservoir, the idea that what you know is worth more moving through you to other people than stored up where it stops. He turned sixty and decided out loud to keep teaching until he is a hundred and ten, and the conversation holds up whether you run a classroom or a company. Listen here:
Listen →
I don't just talk about Agile and leadership, I put it into practice in my Skool, building and shipping real products with Claude every Thursday at 12:30pm EST.
Join us →This week's piece on the unowned AI mandate is the same problem James Wright and I open Agile Sucks! (When You Do It Wrong) with, in a chapter set in a combat zone where five civilian contractors cut hostile incidents from more than eight hundred a month to under fifty in a year by giving accountability a single owner at every level. Distributed authority worked there because someone's name sat on each result. You can find the book here:
Get the book →A short read every Saturday morning. No spam, unsubscribe anytime.
Subscribe