The AI Value Gap Is a Traceability Problem
Three numbers from the same survey, in the same year, describing the same organizations.
Eighty percent of people who use AI in their work say it improved their individual productivity. Thirty-seven percent of their organizations can attribute any earnings impact to AI at all. Six percent qualify as high performers — meaning they attribute at least five percent of EBIT to AI and describe the impact as significant.
Those figures come from McKinsey's The state of AI in 2026: On the road to ROI, published 25 August 2026, fielded 4 May to 8 June with 1,719 respondents across 97 nations, weighted by each nation's contribution to global GDP. The 37 percent is essentially unchanged from 2025. So is the six percent.
One clarification before going further, because the middle figure is routinely misread: 37 percent does not mean AI raised EBIT by 37 percent. It means 37 percent of respondents attribute at least some EBIT impact to AI. This is a respondent-reported survey, not a controlled causal study. That makes the numbers evidence of what leaders can currently see and attribute — which, for an argument about traceability, is precisely the relevant measurement.
The interesting part is not any single number. It is the shape of the collapse: 80 to 37 to 6. Something real is happening at the level of the individual desk, and almost none of it is arriving at the income statement. A fourth figure from the same survey sharpens the point — thirty-two percent of organizations declined to purchase a software product last year because they built the capability in-house with agentic coding tools. Build capacity is demonstrably not the constraint.
I want to argue that this is not an AI problem. It is a traceability problem, and it belongs to a discipline that already exists.
The Gap Is Not a Capability Gap
The instinctive reading of 80 versus 6 is that the productivity gains are illusory — that people feel faster without being faster, and the earnings data is the honest measurement catching up with the enthusiasm.
That reading is available, and it is probably true in places. But it cannot be the whole story, because the same survey shows organizations making hard commercial decisions on the strength of those gains. You do not cancel a software purchase because your engineers feel productive. Thirty-two percent of respondents did exactly that.
The more defensible reading is that the productivity is real and local. It accrues to a task. It shows up as an engineer shipping a feature in two days instead of five, an analyst reading forty documents instead of eight, a proposal manager producing a first draft before lunch. Each of those is genuine. None of them is automatically a financial outcome.
Between a faster task and a changed P&L sits a chain of questions that no model answers: was that task on the critical path of anything? Did the time saved get reallocated, or absorbed? Did the throughput increase reach a constraint that was actually binding? Was anyone accountable for converting the slack into an outcome?
Those are not AI questions. They are line-of-sight questions.
What Separates the Six Percent
If the constraint were model capability, the high performers would be the organizations using the most AI. They are not.
The differentiator McKinsey reports is structural. Nearly three-quarters of high performers say they have fundamentally redesigned workflows around AI. Among everyone else, roughly one-quarter. The same group is more likely to have defined processes for measuring AI's impact, to pursue growth and innovation rather than efficiency alone, to show senior-leader commitment, and to actively manage AI-related risk.
Read that list again and notice what is absent from it. Not better models. Not more tooling. Not more pilots. Every item is a statement about how work and accountability are organized.
The inverse is the more useful finding: the majority are inserting AI into workflows that were designed around human constraints — sequencing, batching, handoff points, review gates that exist because a person needed to sleep. Automating a step inside a structure built for different physics makes that step faster and the structure no cheaper. The gains stay local. That is the 80 percent, stranded.
Which means the gap has an empirical shape, not just a theoretical one, and it is a gap in how work is structured rather than in what the technology can do.
Where the Line of Sight Breaks
A strategic AI plan states intent. It names ambitions — improve customer responsiveness, reduce cycle time, strengthen risk posture — and usually attaches a budget and a set of platform decisions.
Value, however, is not produced by intent. It is produced by work. And work has structure well below the level at which most AI strategies operate:
Functions contain workflows. Workflows contain activities. Activities decompose into tasks. Tasks contain decisions — the points where judgment is exercised — and they are connected by handoffs, where work and accountability change hands.
Most AI plans name the intent at the top and the tooling at the bottom, and skip everything in between. The result is a document that is simultaneously true and unusable: nothing in it is wrong, and nothing in it tells you which of the four thousand things your organization does each week should be touched first.
That missing middle is where the 80 percent goes to die. Individual productivity is a task-level phenomenon. Enterprise value is a function-level phenomenon. If nothing connects the two, the gains are real at one end and invisible at the other, and no amount of additional model capability closes the distance.
Decomposition Reveals. Prioritization Decides.
Decompose a function honestly and the first thing you discover is abundance. A mid-sized accounts payable operation will surface somewhere between eighty and two hundred discrete tasks. A capture and proposal function, more. Each is a candidate for some degree of AI involvement.
This is where a lot of transformation programmes go wrong, because abundance reads as opportunity. It is not. It is a queue, and queues need an ordering principle.
The candidates are not equally valuable, not equally feasible, not equally ready, and not equally appropriate. Some sit on data that does not exist in retrievable form. Some carry regulatory exposure that makes autonomy a bad trade at any accuracy. Some would work beautifully and save four hours a month. Some depend on three other things being true first.
So decomposition has to be followed by scoring, against factors that are explicit and applied uniformly:
Strategic alignment · business or mission value · feasibility · data readiness · risk · implementation complexity · human impact · dependencies · time-to-value.
The factors matter less than the fact that they are written down before the scoring starts. A prioritization that cannot be reconstructed is a preference with a spreadsheet attached — and it will not survive contact with the first executive who disagrees with the ranking.
Decomposition reveals where AI could be applied. Prioritization determines where it should be applied first. Those are different activities and they fail differently. Skipping the first produces a plan with no purchase on the actual work. Skipping the second produces a hundred pilots and no portfolio.
A Worked Example
Abstract arguments about traceability are cheap. Here is the concrete version, at the smallest scale that still shows the mechanism.
Strategic objective: reduce working-capital volatility.
Function: Accounts Payable.
Workflow: invoice-to-pay exception handling.
Activities: intake exception → classify cause → gather supporting evidence → decide disposition → post adjustment → notify supplier.
Decomposing just the middle of that workflow produces three plausible AI candidates. Scored against the factors above, abbreviated to the five that separate them:
| Candidate task | Value | Data readiness | Risk | Time-to-value | Verdict |
|---|---|---|---|---|---|
| Classify exception cause | High — gates everything downstream | High — 3 years of labelled history in the ERP | Low — misclassification is caught at disposition | Weeks | First |
| Draft supplier correspondence | Moderate — saves minutes, not cycle time | Moderate | Moderate — goes to an external party | Days | Later |
| Approve payment release | High | High | Severe — irreversible, audited, segregation-of-duties control | n/a | Not a candidate |
Note that the selected task is not an overlay on the existing process — classifying the cause earlier changes what the downstream activities receive, which is what workflow redesign means in practice at this scale.
Note also what the third row does. It is the highest-value candidate in the set and it is excluded — not because the technology would fail, but because the control exists for a reason that predates AI. A prioritization method that cannot produce a confident "no" is not a method.
Now the part that matters. The selected task carries a trace in both directions:
Upward: classify exception cause → gather evidence activity → invoice-to-pay exception workflow → Accounts Payable → reduce working-capital volatility. An engineer maintaining that classifier can see, in one line, which strategic objective their work serves and who owns the outcome.
Downward: reduce working-capital volatility → Accounts Payable → exception handling → classify cause, with two candidates deliberately deferred and one deliberately refused, each with a recorded reason. A CFO asking "what are we doing about working capital, and why that and not something else" gets an answer with evidence rather than an anecdote.
The measurement follows from the trace. Because the objective is named, the metric is not invented afterwards: exception ageing and days-payable variance were already the measures of that objective. The AI initiative inherits them instead of manufacturing a proxy like "hours saved" that nobody can convert into money.
Traceability Has to Run Both Ways
One direction alone is not traceability.
Strategy that traces only downward is a cascade — a reporting structure that pushes targets to the edge and measures compliance. Implementation that traces only upward is a justification exercise, assembled after the fact to explain work already underway.
What makes the trace load-bearing is that both ends can be walked by different people for different reasons. Leadership reads down: here is the intent, here is the prioritized execution, here is what we chose not to do. Operators and engineers read up: here is the thing I am building, here is the objective it serves, here is who is accountable if it does not.
When the trace breaks — and it breaks constantly, through reorganizations, pivots, and the ordinary churn of work — the break is itself the finding. An initiative that no longer traces to a live objective is not necessarily wrong, but it now requires a decision rather than a budget line.
This Is a Systems-Engineering Problem
Every element of the above is borrowed. Requirements decomposition, allocation to components, verification against the originating requirement, and a maintained bidirectional trace matrix are not new ideas. They are ordinary practice in domains where the cost of an untraceable requirement is measured in hulls or airframes.
The discipline exists. It has simply not been pointed at AI adoption, because AI adoption arrived through a different door — as a technology procurement and a capability question, rather than as a systems problem about how work is structured and where judgment lives.
That framing is what I have been exploring through ARKONA Research, and it is the thesis underneath COMET: decompose the work to the task level, classify where human judgment must remain, score what should move and in what order, and keep the trace from strategic intent through to a measured outcome — auditable in both directions.
It is a less exciting proposition than autonomous everything. It has the advantage of being checkable.
Key Takeaway
The 80-to-6 collapse is the most important measurement in enterprise AI right now, and it is widely misread as evidence that the technology underdelivers. The technology is delivering — to individuals, at the task level, exactly where the survey says it is.
What is missing is the connective structure that turns a faster task into a financial outcome, and the accountability that makes someone responsible for the conversion. The organizations that found it did not find it by adopting harder — they found it by redesigning the work, three-quarters of them against one-quarter of everyone else. That structure is buildable. It is also unglamorous, and it looks much more like requirements engineering than like AI.
Deploying AI is not the finish line. Measurable value is — and you cannot measure what you cannot trace.
Source. McKinsey & Company / QuantumBlack, The state of AI in 2026: On the road to ROI, 25 August 2026. Online survey fielded 4 May – 8 June 2026; 1,719 respondents across 97 nations, weighted by each nation's contribution to global GDP. Figures are respondent-reported. The report does not evaluate or endorse ARKONA Research, COMET, or any methodology discussed here; the connections drawn above are mine. Report page.