The AI Bill
Buys Three Minutes
A firm can buy its people about three minutes a week. Erik Brynjolfsson thinks the national accounts are quiet because the reorganisation has not been counted yet. Daron Acemoglu thinks they are quiet because the minutes were never going to move them. The bill is not in dispute. Where it goes is.
On 2 May 2018, Elon Musk sat on Tesla's quarterly earnings call and described a robot with one job. It stuck a mat of fibreglass fluff to the top of a battery pack, to quiet the cabin. He called it a flufferbot. It could not pick the fluff up. When it managed to, it put the mat down in a random place, and the production line stopped.
He asked whether the car needed the mat. They built one with it and one without. The noise in the cabin did not change. The part was unnecessary.
The machine built to attach it had been the thing breaking the factory.
Three weeks earlier he had written, in public, that excessive automation at Tesla was a mistake. To be precise, his mistake. Humans, he said, were underrated.
That is a factory in Fremont. The argument in front of you is an office argument, and it has the same shape. A company pays a few dollars a month. A person finishes a task a little faster. The minutes come back. Then they have to get past everything the firm did not think to delete: the meeting, the sign-off, the week spent waiting for someone else to say yes.
One reading says those minutes are the start of a reorganisation the statistics cannot see yet. The other says the minutes are real and the national accounts will barely move, because most of the economy is not the task that got faster. Both readings start from the same bill.
The question is where a gain goes between the person who received it and the profit and loss account. A late arrival and a small number are different forecasts, and they imply different years. A firm has already paid the bill. A finance ministry has already written a number into a forecast. What the piece owes you is not a mood. It is the observation, by the end of 2028, that would show which of those had been the wrong one.
The bill, on one ruler
In August 2026 the payments company Ramp counted what American businesses actually paid for artificial intelligence, the software that generates text, code, images or decisions, rather than what they told a survey. The measure is card charges and supplier bills. Ramp has put the wider transaction panel behind its spending reports at more than 70,000 businesses. A firm counts if it paid an AI vendor that month. Free tools and personal accounts are invisible, so use is wider than spend.
The median and the top tenth, on the published chart, are a next-day reading, not a sentence Ramp's economist typed. The figure he does state is the top hundredth. He warns that this last number is a small group and moves when late invoices arrive. All three are the firm's bill divided by headcount, not by the number of people who open the tool.
Median
Per person, per month
Top tenth
About fifty-four times the median
Top hundredth
One month
$7,205 is one employee's bill for one month at a firm in that top hundredth. At $12.50, the typical firm on this panel would take forty-eight years to spend the same amount on that person. Not a forecast. If the median bill never rose, that is how long the ordinary firm takes to match what the heavy spender did between two pay cycles.
Set the median against a wage. In June 2026 the US Bureau of Labor Statistics put the full cost of a full-time private-sector worker at $54 an hour. At forty hours a week, that worker costs $9,360 a month. Twelve dollars and fifty cents is about thirteen minutes of that month, which is about three minutes a week. A British full-time median salary that spring was £39,039, and that is gross pay, not the employer's full cost. Even against the smaller British number, the bill is loose change.
Three minutes is not a finding about productivity. It is the price of the ticket.
The people who sign the bill are barely in the tool. Between November 2025 and January 2026, Nicholas Bloom, Steven Davis, Ivan Yotzov and their co-authors asked nearly 6,000 senior executives what had happened at their own firms. Britain was in the sample. The Bank of England fielded the British questions. The executives averaged an hour and a half a week on the tools. About nine in ten reported no impact on employment or on productivity over three years.
The ticket
3 min
A week of a full-time American wage, at $54 an hour, against a $12.50 bill.
The signature
90 min
A week in the tool, for the senior executives who were asked.
On 8 May 2025, at Klarna's headquarters in Stockholm, Sebastian Siemiatkowski told Bloomberg that the cost-cutting in customer service had gone too far. The firm had spent the previous year saying a chatbot was doing the work of 700 agents. He wanted a customer to be able to reach a person. "As cost unfortunately seems to have been a too predominant evaluation factor," he said, "what you end up having is lower quality." He was not retiring the software. He was putting a human back on a line the software had been left to finish.
What both sides can already see
In Britain, by June 2026, about 35 per cent of firms with ten or more staff told the Office for National Statistics they were using at least one of these tools, up from about 12 per cent in late 2023. Among larger firms the most common purpose was the operation they already ran. More than 60 per cent said so. A coat of paint on the existing week, not a new factory. A survey of senior managers finds a higher share of British firms using the tools. The questions are not the same, so the shares are not the same number.
The national number is quiet. American output per hour rose 2.1 per cent in 2025, and the current cycle was still running at 2.1 per cent through mid-2026, which is also the average since 1947.
On 12 July 1987, in the New York Times Book Review, Robert Solow wrote the sentence that has been doing unpaid work ever since. "You can see the computer age everywhere but in the productivity statistics."
12 July 1987
Solow
The computer age is everywhere but in the productivity statistics.
Late 1973 to late 1995
1.5% a year
American labour productivity, on the New York Fed's reading.
Late 1995 to mid-2004
3.1% a year
The same series, after the work was rebuilt around the machines.
2025
2.1%
The current American cycle. The long-run average since 1947, not the late 1990s.
End of 2028
The exit
If the gain is still only in the top tenth, the lag has not turned.
+0.7%
Tax records
−0.2%
Household survey
The Office for Budget Responsibility, in November 2025, assumes a milder version and does not date the acceleration. On their central path these tools add about 0.2 percentage points to annual productivity growth by around 2030. Over a decade, that is around 2½ per cent on the level. They think the slower shape is the likely one, and they call the assumptions behind a scenario of about 0.8 percentage points a year fairly unlikely.
But a quiet number has two honest readings, and they are not the same decade.
The strongest version of the argument, as Erik Brynjolfsson would want it put
A general purpose technology is one that shows up in every industry, the way electricity did, and then computers. Brynjolfsson, with Daniel Rock and Chad Syverson, argued in the American Economic Journal in 2021 that these technologies do not pay when you buy them. They pay when you rebuild the firm around them. While you are spending on things the statistician cannot see, measured productivity looks worse than the underlying truth. Later, when that spending has become a way of working, the same statistic overstates the gain.
They ran it backwards on computers. Once the intangible investment tied to hardware and software is given a value, the level of American total factor productivity, the part of growth that is not just more staff or more machines, sat 15.9 per cent above the official measure by the end of 2017. A level, not a yearly rate. The official series had been missing a storey of the building.
On this reading, Solow was early, not wrong. The Office for Budget Responsibility agrees about the shape, if not the height.
It is not enough to take a person out and plug the model in where the person was. You have to rethink the organisation.
He said that on 26 March 2025, to Bill Kerr, on Harvard Business School's Managing the Future of Work. While the spending goes in and nothing comes out, the accounts look worse. He thinks the economy is near that trough, and that this J will be shorter than electricity's. On the same tape he said electricity took thirty to forty years: the first factories put the new motors where the steam engine had been, and only a later generation, redesigning the floor around the work, doubled productivity. He cited the American print for the last quarter of 2024, about 1.2 per cent, and said it would tick up. In 2025 it did, to 2.1 per cent. That is the average since 1947, not the pace of the late 1990s. On 17 March 2026, at the Hoover Institution, he told Steven Davis he was betting productivity would come in substantially higher than the official forecasts, because he could see the gains on the ground and thought they would spread. The bet is his. The print he is waiting for is not in yet.
The cleanest evidence that the reorganisation is the bottleneck, rather than the tool, is a 2026 experiment by Hyunjin Kim, Dahyeon Kim and Rembrand Koning. They took 515 high-growth startups. The treated half were shown how other firms had rebuilt production around the technology. Treated firms found 44 per cent more uses, concentrated in product development and strategy. They completed 12 per cent more tasks. They were 18 per cent more likely to land a paying customer. They generated 1.9 times the revenue of the control group, asked for 39.5 per cent less outside capital, just over $220,000 less in the authors' seminar account, and did not hire to get there. The tools were held still. The information about the work was what moved.
What this case concedes is the map's edge. These are young firms in one programme, not a hospital with a century of meetings on the calendar. Nothing in the J-curve paper names the year the curve turns. The British path this side is glad to cite, for its shape, is still measured in tenths of a point. A lag is a reason to wait. It is not a receipt. The profit and loss moved. For whom is the next question.
The strongest version of the argument, as Daron Acemoglu would want it put
Acemoglu's reply is not that the studies are fake. He uses them. A saving on a task becomes a saving for the economy only after you multiply by the share of the economy that task actually is.
It halves the memo, and memos are not the economy.
In a paper for the National Bureau of Economic Research in May 2024, he puts numbers on that seatbelt. If the remaining tasks are the hard ones, the ones without a clean right answer, he puts the decade under 0.53 per cent.
Saving on the task
From the early studies
Share reached in ten years
Not every task the tools can touch
0.66%
Over the decade
The study he feeds in is not a hostile one. Brynjolfsson, Danielle Li and Lindsey Raymond found that a generative assistant raised cases resolved per hour in customer service by about 14 per cent. Acemoglu records that the gain sat with the less experienced staff, and treats the result as real: the sort of easy task the first studies were always going to find. Easy tasks have a scoreboard. The decisions a manager is paid for often do not.
The concrete case is Procter & Gamble, and it is a good case, which is why it cannot be waved into a national forecast. Dell'Acqua and colleagues, in Organization Science in 2025, put 791 experienced people on real product proposals for their own business units. A person with the model matched a two-person team without it. Ethan Mollick, a co-author, describes the sessions as a single day. The paper adds what the headline leaves out: the model lifted the ideas, and human judgement still did the choosing. One day. Ideas. A judge at the end. Not a case of shampoo sold.
The startup average belongs here too, read the other way up. The authors attach a sentence to the 1.9 times: the revenue and investment gains are largest at the 90th percentile and above. A higher ceiling, not a modest lift for the firm in the middle. The firms were selected and the horizon was short. That is not a reason to ignore 1.9 times. It is a reason not to paste it onto the national statistics.
So does the executive survey, read as a want rather than a measurement. The same people who report no effect over three years predict that productivity at their own firm will rise 1.4 per cent over the next three. They also expect employment at the firm to fall by 0.7 per cent. Employees, asked separately, expect employment at their firms to rise. A gain of 1.4 per cent over three years is under half a point a year, filed by people who report no effect so far and who spend an hour and a half a week on the tools. It is not a measurement of a gain that has arrived.
The concession is expensive, so it counts. The task gains are not a mirage. Fourteen per cent in a call centre, and a matched team at Procter & Gamble, happened. The Office for Budget Responsibility built its number with the same kind of task-by-task sum, and still published a larger one than his. The rulers are not the same: his 0.66 per cent is total factor productivity, and their 2½ per cent is labour productivity. They said their figure was higher than his. They did not say it was the same sum. On that ruler the next decade is a few extra tenths, not a new 1995.
The arguments in circulation that neither side should be making
The two cases above can both be held by a careful person. A lot of what travels between them cannot.
| The claim | What it is actually measuring |
|---|---|
| The bill is so small that the return is automatic. Three minutes, and you have paid for the year. | The price of the software. Not whether the minutes survive the next meeting, and not whether the task was worth doing. |
| Generative AI could add $2.6 trillion to $4.4 trillion a year. McKinsey Global Institute, June 2023. The top of the range was larger than the British economy in 2021. | A sum of potential, if the use cases are adopted, with no clock and no subtraction for the work that should have been deleted. |
| MIT showed that 95 per cent of AI pilots fail. | A preliminary Project NANDA paper, July 2025. The executive summary says 95 per cent of organisations are getting zero return. The chart underneath is a narrower funnel: 60 per cent looked, 20 per cent piloted, 5 per cent reached production. Several fractions were allowed to become one sentence, and the sentence was allowed to become MIT. |
| The productivity statistics never moved for computers, so they will not move for this. | Solow's sentence in 1987, and a refusal to turn the page. The same American series ran at about 1.5 per cent a year and then at about 3.1 per cent. The analogy includes a second act people leave off. |
Two claims promise a transformation the evidence has not scheduled. Two promise a failure it has not delivered.
Ramp publishes a method: paid transactions, anonymised, blind to free tools, volatile at the very top. The statistical offices publish methods because they are public bodies. The academic papers name authors and venues. McKinsey sells the transformation it sizes. Project NANDA builds agent infrastructure and labelled its paper preliminary. That does not make a number true or false. It tells you which ones you can recompute.
The country forecast, then the firm
The economists are answering a question about the country. The bill asks a question about the firm. The path that keeps both papers is the Office for Budget Responsibility's: around 2½ per cent on the level of output per hour over a decade, about 0.2 percentage points a year by around 2030. Brynjolfsson's J says this could be early. Acemoglu's ceiling says it could be high. Neither is a date, and neither is a doubling.
Where the gain disappears is in the work that was not worth doing. The three minutes come back. They die in the meeting the fluff would have recognised, the same way Klarna's cheaper support came out worse. The firms that pull away delete the mat before they buy the robot. That is why the same tools produce a 1.9 times average and a quiet national statistic.
The step that fails a noise test does not get faster. It goes.
- 01Make the requirement less dumb
- 02Delete the part or the process
- 03Only then simplify
- 04Only then go faster
- 05Automate last
On 15 June 2026, Aaron Levie, who runs Box, told Michael Krigsman why the agents that write code moved and the rest of the office did not. Code is text. You can test whether it still works. The person using it already knows how to steer it.
If he had not reviewed a report, and the agent had pulled the wrong data, he would have reached exactly the wrong conclusion.
A contract, he said, cannot be computed as correct. It has to go through the other party's red lines. A chief executive is the furthest person from the real work, which is why it is so easy to decide that an engineer, or a marketing campaign, can be automated. He calls that first rush an AI psychosis. The other side of it is the discovery that a person still has to read the step. A 2026 paper by Demirer, Horton and co-authors models the same thing as one check at the end of a run, not a check after every sentence. The paper is a model. Levie's unread report is the case. The Procter & Gamble judges still chose. A draft is not a payment, and it is not a diagnosis.
A cheaper model is enough when nobody has to stand behind the step. By September 2026 the price of a million tokens, the units the model companies bill, had fallen 41 per cent from the March peak, to 68 cents, and firms were already pushing people off the expensive models.
The exit is a measurement. If, by the end of 2028, a study of ordinary firms finds the gain at the median and not only above the 90th percentile, and the Office for National Statistics shows output per hour in the industries using these tools growing faster than the rest of the economy for four quarters on a single method, the ceiling was too low. If the new evidence still looks like Kim, Kim and Koning, a higher ceiling and a flat middle, and British output per hour is still inside the post-2019 range the two official methods already disagree about, the J has not turned. The test is the middle of the distribution. The clock is the end of 2028.
You do not need the robot to know the mat was useless. You need the two cars, and someone willing to sit in both.
The three minutes are that test, run on a wage. They fit on a receipt. They also fit inside a meeting that was already in the diary, and they do not survive it unless someone asks of the meeting what was asked of the mat.
In Stockholm, in May 2025, a chief executive asked it of a customer calling about money. The cheaper line was worse. He wanted a person back on the phone.
The bill has already been paid. The meeting has not been cancelled.
Ramp, Ara Kharazian, September 2026 AI Index, and the August index letter. Accounts Recovery, 10 September 2026, for the median and top-tenth chart readings. US Bureau of Labor Statistics, Employer Costs for Employee Compensation, June 2026, and Productivity and Costs for 2025. Office for National Statistics, earnings 2025, AI in UK businesses (20 July 2026), and the productivity flash of 18 August 2026. Robert Solow, New York Times Book Review, 12 July 1987. Jorgenson, Ho and Stiroh, Federal Reserve Bank of New York, 2004. Office for Budget Responsibility, Briefing paper No. 9, November 2025. Brynjolfsson, Rock and Syverson, American Economic Journal: Macroeconomics, 2021. Kim, Kim and Koning, INSEAD working paper 2026/20/STR. Dell'Acqua and co-authors, Organization Science, 2025. Acemoglu, NBER working paper 32487. Yotzov, Barrero, Bloom, Davis and co-authors, NBER working paper 34836. Brynjolfsson, Managing the Future of Work, 26 March 2025, and the Hoover panel of 17 March 2026. Aaron Levie, CXOTalk, 15 June 2026. Tesla earnings call, 2 May 2018. Elon Musk with Tim Dodd, summer 2021. Sebastian Siemiatkowski, Bloomberg, 8 May 2025. McKinsey Global Institute, June 2023. Project NANDA, July 2025.