The Control Group Is Gone · How to Prove AI Worked When Nobody Will Work Without It

A long, dark industrial corridor of identical doorways receding into the distance, with the two nearest doorways on the left glowing amber while the matching doorways on the right and farther down stay dark, representing a staggered rollout where teams not yet live serve as the comparison group.

Your operations lead says the AI assistant has been a hit. Ask the team and you get the same answer in different words: they would not go back. Then your CFO, who has no position to defend and would like the number to hold, asks the question that matters. Compared to what? Nobody in the room can answer, because everyone has been using the tool for months and nobody is left doing the work the old way.

That is not a failure of your measurement. It is a new condition, and the people best equipped to measure AI’s effect on work have just run into it in public.

 

The best experiment in the field lost its control group

In July 2025, the research nonprofit METR published a randomized controlled trial of AI coding tools. Sixteen experienced open-source developers completed 246 real issues, each randomly assigned to be done with or without AI assistance. The tasks averaged about two hours, in repositories the developers had worked on for years. The result got attention for its direction: with AI allowed, developers took 19 percent longer to finish. Before the study they had forecast a 24 percent speedup, and after it they still believed AI had sped them up by 20 percent.

Seven months later, in February 2026, METR published a follow-up with late-2025 tools: 57 developers, more than 800 tasks, 143 repositories. The new estimates pointed toward a speedup, but the confidence intervals were wide enough to include no effect at all. And the update said something more interesting than the numbers. METR reported that 30 to 50 percent of developers told them they were choosing not to submit some tasks because they did not want to do them without AI. An increased share said they would not want to do half of their work without it, even at $50 an hour. Their own conclusion: because of those selection effects, the data is “only very weak evidence” for the size of the speedup. They are changing the experiment design rather than publishing a number they do not trust.

The durable lesson is about measurement, not magnitude.

Notice what happened. The experiment did not fail because the tool stopped working. It failed because the tool became too useful to give up, and a comparison needs people willing to be on the other side of it. METR’s own post says its early-2025 result no longer reflects current tools, so do not carry the 19 percent figure around as a fact about AI today. The durable lesson is about measurement, not magnitude.

 

Why this lands on your desk

You have the same problem with none of the research budget. The moment an AI tool becomes part of how a team works, three things quietly stop being measurable.

The first is perception. The developers in the 2025 study guessed wrong about their own speed, and in the direction that flattered the tool. This is not dishonesty. Work that feels smoother feels faster, and a person is a poor stopwatch for their own afternoon. The second is selection. When people decide which tasks go through the assistant, they pick the ones it handles well, and the results look better than the work as a whole. The third is where the gain goes. We wrote in Your AI Dividend Is Going to the Wrong Bank Account about how the hours AI frees tend to be captured by the person who saved them, which means the person closest to the tool is also the one with the least reason to report what it did to the cost line.

Executive surveys inherit the same limit. PwC’s 29th Global CEO Survey, fielded from September 30 to November 10, 2025 across 4,454 CEOs, found that 56 percent reported neither revenue nor cost benefits from AI in the prior twelve months. PwC’s mid-year snapshot of 351 of those CEOs, fielded May 15 to June 22, 2026, found that 39 percent now report positive AI outcomes. Both are useful. Both are also executives describing their own companies, which is the same kind of evidence as a team saying it would not go back. That is a signal, and it is not proof.

 

Build the comparison group on purpose

If the control group will not appear on its own, you create one, and the cheapest way is sequence. Do not turn the tool on everywhere at once. Pick teams, sites, branches, or queues that do broadly similar work, and bring them online in order. For the length of each stage, the teams that have not started yet are your comparison group, doing the same work in the same quarter under the same market conditions, which is exactly what a before-and-after comparison cannot give you.

One rule makes or breaks this. You assign the order, and you assign it before anyone volunteers. The team that begs to go first is the team with the most enthusiasm, the most capable lead, and the best chance of improving no matter what you hand them. If the eager team goes first, you will measure eagerness. Draw the sequence from a calendar, a coin, or a list sorted by something unrelated to the outcome, write it down, and hold to it.

This has a cost, and it is worth saying plainly. Some teams wait a quarter for a tool they would have liked sooner. Against that, you get an answer your CFO can challenge and watch survive, which is worth more than a faster rollout that ends in a number nobody certifies. It also fits naturally inside a sequenced, governed install like Project Vista, where the order of deployment is already a decision someone is making anyway.

 

Count the work at the door

Measure only the work that went through the tool and you have measured the tool on its best day.

The METR developers withheld tasks they did not want to do without AI. Your people will do the equivalent without meaning to. The support agent runs the easy tickets through the assistant and handles the thorny ones by hand. The analyst drafts with it when the data is clean and skips it when the data is a mess. Measure only the work that went through the tool and you have measured the tool on its best day.

The fix is to define the unit of work at intake rather than at use. Every ticket that enters the queue, every invoice that arrives, every claim that opens, every order that is placed counts, whether or not anyone judged it a good fit for AI. The denominator is whatever came through the door. That one decision removes the largest source of flattering error, and it costs nothing but a tally taken one step earlier in the process.

The metric itself should pass the tests laid out in Proving AI EBITDA: per unit, attributable, defined in writing before launch, and baselined. A comparison group does not replace those. It is what lets them hold when the quarter has other things going on in it.

 

Ask the system and the customer, not the user

A second source, one that does not pass through the person using the tool, turns testimony into evidence. Cycle time stamped by the system of record. Rework and error returns. Complaints. Days to collect. Reopened tickets. The new-hire ramp, measured in weeks to full productivity. These move whether or not anyone feels faster, which is the point.

User sentiment is not worthless. Treat it as color, and treat disagreement as information. If the team reports that work is dramatically faster and the system timestamps are flat, you have found a perception gap, and it is more useful to know about on the first quarter than the fourth. If the two sources agree, you have the thing most AI claims lack: two independent measures telling the same story.

 

Write down what would change your mind

Before the first team goes live, put one sentence on paper: the result that would make you stop, slow down, or change the plan. A claim that cannot lose is not a claim. It is a hope with a budget.

We hold ourselves to the same standard in public. Our prediction ledger lists each call with its date, the date it resolves, the source that decides it, and a confidence level, and entries are never deleted. Misses count the same as hits in the denominator. That discipline is uncomfortable on purpose. It is also the only reason a reader has cause to believe the hits, and the same logic applies to the AI project you are about to fund.

 

What this will and will not tell you

A staggered rollout across a handful of teams is not a laboratory. The groups are small, the quarters are noisy, and a good result will come with error bars you should respect. What it gives you is narrower and more useful: a claim built so that someone skeptical can poke at it and find it still standing. That is the standard a board applies, and it is the one that survives a change of CFO.

It will also keep you honest about when to stop measuring. After a tool has been in place long enough that no team remembers the old way, the comparison is gone for good, and what you have is a trend line and a few reference points. Better to build the proof in the first two quarters, while the control group still exists, than to go looking for it in the fourth.

Of course they would not. Ask what the work looks like on the teams that have not started yet.

If you are about to roll out AI across several teams and want to design the sequence before the first one goes live, an independent assessment is a good place to start. Either way, stop asking your people whether they would go back. Of course they would not. Ask what the work looks like on the teams that have not started yet.

More from our blog

The Control Group Is Gone · How to Prove AI Worked When Nobody Will Work Without It

The Control Group Is Gone · How to Prove AI Worked When Nobody Will Work Without It

Your operations lead says the AI assistant has been a hit. Ask the team and you get the same answer…
The 9% Problem · What Your Board Actually Wants to See on AI

The 9% Problem · What Your Board Actually Wants to See on AI

Somewhere in the next board meeting, or the one after that, someone is going to ask you a version of…
Cheap Tools, Scarce Judgment · Why Outside Expertise Has Never Been More Valuable

Cheap Tools, Scarce Judgment · Why Outside Expertise Has Never Been More Valuable

AI tools get cheaper every quarter. The experience to use them well does not. Why outside expertise is worth the…