Proving AI EBITDA · Metrics That Boards Actually Believe

Tracking AI EBITDA is all about choosing a metric

 

The question that kills the number

This quarter you brought the board a real result. Not a pilot count or a license utilization chart, but finished work: the AI your team placed into order entry two quarters ago is performing, the group is measurably faster, and finance has attached a figure to the improvement: two million dollars. For a moment it is the strongest number in the room.

Then your CFO asks her question. She is not hostile, and she has no position to defend; she would prefer that the number hold.

How do we know that was the AI?

The question is not an attack. It is her job. Over the same two quarters, order volume rose, three positions went unfilled, freight normalized, and a price increase took effect in March. Each of those moves a line worth two million dollars. So which portion of the two million belongs to the software, and by what method would anyone in the room separate it?

You have no answer, because the number was never built to produce one. The board does not conclude that you are wrong. It concludes that the result is unproven, a different condition from false that lands in the same place. The item moves to next quarter, and next quarter will be harder, because by then the before-picture will be four quarters old and reconstructed mostly from memory.

We wrote about the underlying shift in From AI Strategy to AI EBITDA: boards stopped asking what your AI strategy is and began asking where the earnings are. That article named the change; this one supplies the measurement.

What failed was not the AI, but the lack of a metric encapsulating the benefits.

The detail worth pausing on is this: in most of these meetings, the result is real. The work happened, the hours moved, the money is somewhere in the P&L. What failed was not the AI, but the lack of a metric encapsulating the benefits.

 

Why aggregate numbers fail, and what survives

An aggregate claim is a number without a denominator, and a number without a denominator cannot be checked. “We saved two million dollars” and “productivity is up thirty percent” are not falsehoods; they are unfalsifiable in a room where a dozen other variables moved over the same period. Your CFO is not refusing to believe you. She is declining to certify a number whose causes she cannot separate, which is the correct professional instinct and the reason you hired her.

The metrics that survive that room pass the same four tests, and the tests are worth knowing precisely, because they apply to any number anyone hands you, beginning with your own.

The first is that the metric is per-unit rather than aggregate. It carries a denominator that scales with the business: administrative hours per producing employee, cost per order processed, days to onboard a new hire or open a new location. The denominator is what keeps growth from masquerading as efficiency and contraction from masquerading as failure. Total administrative cost falls when you improve; it also falls when you shrink, and the aggregate cannot say which occurred. Cost per order can.

The second is that the metric is attributable. It sits close enough to the thing you changed that no plausible alternative explanation moves it as far as your intervention does. Gross margin has forty parents: pricing, mix, freight, rebates, one strong quarter in one region. Administrative hours per order processed has two or three, and you can name each of them. Distance from the intervention is where attribution goes to die.

Cleanliness is the third test, and the least expensive to pass. The metric is defined tightly, in writing, before launch: what counts as an order; whether a canceled order counts; whether the time a supervisor spends resolving exceptions sits inside the measure or outside it. Settled in advance, these are decisions that take an afternoon. Settled ninety days later, with a number already on the table, they become arguments, and everyone involved can see which definition favors whom.

The last test is the baseline: measured before the intervention rather than reconstructed after it. A baseline assembled from memory is a negotiation, not a measurement, and the room always knows. Reconstructed baselines carry a signature: they are round, they are generous, and no one can produce the query behind them.

Four tests, and the final one is the expensive one, because it is the only one that cannot be passed retroactively.

 

One number, carried all the way through

Consider an illustrative company: a regional distributor, roughly $120 million in revenue, processing about 120,000 orders a year at current volume, averaging a thousand dollars each. Inside sales and customer service handle order entry, order management, status inquiries, and exceptions. Two quarters ago the company introduced AI into that workflow.

Before launch it wrote the definition down, then spent six weeks measuring against it: administrative hours per order processed, counting the four activities above and excluding quote preparation, outbound selling, and collections. The baseline came in at 0.40 hours per order, twenty-four minutes. Across 120,000 orders, that is 48,000 hours a year, the equivalent of roughly twenty-three full-time positions.

Two quarters after launch, the same measure reads 0.28 hours per order, just under seventeen minutes. The per-unit delta is 0.12 hours, roughly seven minutes an order.

The arithmetic from here is deliberately plain. At a loaded labor cost of $42 an hour (wages, payroll taxes, benefits), 0.12 hours is $5.04 an order. Across 120,000 orders a year, that is $604,800, or 14,400 hours, about seven full-time equivalents at 2,080 hours each.

That is the per-unit value created, and it is not yet EBITDA; you should be the one to say so. Subtract the license and implementation cost, call it $180,000 in year one, and $424,800 remains. Then trace where the hours went, because hours become earnings only when they leave the cost base. At a distributor like this, three positions were retired rather than backfilled, roughly $262,000 of actual spend removed. The remainder, close to four positions’ worth, is cost avoided against a hiring plan that rising volume would have justified. Both are real, and only the first reduces spend against last year, so year one lands nearer $82,000 while the run rate points toward the larger figure. Show the board all three, labeled: value created, spend removed, spend avoided. A board that can see the entire derivation stops wondering what it is not being shown.

Credibility never came from the size of the improvement. It comes from the ability to measure and attribute.

Note that this is a smaller claim than the aggregate that opened this article, not a larger one. Credibility never came from the size of the improvement. It comes from the ability to measure and attribute.

Now consider the objection a skeptical CFO raises here, which concerns mix rather than software. Did the orders simply become easier: more electronic reorders, fewer manual ones, so that hours per order fell without anyone improving anything? It is the right question, and it is fatal to a per-unit metric that cannot answer it.

This one can, because the last two tests were passed months earlier: a written definition and a genuine before-picture allow the same measure to be split by segment on both sides of the launch. Electronic reorders went from 0.22 hours to 0.16. Manual and phone orders went from 0.67 to 0.46. Both segments fell, and by comparable proportions, while the mix between them held at sixty percent electronic in both windows. If mix were doing the work, the segments would have sat still while the blended number moved; that is not what happened. The mix explanation dies in a single slide, because of a decision made before launch rather than an analysis performed under pressure.

Two other candidates deserve a sentence each, because your CFO will raise them and you should raise them first. Volume rose over the period, and the denominator has already absorbed it. Thin staffing points the wrong way: when a team runs short, hours per order rise or service slips, and here accuracy and cycle time held.

 

The revenue side, where measurement runs out

The same deployment cut quote turnaround from three-to-five days to the same afternoon. Everyone involved believes that wins business. No one can measure it, because a won order has too many parents: the representative, the relationship, the price, the competitor’s stumble, the customer’s timing.

There are two undisciplined responses, and you have watched both. The first asserts a point estimate: same-day quoting produced $1.2 million in incremental revenue. That is precision without derivation, and a competent CFO finds the seam in under a minute. The second says nothing, on the theory that unmeasurable means unmentionable. That is the more respectable failure, and still a failure, because it leaves the largest portion of the value unreported.

There is a third response, older than any of this: break the number no one can measure into factors someone can. It is the physicist’s habit of estimating an unknowable quantity from knowable pieces, the one Fermi was known for.

For this distributor, the revenue claim decomposes into four factors. Quotes issued per year: 18,000, and the system knows it exactly. Current win rate: 26 percent, known with equal precision. Average gross margin on a won quote: about $1,320, from an average quoted order near $6,000 at roughly 22 percent, since quoted business runs larger than the reorder book; finance can confirm that one. Then the fourth factor, the plausible improvement in win rate from answering in hours rather than days, which no one knows.

Three facts and one judgment. So you state the judgment as a judgment: one to three points of win-rate lift. Applied to 18,000 quotes, that is 180 to 540 additional wins a year, and at $1,320 each, roughly $240,000 to $710,000 of additional gross margin. Gross margin, not EBITDA: incremental selling and processing costs come off it, and quote preparation was excluded from the cost metric deliberately, so nothing is counted twice.

You then present the range to the board as a range, labeled: three of these numbers come from our systems, the fourth is our estimate, and here is why it is wide.

What earns the trust comes next: a commitment to narrow it. One quarter later, 4,500 quotes have gone out under the new turnaround, and the win rate reads 27.4 percent, 1.4 points above baseline, holding in both regions with no price change underneath it. One quarter is one season, not proof by itself. But it is evidence, and evidence does specific work: it makes the top of the range less likely. The claim becomes one to two points, roughly $240,000 to $475,000. The floor does not move yet, and you say why: one quarter cannot rule out the low end, and the board should hear that from you.

Nothing exotic is happening. Updating a range as evidence arrives is all Bayesian reasoning has ever meant, and your CFO performs a version of it every time she revises a forecast. The difference is that you are doing it before the board deliberately, and announcing in advance that you will.

A range you can defend beats a point you cannot. A range that narrows beats both.

 

Evidence on a cadence

None of this works as an event. It works as a rhythm.

The per-unit metric appears on the same slide, in the same words, every quarter, whether the quarter was good or not. That last clause carries the entire discipline. A metric that appears only when it flatters you is marketing, and boards learn quickly to read it as such. Report the quarter in which hours per order ticked back up because you onboarded two large customers whose data arrived in poor shape, explain why, and one uncomfortable meeting will have purchased years of credibility.

We have written elsewhere about what the first ninety days of an AI deployment should produce, and that review rhythm is the one that carries these metrics. Ninety days is short enough that memory still works and long enough for a signal to separate itself from noise.

The compounding is the point. Four claims made in four quarters are four things a board must take on faith. The same per-unit metric reported four quarters running, under an unchanged definition, is a trend, and no one takes a trend on faith. The first quarter is a claim. The fourth is a record. Expect the disciplined number to come in smaller than the aggregate you would have claimed; it is also the only number that survives your CFO, which is what makes it worth more.

 

What this article cannot tell you

The four tests establish how to judge a metric. They do not establish which metric is yours.

That is the honest limit of any general writing on this subject, and it is better stated plainly. In a specific business, technology value hides in a dozen places, and which of them holds the most depends on things no framework knows: where the real constraint sits, what the operators quietly work around, which cost is structural and which is only habit. Two distributors of identical size and product mix will have different answers, because one is bleeding in quote turnaround and the other in returns processing, and neither could have guessed which without looking. Selecting the number is diagnosis and judgment. It is never the same twice, and it is not checklist work.

We have said elsewhere why finding that number is the craft, and we take it seriously enough to put our own fees on the outcome for the right engagements.

One step is available to you this week regardless. Take the single AI workflow already running in your company, the one you would speak about if the board asked today. Write down what it should improve, in per-unit terms, with a denominator that scales. Define it tightly enough that no one can argue about it ninety days from now. Then capture the baseline, this week, from your systems rather than from anyone’s recollection.

A baseline cannot be captured retroactively; it can only be negotiated. Every quarter that workflow runs unmeasured converts a measurement you could have had into a memory you will have to defend.

The result you report a year from now will be exactly as believable as the before-picture you take this week.

The result you report a year from now will be exactly as believable as the before-picture you take this week.

More from our blog

Proving AI EBITDA · Metrics That Boards Actually Believe

Proving AI EBITDA · Metrics That Boards Actually Believe

The question that kills the number This quarter you brought the board a real result. Not a pilot count or…
NAIC Just Wrote Your AI Governance Job Description · Solving Without a Full-Time Hire

NAIC Just Wrote Your AI Governance Job Description · Solving Without a Full-Time Hire

Market conduct exam letters do not announce themes. They enumerate. Somewhere between the request for complaint-handling logs and the request…
Team "Still Not Ready for AI"? · Readiness Is Designed, Not Awaited

Team "Still Not Ready for AI"? · Readiness Is Designed, Not Awaited

  The Meeting Where the Question Settled Itself The scene has repeated itself in mid-market boardrooms with remarkable consistency over…