Skip to content
infoloop

The metrics that prove an AI copilot is working

01 JUNE 2026Operations5 min read

A demo proves nothing. Four numbers tell you whether an AI copilot is paying for itself, and a few popular ones tell you nothing at all. Here is how we measure the copilots we build.

A copilot is software that sits beside your team and does part of their work: answers a question, drafts the reply, fills in the form. Once it is live, somebody will ask whether it was worth the money. This is how to answer that in one page, with numbers you would be happy to defend.

Write down how the job runs today

A percentage with nothing to compare it against is decoration. Before the copilot touches a single piece of work, write down how the job runs now.

How many enquiries arrive in a normal week. How long one takes from arriving to being finished. How many a person gets through in a day. How often the answer has to be corrected afterwards. Take the figures over a stretch of ordinary trading, not your quietest fortnight and not the week the phones went mad.

This step is dull and it is the whole game. Every claim you make later is measured against this list. Teams who skip it end up arguing about whether things feel better, which is an argument nobody wins.

One: how much work it handles on its own

The share of the work the copilot finishes, or drafts well enough to send, without a person stepping in. This is the headline.

Two things keep the number honest. Only count work you actually pointed it at, so a copilot built for delivery questions is not marked down for the complaint that arrived on Thursday. And read the trend across weeks, because a single good day tells you nothing.

Then split it by type of request. It is perfectly normal for a copilot to be excellent at “where is my order” and poor at “I want to complain about the fitter who did my brakes”. That split is not a failure. It is your list of what to improve next, in order.

Two: how long that work now takes

Compare like with like against what you wrote down at the start: time to first reply, and time from arriving to finished, for the tasks the copilot covers.

If those are not falling, the copilot is adding steps rather than removing them. That happens more often than anyone selling one will admit, usually because somebody now has to read and approve every draft. A copilot that creates a checking job and saves nothing is a cost with a nice screen on the front.

Look at the slow cases too, not just the average. An average that improves while your worst cases get worse is how a business ends up with better reports and angrier customers.

Pick the one number that matters most to your business, tie the copilot to it before you build, and report it every month.

Three: the hours it hands back

Turn the first two numbers into hours a week. This is the one your finance team understands, because hours become either lower cost or the same team absorbing more work without hiring anybody.

Be honest about the sum. Take off the time people spend checking its drafts, correcting them, and keeping it fed with up-to-date information. The number that survives that subtraction is the real one, and it is the only one worth putting in front of a board.

On a support assistant we built for a fintech, the measure that mattered was how much manual ticket work was left for the team. It came down by 72%. That is the shape of number to aim for: one specific job, before and after, counted the same way both times.

Four: whether quality held

Speed with worse answers is not a saving. It is a complaint you receive later, from somebody who is now less patient.

Track how often a person had to correct the copilot, how often it passed a job to a human, and whether customers rate the experience any worse than before. Handing over is a good sign rather than a bad one, as long as it happens for the right reasons. A copilot that knows what it does not know is worth far more than a confident one that guesses.

The numbers that look impressive and mean nothing

Four you will be shown, and why to wave them away.

Total questions asked, or conversations held. That measures curiosity in week one, not value in month six.

A score on a test set. Exam conditions are not your Tuesday inbox, where people spell things wrong, attach photographs and ask two questions at once.

Hours saved, worked out by multiplying the number of tasks by a guessed minutes-per-task. That is an assumption wearing a number’s clothes.

How much your team likes the tool. Pleasant to hear, and evidence of nothing.

Reporting it so it survives a finance review

One page. Monthly. The same four numbers, worked out the same way every time, with your starting figures printed next to them so the comparison is visible without anybody having to remember.

Add a line of plain writing saying what changed and why. When a number dips, say so. A report that only ever goes up stops being read, and the honest version is the one that keeps the copilot funded.

What it costs to skip all this

You end up in a meeting defending a piece of software on instinct. Nobody can say whether it helped, so the argument turns into who feels strongly about it. Good tools get switched off that way, and bad ones survive for years. Writing down the starting figures before you begin is what prevents that conversation entirely.

In short

Write down where you are starting from. Then follow four numbers: how much work it handles alone, how long that work takes, the hours it hands back, and whether quality held. One page, once a month, with the starting figures printed alongside. That is how we run the copilots we build, and how we know whether they are earning their keep.

Frequently asked questions

  • How do you measure whether an AI copilot is working?

    Write down how the job runs today, then follow four numbers over weeks rather than days: how much work the copilot finishes on its own, how long that work now takes, the hours it hands back, and whether quality held. Report them on one page every month with the starting figures alongside.

  • How do you work out the hours an AI copilot really saves?

    Turn the share of work it handles alone and how much quicker that work now gets done into hours a week, then take off the time people spend checking its drafts, correcting them, and keeping it fed with up-to-date information. The number left after that subtraction is the only one worth putting in front of a board.

  • What AI copilot metrics look impressive but mean nothing?

    Wave away total questions asked, scores from a test set, hours saved worked out by multiplying tasks by guessed minutes per task, and how much your team likes the tool. None of them is evidence of value: they measure curiosity, exam conditions, guesses and goodwill rather than what happened on your real work.

Infoloop team

Operations

The people at Infoloop who build and run the software described here: engineers, designers and the team that stays after launch.

More

An interesting read? Here is more related to it.