Skip to content
infoloop
infoloop

The metrics that prove an AI copilot is working

Updated AIOperations5 min read

Two panels compare AI copilot numbers. Ignore four that mean nothing: total questions asked, a score on a test set, hours saved from guesses, and how much your team likes it. Measure four numbers against your starting figures: work it handles alone, how long that work takes, the hours it hands back, and whether quality held.

A demo proves nothing. Four numbers tell you whether an AI copilot is paying for itself, and a few popular ones tell you nothing at all. Here is how we measure the copilots we deliver.

Four numbers prove an AI copilot is working: the share of work it finishes without a person stepping in, how long that work now takes, the hours it hands back after checking time, and whether quality held. Each one only counts when you measure it against how the job ran before the copilot went live.

Key takeaways

  • Write down how the job runs today before you switch anything on, or every number you quote afterwards means nothing.
  • Four numbers matter: how much work it handles alone, how long that work takes, the hours it hands back, and whether quality held.
  • Count only the work you actually pointed it at, and read the trend rather than one good day.
  • Take off the time people spend checking and correcting it before you claim any hours saved.
  • Ignore total questions asked, scores from a test set, and hours saved worked out by multiplying guesses.

A copilot is software that sits beside your team and does part of their work: it answers a question, drafts the reply or fills in the form. Once it is live, somebody will ask whether it was worth the money. This is how to answer that in one page, with numbers you would be happy to defend.

Write down how the job runs today

A percentage with nothing to compare it against is decoration. Before the copilot touches a single piece of work, write down how the job runs now.

  • How many inquiries arrive in a normal week.
  • How long one takes from arriving to being finished.
  • How many one person gets through in a day.
  • How often the answer has to be corrected afterwards.

Take the figures from a run of normal weeks, not your quietest two weeks and not the week the phones never stopped.

This step is dull and it is the whole game. Every claim you make later is measured against this list. Teams who skip it end up arguing about whether things feel better, which is an argument nobody wins. If the copilot is not built yet, write this list alongside the five decisions we make before an AI agent goes live.

One: how much work it handles on its own

The share of the work the copilot finishes, or drafts well enough to send, without a person stepping in. This is the headline.

Two things keep the number honest. Only count work you actually pointed it at, so a copilot built for delivery questions is not marked down for the complaint that arrived on Thursday. And read the trend across weeks, because a single good day tells you nothing.

Then split it by type of request. It is perfectly normal for a copilot to be excellent at “where is my order” and poor at “I want to complain about the mechanic who did my brakes”. That split is not a failure. It is your list of what to improve next, in order.

Two: how long that work now takes

Compare like with like against what you wrote down at the start: time to first reply, and time from arriving to finished, for the tasks the copilot covers.

If those are not falling, the copilot is adding steps rather than removing them. That happens more often than anyone selling one will admit, usually because somebody now has to read and approve every draft. A copilot that creates a checking job and saves nothing is a cost with a nice screen on the front.

Look at the slow cases too, not just the average. An average that improves while your worst cases get worse is how a business ends up with better reports and angrier customers.

Pick the one number that matters most to your business, tie the copilot to it before you build, and report it every month.

Three: the hours it hands back

Turn the first two numbers into hours a week. This is the one your finance team understands, because hours become either lower cost or the same team absorbing more work without hiring anybody.

Be honest about the sum. Take off the time people spend checking its drafts, correcting them, and keeping it fed with up-to-date information. The number that survives that subtraction is the real one, and it is the only one worth putting in front of a board.

On a support assistant we built for a fintech scale-up, the measure that mattered was manual handling of the ticket categories it was built for. Within a quarter it fell 72%, and first response time went from hours to under two minutes. That is the shape of number to aim for: one specific job, before and after, counted the same way both times.

Four: whether quality held

Speed with worse answers is not a saving. It is a complaint you receive later, from somebody who is now less patient.

Track how often a person had to correct the copilot, how often it passed a job to a human, and whether customers rate the experience any worse than before. Handing over is a good sign rather than a bad one, as long as it happens for the right reasons. A copilot that knows what it does not know is worth far more than a confident one that guesses. If it serves customers in regulated markets, quality also means evidence that it is controlled, and our AI governance framework checklist covers what to log.

The numbers that look impressive and mean nothing

Four you will be shown, and why to wave them away.

Total questions asked, or conversations held. That measures curiosity in week one, not value in month six.

A score on a test set. Exam conditions are not your Tuesday inbox, where people spell things wrong, attach photographs and ask two questions at once.

Hours saved, worked out by multiplying the number of tasks by a guessed minutes-per-task. That is an assumption wearing a number’s clothes.

How much your team likes the tool. Pleasant to hear, and evidence of nothing.

Reporting it so it survives a finance review

One page. Monthly. The same four numbers, worked out the same way every time, with your starting figures printed next to them so the comparison is visible without anybody having to remember.

Add a line of plain writing saying what changed and why. When a number dips, say so. A report that only ever goes up stops being read, and the honest version is the one that keeps the copilot funded.

What it costs to skip all this

You end up in a meeting defending a piece of software on instinct. Nobody can say whether it helped, so the argument turns into who feels strongly about it. Good tools get switched off that way, and bad ones survive for years. Writing down the starting figures before you begin is what prevents that conversation entirely.

In short

Write down where you are starting from. Then follow four numbers: how much work it handles alone, how long that work takes, the hours it hands back, and whether quality held. One page, once a month, with the starting figures printed alongside. That is how we measure every copilot we deliver, and how you can tell whether yours is earning its keep.

Book an AI copilot discovery call if you are planning a copilot, or already run one and cannot fill in that page yet. Bring the job you want it to take on, and in 30 minutes we will talk through the starting figures and the four numbers with you. You can also see how we build AI assistants and agents that work from your own business data.

Frequently asked questions

  • How do you measure whether an AI copilot is working?

    Write down how the job runs today, then follow four numbers over weeks rather than days. The numbers are how much work it finishes alone, how long that work takes, the hours it hands back and whether quality held. Report them on one page every month, with the starting figures alongside.

  • How do you work out the hours an AI copilot really saves?

    Turn the share of work it handles alone, and how much quicker that work gets done, into hours a week. Then take off the time people spend checking its drafts, correcting them and keeping it fed with up-to-date information. The number left after that subtraction is the only one worth putting in front of a board.

  • What AI copilot metrics look impressive but mean nothing?

    Ignore total questions asked, test-set scores, hours saved from guessed minutes per task, and how much your team likes the tool. None of them is evidence of value. They measure curiosity, exam conditions, guesses and goodwill rather than what happened on your real work.

Rahul Kaneria

Co-founder and CTO

Rahul is Infoloop's CTO. He sets the architecture for every client build and leads the engineers who ship and support it.

Further reading

More articles on this topic from the Infoloop team.