The Only Five Metrics That Tell You If Your AI Is Working
Six weeks after launch, someone asks whether the AI is worth keeping. The dashboard says 412 conversations. Nobody in the room can say whether that is good, because nobody knows how many enquiries the business used to get, or how many of them turned into work.
That conversation happens constantly, and it is why choosing your AI KPIs for business matters more before launch than after. Five numbers cover almost everything you need. The harder part is that four of the five have to be measured while the old system is still running, and once the AI is live, the old system is gone.
Record the baseline before anything goes live#
Give yourself a fortnight of deliberate counting. It is tedious and it is the difference between a renewal decision based on evidence and one based on a feeling.
Pull the missed call count from your phone system, or from the call log on the mobile if that is where enquiries land. Count enquiries by channel for two weeks, including the ones that arrive as an Instagram DM or a text to your personal number. Note roughly how long you take to reply to each, and how many of those enquiries became paid work. The method for the call side is in what missed calls are really costing you.
If you are already live and skipped this, reconstruct what you can. Your phone provider holds call records, your email holds timestamps, and your accounting software holds job counts by month. A reconstructed baseline is imperfect and still far better than arguing from memory.
The five metrics worth tracking#
Containment rate#
The share of conversations the AI finished on its own, with no human touching it. This is the closest thing to a workload number, and it tells you how much of your inbox the system genuinely removed.
Read it carefully, because it can lie in both directions. A low rate might mean the agent lacks knowledge it should have, which is a fixable content problem rather than a technology one. A very high rate can mean the agent is stubbornly refusing to escalate, so check the outcomes of contained conversations rather than trusting the percentage alone.
Baseline equivalent: the share of enquiries your team currently resolves in one reply, without a callback or a second exchange.
Response time#
Median time from an enquiry arriving to a substantive first reply. Use the median rather than the average, because one enquiry you answered nine days late will drag an average into meaninglessness.
Response time is the metric most likely to move dramatically, because the comparison is often between a few hours and a few seconds. It is also the one most closely tied to whether you win the job, since a customer who has messaged three businesses tends to engage with whoever replies first.
Baseline equivalent: the same measurement across your last fortnight of enquiries, taken from timestamps you already have.
Qualified enquiry rate#
Of the conversations the AI handled, how many produced an enquiry worth your time. That means a real job in your service area, with a budget signal and a name and number attached.
This is the metric that separates a busy system from a useful one. Volume rising while qualified enquiries stay flat means the agent is filling your day with conversations that were never going anywhere. Decide what qualified means for your business before launch and write the definition down, because it will drift otherwise. The dimensions worth using are covered in qualifying leads on your website.
Baseline equivalent: of the enquiries you received last fortnight, how many did you consider worth quoting.
Escalation quality#
Of the conversations that went to a human, how many arrived with enough context that the customer did not have to start again. This is a quality measure rather than a volume one, and it is the only one on this list you assess by reading rather than counting.
Take ten escalated conversations a month and read them end to end. You are looking for whether the handoff happened at the right moment and whether the transcript came across intact. A system that escalates late, after the customer has already asked twice for a person, is doing measurable damage even when every other number looks healthy. Getting the handoff right sets out the triggers.
Baseline equivalent: how often a customer currently has to repeat their story when a call gets passed between people in your business.
Cost per resolved enquiry#
Total monthly spend on the system, including usage charges above the base plan, divided by the number of enquiries it resolved without human involvement. One number, easy to compare, and it is the figure that should decide renewal.
Compare it against what the same work costs in staff time. If an admin hour costs you $35 and a resolved enquiry takes eight minutes of that hour, you have your comparison. The full arithmetic, including a worked example where the AI does not pay for itself, sits in working out the ROI of AI automation.
Baseline equivalent: your current cost per handled enquiry, which is staff hours spent on enquiry handling multiplied by the loaded hourly rate, divided by enquiries handled.
Be careful about the enquiries the old process never handled at all. The calls that rang out at seven in the evening cost you nothing in staff time, so they make your historical cost per enquiry look flattering. Those are also the enquiries the AI is most likely to convert, since the alternative was voicemail. Count them separately rather than folding them into the comparison, or you will understate what the system is doing for you.
What to record, and where the number comes from#
| Metric | Where you get it | When to baseline |
|---|---|---|
| Containment rate | AI platform reporting | Two weeks before launch, from your own resolution habits |
| Response time | Timestamps on calls, email and DMs | Two weeks before launch |
| Qualified enquiry rate | Your own definition applied to transcripts | Two weeks before launch |
| Escalation quality | Reading ten transcripts a month | After launch only |
| Cost per resolved enquiry | Invoice divided by platform reporting | Two weeks before launch, using staff time |
Four of the five need work before you switch anything on. That is the whole argument for spending a fortnight on this while the system is still in shadow mode, which is where the 90-day implementation roadmap puts it.
The monthly review that takes half an hour#
Put a recurring half hour in one person's calendar. Nobody owning the numbers is one of the most common AI mistakes small businesses make, and it is the reason systems quietly drift into quoting last year's prices.
In that half hour, look at the five numbers against the baseline, then read ten conversations. The reading matters more than the dashboard. Patterns show up in transcripts weeks before they show up in a percentage, and most of what you find is a content gap you can fix in ten minutes by adding an answer the agent did not have.
Write down one change per month and make it. Chasing five improvements at once means you cannot attribute any movement afterwards, which puts you back where you started.
Where to begin#
If you have not launched yet, start the fortnight of counting today. Open a note on your phone and record every enquiry as it arrives, with the channel and the time. Two weeks of that gives you three of the five metrics and costs you a few seconds a day.
If you are already live, spend an hour reconstructing what you can from your phone records and your accounting software, then set the monthly review. An imperfect baseline built in an afternoon still lets you answer the question that eventually gets asked in every business, which is whether this thing is earning its keep. For how these numbers fit into a broader plan, the overview of AI solutions for small business covers the sequencing.
Common questions
What is a good containment rate for an AI agent?
Containment is the share of conversations the AI finishes without a human. For a well-configured agent handling routine enquiries, somewhere between half and three quarters is realistic. A very high rate can mean the agent is refusing to escalate, so always read it alongside how those conversations ended.
How do I measure AI performance if I have no baseline?
Reconstruct one from what you already hold. Phone systems keep missed call logs, email keeps timestamps, and your accounting software holds job counts and values by month. A rough baseline built from records you already have beats no baseline, though a fortnight of deliberate counting is better.
What is cost per resolved enquiry?
Total monthly cost of the AI system, including subscription and usage charges, divided by the number of enquiries it resolved without human involvement. It gives you a single figure to compare against what the same work costs in staff time, which is the comparison that decides renewal.