Resolution rate, first-reply speed, and the split between what your AI closed alone and what a person finished, in plain outcome numbers. Then replay any change against your own past conversations first.
A reply is not a result. What you want to know is how often a customer arrives with something to sort out and leaves with it sorted, no second round, no quiet drop-off. The resolution rate puts that number in front of you, tracked over time, so you can tell at a glance whether the agent is closing the work or just touching it. When the number moves, you know something changed and you can go look.
The gap between a customer reaching out and a first answer landing is the part they feel most, and on most teams nobody is watching it. Here it is a number you can see, measured across every channel the agent covers. A slow stretch shows up as a number going the wrong way instead of as a complaint days later. Fast is the default when the agent answers, and you can prove it rather than assume it.
The honest measure of an AI agent is how much it finishes without a person stepping in. This view splits the two cleanly: the conversations the agent carried end to end, and the ones it passed to your team. It also tells you why it passed them, because the handoffs are not random. The agent hands up when it hits a limit you set and will not cross, or when it reads a high-risk signal like a customer sounding ready to leave. You see the real division of labor, week over week. The agent kept what it could safely keep and escalated the rest on purpose.
Fast and resolved is only half of good. The other half is consistent: the same kind of request handled the same way every time, in line with the calls your team has already made. As the agent works, it leans on past decisions, the close matches from your own history. You can see how often its answers agree with them. A reply backed by a dozen similar cases that mostly went well is a reply you can stand behind. A topic where the agent keeps diverging from precedent is exactly the place to look before it becomes a pattern of its own.
“Your unit is still under warranty, so I have booked the repair at no charge.”
Changing how the agent answers is usually a guess you only grade later, in production, on real customers. This view lets you ask the question first: take a change you are considering, with the policies and precedents that shape how the agent decides, and run it against conversations that already happened. Then read the projected result before it ever reaches a live customer. You decide with evidence from your own history instead of shipping a hunch and hoping.
Decide with evidence from your own history, not a hunch.
Put your AI on your conversations and watch resolution rate, reply speed, and the AI-versus-human split move in plain numbers. Test the next change against your own history first. Analytics that show the work and the judgment behind it.