Measurement guide · Written for the person reporting the number
An AI agent containing 90% of your conversations might be your worst-performing employee
Containment is the number every vendor leads with, every board pack quotes, and almost nobody reads correctly. It is also the easiest metric in the category to game — because an agent that never escalates contains everything, including the conversations it is quietly ruining.
Eight numbers actually matter: containment rate, true resolution rate, CSAT split between contained and escalated conversations, time to human, ranked escalation reasons, first-response time, cost per conversation, and revenue influenced. The last one is the most under-measured metric in customer service, which is precisely why most businesses undervalue what an agent is worth to them — usually by more than the entire software bill.
None of this needs a new tool or an analyst. It needs four numbers read together, a weekly thirty-minute habit, and the willingness to compute the half of the ROI equation everybody skips.
Start by putting your own four numbers in and seeing what they say about each other.
By Jugl·12 min read·Meta Business Partner·1,000+ businesses
Four numbers, read the way an operator reads them
No single metric means anything alone · drag any slider and watch the diagnosis change
Conversations handled entirely by the AI. The headline number, and the easiest one in the category to game.
Contained conversations where the issue was actually solved — no follow-up contact within 48 hours. Always lower than containment.
Asked after conversations the AI handled alone. Below 3.6 is a red flag regardless of volume.
Asked after conversations that reached a human. If this is much lower, the problem is the handoff, not the agent.
72% contained against 48% actually resolved is a 24-point gap. Those conversations ended without a human — but they also ended without an answer. Somebody gave up, re-contacted through another channel, or simply went elsewhere. The gap between these two numbers is the single most diagnostic figure you have, and almost nobody computes it.
Do this next: Measure true resolution with a no-follow-up-within-48-hours proxy, then read twenty of the conversations that failed it.
The diagnosis is a reading of your four inputs against the benchmark ranges below, not an audit of your actual deployment. Its only job is to demonstrate the habit: never read containment on its own. An agent containing 90% of conversations by refusing to escalate is failing loudly and reporting success.
The 60-second version
The eight metrics that matter for an AI customer agent are containment rate (50–75% is realistic), true resolution rate, CSAT split between contained and escalated conversations, time to human, ranked escalation reasons, first-response time, cost per conversation, and revenue influenced.
The trap: containment alone is the most-quoted and most-misleading number in the category. An agent that contains 90% of conversations by refusing to escalate is failing, not succeeding. Always read it next to true resolution rate and contained CSAT.
ROI, properly: (contained conversations × handling minutes × loaded hourly cost) + (after-hours conversations captured × conversion rate × average order value) − software, channel fees and maintenance. Most businesses compute only the first term, and for consumer businesses the second is frequently larger.
The habit that beats every dashboard: someone reads twenty real conversations a week. That is where you find the agent confidently quoting a policy that changed three months ago — which no chart will ever show you.
The eight metrics that matter
Ordered roughly by how often they are quoted, which is almost exactly the inverse of how useful they are. The last one is the one that changes the business case.
One structural point about metric five. Escalation reasons are not a performance metric at all — they are a work queue, delivered to you weekly by your own customers, ranked by frequency, for free. Teams that read them see containment rise 15–25 points in the first quarter. Teams that do not see containment plateau in month two and conclude the product underperforms. The difference is thirty minutes a week. How those escalations should be designed in the first place is in AI-to-human handoff.
Realistic benchmarks
Starting ranges rather than standards. Vary them by category: a technical product with a complex catalogue will contain less than a fashion retailer answering sizing questions, and that is the shape of the problem rather than a failure.
| Metric | Weak | Typical | Strong |
|---|---|---|---|
| Containment rate | Under 40% | 50–65% | 70–80% |
| True resolution rate | Under 35% | 45–60% | 65–75% |
| Contained CSAT | Under 3.5 / 5 | 3.8–4.2 | 4.3+ |
| Escalated CSAT | Under 3.5 / 5 | 4.0–4.4 | 4.5+ |
| Time to human | Over 30 min | 5–15 min | Under 5 min |
| First response (AI) | Over 30 sec | 2–10 sec | Under 2 sec |
| Repeat-explanation rate | Over 30% | 10–20% | Under 5% |
Note the last row, because no platform will report it for you. Repeat-explanation rate — how often a customer restates their issue to the human after escalation — has to be measured by sampling transcripts. Almost every team is wrong about their own number the first time they check, and it is the clearest single signal of whether your handoff is designed or merely switched on.
The five metrics to ignore
None of these are lies. They simply do not tell you anything you can act on, and they crowd out the numbers that do — particularly in vendor decks, where they appear for exactly that reason.
Calculating ROI properly — both sides
Two terms, and most businesses only compute the first. The cost side is easy to model and systematically overstated in importance. The revenue side is harder to model, usually larger, and almost always ignored.
Both sides of the ledger, not just the cheap one
Cost saved · revenue captured · maintenance included, because it always arrives
Across every channel — WhatsApp, Instagram, Messenger, web chat, email.
50–75% is realistic for a well-grounded agent. Use your real number, not the vendor's.
How long one of these conversations takes a human, including the context-switching around it.
Salary plus employment costs, tools and management overhead — not the headline wage.
Evenings and weekends. For consumer businesses this is routinely a third of all inbound, and almost nobody has measured it.
Of the conversations you would otherwise have lost overnight, the share that turns into an order once answered.
Or average booking value, or first-order value for subscription businesses.
Subscription plus channel fees. Jugl's published tiers are Free, $31, $119 and $390 with the AI included.
Total cost includes $300 of maintenance — roughly 12 hours a month of someone reading conversations and updating the knowledge base, valued at your own hourly cost. Teams that leave this out of the model are the same teams whose containment plateaus in month two and who then conclude the product underperformed.
Directional modelling from your own inputs, not a quote, forecast or guarantee of results. Deflection savings assume contained conversations would otherwise have consumed a human's time at your stated handling time and loaded cost; capture revenue assumes after-hours conversations answered immediately convert at your stated rate. Both are estimates — the value of running them is that it forces the after-hours number into the conversation, where it usually belongs and rarely appears.
Include maintenance honestly
Two to four hours a week of someone reading conversations and updating the knowledge base, especially in the first quarter. Leaving it out of the model does not make it disappear; it just means it never gets scheduled, which is how containment plateaus and a perfectly good deployment gets written off. It is the cheapest input in the whole system and the one with the largest effect on the output.
And be honest about the baseline. The comparison is not the agent against zero — it is the agent against what your current setup already costs in late replies, after-hours silence and conversations answered by whoever got there first. If you have never priced that, the pricing question will feel much harder than it is; the pricing guide opens with exactly that calculation.
Reporting cadence that people actually keep
How to read a vendor’s performance numbers
Every number in a vendor deck is true and selected. These six questions convert a flattering headline into something you can compare — and they are far more useful in a sales call than negotiating 10% off.
The last question is the one buyers underrate. Pricing model and measurement are entangled: on per-resolution billing, every point of containment you gain also raises your invoice, which quietly discourages the exact optimisation you are trying to run. The four billing models and what each one does to your behaviour are set out in AI customer service pricing, and the vendor-by-vendor view is in best AI agent for business.
What Jugl makes easy to measure
Measurement is mostly an architecture question rather than a dashboard question. Four separate tools produce four sets of numbers that cannot be honestly added together; one agent across every channel produces one. That is the substance of what follows.
Worth stating the limit plainly: no platform reads your transcripts for you, and no dashboard replaces the weekly thirty minutes. What good architecture does is make the numbers comparable and the transcripts easy to reach, so the habit is cheap enough that someone actually keeps it. Full product detail is on what is Jugl.
What measuring nothing costs you
There is a specific cost to postponing this, and it is not the software. It is that the improvement curve does not start until the measurement does. Deployments improve steeply in weeks two to eight because escalation reasons show you what to fix — but only if someone is collecting them. A quarter without measurement is not a quarter saved; it is a quarter of the compounding curve you never got.
The second cost is the one nobody notices: without the revenue side of the model, this purchase looks like a cost centre, gets budgeted like a cost centre, and gets scoped like one — reactive deflection only, no proactive messaging, no selling, no booking. Businesses that measure revenue influenced end up with a materially more capable deployment, because they gave it a job worth expanding. Measurement is not reporting. It is what determines how big the thing is allowed to get.
Questions teams ask about the numbers
What is a good containment rate for an AI agent?
What is the difference between containment and resolution?
How do I measure AI agent performance?
How long before AI agent performance stabilises?
What if containment is high but CSAT is low?
How do I calculate the ROI of an AI customer service agent?
Should I survey every conversation for CSAT?
Which AI agent metrics are vanity metrics?
How do I measure revenue influenced by an AI agent?
What should I ask a vendor about their performance numbers?
How often should I review AI agent performance?
Do these benchmarks apply to every industry?
Benchmarks are someone else’s numbers. Get your own.
One week on a free tier, pointed at your own website, catalogue and policies, produces a real containment rate on real questions from real customers. That is worth more than every benchmark table on the internet, including the one further up this page.
No card. No developer. No implementation project. And no meter running while you measure.
The improvement curve starts the week the measurement does — and not before. Every quarter you wait is a quarter you do not get back.
SOC 2 Type 2 · HIPAA compliant · Meta Business Partner · NVIDIA Inception · 1000+ businesses
Keep reading
Sources: Jugl deployment experience and published product documentation. Benchmark ranges for containment, true resolution, CSAT, time to human and repeat-explanation rate are directional figures drawn from typical deployments rather than industry standards or guarantees, and vary materially by category, channel mix and knowledge base quality. Diagnostic and ROI outputs are estimates generated from your own inputs, not quotes, forecasts or guarantees of results. Meta, WhatsApp, Messenger, Instagram and Facebook are trademarks of Meta Platforms, Inc.; Jugl is a Meta Business Partner and this guide is published by Jugl and is not endorsed by or affiliated with Meta Platforms, Inc. All other product names are trademarks of their respective owners.