Benchmark report · Written by a vendor, and honest about it
AI agent benchmarks: what good actually looks like, and how to spot an inflated number
Median tier-1 automation across programmes sits at about 41%, with the top quartile near 59%. New deployments launch at 40–50% and pass 60% after six to twelve months of tuning.
Treat 60–67% as solid, 70–75% as strong, and 80%+ as best-in-class — achievable only on unusually structured workloads. If a vendor quotes you 90%, the useful response is not scepticism. It is a question: what is the denominator?
This page exists because the single most expensive thing in this category is not a bad model. It is a well-run business making decisions off a number that means something different from what everyone in the room assumes it means.
By Jugl11 min read8 benchmark metricsInteractive scorer
The 60-second version
Containment: 40–50% at launch, 55–60% median once tuned, 65–72% strong, 80%+ best-in-class on structured workloads only. The cross-programme tier-1 automation median is ~41% with the top quartile at ~59% (Aissist.io).
CSAT: 4.32–4.41 out of 5 on well-structured intents, but 3.34 on complaint handling (Zendesk). That single split explains most negative sentiment about AI support.
Re-contact: 11.3% on AI-resolved conversations against 8.7% human-resolved (Zendesk). Above 15% means your containment number is overstating what is genuinely being resolved.
The definitions matter more than the numbers. Deflection counts abandonment as a win; resolution does not. Industry-wide deflection is 45%+ while genuine self-service resolution is ~14% (Gartner). Ask what the denominator is before you compare anything.
- What is an AI agent benchmark?
- At a glance
- The definitions problem
- The benchmark table
- The three numbers that should worry you
- What separates a 41% deployment from a 70% one
- How to benchmark your own deployment
- How to tell when a vendor is inflating
- The questions behind the benchmarks
- Methodology and disclosure
- FAQ — 20 questions
- People also ask
Definition
What is an AI agent benchmark?
An AI agent benchmark is a published performance band for a customer-service AI deployment, expressed across a small set of metrics: containment or resolution rate, CSAT, escalation rate, re-contact rate and cost per resolution. Benchmarks exist because raw numbers are meaningless in isolation — 48% containment is either good or poor depending entirely on workload structure and time since launch. Cross-programme medians place tier-1 automation near 41%, with the top quartile around 59% (Aissist.io); mature deployments contain 55–72%, and CSAT runs 4.32–4.41 out of 5 on structured intents against 3.34 on complaints (Zendesk). The critical caveat is definitional: deflection, containment and resolution are not interchangeable, and most inflated vendor claims live in that ambiguity.
Definition maintained by the Jugl Editorial Team. Bands are compiled from Zendesk, Gartner, Aissist.io, Lorikeet and Salesforce State of Service.
AI agent benchmarks at a glance
At a glance
- What it measures
- Whether an AI deployment is genuinely resolving customer issues, or merely avoiding humans.
- Containment bands
- 40–50% at launch · 55–60% median tuned · 65–72% strong · 80%+ best-in-class.
- Cross-programme median
- ~41% tier-1 automation; top quartile ~59% (Aissist.io).
- CSAT bands
- 4.32–4.41/5 structured intents · 3.34/5 complaints (Zendesk).
- Escalation rate
- 20–35% normal; below 15% usually means the agent is overreaching.
- Re-contact at 48h
- 11.3% AI-resolved vs 8.7% human-resolved (Zendesk). Above 15% invalidates containment.
- Cost per resolution
- $2–3 median · $1–2 strong · under $1 best-in-class.
- Time to the strong band
- Six to twelve months, with the steepest gains in months two to four.
- Sample needed
- ~200–300 conversations before containment and CSAT stabilise enough to act on.
- Biggest single lever
- Knowledge grounding — ~85% more accurate than ungrounded systems.
- Most misused metric
- Deflection, quoted as though it were resolution. It counts abandonment as success.
- Who should measure this
- Anyone running an AI agent past month one. Five numbers, reviewed weekly and monthly.
The definitions problem
Before any benchmark means anything, four terms need pinning down, because vendors use them interchangeably and the differences are enormous.
| Term | What it means | How it gets inflated |
|---|---|---|
| Deflection | The conversation did not reach a human | Includes customers who gave up and left |
| Containment | The AI handled the conversation end to end | Same problem, slightly narrower |
| Resolution | The customer's issue was actually solved | Hard to measure, so often quietly substituted |
| First-contact resolution | Solved on the first interaction, no re-contact | The honest one, and the rarest to be quoted |
Aissist's benchmark makes the point directly: deflection counts any conversation that did not reach a human, including customers who abandoned. That is why deflection numbers look so much better than resolution numbers, and why the industry-wide picture is 45%+ deflection against roughly 14% actual self-service resolution (Gartner).
The benchmark table
| Metric | Weak | Median | Strong | Best-in-class |
|---|---|---|---|---|
| Tier-1 automation | <30% | ~41% | 59–67% | 75%+ |
| Containment (mature, 6–12mo) | <45% | 55–60% | 65–72% | 80%+ |
| First-contact resolution | <40% | 50% | 55–70% | 75%+ |
| CSAT, structured intents | <3.8 | 4.0 | 4.32–4.41 | 4.5+ |
| CSAT, complaints | <3.0 | 3.34 | 3.8 | 4.2+ |
| Re-contact rate | >15% | 11.3% | <9% | <7% |
| Escalation rate | >50% | 22–35% | 20–25% | 15–22% |
| Cost per resolution | >$5 | $2–3 | $1–2 | <$1 |
Sources: Zendesk · Gartner · Aissist.io · Lorikeet · Salesforce State of Service
Score your own deployment against the bands
Four numbers · the diagnosis matters more than the grade
Resolution, not deflection. Currently: Launch band.
Out of 5, structured intents. Currently: Median.
The honesty check on containment. Currently: Weak.
Performance climbs for six to twelve months. Judging a deployment at week three tells you nothing.
The three numbers that should worry you
What separates a 41% deployment from a 70% one
It is not the model. Based on the benchmark data and how these deployments actually progress, five things account for nearly all of the variance — and every one of them is a decision you control.
How to benchmark your own deployment
Month one, measure five things: containment, escalation rate, escalation reasons, CSAT split by contained versus escalated, and re-contact rate at 48 hours. That is the whole instrument panel. Then read it like this:
- Containment — AI handled end to end, abandonment excluded from the numerator
- Escalation rate — 20–35% normal; below 15% means the agent is overreaching
- Escalation reasons — a free, pre-labelled backlog of what the agent cannot do
- CSAT split by contained versus escalated — never blended
- Re-contact at 48 hours — the honesty check on every other number
The operational version of this — what to instrument, what to review weekly, and what to ignore — is in how to measure AI agent performance.
How to tell when a vendor is inflating
Where Jugl sits against these benchmarks
Stated plainly, since this page is published by a vendor. Jugl customers typically see around 73% fewer tickets reaching a human at roughly 94% satisfaction — which lands in the strong band on the table above, not the best-in-class one, and which is a customer-reported typical result rather than a guarantee. The design decisions behind it are the five in section 04: the agent is grounded in your own catalogue and policies, it can take real actions like looking up an order or booking an appointment rather than only describing them, and complaints hand over to a human with the full thread attached.
The questions behind the benchmarks
What containment rate should my deployment be hitting?
Short answer
40–50% in the launch band, 55–60% once tuned, 65–72% strong, and 80%+ only on unusually structured workloads. The cross-programme median for tier-1 automation is about 41% with the top quartile near 59% (Aissist.io) — so if you are at 55% you are already ahead of most deployments, not behind the marketing.
Example
What should change as a deployment matures?
Short answer
The constraint moves. In the launch band the limiting factor is knowledge coverage; in the middle band it is intent triage — deciding what the agent should not attempt; in the strong band it is action capability, meaning what the agent is allowed to actually do rather than describe. Chasing the wrong constraint is why deployments stall.
| Metric | Launch (0–3 mo) | Tuning (3–9 mo) | Mature (9 mo+) |
|---|---|---|---|
| Containment | 40–50% | 55–60% | 65–72% |
| CSAT, structured intents | 3.9–4.1 | 4.1–4.3 | 4.32–4.41 |
| Escalation rate | 35–50% | 25–35% | 20–25% |
| Re-contact at 48h | 14–18% | 11–13% | <9% |
| Main constraint | Knowledge coverage | Intent triage | Action capability |
| What to do next | Read every escalation | Route complaints to humans | Give the agent more real actions |
Which metrics can actually be trusted?
Short answer
Re-contact rate at 48 hours and CSAT split by tier are the two hardest to game, which is why they are the two you should demand. Deflection is the easiest to inflate because it counts abandonment as success, and it is consequently the metric vendors quote most often.
| Metric | Trustworthiness | Why |
|---|---|---|
| Deflection rate | Low | Counts abandonment as success — inflates by design |
| Containment rate | Medium | Honest if abandonment is excluded from the numerator |
| Resolution rate | High | Hard to measure, which is why it gets substituted |
| First-contact resolution | Highest | The one worth quoting, and the rarest to be quoted |
| CSAT (blended) | Low | Hides the complaint-handling collapse to 3.34 |
| CSAT (split by tier) | High | Shows you exactly which conversations to stop automating |
| Re-contact at 48h | Highest | The honesty check on every other number |
Benchmarking: what it tells you and what it cannot
- ✓Whether your number is genuinely poor or simply early
- ✓Which constraint to fix next, from the shape of the gap
- ✓A defensible basis for a vendor conversation
- ✓Early warning when containment rises but quality falls
- ✓A way to value tuning effort against other work
- ×Whether your workload is comparable to the sample — it often is not
- ×What your specific customers will tolerate
- ×Whether the vendor measured the same thing you did
- ×Anything useful below ~200 conversations in the period
- ×Voice performance — the published data is predominantly chat and messaging
Methodology and disclosure
Written by
Jugl Editorial TeamJugl Inc., Frisco, Texas — an AI customer agent platform used by 1,000+ businesses.
Reviewed by
Jugl product & customer operationsChecked against live deployment data and current vendor documentation.
Methodology & disclosure
How the bands were compiled. Values are drawn from named third-party research — Zendesk CX benchmark data for CSAT and re-contact, Aissist.io for tier-1 automation medians and quartiles, Gartner for cost per contact and self-service resolution, Lorikeet for AI-native first-contact resolution, Salesforce State of Service for loaded agent cost. Where sources express a metric differently, the band is widened rather than reconciled into false precision.
Why the definitions section comes first. Deflection, containment, resolution and first-contact resolution are measured differently and quoted interchangeably. A benchmark table built on top of that ambiguity is worse than no table, because it lends false authority to a comparison between two things that were never the same measurement. Pinning the terms down is what makes the rest of the page usable.
Conflict of interest. This report is published by Jugl, which sells an AI customer agent and is therefore an interested party. Jugl's own figures — around 73% fewer tickets reaching a human at roughly 94% satisfaction — are customer-reported and typical rather than guaranteed, and are placed in the strong band rather than the best-in-class one, in the same table as everyone else's.
How this page is maintained. Bands are reviewed against their originating sources and against live deployment data. The page carries no year stamp because a dated benchmark misleads the moment it ages. Scorer outputs are diagnostic estimates generated from your own inputs — not quotes, forecasts or guarantees.
AI agent benchmarks: 20 questions answered
What is a good AI containment rate?
Why is my resolution rate lower than the vendor promised?
What is a normal escalation rate for an AI agent?
How long until AI agent performance stabilises?
What is the difference between deflection, containment and resolution?
What is a good CSAT for an AI agent?
What re-contact rate should I expect?
How do I know if a vendor is inflating their numbers?
Which metrics should a small business actually track?
What is a benchmark actually worth if my business is unusual?
How do I calculate containment rate correctly?
Should I measure cost per resolution or cost per contact?
What is a good first response time for an AI agent?
How many conversations do I need before benchmarks mean anything?
Do these benchmarks apply to voice AI as well as chat?
What benchmark should I hold a vendor to contractually?
How often should I review AI agent performance?
Does a better AI model improve these benchmarks?
What does Jugl benchmark at?
People also ask
The gap between 41% and 70% is six months of somebody paying attention
That is the finding underneath every number on this page. The top quartile did not buy a better model — they started earlier and read their escalation logs. The steepest gains land in months two through four, which means the cost of waiting a quarter is not zero. It is the quarter.
Point Jugl at your website, catalogue and policies, connect a channel, and start generating your own numbers instead of comparing other people's. The free tier is permanent, needs no card, and gives you the only benchmark that ever really mattered: your own transcripts.
Anyone can quote a containment rate. Very few will show you the denominator.
SOC 2 Type 2 · HIPAA compliant · Meta Business Partner · NVIDIA Inception · 1000+ businesses
Keep reading
Sources: Zendesk CX benchmark data (CSAT by intent tier, re-contact rates); Gartner customer service and support research (self-service resolution, cost per contact, agentic resolution projections); Aissist.io AI service benchmark (tier-1 automation median and quartiles, deflection definition); Lorikeet AI customer service analysis (first-contact resolution and cost per resolution for AI-native platforms); Salesforce State of Service (loaded agent cost, case resolution projections). Benchmark bands are compiled from these sources and are directional rather than certified; definitions of deflection, containment and resolution differ between them and are distinguished explicitly in the text. This report is published by Jugl, which sells an AI customer agent and is therefore an interested party — Jugl's own outcome figures are customer-reported and typical rather than guaranteed, and are stated in the same table bands as everyone else's. Scorer outputs are diagnostic estimates generated from your own inputs, not quotes, forecasts or guarantees. Meta, WhatsApp, Messenger, Instagram and Facebook are trademarks of Meta Platforms, Inc.; Jugl is a Meta Business Partner and this page is published by Jugl and is not endorsed by or affiliated with Meta Platforms, Inc. All other product names are trademarks of their respective owners.
Start free at Jugl · No card required · Permanent free tier