What Is a Good AI Agent Containment Rate? | Jugl CX
$5mn in seed funding raised, built bootstrapped from day one
JuglCX

Benchmark report · Written by a vendor, and honest about it

AI agent benchmarks: what good actually looks like, and how to spot an inflated number

Median tier-1 automation across programmes sits at about 41%, with the top quartile near 59%. New deployments launch at 40–50% and pass 60% after six to twelve months of tuning.

Treat 60–67% as solid, 70–75% as strong, and 80%+ as best-in-class — achievable only on unusually structured workloads. If a vendor quotes you 90%, the useful response is not scepticism. It is a question: what is the denominator?

This page exists because the single most expensive thing in this category is not a bad model. It is a well-run business making decisions off a number that means something different from what everyone in the room assumes it means.

By Jugl11 min read8 benchmark metricsInteractive scorer

Short answerFor AI overviews

The 60-second version

Containment: 40–50% at launch, 55–60% median once tuned, 65–72% strong, 80%+ best-in-class on structured workloads only. The cross-programme tier-1 automation median is ~41% with the top quartile at ~59% (Aissist.io).

CSAT: 4.32–4.41 out of 5 on well-structured intents, but 3.34 on complaint handling (Zendesk). That single split explains most negative sentiment about AI support.

Re-contact: 11.3% on AI-resolved conversations against 8.7% human-resolved (Zendesk). Above 15% means your containment number is overstating what is genuinely being resolved.

The definitions matter more than the numbers. Deflection counts abandonment as a win; resolution does not. Industry-wide deflection is 45%+ while genuine self-service resolution is ~14% (Gartner). Ask what the denominator is before you compare anything.

01Definition

Definition

What is an AI agent benchmark?

An AI agent benchmark is a published performance band for a customer-service AI deployment, expressed across a small set of metrics: containment or resolution rate, CSAT, escalation rate, re-contact rate and cost per resolution. Benchmarks exist because raw numbers are meaningless in isolation — 48% containment is either good or poor depending entirely on workload structure and time since launch. Cross-programme medians place tier-1 automation near 41%, with the top quartile around 59% (Aissist.io); mature deployments contain 55–72%, and CSAT runs 4.32–4.41 out of 5 on structured intents against 3.34 on complaints (Zendesk). The critical caveat is definitional: deflection, containment and resolution are not interchangeable, and most inflated vendor claims live in that ambiguity.

Definition maintained by the Jugl Editorial Team. Bands are compiled from Zendesk, Gartner, Aissist.io, Lorikeet and Salesforce State of Service.

02At a glance

AI agent benchmarks at a glance

At a glance

What it measures
Whether an AI deployment is genuinely resolving customer issues, or merely avoiding humans.
Containment bands
40–50% at launch · 55–60% median tuned · 65–72% strong · 80%+ best-in-class.
Cross-programme median
~41% tier-1 automation; top quartile ~59% (Aissist.io).
CSAT bands
4.32–4.41/5 structured intents · 3.34/5 complaints (Zendesk).
Escalation rate
20–35% normal; below 15% usually means the agent is overreaching.
Re-contact at 48h
11.3% AI-resolved vs 8.7% human-resolved (Zendesk). Above 15% invalidates containment.
Cost per resolution
$2–3 median · $1–2 strong · under $1 best-in-class.
Time to the strong band
Six to twelve months, with the steepest gains in months two to four.
Sample needed
~200–300 conversations before containment and CSAT stabilise enough to act on.
Biggest single lever
Knowledge grounding — ~85% more accurate than ungrounded systems.
Most misused metric
Deflection, quoted as though it were resolution. It counts abandonment as success.
Who should measure this
Anyone running an AI agent past month one. Five numbers, reviewed weekly and monthly.
SOC 2 Type 2certified
HIPAAcompliant
MetaBusiness Partner
1,000+businesses
03Definitions

The definitions problem

Before any benchmark means anything, four terms need pinning down, because vendors use them interchangeably and the differences are enormous.

TermWhat it meansHow it gets inflated
DeflectionThe conversation did not reach a humanIncludes customers who gave up and left
ContainmentThe AI handled the conversation end to endSame problem, slightly narrower
ResolutionThe customer's issue was actually solvedHard to measure, so often quietly substituted
First-contact resolutionSolved on the first interaction, no re-contactThe honest one, and the rarest to be quoted

Aissist's benchmark makes the point directly: deflection counts any conversation that did not reach a human, including customers who abandoned. That is why deflection numbers look so much better than resolution numbers, and why the industry-wide picture is 45%+ deflection against roughly 14% actual self-service resolution (Gartner).

Three questions for any vendor quoting you a number. What is the denominator? Does an abandoned conversation count as a success? How do you measure whether the issue was actually solved? Those three questions have ended more sales processes than any feature comparison, which is exactly why they are worth asking on the first call rather than the fourth.
04The table

The benchmark table

MetricWeakMedianStrongBest-in-class
Tier-1 automation<30%~41%59–67%75%+
Containment (mature, 6–12mo)<45%55–60%65–72%80%+
First-contact resolution<40%50%55–70%75%+
CSAT, structured intents<3.84.04.32–4.414.5+
CSAT, complaints<3.03.343.84.2+
Re-contact rate>15%11.3%<9%<7%
Escalation rate>50%22–35%20–25%15–22%
Cost per resolution>$5$2–3$1–2<$1

Sources: Zendesk · Gartner · Aissist.io · Lorikeet · Salesforce State of Service

Score your own deployment against the bands

Four numbers · the diagnosis matters more than the grade

Containment (resolved end to end)48%

Resolution, not deflection. Currently: Launch band.

CSAT on AI-handled conversations4.00

Out of 5, structured intents. Currently: Median.

Re-contact within 48 hours13%

The honesty check on containment. Currently: Weak.

Months since launch4 mo

Performance climbs for six to twelve months. Judging a deployment at week three tells you nothing.

ContainmentLaunch band48% vs 55–60% median
CSATMedian4.00 vs 4.32–4.41 strong
Re-contactWeak13% vs 11.3% AI benchmark
What your numbers actually sayMiddle of the field — roughly where the 41% median sits, and where most deployments stall permanently. The businesses that climb out do one unglamorous thing: read every escalation for a week, and fix the top five reasons. That is usually worth 10–15 points of containment on its own.
05Red flags

The three numbers that should worry you

Complaint-handling CSAT of 3.34Against 4.32–4.41 for structured queries (Zendesk). Nothing in the dataset is more diagnostic. AI is good at answering questions and bad at absorbing anger, and every business routing complaints into an AI agent is personally manufacturing the 64% of customers who say they would rather companies did not use AI at all. The fix is not a better model — it is a routing rule, and it takes ten minutes. See the handoff guide.
Re-contact 11.3% on AI vs 8.7% on humanA 2.6-point quality gap: small enough to be acceptable, large enough to be the honesty check on everything else. If your AI-resolved re-contact rate is 20%, your resolution number is fiction and no other metric on your dashboard means anything until you fix it.
The 41% medianMost deployments are well short of what vendors advertise. That is not because the technology does not work — it is because most deployments never get tuned after launch. The gap between 41% and 70% is almost entirely operational effort, and it is effort measured in hours a week, not headcount.
06The gap

What separates a 41% deployment from a 70% one

It is not the model. Based on the benchmark data and how these deployments actually progress, five things account for nearly all of the variance — and every one of them is a decision you control.

1
Knowledge base coverageGrounded systems run roughly 85% more accurate than ungrounded ones. The difference between deployments is almost entirely the quality and coverage of what has been fed in — your real prices, your real policies, your real catalogue. This is where grounding stops being a buzzword and starts being 30 points of containment.
2
Post-launch tuningDeployments climb from 40–50% to 60%+ over six to twelve months. That climb does not happen on its own. It happens because somebody reads the escalations weekly and closes the gaps they expose. An hour a week is the whole intervention.
3
Intent triageSending complaints to AI drags CSAT to 3.34 and generates re-contacts. Routing them straight to humans raises both numbers at once — one of the very few changes in this category that improves two metrics without a trade-off.
4
Action capabilityAn agent that can look up a real order resolves. One that can only describe policy deflects. This is the difference between the 45% deflection number and the 14% resolution number, stated in one sentence — and it is the question to ask a vendor before any other. See AI agent vs chatbot.
5
Escalation speedFast handoff protects CSAT on exactly the conversations AI cannot handle, which is where all the sentiment damage happens. Slow handoff turns a 30-second annoyance into a review.
07Your numbers

How to benchmark your own deployment

Month one, measure five things: containment, escalation rate, escalation reasons, CSAT split by contained versus escalated, and re-contact rate at 48 hours. That is the whole instrument panel. Then read it like this:

Containment below 45% after three monthsYour knowledge base has gaps. Read the escalation log — it is a free, pre-labelled list of every question your agent could not answer, ranked by frequency, generated by your own customers.
Containment high, CSAT lowThe agent is holding conversations it should hand over. Tighten the escalation triggers, especially on anything with emotional content.
Re-contact above 15%Your containment number is overstating what is genuinely being resolved. Treat the containment figure as unreliable until this comes down.
Complaint CSAT near 3.3Route complaints to humans immediately and watch overall CSAT rise within a fortnight. This is the highest-return single change available to most deployments.
Your month-one instrument panel
  • Containment — AI handled end to end, abandonment excluded from the numerator
  • Escalation rate — 20–35% normal; below 15% means the agent is overreaching
  • Escalation reasons — a free, pre-labelled backlog of what the agent cannot do
  • CSAT split by contained versus escalated — never blended
  • Re-contact at 48 hours — the honesty check on every other number

The operational version of this — what to instrument, what to review weekly, and what to ignore — is in how to measure AI agent performance.

08Vendor test

How to tell when a vendor is inflating

1
Ask for the denominator, in writing"90% containment" of what? All inbound? Only conversations the bot chose to engage? Only one intent category? The answer usually arrives with several qualifiers attached, and the qualifiers are the actual product.
2
Ask them to demo a failureEvery vendor can demo a success. Ask to see what happens when the agent does not know, when the customer is angry, and when the question contains two requests at once. That is 30% of your real volume and it is never in the demo script.
3
Bring your five ugliest real messagesNot the tidy example. The multi-part one, the furious one, the one with a typo'd order number and a question about two products. A rehearsed demo cannot survive them, and a good agent does not need to.
4
Model the invoice at three times your volumePer-resolution and per-token pricing scale linearly with success, which means the better the AI performs the more you pay. Ask what the bill looks like when the deployment works — not when it is small. The four pricing shapes are broken down in the pricing guide.

Where Jugl sits against these benchmarks

Stated plainly, since this page is published by a vendor. Jugl customers typically see around 73% fewer tickets reaching a human at roughly 94% satisfaction — which lands in the strong band on the table above, not the best-in-class one, and which is a customer-reported typical result rather than a guarantee. The design decisions behind it are the five in section 04: the agent is grounded in your own catalogue and policies, it can take real actions like looking up an order or booking an appointment rather than only describing them, and complaints hand over to a human with the full thread attached.

The commercial part that affects your benchmarks. Jugl publishes flat tiers — Free, $31, $119 and $390 a month — with the AI included and nothing metered per message, per token or per resolution. That matters here rather than only on the pricing page: under per-resolution billing, every point of containment you gain increases your invoice, which quietly discourages the exact tuning that moves you from 41% to 70%. Flat pricing means improving the deployment is free. Start at what is Jugl.
09Direct answers

The questions behind the benchmarks

What containment rate should my deployment be hitting?

Short answer

40–50% in the launch band, 55–60% once tuned, 65–72% strong, and 80%+ only on unusually structured workloads. The cross-programme median for tier-1 automation is about 41% with the top quartile near 59% (Aissist.io) — so if you are at 55% you are already ahead of most deployments, not behind the marketing.

Example

A deployment at 48% after four months looks disappointing against a vendor's 80% slide and is in fact perfectly normal — it is inside the launch band and about to enter the steepest part of the improvement curve, provided somebody is reading the escalation log.
Key takeawayBenchmark against the published bands and your own months-since-launch, never against a sales deck. Most 'underperforming' deployments are on schedule.

What should change as a deployment matures?

Short answer

The constraint moves. In the launch band the limiting factor is knowledge coverage; in the middle band it is intent triage — deciding what the agent should not attempt; in the strong band it is action capability, meaning what the agent is allowed to actually do rather than describe. Chasing the wrong constraint is why deployments stall.

MetricLaunch (0–3 mo)Tuning (3–9 mo)Mature (9 mo+)
Containment40–50%55–60%65–72%
CSAT, structured intents3.9–4.14.1–4.34.32–4.41
Escalation rate35–50%25–35%20–25%
Re-contact at 48h14–18%11–13%<9%
Main constraintKnowledge coverageIntent triageAction capability
What to do nextRead every escalationRoute complaints to humansGive the agent more real actions
Key takeawayDiagnose which band you are in before deciding what to fix. Adding more knowledge to a deployment whose real problem is routing produces no improvement at all.

Which metrics can actually be trusted?

Short answer

Re-contact rate at 48 hours and CSAT split by tier are the two hardest to game, which is why they are the two you should demand. Deflection is the easiest to inflate because it counts abandonment as success, and it is consequently the metric vendors quote most often.

MetricTrustworthinessWhy
Deflection rateLowCounts abandonment as success — inflates by design
Containment rateMediumHonest if abandonment is excluded from the numerator
Resolution rateHighHard to measure, which is why it gets substituted
First-contact resolutionHighestThe one worth quoting, and the rarest to be quoted
CSAT (blended)LowHides the complaint-handling collapse to 3.34
CSAT (split by tier)HighShows you exactly which conversations to stop automating
Re-contact at 48hHighestThe honesty check on every other number
Key takeawayIf you can only track two numbers, track re-contact at 48 hours and CSAT split by contained versus escalated. Between them they expose almost every way a containment figure can lie.

Benchmarking: what it tells you and what it cannot

What benchmarking gives you
  • Whether your number is genuinely poor or simply early
  • Which constraint to fix next, from the shape of the gap
  • A defensible basis for a vendor conversation
  • Early warning when containment rises but quality falls
  • A way to value tuning effort against other work
What benchmarks will not tell you
  • Whether your workload is comparable to the sample — it often is not
  • What your specific customers will tolerate
  • Whether the vendor measured the same thing you did
  • Anything useful below ~200 conversations in the period
  • Voice performance — the published data is predominantly chat and messaging
Get your own numbers before you benchmark themThe free conversation audit shows where your current setup is losing conversations — the raw material for every metric on this page.
Get the free auditNo card required
10EEAT

Methodology and disclosure

Written by

Jugl Editorial Team

Jugl Inc., Frisco, Texas — an AI customer agent platform used by 1,000+ businesses.

Reviewed by

Jugl product & customer operations

Checked against live deployment data and current vendor documentation.

Methodology & disclosure

How the bands were compiled. Values are drawn from named third-party research — Zendesk CX benchmark data for CSAT and re-contact, Aissist.io for tier-1 automation medians and quartiles, Gartner for cost per contact and self-service resolution, Lorikeet for AI-native first-contact resolution, Salesforce State of Service for loaded agent cost. Where sources express a metric differently, the band is widened rather than reconciled into false precision.

Why the definitions section comes first. Deflection, containment, resolution and first-contact resolution are measured differently and quoted interchangeably. A benchmark table built on top of that ambiguity is worse than no table, because it lends false authority to a comparison between two things that were never the same measurement. Pinning the terms down is what makes the rest of the page usable.

Conflict of interest. This report is published by Jugl, which sells an AI customer agent and is therefore an interested party. Jugl's own figures — around 73% fewer tickets reaching a human at roughly 94% satisfaction — are customer-reported and typical rather than guaranteed, and are placed in the strong band rather than the best-in-class one, in the same table as everyone else's.

How this page is maintained. Bands are reviewed against their originating sources and against live deployment data. The page carries no year stamp because a dated benchmark misleads the moment it ages. Scorer outputs are diagnostic estimates generated from your own inputs — not quotes, forecasts or guarantees.

11FAQ

AI agent benchmarks: 20 questions answered

What is a good AI containment rate?
55–60% is median for a tuned deployment, 65–72% is strong, and 80%+ is best-in-class but only achievable on unusually structured workloads. Launch expectations should be 40–50%. The cross-programme median for tier-1 automation is about 41% with the top quartile near 59% (Aissist.io), which tells you most deployments stall rather than climb.
Why is my resolution rate lower than the vendor promised?
Almost always one of two reasons. Either their number was deflection and yours is resolution — deflection counts every conversation that did not reach a human, including customers who gave up — or the deployment has not been tuned since launch. Both are fixable, and the second one is fixable in an afternoon a week.
What is a normal escalation rate for an AI agent?
20–35% is the normal band, with 20–25% considered strong. Below 15% often means the agent is overreaching rather than excelling: it is holding conversations it should be handing over, which shows up two weeks later as soft CSAT and a rising re-contact rate.
How long until AI agent performance stabilises?
Six to twelve months to reach the 60%+ band, with the steepest gains in months two through four. Judging a deployment at week three tells you nothing except how good the initial knowledge base was.
What is the difference between deflection, containment and resolution?
Deflection means the conversation did not reach a human, which includes abandonment. Containment means the AI handled it end to end. Resolution means the customer's problem was actually solved. Industry-wide, deflection runs 45%+ while genuine self-service resolution sits near 14% (Gartner) — that gap is mostly abandoned conversations counted as successes.
What is a good CSAT for an AI agent?
4.32–4.41 out of 5 on well-structured intents is strong (Zendesk); 4.0 is median and 4.5+ is best-in-class. Complaint handling is the exception and benchmarks far lower at 3.34, which is why the highest-return design decision available is routing complaints to a human immediately rather than trying to improve the bot.
What re-contact rate should I expect?
11.3% on AI-resolved conversations against 8.7% human-resolved (Zendesk) — a 2.6-point quality gap that is small enough to accept. Above 15% means your containment number is overstating what is genuinely being resolved, and no other metric on your dashboard is trustworthy until you fix it.
How do I know if a vendor is inflating their numbers?
Ask three questions: what is the denominator, does an abandoned conversation count as a success, and how do you measure whether the issue was actually solved? Then ask them to demo a failure rather than a success. A vendor quoting 90% containment across a mixed workload is either counting deflection or serving a very narrow set of intents.
Which metrics should a small business actually track?
Five, and no more: containment, escalation rate, escalation reasons, CSAT split by contained versus escalated, and re-contact rate at 48 hours. Escalation reasons is the one that pays for itself, because it is a free, pre-labelled list of exactly what your agent cannot yet do.
What is a benchmark actually worth if my business is unusual?
It is worth a direction, not a target. Benchmarks tell you whether 48% containment means you are behind or ahead — and the answer differs enormously by workload structure. A store whose inbound is 60% order-status questions should beat these bands comfortably; a consultancy fielding bespoke scoping questions will sit below them and should. Use the bands to decide whether to investigate, not to set an OKR.
How do I calculate containment rate correctly?
Conversations the AI handled end to end, divided by all inbound conversations in the period — with abandoned sessions excluded from the numerator. That exclusion is the entire difference between an honest number and a vendor number. If a customer opened the chat, read one unhelpful answer and left, that is not containment; counting it is how the industry produces 45%+ deflection against 14% genuine self-service resolution.
Should I measure cost per resolution or cost per contact?
Both, because they answer different questions. Cost per contact tells you what the channel costs to run; cost per resolution tells you what outcomes cost. Median benchmarks are $1.84 per self-service contact against $13.50 agent-assisted (Gartner), and $1–2 per resolution is the strong band. The trap is comparing your cost per resolution against a vendor’s cost per contact — a mismatch that flatters whoever is selling.
What is a good first response time for an AI agent?
Seconds, and this is one of the few metrics where AI simply wins. The benchmark worth setting is not the AI’s response time but your escalated response time — how long a customer waits after the handover. That is where the satisfaction damage happens, and it is the number most dashboards never show.
How many conversations do I need before benchmarks mean anything?
Roughly 200–300 conversations in a period before containment and CSAT stabilise enough to act on. Below that you are reading noise, and the most common consequence is switching off a deployment in week three based on a sample of forty. Escalation reasons, by contrast, are useful immediately — even ten escalations tell you something specific to fix.
Do these benchmarks apply to voice AI as well as chat?
Partially. Containment and resolution translate; CSAT and re-contact do not translate cleanly because voice carries different expectations and different abandonment behaviour. Voice AI now handles about 19% of inbound contact-centre volume (Forrester Wave), roughly tripling in two years, but the published benchmark data in this category is predominantly chat and messaging. Treat voice numbers as a separate dataset.
What benchmark should I hold a vendor to contractually?
Resolution with a stated denominator, plus re-contact at 48 hours, plus CSAT split by contained versus escalated. Those three together are almost impossible to game. A commitment expressed only as deflection or containment can be met by a deployment that is actively annoying your customers, which is why it is the shape vendors prefer.
How often should I review AI agent performance?
Escalation reasons weekly, the full metric set monthly, and the vendor relationship quarterly. The weekly review is the one that produces the results — deployments climb from 40–50% to 60%+ over six to twelve months, and that climb is made almost entirely of somebody reading escalations and closing the top few gaps each week.
Does a better AI model improve these benchmarks?
Less than you would expect. Grounding is the dominant variable: systems grounded in real business data run roughly 85% more accurate than ungrounded ones, and action capability — whether the agent can look up a real order rather than describe a policy — accounts for most of the gap between deflection and resolution. Model upgrades move the number a little; knowledge coverage and integrations move it a lot.
What does Jugl benchmark at?
Jugl customers typically see around 73% fewer tickets reaching a human at roughly 94% satisfaction — which lands in the strong band on the table above rather than the best-in-class one, and which is a customer-reported typical result rather than a guarantee. The commercially relevant detail is that Jugl’s flat published tiers mean improving those numbers does not increase your invoice, whereas per-resolution pricing charges you more for every point of containment you earn.
12People also ask

People also ask

What is a good AI agent containment rate?55–60% is median once tuned, 65–72% is strong, and 80%+ is best-in-class on unusually structured workloads. Launch expectations should be 40–50%.
What is the difference between containment and deflection?Containment means the AI handled the conversation end to end. Deflection only means it never reached a human — which includes customers who gave up and left.
What CSAT should an AI agent achieve?4.32–4.41 out of 5 on structured intents is strong (Zendesk). Complaint handling benchmarks far lower at 3.34, which is a boundary rather than a tuning problem.
Is a 90% containment rate realistic?Not across a mixed workload. Nobody credible in the published data is at 90%. Ask what the denominator is and whether abandoned conversations count.
What is a normal escalation rate?20–35%, with 20–25% considered strong. Below 15% usually means the agent is holding conversations it should hand over.
How long until an AI agent performs well?Six to twelve months to reach the 60%+ band, with the steepest gains in months two through four — and only if somebody reviews escalations weekly.
What re-contact rate is acceptable?11.3% on AI-resolved conversations against 8.7% human-resolved (Zendesk). Above 15% means your containment figure is overstating genuine resolution.
Why is my AI agent underperforming the vendor’s claim?Usually because their number was deflection and yours is resolution, or because nobody has tuned the deployment since launch. Both are fixable.
NextStart free

The gap between 41% and 70% is six months of somebody paying attention

That is the finding underneath every number on this page. The top quartile did not buy a better model — they started earlier and read their escalation logs. The steepest gains land in months two through four, which means the cost of waiting a quarter is not zero. It is the quarter.

Point Jugl at your website, catalogue and policies, connect a channel, and start generating your own numbers instead of comparing other people's. The free tier is permanent, needs no card, and gives you the only benchmark that ever really mattered: your own transcripts.

WhatsApp, Instagram, Facebook, web chat, email and SMSOne agent, one brain, one shared customer historyBooks, sells, looks up orders, qualifies leadsTrained on your data — your prices, your policiesFull-context handover to a real human, by designPublished flat tiers — nothing metered per resolution

Anyone can quote a containment rate. Very few will show you the denominator.

SOC 2 Type 2 · HIPAA compliant · Meta Business Partner · NVIDIA Inception · 1000+ businesses

Keep reading

AI customer service statistics42 sourced numbers on resolution, cost, trust and channels.How to measure AI agent performanceThe operational version: what to instrument and review weekly.AI-to-human handoffThe routing decision that decides whether CSAT reads 4.3 or 3.34.AI customer service pricingWhy per-resolution billing taxes the improvement you are trying to make.Best AI agent for businessA weighted vendor scorecard and the five ways this purchase goes wrong.AI agent vs chatbotAction capability, which is the difference between deflection and resolution.Connecting your knowledgeGrounding is worth ~85% more accuracy. Here is how it actually works.WhatsApp Business statistics38 sourced numbers on the channel where most inbound now arrives.Intercom alternativesWhat $0.99 per resolution costs once the AI starts succeeding.What is Jugl?The full product overview — capabilities, fit, pricing, and who should walk away.Free conversation auditWhere your current setup is losing conversations, mapped before you buy.Jugl pricingFour published flat tiers with the AI included. Free forever, no card.

Sources: Zendesk CX benchmark data (CSAT by intent tier, re-contact rates); Gartner customer service and support research (self-service resolution, cost per contact, agentic resolution projections); Aissist.io AI service benchmark (tier-1 automation median and quartiles, deflection definition); Lorikeet AI customer service analysis (first-contact resolution and cost per resolution for AI-native platforms); Salesforce State of Service (loaded agent cost, case resolution projections). Benchmark bands are compiled from these sources and are directional rather than certified; definitions of deflection, containment and resolution differ between them and are distinguished explicitly in the text. This report is published by Jugl, which sells an AI customer agent and is therefore an interested party — Jugl's own outcome figures are customer-reported and typical rather than guaranteed, and are stated in the same table bands as everyone else's. Scorer outputs are diagnostic estimates generated from your own inputs, not quotes, forecasts or guarantees. Meta, WhatsApp, Messenger, Instagram and Facebook are trademarks of Meta Platforms, Inc.; Jugl is a Meta Business Partner and this page is published by Jugl and is not endorsed by or affiliated with Meta Platforms, Inc. All other product names are trademarks of their respective owners.

Start free at Jugl · No card required · Permanent free tier