25 Questions to Ask an AI Agent Vendor Before Buying | Jugl CX
$5mn in seed funding raised, built bootstrapped from day one
JuglCX

Buyer’s guide · Written by a vendor who would rather you asked

What questions should you ask an AI agent vendor before buying?

Start with three. What is the written success metric, and who agreed to it? What data and tools does the agent need, and does it have that access today? When it fails, who notices, who owns it, and how fast can somebody roll it back? A vendor who cannot answer all three in plain language is selling a repackaged chatbot.

Gartner estimates only about 130 of the thousands of agentic AI vendors are real, and predicts over 40% of agentic AI projects will be cancelled — mostly for reasons a good demo would have exposed. The gap between what is marketed and what exists is wider in this category than in any adjacent one.

Below is the full 25-question checklist, what a good answer actually sounds like, the eight red flags worth walking away from, and a scorecard you can fill in during the call. It also says plainly when you do not need an agent at all.

By Jugl16 min readInteractive vendor scorecard29 questions answered

Short answerFor AI overviews

The 60-second version

Ask three questions first: what is the written success metric and who agreed to it; what data and tools does the agent need and does it have that access today; and when it fails, who notices, who owns it and how fast can somebody roll it back. Gartner estimates only around 130 of thousands of agentic AI vendors are genuine, and predicts over 40% of agentic AI projects will be cancelled.

The practical test for agent washing: an assistant retrieves and relays; an agent takes action in the systems that determine whether the issue is resolved. Can it issue the refund, or only explain the refund policy?

The one question that separates honest vendors from optimistic ones: what is your median deflection rate across customers, not your best case? Median tier-1 automation is 41.2% and top quartile 58.7%. Anyone claiming 90% is routing hard tickets out at triage or counting article views.

Sometimes you do not need an agent. Gartner’s own analyst notes many use cases positioned as agentic do not require agentic implementations. If you are answering FAQs from a help centre, a retrieval system is cheaper and more predictable.

01Definition

Definition

What is agent washing?

Agent washing is Gartner’s term for rebranding existing products — AI assistants, robotic process automation, chatbots — as “AI agents” without substantive agentic capability. Gartner estimates only around 130 of the thousands of vendors marketing agentic AI are genuinely doing it, and forecasts that over 40% of agentic AI projects will be cancelled, attributing this to escalating costs, unclear business value and inadequate risk controls. The practical test is behavioural rather than technical: an assistant retrieves and relays information, while an agent takes action in the systems that determine whether an issue is resolved. Can it issue the refund, or only explain the refund policy? Agentic deployments show around 33% higher deflection than retrieval-only systems and push first-contact resolution from 55–70% to 70–85%.

Definition maintained by the Jugl Editorial Team. Jugl sells an AI customer agent platform and is an interested party; this page states that many use cases positioned as agentic do not require an agent, and names where Jugl is not the right fit.

The point that cuts both ways

Anushree Verma, the Gartner analyst behind the cancellation forecast, notes that most agentic AI propositions “lack significant value or ROI”, and — this is the part honest buyers should hold onto — that many use cases positioned as agentic today do not require agentic implementations at all.

Sometimes you do not need an agent. If your requirement is answering frequently asked questions from a help centre, a good retrieval system will do it more cheaply and more predictably, with fewer integration points and less to go wrong. Paying agent prices for assistant work is its own failure mode, and in our experience it is more common than the reverse. The distinction is set out in full on AI agent vs chatbot.

Signs you are looking at a real agent
  • It can demonstrate an action in a connected system, live, in your demo
  • The vendor volunteers a median rather than a best case
  • They describe a degraded-but-useful mode if an integration is not ready
  • They can name their worst intent types without being pushed
  • They have an audit trail and can show what the agent did and why
  • They name what is still evolving rather than claiming everything is solved
Signs you are looking at agent washing
  • Capability questions get answered by naming the underlying model
  • Only best-case deflection figures, repeated when you ask for the median
  • No permission scoping model — they have not thought about blast radius
  • Model-level safety presented as the primary control
  • Will not quote against your peak month
  • No documented rollback path or kill switch owner
02At a glance

The due diligence picture at a glance

At a glance

The three questions that matter most
Written success metric · integration access today · rollback when it fails
Agentic AI projects predicted cancelled
Over 40% (Gartner)
Agentic AI vendors Gartner considers real
~130 of thousands
CEOs seeing both revenue growth and cost reduction from AI
12% (PwC)
CEOs seeing no significant financial benefit yet
56% (PwC)
Typical cost overrun against estimate
2–3×
Median tier-1 deflection rate
41.2%
Top-quartile deflection
58.7%
Deflection lift, agentic vs retrieval-only
~33%
First-contact resolution: retrieval vs agentic
55–70% vs 70–85%
Organisations reporting AI agent security incidents
88%
Agents going live with full security approval
14.4%
Organisations with no human-in-the-loop oversight
41–44%
Lacking purpose binding, kill switches or network isolation
55–63%
The practical agent test
Can it issue the refund, or only explain the policy?
Do before any demo
Write the metric · pull 100–200 real conversations · model peak month
When you do not need an agent
FAQ answering from a help centre — retrieval is cheaper and steadier
SOC 2 Type 2certified
HIPAAcompliant
MetaBusiness Partner
1,000+businesses
03The big three

The three questions that matter most

~130agentic AI vendors Gartner calls real
40%+of agentic projects predicted cancelled
2–3×typical cost overrun on failed projects
14.4%of agents live with full security approval

These work because they are impossible to answer with marketing language. Vendors building real systems will have answers. The ones repackaging a chatbot will change the subject to the model.

1
“What is the written success metric, and who agreed to it?”Not “improve customer experience.” A number, a baseline, a date, and a named owner on your side who signed up to it. Projects without this cannot be evaluated, so they drift until somebody cancels them. This is the single most common cause of the 40%.
2
“What data and tools does the agent need — and does it have that access today?”Most complexity failures are integration failures wearing a costume: the AI understood the request perfectly and had no way to act on it. Ask the follow-up nobody asks — what breaks if the integration is not ready on day one?
3
“When it fails, who notices, who owns it, and how fast can somebody roll it back?”Failure is a certainty, not a risk. 41–44% of organisations have no human-in-the-loop oversight, and 55–63% lack purpose binding, kill switches or network isolation. If the vendor has not thought about rollback, you will be building it during an incident.
04Questions 1–5

Capability — is it actually an agent?

Questions 1–5
  • Can the agent take actions in my systems, or only retrieve information? Show me one.
  • Which of my systems can it read from, and which can it write to?
  • What happens when it is unsure? Walk me through the escalation live.
  • Can it handle a multi-step request end to end — cancel, refund, reorder?
  • What does it do with a request outside its trained domain?

Question 1 is the whole category test, and it should be answered by a demonstration rather than a description. Ask them to show one action in a sandbox connected to a system resembling yours — and watch specifically what happens when the API returns an error, because that is where the difference between a product and a prototype shows up.

Question 3 is the one we would most want you to ask us. Escalation quality is the difference between a 5–10 point satisfaction penalty and effective parity, and it is almost never shown unprompted in a demo. Ask for a person, count the steps, and then look at what the human receives — the design detail is on the handoff guide.

05Questions 6–10

Performance — what is the honest number?

Questions 6–10
  • What is the median deflection or resolution rate across your customers — not your best case?
  • What is the rate at 30 days versus 12 months?
  • Which intent types does it handle worst?
  • How do you define “resolved”? Does a customer abandoning count?
  • Can I see CSAT split by AI-handled versus escalated conversations?
Question 6 separates honest vendors from optimistic ones. Median tier-1 deflection across enterprise programmes sits at 41.2%, with a top quartile of 58.7%. Anyone claiming 90% on general volume is either routing hard tickets out at triage — so the denominator excludes the difficult work — or counting self-service article views as resolutions. Ask which. If you do not get a median after asking three times, assume it is bad.

Question 9 is the quiet one. If a customer abandoning a conversation counts as a resolution, every number the vendor has given you means something different from what you assumed — and you will be optimising for making people give up. At roughly 2.3 contacts per issue, that costs more than no automation at all. The measurement framework is on the ROI measurement page.

06Questions 11–16

Cost — what will I actually pay?

Questions 11–16
  • What is my total bill at 1.5× my average monthly volume?
  • Does an AI-resolved conversation also consume ticket or conversation quota?
  • Are AI conversations metered separately from base conversations?
  • What are the real seat limits on this tier — not the marketing number?
  • What are the overage rates, and when do they trigger?
  • What is included in setup, and what is billed separately?

Costs on failed projects commonly balloon two to three times beyond estimates, and the mechanism is almost always metering behaviour that only becomes visible at volume. Model your peak month, not your average. A vendor who will not quote at peak knows the number is ugly.

Question 12 catches the most common double-charge in this category, and it is easy to miss in a quote: an AI-resolved conversation consuming ticket quota means you pay for the automation and for the ticket it replaced. The four pricing models and how each behaves at volume are decoded on the pricing page, and setup specifically on the setup cost page.

07Questions 17–20

Implementation — what does this cost me in time?

Questions 17–20
  • How many hours of my team’s time does deployment require?
  • What does it train on, and who prepares that content?
  • What happens if my documentation is out of date?
  • Who owns this after launch, and how much maintenance does it need?

Question 17 uncovers the largest line item that appears on no invoice: budget 15–40 hours for a support deployment and 40–120 for sales. Question 19 is the one vendors most often answer optimistically — a stale knowledge base caps deflection at 40–55% regardless of platform, and a vendor who says otherwise is either not measuring or not telling you.

Question 20 is the difference between a deployment that climbs from 45% to 60% resolution and one that stays at 45%. A vendor who tells you maintenance is unnecessary is describing the thing that will cap your return. The seven-step method, including what the review actually involves, is on the training guide.

08Questions 21–25

Security and governance

Questions 21–25
  • What can the agent access, and how is that scoped?
  • Is there an audit trail? Can I see what it did and why?
  • Is there a kill switch, and who can pull it?
  • Where is conversation data stored, for how long, and who can see it?
  • Will you sign a BAA or DPA? (Non-negotiable if you handle health or regulated data.)

Context for why this section matters: 88% of organisations reported confirmed or suspected AI agent security incidents in the past year, and only 14.4% of agents go live with full security approval. For most mid-market businesses the exposure is not a sophisticated attack — it is an over-permissioned agent deployed without review, where somebody granted write access to make a demo work and nobody removed it.

Question 22 is the one to insist on at any company size. Without an audit trail you cannot investigate anything, and “we do not know what it did” is a considerably worse position to be in than “it did something wrong.” If you handle health data, question 25 comes before everything else on this page — see the healthcare guide.

09The scorecard

Score the vendor in front of you

Eight inputs — five capability groups and the three hard gates. Fill it in during the call rather than afterwards, because the answers get more generous in memory. Outputs are a structured opinion generated from your own scoring, not a rating of any vendor.

Score the vendor in front of you

Five capability groups, three hard gates, and the red flags that are firing

Capability — can it act?3/5

Questions 1–5. Did they demonstrate an action in a system, live, or only describe one? An agent issues the refund; an assistant explains the refund policy.

Performance honesty3/5

Questions 6–10. Did they give you a median rather than a best case, name their worst intent types, and define what counts as resolved?

Cost transparency3/5

Questions 11–16. Did they quote at 1.5× your average volume, and explain metering, seat limits and overage triggers without being chased?

Implementation realism3/5

Questions 17–20. Did they name the hours of your team's time honestly, and say what happens if your documentation is out of date?

Security and governance3/5

Questions 21–25. Permission scoping, audit trail, kill switch, data location and retention, and whether they will sign a BAA or DPA.

Written success metric agreedNo

A number, a baseline, a date, and a named owner on your side. Projects without this cannot be evaluated, so they drift until somebody cancels them.

Rollback path existsNo

When it fails — and it will — who notices, who owns the outcome, and how fast can somebody roll it back? If the vendor has not thought about this, you will build it during an incident.

Median rate disclosedNo

Their median deflection or resolution rate across customers, not their best case. This is the single question that separates honest vendors from optimistic ones.

Vendor score42/100weighted across five groups
Red flags firing3see below
Hard gates passed0/3metric, rollback, median
VerdictWalk awaybefore signing anything
Weakest groupCapabilityfix this before the next call
3 red flags firingNo written success metric with a named owner. No rollback path or kill switch. Would not give a median deflection rate. Each of these is answerable in plain language by a vendor building a real system, and each is reliably deflected by a vendor repackaging a chatbot. Ask again and listen for whether the answer changes subject to the underlying model — which frontier model they use tells you almost nothing about whether their agent can issue a refund in your system.
Know what good looks like before the first demoSet up a free agent trained on your own content and run last month's real questions through it. You will walk into every vendor call — including ours — with a baseline instead of a wish list.
Start freeNo card required
10Red flags

The red flags

Walk away, or at least slow down, if a vendor:

Cannot describe their permission scoping modelThey have not thought about blast radius. This is not a documentation gap, it is an architecture gap.
Has no audit trail capabilityYou will not be able to investigate an incident, which means you will not be able to explain one either.
Relies on model-level safety as the primary control“The model will not do that” is not a control. It is a hope with a confidence interval.
Claims to have solved all agent security challengesNobody has. Mature builders name what is still evolving, and the naming is the credential.
Quotes only best-case deflectionAsk for the median three times. If you do not get it, assume it is bad — and assume they know it is.
Deflects capability questions toward the underlying modelWhich frontier model they use tells you almost nothing about whether their agent can issue a refund in your system.
Will not quote at your peak volumeSomebody on their side knows the number is ugly, and you will find it out in your busiest month.
Has no clear exit pathAsk about data export before signing, not after. The answer is much harder to get later.
11Prioritisation

Which questions matter most for your situation

If you are…Prioritise questions
A small business with low volume11–16 (cost) and 17–20 (your own time)
Replacing an existing tool6–10 (performance) and 25 (exit path)
In a regulated industry21–25 (security), plus BAA or DPA before anything else
Buying primarily for sales1–5 (can it act), and 2 specifically for catalogue and CRM access
Scaling an existing deployment7 (the 12-month rate) and 20 (maintenance ownership)

The weighting matters because the checklist is long enough that asking all of it evenly produces a tired conversation and a shallow read. Pick your five, ask them properly, and use the rest as a written follow-up. Vendors who answer written questions carefully are telling you something about what support will feel like later.

12Preparation

What to do before the demo

1
Pull 100–200 real historical conversationsOnes your team already resolved. Ask the vendor to run them. Vendor demos use vendor-friendly questions; yours will not be, and twenty minutes on your own hard conversations tells you more than an hour of scripted walkthrough.
2
Write your success metric down firstA number, a baseline, a date, a named owner. If you arrive without one, you will adopt theirs — and theirs will be the one their product happens to be good at.
3
Model your peak monthBring the number and ask them to quote against it. This single move prevents the most common form of cost overrun in the category, and it takes ten minutes with your own volume data.
Buyers who do these three things ask fundamentally better questions than buyers who do not — and vendors notice within about five minutes. That change in the conversation is worth more than any individual answer, because it moves you from being sold to, to evaluating.
13Direct answers

The five questions behind every vendor evaluation

How do I know if it is a real agent?

Short answer

One test: can it take an action in your systems, or only retrieve information? An agent issues the refund; an assistant explains the refund policy. Agentic deployments show around 33% higher deflection and push first-contact resolution from 55–70% to 70–85%.

Example

Ask for one action, live, in a sandbox connected to a system resembling yours — and watch what happens when the API returns an error. That moment separates a product from a prototype more reliably than any feature list.
Key takeawayGartner estimates only around 130 of thousands of agentic AI vendors are genuine. The demonstration takes five minutes and settles the question.

What performance number should I insist on?

Short answer

The median deflection or resolution rate across their customers, not the best case. Median tier-1 automation is 41.2% and top quartile 58.7%. Anyone claiming 90% on general volume is routing hard tickets out at triage or counting self-service article views.

Example

A good answer sounds specific and unflattering: “mid-forties at 30 days, low sixties at twelve months for customers running a weekly review; our worst intents are complaints and billing exceptions, which we route rather than attempt.”
Key takeawayAsk three times if you have to. Then ask how they define resolved, and whether a customer abandoning counts — that answer changes the meaning of every figure they gave you.

How do I avoid the cost overrun?

Short answer

Quote at 1.5× your average monthly volume, not at average. Costs on failed projects commonly balloon two to three times beyond estimates, and the cause is almost always metering behaviour that only becomes visible at volume.

Example

The specific trap: an AI-resolved conversation that also consumes ticket quota means you pay for the automation and for the ticket it replaced. It is invisible in a headline price and obvious in a December invoice.
Key takeawayA vendor who will not quote at your peak month knows the number is ugly. That refusal is itself the answer to the question you asked.

What happens when it goes wrong?

Short answer

Ask who notices, who owns the outcome, and how fast somebody can roll it back. Failure is a certainty rather than a risk — 41–44% of organisations have no human-in-the-loop oversight and 55–63% lack kill switches, purpose binding or network isolation.

Example

If the vendor has not thought about rollback, you will be building it during an incident, at the worst possible moment, with your customers watching. Ask who on their side can pull a kill switch, who on yours, and how long it takes to take effect.
Key takeawayInsist on an audit trail at any company size. 'We do not know what it did' is a worse position than 'it did something wrong', and it is the one you inherit without logging.

Do I even need an agent?

Short answer

Possibly not. Gartner's own analyst notes many use cases positioned as agentic do not require agentic implementations. If you are answering FAQs from a help centre, a good retrieval system is cheaper, more predictable and has fewer integration points to fail.

Example

The sequencing test: do any of your top ten intents require an action in a system — a refund, a booking change, an order lookup that writes something? If none do, you do not need an agent yet, and paying agent prices for assistant work is its own failure mode.
Key takeawayClassify your top ten intents before shortlisting anything. That classification decides the category you are buying in, and it is a decision no vendor can make for you.
14Disclosure

How Jugl answers these questions

We would rather you asked us the hard ones than found out later. Here is where we stand on each group, stated as a disclosure rather than a claim — this is our own product and you should read it that way.

On capability (Q1–5). Jugl’s agents do not just answer — they qualify leads, recommend products, schedule appointments and detect buying intent inside the conversation. On anything requiring judgment, the agent escalates to a real human with the full conversation attached, so nothing is repeated. Ask us to demo that escalation live; it is the part we would most like you to see, and the part most vendors skip.

On implementation (Q17–20). Jugl trains on your existing website, documents and past conversations, and goes live across WhatsApp, Instagram, Facebook, web chat and email without a developer. You are organising content you already have rather than commissioning a project — but the content audit is still yours, and a stale knowledge base will cap your resolution rate on our platform exactly as it would on anybody else’s.

On cost (Q11–16). AI is included rather than metered as a separate per-resolution fee on top of a conversation charge, which means the bill does not rise as the deployment succeeds. Ask us to quote against your peak month — we would rather do that than have you discover it in November. Pricing is published on the pricing page and worked bills against named competitors are on the comparison hub.

Where we will tell you we are not the fit. If you need phone answering, deep POS or property management system integration, or PHI handling without a BAA in place, we will say so. Jugl is a messaging and chat platform — that is the thing it does properly. Bring your three big questions and your 100 real historical conversations. If Jugl does not handle them well, you should know before you buy, and we would rather lose a bad-fit deal than win one. Jugl is used by 1,000+ businesses and is a Meta Business Partner.

15EEAT

Methodology and disclosure

Written by

Jugl Editorial Team

Jugl Inc., Frisco, Texas — an AI customer agent platform used by 1,000+ businesses.

Reviewed by

Jugl product & customer operations

Checked against live deployment data and current vendor documentation.

Methodology & disclosure

Where the figures come from. The agent washing definition, the estimate of genuine agentic vendors, the project cancellation forecast and the attribution of its causes are Gartner, including published commentary from the analyst behind the forecast. CEO figures on revenue growth and cost reduction from AI are PwC’s Global CEO Survey. Median and top-quartile tier-1 deflection rates, the agentic versus retrieval deflection lift and first-contact resolution ranges are from published enterprise CX research and programme analysis. AI agent security incident rates, security approval rates, human-in-the-loop oversight gaps and control gaps are from published AI governance research. Cost overrun ranges are from published implementation analysis. Jugl pricing is our own published price list.

How the scorecard works. Five capability groups are weighted — capability 26%, performance honesty 24%, cost transparency 18%, security and governance 18%, implementation realism 14% — and scaled to 70 points. The three hard gates (written success metric, rollback path, median rate disclosed) contribute the remaining 30, weighted 12, 9 and 9 respectively, because each is a disqualifier rather than a preference. Red flags fire on any group scoring two or below and on any failed gate. The output is a structured opinion generated from your own scoring of a conversation you had — it is not a rating of any vendor, ours included, and we have no visibility into what you enter.

Conflict of interest, stated plainly. Jugl sells an AI customer agent platform, so a page telling you how to interrogate AI agent vendors is published by one of the vendors you would be interrogating. Three things are included specifically because they cut against that interest: the page states that many use cases positioned as agentic do not require an agent at all, which is an argument against buying our category; it names the situations in which Jugl is not the right fit; and the checklist is written to be used on us as readily as on anybody else.

How this page is maintained. Reviewed against current published research and revised when sources update. Deliberately evergreen — no publish date and no year stamps — because a dated buyer’s checklist misleads the moment it ages, while the questions that expose a repackaged chatbot have not changed.

16FAQ

Evaluating AI agent vendors: 21 questions answered

What questions should I ask an AI agent vendor before buying?
Start with three, because they are impossible to answer with marketing language. What is the written success metric, and who agreed to it? What data and tools does the agent need, and does it have that access today? When it fails, who notices, who owns it, and how fast can somebody roll it back? If a vendor cannot answer all three in plain language, you are looking at a repackaged chatbot. Then work through the full checklist by group: capability (can it act, or only retrieve), performance honesty (median rather than best case), cost (quote at 1.5× your average volume), implementation (hours of your team's time), and security and governance (scoping, audit trail, kill switch, BAA or DPA).
What is agent washing, and why does it matter?
Gartner coined the term for a specific and widespread practice: rebranding existing products — AI assistants, RPA, chatbots — as "agents" without substantive agentic capability. The scale is remarkable. Gartner estimates only around 130 of the thousands of vendors marketing agentic AI are genuinely doing it, and predicts over 40% of agentic AI projects will be cancelled, with the analyst behind that forecast noting that most agentic AI propositions lack significant value or ROI. It matters commercially because agent pricing is materially higher than assistant pricing, and because a deployment scoped as agentic that turns out to be retrieval-only will miss its resolution target and take the blame for a capability it never had.
What is the practical test for whether something is really an agent?
An assistant retrieves and relays. An agent takes action in the systems that determine whether the issue is resolved. Can it issue the refund, or only explain the refund policy? That single question separates the two categories more reliably than any feature list, and it is answerable in a live demo rather than a document. The performance difference is measurable: agentic deployments show around 33% higher deflection than retrieval-only systems, and push first-contact resolution from 55–70% to 70–85%. Ask the vendor to show you one action, in a sandbox connected to a system resembling yours, and watch what they do when the API returns an error.
Do I always need an agent rather than a chatbot?
No, and this cuts against our own commercial interest to say. Gartner's own analyst notes that many use cases positioned as agentic today do not require agentic implementations at all. If your requirement is answering frequently asked questions from a help centre, a good retrieval system will do it more cheaply and more predictably than an agent, with fewer integration points and less to go wrong. Paying agent prices for assistant work is its own failure mode, and it is more common than the reverse. The honest sequencing question is whether any of your top ten intents require an action in a system. If none do, you do not need an agent yet.
Why is the written success metric the most important question?
Because projects without one cannot be evaluated, so they drift until somebody cancels them — and that is the single most common cause of the 40% cancellation forecast. "Improve customer experience" is not a success metric. A number, a baseline, a date, and a named owner on your side who signed up to it is. The named owner matters as much as the number: a metric nobody in your organisation has agreed to is a metric that will be reinterpreted at every review until it means whatever the current result happens to be. Write it down before the first demo, or you will adopt the vendor's framing by default.
What should I ask about integration access?
What data and tools does the agent actually need to reach, and does it have that access today? Most complexity failures are integration failures wearing a costume: the AI understood the request perfectly and had no way to act on it. The follow-up question is the one that gets skipped — ask what breaks if the integration is not ready on day one. A vendor with a real answer will describe a degraded but useful mode; a vendor without one will describe a timeline. Also ask which systems it can read from and which it can write to, because those are different permissions, different risk profiles, and frequently different phases of your project.
What should I ask about failure and rollback?
When it fails, who notices, who owns the outcome, and how fast can somebody roll it back? Failure is a certainty, not a risk, and this question exposes whether the vendor has operated a real deployment or only sold one. The context makes it urgent: 41–44% of organisations have no human-in-the-loop oversight, and 55–63% lack purpose binding, kill switches or network isolation. If the vendor has not thought about rollback, you will be building it during an incident — at the worst possible moment, with your customers watching. Ask specifically who on their side can pull a kill switch, who on yours, and how long it takes to take effect.
What is the single question that separates honest vendors from optimistic ones?
What is your median deflection or resolution rate across your customers — not your best case? Median tier-1 deflection across enterprise programmes sits at 41.2%, with a top quartile of 58.7%. Anyone claiming 90% on general volume is either routing hard tickets out at triage — so the denominator excludes the difficult work — or counting self-service article views as resolutions. Ask which. Then ask the follow-up: what is the rate at 30 days versus 12 months, since deployments launch at 40–50% and climb past 60% only with active tuning. A vendor who gives you both numbers without being chased is telling you something useful about how they operate.
How should I ask about cost?
Model your peak month, not your average, and ask for a quote at 1.5× your average monthly volume. Costs on failed projects commonly balloon two to three times beyond estimates, and the mechanism is almost always metering behaviour that only becomes visible at volume. Six specific questions: what is my total bill at 1.5× average volume; does an AI-resolved conversation also consume ticket or conversation quota; are AI conversations metered separately from base conversations; what are the real seat limits on this tier rather than the marketing number; what are the overage rates and when do they trigger; and what is included in setup versus billed separately. A vendor who will not quote at peak knows the number is ugly.
What should I ask about implementation time?
How many hours of my team's time does deployment require? That figure is routinely larger than the software cost and appears on no invoice — budget 15–40 hours for a support deployment and 40–120 for sales. Then: what does it train on, and who prepares that content? What happens if my documentation is out of date? And who owns this after launch, and how much maintenance does it need? That last one is the difference between a deployment that climbs from 45% to 60% resolution and one that stays at 45%. A vendor who says maintenance is unnecessary is describing the thing that will cap your return.
What are the security and governance questions?
Five. What can the agent access, and how is that scoped? Is there an audit trail, and can I see what it did and why? Is there a kill switch, and who can pull it? Where is conversation data stored, for how long, and who can see it? And will you sign a BAA or DPA — non-negotiable if you handle health or regulated data. The context for why this matters: 88% of organisations reported confirmed or suspected AI agent security incidents in the past year, and only 14.4% of agents go live with full security approval. These five questions take about five minutes and are worth asking at any company size.
Should a small business worry about AI agent security?
Yes, and the reason is not the one most people assume. Mid-market and small-business exposure typically comes from over-permissioned agents deployed without review rather than from sophisticated attacks — somebody grants write access to a system to make a demo work and nobody removes it afterwards. That is a governance failure with a five-minute fix and a long tail of consequences. Questions 21 to 24 take five minutes to ask and are worth it at any size. The one that matters most for a small team is the audit trail: without it you cannot investigate anything, and "we do not know what it did" is a considerably worse position than "it did something wrong".
What are the red flags in a vendor demo?
Eight. They cannot describe their permission scoping model, which means they have not thought about blast radius. They have no audit trail capability, so you will not be able to investigate an incident. They rely on model-level safety as their primary control — "the model will not do that" is not a control. They claim to have solved all agent security challenges; nobody has, and mature builders name what is still evolving. They quote only best-case deflection. They deflect capability questions toward the underlying model, which tells you almost nothing about whether their agent can act in your systems. They will not quote at your peak volume. And they have no clear exit path — ask about data export before signing, not after.
Which questions should I prioritise for my situation?
Weight the checklist by what you are actually buying. A small business with low volume should prioritise questions 11–16 on cost and 17–20 on your own time, because those are the two things that sink small deployments. Replacing an existing tool means prioritising 6–10 on performance and 25 on the exit path, since you already know the category and the risk is switching cost. In a regulated industry, 21–25 on security come first, with the BAA or DPA settled before anything else is discussed. Buying primarily for sales means 1–5 on capability and specifically question 2 on catalogue and CRM access. Scaling an existing deployment means question 7 on the 12-month rate and question 20 on maintenance ownership.
What should I do before the demo?
Three things, and they change the conversation entirely. Pull 100–200 real historical conversations your team already resolved and ask the vendor to run them — vendor demos use vendor-friendly questions, and yours will not be. Write your success metric down first, because if you arrive without one you will adopt theirs. And model your peak month, bring the number, and ask them to quote against it. Buyers who do these three things ask fundamentally better questions than buyers who do not, and vendors notice within about five minutes. It also changes what you learn: a demo on your own hard conversations tells you more in twenty minutes than a scripted walkthrough tells you in an hour.
How do I test escalation in a demo?
Ask them to break it, live. Specifically: give the agent a question outside its trained domain and watch what it does; give it an ambiguous request and see whether it asks a clarifying question or guesses; type "I want to speak to a person" and count the steps; and then look at what the human receives when the conversation transfers. That last one is the part vendors rarely volunteer, and it is the difference between a 5–10 point satisfaction penalty and effective parity. If the handoff arrives without the transcript, the agent's understanding of the problem, and what it already tried, the customer will re-explain — and you have converted one contact into two.
What does a good answer sound like on the performance question?
Specific, unflattering and volunteered. Something like: "Median across our customers is in the mid-forties at 30 days, low sixties at twelve months for customers who run a weekly escalation review. Our worst intent types are complaints and billing exceptions, which we route rather than attempt. We count a resolution as a conversation with no repeat contact within seven days, so abandonment does not count." A vendor who says that is telling you they measure honestly and know their weaknesses. A vendor who answers with a single number above 80% and no denominator is either measuring something else or hoping you will not ask what.
Should I run a paid pilot or a free trial?
A free tier is the better first step and a paid pilot is the better second one, for different reasons. Start on a free tier because it lets you test on your own content and your own questions with no commitment and no procurement process — an afternoon of that tells you more than a month of evaluation documents. Then run a paid pilot with a written success metric, a named owner and a defined end date, because a pilot without those three is an indefinite trial that nobody ever concludes. Set the stop condition in advance: what result at the end of the pilot would mean this was the wrong choice? Agreeing that before you start is what makes the pilot a decision rather than a habit.
How do I compare quotes across vendors fairly?
Normalise on twelve-month total cost at your realistic volume and resolution rate, not on headline subscription. Three things distort comparisons badly. Per-resolution metering means your bill rises as the deployment succeeds — a programme improving from 41% to 65% resolution increases that invoice by more than half. Seat limits are frequently lower than the marketing tier implies. And setup varies from zero to five figures depending on how much manual configuration the product requires. Build one spreadsheet, put every vendor through the same peak-month scenario, and include your own internal hours. The four pricing models in this category are decoded on our pricing page.
What if the vendor cannot answer some of these questions?
It depends which ones. Not knowing the exact overage rate off the top of their head is fine — ask for it in writing. Not being able to describe their permission scoping model, produce a median resolution rate, or explain what happens when the agent fails is not fine, because those are not lookups, they are architecture. The tell is what they do with the gap: a mature vendor says "I do not have that number, I will get it to you by Thursday" and then does. A vendor repackaging a chatbot changes the subject to the underlying model. Both responses are informative, and the second one is the more useful signal.
How does Jugl answer these questions?
On capability, Jugl's agents qualify leads, recommend products, schedule appointments and detect buying intent inside the conversation, and escalate to a real human with the full conversation attached on anything requiring judgment — ask us to demo that escalation live, because it is the part we would most like you to see. On implementation, Jugl trains on your existing website, documents and past conversations and goes live across WhatsApp, Instagram, Facebook, web chat and email without a developer. On cost, AI is included rather than metered as a separate per-resolution fee on top of a conversation charge, and we would rather quote against your peak month than have you discover it later. Where we are not the fit — phone answering, deep POS or PMS integration, PHI handling without a BAA — we will say so.
17People also ask

People also ask

What is agent washing?Gartner's term for rebranding existing products — AI assistants, RPA, chatbots — as "agents" without substantive agentic capability. Gartner estimates only around 130 of the thousands of vendors marketing agentic AI are genuinely doing it.
How do I know if an AI agent is real?One test: can it take an action in your systems, or only retrieve information? An agent issues the refund; an assistant explains the refund policy. Agentic deployments show around 33% higher deflection and push first-contact resolution from 55–70% to 70–85%.
What deflection rate should a vendor be able to promise?Median tier-1 automation is 41.2% and top quartile 58.7%. Treat 60–67% as strong and 80%+ as best-in-class on highly structured workloads only. Ask for the median, not the best case, and ask three times if you have to.
Do I always need an agent rather than a chatbot?No. Gartner's own analyst notes many use cases positioned as agentic do not require agentic implementations. If you are answering FAQs from a help centre, a good retrieval system is cheaper and more predictable than an agent.
What is the biggest reason AI agent projects fail?No written success metric with a named owner. Gartner attributes its forecast that over 40% of agentic AI projects will be cancelled to escalating costs, unclear business value and inadequate risk controls — all downstream of not defining success first.
What are red flags in an AI vendor demo?Cannot describe permission scoping, has no audit trail, relies on model-level safety as the primary control, quotes only best-case deflection, will not quote at your peak volume, or deflects capability questions toward the underlying model.
Should a small business ask about AI security?Yes. Mid-market exposure typically comes from over-permissioned agents deployed without review rather than sophisticated attacks. 88% of organisations reported AI agent security incidents and only 14.4% of agents go live with full security approval.
How much do AI agent projects overrun?Costs on failed projects commonly balloon two to three times beyond estimates. The usual cause is quoting against average volume rather than peak, and discovering metering behaviour — separate AI resolution charges, overages — in your busiest month.
NextStart free

Test it before you talk to anyone

Set up a free agent trained on your own content, run last month’s real questions through it, and go into every vendor demo — including ours — knowing what good looks like for your business. A buyer with a baseline asks completely different questions from a buyer with a wish list, and the difference shows up in the first five minutes.

Bring the three big questions. Bring your hundred real historical conversations. If a platform does not handle them well, you should find that out in an afternoon rather than in month three of a contract — and that is as true of us as of anybody else on your shortlist.

Free tier that stays free — no card, live the same dayTrains on your website, documents and past conversationsAsk us to demo the escalation live — it is the part we want you to seeAI included, not metered as a separate per-resolution feeWe will tell you where we are not the fitFlat published tiers, quotable against your peak month

Over 40% of agentic AI projects are predicted to be cancelled, mostly for reasons a good demo would have exposed. Three questions and an afternoon is what it takes to not be one of them.

SOC 2 Type 2 · HIPAA compliant · Meta Business Partner · NVIDIA Inception · 1000+ businesses

Keep reading

Best AI agent for businessThe category ranked, with 12 weighted checks.Measuring AI agent ROIThe four-metric framework, and why deflection alone erodes savings.AI agent vs chatbotThe nine technical differences behind the agent washing test.Do customers trust AI agents?How to test escalation design, and why it decides everything.AI customer service pricingThe four pricing models, and how each behaves at peak volume.AI setup costThe hours of your own time that appear on no invoice.Train an AI agent on your dataWhat happens if your documentation is out of date.AI-to-human handoffWhat to look for when you break the escalation on purpose.AI and complex problemsWhy most capability failures are integration failures.11 AI support mistakesThe failure modes behind the cancellation statistic.AI agents for healthcareWhy the BAA question comes before everything else.Compare JuglWorked bills against named competitors at three volumes.What is Jugl?Capabilities, fit, pricing, and who should walk away.Jugl pricingFour published flat tiers with the AI included. Free forever, no card.

Sources: Gartner (the agent washing definition, the estimate of genuinely agentic vendors, the forecast that over 40% of agentic AI projects will be cancelled, and published analyst commentary on causes and on use cases that do not require agentic implementations); PwC’s Global CEO Survey (CEOs reporting revenue growth and cost reduction from AI, and those reporting no significant financial benefit); published enterprise CX research and programme analysis (median and top-quartile tier-1 deflection, the agentic versus retrieval-only deflection lift, and first-contact resolution ranges); published AI governance research (AI agent security incident rates, security approval rates, human-in-the-loop oversight gaps, and gaps in purpose binding, kill switches and network isolation); published implementation analysis (cost overrun ranges); and Jugl’s published price list. This page is published by Jugl, which sells an AI customer agent platform and is therefore one of the vendors this checklist is designed to interrogate; it states that many use cases positioned as agentic do not require an agent at all, and names the situations in which Jugl is not the right fit. Scorecard outputs are a structured opinion generated from your own scoring, not a rating of any vendor. Jugl’s outcome figures are customer-reported and typical rather than guaranteed. Meta, WhatsApp, Messenger, Instagram and Facebook are trademarks of Meta Platforms, Inc.; Jugl is a Meta Business Partner and this page is published by Jugl and is not endorsed by or affiliated with Meta Platforms, Inc. All other product names are trademarks of their respective owners.

Start free at Jugl · No card required · Permanent free tier