Buyer’s guide · Written by a vendor who would rather you asked
What questions should you ask an AI agent vendor before buying?
Start with three. What is the written success metric, and who agreed to it? What data and tools does the agent need, and does it have that access today? When it fails, who notices, who owns it, and how fast can somebody roll it back? A vendor who cannot answer all three in plain language is selling a repackaged chatbot.
Gartner estimates only about 130 of the thousands of agentic AI vendors are real, and predicts over 40% of agentic AI projects will be cancelled — mostly for reasons a good demo would have exposed. The gap between what is marketed and what exists is wider in this category than in any adjacent one.
Below is the full 25-question checklist, what a good answer actually sounds like, the eight red flags worth walking away from, and a scorecard you can fill in during the call. It also says plainly when you do not need an agent at all.
By Jugl16 min readInteractive vendor scorecard29 questions answered
The 60-second version
Ask three questions first: what is the written success metric and who agreed to it; what data and tools does the agent need and does it have that access today; and when it fails, who notices, who owns it and how fast can somebody roll it back. Gartner estimates only around 130 of thousands of agentic AI vendors are genuine, and predicts over 40% of agentic AI projects will be cancelled.
The practical test for agent washing: an assistant retrieves and relays; an agent takes action in the systems that determine whether the issue is resolved. Can it issue the refund, or only explain the refund policy?
The one question that separates honest vendors from optimistic ones: what is your median deflection rate across customers, not your best case? Median tier-1 automation is 41.2% and top quartile 58.7%. Anyone claiming 90% is routing hard tickets out at triage or counting article views.
Sometimes you do not need an agent. Gartner’s own analyst notes many use cases positioned as agentic do not require agentic implementations. If you are answering FAQs from a help centre, a retrieval system is cheaper and more predictable.
- What agent washing is, and why it matters
- The due diligence picture at a glance
- The three questions that matter most
- Questions 1–5: is it actually an agent?
- Questions 6–10: what is the honest number?
- Questions 11–16: what will I actually pay?
- Questions 17–20: what does this cost me in time?
- Questions 21–25: security and governance
- Score the vendor in front of you
- The red flags
- Which questions matter most for your situation
- What to do before the demo
- The five questions behind every vendor evaluation
- How Jugl answers these questions
- Methodology and disclosure
- FAQ — 21 questions answered
- People also ask
Definition
What is agent washing?
Agent washing is Gartner’s term for rebranding existing products — AI assistants, robotic process automation, chatbots — as “AI agents” without substantive agentic capability. Gartner estimates only around 130 of the thousands of vendors marketing agentic AI are genuinely doing it, and forecasts that over 40% of agentic AI projects will be cancelled, attributing this to escalating costs, unclear business value and inadequate risk controls. The practical test is behavioural rather than technical: an assistant retrieves and relays information, while an agent takes action in the systems that determine whether an issue is resolved. Can it issue the refund, or only explain the refund policy? Agentic deployments show around 33% higher deflection than retrieval-only systems and push first-contact resolution from 55–70% to 70–85%.
Definition maintained by the Jugl Editorial Team. Jugl sells an AI customer agent platform and is an interested party; this page states that many use cases positioned as agentic do not require an agent, and names where Jugl is not the right fit.
The point that cuts both ways
Anushree Verma, the Gartner analyst behind the cancellation forecast, notes that most agentic AI propositions “lack significant value or ROI”, and — this is the part honest buyers should hold onto — that many use cases positioned as agentic today do not require agentic implementations at all.
Sometimes you do not need an agent. If your requirement is answering frequently asked questions from a help centre, a good retrieval system will do it more cheaply and more predictably, with fewer integration points and less to go wrong. Paying agent prices for assistant work is its own failure mode, and in our experience it is more common than the reverse. The distinction is set out in full on AI agent vs chatbot.
- ✓It can demonstrate an action in a connected system, live, in your demo
- ✓The vendor volunteers a median rather than a best case
- ✓They describe a degraded-but-useful mode if an integration is not ready
- ✓They can name their worst intent types without being pushed
- ✓They have an audit trail and can show what the agent did and why
- ✓They name what is still evolving rather than claiming everything is solved
- ×Capability questions get answered by naming the underlying model
- ×Only best-case deflection figures, repeated when you ask for the median
- ×No permission scoping model — they have not thought about blast radius
- ×Model-level safety presented as the primary control
- ×Will not quote against your peak month
- ×No documented rollback path or kill switch owner
The due diligence picture at a glance
At a glance
- The three questions that matter most
- Written success metric · integration access today · rollback when it fails
- Agentic AI projects predicted cancelled
- Over 40% (Gartner)
- Agentic AI vendors Gartner considers real
- ~130 of thousands
- CEOs seeing both revenue growth and cost reduction from AI
- 12% (PwC)
- CEOs seeing no significant financial benefit yet
- 56% (PwC)
- Typical cost overrun against estimate
- 2–3×
- Median tier-1 deflection rate
- 41.2%
- Top-quartile deflection
- 58.7%
- Deflection lift, agentic vs retrieval-only
- ~33%
- First-contact resolution: retrieval vs agentic
- 55–70% vs 70–85%
- Organisations reporting AI agent security incidents
- 88%
- Agents going live with full security approval
- 14.4%
- Organisations with no human-in-the-loop oversight
- 41–44%
- Lacking purpose binding, kill switches or network isolation
- 55–63%
- The practical agent test
- Can it issue the refund, or only explain the policy?
- Do before any demo
- Write the metric · pull 100–200 real conversations · model peak month
- When you do not need an agent
- FAQ answering from a help centre — retrieval is cheaper and steadier
The three questions that matter most
These work because they are impossible to answer with marketing language. Vendors building real systems will have answers. The ones repackaging a chatbot will change the subject to the model.
Capability — is it actually an agent?
- Can the agent take actions in my systems, or only retrieve information? Show me one.
- Which of my systems can it read from, and which can it write to?
- What happens when it is unsure? Walk me through the escalation live.
- Can it handle a multi-step request end to end — cancel, refund, reorder?
- What does it do with a request outside its trained domain?
Question 1 is the whole category test, and it should be answered by a demonstration rather than a description. Ask them to show one action in a sandbox connected to a system resembling yours — and watch specifically what happens when the API returns an error, because that is where the difference between a product and a prototype shows up.
Question 3 is the one we would most want you to ask us. Escalation quality is the difference between a 5–10 point satisfaction penalty and effective parity, and it is almost never shown unprompted in a demo. Ask for a person, count the steps, and then look at what the human receives — the design detail is on the handoff guide.
Performance — what is the honest number?
- What is the median deflection or resolution rate across your customers — not your best case?
- What is the rate at 30 days versus 12 months?
- Which intent types does it handle worst?
- How do you define “resolved”? Does a customer abandoning count?
- Can I see CSAT split by AI-handled versus escalated conversations?
Question 9 is the quiet one. If a customer abandoning a conversation counts as a resolution, every number the vendor has given you means something different from what you assumed — and you will be optimising for making people give up. At roughly 2.3 contacts per issue, that costs more than no automation at all. The measurement framework is on the ROI measurement page.
Cost — what will I actually pay?
- What is my total bill at 1.5× my average monthly volume?
- Does an AI-resolved conversation also consume ticket or conversation quota?
- Are AI conversations metered separately from base conversations?
- What are the real seat limits on this tier — not the marketing number?
- What are the overage rates, and when do they trigger?
- What is included in setup, and what is billed separately?
Costs on failed projects commonly balloon two to three times beyond estimates, and the mechanism is almost always metering behaviour that only becomes visible at volume. Model your peak month, not your average. A vendor who will not quote at peak knows the number is ugly.
Question 12 catches the most common double-charge in this category, and it is easy to miss in a quote: an AI-resolved conversation consuming ticket quota means you pay for the automation and for the ticket it replaced. The four pricing models and how each behaves at volume are decoded on the pricing page, and setup specifically on the setup cost page.
Implementation — what does this cost me in time?
- How many hours of my team’s time does deployment require?
- What does it train on, and who prepares that content?
- What happens if my documentation is out of date?
- Who owns this after launch, and how much maintenance does it need?
Question 17 uncovers the largest line item that appears on no invoice: budget 15–40 hours for a support deployment and 40–120 for sales. Question 19 is the one vendors most often answer optimistically — a stale knowledge base caps deflection at 40–55% regardless of platform, and a vendor who says otherwise is either not measuring or not telling you.
Question 20 is the difference between a deployment that climbs from 45% to 60% resolution and one that stays at 45%. A vendor who tells you maintenance is unnecessary is describing the thing that will cap your return. The seven-step method, including what the review actually involves, is on the training guide.
Security and governance
- What can the agent access, and how is that scoped?
- Is there an audit trail? Can I see what it did and why?
- Is there a kill switch, and who can pull it?
- Where is conversation data stored, for how long, and who can see it?
- Will you sign a BAA or DPA? (Non-negotiable if you handle health or regulated data.)
Context for why this section matters: 88% of organisations reported confirmed or suspected AI agent security incidents in the past year, and only 14.4% of agents go live with full security approval. For most mid-market businesses the exposure is not a sophisticated attack — it is an over-permissioned agent deployed without review, where somebody granted write access to make a demo work and nobody removed it.
Question 22 is the one to insist on at any company size. Without an audit trail you cannot investigate anything, and “we do not know what it did” is a considerably worse position to be in than “it did something wrong.” If you handle health data, question 25 comes before everything else on this page — see the healthcare guide.
Score the vendor in front of you
Eight inputs — five capability groups and the three hard gates. Fill it in during the call rather than afterwards, because the answers get more generous in memory. Outputs are a structured opinion generated from your own scoring, not a rating of any vendor.
Score the vendor in front of you
Five capability groups, three hard gates, and the red flags that are firing
Questions 1–5. Did they demonstrate an action in a system, live, or only describe one? An agent issues the refund; an assistant explains the refund policy.
Questions 6–10. Did they give you a median rather than a best case, name their worst intent types, and define what counts as resolved?
Questions 11–16. Did they quote at 1.5× your average volume, and explain metering, seat limits and overage triggers without being chased?
Questions 17–20. Did they name the hours of your team's time honestly, and say what happens if your documentation is out of date?
Questions 21–25. Permission scoping, audit trail, kill switch, data location and retention, and whether they will sign a BAA or DPA.
A number, a baseline, a date, and a named owner on your side. Projects without this cannot be evaluated, so they drift until somebody cancels them.
When it fails — and it will — who notices, who owns the outcome, and how fast can somebody roll it back? If the vendor has not thought about this, you will build it during an incident.
Their median deflection or resolution rate across customers, not their best case. This is the single question that separates honest vendors from optimistic ones.
The red flags
Walk away, or at least slow down, if a vendor:
Which questions matter most for your situation
| If you are… | Prioritise questions |
|---|---|
| A small business with low volume | 11–16 (cost) and 17–20 (your own time) |
| Replacing an existing tool | 6–10 (performance) and 25 (exit path) |
| In a regulated industry | 21–25 (security), plus BAA or DPA before anything else |
| Buying primarily for sales | 1–5 (can it act), and 2 specifically for catalogue and CRM access |
| Scaling an existing deployment | 7 (the 12-month rate) and 20 (maintenance ownership) |
The weighting matters because the checklist is long enough that asking all of it evenly produces a tired conversation and a shallow read. Pick your five, ask them properly, and use the rest as a written follow-up. Vendors who answer written questions carefully are telling you something about what support will feel like later.
What to do before the demo
The five questions behind every vendor evaluation
How do I know if it is a real agent?
Short answer
One test: can it take an action in your systems, or only retrieve information? An agent issues the refund; an assistant explains the refund policy. Agentic deployments show around 33% higher deflection and push first-contact resolution from 55–70% to 70–85%.
Example
What performance number should I insist on?
Short answer
The median deflection or resolution rate across their customers, not the best case. Median tier-1 automation is 41.2% and top quartile 58.7%. Anyone claiming 90% on general volume is routing hard tickets out at triage or counting self-service article views.
Example
How do I avoid the cost overrun?
Short answer
Quote at 1.5× your average monthly volume, not at average. Costs on failed projects commonly balloon two to three times beyond estimates, and the cause is almost always metering behaviour that only becomes visible at volume.
Example
What happens when it goes wrong?
Short answer
Ask who notices, who owns the outcome, and how fast somebody can roll it back. Failure is a certainty rather than a risk — 41–44% of organisations have no human-in-the-loop oversight and 55–63% lack kill switches, purpose binding or network isolation.
Example
Do I even need an agent?
Short answer
Possibly not. Gartner's own analyst notes many use cases positioned as agentic do not require agentic implementations. If you are answering FAQs from a help centre, a good retrieval system is cheaper, more predictable and has fewer integration points to fail.
Example
How Jugl answers these questions
We would rather you asked us the hard ones than found out later. Here is where we stand on each group, stated as a disclosure rather than a claim — this is our own product and you should read it that way.
On capability (Q1–5). Jugl’s agents do not just answer — they qualify leads, recommend products, schedule appointments and detect buying intent inside the conversation. On anything requiring judgment, the agent escalates to a real human with the full conversation attached, so nothing is repeated. Ask us to demo that escalation live; it is the part we would most like you to see, and the part most vendors skip.
On implementation (Q17–20). Jugl trains on your existing website, documents and past conversations, and goes live across WhatsApp, Instagram, Facebook, web chat and email without a developer. You are organising content you already have rather than commissioning a project — but the content audit is still yours, and a stale knowledge base will cap your resolution rate on our platform exactly as it would on anybody else’s.
On cost (Q11–16). AI is included rather than metered as a separate per-resolution fee on top of a conversation charge, which means the bill does not rise as the deployment succeeds. Ask us to quote against your peak month — we would rather do that than have you discover it in November. Pricing is published on the pricing page and worked bills against named competitors are on the comparison hub.
Where we will tell you we are not the fit. If you need phone answering, deep POS or property management system integration, or PHI handling without a BAA in place, we will say so. Jugl is a messaging and chat platform — that is the thing it does properly. Bring your three big questions and your 100 real historical conversations. If Jugl does not handle them well, you should know before you buy, and we would rather lose a bad-fit deal than win one. Jugl is used by 1,000+ businesses and is a Meta Business Partner.
Methodology and disclosure
Written by
Jugl Editorial TeamJugl Inc., Frisco, Texas — an AI customer agent platform used by 1,000+ businesses.
Reviewed by
Jugl product & customer operationsChecked against live deployment data and current vendor documentation.
Methodology & disclosure
Where the figures come from. The agent washing definition, the estimate of genuine agentic vendors, the project cancellation forecast and the attribution of its causes are Gartner, including published commentary from the analyst behind the forecast. CEO figures on revenue growth and cost reduction from AI are PwC’s Global CEO Survey. Median and top-quartile tier-1 deflection rates, the agentic versus retrieval deflection lift and first-contact resolution ranges are from published enterprise CX research and programme analysis. AI agent security incident rates, security approval rates, human-in-the-loop oversight gaps and control gaps are from published AI governance research. Cost overrun ranges are from published implementation analysis. Jugl pricing is our own published price list.
How the scorecard works. Five capability groups are weighted — capability 26%, performance honesty 24%, cost transparency 18%, security and governance 18%, implementation realism 14% — and scaled to 70 points. The three hard gates (written success metric, rollback path, median rate disclosed) contribute the remaining 30, weighted 12, 9 and 9 respectively, because each is a disqualifier rather than a preference. Red flags fire on any group scoring two or below and on any failed gate. The output is a structured opinion generated from your own scoring of a conversation you had — it is not a rating of any vendor, ours included, and we have no visibility into what you enter.
Conflict of interest, stated plainly. Jugl sells an AI customer agent platform, so a page telling you how to interrogate AI agent vendors is published by one of the vendors you would be interrogating. Three things are included specifically because they cut against that interest: the page states that many use cases positioned as agentic do not require an agent at all, which is an argument against buying our category; it names the situations in which Jugl is not the right fit; and the checklist is written to be used on us as readily as on anybody else.
How this page is maintained. Reviewed against current published research and revised when sources update. Deliberately evergreen — no publish date and no year stamps — because a dated buyer’s checklist misleads the moment it ages, while the questions that expose a repackaged chatbot have not changed.
Evaluating AI agent vendors: 21 questions answered
What questions should I ask an AI agent vendor before buying?
What is agent washing, and why does it matter?
What is the practical test for whether something is really an agent?
Do I always need an agent rather than a chatbot?
Why is the written success metric the most important question?
What should I ask about integration access?
What should I ask about failure and rollback?
What is the single question that separates honest vendors from optimistic ones?
How should I ask about cost?
What should I ask about implementation time?
What are the security and governance questions?
Should a small business worry about AI agent security?
What are the red flags in a vendor demo?
Which questions should I prioritise for my situation?
What should I do before the demo?
How do I test escalation in a demo?
What does a good answer sound like on the performance question?
Should I run a paid pilot or a free trial?
How do I compare quotes across vendors fairly?
What if the vendor cannot answer some of these questions?
How does Jugl answer these questions?
People also ask
Test it before you talk to anyone
Set up a free agent trained on your own content, run last month’s real questions through it, and go into every vendor demo — including ours — knowing what good looks like for your business. A buyer with a baseline asks completely different questions from a buyer with a wish list, and the difference shows up in the first five minutes.
Bring the three big questions. Bring your hundred real historical conversations. If a platform does not handle them well, you should find that out in an afternoon rather than in month three of a contract — and that is as true of us as of anybody else on your shortlist.
Over 40% of agentic AI projects are predicted to be cancelled, mostly for reasons a good demo would have exposed. Three questions and an afternoon is what it takes to not be one of them.
SOC 2 Type 2 · HIPAA compliant · Meta Business Partner · NVIDIA Inception · 1000+ businesses
Keep reading
Sources: Gartner (the agent washing definition, the estimate of genuinely agentic vendors, the forecast that over 40% of agentic AI projects will be cancelled, and published analyst commentary on causes and on use cases that do not require agentic implementations); PwC’s Global CEO Survey (CEOs reporting revenue growth and cost reduction from AI, and those reporting no significant financial benefit); published enterprise CX research and programme analysis (median and top-quartile tier-1 deflection, the agentic versus retrieval-only deflection lift, and first-contact resolution ranges); published AI governance research (AI agent security incident rates, security approval rates, human-in-the-loop oversight gaps, and gaps in purpose binding, kill switches and network isolation); published implementation analysis (cost overrun ranges); and Jugl’s published price list. This page is published by Jugl, which sells an AI customer agent platform and is therefore one of the vendors this checklist is designed to interrogate; it states that many use cases positioned as agentic do not require an agent at all, and names the situations in which Jugl is not the right fit. Scorecard outputs are a structured opinion generated from your own scoring, not a rating of any vendor. Jugl’s outcome figures are customer-reported and typical rather than guaranteed. Meta, WhatsApp, Messenger, Instagram and Facebook are trademarks of Meta Platforms, Inc.; Jugl is a Meta Business Partner and this page is published by Jugl and is not endorsed by or affiliated with Meta Platforms, Inc. All other product names are trademarks of their respective owners.
Start free at Jugl · No card required · Permanent free tier