AI in customer support: Copilot, AI agents, and what actually works
Past the hype: what Copilot and AI agents each actually do, what neither should be trusted to do alone, and how to tell a genuinely useful vendor from a demo that only works in the demo.
- The state of AI in support
- Two different things
- What AI should not do alone
- Why grounding matters
- Evaluating a vendor
- Where the ROI actually is
The state of AI in support
Nearly every support vendor now claims "AI-powered" somewhere on their homepage. Underneath the marketing, the actual capability varies enormously — from a genuinely useful drafting assistant to a chatbot that loops customers through the same three unhelpful answers. The label tells you almost nothing; the architecture underneath does.
Two different things
Copilot assists a human: it summarizes threads, drafts replies, and answers questions from your knowledge base, but nothing reaches a customer until an agent approves it. AI agents are a different product entirely — configured, autonomous actors that reply to customers directly, with no human approving each individual message. Buying one when you needed the other is the single most common source of disappointment we hear from teams evaluating AI support tools. Our Copilot vs. AI agents buyer's guide goes deeper on picking the right one.
What AI should not do alone
Refunds, account closures, policy exceptions, and anything irreversible should keep a human in the loop until a team has real evidence — not just vendor confidence — that it is safe to remove one. This is not a temporary limitation waiting for a model upgrade; it is a boundary about where autonomous judgment is appropriate at all, and a good AI agent architecture treats it as a hard rule enforced by scoped tools, not a soft guideline a prompt can override.
Why grounding matters
A responsible AI answer comes from your knowledge base and your actual data — order status, documented policy — not from a model's general training, which is a plausible-sounding guess dressed up as an answer. This is what separates an AI agent that says "your order ships in 3-5 business days" because that is your real policy, from one that says it because it is a common sentence about shipping in general. See our note on RAG in the glossary for the mechanism behind this.
Evaluating a vendor
- Is this grounded in our own knowledge base, or generating from general training data?
- Can we simulate a flow before it goes live, and see exactly what it will say?
- Is there a mandatory handoff to a human, and does it carry a summary — or does the customer start over?
- Are versions tracked, so a bad change can be rolled back?
- What happens when the model is not confident — does it guess, or hand off?
- Is customer PII masked or redacted before it reaches the model?
Where the ROI actually is
Two separate effects, usually conflated into one vague "AI saves time" claim: full end-to-end resolution for the questions that do not need a human (deflection), and reduced handle time on the tickets a human still handles (Copilot assist). They compound, but they come from different mechanisms and should be measured separately. Our ROI calculator models both against your own ticket volume and handle time, rather than a generic industry benchmark.
See this running in Samvaads
Everything described here is a real, shipping part of the product — not a roadmap item.
