Ticket classification
Route by intent, urgency, product, language, or required team. Use a human-reviewed sample to check routing errors and make sure labels do not conceal vulnerable or high-impact cases.
Don't stop here
Hand-picked guides our readers explore right after this one.
AI workflows for QBR prep, account summaries, executive narratives, risks, and follow-up actions
Read the guideStunning image generation with Midjourney prompt mastery
Read the guideBuild a compelling resume with AI assistance
Read the guideCustomer support workflow guide - United States
I would not measure support AI by how many conversations disappear. I would ask whether customers get the right answer, agents get better context, and complex cases reach a capable person with less repetition and rework.
Michael Okeje
AI service workflow and customer experience research Β· Last updated August 13, 2026
Support teams have several different jobs hiding behind the word automation. Some need faster search. Some need cleaner routing. Some need a reliable summary after a long call. Some need self-service for a narrow question. Buying one AI agent for all of them creates an evaluation problem and often gives a system too much authority too early.
I would begin with an assistive workflow that is visible, reversible, and easy to compare with the current process. Show the agent the source behind a suggested answer. Keep the final case record in the approved support platform. Give the customer a real human path. Then measure whether the change improved the outcome rather than merely lowering the number of conversations that reached an agent.
Route by intent, urgency, product, language, or required team. Use a human-reviewed sample to check routing errors and make sure labels do not conceal vulnerable or high-impact cases.
Find the approved article or policy passage behind a response. Show the source to the agent and record when content is missing, stale, contradictory, or outside the customer's plan.
Create a concise draft in the correct tone and channel. Keep the agent responsible for accuracy, empathy, policy exceptions, and the decision to send.
Turn a long thread into issue, steps tried, customer goal, evidence, promised follow-up, and next owner. Preserve uncertainty and keep the original conversation available.
Sample conversations for policy adherence, unsupported claims, missing disclosure, escalation, and customer effort. Use the result for coaching and evaluation, not automatic punishment.
Find repeated questions, article gaps, conflicting instructions, and customer terminology. A subject-matter owner still approves the replacement article and its effective date.
Handle narrow, documented intents such as status checks or basic instructions. Require authentication and authorization before any account action, and make the handoff obvious.
Summarize approved calls or visits into facts, actions, commitments, and follow-up. Check consent, retention, speaker errors, and whether the summary changed the case meaning.
Start with an approved email or help-desk assistant that drafts replies and tags requests. Add a short article library and a human review rule before exploring an autonomous chatbot.
Prioritize product-aware retrieval, bug and feature tagging, account context, and an escalation path into engineering. Measure whether suggestions are grounded in current release notes and documentation.
Start with order status, returns, delivery questions, and agent summaries. Keep payment, fraud, refund exceptions, and vulnerable-customer cases behind authentication and human approval.
Use a controlled workspace, strict data minimization, source-visible drafts, review sampling, and explicit stop conditions. A public chatbot should not be the system of record or decision-maker.
Map intents, channel costs, staffing, and escalation capacity before buying an agent. Pilot one queue and check whether the human queue becomes harder rather than smaller.
When a support leader asks me which AI tool to buy, I start with the queue, not the brand. Is the problem that agents cannot find the right article? Are tickets routed to the wrong team? Are customers repeating the same facts after a handoff? Are agents spending their shift writing summaries? Is the knowledge base stale? Each problem has a different first tool and a different failure mode.
The best early support AI is often assistive. It can retrieve an approved answer, suggest a draft, summarize a case, classify an intent, or identify an article gap while a person checks the result. That work is visible and reversible. It creates evidence before the team gives a system permission to send messages, change records, approve refunds, or make promises.
Salesforce's Seventh State of Service, based on 6,500 service professionals, reports that respondents estimated AI handled 30% of cases in 2025 and expected 50% by 2027. That is a survey estimate, not a universal measurement. Gartner's survey of 321 leaders found 20% reported AI-driven headcount reduction, while 55% reported stable staffing while handling higher volumes. The real story is augmentation, redesign, and rising expectations.
I use a short diagnostic before comparing vendors. Name the request volume and channel. Identify whether the work is retrieval, classification, drafting, decision support, or an external action. Then describe what happens when the tool is wrong. A wrong article suggestion may cost minutes. A wrong refund, account change, medical statement, or legal promise can create material harm.
Ask where the source of truth lives. If it is a knowledge base, the AI project includes article ownership, versioning, review dates, and a way to handle contradictory policy. If it is the CRM or order system, the project includes permissions, authentication, and safe tool calls. If it is in experienced agents' heads, the project includes knowledge capture rather than just a model subscription.
Finally, name the customer outcome. Deflecting more tickets is an internal metric. Resolving a delivery-status question without repeat contact is a customer outcome. Drafting replies faster is a productivity metric. Giving agents more time for complex cases without lowering quality is a service outcome. A better definition prevents the dashboard from rewarding the wrong behavior.
Triage is a reasonable first workflow because it is frequent and its errors can be inspected. An AI classifier can suggest intent, urgency, language, product, sentiment, or destination. Do not assume sentiment equals urgency or that a confident label is correct. A calm message can describe a major outage, while an angry message can concern a routine issue.
Build a label policy before switching on an AI router. Define each label in plain language, give positive and negative examples, name escalation conditions, and record what happens when confidence is low. Keep a human queue for ambiguous cases. Sample routed tickets by intent and compare the AI label with a trained reviewer's label. Investigate false routing, not only overall accuracy.
Look for systematic routing problems. Does the model send non-native English messages to a lower-priority queue? Does it confuse a billing dispute with a cancellation request? Does it treat an accessibility request as general feedback? A single accuracy number cannot answer those questions. Labels and review samples need to reflect the service's real risk.
Agent assist works best when it reduces search burden without hiding evidence. A useful suggestion shows the article or policy passage, its version where practical, and the gaps that require a human answer. If the tool gives an answer without showing the source, the agent has to trust a paragraph instead of checking policy.
Test retrieval with historical cases, including badly worded questions, customer terminology, outdated articles, and cases where two policies conflict. Score whether the right source was retrieved, whether the draft stayed within the source, whether it acknowledged missing information, and whether the next step was accurate. A fluent answer that cites the wrong policy is a high-severity error.
The knowledge base needs an owner. AI can reveal gaps and repeated questions, but it should not silently write a new policy. Route proposed article changes to a subject-matter owner, add an effective date, archive old versions carefully, and rerun the evaluation set after changes. Support AI is only as reliable as the content and permissions behind it.
A good summary gives the next agent the customer's goal, relevant account or order context, steps already tried, evidence, promises made, unresolved questions, and next action. It does not merely compress every sentence into a shorter transcript. The summary should preserve uncertainty and distinguish what the customer said from what the system inferred.
Track time to understand a transferred case, the number of times a customer repeats information, missing commitments, and correction rate. Compare AI summaries with normal handoffs. If the summary is shorter but removes the detail that explains why the customer is upset or why normal policy does not fit, it is not an improvement.
For voice support, add consent, transcription accuracy, speaker attribution, retention, and accessibility to the evaluation. A transcript can mishear a product code, name, amount, or address. The final case record should be checked by the agent where the error matters. Never let a summary become an invisible decision that no one can challenge.
A customer-facing AI agent should earn authority gradually. Start with an information request that has a current source and a clear answer. Then consider a reversible authenticated action. Keep exceptions, high-value accounts, payment disputes, identity changes, safety issues, legal questions, and emotionally sensitive cases on a human path unless the organization has stronger controls.
The handoff is part of the product. Tell the customer when they are interacting with AI, explain what happens next, preserve the conversation, and give the human agent the attempted steps and sources. A handoff that says please repeat your issue is not successful escalation. Measure cases that should have escalated earlier and the quality of context received by the human.
Do not use containment as the only success metric. Containment can rise because the system resolved more cases, or because customers gave up, left the channel, or could not reach a person. Pair it with repeat contact, complaint rate, resolution quality, customer effort, accessibility outcomes, corrections, and human queue health. The real win is the right customer outcome.
Gartner's December 2025 survey separates staffing from volume. Twenty percent of surveyed leaders reported AI-driven headcount reduction, while 55% reported stable staffing as they handled higher customer volumes. Gartner also reported that 42% of organizations were hiring specialized roles such as AI strategists, conversational AI designers, and automation analysts. That is a workforce redesign story.
When routine work moves to AI, agents may spend more time on investigation, exceptions, retention, technical diagnosis, accessibility, and emotionally difficult conversations. Those jobs require better context and tools. If the organization measures only average handling time, it may push complexity onto people without investing in training or decision rights.
Involve agents in the pilot. Ask which searches waste time, which policies are hard to interpret, which handoffs fail, and which suggestions would make the work worse. Give them a way to flag an unsafe or unsupported answer. Their corrections are evaluation data that tells the team where a model or knowledge base fails.
Before a support team sends data to an AI tool, I want five answers: what data leaves the support system, whether it is used for training, how long it is retained, who can access it, and how it can be deleted or corrected. I also want to know whether a model can call tools, what permissions those calls use, and whether every action is logged. If the vendor answer is vague, limit the workflow to public or minimized data.
Remove unnecessary payment information, authentication secrets, identity documents, health details, and other sensitive data before using a drafting or research workflow. Consent and disclosure requirements vary by context. Involve security, privacy, legal, and the service owner before launching a customer-facing system. NIST's AI RMF is a useful way to organize work around governing, mapping, measuring, and managing risk.
Trust also includes honesty. Do not say a human reviewed an answer when a human did not. Do not claim an AI agent can solve every issue if it transfers difficult cases. Do not hide a limitation behind a friendly tone. Customers accept a boundary more easily when the system explains it clearly and gives them a useful next step.
Week one is the baseline and design. Choose one intent family, such as order-status requests, password resets, or internal IT tickets. Record volume, channel, resolution, repeat contact, escalation, time, quality, and customer effort. Define source articles, excluded cases, review sample, data boundary, owner, and stop conditions.
Week two is shadow mode. Let the system classify, retrieve, or draft without sending a response or changing a record. Compare output with trained reviewers. Label errors as source error, retrieval error, instruction error, policy gap, transcription error, or human-review error. Severity matters more than a simple pass rate.
Weeks three and four are supervised assist. Agents use suggestions but review every output. Track edit distance, unsupported claims, time saved, repeat contact, escalations, and agent feedback. At the end, decide whether to expand, remain assistive, redesign, or stop. Add authentication, authorization, rollback, and an explicit human path before granting customer-facing action rights.
A small business with a shared inbox should begin with a help-desk or email assistant that tags requests and drafts replies from a short approved library. It does not need an autonomous agent on day one. The first win is a consistent response process and fewer missed messages.
A SaaS company should prioritize retrieval connected to product documentation, release notes, account context, and an engineering escalation path. A reply assistant that does not know which version a customer uses can create more support work. The tool must expose uncertainty and route bugs rather than invent workarounds.
An ecommerce team should start with order status, returns instructions, delivery information, and agent summaries. Payment disputes, fraud, refunds outside policy, and identity changes need stronger controls. A high-volume contact center should understand its intent mix and staffing constraints, then pilot one queue. The decision is about service quality and operating control, not the number of AI features in a demo.
Using only the approved policy passages and case facts below, draft a reply. First list the source passage used. Then state what is known, unknown, the next action, and the expected time. Do not invent a policy, refund, delivery date, exception, or commitment. If the sources do not answer the question, say that the case needs human review. Case: [paste minimized case]. Sources: [paste approved passages].
Turn this support thread into an escalation brief with customer goal, exact issue, context, steps already tried, evidence, promised actions, unresolved questions, severity, suggested owner, and what the next agent must not repeat. Preserve the customer's meaning. Mark inferred detail as an assumption and do not promise a resolution. Thread: [paste approved content].
Review these anonymized support cases against the approved knowledge articles. Identify repeated questions, missing articles, conflicting instructions, outdated passages, customer terminology, and cases that should not be automated. Cite the case pattern and name the subject-matter owner who should approve each change. Cases: [paste]. Articles: [paste].
Score this response for factual accuracy, source grounding, policy adherence, completeness, tone, accessibility, next-step clarity, escalation judgment, and customer effort. Give evidence for each score and quote unsupported claims. Separate coaching advice from an employment decision. Response: [paste]. Rubric: [paste].
Baseline
Choose one queue or intent family and record volume, resolution, repeat contact, escalation, customer effort, handling time, quality, and cost.
Shadow
Let the tool retrieve, classify, or draft without sending or changing records. Compare output with trained reviewers and label errors by cause and severity.
Assist
Give the tool a limited role with trained agents. Require review for every output and record edits, omissions, unsupported claims, and checking time.
Bounded action
If evidence supports it, allow a narrow reversible action such as tagging or a standard status update. Keep refunds and exceptions controlled.
Decision
Expand, pause, narrow, or stop. Publish the result by intent and convert confirmed failures into regression tests.
Resolution quality by intent
Repeat contacts and customer effort
Escalation timing and handoff quality
Agent edit and correction rate
Unsupported or unsafe answers
Source freshness and article gaps
Queue health after automation
Cost including review and failure handling
The workflow recommendations are editorial guidance. The linked reports provide survey context, not universal benchmarks. Preserve each source's population, date, definition, and limitations before quoting a statistic in a business case.
A survey of 321 leaders found 20% reported AI-driven staffing reduction, while 55% reported stable staffing with higher customer volumes.
Open sourceGartner describes priorities around agentic AI, customer experience, ROI, workforce skills, and value-centered service.
Open sourceSalesforce reports survey findings from 6,500 service professionals, including estimates about current and expected AI case handling and security concerns.
Open sourceA Salesforce report distinguishing current AI use, planned use, and the expected role of automation in service work.
Open sourceVendor research about AI copilots, autonomous service, trust, and customer expectations. Treat it as vendor survey evidence, not a universal census.
Open sourceA voluntary framework for managing trustworthiness considerations across AI design, development, use, and evaluation.
Open sourceThe best first tool depends on the support bottleneck. Use an AI copilot for agent replies and knowledge retrieval, an approved ticketing platform for workflow and case records, a meeting or voice tool for call summaries, and a customer-facing agent only for narrow, well-supported intents with a reliable human escalation path. Choose the workflow before the vendor.
AI can reduce repetitive triage, lookup, drafting, and summarization work, but it does not remove the need for people in complex, emotional, high-risk, or exception cases. Gartner reported that 20% of surveyed service leaders had reduced staffing due to AI, while 55% had stable staffing while handling higher volumes. The practical shift is usually task redesign and augmentation.
Use one when the intent is narrow, the source content is current, the action is reversible or controlled, and the customer can reach a capable human without starting over. Do not measure success by deflection alone. Check correct resolution, repeat contacts, escalation quality, customer effort, accessibility, and unsafe or unsupported answers.
Only data permitted by the employer's policy, customer commitments, vendor agreement, and applicable requirements should be sent. Minimize personal and payment information, confirm retention and training settings, restrict access, and keep the final case record in the approved support system.
Start with a baseline for resolution quality, repeat contact, escalation, customer effort, handling time, agent correction, cost, and response latency. Report results by intent and channel. An average containment rate can look good while the AI frustrates customers or transfers the hardest cases to an overloaded team.
Agent assist is usually a strong starting point: retrieve an approved article, suggest a reply, summarize the case, or identify missing information while a human reviews the result. It is visible, reversible, and easier to evaluate than allowing an autonomous system to change an account or promise an exception.