85%
Service organizations using at least one form of AI in Salesforce’s 2026 research
Salesforce report; vendor research with its own sample and definition of service AI.
Don't stop here
Hand-picked guides our readers explore right after this one.
Prompts for building and using AI chatbots for customer service, sales, support automation, and conversational design
Read the guideExpert guide to Claude prompts with XML tags, artifacts, and complex reasoning
Read the guideUnlock Google's Gemini with multimodal prompting strategies
Read the guideEvidence center · Customer service
Adoption is rising, but the useful story is not a race to remove people. It is a redesign of routine work, human escalation, customer trust, and the systems that support both.
Michael Okeje
Primary-source research and service-workflow analysis · Last updated August 13, 2026
85%
Salesforce report; vendor research with its own sample and definition of service AI.
66%
Salesforce reported this rose from 39% in its 2025 comparison; do not treat it as a universal market census.
20%
Gartner survey of 321 customer service and support leaders conducted in October 2025.
55%
Gartner’s 2025 customer service leadership survey; the result supports an augmentation reading.
85%
Gartner survey of 321 service and support leaders worldwide, September–October 2025.
54% vs 32%
Gartner survey of 5,801 U.S. customers conducted January–February 2025.
Customer service AI is often summarized as a race toward agentless support. That framing is too narrow for the evidence. The more useful questions are: which tasks are being automated, which are being assisted, what happens when the system is uncertain, and whether human agents are being removed or moved toward more complex work.
The latest public figures point in several directions at once. Salesforce reported strong growth in agentic AI adoption in its 2026 research. Gartner reported that only 20% of surveyed customer service leaders had reduced agent staffing due to AI, while 55% had stable staffing while handling higher customer volumes. Another Gartner release reported that 85% of service leaders were expanding human-agent responsibilities. Adoption and workforce redesign can happen together.
Customer trust adds another boundary. Gartner reported that 54% of surveyed U.S. customers trusted human agents more than AI for product or service recommendations, while 32% trusted AI more. This does not mean customers reject automation. It means the appropriate channel depends on the consequence and the kind of judgment the interaction requires.
I treat the numbers below as evidence about operating choices, not as permission to make a universal prediction. A support leader should use them to ask where AI can safely reduce effort, where it can improve an agent's context, and where a human relationship is part of the product experience.
Salesforce's 2026 report said 85% of service organizations used at least one form of AI and 66% used agentic AI, up from 39% in its 2025 comparison. Those figures describe Salesforce's research and definitions. They are useful for seeing the direction of change in that population; they are not a census of every contact center, help desk, or local business.
Gartner's 2025 research looked at different questions. Its survey of 265 customer service and support leaders found that AI agents were rising in value but remained outside the top ten technologies leaders ranked as most valuable at that time, while Gartner predicted 73% of customer service organizations would implement agent-assist solutions by the end of 2025. A prediction and an observed adoption figure should remain visibly separate.
The term agentic also needs care. A system that suggests a reply, retrieves an article, summarizes a case, or routes a ticket may be called an AI agent by one vendor and agent assist by another. Ask what the system can actually do: read, recommend, act, modify records, send communications, or resolve a case without approval.
A support team should publish its own adoption definition. For example: assistive AI means a human reviews every output; automated self-service means the system completes a bounded request; autonomous action means the system changes a record or triggers an external effect. Internal metrics become much more useful when the labels are stable.
Gartner reported in December 2025 that 20% of surveyed customer service leaders had reduced agent staffing due to AI, while 55% reported stable staffing while handling higher customer volumes. That is a more nuanced picture than the headline prediction that AI will remove most service jobs. It suggests that increased demand, new channels, and harder cases can absorb efficiency gains.
In April 2026, Gartner reported that 85% of service and support leaders were expanding human-agent responsibilities and 75% were shifting agents into entirely new roles. The implication for workforce planning is practical: if AI handles routine requests, people may need stronger judgment, escalation, product expertise, investigation, quality review, and relationship skills.
A team should measure task mix, not only agent count. Record the percentage of conversations that are routine, the share escalated, the time agents spend reviewing AI output, the amount of context available at handoff, and the kinds of decisions humans make after deployment. A lower contact volume can coexist with higher human complexity.
Do not call a workforce change a productivity improvement until you define who benefits. A shorter average handling time may help the business but increase customer repetition or agent stress. A strong evaluation includes customer effort, repeat contacts, quality, and employee experience alongside cost.
The Gartner customer survey is useful because it separates the customer from the service leader. Among 5,801 U.S. customers surveyed in early 2025, Gartner reported that 54% trusted human agents more than AI for product or service recommendations, compared with 32% who trusted AI more. Recommendations are a useful test case because they involve judgment, context, and the possibility of financial or practical harm.
A customer may happily use self-service to check a delivery status, reset a password, or find a published return window. The same person may want a human when a payment is disputed, a medical issue is involved, a service has failed repeatedly, or the answer requires an exception. Design the escalation around the customer's risk and frustration, not around an arbitrary deflection target.
Disclose what the system is doing in plain language. Do not imply that a human reviewed a recommendation when the system did not. Show the source or policy where appropriate, give the customer a way to correct important information, and preserve the conversation context when handing off. A human handoff that forces the customer to start again is a quality failure even if the AI contained the first contact.
Track escalation quality. An escalation rate can rise because the AI is failing, or because it is correctly identifying complex cases earlier. The useful questions are whether the right cases are escalated, whether the human receives the evidence and attempted steps, and whether the customer feels the handoff solved the problem.
The safest first uses are often agent assist, summarization, knowledge retrieval, translation, classification, and draft generation. These can reduce search and writing effort while keeping a trained person responsible for the answer. They still need source quality, privacy controls, and review because a confident summary can omit a critical exception.
Customer-facing self-service can work well for narrow, well-documented questions with a clear fallback. The system should know which sources are current, identify when the answer is unsupported, and avoid making promises beyond policy. The definition of resolution should include whether the customer got the right outcome, not just whether the chat ended.
High-impact actions need stronger controls. Refunds, cancellations, account changes, identity decisions, legal commitments, medical guidance, and recommendations tied to material consequences should have authentication, authorization, validation, and an approval or escalation path. A model instruction is not a security boundary.
Use [AI evaluation](/ai-evaluation) to test outcomes, trajectories, tool calls, refusals, and escalation. Use [AI governance](/ai-governance) to document ownership, risk, monitoring, and the conditions under which the system is paused. Customer service is a strong example of why quality and governance belong together.
Choose one queue or intent family and record a baseline for four weeks if practical. Measure resolution, repeat contacts, transfers, customer effort, time to resolution, agent handling time, review time, cost, and quality. Identify which cases are excluded and why. The baseline makes it possible to tell whether AI changed the work or simply changed the dashboard.
Run an assisted pilot before autonomous resolution. Compare AI suggestions with normal work, record corrections, and have quality reviewers score accuracy, policy adherence, empathy, completeness, and escalation. For retrieval-based systems, check whether the source passage actually supports the answer.
For a customer-facing system, add a holdout or staged rollout where appropriate. Track customer outcomes, not just containment. Review complaints, repeat contacts, human transfers, and the cases where the system should have escalated earlier. An AI that contains a conversation by frustrating the customer is not creating durable value.
Report the result by intent, not only as a global average. A model may perform well on order status and poorly on billing disputes. The average can conceal the exact category where the business risk lives. Keep an incident log and turn confirmed failures into new evaluation cases.
Define what counts as AI, assistive, and autonomous.
Carry the survey population with each statistic.
Measure resolution quality, not only containment.
Track repeat contacts and customer effort.
Record human correction and escalation quality.
Keep humans accountable for consequential decisions.
Test source grounding and policy adherence.
Give customers an honest path to a person.
Review results by intent instead of only global averages.
Turn incidents into regression tests before expanding scope.
Human-agent responsibility, staffing plans, and U.S. customer trust in human versus AI recommendations.
Open sourceStaffing stability, higher volumes, and the difference between augmentation and replacement.
Open sourceAI and agentic-AI adoption figures, with Salesforce's survey context.
Open sourceExpected AI case handling and the global service-professional survey.
Open sourceThe answer depends on what counts as AI and who was surveyed. Salesforce reported in 2026 that 85% of service organizations in its research used at least one form of AI, while 66% used agentic AI, up from 39% in its 2025 comparison. Gartner's surveys measure different populations and technologies, so the figures should not be combined into one market average.
The evidence points more strongly to augmentation and role redesign than a simple replacement rate. Gartner reported that 20% of surveyed leaders had reduced agent staffing due to AI, while 55% reported stable staffing while handling higher customer volumes. Another Gartner release said 85% of leaders were expanding human-agent responsibilities.
Salesforce reported that its 2025 State of Service research found AI was expected to handle half of all service cases by 2027, up from 30% at the time of the report. That is an expectation from a survey, not a universal measured forecast. Actual coverage depends on the product, data, channel, policy, and escalation design.
Preference depends on the task. Gartner reported that 54% of surveyed U.S. customers trusted human agents more than AI for product or service recommendations, while 32% trusted AI more. Customers may prefer speed and self-service for simple requests but want a human for high-stakes, emotional, or advisory interactions.
Measure resolution quality, customer effort, escalation quality, repeat contacts, time to resolution, agent correction, cost, latency, accessibility, and the rate of unsafe or unsupported answers. Compare those outcomes with a baseline and inspect whether AI shifts difficult work to humans rather than simply reducing contact volume.