I wrote about AI chatbots in ecommerce support in April 2024. Rereading it now, it is an advert. Every paragraph says the technology is transformative and nothing in it could ever have been wrong, because it made no claim specific enough to fail. It also cited eBay ShopBot as a live product. eBay killed ShopBot in September 2018, six years before I published.
So this is the rewrite. Same question, actual numbers, and the numbers are more interesting than the hype was.
Most stores still have not done it
The story everyone tells is that AI support is everywhere. Gorgias runs support for a large number of ecommerce brands, so they can count instead of guess. They looked at a year of their own production traffic and found that roughly one in five brands had put AI in front of a customer at all.
Cross-industry surveys report much higher numbers. Salesforce asked 3,075 service professionals in March and April 2026 and found AI agent use jumped from 39% to 66% in a year. Both can be true. One counts what people say in a survey, the other counts what is actually wired into a live support queue. When those two disagree, believe the traffic.
The trade-off nobody advertises
Here is the part my 2024 version could not have told you, because it never looked. Automation does make support dramatically faster. It also makes customers slightly less happy. Both effects show up in the same dataset.
Automate more, answer faster, satisfy slightly less
Drag through the three automation levels Gorgias actually measured.
Gorgias reported response time at 0% and 30% automation, and CSAT at 0% and 20%. Blank tiles are levels they did not publish, not zeroes. Source: Gorgias, published 1 April 2026.
Median first response falls from over twelve hours to about eighty minutes. That is a real win and not a small one. But satisfaction moves the other way, from 90.3% down to 87.9%. Two and a half points is not a catastrophe. It is also not the universally better experience the sales decks promise, and it is the number that never makes it into a case study.
What happened when someone ran the experiment properly
Surveys and vendor dashboards both have obvious incentives. Randomised field experiments do not. Two of them ran inside Alibaba's live customer service operation and were published in February and May 2026.
The headline result is that generative AI improved speed and, overall, ratings. The second result is the one worth sitting with: for the chats the AI was eligible to handle, ratings dropped substantially, and the worst outcomes came when an already-frustrated customer got escalated from the bot to a human.
Three quarters have already pulled one
The clearest signal that this is harder than it looks is how often it gets undone. Sinch data reported in May 2026 found 62% of companies had customer-communication agents live in production. In the same population, 74% had rolled back or shut down at least one of them on governance grounds.
More companies have retired an agent than currently run one. That is not a technology that does not work. That is a technology being deployed faster than anyone has figured out how to supervise it.
What the winners actually built
The pattern that survives contact with production is boring: let the machine own the queries with a correct answer, and route everything else to a person quickly and without making the customer repeat themselves.
- Order tracking, returns, refunds and cancellations. Bounded questions with a definite answer that lives in a database.
- Triage. Reading the message, tagging it and routing it, without attempting a reply.
- Drafting. The machine writes, a human sends. Most of the speed, none of the unsupervised risk.
- Handing over with the full transcript attached, so the human starts where the bot stopped.
Amazon is the loudest counterexample and worth reading carefully. Rufus, renamed Alexa for Shopping in May 2026, does handle tracking, returns, refunds and billing, and Amazon says over 300 million people used it in 2025 and it drove close to $12 billion in incremental sales. Those are company-reported figures with no independent audit, and they measure shopping, not support quality. Amazon also has a catalogue and an order graph almost nobody else has. Their result does not port to your Shopify store.
A note on how the first version got written
The 2024 post opened with 'In today's fast-paced digital age'. It used 'leveraging the power of' twice. It promised a 'significant competitive advantage' without saying over whom. That is not a writing style, it is the absence of one, and it is what you get when you let a model produce 1,800 words on a topic you have not researched.
The tell is not any single phrase. It is that you can delete most of the sentences and lose no information. Run the opening of the old version through this and see for yourself.
Slop detector
Paste anything. It counts load-bearing nothing. Loaded with the opening of the 2024 version of this post.
- Throat-clearing 3
- Empty intensifiers 3
- Consultant filler 2
This matches a fixed phrase list, so it proves nothing about authorship. Humans write like this too, and any model told to avoid these phrases will. Treat a high score as a prompt to cut, not as a verdict.
The fix was never better vocabulary. It was going and finding out what the numbers say, then reporting them including the ones that undercut the argument. The CSAT drop and the 74% rollback rate are the two most useful facts here, and a promotional version of this post would have contained neither.
The short version
AI support in 2026 is real, narrower than advertised, and genuinely good at a specific job: fast answers to bounded questions. It is measurably worse at making people feel looked after, and the moment it hands over to a human badly, it costs you more goodwill than it saved you in response time.
Build the handoff first. Everything else is easier than that part, and that part is where the satisfaction is won or lost.