Six people. About 2,400 tickets a month. That was the support desk at a mid-sized e-commerce platform I sat with last spring, and the backlog had stopped being a scheduling problem and started being a morale problem. Two agents spent most of every Monday answering the same question about where an order was, or how to apply a discount code that had already expired.
We set up Gleen AI over a two-week window. Sixty days later, roughly 38% of incoming conversations were resolved without a human touching them, first response time on chat dropped from four minutes to under a minute, and CSAT held steady at 4.6 out of 5. Nobody got laid off. Two people got moved onto writing documentation and chasing product bugs, which was a better use of their week anyway.
What follows is the sequence I would run again, in order, with the actual numbers and examples from that rollout. Skip steps and you will get a chatbot that confidently invents a refund policy. Follow them and the thing mostly behaves.
Step 0: Mine your own ticket history before you touch the tool
The single most common mistake is logging into a new AI support platform and immediately pointing it at the help center. The help center is not where your customer questions live. Your inbox is.
Pull 90 days of closed tickets
Export them with tags, resolution notes, and the agent’s final reply. For that e-commerce team, 90 days came to just over 7,100 tickets. We tagged them by hand for two days (tedious, worth it) and found that 19 question types accounted for 61% of the volume. Order tracking, discount codes, return eligibility, delivery address changes, and invoice requests were the top five on their own.
Build a one-page inventory
You want a list of what people actually ask, how often, and whether the answer already exists somewhere in writing. Collect these before configuration starts:
- Top 20 questions by volume, with rough counts
- Every answer that only exists in an agent’s head or in Slack
- Your refund, cancellation, and escalation policies as written documents
- The five cases where an agent must involve a human, and why
- Any product quirk that generates “why does it do that” tickets
That fourth item matters more than the rest. Boundaries are a knowledge input, not an afterthought.
Step 1: Connect the sources that reflect reality
Gleen ingests from help center articles, Notion or Confluence spaces, PDFs of internal policy, and past resolved replies. Connect all of them, then check what it actually pulled in. In our case, the ingestion surfaced a 2022 shipping policy that contradicted the current one. That kind of conflict is exactly what you want to find on day one rather than in a customer conversation.
This is where the system’s underlying approach matters. Gleen AI learns your products from the sources you feed it, rather than relying on a generic model that has read the internet and none of your internal wiki. If your documentation is thin, the answers will be thin. There is no prompt that fixes a missing policy page.
Spend an hour deleting or dating stale articles. It is the highest-leverage hour in the whole project.
Step 2: Write your first three playbooks, not thirty
A playbook is the instruction set for a category of conversation: what triggers it, what the agent is allowed to do, how it should sound, and when it hands off. Resist building for every scenario. Three playbooks covering your top 40% of volume will teach you more in a week than a hundred built blind.
A worked example
One playbook we wrote handled subscription cancellations for a SaaS client. It looked roughly like this:
- Triggers: “cancel my plan”, “stop billing”, “I don’t want to renew”
- Allowed actions: explain cancellation terms, confirm the effective date, offer a pause option once, process the cancellation
- Required confirmation: read back the account email and the final billing date before acting
- Tone: matter-of-fact. No guilt-tripping, no three-paragraph retention pitch
- Escalate when: the account is on an annual contract, the customer mentions a chargeback, or they ask about a refund for a past period
That last line prevented at least two dozen messy conversations in the first month. The agent handled monthly cancellations cleanly and quietly handed annual contracts to a human, which is precisely the correct division of labour.
Step 3: Define hard limits in writing
Give the system a list of things it must never do, and put that list somewhere a human reviews it quarterly. Ours included:
- Never issue a refund above $150 or outside a 30-day window
- Never change account ownership, email address, or payment method
- Never comment on roadmap, unreleased features, or competitor comparisons
- Never speculate about the cause of an outage
- Never continue a conversation where the customer has asked for a human twice
Two escalations, not one, is a deliberate rule. People sometimes ask for a person out of frustration and get a good automated answer that resolves things. But if they ask twice, they mean it.
Step 4: Test against 50 real tickets before anyone sees it live
Take 50 closed tickets you already know the correct answers to. Run them through the agent. Grade each reply on three things: is it factually correct, is it complete enough that a customer would stop asking, and does it sound like your brand.
A useful benchmark: if fewer than 70% pass all three, do not launch. Go back and fix documentation or the playbook. On our first pass, 31 of 50 passed. The failures clustered around two topics where the internal docs were vague, and after rewriting those two pages, the second pass hit 41 of 50.
Keep a written record of every failure and what you changed. That file becomes the most valuable document on the support team.
Step 5: Roll out in stages, starting with shadow mode
Shadow mode means the agent drafts replies that nobody sends. Agents see the draft next to their own answer and flag anything wrong. Two weeks of shadow mode on a few hundred conversations catches the weird stuff: the customer who writes in Portuguese, the one who pastes a screenshot into the chat, the one asking about a product you discontinued in 2021.
After shadow mode, route roughly 15% of chat traffic to the agent with a human able to jump in. Hold there for a week and watch CSAT and escalation rate. Then expand to all chat, then to email, which is slower and less forgiving because replies need more structure. Full rollout for that e-commerce team took 19 days from first login.
Step 6: Review weekly and fix the biggest failure
A deflection rate is not a launch metric, it is an ongoing practice. Every Monday we pulled the conversations where a human had to correct or completely rewrite the agent’s reply, sorted them by topic, and fixed the single largest cluster.
One week the top cluster was proration. The agent was calculating mid-cycle upgrades correctly on monthly plans and badly wrong on annual ones. The fix was two sentences added to the billing playbook plus a new help article with worked examples. The cluster went from 40 conversations to four the following week.
The other thing weekly review surfaces is demand you never noticed. Twelve customers asked about a new API rate limit that no one had documented anywhere. That became a product announcement, not a support answer.
What 60 days of this actually looks like
Deflection settled around 38% of inbound conversations, and it climbed slowly rather than in a jump. Median first response on chat went from four minutes to 51 seconds. Handle time per ticket for the humans dropped 22% because they stopped re-reading policy PDFs. The team’s headcount stayed at six while volume grew 12%, which is the outcome most support leaders are quietly hoping for and rarely say out loud.
The thing that surprised me was where the hours went. The agents who used to answer tracking questions spent their afternoons writing the documentation that made the agent better, which made fewer tickets reach a human, which gave them more time to write. Set the loop up properly and it compounds. Set it up sloppily and you spend your Mondays apologising to customers for a bot that promised free shipping to Brazil.

