← All articles

When a chatbot is the wrong answer

Half the businesses that ask us for a chatbot have a documentation problem, not a conversation problem.

A support agent at a busy desk surrounded by handwritten notes and a policy manual, with a chatbot window open on screen

The short version

  • A bot can only answer from what exists. Undocumented policy becomes invented policy.
  • Deflection is not resolution — the gap between the two is where churn hides.
  • Write and publish your twenty most-asked answers first. Often the bot stops being urgent.
  • Where bots do work, it is because they are wired to a real system, not to a vibe.

A chatbot can only answer from what exists. If your policies live in three people’s heads, a bot will confidently invent the rest, and you will have created a new problem with a friendly interface on it.

Half the businesses that ask us for a chatbot have a documentation problem, not a conversation problem. That is not a criticism — it is the normal state of a company that grew by knowing things rather than writing them down.

Deflection is not resolution

The metric most bots are sold on is deflection: the share of enquiries that never reached a human. It is a flattering number, because an abandoned conversation and a confidently wrong answer both count as successes.

Gartner's research is blunt about the gap. AI deflects more than 45% of customer queries, but only around 14% of issues reach genuine self-service resolution — the other thirty-odd points are customers who came back through a different channel, or gave up. Even for issues customers themselves described as very simple, the resolution rate reached only 36%.

A bot with a 90% deflection rate can sit on a 40% resolution rate. The number worth running the business on is the one that confirms the problem went away.

When self-service fails, the causes are consistent and they are not model quality: around 43% of failures happen because the customer cannot find content relevant to their issue, and about 45% because the system did not understand what they were trying to do. Both are content and scope problems.

What we suggest instead

Write the twenty questions your customers actually ask, answer them properly once, and put them somewhere a person can find. Concretely:

  1. Pull the last two hundred enquiries from email, WhatsApp and the phone log. Do not guess the list — the real one is always different from the imagined one.
  2. Cluster them. You will usually find that twelve to twenty question types cover eighty percent of volume.
  3. Write one clear answer each, approved by whoever actually owns that policy. Include the awkward ones: refunds, delays, pricing exceptions.
  4. Publish them where customers already look, and where staff can copy from them.
  5. Measure again in a month. Enquiry volume usually falls enough that the bot is no longer the urgent purchase.

If it still is, now it has something true to work from — and the deflection numbers move accordingly. Teams whose help content was updated in the last thirty days consistently outperform teams whose content has not been reviewed in six months, by more than a factor of two.

Where bots do earn their keep

If you do build one, build it properly

Measure it honestly

Two numbers, reviewed monthly. Containment: the share of conversations that ended without escalation. True deflection: containment minus anyone who came back through another channel within seven days. The second number is typically thirty to forty percent lower than the first, and it is the only one worth reporting to a board.

Set a floor as well as a target: if containment is below thirty percent after a month, the answer is not a better model. It is missing content.

Frequently asked questions

Will a chatbot reduce our headcount?

In a business under fifty people, almost never. What it changes is what the same people spend their day on — the same pattern we described in the invoice automation write-up.

Can it use our existing documents?

Yes, and that is the right approach — retrieval over your own material rather than a model answering from general knowledge. The quality ceiling is set by the documents, which is why we start there.

What does it cost to run?

Per-conversation costs are small; the ongoing cost that matters is the half-day a month someone spends reading transcripts and updating answers. Budget for that or the bot will decay.

Can you just tell us whether we need one?

That is what the free review is for. An hour, an honest answer, and you keep the write-up either way. Book a call.

Sources

  • Gartner customer service research, 2025 — self-service resolution rate (14%), deflection above 45%, and causes of self-service failure.
  • Gartner 2025 Customer Service Technology Survey — containment benchmarks for retrieval-based versus rule-based deployments.
  • HubSpot State of Service and industry deflection benchmark syntheses, 2025–2026 — effect of knowledge base freshness on deflection.

Recognise your own setup in this?

The free audit takes an hour and you keep the write-up.

Book a call

Keep reading