Read conversation quality
Success rate, resolution rate, thumbs and answer outcomes measure different things. Here is how to read all four together.
Browse topics
What it is
There is no single "quality score" in Agentency. Nothing hidden is being computed and withheld from you. Instead, four separate measurements each answer a different question, and reading them together is what tells you whether the chatbot is any good.
| Measurement | The question it answers | Where you see it |
|---|---|---|
| Successful response rate | Did the chatbot produce a real answer? | Dashboard card |
| Answer Quality split | How many messages were answered fully, partly, or not at all? | Dashboard chart |
| Resolution Rate | Did whole conversations end with everything answered? | Dashboard card |
| User Feedback | Did visitors say the answers were useful? | Dashboard chart, and thumbs on transcripts |
None of them measures whether an answer was true. That still requires reading transcripts, and this article is largely about doing that efficiently.
When you would use it
When someone asks "is the chatbot working?" and deserves a real answer. Before deciding whether to expand it to another part of your site. After adding knowledge, to check whether the gap you were chasing actually closed.
Where to find it
The four measurements are on the Dashboard, under the sections labelled Quality and performance and Breakdowns.
The transcripts are on Chatbots → your chatbot → Conversations.
Steps
- Open Dashboard and pick a range that contains enough traffic to mean something. Twenty conversations is not a sample.
- Read Successful response rate and Resolution Rate together. The gap between them is where the story is.
- Look at the Answer Quality doughnut. Watch the size of the No answer slice — that is your knowledge gap made visible.
- Look at User Feedback. A handful of thumbs-down is worth more than any percentage on the page, because a visitor bothered to tell you.
- Check Top Chatbots if you run several. One bad chatbot can pull the whole account's average down while the rest are fine.
- Now open Conversations on the weakest chatbot and read. Sort towards the unhappy end: thumbs-down turns, and answers marked as no-answer.
- Fix what you find in knowledge or instructions, not in settings. See Knowledge overview and Write instructions.
- Come back a week later at the same range and compare.
Reading the four together
High success, high resolution. The chatbot is answering, and answering completely. Check a few transcripts for accuracy and then leave it alone.
High success, low resolution. It answers most things but reliably fails at one kind of question. Find the pattern — it is usually a single topic missing from the knowledge base. This combination is the most actionable one on the page.
Low success, lots of "No answer". Either the knowledge does not cover what people ask, or it never finished training. Check the chatbot's knowledge status first — see Dataset status.
Good numbers, bad feedback. The chatbot is confidently answering incorrectly. Nothing on the dashboard will show you this; only the transcripts will. A confident wrong answer counts as a successful one, and that is the honest limitation of any automated quality figure.
Everything looks fine but people escalate constantly. Read the handoff conversations. Sometimes the chatbot is correctly recognising something it should not attempt, which is good. Sometimes it is bailing out of questions it could answer, which means the instructions are too cautious.
What "successful" actually requires
A message counts as successful when the chatbot answered it and either the answer was high-confidence or the visitor gave it a thumbs-up. A real thumbs-up always counts, whatever the confidence — a human saying "this helped" is better evidence than the chatbot's own estimate of itself.
One definition is used everywhere the figure appears, so the dashboard and any per-chatbot view will never disagree about what the word means.
What it does not include: an answer that was refused, an answer that arrived empty, or a turn the chatbot could not complete. Those show as Partial answer or No answer in the quality split.
Thumbs up and thumbs down
Feedback is the only signal that comes from your actual visitors, so it deserves more attention than its volume suggests.
Visitors can rate answers on the website widget and on the hosted page. In the owner's test chat drawer, the thumbs appear in the production-preview mode; the testing mode replaces them with per-turn detail instead.
Two things to know:
- Feedback costs nothing. Rating an answer does not spend a message credit, so you can encourage it freely.
- Volume is always low. Most visitors never rate anything. Treat every thumbs-down as a bug report worth reading in full, rather than as one data point in a percentage.
The User Feedback chart stays empty until someone rates something: "Thumbs up/down from your visitors will appear here." An empty chart means no ratings, not zero-rated answers.
A weekly quality routine
Fifteen minutes, once a week, beats a dashboard nobody reads:
- Dashboard on 7 days. Note the two rates and the size of the No answer slice.
- Open Conversations and read every thumbs-down. There will not be many.
- Read three no-answer conversations at random.
- Write down the questions the chatbot could not handle.
- Add one piece of knowledge that answers the most common one, and retrain.
- Next week, check whether that slice shrank.
Limits and plan notes
Every plan gets the same measurements. There is no premium analytics tier.
Test chat is included in all of it. Your own testing raises message counts and shifts the rates, which matters most in launch week when your own messages may outnumber real ones. Bear it in mind before reporting the first week's figures to anyone.
Costs and model internals are not shown. These are product measurements, not billing. Usage and spending live on Settings → Billing — see How message credits work.
Empty conversations are excluded from every conversation-level figure, so a visitor who opened the bubble and left changes nothing.
Common problems
Every figure is zero after a busy day.
Check the range and your timezone before anything else, then work through Analytics shows no data yet.
Success is high and I still get complaints.
Read transcripts. A confident wrong answer is counted as a success, and no automated metric will catch it for you.
No feedback at all.
Normal. Most visitors do not rate answers. Look at the answer-quality split instead, which covers every message rather than only rated ones.
One chatbot is dragging the average down.
That is what Top Chatbots is for — each entry carries its own success rate. Work on that one rather than on the account average.
Quality dropped the day after I added knowledge.
Check that training actually finished. A chatbot mid-retrain, or one whose training failed, answers from what it had before — see Knowledge is stuck processing.
Common questions
Where do I find the questions my chatbot handled badly?
Open the chatbot's Overview tab and use the Needs attention panel. It filters between unanswered questions and replies that got a thumbs-down, and links straight into the transcript.
Is it bad that the chatbot says it does not know?
It is far better than inventing an answer. A refusal is not counted as a success, and a cluster of them on one topic is a precise list of the material you have not written yet.
Where can visitors rate an answer?
On the website widget and the hosted chatbot page. Inside Test Chatbot, the production preview mode shows the same thumbs so you can see what a visitor sees.
Almost nobody leaves feedback. Is that normal?
Yes. Most visitors never rate anything. Judge quality from the answered, partial and unanswered split and from reading conversations, and treat thumbs as a bonus signal.
Quality dropped after I added a lot of content. Why?
More material means more competition for each question. Remove duplicated or obsolete sources rather than adding more, and keep one authoritative source per topic.
Was this article helpful?
Related articles
Understand dashboard KPIs
Conversations, messages, answer quality, response time, resolution and handoff — what each dashboard card really counts.
Analytics shows no data yet
Zeros usually mean the wrong account, the wrong window, or no visitor messages yet. Work down this list in order.
Chatbot readiness
Two independent lights: active versus draft, and trained versus not. Visitors need the first; good answers need both.
The chatbot will not answer
Read the exact message the visitor sees — each cause produces a different one. Credits, training, pausing, and the widget all look alike otherwise.
Ready to try it on your own content?
Create a free workspace, add a document, and ask the questions your team is tired of answering.