# How to Audit Your AI Support Agent's Conversations: A Practical QA Checklist

> Most teams find out their AI agent gave a bad answer when a customer complains, not when someone reads the transcript. A weekly conversation-audit playbook: what to sample, a five-point rubric, and how to close the loop from a bad answer to a fixed one.

*By Mithun · Published July 11, 2026 · 11 min read*

Category: Customer support ops

Tags: QA, Customer support ops, Citations, Human handoff, AI customer support

{/* Image note: generated magnifying-glass-over-chat-bubbles illustration sits above the title. No product screenshots needed — this is an operations playbook, not a comparison post. */}

Most teams find out their AI support agent gave a bad answer when a customer complains loudly enough to get noticed — not because someone read the transcript. That gap between what a dashboard reports and what actually happened in the conversation is bigger, and more expensive, than most support leads assume.

A July 2026 industry survey put a number on it: two-thirds of CX leaders call their most recent AI project a success, even though 53% went over budget and 43% are delayed or stalled. A separate survey found three-quarters of enterprises have rolled back an AI deployment, and the top reasons were not "it didn't sound human enough" — they were customer data exposure, hallucination or brand risk, and the inability to diagnose what went wrong. ([Laivly 2026 AI Deployment Index and Sinch survey, via CX Dive, July 2026](https://www.customerexperiencedive.com/news/behind-the-disconnect-how-cx-leaders-view-ai-projects-and-results/824521/)) You cannot diagnose what you never looked at.

This is a playbook for the part most teams skip: actually reading a sample of your AI agent's real conversations, on a schedule, against a rubric, and turning what you find into fixes. It assumes you already have an agent live — if you have not launched yet, read the [pre-launch testing playbook](/blog/test-ai-support-agent-before-launch/) first, and if you want the metrics to track over time, see the [post-launch scorecard](/blog/ai-customer-service-metrics/). This post is neither of those. It is the recurring human habit that catches what a metrics dashboard structurally cannot.

## Why a metrics dashboard isn't an audit

A resolution-rate or deflection-rate number tells you *how many* conversations ended a certain way. It cannot tell you *whether the agent was right*. Those are different questions, and conflating them is how confidently wrong answers survive for months.

CSAT is the closest most teams get to a quality signal, and it is thinner than it looks. Intercom's own measurement research, published July 1, 2026, found that CSAT surveys capture under 10% of conversations, skew toward people who are unusually happy or unusually angry, and compress several distinct problems into one score — so a 4-out-of-5 can hide a wrong answer that the customer just didn't bother to flag. ([Intercom, "How to measure the customer experience as AI scales," July 2026](https://www.intercom.com/blog/how-to-measure-customer-experience-as-ai-scales/))

Manual QA in traditional contact centers has the same shape of problem. Verint (formerly Calabrio) puts typical manual review coverage at 1–3% of interactions — a sample so small that a systemic issue has to happen very often before a reviewer stumbles onto it by chance. ([Verint / Calabrio call center QA guide](https://www.verint.com/guides/call-center-quality-assurance-best-practices/)) An AI agent that answers thousands of conversations a month can develop a specific, repeatable failure — a stale price, a policy the agent keeps guessing at, a tone that reads as dismissive on refund questions — and none of it shows up in a dashboard tuned to count outcomes rather than read them.

## What reading a transcript catches that a number can't

A rubric-based read of an actual conversation surfaces four failure classes that resolution rate, deflection rate, and CSAT are structurally blind to:

- **Confidently wrong, technically cited.** The agent names a source, but the source doesn't actually say what the agent claimed — stale content, a citation that supports an adjacent fact, or a case where retrieval pulled the wrong paragraph. This looks identical to a good answer in every metric except accuracy.
- **Resolved by the numbers, unresolved for the customer.** The conversation ended without an escalation, so it counts as a win — but the last message was the customer giving up, not the customer being satisfied.
- **Policy or tone drift.** A refund question gets an answer that's technically correct but reads as curt, or the agent starts offering something it isn't authorized to offer. Nothing breaks; it just quietly stops matching the brand.
- **Escalation timing that's wrong in either direction.** Handing off a question the agent could easily have answered wastes an operator's time; not handing off a question the agent clearly couldn't answer creates the first failure class above.

None of these show up as an anomaly in aggregate numbers. They show up when a person reads the actual back-and-forth.

## Build a sample that finds problems, not just conversations

Random sampling alone wastes most of your review time on conversations that were already fine. A stratified sample finds more per hour:

1. **Every escalated or handed-off conversation**, or a defensible slice of them if volume is high. These are pre-selected for risk — the AI already signaled it couldn't finish the job, so check whether it stopped at the right point and passed along the right context.
2. **Every low-confidence or negatively-rated conversation.** If your tool surfaces conversations flagged by a thumbs-down or a low-confidence signal, review all of them — this bucket is small by construction and highest-yield.
3. **Your top 10–20 questions by volume**, a handful each. These carry the most exposure if something is systematically wrong, because the same mistake repeats at scale.
4. **A small random baseline** — even 10–15 conversations a week — specifically to catch the failure mode nobody flagged. Escalation and low-confidence signals only catch what the *system* noticed was uncertain; a confidently wrong answer, by definition, didn't trip either alarm.

Thirty to fifty conversations a week across these four buckets is enough to catch a real, repeating problem within days rather than months — and it's a fraction of the volume most teams assume "proper QA" requires.

## A five-point rubric, not a vibe check

Reading transcripts without a rubric turns into skimming for anything that feels off, which is inconsistent between reviewers and forgets last week's findings. Score each sampled conversation against the same five checks:

| Check | What you're verifying |
| --- | --- |
| **Grounding** | Click through the cited source. Does it actually say what the agent claimed, not just something adjacent? |
| **Tone and policy fit** | Does the reply match your brand voice and current policy, especially on refunds, discounts, and anything with legal or financial weight? |
| **Escalation timing** | Did the agent hand off too late (after guessing), too early (on something it could have answered), or not at all when it should have? |
| **Resolution truth** | Read the last two messages. Did the customer actually get what they needed, or did the conversation just end? |
| **Repeat pattern** | Have you seen this exact failure in a prior audit? A repeat means the last fix didn't hold, or the fix never happened. |

A conversation that fails "grounding" needs a content fix. One that fails "tone" needs an instructions fix. One that fails "escalation timing" needs a handoff-rule fix — see the [handoff playbook](/blog/ai-support-handoff/) for where those rules usually live. Keeping the rubric this specific is what turns "read some chats" into something you can hand to a new hire and get consistent results from.

## Close the loop, or the audit is just reading

An audit that doesn't produce a fix is a wasted hour. Every failed check should land in one of three places by the end of the session:

- **A corrected answer**, saved as canonical wording so the same question doesn't fail again. If your tool supports overwriting a bad AI reply and pinning the correct version, use it — this is the fastest way to fix a "confidently wrong" answer for good, versus hoping the underlying source gets noticed and fixed on its own schedule.
- **A source-freshness ticket**, if the citation was accurate months ago but the underlying page changed. This is a knowledge base problem, not an agent problem — see the [knowledge base maintenance playbook](/blog/ai-knowledge-base-maintenance/).
- **An instructions or handoff-rule edit**, if the failure is about tone, scope, or timing rather than facts.

Track one number across audits: **repeat-failure rate** — the share of this week's findings that are the same issue as last week's. A flat or rising repeat rate means the audit is finding problems but nobody's fixing them, which is worse than not auditing at all, because it creates a paper trail of known issues nobody acted on.

## Make it thirty minutes, not a quarterly project

The audit that survives is the one sized to fit inside a normal week, with a named owner:

- **Weekly, not monthly.** A monthly cadence lets a bad pattern run for weeks before anyone catches it — by the time you review, the wrong answer may have gone out hundreds of times.
- **One named owner**, even if it rotates. "The team will look at this" reliably means nobody does.
- **Thirty to sixty minutes.** If the review is taking longer, the sample is too big or the rubric is too vague — tighten one of them rather than letting the session balloon until it gets skipped under deadline pressure.
- **Higher-risk topics get a tighter cadence.** Refunds, cancellations, anything with legal exposure, or a newly launched product line warrants daily spot checks for the first week or two, then steps down to weekly once it's stable.

This is deliberately smaller than a full call-center QA program with scorecards and calibration sessions. For most small and mid-size support teams, a disciplined thirty-minute weekly habit catches the failures that matter more reliably than a heavyweight process that gets deprioritized the first busy week.

## Where Owlish fits

Owlish's console has real building blocks for this workflow, not just a metrics dashboard. **Chat Logs** is a per-agent browser of every past session, and every AI reply shows a "Used N sources" toggle that expands to the actual cited source — so checking the "grounding" row of the rubric means clicking a link inside the transcript, not hunting through your knowledge base separately. The analytics dashboard's **"Conversations needing review"** list is already the stratification work for bucket 2 above: it auto-surfaces sessions flagged by a thumbs-down or low AI confidence, so that highest-yield bucket is pre-built rather than something you assemble by hand. The **Helpdesk Inbox** is where every escalated or handed-off conversation lands, covering bucket 1. And when a transcript fails the grounding check, **Revise Answer** lets an operator correct the reply directly from Chat Logs and save it as a Direct Response — a real "fix and pin" loop, not a suggestion to go update a document somewhere else.

What Owlish doesn't have today: a CSAT survey or star-rating capture (feedback is thumbs up/down on individual replies), a manual tag-or-flag-for-review workflow beyond the automatic thumbs-down/low-confidence surfacing, or a built-in scoring rubric UI — the five-point checklist above is still something you run yourself, on paper or a spreadsheet, using what Chat Logs and the Inbox surface. Full-coverage AI-graded scoring of every conversation, the direction Intercom's research points toward, isn't something Owlish or most tools in this price range do today; treat this playbook as the practical middle ground between a 1–3%-of-everything legacy sample and that not-yet-mainstream ideal.

## FAQ

**How many conversations should I audit each week?**
Thirty to fifty is a reasonable range for most small support teams — enough across the four sample buckets to catch a repeating problem within a week or two, without turning into a part-time job. Scale it to your volume: if you handle a few hundred conversations a week, thirty is proportionally a much bigger sample than if you handle tens of thousands.

**Is this different from testing before launch?**
Yes. Pre-launch testing runs a fixed scenario bank against an agent that hasn't talked to customers yet, to catch failures before anyone sees them — see the [pre-launch QA playbook](/blog/test-ai-support-agent-before-launch/). A conversation audit reviews real, already-happened conversations on an ongoing basis, because production traffic finds edge cases no pre-launch scenario bank anticipated.

**Can't the AI just grade its own conversations?**
Automated, full-coverage AI scoring is a real and growing pattern — Intercom's July 2026 research reports roughly 5x the conversation coverage of CSAT alone using this approach. It's a meaningful complement once you have volume to justify the tooling, but it doesn't replace a human periodically reading real transcripts against a rubric, especially for the "does this citation actually support the claim" and "does this match our brand" checks, which still need a human judgment call.

**What if we don't have time for a formal QA program?**
That's the point of sizing this to thirty to sixty minutes a week rather than a calibration-session-and-scorecard program built for a 200-agent contact center. A small, consistent habit with a named owner beats an ambitious program that gets skipped the first time things get busy.

**What should trigger an off-cycle audit instead of waiting for the weekly review?**
A new product launch, a policy change (pricing, refunds, terms), a spike in escalations on one topic, or any customer complaint that reached you directly. Any of those warrants pulling the relevant conversations immediately rather than waiting for the scheduled sample.

---

If you're running an AI support agent and want the citation trail and the flagged-conversation list already built into the review workflow, [see how Owlish's Chat Logs and Helpdesk Inbox work](/what-it-is/), or [start free](/pricing/) and audit your first week of conversations against the checklist above.

---

Source: https://owlish.bot/blog/ai-support-agent-conversation-audit/
