Back to all posts
Customer support ops

How to Audit Your AI Support Agent's Conversations: A Practical QA Checklist

Most teams find out their AI agent gave a bad answer when a customer complains, not when someone reads the transcript. A weekly conversation-audit playbook: what to sample, a five-point rubric, and how to close the loop from a bad answer to a fixed one.

11 min read
QA Customer support ops Citations Human handoff AI customer support
Editorial illustration of a large magnifying glass hovering over a grid of chat bubble icons, with three bubbles glowing amber inside the lens next to a checklist clipboard with checkmarks, representing a sampled review of AI support conversations.

Most teams find out their AI support agent gave a bad answer when a customer complains loudly enough to get noticed — not because someone read the transcript. That gap between what a dashboard reports and what actually happened in the conversation is bigger, and more expensive, than most support leads assume.

A July 2026 industry survey put a number on it: two-thirds of CX leaders call their most recent AI project a success, even though 53% went over budget and 43% are delayed or stalled. A separate survey found three-quarters of enterprises have rolled back an AI deployment, and the top reasons were not “it didn’t sound human enough” — they were customer data exposure, hallucination or brand risk, and the inability to diagnose what went wrong. (Laivly 2026 AI Deployment Index and Sinch survey, via CX Dive, July 2026) You cannot diagnose what you never looked at.

This is a playbook for the part most teams skip: actually reading a sample of your AI agent’s real conversations, on a schedule, against a rubric, and turning what you find into fixes. It assumes you already have an agent live — if you have not launched yet, read the pre-launch testing playbook first, and if you want the metrics to track over time, see the post-launch scorecard. This post is neither of those. It is the recurring human habit that catches what a metrics dashboard structurally cannot.

Why a metrics dashboard isn’t an audit

A resolution-rate or deflection-rate number tells you how many conversations ended a certain way. It cannot tell you whether the agent was right. Those are different questions, and conflating them is how confidently wrong answers survive for months.

CSAT is the closest most teams get to a quality signal, and it is thinner than it looks. Intercom’s own measurement research, published July 1, 2026, found that CSAT surveys capture under 10% of conversations, skew toward people who are unusually happy or unusually angry, and compress several distinct problems into one score — so a 4-out-of-5 can hide a wrong answer that the customer just didn’t bother to flag. (Intercom, “How to measure the customer experience as AI scales,” July 2026)

Manual QA in traditional contact centers has the same shape of problem. Verint (formerly Calabrio) puts typical manual review coverage at 1–3% of interactions — a sample so small that a systemic issue has to happen very often before a reviewer stumbles onto it by chance. (Verint / Calabrio call center QA guide) An AI agent that answers thousands of conversations a month can develop a specific, repeatable failure — a stale price, a policy the agent keeps guessing at, a tone that reads as dismissive on refund questions — and none of it shows up in a dashboard tuned to count outcomes rather than read them.

What reading a transcript catches that a number can’t

A rubric-based read of an actual conversation surfaces four failure classes that resolution rate, deflection rate, and CSAT are structurally blind to:

None of these show up as an anomaly in aggregate numbers. They show up when a person reads the actual back-and-forth.

Build a sample that finds problems, not just conversations

Random sampling alone wastes most of your review time on conversations that were already fine. A stratified sample finds more per hour:

  1. Every escalated or handed-off conversation, or a defensible slice of them if volume is high. These are pre-selected for risk — the AI already signaled it couldn’t finish the job, so check whether it stopped at the right point and passed along the right context.
  2. Every low-confidence or negatively-rated conversation. If your tool surfaces conversations flagged by a thumbs-down or a low-confidence signal, review all of them — this bucket is small by construction and highest-yield.
  3. Your top 10–20 questions by volume, a handful each. These carry the most exposure if something is systematically wrong, because the same mistake repeats at scale.
  4. A small random baseline — even 10–15 conversations a week — specifically to catch the failure mode nobody flagged. Escalation and low-confidence signals only catch what the system noticed was uncertain; a confidently wrong answer, by definition, didn’t trip either alarm.

Thirty to fifty conversations a week across these four buckets is enough to catch a real, repeating problem within days rather than months — and it’s a fraction of the volume most teams assume “proper QA” requires.

A five-point rubric, not a vibe check

Reading transcripts without a rubric turns into skimming for anything that feels off, which is inconsistent between reviewers and forgets last week’s findings. Score each sampled conversation against the same five checks:

CheckWhat you’re verifying
GroundingClick through the cited source. Does it actually say what the agent claimed, not just something adjacent?
Tone and policy fitDoes the reply match your brand voice and current policy, especially on refunds, discounts, and anything with legal or financial weight?
Escalation timingDid the agent hand off too late (after guessing), too early (on something it could have answered), or not at all when it should have?
Resolution truthRead the last two messages. Did the customer actually get what they needed, or did the conversation just end?
Repeat patternHave you seen this exact failure in a prior audit? A repeat means the last fix didn’t hold, or the fix never happened.

A conversation that fails “grounding” needs a content fix. One that fails “tone” needs an instructions fix. One that fails “escalation timing” needs a handoff-rule fix — see the handoff playbook for where those rules usually live. Keeping the rubric this specific is what turns “read some chats” into something you can hand to a new hire and get consistent results from.

Close the loop, or the audit is just reading

An audit that doesn’t produce a fix is a wasted hour. Every failed check should land in one of three places by the end of the session:

Track one number across audits: repeat-failure rate — the share of this week’s findings that are the same issue as last week’s. A flat or rising repeat rate means the audit is finding problems but nobody’s fixing them, which is worse than not auditing at all, because it creates a paper trail of known issues nobody acted on.

Make it thirty minutes, not a quarterly project

The audit that survives is the one sized to fit inside a normal week, with a named owner:

This is deliberately smaller than a full call-center QA program with scorecards and calibration sessions. For most small and mid-size support teams, a disciplined thirty-minute weekly habit catches the failures that matter more reliably than a heavyweight process that gets deprioritized the first busy week.

Where Owlish fits

Owlish’s console has real building blocks for this workflow, not just a metrics dashboard. Chat Logs is a per-agent browser of every past session, and every AI reply shows a “Used N sources” toggle that expands to the actual cited source — so checking the “grounding” row of the rubric means clicking a link inside the transcript, not hunting through your knowledge base separately. The analytics dashboard’s “Conversations needing review” list is already the stratification work for bucket 2 above: it auto-surfaces sessions flagged by a thumbs-down or low AI confidence, so that highest-yield bucket is pre-built rather than something you assemble by hand. The Helpdesk Inbox is where every escalated or handed-off conversation lands, covering bucket 1. And when a transcript fails the grounding check, Revise Answer lets an operator correct the reply directly from Chat Logs and save it as a Direct Response — a real “fix and pin” loop, not a suggestion to go update a document somewhere else.

What Owlish doesn’t have today: a CSAT survey or star-rating capture (feedback is thumbs up/down on individual replies), a manual tag-or-flag-for-review workflow beyond the automatic thumbs-down/low-confidence surfacing, or a built-in scoring rubric UI — the five-point checklist above is still something you run yourself, on paper or a spreadsheet, using what Chat Logs and the Inbox surface. Full-coverage AI-graded scoring of every conversation, the direction Intercom’s research points toward, isn’t something Owlish or most tools in this price range do today; treat this playbook as the practical middle ground between a 1–3%-of-everything legacy sample and that not-yet-mainstream ideal.

FAQ

How many conversations should I audit each week? Thirty to fifty is a reasonable range for most small support teams — enough across the four sample buckets to catch a repeating problem within a week or two, without turning into a part-time job. Scale it to your volume: if you handle a few hundred conversations a week, thirty is proportionally a much bigger sample than if you handle tens of thousands.

Is this different from testing before launch? Yes. Pre-launch testing runs a fixed scenario bank against an agent that hasn’t talked to customers yet, to catch failures before anyone sees them — see the pre-launch QA playbook. A conversation audit reviews real, already-happened conversations on an ongoing basis, because production traffic finds edge cases no pre-launch scenario bank anticipated.

Can’t the AI just grade its own conversations? Automated, full-coverage AI scoring is a real and growing pattern — Intercom’s July 2026 research reports roughly 5x the conversation coverage of CSAT alone using this approach. It’s a meaningful complement once you have volume to justify the tooling, but it doesn’t replace a human periodically reading real transcripts against a rubric, especially for the “does this citation actually support the claim” and “does this match our brand” checks, which still need a human judgment call.

What if we don’t have time for a formal QA program? That’s the point of sizing this to thirty to sixty minutes a week rather than a calibration-session-and-scorecard program built for a 200-agent contact center. A small, consistent habit with a named owner beats an ambitious program that gets skipped the first time things get busy.

What should trigger an off-cycle audit instead of waiting for the weekly review? A new product launch, a policy change (pricing, refunds, terms), a spike in escalations on one topic, or any customer complaint that reached you directly. Any of those warrants pulling the relevant conversations immediately rather than waiting for the scheduled sample.


If you’re running an AI support agent and want the citation trail and the flagged-conversation list already built into the review workflow, see how Owlish’s Chat Logs and Helpdesk Inbox work, or start free and audit your first week of conversations against the checklist above.

Keep reading

Related posts

Try Owlish

Build a support agent your operators actually trust.

Start Free without a card. Source-cited answers. Hand off to a human the moment the agent isn't sure.