Metricsense
ComplianceQuality AssuranceInsurance

Your QA score only measures the calls you chose to hear

A 94% compliance rate sounds like a fact about your call floor. It is a fact about your sample. In a regulated contact centre, that gap is where the fines live.

Metricsense Team
21 July 2026 · 7 min read
Share

It is Monday. The compliance report is due. Somewhere in it is a number you will stand behind in front of the board, and maybe one day in front of a regulator: last month you were 94% compliant.

Here is the quiet problem with that number. It was calculated from the calls a QA analyst had time to listen to. At 500 calls a week, that is roughly 2,000 calls a month, and a human reviewing a handful per agent covers maybe 2 to 5% of them. The other 95% sat in your recording system, untouched. Your 94% is not a measurement of your call floor. It is a measurement of the slice you sampled, projected onto the whole with a confidence nobody actually has.

That is the reframe worth sitting with: a sampled compliance rate is not a compliance rate. It is an estimate of one, taken from the calls you chose to hear.

Call it the 5% Lie. Not because anyone is lying, but because the number behaves like a truth about all your calls when it is only a truth about a few of them.

Why sampling felt like the honest choice

The 5% was never negligence. It was arithmetic. Nobody can listen to 2,000 calls a month, so QA does the only thing a human process can: it takes a representative sample, scores it carefully, and extrapolates. That is how survey research works, how exit polls work, how quality control on a production line works. For decades it was the best available method, and a diligent QA team ran it diligently.

The trouble is that a contact centre is not a production line, and a compliance breach is not a random defect. Sampling assumes the thing you are measuring is spread evenly through the population, so a small slice represents the whole. Compliance risk is the opposite. It is rare, clustered, and it hides in exactly the calls a sample is least likely to contain.

The call that gets you fined is not in your sample

There is a famous story from the Second World War. Analysts wanted to armour the bombers that came back riddled with bullet holes, until the statistician Abraham Wald pointed out the obvious thing everyone had missed: the planes that came back were the ones that survived. The holes showed where a bomber could take damage and still fly home. The armour belonged on the parts with no holes, because the planes hit there never came back to be studied.

Your QA sample is the planes that came back. The calls a busy analyst reaches are disproportionately the ordinary ones. The single call most likely to put a licence at risk, an agent who slips from general advice into personal advice, or a distressed customer who should have been escalated, is almost always in the 95% nobody listened to. It does not announce itself. It surfaces weeks later as an AFCA complaint or an ASIC question, long after coaching could have fixed it.

The call that becomes a regulator complaint is, by definition, the one your sample did not include.

We can put a number on how thin that needle is. When Life Insurance Direct, an Australian life-insurance broker running on AWS Connect under a general-advice licence, analysed a full window of 5,438 calls, exactly 16 were flagged as High Concern or Critical Risk. One of them was a 73-year-old carer on a pension whose premiums were climbing $50 a month every year, who mentioned she owed money to family. Price shock, financial vulnerability, and a fixed income, a genuine escalation that needed senior handling.

16 of 5,438

calls flagged High Concern or Critical Risk. A 2 to 5% sample never reliably finds sixteen calls in five thousand. That is not a QA failure. It is a mathematical certainty of sampling.

The sample is not just small. It is biased.

Even the calls you do review are not a fair picture, because your agents know the game. When people know which calls are scored, they perform for the scored ones. The disclosure is crisper, the pace is calmer, the boxes get ticked. Score goes up. Complaints, drawn from the unscored calls, stay exactly where they were.

This is the observer effect, and it means a sampled score does not just miss the risky calls. It systematically flatters the ones it sees. You are not measuring performance. You are measuring performance-under-observation, which is a different and more optimistic thing.

What 100% actually changes (and what it does not)

The obvious answer is to stop sampling and read every call. It is now technically possible: modern analysis can score 100% of conversations against whatever checks you define, and link every finding to the exact moment in the transcript. Life Insurance Direct connected their call feed and had structured results across all 5,438 calls, on 15 separate dimensions, in under an hour, inside their own AWS account, with PII stripped before any transcript reached a model.

The part most people get backwards

Reading 100% of calls does not bury you in failures. It separates the failures from the noise.

Take the most basic, most auditable obligation a general-advice licensee has: the General Advice Warning, required under s949A of the Corporations Act whenever general advice is given. Of the 1,597 calls where an advice context applied, a naive scan flagged 557 as having no warning captured. A lesser tool would have headlined “557 breaches” and handed a reviewer an unusable pile.

The truth was more useful and more honest. Most of those 557 were calls transferred before any advice was given, or calls where no product information was ultimately provided, so no warning was ever triggered. The genuine exposure was the far narrower slice inside it: calls where product information *was* given but no warning followed. That is the subset a human actually needs to review, and isolating it is the whole value. Not “we found 557 breaches”, but “we separated the calls a person needs to look at from the hundreds that are routine”.

The flip side is just as valuable. The same pass confirmed 281 calls as clean general advice with the warning properly delivered, and caught 21 calls where an agent did the hardest thing in the job correctly: a customer asked “so which one is best for me?” and the agent held the general-advice line instead of crossing into personal advice. And when an agent did cross it, the evidence was right there:

CallerSo which one would actually be the best for me?
AgentHonestly, I'd go with TAL if I were you, that's the best one for your situation.
FlagLikely personal advice, given under a general-advice licence.

That is what a defensible compliance signal looks like. Not a score you have to trust, the exact sentence, on the exact call, ready for a human to confirm.

The honest limits

Full coverage is not magic, and anyone who tells you it is should worry you more than the 5% Lie does. Three limits worth stating plainly.

First, an automated flag is a candidate issue with evidence, not a finding of breach. A human QA reviewer confirms it. The tool reports nothing to a regulator on its own. You keep your scorecards and your judgement; you just stop reviewing blind.

Second, analysis runs on transcripts after the call ends. It is not live in-call monitoring, and it will not whisper in an agent's ear mid-call. Its job is to make the week's coaching precise, not to referee the call as it happens.

Third, it reads what was said. Correlating a call to a downstream outcome that lives in another system, a chargeback, a lapse, a no-show, needs your data joined to theirs. The call content is the part that was invisible; the joins are still yours to make.

None of that dilutes the core point. It sharpens it. The goal was never to replace human judgement with a black box. It was to stop asking human judgement to work from a 5% sample.

What this is really about

Notice what changes for the compliance lead when the sample goes away. The fear was never the 94%. It was the unspoken question underneath it: what is in the 95% I never heard? For most teams that question has no answer, only a hope. Full coverage replaces the hope with a list. And often the same pass hands you something you were not even looking for. When Life Insurance Direct read every call, they also found that on 4,692 of them, no next-step recommendation was made at all, a quantified revenue gap, surfaced free, on a business where the phone call is the sale.

That is the shift. Recording is the infrastructure. Analysis is the intelligence. You already have the raw material sitting in your recordings. The only question is whether you are still reporting a number about the calls you chose to hear.

Metricsense Team

Evidence Intelligence, by Avesta Labs

The Metricsense team reads 100% of customer conversations, calls, reviews, tickets and surveys, and links every finding to the exact source quote. We write about what the full record shows once you stop sampling: compliance, quality, and the calls that never make it into a QA report.

See what is in the 95% you have never heard

Metricsense can run a week of your own calls, in your own cloud, and show you the advice-boundary split, the escalation needles, and the genuine missing-warning gap separated from the routine noise, usually within a day.

Make sense with Metricsense