Metricsense
ComplianceQuality AssuranceInsurance

Your QA score only measures the calls you chose to hear

A 93% pass rate is a fact about a few hundred calls somebody selected. ASIC put the problem in writing in August 2025, and named the fix in the same letter.

Gaurav Soni
Gaurav Soni
11 August 2026 · 5 min read
Share

In August 2025 the regulator wrote to life insurers and their distributors and listed, as a deficiency, that some life companies were reviewing only a small percentage of their sales calls.

Why it matters: that sentence quietly moves your sampling rate out of the operations budget and into the conduct file, because it is now something ASIC has described in writing, to your industry, in its own words.

The same letter names the improvement. Several life companies now quality assure all sales calls instead of a small sample, using technology such as AI-powered speech analytics. The regulator is not speculating about what good looks like. It has seen it and written it down.

I lead product at Metricsense. The corpus behind the arithmetic below is 9,218 calls from one anonymised Australian life-insurance business, 16 April to 10 August 2026, every one of them read against the same checks. What you should have by the end is the actual number of calls on your own floor that nobody has ever reviewed, and six questions to take into your next QA meeting.

The regulator put this in writing in August 2025

The document is ASIC's letter on improving the direct sale of life insurance, dated 18 August 2025 and addressed to life insurers and distributors. It is short, it is public, and most compliance leads I speak to have read it.

Several life companies now quality assure all sales calls, instead of a small sample.

Notice what that does to the conversation. You are no longer arguing that full coverage is a good idea. You are explaining why your business is on the side of the comparison the regulator described as the deficiency, which is a much harder paper to write.

This is also why I will not quote you the widely circulated claim that only 1 to 3% of calls get reviewed anywhere. I cannot find a primary source for it, and the reader I am writing for is exactly the person who would go and check.

Sampling arrived because it was possible, not because it worked

Be fair to the method before dismantling it. When a QA analyst had to sit with a headset and a scoring sheet, listening to eight calls a day, a 2 to 5% sample was not a compromise. It was the ceiling of what a human being could physically do.

Everything built on top of it inherited that ceiling: the calibration sessions, the monthly scorecard, the pass rate on the board pack. None of it was designed around what gives you assurance. It was designed around what one person could listen to in a day, and then it hardened into a standard.

The constraint is gone. The standard is still here, which is the ordinary way institutions end up defending a number that no longer means what it used to.

The arithmetic on 9,218 calls takes one line

Take the corpus I have been reading: 9,218 calls across the window. At a 3% sample that is 277 calls reviewed, which is a real amount of listening and still almost none of the floor.

19 of 8,460

scored calls were recorded as high concern or critical escalation risk. A 3% sample expects to reach 0.57 of them.

Here is the working, and you can check it on a phone. Each of the 19 has a 3% chance of being in the sample, so the expected number reached is 19 times 0.03, which is 0.57. The chance of reaching none of them at all is 0.97 to the power of 19, which is 0.56. More likely than not, a diligent 3% sample run across four months touches none of the calls that carried the most risk.

Run the same line on your own volumes before you accept mine. You need three numbers you already have: calls per week, your sample rate, and how many calls you would call genuinely serious in a quarter.

The name for it

The unreviewed floor: the count of calls in a period that no human and no system has assessed against any check. It is not a rate, it is a number of calls, and it is usually the largest number in the QA pack that nobody has ever written down.

Your score measures your selection, not your floor

A 93% QA pass rate is a true statement about 277 calls. It is not a statement about 9,218, and the gap between those two things is where the assurance quietly disappears.

It gets worse if the sample is not random, and almost none are. Analysts pick calls that are available, that are the right length, that belong to advisers already under review, or that landed in a week when somebody had time. Every one of those is a reasonable operational choice, and every one of them makes the score a fact about the choosing.

Full coverage produces a shorter list, not a longer one

1
Screened100% of screened

every call in the window, read against the same checks

2
Obligation not engaged62% of screened

transferred, no product information given, caller rang off

3
Candidate for review22% of screened

the check's conditions were met, so a person needs to look

4
Confirmed by a reviewer9% of screened

a person looked at the evidence and agreed

5
Reportable situation assessment4% of screened

a legal judgement, made by the people who already make it

no human judgement requireda person decidesyour people, your scorecards

Illustrative proportions. The shape of the funnel is the claim, not these percentages.

Full coverage does not hand your reviewers every call to listen to. It hands them a filtered list, and most of a corpus never reaches a human at all.Illustrative proportions. The shape of the funnel is the claim, not these percentages.

Exposure arrives one call at a time, never in aggregate

This is the part the aggregate score is structurally unable to represent. Nobody complains about your pass rate.

A complaint arrives attached to one conversation, on one afternoon, between one customer and one adviser. AFCA will ask what was said on that call. Your answer is either the recording plus a record of what your process assessed, or it is a monthly average that has nothing to say about the specific twelve minutes in question.

That asymmetry is the whole argument. Assurance is measured in aggregate and consumed one call at a time.

Where this argument does not apply

Coverage is not accuracy. Reading 100% of calls badly is worse than reading 3% of them well, because it produces confident numbers with nothing behind them. If you cannot open the transcript moment behind a score, more coverage just means more unfounded scores.

Every count above is model output. The 19 are calls our scoring recorded as high concern or critical, not calls a human has confirmed carried real risk, and certainly not breaches. That determination belongs to your compliance process, and to the regulator, and to nobody else.

One business, one window. These are 9,218 calls from a single Australian life-insurance business between 16 April and 10 August 2026. The arithmetic transfers. The rates almost certainly do not, which is why the useful version of this post is the one you run on your own numbers.

Six questions for your next QA review

  1. How many calls did we take last quarter, and how many did anybody assess? The difference is your unreviewed floor. Write it on the board pack.
  2. Is our sample actually random, or is it whatever was convenient? Ask the analyst who pulls it, not the person who reports it.
  3. If AFCA asks about one specific call from March, what can we produce within a day?
  4. What is our pass rate a rate of, precisely? Name the population out loud.
  5. Which advisers had zero calls reviewed last quarter? There will be some.
  6. If we found something serious on an unreviewed call from four months ago, what would we do with it? Decide that before you go looking.

What I would do first, and it costs nothing

Do not buy anything yet. Pull your last 40 scored calls and count how many of the scores link to a specific quoted moment in the transcript. If it is under half, your QA pack is a collection of opinions with a number attached, and full coverage would only give you more of them.

Fix that first. The coverage argument only becomes worth having once each score can be traced back to the thing that produced it.

Gaurav Soni
Gaurav Soni

Head of Product, Metricsense

Gaurav Soni leads product for Metricsense. He has spent 13 years in product at the intersection of data, analytics and AI, starting in engineering, building APIs, data pipelines and analytics products, and moving steadily closer to the decisions leaders actually make. The thread through all of it is the same: turning messy, unstructured conversation into evidence a team can act on. At Metricsense that means reading 100% of a company's calls, reviews and tickets, and linking every finding to the exact quote.

Measure your own unreviewed floor

We will read a window of your own calls, in your own cloud and region, and give you the count of calls nobody has ever assessed alongside what the checks found in them. Relevant if you hold an Australian advice or distribution licence and your QA pack currently reports a rate rather than a population. If your board is comfortable with the sample you have, this will not be a comfortable conversation.

Make sense with Metricsense