Metricsense
Method against method

AI call monitoring vs manual call QA sampling for Australian financial services licensees

Manual call QA sampling reviews a hand-picked share of calls. AI call monitoring reads all of them and links every score to the transcript moment that produced it. The difference is not accuracy. It is selection: which calls anyone hears, and what you can evidence about the rest.

Sampling is a selection problem before it is a volume problem. That is the argument here. No company is named on this page: the villain is the method, and one question below is about when the method beats the software.

When is sampling still the right choice?

More often than a vendor page will admit. If a reviewer can hear a meaningful share of your calls, sampling is adequate and cheaper, and software buys paperwork rather than coverage. If your judgement calls are subtle and your reviewers experienced, a human on a sample beats a model on everything.

Where sampling wins, said plainly.

  • Low volume. A few hundred calls a year is a corpus one person can work through. Coverage is not your problem.
  • A brand new check. Before a check is worth running across every call, someone has to listen and work out what they are looking for. That work is human and it comes first.
  • Nuance you cannot write down. If you cannot write the standard down, no system can run it, and the honest next step is more listening, not more automation.

If you are in one of those three, we would rather say so than sell you something. The rest of this page is about selection at volume.

Has the regulator said anything about sampling sales calls?

It has, in writing, to chief executives, and it named sampling as the deficiency. In August 2025 ASIC wrote to the chief executives of Australian life insurers and distributors about the direct sale of life insurance.

Some life companies only reviewed a small percentage of sales calls, relying on manual reviews to identify compliance issues.

ASIC, "Improving the direct sale of life insurance". Dear CEO letter, 18 August 2025. Read the letter (PDF).

What changes when every call is read instead of a sample?

The selection rule disappears, and with it the argument about whether the right calls were chosen. A recurring problem stops looking like a run of individual mistakes and starts looking like one wording problem repeated across a floor.

Manual QA samplingReading every call
How calls are chosenA selection rule, usually availabilityNo selection rule to defend
Non-converting callsRarely in the setIn the set by default
Adding a new conceptRetrain reviewers, rescore by handBuild the concept, rerun the window
Where sampling is betterLow volume, a new check, or judgement you cannot yet write downOverkill: paperwork, not coverage
  • Most selection rules are some version of availability: who had capacity, which team was under review, which calls a complaint already reached. The call nobody complained about, that did not convert and that nobody had time for, is invisible three times over.
  • Where a concept comes from: you document how your team judges a call, we assess that document and build the concept from it, and you validate the output.

What does reading every call not do?

Four things, and a vendor page that lists none of them is selling the wrong picture. Coverage is not accuracy: reading every call badly is worse than reading a handful well.

The limits.

  • A score is a model output, not a fact about what an adviser did. We can say what was recorded, not what a person meant by it.
  • We publish no accuracy figure. It gets measured on your calls, in the pilot, against your own QA.
  • Full coverage decides nothing. Metricsense makes no compliance determination and reports nothing to a regulator.
  • Language coverage beyond English is validated during the pilot, rather than claimed up front.

What should you do next, whether or not you talk to us?

Write down how last quarter's reviewed calls were chosen, then count how many of them did not convert. Those two numbers are this argument measured on your own floor, and neither needs a vendor.

Life Insurance Direct, an Australian life-insurance broker holding an AFS licence, moved from hand-picked reviews to full coverage: 10,000+ calls read and scored on both sides, inside their own sovereign cloud.

Run it against your own last quarter

The pilot runs on around 50 calls you already hold, with a readout inside a week. If it does not hold up against your own QA, you keep the readout and we go away.

The calls stay in your own sovereign cloud, in your own region and tenancy, and PII is stripped before any transcript reaches the model. Your security reviewer will want data governance, form-free.

Feedback intelligence
Make sense with MetricsenseAICPA SOC
© 2026 Metricsense. Built and operated by Avesta Labs.