Metricsense
ComplianceConduct RiskInsurance

The call that becomes a complaint is the call nobody chose to review

QA samples are not random. They are convenient. And convenience selects against exactly the calls that create conduct risk.

Gaurav Soni
Gaurav Soni
22 July 2026 · 5 min read
Share

Ask a QA lead whether their sample is random and they will say yes. Ask them how a call gets into the sample and the answer is never random.

It is the calls from the start of the week, because that is when the reviewer had time. It is the calls that are eight minutes long, because a thirty-four minute call eats a whole review slot. It is the adviser already on a coaching plan, because there is a form to complete. It is whatever came back from a complaint, because somebody upstairs asked. Every one of those is a sensible operational reason. None of them has anything to do with risk, and several of them point away from it.

I lead product at Metricsense. We read 100% of a business's customer conversations and tie every score to the exact moment in the transcript that produced it. Over four months I have been working through 7,715 analysed calls from one Australian life-insurance business, 1 April to 24 July 2026. The previous post in this series argued that a sampled QA score describes your reviewers rather than your advisers. This one is about something narrower and worse: the direction of the error.

A sample that is merely small is a manageable problem

If your sample were genuinely random, small coverage would be an honest limitation with known mathematics. You could state a confidence interval, admit that rare events are invisible to you, and get on with it. Auditors accept that. Statisticians accept that.

The problem is that a convenient sample is not a small random sample. It is a *differently shaped* one. And once the shape is wrong, more coverage does not fix it. Reviewing 4% instead of 2% of the same easy calls buys you a more precise measurement of the easy calls.

The name for it

The tendency for review capacity to flow toward conversations that are short, recent, tidy and quick to score, which are systematically not the conversations that generate conduct risk, is the easy-call bias.

ASIC has described this exact failure

In the 18 August 2025 letter closing its review of direct life insurance sales, ASIC records that many entities focused their sales and pay practice reviews on successful sales calls and failed to review a targeted number of non-converting calls. It separately notes that most life companies in the review ran fewer quality assurance checks on retention calls than on sales calls.

Read those as statements about selection and they are damning in a specific way. Reviewing successful sales tells you how your advisers behave when the customer said yes. It is structurally blind to two things: the pressure that was applied when the customer was hesitant, and the adviser who over-corrected, refused to answer a fair question, and lost a sale that should have closed. The non-converting call is where both of those live, and it is the call least likely to be reviewed, because nobody has a reason to pull it.

The same logic applies to retention. A retention call is, by construction, a conversation with an unhappy customer who is trying to leave. If you review fewer of those than sales calls, you have aimed your assurance away from your most emotionally loaded conversations.

What the difficult calls actually look like

Across 7,016 calls where the Escalation Risk check returned a result, eighteen were flagged High Concern or Critical Risk: eleven and seven. That is 0.26% of the scored floor. Separately, when the Recurring Issues check grouped what customers were actually calling about, the largest cluster by a wide margin was Medical and Underwriting.

Five themes, not 4,183 unrelated adviser failures

Medical & underwriting
1,435
Process & system issues
1,190
Price & value concerns
692
Policy & coverage decisions
472
Product understanding
394

Calls touching each theme, from one anonymised Australian life-insurance window of 7,715 calls (1 April to 24 July 2026). Excludes 3,133 calls with no recurring issue recorded and 399 marked none. A call may touch more than one theme, so the bars do not sum to the window.

Sorted by cause rather than by name, the coaching queue turns into a short list of documents to fix. But look at the top bar and ask a different question: how many of those did anyone review? Medical and underwriting calls are the longest, most technical and most emotionally difficult conversations on the floor, and in any manual operation they are among the least likely to be pulled, for the entirely human reason that they take three times as long to score.Calls touching each theme, 1 April to 24 July 2026, 7,715 calls analysed. A call may touch more than one theme, so the bars do not sum to the window.
The largest category of customer difficulty and the review process are pointed in opposite directions. That is the easy-call bias in one sentence.

Why "we review the complaints" is not an answer

The most common objection I get is that the serious calls surface anyway, through complaints.

The FCA looked directly at this in its June 2024 multi-firm review of outcomes monitoring under the Consumer Duty, across twenty larger insurance firms. It found undue reliance on complaints data and, more pointedly, that few firms could provide clear evidence of monitoring having led to proactive action.

Complaints are the tail of the tail. They require a customer to be dissatisfied enough, articulate enough, and persistent enough to escalate. A vulnerable customer who was confused by a premium explanation and quietly lapsed six months later files nothing. AFCA's calendar-year 2025 data is worth sitting with here: 111,373 complaints received, up 14% on the previous year, with complaints about misleading product or service information up 110%. The fastest-growing complaint category in the country is, almost by definition, a dispute about what somebody was told. By the time it reaches AFCA, the conversation is eighteen months old and somebody is reconstructing it from memory.

Where this argument does not apply

Bias is not conspiracy. Nobody is hiding calls. The easy-call bias is what happens when a finite review team meets an infinite queue, and it would happen in any well-intentioned operation. Naming it is not an accusation.

Coverage does not remove judgement, it relocates it. Reading every call replaces a sampling problem with a measurement problem. Every flag is model output against a definition somebody wrote, with a false positive rate and a false negative rate. It is a candidate for human review, never a finding of breach. Whether any of those eighteen represented an actual failure is a judgement for the licensee, not for a model and not for us.

One business is not a market. These are results from a single analysis window at a single Australian life-insurance business. The shape of the argument transfers. The rates do not.

The frame: risk-weighted review, built with what you have

You do not need to buy anything to fix the direction of your sample. You need to stop treating review capacity as something to spread evenly. Split next month's review budget into three tranches instead of one.

  • Tranche one, the calibration sample (say 40% of capacity). Genuinely random. Its only job is to give you an unbiased read on common behaviours. Draw it with a random number, not a reviewer's judgement, and do not let anyone substitute a call because it is inconvenient.
  • Tranche two, the structural sample (40%). Deliberately quota'd against the categories your current process avoids. Longest decile by duration. Non-converting sales calls. Retention calls. Whatever your equivalent of the medical and underwriting cluster is. You are not looking for failures here. You are looking at the part of the floor you have never seen.
  • Tranche three, the reactive sample (20%). Complaints, escalations, anything anyone raised. This is what most teams currently spend most of their capacity on, and it should be the smallest slice, because it is the only one guaranteed to be backward-looking.

Then run the one diagnostic that makes the whole thing worth it: compare the pass rate of tranche two against tranche one. If your structural sample scores materially worse than your random sample, you have measured your own blind spot, in your own data, with your own people. That number is the most useful thing a QA function can put in front of a risk committee, and it does not require a vendor.

Four questions for the next risk meeting

  1. Of the calls we reviewed last month, what share were chosen by a rule rather than by a person?
  2. What is our review rate on non-converting sales calls compared with successful ones?
  3. What is our review rate on retention calls compared with sales calls?
  4. What is the longest call anyone in this business has reviewed end to end in the last quarter?
Gaurav Soni
Gaurav Soni

Head of Product, Metricsense

Gaurav Soni leads product for Metricsense. He has spent 13 years in product at the intersection of data, analytics and AI, starting in engineering, building APIs, data pipelines and analytics products, and moving steadily closer to the decisions leaders actually make. The thread through all of it is the same: turning messy, unstructured conversation into evidence a team can act on. At Metricsense that means reading 100% of a company's calls, reviews and tickets, and linking every finding to the exact quote.

See the part of the floor you have never reviewed

Metricsense reads every conversation rather than a sample, scores both the customer side and the adviser side against metrics you write in plain English, and links every score to the exact transcript moment behind it. It runs in your own cloud and region. A pilot is worth your time if you are a licensee or distributor with more than roughly 1,000 recorded conversations a month, you already have a compliance or quality question you cannot currently answer, and someone can spend an hour a week confirming flags. Below that volume, the three-tranche frame above will get you most of the way and I would rather you just ran it.

Make sense with Metricsense