Metricsense
ComplianceConduct RiskInsurance

The complaint call is the one nobody chose to review

QA samples are not random. They are convenient. And convenience selects against exactly the calls that carry conduct risk.

Gaurav Soni
Gaurav Soni
11 August 2026 · 5 min read
Share

Your QA sample is not random, it is convenient, and convenience has a direction.

That matters because a sample which is merely small can be fixed with arithmetic and a bigger budget. A sample biased against the calls you most need to see cannot be fixed by making it bigger.

Nobody designed it this way. It is the accumulated residue of a hundred sensible decisions about which call to open when you have forty minutes before the calibration meeting.

I lead product at Metricsense, and for the past four months I have been reading a corpus of 9,218 calls from one anonymised Australian life-insurance business, 16 April to 10 August 2026. This post has no new headline number in it. What it has is a mechanism you can check against your own review process this afternoon.

A small sample is manageable. A biased one is not.

Sampling theory is generous to you when the sample is random. A random 3% still misses most of what happens, but what it catches is representative, and you can put honest error bars around it.

Bias breaks that guarantee completely. If the selection process is systematically less likely to pick difficult calls, then increasing the sample from 3% to 6% does not halve your blind spot. It gives you twice as many of the calls you were already seeing.

The name for it

The easy-call bias: the tendency of any human-selected review sample to over-represent short, clean, well-recorded, recently-closed calls, because those are the ones that are cheapest to review. It is a property of the selection step, not of the reviewer, and no amount of reviewer diligence removes it.

19 calls in 8,460 were scored high risk, and they hide

In the window I have been reading, 19 of 8,460 scored calls came back as high concern or critical escalation risk. That is roughly one call in 445.

Two things follow, and the second one is the awkward one. A rate that low means a sample is unlikely to reach them: at 3% the expected number found is 0.57. And a rate that low also means I cannot tell you much about them as a group, because 19 is a small number and I am not going to pretend otherwise. What I can tell you is that they did not cluster in the places a reviewer would have looked.

They are also flags, not findings. Our scoring recorded escalation risk on those 19 calls. Whether any of them represents an actual problem is a judgement for the business that owns the licence, and it is a judgement I have deliberately not made here.

Short, clean and early is what a reviewer reaches for

Watch how a review sample actually gets built and the bias stops being theoretical. The analyst has a list, a scoring sheet and a finite afternoon.

  • Length. A four-minute call gets reviewed. A thirty-one-minute call with a confused customer, two hold periods and a transfer gets skipped, and it is skipped by the same person every week.
  • Audio quality. A call with a poor line or heavy background noise is genuinely painful to score, so it goes back on the list. Difficult conversations disproportionately happen on difficult lines.
  • Recency. Reviewers work from the top of a list sorted by date. The oldest unreviewed calls in a quarter stay unreviewed, permanently.
  • Who took it. Once an adviser is on a watchlist, their calls get pulled repeatedly, which mechanically reduces coverage of everyone else.

Random sampling misses the risky call because it is rare. Targeted sampling misses it because it has no signal to aim at.

Random 5% sample

0 of 5 caught

Targeted sample

2 of 5 caught

Read every call

5 of 5 caught

ordinary call (n=145)call scored high risk (n=5)reviewed
Random sampling misses the risky call because it is rare. Targeted sampling misses it because the calls that carry risk often carry no signal to aim at.Illustrative, and deliberately exaggerated: the grid shows 5 risky calls in 150 so the shape is visible. The measured rate in this corpus was 19 of 8,460 scored calls, about one in 445, which would be less than half a dot.

Reviewing complaints means reviewing the ones that already failed

The standard defence is that the difficult calls do get looked at, because complaints trigger a review. It is a fair point and it does not survive contact with the timeline.

A complaint-triggered review happens after the customer has already been harmed enough to complain, and after the licensee has already been exposed. It is a fine control, but it is a remediation control. Calling it assurance is the category error, and it is the most common one I meet in a risk meeting.

Complaints tell you about the calls that failed loudly. Conduct risk lives in the ones that failed quietly.

The quiet failure is the customer who was told something not quite right, accepted it, bought the policy and went away satisfied. Nothing about that call will ever put it on a review list. It surfaces three years later at claim time, which is precisely when it is most expensive.

One segment where coverage is zero, not thin

There is a category of call that no sampling regime I have looked at reaches at all, and it is not a small one: the calls that produced no sale, no complaint and no follow-up.

They are invisible to every trigger you have. There is no application to audit, no customer to survey, no complaint to investigate, and no revenue attached to make anyone curious. So they are not under-sampled. They are structurally unsampled, and in most books they are a large fraction of the floor.

This is the category I would look at first if I were you, and not for compliance reasons. A call that ends in nothing is the cheapest possible place to find out that your explanation of a product does not work, because nobody has bought anything yet and nobody is upset. It is the only part of the floor where a defect and a fix can occupy the same week.

I do not know what is in yours. That is rather the point.

Risk-weighted review, built with what you already own

You do not need to buy anything to reduce the bias. You need to stop letting convenience choose, which mostly means writing the selection rule down so it can be argued with.

  1. Stratify by length. Reserve a fixed quota of the sample for the longest decile of calls. It will be unpopular with whoever reviews them, and it is the single highest-yield change on this list.
  2. Reserve a quota for zero-outcome calls. No sale, no complaint, no follow-up. Nobody looks at these, so make somebody look at a few.
  3. Sample the oldest unreviewed calls, not the newest. Work the list from the wrong end once a month.
  4. Cap per-adviser reviews. Once an adviser has had four calls reviewed in a month, the fifth slot goes to someone with zero.
  5. Record why each call was selected. One field, one word. Within a quarter you will be able to see your own selection bias in a pivot table, which is far more persuasive internally than this post.

Six questions for the next risk meeting

  1. Who chooses which calls get reviewed, and what is written down about how?
  2. What is the median length of a reviewed call, against the median length of all calls? If reviewed calls are shorter, you have measured your own bias.
  3. How many calls with no sale and no complaint did we review last quarter?
  4. What is the oldest call in the last quarter that nobody has ever assessed?
  5. If a call from four months ago turned out to carry a real problem, who would we tell, and by when?
  6. What would change on Monday if we found one? Answer this before you go looking, not after.

You cannot sample your way to knowing what you missed

Every argument in this post is really one observation: a sample can only ever describe itself. The calls you did not pick are not a smaller version of the calls you did, because the picking was never neutral.

The check to run is not on the calls. It is on the choosing.

Gaurav Soni
Gaurav Soni

Head of Product, Metricsense

Gaurav Soni leads product for Metricsense. He has spent 13 years in product at the intersection of data, analytics and AI, starting in engineering, building APIs, data pipelines and analytics products, and moving steadily closer to the decisions leaders actually make. The thread through all of it is the same: turning messy, unstructured conversation into evidence a team can act on. At Metricsense that means reading 100% of a company's calls, reviews and tickets, and linking every finding to the exact quote.

See the part of the floor you have never reviewed

A pilot reads a window of your own calls, in your own cloud and region, and shows you what is in the calls your sample has never reached, with the transcript moment attached to every flag. Useful if you already suspect your review list is chosen by convenience. If you have run the stratification changes above and they closed the gap, you may not need us at all.

Make sense with Metricsense