Metricsense
ComplianceQuality AssuranceInsurance

Your QA sample has an adverse selection problem

The calls most likely to contain a compliance problem are the least likely to be in your sample. That is not bad luck. It is adverse selection, and every insurer already knows exactly how it works.

Gaurav Soni
Gaurav Soni
23 July 2026 · 6 min read
Share

Every insurer understands adverse selection in their sleep. It is the first risk you underwrite against: the customer most eager to buy income protection is, a little too often, the one who already suspects they will need it. Price for the average life and the higher risks quietly find their way onto your book. It is not bad luck. It is the structure of the thing.

Now turn that same lens on your own quality assurance, because the identical mechanism is running there, and almost nobody has named it. When a QA analyst reviews a handful of calls per agent each month, the calls most likely to contain a compliance problem are the calls least likely to be picked. The risk self-selects out of your sample, for reasons every bit as structural as the ones your underwriters price for.

The reframe: your QA sample is not just too small. It is biased against the exact calls that matter. The features that make a call risky are, over and over, the same features that make a human reviewer skip it.

Call it what your own underwriters would call it: adverse selection in the QA sample. A process that feels neutral, pull a few calls, score them, extrapolate, quietly loaded against the calls you most need to hear.

Why the risky call skips the sample

Start with what a reviewer actually reaches for. Consciously or not, people pull the calls that look normal: a familiar queue, average length, clean audio, an agent they already have a feel for. The risky call is frequently the outlier, the twenty-two-minute struggle, the transfer, the callback of a callback, the one that looked like nothing on the surface metrics, short handle time, customer said “fine”, ticket closed, so nobody gave it a second listen.

Then add the part that makes it worse: the agent who made the error rarely raises their hand. An agent who confidently drifts from general advice into personal advice does not think they did anything wrong, so they never nominate that call for review. The calls that get self-reported are the ones where the agent already knew there was trouble. The quiet, confident mistake, the most dangerous kind, is exactly the one nobody flags.

The segment your sample covers zero percent

There is one slice where the selection is not partial but total. Every multilingual insurer has it: calls conducted in a language the QA reviewer does not speak. Those calls are not under-sampled. They are not sampled at all. The coverage rate for them is not 2 to 5%. It is zero.

And it falls to zero in precisely the place compliance risk runs highest: first-language-not-English customers, often from migrant communities, navigating stepped premiums, duty of disclosure and general-advice warnings for the first time. If an agent handling those calls is mishandling disclosure, nobody on the QA team can even read the transcript to find out. There the blind spot is not thin. It is a wall. Reading every call, in the languages your customers actually use, is the only thing that turns that wall back into a window (validated on your own languages during a pilot).

0%

the QA coverage rate for calls in a language your reviewers do not speak. Not a thin sample, no sample at all, in exactly the segment where disclosure risk tends to run highest.

A random sample does not save you either

The natural reply is: fine, we will sample randomly, and then it is representative. But adverse selection defeats random sampling too, for a plain reason, the risky call is rare. In a full window of 5,438 calls, only a handful, on the order of sixteen, carry genuine high or critical escalation risk. A truly random 5% sample would, on average, be expected to catch fewer than one of them. Representative of your ordinary call, blind to the exceptional one. And human sampling is rarely truly random in any case, because reviewers are people, and people quietly avoid the awkward twenty-five-minute call in a language they only half-follow.

A representative sample is representative of your ordinary calls. The call that gets you fined is not ordinary.

What reading everything actually surfaces

Read every call and adverse selection has nowhere to hide, because there is no selection left to be adverse. The rare escalation is caught not because someone guessed to pull it, but because everything is read the same way, against the same checks. And the patterns a sample scatters into noise resolve into a map. Across a single window of 5,438 calls, the recurring issues grouped into a few themes, Medical and Underwriting on 970 calls, Process and System friction on 893. That is not nearly 1,900 individual failures. It is a small number of systemic gaps repeated at scale, the kind of pattern a sample can never assemble because it only ever holds fragments.

And each flag arrives with the receipt attached. Here is the kind of call full coverage catches and a sample skips, the sort that looks completely fine on every surface metric, an illustration rather than a real call:

CallerHonestly the premiums have gone up so much, I have had to borrow from my daughter just to keep it going.
AgentI understand. I will email you the updated schedule and we will leave it there for today.
FlagCandidate: financial-vulnerability signal not escalated. Short, polite, closed on time, and exactly the call a sample would pass over. For human review.

Nothing about that call trips a surface metric. It is short, courteous, resolved on time, and it would sail past any reviewer scanning for the obvious. The content is the only place the risk lives, and the content is the one thing a sample almost never reaches.

The honest limits

Reading everything is not a magic wand, and three limits are worth stating plainly, because a claim without its caveats is the kind of overselling a compliance buyer sees through in a second.

First, every flag is a candidate issue with evidence, not a finding of breach. A human reviewer confirms it; the tool reports nothing to a regulator on its own. You keep your scorecards and your judgement, you just stop choosing which calls to trust blind.

Second, analysis reads transcripts after the call, not live. It makes the next coaching conversation precise; it does not referee the call in the moment.

Third, it reads what was said. Tying a specific call to a downstream outcome in another system, a lapse, a complaint, a chargeback, needs your own data joined to the call content, and that join stays yours to make.

You already know how to think about this

Adverse selection is not a new enemy. It is one your business was built to reason about, and you already know the answer, because you use it every day: you do not solve adverse selection by pricing the average customer harder, you solve it by seeing the whole pool. The same is true of your calls. The fix for a sample that is structurally biased against the risky call is not a cleverer sample. It is no sample at all.

Recording is the infrastructure. Analysis is the intelligence. The calls are already sitting in your system, the risky ones included. The only question left is whether you keep reviewing the pool through a keyhole that, by its very nature, looks away from the calls that matter most.

Gaurav Soni
Gaurav Soni

Head of Product, Metricsense

Gaurav Soni leads product for Metricsense. He has spent 13 years in product at the intersection of data, analytics and AI, starting in engineering, building APIs, data pipelines and analytics products, and moving steadily closer to the decisions leaders actually make. The thread through all of it is the same: turning messy, unstructured conversation into evidence a team can act on. At Metricsense that means reading 100% of a company's calls, reviews and tickets, and linking every finding to the exact quote.

Find the calls your sample selects against

Metricsense can read a week of your own calls, in your own cloud and in the languages your customers actually use, and surface the rare escalation, the disclosure gap and the confident mistake that a sample is built to miss, usually within a day.

Make sense with Metricsense