When everyone makes the same mistake, it is the script
One adviser getting it wrong is a coaching conversation. The same gap across the whole floor is a wording problem, and coaching will not touch it.
Read a few hundred of these calls back to back and the pattern that stands out is not variation between advisers. It is how alike the calls are: the same subject, arriving at the same point, in almost the same words.
Why it matters: every QA process I have seen routes that pattern to individual coaching, which is the one intervention guaranteed not to fix it.
I am the tech lead at Metricsense. I built a lot of the scoring that reads these conversations, and the corpus here is 9,218 calls from one anonymised Australian life-insurance business, 16 April to 10 August 2026. What you should take away is a two-minute test that tells you which kind of problem you have, before you book anybody a coaching session.
One adviser wrong is coaching. Everyone wrong is wording.
These are different problems with different fixes and they are almost impossible to tell apart one call at a time, which is exactly how QA looks at them.
A skill gap belongs to a person. It has a cause you can name, it responds to practice, and it shows up unevenly: some people have it and some people do not. A script gap belongs to the wording everyone was trained on. It shows up evenly, it is invisible to the person making it, and no amount of practice removes it, because everybody is executing the instruction correctly.
The tell is the distribution, and a sample of eight calls a week cannot show you a distribution.
Five themes carried 5,047 calls, not 5,047 separate failures
When we group what these conversations were actually about, the long tail people expect does not appear. The corpus concentrates.
Five themes carry 5,047 calls, not 5,047 separate problems
Calls by recurring theme, one anonymised Australian life-insurance business, 16 April to 10 August 2026. Scored population 5,736, of which 689 were recorded as having no recurring issue. A further 3,475 calls in the window were not scored on this dimension. Queried 11 August 2026. Model-derived groupings, not confirmed findings.
scored calls circled medical and underwriting. That is a queue about one subject, not a thousand separate one-off problems.
Concentration like this is the first evidence that you are looking at a system rather than at people. The interesting question is not which advisers handled those calls badly. It is what your medical and underwriting explanation actually says, and whether it can be followed on the phone by somebody who is anxious and has never bought life insurance before.
The script signature: a defect whose rate barely varies across the people executing it. If your worst performer and your best performer produce nearly the same rate on a check, you are not looking at a skill distribution. You are looking at the instruction they share.
A people problem has a spread. A script problem is flat.
This is the whole diagnostic, and it needs no new technology to run.
On a genuine skill gap, per-person rates fan out. Somebody is near zero, somebody is much worse than the average, and the shape of that fan is where your coaching budget should go. On a script gap, the rates bunch. Your top adviser and your median adviser land within a few points of each other, and critically, nobody is clean. That floor is the signature.
If nobody on the floor is near zero, you are not measuring the floor. You are measuring the wording it was given.
The two-minute test on your own QA sheet
You can run this today with a spreadsheet and whatever review data you already have. It is crude, and crude is enough to tell the two cases apart.
- Pick one check that keeps failing. One. The one your QA lead complains about most.
- For each adviser with at least ten reviewed calls, work out their failure rate on that check alone.
- Sort the rates and look at the lowest one, not the average. The lowest rate is the number that matters.
- If the best person on the floor is near zero, it is a skill distribution: coach it, and start with the people furthest from that floor.
- If the best person on the floor is nowhere near zero, stop coaching and go and read the script. Everybody is doing what they were told.
The honest caveat on your own version of this test: with ten reviewed calls per adviser the rates are noisy, and you should treat a difference of a few points as nothing at all. You are looking for the difference between a fan and a huddle, not for a ranking.
Coaching a script problem is how you lose good advisers
This is the cost people underestimate, and it is the reason I care about the distinction more than the measurement.
Put an adviser in a coaching session for a failure that every one of their colleagues also produces, and you have told a competent person they are deficient at something they did correctly. Do it twice and they stop trusting the QA process. Do it across a floor and you have taught your best people that the scorecard is arbitrary, which is a much more expensive problem than the original defect.
There is a second cost, quieter and worse. Coached hard enough on a compliance-shaped check, advisers stop explaining things rather than risk explaining them wrong. Over-blocking does not show up on a compliance report, because nothing went wrong. It shows up in the sale that did not happen.
What this does not tell you, including about our numbers
The themes above are model-derived groupings. They are what our scoring recorded a call to be about, not a human-confirmed classification, and the boundary between "product understanding" and "policy and coverage decisions" is a judgement the model is making on your behalf. Read the grouping as a map, not as a finding.
A flat rate is evidence, not proof. A defect can also look flat because every adviser was trained by the same person, or because the underlying task is genuinely hard for everyone. The test tells you to go and read the script. It does not tell you the script is wrong.
One business, one window, one product line. 9,218 calls from a single Australian life-insurance business over four months. The diagnostic transfers to any floor with a script. The concentration almost certainly does not.
Four questions for your next operations meeting
- On our most-failed check, what is the failure rate of our best adviser? If nobody can answer, that is the finding.
- When did anyone last read the actual wording of that part of the script out loud, on a phone, to somebody who does not work here?
- How many of last quarter's coaching sessions were about a defect that the whole floor produces?
- Where might our advisers be over-explaining or refusing to explain, because coaching made the check feel dangerous? Nobody reports this, so you have to ask.
Fix the sentence, then coach the difference
The order matters more than either step. Rewrite the wording first, let it run for a month, and then look at what is left. What survives a script fix is the real skill distribution, and it is usually much smaller and much more coachable than the original number suggested.
One thing I have not worked out: how long a rewritten script takes to show up in the calls. I expected a week. On the changes I have watched, it has been closer to a month, and I do not yet know whether that is the retraining, the habit, or something about how these conversations get started.

Tech Lead, Metricsense
Bhautik Desai is the tech lead at Metricsense, where he builds the systems that read customer conversations and link each finding back to the exact quote. He writes about how the product works under the hood and what it takes to run this kind of analysis on real data.
Find out whether it is your people or your wording
A pilot reads a window of your own calls and gives you the per-check rate across everyone on the floor, with the transcript moment behind each one, so the fan-or-huddle question answers itself. Worth doing if your coaching log keeps repeating the same defect. If you can already run the two-minute test on your existing QA data, run that first and save yourself a procurement cycle.
Keep reading
Avesta Labs is an OpenAI Select Partner. Your calls stay put.
The company that builds Metricsense has joined the OpenAI Partner Network. Here is the short list of what that changes for the recordings you have trusted us with, and the longer list of what it does not.
A confidence score is not evidence
0.94 tells you how sure a model is. It tells you nothing about what your adviser said. Only one of those survives a regulator.
An ask was recorded on 34 of 8,479 scored calls
Read end to end rather than sampled, the clearest revenue pattern in this corpus was not a bad pitch or a hard objection. It was a conversation that simply stopped.