Is it an adviser problem, or is the script teaching everyone the same mistake?
When the same mistake shows up on agent after agent, coaching each of them is the expensive way to fix nothing. The cause is usually upstream, in the script, not in the people.
It is the weekly QA review, and a team lead is looking at the same red flag for the ninth time. Nine different agents, the identical miss: none of them told the customer that a stepped premium keeps climbing every year. The instinct, the one every scorecard trains, is to book nine coaching conversations.
Before booking a single one, there is a question worth asking, and it changes everything: if nine agents made the same mistake, is that nine agent problems, or one?
The reframe: a shared mistake is not a pile of individual failures. It is a single systemic fault, repeated. Coaching the individuals for it is the expensive way to fix nothing.
Quality control settled this argument decades ago, and the discipline your QA team descends from already has the words for it. W. Edwards Deming drew a hard line between common-cause variation, faults built into the system that show up for everyone, and special-cause variation, the genuine one-off, the individual outlier. His warning was blunt: blame a worker for a common-cause fault and you will coach forever and change nothing, because the fault was never theirs to fix.
The signpost, not the drivers
Picture a junction where nine drivers in a row take the wrong exit. You could pull each one aside and retrain them on directions. Or you could walk up to the signpost, notice it is pointing the wrong way, and fix it once. Nine mistakes, one cause. Retraining the drivers is not just slower, it is a category error: they read the sign correctly. The sign was wrong.
A script that tells agents to describe a premium without ever mentioning it can rise is that signpost. Every agent who follows it faithfully produces the same gap. They are not failing; they are complying, with the wrong instruction. Coaching them one by one treats a signpost problem as a driver problem, and the signpost keeps pointing the wrong way the whole time.
Why the scorecard makes you blame the person
The reason teams reach for individual coaching is not carelessness. It is the shape of the tool. A scorecard is built per agent: one row per person, a number in each. Look at the world through per-agent rows and every failure looks personal, because the format has no column for the script. You cannot see a systemic cause in an instrument designed to grade individuals. The pattern stays invisible until you step back from the people and look at the calls themselves.
You can only tell cause from person at 100%
This is where coverage stops being an abstraction. With a 2 to 5% sample, you cannot tell the two apart. A handful of calls per agent is far too few to know whether a mistake is one person's habit or the whole floor's script; the numbers are too small to separate the pattern from the person. So teams guess, and guessing usually means blaming whoever is in the row.
Read every call and the distinction resolves. Across a single window of 5,438 calls, the recurring issues did not scatter into thousands of individual failings; they grouped into a handful of themes, Medical and Underwriting on 970 calls, Process and System friction on 893, with product-understanding gaps concentrated again on their own set. That shape is the tell. A problem that lands on hundreds of calls across many agents is common cause: a script, a process, a training gap. A problem that lands on two agents and nobody else is special cause: coach those two. Same data, and for the first time you can see which is which.
nine agents, one shared miss. Read as nine coaching problems, it costs nine sessions and fixes little. Read as one script gap, it is a single change that lifts the whole floor at once.
The coaching economics
Getting this backwards is not an abstract error, it is spent selling time. Treat a common-cause script gap as an individual failing and you book coaching for every agent who ever hit it, hours of the floor pulled off the phones, most of it re-teaching the agents who already had it right. Fix the script instead and the gap closes for everyone in one change. It runs the other way too: pull the whole floor into a session to fix what is really two agents' habit, and you have spent everyone's morning on a two-person problem. Precision is the entire game. The same mistake in twenty mouths is one fix; a different mistake in one mouth is another; you have to be able to tell them apart to spend your coaching where it actually pays.
The same mistake in twenty mouths is one problem, not twenty. Coach it twenty times and you have fixed it none.
The honest limits
Three limits, because the reframe cuts both ways.
First, not every shared pattern is a script fault, and not every individual flag is a system fault. Some agents genuinely are the outlier and need direct coaching. The point is not to stop coaching people; it is to stop coaching people for the system's mistakes. Telling the two apart is the real work, and it needs the per-agent, per-call evidence, not a hunch.
Second, the tool surfaces the pattern; it does not decide the fix. Whether a common-cause gap is best closed by a script change, a training module or a process tweak is a human judgement. Metricsense shows you that nine agents share one gap; you decide what to do about the signpost.
Third, every flag is a candidate for review, not a verdict, and tying a coaching change to a downstream result, better conversion or fewer complaints, needs your own outcome data joined to the call content.
Person, or pattern?
So the next time the same red flag lands on agent after agent, ask the question before you book the coaching: person, or pattern? Common cause, or special? Get it right and one script fix quietly lifts the whole floor. Get it wrong and you will hold the same coaching conversation forever, wondering why the number never moves.
Recording is the infrastructure. Analysis is the intelligence. The calls already hold the answer to whether it is your agents or your script. The only thing missing is a way to read all of them, so the pattern can finally show itself.

Tech Lead, Metricsense
Bhautik Desai is the tech lead at Metricsense, where he builds the systems that read customer conversations and link each finding back to the exact quote. He writes about how the product works under the hood and what it takes to run this kind of analysis on real data.
Tell the script problem from the agent problem
Metricsense reads a week of your own calls, in your own cloud, and shows you which mistakes are one agent's habit and which are the whole floor reading the same script, so you coach where it actually pays, usually within a day.
Keep reading
A confidence score is not evidence
An AI can tell you a call was 92% compliant. It cannot, on its own, tell a regulator what the agent actually said. In a dispute, only one of those survives.
Your QA sample has an adverse selection problem
The calls most likely to contain a compliance problem are the least likely to be in your sample. That is not bad luck. It is adverse selection, and every insurer already knows exactly how it works.
Your QA score only measures the calls you chose to hear
A 94% compliance rate sounds like a fact about your call floor. It is a fact about your sample. In a regulated contact centre, that gap is where the fines live.