A confidence score is not evidence
An AI can tell you a call was 92% compliant. It cannot, on its own, tell a regulator what the agent actually said. In a dispute, only one of those survives.
A customer rings the ombudsman. Six weeks ago, they say, your agent told them their income protection would pay out if they were made redundant. It does not; that is not what the cover is for. Now there is a complaint, an eight-week clock, and one simple question on the table: what did your agent actually say?
You open your QA tool. It has scored the call: 92% compliant, green tick, no flags. Reassuring, until you notice it does not answer the only question that matters. It cannot show you the sentence. It graded the call and threw the evidence away.
The reframe: a confidence score is an opinion about the call. Evidence is the call. In a regulated business, the difference between the two is the difference between a number you can report and a fact you can defend.
There is a name worth giving the space between them. The Evidence Gap is the distance between a number that says something happened and proof that it did. A confidence score, however clever the model behind it, lives on the wrong side of that gap.
Why a score feels like the answer
Scores are irresistible, for good reasons. They are tidy. They fit in a dashboard, sort into a league table, roll up into a board slide. A number feels like measurement, like control, like you finally have a handle on twenty thousand conversations. An AI that reads every call and returns a compliance percentage per agent looks, at a glance, like exactly what you always wanted.
But a score is a compression. To hand you one clean number, the model throws away the thing the number was made from: the words. And in a regulated business, the words are the only part that was ever actually worth anything.
A score is testimony. Evidence is the exhibit.
Think about how any dispute is actually settled, at AFCA, at the Financial Ombudsman Service, in front of ASIC. Nobody accepts a score. A tribunal does not care that your system rated the call 92%. It wants the recording, the transcript, the exact words. A confidence score is testimony: it says trust me, this call was fine. Quote-linked evidence is the exhibit: here is precisely what was said, at this second, on this call. In any forum that matters, the exhibit wins and the testimony is just noise.
A confidence score says trust me. Evidence says here is what was said. Only one of those survives a dispute.
The question a score cannot survive
Every compliance number faces the same test eventually, and it is one word: show me. Show me the call behind the 92%. Show me the sentence that made this one a fail. Show me the moment the agent crossed the line. A score cannot answer any of them, because it discarded the evidence in order to become a score. So the instant a real question arrives, from a regulator, an ombudsman, or your own board, an unlinked score is worth precisely nothing, and you are back to pulling the recording by hand, which is where you started.
This is why Metricsense pins every finding to the exact moment in the transcript. When Life Insurance Direct ran a full window of 5,438 calls, each finding came back attached to the sentence that triggered it, not a rating to be trusted but a quote to be read. Here is what that looks like, an illustration rather than a real call:
A score would have compressed that exchange into a number and lost the sentence. The finding keeps the sentence, because the sentence is the only thing you can coach on, correct, or defend.
The confident score is the most dangerous kind
Here is the part that should unsettle anyone leaning on scores, and it gets sharper as AI agents start handling calls themselves. A model can be confidently wrong. A high number attached to a mistaken judgement is more dangerous than no number at all, because it invites you to stop looking. A human agent can state a false thing with total assurance; an AI agent can do the same and call it a fact. In both cases the confidence is not the same as being right, and a confidence figure you cannot open is not proof that anyone was. Score the agent, human or AI, by all means. Just insist on the evidence underneath the score.
The honest limits
Two caveats, because a piece arguing against unearned trust should not ask for any.
First, scores are not useless. A number is a fine way to triage, to sort twenty thousand calls so a human starts with the twenty that matter most. The argument is not no scores ever. It is no score without its evidence attached, because the score sorts and the evidence decides.
Second, evidence-linking is only as strong as the transcript beneath it. Garbled audio makes a weaker exhibit. The point is not that a quote is infallible; it is that a quote can be checked by a person, and a bare score cannot. And the standing limit holds: every flag is a candidate for human review, not a verdict. The evidence makes the reviewer fast; it does not replace them, and the tool reports nothing to a regulator on its own.
The objection you are already forming
Is a 95%-confident flag not good enough? Confident in what, though. The model's confidence is a statement about the model, not about the call. Ninety-five percent sure and completely unable to show you the sentence is still zero percent defensible. And the question every compliance buyer asks next: where does the call data go? Your own cloud, your own region, with PII stripped before any transcript reaches a model. The evidence never leaves your environment.
Keep the score. Stop mistaking it for an answer.
None of this means throw the scores away. Keep them; they are useful for sorting. Just stop treating them as conclusions. A number tells you where to look. Only the words tell you what happened, and only the words will stand up when someone asks you to prove it.
Recording is the infrastructure. Analysis is the intelligence. A score is the headline; the evidence is the story, and in a regulated business the story is the only part that counts. The one compliance score worth trusting is the one you can open.

AI Product Engineer, Metricsense
Krunal Ambaliya is an AI product engineer at Metricsense. He builds the data pipelines, cloud infrastructure, and Agentic AI systems that power the platform, turning thousands of unstructured reviews, tickets, and call transcripts into structured, actionable intelligence. His focus: making sure every insight links back to the exact customer quote so product teams can act on evidence, not assumptions.
See findings you can open, not scores you have to trust
Metricsense reads a week of your own calls, in your own cloud, and returns findings pinned to the exact sentence, the quote a QA lead or a regulator can actually read, usually within a day.
Keep reading
Is it an adviser problem, or is the script teaching everyone the same mistake?
When the same mistake shows up on agent after agent, coaching each of them is the expensive way to fix nothing. The cause is usually upstream, in the script, not in the people.
Your QA sample has an adverse selection problem
The calls most likely to contain a compliance problem are the least likely to be in your sample. That is not bad luck. It is adverse selection, and every insurer already knows exactly how it works.
Your QA score only measures the calls you chose to hear
A 94% compliance rate sounds like a fact about your call floor. It is a fact about your sample. In a regulated contact centre, that gap is where the fines live.