A confidence score is not evidence
0.94 tells you how sure a model is. It tells you nothing about what your adviser said. Only one of those survives a regulator.
If a vendor puts 0.94 next to a compliance flag, ask them what the 0.94 is a property of.
The answer is almost always "the model's own certainty". That is a fact about software, and you are about to file it somewhere that is supposed to contain facts about conduct.
I am an AI product engineer at Metricsense and I build the scoring this post is complaining about. Everything below applies to us, and two of the sections are about defects in our own output that I found this week. If you only take one thing away, take the five questions near the end, and ask us them first.
Three different claims get collapsed into one number
When a system tells you a call is 0.94 likely to contain a problem, three separate assertions are being folded together, and they have wildly different evidentiary weight.
- A model produced a high output on this input. True, checkable, and almost entirely uninteresting to a compliance officer.
- Something in this call matches the pattern the check describes. Sometimes true, and worth a human's attention.
- Your adviser said something they should not have. A claim about a person's conduct, which no model output establishes on its own.
The first is what the number measures. The third is what it gets read as. In a board pack, on a Thursday, under time pressure, nobody is holding those apart, and the number is doing the work of a receipt it cannot produce.
The unsourced score: a number attached to a conduct claim with no way to get from the number back to the words that produced it. You can accept it or override it on instinct. You cannot examine it, and neither can the person it is about.
A score asks for trust. A quote with a timestamp lets the reviewer disagree.
What a score gives the reviewer
0.87
advice-boundary risk
- No wording to examine
- No call to open
- Accept it, or override it on instinct
Nothing here can be checked.
What an evidence-linked flag gives them
“Honestly, if it were me, I’d go with the first option, it’s clearly the best one for your situation.”
call ref · 06:12 · adviser turn 38
- The exact wording, in context
- The call and the moment
- A reviewer can clear it, and say why
Candidate, not a finding. A human decides.
Constructed example. The quote is invented, not a real call.
Our own labels came back in three spellings this week
Here is a concrete failure from our own output, found on 11 August 2026 while I was re-querying the corpus for these posts.
The customer sentiment check on one anonymised Australian life-insurance corpus returned "Mild Concern" on 821 of 8,461 scored calls. It also returned "Mild_Concern" on 37, and "Mild concern" on 8. Same category, three spellings, three separate rows in the distribution.
calls carried mild concern. Reading the top row of our own chart would have undercounted it by 45 calls, about 5%.
Nobody would have noticed. The chart looked right, the top row was the largest, and the two stragglers sat far enough down the list to read as separate small categories. It is a trivial bug with an untrivial property: it produces a number that is wrong in a way no confidence score would ever flag, because the model was perfectly confident about each of the three labels it invented.
Every denominator we published in July has already moved
The second one is worse, because it had already shipped.
The earlier versions of these posts cited a corpus of 7,715 calls and dimension populations of 7,020 and 7,016, queried on 27 July 2026. Re-querying on 11 August, the corpus is 9,218 and those populations are 8,479 and 8,460. New calls arrived and the analysis was re-run on 10 August, so every published denominator moved underneath copy that was already live.
Worse than the drift: the window itself was wrong. Those posts described the corpus as starting on 1 April 2026. Searching for the actual first row this week, the workspace holds nothing before 15 April. We had been publishing a two-week overstatement of our own observation window for a month, and no confidence score anywhere in the system had any view about it.
That is the whole argument of this post arriving as a bill. A number is not defensible because a model was sure. It is defensible because you can say when you ran it, on which population, and go and run it again.
Traceability is not validation, and I will not pretend it is
Our position is that every score should link to the transcript moment behind it. I believe that, and it is not enough.
Evidence-linking fixes the second half of the problem: once a flag arrives with the quote attached, a reviewer can read the words, disagree, and clear it with a reason. It does nothing about the first half. It tells you nothing about the flags that were never raised, and a check that quietly misses a whole class of problem produces a clean report and a linked quote for everything it did find.
Traceability tells you why the system said something. It cannot tell you what the system never said.
Measuring what a check misses needs a labelled sample and a person to build it. It is slow, boring work, and any vendor who claims high recall without having done it is telling you about their confidence, not their coverage.
What we have, and precisely what it is not
We have coverage. Every call in a window is read against the same checks, so a flag is not a statement about who got selected.
We have the transcript moment, always. A flag arrives with the exact wording and its position in the call. If we cannot produce that, we do not raise the flag, and we would rather show you nothing than show you a made-up answer.
We do not have a published precision and recall figure per check on your calls. We have not run that validation on a client corpus, and until we have, the honest description of a Metricsense flag is a candidate issue with evidence attached, adjudicated by your people. Not a finding. Never a breach: that word belongs to the regulator and to your own compliance process, and we do not report anything to anyone on our own.
Five questions for any AI vendor, including us
- When I click a flag, what do I see? If the answer is a score, a category and a summary, but not the customer's and the adviser's actual words with a timestamp, stop there.
- What are the precision and recall of this check on calls like mine, and how were they measured? "High accuracy" is not an answer on an imbalanced class: a check that labels everything as the majority class scores brilliantly and knows nothing. Ask for the base rate and the confusion matrix.
- Was the check frozen before it was validated? Validating on the same calls used to tune the wording inflates the reported number, and this question separates the vendors who have done the work from the ones who have done a demonstration.
- If the number moves in six months, how do I know whether my advisers changed or your model did? You are asking whether they pin and publish model and prompt versions against every figure. Most do not. Chen, Zaharia and Zou tracked one commercial model whose accuracy on a single task fell from 84% to 51% over three months, with nothing changing on the buyer's side.
- Who turns a flag into a finding, and can that ever happen automatically? If a system can put a mark against an adviser without a person confirming it, you have bought a liability rather than a control.
The flag and the proof should arrive together
That is the whole standard, and it is deliberately low enough that the category should have cleared it years ago.
A number on its own asks you to trust a vendor. A number with the words attached asks you to make a judgement, which is the thing you are actually accountable for and the thing a regulator will ask you about. One of those is a product. The other is an argument you will be having alone.
I have not resolved where the line sits on publishing our own validation before it is complete. Publishing a weak precision figure invites a competitor to quote it out of context; not publishing means asking you to take the same thing on trust that this post tells you not to. We are still arguing about it internally, and I would rather you knew that than read a page claiming we had it worked out.

Head of Product, Metricsense
Gaurav Soni leads product for Metricsense. He has spent 13 years in product at the intersection of data, analytics and AI, starting in engineering, building APIs, data pipelines and analytics products, and moving steadily closer to the decisions leaders actually make. The thread through all of it is the same: turning messy, unstructured conversation into evidence a team can act on. At Metricsense that means reading 100% of a company's calls, reviews and tickets, and linking every finding to the exact quote.
Run the receipt test on real calls
A pilot runs on your own conversations, in your own cloud and region, with PII stripped before any transcript reaches a model, and you keep the evidence trail. Sensible if you hold an Australian advice or distribution licence and somebody in your business has to sign off on model risk. If nobody internally owns AI governance yet, do that first. It is the more valuable piece of work and it is not one we can sell you.
Keep reading
Avesta Labs is an OpenAI Select Partner. Your calls stay put.
The company that builds Metricsense has joined the OpenAI Partner Network. Here is the short list of what that changes for the recordings you have trusted us with, and the longer list of what it does not.
An ask was recorded on 34 of 8,479 scored calls
Read end to end rather than sampled, the clearest revenue pattern in this corpus was not a bad pitch or a hard objection. It was a conversation that simply stopped.
When everyone makes the same mistake, it is the script
One adviser getting it wrong is a coaching conversation. The same gap across the whole floor is a wording problem, and coaching will not touch it.