I recently saw a click rejected because it was 73% invalid.
Not invalid. 73% invalid.
I have been 73% sure I locked the front door. I still went back to check.
A confidence score is a strange thing to hand a buyer of traffic. You paid for a click. Someone else decided it was probably not a click. And the evidence you are given is a percentage.
A probability is not evidence.
A score tells you how strongly the system believes its own conclusion. It does not tell you what the click did. It does not tell you which signal mattered, whether the browser contradicted itself, whether the IP had a history, or whether the model simply woke up with a prior.
Those are different conversations.
A score can be perfectly calibrated and still be useless for your purposes. Calibration tells you that among all clicks scored 73%, about 73% were invalid. It does not tell you which one you are holding.
If the vendor says "invalid, 73% confidence," what exactly are you supposed to take to the partner who sold you the traffic?
"Your click was probably bad."
That will go well.
The problem is not that the score is false. It is that, without the reasoning behind it, the score is uncheckable. And when the reasoning stays on the vendor's side, the buyer cannot check it. Nobody has to intend that for it to be the outcome.
Adding a confidence level does not make a black box transparent. It makes it a black box with a number on the side.
Worse, the score usually arrives with a dial. You, the buyer, set the threshold. Block everything above 70%. Or 80%. Or whatever feels responsible.
Once the threshold is your setting, every argument about a wrong call becomes an argument about your configuration.
That is a structural consequence, not a plot.
The vendor did not reject the click. You did, by choosing 70 instead of 60. The vendor just provided the number. If the number was wrong, well, models are probabilistic. You were warned.
But here is what the buyer of traffic actually needs to know.
A click either came from a browser lying about what it is, or it didn't. That is a fact, not a probability. The probability is the vendor's uncertainty about the fact. Those are not the same thing.
A score of 73% does not tell you whether the browser reported itself as Safari on iOS and then answered a question only Chromium can answer. It does not tell you whether the session held together or fell apart. It does not tell you whether thousands of supposedly unrelated users from the same subnet clicked at the same improbable interval.
It tells you the model is 73% sure that something, somewhere, looked off.
That is not something you can take to a partner. It is not something you can audit. It is not something you can be held to.
The demand should not be "give me a confidence level."
The demand should be:
What did you observe, and why does it mean what you say it means?
What signals were observed?
What specifically contradicted what?
Can I check the reasoning myself?
Can I take it to the counterparty and have the argument?
Because the buyer of traffic is not buying a feeling. You are buying a fact about a click. If the vendor cannot produce the fact, the score is not a substitute. It is a way of not saying what was found, even when the vendor is being entirely sincere.
Next time a click is rejected, ask one question.
What did the click do?
If the answer is a number, you have not been told anything.