There is a phrase I keep seeing from anti-ad-fraud vendors that sounds incredibly impressive:
"We block 99.9% of invalid traffic."
Fantastic.
My dishwasher removes 99.9% of something too. I am just not entirely sure what.
The problem with the claim isn't necessarily that it is false. The problem is that, without context, it is almost completely meaningless.
99.9% of what?
Known bots?
Known data-center IPs?
Previously identified proxies?
Headless browsers?
Traffic that violates a vendor's proprietary definition of invalid?
Traffic their own system labeled invalid and then successfully blocked?
Because that last one is particularly convenient. If I create the definition, administer the test, grade the exam and announce the results, I have a sneaking suspicion I am going to graduate summa cum laude.
The real problem in invalid-traffic detection is not catching the obvious stuff. A bot announcing itself from a suspicious data center with an impossible browser configuration is not exactly Professor Moriarty.
The hard part is distinguishing between sophisticated invalid traffic and unusual-but-legitimate human behavior.
And that is where "99.9% blocked" tells you virtually nothing.
Suppose an anti-fraud vendor reviews one million clicks and blocks 100,000 of them.
Were those 100,000 actually fraudulent?
Were 95,000 fraudulent?
Were 70,000?
How many fraudulent clicks did the system fail to detect?
More importantly, how many actual humans did it throw into the garbage disposal along the way?
That last question tends to receive considerably less attention.
In fraud detection, blocking more traffic is easy. I can build a fraud-detection system this afternoon that achieves a truly magnificent fraud rate:
Block 100% of traffic.
Fraud problem solved.
Revenue problem introduced.
What matters is not simply how much traffic gets blocked. What matters is the accuracy of the classification.
You need to understand false negatives — fraud that gets through — and false positives — legitimate users incorrectly condemned as bots, proxies, automation, "junk," or whatever other wonderfully vague label the system produces.
Without an independently validated ground truth, claiming "99.9% effectiveness" is marketing, not measurement.
A meaningful claim would look very different.
Tell me the dataset.
Tell me how truth was established.
Tell me the false-positive rate.
Tell me the false-negative rate.
Tell me how the system performs against previously unseen traffic patterns.
Tell me what happens when a completely legitimate user behaves in a way your model has never seen before.
That is especially important because the internet is messy. Humans behave strangely. Networks behave strangely. Browsers behave strangely. Corporate VPNs, privacy tools, mobile carriers, shared IP addresses and increasingly complicated infrastructure create signals that do not always fit neatly into a fraud vendor's binary worldview.
And yet somehow we have arrived at:
99.9%.
Three beautiful decimal places of reassurance.
The anti-fraud industry absolutely provides tremendous value. We need these systems.
But sophisticated buyers should stop accepting impressive percentages without asking what exactly was measured.
Because "we block 99.9% of invalid traffic" sounds scientific.
Until you ask the most annoying question in advertising:
How do you know?