Ask any competitive pricing vendor how accurate their data is and you'll get a confident number back. The trouble starts when a category manager pulls up a report, spots a price that's clearly wrong and realizes the accuracy on the pitch deck and the accuracy they're receiving aren't the same thing.
That gap is critical. A pricing decision built on a handful of bad matches doesn't just produce one wrong number. It produces a wrong call on margin, a MAP violation that gets missed or a repricing rule that fires on the wrong product entirely. And once a team stops trusting the data, they start double-checking everything by hand, which defeats the point of buying the data in the first place.
Two vendors can both claim 98% accuracy and mean different things by it.
Accuracy answers one question: of the matches we found, how many are genuinely the same product?
Completeness answers a different one: of all the matches that should exist, how many did we actually find?
Match rate answers a third: of the products in your catalog, how many have at least one competitive match at all?
A vendor can hit 98% accuracy but still miss half of the matches out there, or they can get 100% match rate, but most of those matches are wrong.
Sample size matters just as much. Taking a small handful of SKUs for a human to manually review isn't giving you a reliable number. Data quality should be measured at a scale appropriate to your own with a sample size large enough to produce a confident, significant result. r gets less reliable exactly as your business gets bigger and the stakes get higher.
The number itself is less useful than the answer to how they got it. A few questions worth asking before signing anything:
Do you report accuracy, completeness and match rate as separate numbers, or is one figure blending them together?
How do you measure accuracy and completeness and at what confidence levels?
Can you show accuracy and completeness broken down by category or retailer, not just a single number?
What's the underlying matching approach and how do you decide when a match is uncertain rather than automatically confirmed?
How often do you measure data quality and is the sample randomized each time?
What happens when a match is wrong? (This does happen, so worth knowing what the process is!)
Check out our Data Quality questionnaire to bring to all your conversations!