What match rate actually measures
A match rate is a fraction. The numerator is the number of visitor sessions an identification tool could put a name, company, or hashed identity to. The denominator is whatever the vendor decided to count as the pool of visitors in the first place, and that second number is where most published match rates fall apart. Some vendors call this an identity resolution match rate, but it is the same fraction under a longer name, and the same denominator problem applies.
Two tools can process the exact same traffic and report a 4% match rate and a 55% match rate. Both numbers are true. They are just dividing by different things. One counts every session that hit the site, including bots, ad crawlers, and visitors on carrier grade NAT IPs that were never going to match anything in the first place. The other only counts sessions that already passed a pre-filter built to exclude that unmatchable traffic before the math even starts.
Why the same traffic produces wildly different numbers
Most of the gap between a lowball rate and an impressive one comes down to what gets filtered out before the arithmetic starts. Search engine bots, uptime monitors, and scraper traffic can make up a large share of raw hits on a typical site, and stripping them out before calculating a rate raises the number without identifying one additional real person.
Mobile and VPN traffic gets excluded for a related reason. Mobile carriers route large numbers of phones through a small number of shared IP addresses, known as carrier grade NAT, and VPN or corporate proxy traffic does something similar on a smaller scale. IP based methods cannot reliably resolve that traffic to one household or company. A tool that quietly drops it from the denominator looks far more accurate on the traffic that remains, even though nothing about its actual matching improved.
Then there is repeat visitor counting, the easiest lever to pull and the one talked about least. A tool that recognized a visitor on their first visit keeps recognizing them on visit five, six, and seven without doing any new work. Reporting the rate across all visits rather than unique visitors pads the number with those easy repeats.
The question that matters more than the number
Before trusting any published match rate, ask what it was measured against. A rate calculated against all sessions in a period is the only version worth comparing across vendors, and it is the version vendors are least likely to publish voluntarily.
| Denominator used | What it includes | Effect on the published number |
|---|---|---|
| All sessions in a period | Bots, VPN and mobile carrier traffic, one time visitors, returning visitors | Lowest, and most representative of what a buyer actually gets |
| "Matchable" sessions only | Sessions after excluding bot, VPN, and carrier NAT traffic | Higher, and defensible only if the exclusion is disclosed |
| Previously identified visitors | Return visits from people already recognized on a prior visit | Highest, and tells you almost nothing about new traffic |
| Page views rather than unique visitors | Every view, so a single identified person can count many times | Inflated by session depth, not by identification skill |
What actually drives a realistic match rate
Traffic mix matters more than any vendor's technology. B2B software sites with desktop, office network traffic identify at a meaningfully higher rate than consumer sites where most visits come from phones on carrier networks. The underlying IP signal is simply more stable there, and more likely to map to a single company. Run that exact same identification technology on a retail site pulling mostly Instagram and TikTok traffic through in app mobile browsers, and the rate drops, not because the tool got worse but because the traffic itself carries less signal. Comparing match rates across categories of website is close to comparing conversion rates across industries. The number without the context is not useful.
Questions worth asking before you trust one
- Is the rate calculated against all sessions, or a filtered subset?
- Is it unique visitors, or total page views?
- What traffic mix was it measured on, and does it resemble mine?
- Does the vendor define a "match" as a full name and email, a company, or a hashed identifier with a confidence score?
That last question matters as much as the arithmetic. A match that resolves to a company name is a different claim than a match that resolves to a named individual, and both get marketed using the same word.
How Raydar handles this
Raydar's visitor identification works from a hashed IP signal with an explicit high confidence and low confidence split, not device fingerprinting and not a data broker lookup, and it does not identify every visitor. We do not publish a headline match rate, for the reason covered above: any single number would be true for one slice of traffic and misleading for the rest. What Raydar shows instead is per click detail, including how the IP hash identification works, alongside standard link tracking such as UTM capture and referrer data that answers which post drove a click regardless of whether the individual visitor was identified. If you're comparing this against de-anonymization claims from other vendors, it's worth reading how that term differs from identification at the hash level, and how person level identification differs from company level identification, since the two get conflated in most sales decks.
The honest version of the pitch
If a vendor gives you a match rate with no denominator attached, treat it as a marketing number, not a technical spec. Ask for the rate on your own traffic, calculated against every session including bots and mobile carrier IPs, and ask what fraction of matches are company level versus individual level. A vendor confident in their technology will give you that breakdown without hesitation. One who only wants to close the deal will give you the headline number again.