How to Improve Match Rates with Waterfall Enrichment
Match rate is the metric that decides whether your enrichment budget pays off. Here's the practical playbook we use to push waterfall match rates higher: clean your inputs, order providers by strength, and refresh identity in real time so aged CRM records still resolve.

On this page
Match rate is the percentage of records you successfully enrich. Upload 1,000 contacts, get 700 back with a valid email, and your match rate is 70%. It's the single metric that decides whether your enrichment budget turns into pipeline or gets wasted on blanks. Everything else, cost per credit, provider logos, fancy dashboards, is secondary.
I co-founded Pipecorn, and I've watched teams obsess over price per lookup while ignoring the number that actually matters. A cheap provider with a 45% match rate is more expensive than a pricier one at 80%, because you pay twice: once for the miss, and again for the deal you never opened. This is a practical guide to pushing your waterfall enrichment match rates higher, in the order that actually moves the needle.
How do you improve match rates with waterfall enrichment?
You improve match rates with waterfall enrichment by combining four levers: clean, normalized inputs so providers can resolve identity; more sources in the waterfall so a miss on one is caught by the next; ordering providers by data strength rather than price; and a real-time identity refresh so aged records still resolve against the person's current details.
Those four levers compound. Clean inputs raise the ceiling of every provider you call. More providers give you more shots on goal. Smart ordering means you spend credits where they hit. And real-time refresh is the one most people skip, which is exactly why their match rates crater on any list older than a few months. Let me break each one down into something you can act on this week.
The waterfall model itself is simple: you query provider A, and only if it returns nothing do you fall through to provider B, then C, and so on. You pay per successful result, not per attempt. That design is what makes coverage cumulative instead of a single point of failure. But a waterfall is only as good as the inputs you feed it and the order you stack it in.
Why clean inputs decide your ceiling
Clean inputs are the cheapest match-rate improvement available, and almost nobody does the work. Providers resolve a person to an email by matching identity signals: full name, company domain, and ideally a LinkedIn URL. Feed them a nickname, a stale company, or a marketing subdomain and even a strong provider returns a blank. Garbage in, blank out.
Start with the LinkedIn URL. It's the most stable identity anchor a person has, because it survives job changes, name variations, and email churn. If you can attach a LinkedIn URL to a record, do it, because it lets a provider resolve the current human rather than guessing from a name that three people at the company share.
Then normalize the rest. Strip Inc., Ltd., and GmbH off company names and map them to the root domain. Split full names into clean first and last fields. Resolve the corporate domain, not info@ aliases or the newsletter subdomain. Drop obvious junk rows before you spend a single credit on them.
None of this is glamorous. All of it raises the ceiling on every provider in your waterfall at once.
Order the waterfall by strength, not price
Order your waterfall by data strength for the segment you're enriching, not by which provider is cheapest per credit. The whole point of the model is that the first provider to return a valid result wins the record. If you lead with a weak-but-cheap source, you'll accept its low-confidence answer and never fall through to the provider that actually had the right one.
Here's the trap. Cheap-first ordering looks efficient on your invoice and quietly poisons your match quality. The waterfall stops at the first "hit," so a mediocre provider returning a plausible-but-wrong email blocks the strong provider behind it. You optimized the wrong number: cost per attempt instead of valid results delivered.
Order by fit instead. Different providers are strong on different data types and different geographies. One dominates North American B2B, another is far better in Europe or on mobile numbers. Segment your list, then build a waterfall per segment that leads with the source most likely to be right for that data type and region. If you're not sure who's strong where, run a bake-off on a sample and compare enrichment providers on your data before you commit the ordering. Workflow tools like Clay let you reorder the chain by hand once you know who wins where.

- Segment by geography and data type (email vs. mobile).
- Lead each segment with its strongest provider.
- Fall through to complementary sources, not redundant ones.
- Re-test the ordering quarterly as provider coverage shifts.
Why aged CRM data is where match rates collapse
Refresh identity in real time, because the biggest hidden cause of low match rates is stale data on both sides of the match. B2B contact data decays fast, by roughly 22.5% a year (HubSpot's database-decay research): people change jobs, companies rebrand, emails get deprecated. When your CRM list is six or twelve months old and the provider's database was also last crawled months ago, you're matching one snapshot of a moving target against another. Both are wrong, and the record misses.
This is the difference between a static waterfall and a live one. A static waterfall queries pre-built databases: fast, but only as fresh as the last crawl. A live approach re-fetches the person's current identity at the moment of the lookup, then resolves the email against who they are today, not who they were when the database was last refreshed. On aged lists, that gap is enormous.
Pipecorn improves match rates precisely here. We re-fetch the latest LinkedIn data in real time on every lookup, which is why coverage holds up on 6-to-12-month-old CRM data instead of falling off a cliff. In Pipecorn's benchmark of 567 B2B leads, we saw 82% valid-email coverage versus 71% for FullEnrich, another waterfall tool, an 11-point lift overall that widened to about 18 points on aged CRM data. The older your list, the more a real-time refresh matters.
22.5%share of a B2B contact database that decays every year (HubSpot's database-decay research)
Practical takeaway: if you enrich the same list repeatedly, don't cache and forget. Re-run enrichment against a live source when the data is more than a quarter old, and prioritize providers that resolve identity at query time over ones that only read a static index.
Which waterfall enrichment providers offer match rate guarantees?
Several waterfall enrichment tools advertise "match rate guarantees," but the honest answer is that a guarantee almost always means a pricing model, not a promised percentage. In practice it takes two forms: pay-per-valid-result, where you're only charged for records returned with a verified email, and credit-back, where invalid or bounced results are refunded to your balance.
Read the fine print, because "guarantee" is a marketing word doing a lot of work. Most credible providers, Pipecorn included, use the pay-for-valid model: you don't pay for a miss, so your effective cost tracks your real match rate. That's a genuine protection, and it's the closest thing to a guarantee that holds up. What you should be skeptical of is any vendor promising a fixed match-rate number in the abstract, because match rate depends on your input quality and your segment, not just their database.
So evaluate guarantees like this: does the provider only charge for verified results? Do they refund bounces? And, most importantly, does their coverage hold on a sample of your own aged data? A pay-per-valid model plus a strong benchmark on your list beats any headline percentage on a landing page.
Frequently asked questions
What is a good match rate for waterfall enrichment?
It depends on your data and segment, but on clean B2B lists many teams target 70% or higher for valid emails. On aged CRM data the realistic bar is lower unless you use a real-time refresh. Pipecorn's benchmark of 567 B2B leads showed 82% valid-email coverage versus 71% for FullEnrich, a rival waterfall, and the advantage widened to about 18 points on aged data, which is strong for that scenario.
Does adding more providers always increase match rate?
Not automatically. Adding complementary providers that are strong in different geographies or data types raises coverage. Stacking redundant sources that all crawl the same databases mostly adds cost without new hits. Order by strength and fill gaps deliberately rather than piling on overlapping vendors.
Why do match rates drop on old CRM lists?
Because B2B contact data decays: people switch jobs, companies rebrand, and emails get deprecated within months. When both your list and the provider's index are stale, you're matching two outdated snapshots. A live identity refresh at query time resolves the person as they are today and recovers records static databases miss.
Is a match rate guarantee the same as guaranteed accuracy?
No. Most "guarantees" are pricing models: pay-per-valid-result or credit-back on bounces. They protect your spend, not a fixed percentage. Accuracy still depends on your input quality and segment. Treat a guarantee as a cost safeguard and validate real coverage on a sample of your own data.
Conclusion
Match rate is the number that decides whether enrichment funds pipeline or burns budget. You improve it in a specific order: clean and normalize your inputs, attach LinkedIn URLs, order the waterfall by data strength for each segment, and refresh identity in real time so aged CRM records still resolve. That last lever is where most teams leave the biggest gains on the table.
None of this requires a bigger budget, just a sharper process. Fix your inputs this week, re-test your provider ordering on a real sample, and stop paying for misses. If you want to see how a real-time refresh holds up on your own aged list, run Pipecorn against a sample of your CRM and compare the valid-email coverage yourself. The number will tell you everything.






