The direct answer
Do not choose a European B2B data provider from its database size, a vendor-selected sample, or one headline accuracy rate. Freeze one definition of your target market, test every provider in the same time window, and run two separate evaluations: an account-discovery test and a contact-data test.
In the discovery test, measure whether returned companies actually match your written criteria. In the contact test, classify every requested record as confirmed correct, confirmed incorrect, inconclusive, or not returned. Report coverage, confirmed accuracy, uncertainty, usable yield, cost, and review time separately. Never let a high accuracy percentage hide a low return rate or a large unresolved bucket.
For Europe, add a legal and operational review by country. A provider's privacy documentation does not make your intended processing or outreach lawful. Confirm data sources, controller and processor roles, transparency and objection handling, transfers, suppression, and the rules for the channel and country you will use.
This Guide gives you a reproducible protocol. It does not rank providers because the result should depend on your market, fields, workflow, evidence, and risk threshold—not on who published the comparison.
Scope and evidence date
This method is for teams evaluating a database, enrichment service, sales-intelligence platform, or a combination of providers for European B2B work. It covers company discovery, professional contact data, and adjacent research fields. It does not certify a provider, establish a lawful basis, or replace advice from your legal or data-protection team.
The external sources and regulatory links were reviewed on 9 August 2026. Product terms, data sources, subprocessors, pricing, and national rules can change. Recheck them before a purchase or a new campaign.
The core principle is “fit for purpose.” The UK Government Data Quality Framework separates completeness from accuracy and asks users to prioritise quality dimensions according to the decision the data must support. That distinction matters here: a provider can return many records that are wrong, or very few records that are mostly right. Neither result is captured honestly by a single “data quality” score.
Run two tests, not one
Most provider tests start with a known person and ask vendors to append an email or phone number. That measures enrichment. It does not tell you whether a database can discover the right accounts from your actual market definition.
Keep the tests separate. If contact availability influences which accounts reviewers accept, you no longer know whether discovery found the right companies. If account relevance is ignored, a technically valid email can be counted as success even though the company should never have entered the workflow.
Europe is not one test segment. Define the countries, languages, company types, and roles that matter to your decision before testing. Report each critical subgroup independently; a strong aggregate can conceal a weak country, local-title, or market-segment result.
Step 1: write the field contract before opening a vendor
A fair test begins with a one-page field contract. Write it before seeing results so you cannot change the target to flatter a preferred platform.
Include:
- Decision: what will the team do with a record that passes?
- Market: countries and, where relevant, subregions or languages.
- Account criteria: company type, operating activity, size, ownership, location, exclusions, and edge cases.
- Role criteria: function, seniority, responsibility, acceptable local-language titles, and disqualifying roles.
- Requested fields: only fields necessary for the next decision.
- Freshness rule: what “current” means for each field and which evidence date is acceptable.
- Outcome rules: what counts as correct, incorrect, inconclusive, duplicate, and no result.
- Mandatory gates: legal documentation, export controls, integrations, sources, or review features without which the provider cannot be selected.
- Cost boundary: fees, credits, implementation time, and internal review time included in the comparison.
Avoid vague criteria such as “manufacturing companies” or “decision-makers.” A usable definition might be: “independent manufacturers with an operating site in Germany, 50–500 employees, selling their own physical products, excluding distributors and holding companies; find a person responsible for sales operations or commercial systems, preserving the original German title.”
That wording gives reviewers something they can adjudicate. It also exposes when a provider supports only a broad industry code and cannot express the actual requirement.
Step 2: choose a sample that can answer the question
Use your ICP, not the vendor's strongest segment
Build the sample independently. A provider-selected export can demonstrate the interface, but it cannot estimate performance on your market because the vendor controls which rows appear.
Represent the segments that could change the buying decision:
- country and language;
- company-size band;
- easy and difficult industries;
- common and local-language roles;
- known accounts, edge cases, and negative controls;
- records with and without a strong public footprint.
Do not spread a small sample across too many groups. One hundred total rows split across five countries leaves only twenty observations per country. The overall percentage may look precise while the country-level result remains extremely uncertain.
Treat 100 rows as directional and about 400 as a tighter estimate
For a proportion near 50%, the common normal approximation gives a 95% margin of error of roughly ±9.8 percentage points at n=100 and ±4.9 points at n=400, before accounting for non-random sampling or subgroup analysis. The calculation is 1.96 × √(0.5 × 0.5 ÷ n).
This is not a universal minimum. It explains the trade-off:
- Use roughly 100 eligible targets in one primary segment for an initial screen and workflow test.
- Use roughly 400 eligible targets when a difference of about five percentage points could change a material purchase decision.
- Size important country or persona strata separately. Four hundred total rows do not provide n=400 evidence for each subgroup.
- Report an interval with every important proportion. For small counts or results near zero or one, use an exact binomial or Wilson interval rather than a simplistic symmetric range.
NIST's guidance on confidence intervals for proportions explains why small samples and very low failure counts need more careful intervals. No confidence interval repairs a biased sample, inconsistent review, or a vendor-curated test set.
Keep a sealed answer key
Give each target a stable test ID. Store the expected account status, known company URL, known role information, evidence links, and evidence date in an answer key that is hidden from the provider outputs and, if possible, from the first reviewer.
The answer key must allow “unknown.” Public websites, professional profiles, registries, and your CRM can disagree. If the available evidence cannot decide a field, classify it as inconclusive rather than forcing a correct or incorrect result.
Do not use live unsolicited calls or emails merely to create ground truth. Verification activity is itself processing and may be marketing. Use records and methods your organisation is permitted to use, such as recently confirmed customer data, consented test contacts, authoritative registers, current company pages, and documented manual research.
Step 3: run the account-discovery test
Freeze the intent, not an identical click path
Different products expose different interfaces. One may accept natural language, another fixed filters, and another an API. Requiring identical clicks would favour one interaction model; allowing each vendor to reinterpret the ICP privately would make results incomparable.
Freeze these instead:
- the written account definition;
- the permitted source inputs;
- the maximum number of returned rows;
- the time budget and operator experience;
- any translation from the field contract into product-specific filters;
- the export time and test date.
Retain screenshots or a query log showing how the same intent was represented in every product. If a criterion cannot be expressed, record that as an interface or taxonomy limitation instead of silently dropping it.
Include positive, edge, and negative cases
A discovery sample should contain:
- Clear positives: accounts that plainly satisfy the written criteria.
- Edge cases: subsidiaries, groups, mixed business models, local legal forms, companies with sparse sites, and ambiguous size boundaries.
- Negative controls: distributors when manufacturers are requested, agencies with similar keywords, inactive entities, duplicate branches, or companies outside the geography.
Negative controls test precision. A search that returns every keyword match may appear to have broad coverage while creating expensive review work.
Measure discovery with honest names
Use these measures:
- List precision: accepted matching accounts ÷ reviewed returned accounts.
- Known-set recovery: known eligible accounts recovered ÷ eligible accounts in the sealed reference set.
- Duplicate rate: duplicate entities ÷ returned rows.
- Unresolved rate: accounts that cannot be adjudicated ÷ reviewed returned accounts.
- Review minutes per accepted account: human review time ÷ accepted matching accounts.
- Reason coverage: returned accounts with enough evidence to explain why they match ÷ returned accounts.
Call known-set recovery a proxy, not universal market recall. True recall requires a credible denominator for the entire eligible market. If you do not know that universe, “we found 80% of all companies” is not a result you can support.
Retain the rejection reasons. A provider that repeatedly returns distributors, parent companies, or the wrong local role has a systematic failure that one average percentage can hide.
Step 4: run the contact-data test only on accepted accounts
Create a second frozen file from accounts that passed discovery. Specify the role first, then request only the contact fields the workflow needs. This avoids paying to enrich accounts that should have been rejected.
For each requested record or field, assign exactly one outcome:
- Confirmed correct: independent, current evidence supports the person, company, role, and field at the required confidence.
- Confirmed incorrect: evidence shows the person, company, role, or field is wrong or no longer current.
- Inconclusive: the provider returned a value, but available evidence cannot responsibly confirm or reject it.
- No result: the provider did not return the requested value.
Do not count “syntactically valid,” “SMTP accepted,” or “a phone rang” as complete proof that a field belongs to the right current person. These can be useful signals, but the adjudication rule must match the intended decision. Also record catch-all domains, shared inboxes, headquarters numbers, inferred values, and fields without a verification date as separate attributes where they matter.
Run providers in the same short window. Shuffle outputs and remove provider branding before review where practical. Give reviewers the same evidence sources and written rules. Double-review a subset; if reviewers disagree often, improve the adjudication rule before trusting the vendor comparison.
Calculate a provider test without hiding uncertainty
Enter one provider's raw test counts. The scorecard keeps coverage, confirmed accuracy, uncertainty, usable yield, and cost on separate denominators.
The scorecard caps returned records at the target count and confirmed outcomes at returned records. Every returned record not confirmed correct or incorrect remains inconclusive.
Returned coverage is 72%. Confirmed accuracy is 85.9% across 64 decisive reviews; 8 returned records remain inconclusive. Usable yield across the full sample is 55%.
| Metric | Result | Definition |
|---|---|---|
| Returned coverage | 72% | Records returned ÷ targets evaluated |
| Confirmed accuracy | 85.9% | Confirmed correct ÷ (confirmed correct + confirmed incorrect) |
| Inconclusive share | 11.1% | Inconclusive ÷ records returned |
| Usable yield | 55% | Confirmed correct ÷ targets evaluated |
| Cost per confirmed correct record | €13.09 | Test cost ÷ confirmed correct |
The four outcomes are mutually exclusive and add up to the full target sample. Keeping inconclusive and no-result records visible prevents apparent accuracy from hiding weak yield.
Returned coverage uses every requested target as its denominator. Confirmed accuracy uses only records with a decisive review result. Unknown share uses returned records. Usable yield uses every target, so it cannot be inflated by omitting no-results. Cost per confirmed correct record divides the full test cost by confirmed correct outcomes.
This arithmetic does not establish legal permission, data freshness, statistical significance, or provider superiority. Report raw counts, sample composition, test date, confidence intervals, and the evidence used to adjudicate each field.
Step 5: keep every denominator visible
Suppose a provider is asked for 100 records, returns 72, and review confirms 55 correct, 9 incorrect, and 8 inconclusive. The provider costs €720 for the test.
The useful results are:
- Returned coverage: 72 ÷ 100 = 72%.
- Confirmed accuracy: 55 ÷ (55 + 9) = 85.9% across decisive reviews.
- Inconclusive share: 8 ÷ 72 = 11.1% of returned records.
- Usable yield: 55 ÷ 100 = 55% across the full request.
- Cost per confirmed correct record: €720 ÷ 55 = €13.09, before internal review time.
“85.9% accurate” alone would be misleading. It says nothing about the 28 no-results or 8 unresolved records. Conversely, treating every no-result as an error would confuse availability with correctness.
For each percentage, publish the raw numerator and denominator. Add a confidence interval where the sample is intended to estimate broader performance. Report results by decision-critical strata as well as overall; a strong UK result must not hide weak German coverage if Germany is the primary market.
Do not collapse gates into a weighted beauty score
A single weighted score is convenient but dangerous. A high interface rating should not compensate for a failed legal requirement, and broad email coverage should not compensate for missing the one field your workflow needs.
Use three layers:
- Pass/fail gates: intended-use documentation, required geography, mandatory fields, security, export or integration requirements, and legal review.
- Quality thresholds: minimum discovery precision, usable yield, maximum confirmed error rate, maximum inconclusive share, and evidence freshness.
- Optimisation among passing providers: cost per confirmed usable record, review time, workflow speed, and operator experience.
Set the thresholds before calculating the result. If no provider passes, change the workflow, add a second source, reduce the requested fields, or accept that the target is not supportable. Do not lower the definition after seeing a preferred vendor fail.
Step 6: evaluate the operating workflow, not only the export
A clean sample can still become a poor production system. During a paid pilot, observe:
- whether credits are charged for no-results, uncertain results, duplicates, or retries;
- whether verification dates, sources, confidence, and no-result states survive export;
- whether local-language titles are preserved and mapped without losing the original;
- whether accepted and rejected records retain their reasons;
- whether scheduled refreshes overwrite reviewed values or create a new auditable version;
- whether suppression and correction decisions propagate to exports and integrations;
- whether permissions, approvals, API limits, and CRM sync fit the real team;
- how much manual cleanup occurs after the provider declares the job complete.
Calculate total cost per confirmed usable record. If account readiness rather than a contact field is the outcome you buy, use the sales-ready account cost model instead of silently changing the denominator:
(provider fees + implementation cost + reviewer/rework minutes × loaded hourly cost ÷ 60) ÷ confirmed usable records
Keep provider fees, implementation, reviewer labour, and rework visible as separate rows. The purpose is not to invent a perfect financial model; it is to stop a cheap lookup price from hiding expensive review and cleanup.
Repeat a smaller holdout test after the initial pilot. A vendor-assisted sample can behave differently from routine self-service use. Record the product tier, settings, API version, credits, date, and whether vendor staff configured the query.
Step 7: conduct European privacy and outreach due diligence
This section is a procurement checklist, not legal advice.
Separate provider compliance from your intended use
The European Commission says personal data must be processed lawfully, fairly, transparently, for specified purposes, with data minimisation and reasonable steps to keep it accurate. See its overview of the GDPR principles.
A provider may have a privacy programme and still be unsuitable for your purpose. Your organisation must assess why it processes the data, which fields are necessary, which legal basis it relies on, how people are informed, how objections and corrections are honoured, and which additional rules govern the communication channel.
Ask every provider:
- Which sources and collection methods apply to each field and country?
- What are the respective controller, joint-controller, and processor roles?
- Which lawful bases does the provider rely on, and what must the customer assess independently?
- How and when are people informed when data was not obtained from them?
- How can a person access, correct, object to, or erase data?
- How are opt-outs and suppression lists applied to new imports, refreshes, exports, and subprocessors?
- Which recipients and subprocessors receive the data?
- Where is data processed, and which transfer mechanism and supplementary safeguards apply outside the EEA?
- What retention periods, deletion behaviour, security measures, and incident terms apply?
- Can the provider document field provenance, last verification, and correction history?
Do not assume every vendor is merely your processor. The European Commission's explanation of how the GDPR applies to controller and processor roles notes that roles depend on who determines the purposes and means of processing. A provider may act in different roles for its own database and for customer-supplied enrichment.
Test transparency and objections in practice
The Commission's guidance on third-party data used for marketing says the acquiring organisation must ensure the source could lawfully provide the data, keep the list current, respect objections, provide required information, and comply with ePrivacy rules. The Commission also states that people have an unconditional right to object to processing for direct marketing and must be informed of that right by the first communication.
Ask for the actual notice, objection route, suppression workflow, and correction SLA. Test them with an authorised internal or controlled record. A policy link without an operational path is not enough.
If legitimate interests are being considered, GDPR Article 6(1)(f) requires processing to be necessary for legitimate interests pursued by the controller or a third party, unless those interests are overridden by the person's interests or fundamental rights and freedoms. The EDPB's public-consultation draft of Guidelines 1/2024 sets out a useful three-part assessment: identify an interest, show necessity, and balance it against the person's interests, rights, and freedoms. “B2B” is not a legal basis by itself.
Check country and channel rules separately
GDPR is not the only rule. National laws implementing ePrivacy and regulating unfair commercial practices can impose different conditions on email and phone outreach.
For example, Section 7 of Germany's Act against Unfair Competition distinguishes telephone advertising to other market participants—which requires at least presumed consent—from electronic mail, which generally requires prior express consent subject to a narrow existing-customer exception. Buying a work email or phone number does not answer whether a particular communication is permitted.
Create a country-by-channel matrix with legal review before activation. Include the target, purpose, relationship, lawful basis, source, required notice, objection and suppression process, and the specific marketing rule. Re-review it when the country, channel, audience, or use case changes.
Step 8: make the decision reproducible
Your final report should let a sceptical reader recreate the comparison without access to the vendor's sales presentation.
Publish internally:
- field contract and exclusions;
- providers, tiers, configuration, and assistance received;
- test dates and time window;
- sample source, size, strata, and limitations;
- answer-key method and evidence dates;
- exact outcome definitions;
- raw counts and denominators;
- confidence intervals and subgroup results;
- disagreement and inconclusive rates;
- provider fees, credits, implementation effort, and review hours;
- legal and security gates with reviewer and date;
- known limitations and the next retest trigger.
Choose the provider that passes the mandatory gates and performs best on the fields and segments that drive your actual decision. You may rationally choose different sources for different countries or fields. If you test a waterfall, report incremental lift and incremental cost at each step rather than attributing the combined result to every provider.
Schedule a retest when the provider changes sources, verification policy, pricing, product tier, or material subprocessors; when your ICP moves into a new country or segment; or when observed production quality crosses a threshold. Data quality is a monitored process, not a procurement event.
A blank report template
Use the summary below for the decision record, download the blank row-level CSV template for the underlying observations, and download the companion provider-cost ledger for the pilot total. The stable test ID lets you join discovery, contact, evidence, and review time without using a person's details as the key. Record the provider tier and configuration, test window and run date, source class, plus second-reviewer and reconciliation fields; otherwise a future comparison can look like a provider change when it was a configuration or review change.
Mark confirmed_usable_status as yes in the row-level template only when the record passed the written field contract and human review; use no for all other completed outcomes and leave it blank only while review is incomplete. In the companion ledger, use the controlled categories provider_fee, implementation, reviewer_labour, rework, and, only where relevant, record_specific_direct_cost. For another cost entry, duplicate a non-summary cost row so its entry_total_cost formula remains, then insert it above pilot-summary. Put contract or credit fees and implementation costs into the companion ledger once at whole_pilot_once; never repeat them on every observation. Put reviewer and rework minutes on cost rows linked to their test_id, including rejected and no-result rows so their incurred work remains in the numerator. Populate confirmed_usable_record_id only when it links to a row marked confirmed_usable_status = yes; that identifier explains the denominator link, not whether a cost is included. When record_specific_direct_cost_eur is populated on a row-template record, create exactly one matching companion-ledger row with the same linked_test_id, cost_category = record_specific_direct_cost, and the amount in direct_cost; treat the template value as an audit reference and do not sum it separately. Use one currency in the companion ledger, or document the conversion basis before summing. Finally, copy the count of row-template records marked confirmed_usable_status = yes for the same pilot into the ledger's summary row. Its formula then divides total pilot cost by that count, making total cost and cost per confirmed usable record reproducible without duplicating provider fees.
Copy this summary for each provider:
Attach a row-level file with stable test IDs, outcomes, evidence dates, and rejection reasons. Remove or restrict personal data that is not needed for the evaluation.
The most convincing Leadbase claim is your own test
A vendor can always choose the accounts that make its data look good. Leadbase becomes interesting when you do the opposite: freeze your own ICP, difficult rows, acceptance rule, contact targets, and cost ledger before the product sees the test.
Do not trust Leadbase because this Guide says it is better. Test whether it can build the exact list your current provider's schema cannot express.
Leadbase is built for teams that want to turn a market hypothesis into a reviewable, repeatable operating workflow—not simply purchase a file of contacts. In this protocol, that difference appears at four points:
- Describe the market you actually mean. Use Leadbase account discovery to find companies by meaning rather than forcing the ICP into one industry taxonomy.
- Turn unusual fit criteria into fields. With Leadbase Enrich, research the service, certification, operating model, hiring pattern, or other public condition that decides whether each account belongs.
- Qualify before contact lookup. Apply your own pass, review, and fail logic before adding decision-makers, so the benchmark measures usable targets rather than raw rows.
- Complete and repeat the workflow. Add the relevant people and contact fields, then rerun time-sensitive custom research when the market changes. The shared Sheet, sources, confidence, no-results, and history make each stage inspectable.
The commercial Aha is relevance before reveal. First prove that an account belongs in the market; then spend contact effort only on the accepted rows; finally evaluate the provider on confirmed usable outcomes and total review cost. A larger raw export cannot hide a weak acceptance rate when every denominator remains in the protocol.
This is a product-fit statement, not a shortcut around the test. Leadbase still has to clear your legal, security, coverage, accuracy, uncertainty, cost, and operating gates in the countries and segments that matter to you. The useful next step is to make Leadbase earn the decision with a controlled sample or freeze the evaluation with our team.
What this Guide does not claim
- It does not claim that 100 rows are enough for every purchase decision.
- It does not turn a provider's “verified” label into independent proof.
- It does not treat missing data as inaccurate or uncertain data as correct.
- It does not infer European or country-level performance from a global average.
- It does not say that a privacy policy makes a campaign lawful.
- It does not rank Leadbase or another provider without a dated, reproducible result.
The useful outcome is not a universal winner. It is a documented decision that another person can inspect, challenge, and rerun.




