Methodology · GTM data benchmark

How the Apollo vs Clay comparison was conducted

This document specifies the procedure. It contains no result figures except where a number validates an instrument or justifies a metric choice — the findings are on the results page. Read this first: three of the published metrics have known biases that change how they should be read.

Version
1.0
Run window
2026-09-10 to 2026-09-12
Re-verification
2026-12-12

Conventions

These apply to this document and to the results page, in identical wording.

  • Dates are ISO-8601 in UTC. Every figure carries the date it was established.
  • Currency is USD. No conversions were required.
  • “Verified” means we retrieved this artifact ourselves, via API, on the stated date. It is never applied to a figure a vendor reports about itself. Vendor-reported counts are marked as such or excluded.
  • Conflicts between sources are reported, never silently resolved. Where Apollo and Clay disagree about a person, both values appear. Where a vendor and LinkedIn disagree, both appear and the disagreement is the datum.
  • Volatility. People change jobs. Every accuracy figure is a statement about 2026-09-12 and decays from that date. Employer fields decay fastest.
  • Scope of the comparison. Both vendors were queried through their APIs on the dates stated. Neither vendor was contacted, and neither has reviewed these results.

Scope and intended use

Measured: for five fixed ICP definitions — (a) how many qualifying people each vendor returns and how much those sets overlap, (b) whether returned rows satisfy the stated filter, and (c) whether the fields a vendor asserts about a person are contradicted by that person's LinkedIn profile.

Not measured: email or phone deliverability, intent data, enrichment beyond the four fields listed below, pricing, API latency, or coverage outside these five ICPs. Email accuracy is the most commonly cited axis in this category and is absent here — it would require a separate deliverability study.

Intended reader: someone choosing between the two vendors for outbound prospecting in segments resembling the five below, who wants to know whether the difference between them is volume, accuracy, or both.

Not intended as a statement about either vendor's total database size. Nothing in this design supports a claim about how many qualifying people exist in the world.

Definitions

Every figure on the results page depends on these.

Set
One of five ICP definitions, expressed in English, independent of how any vendor was queried.
Encoding
The query text one vendor was given for one set. Two encodings of the same set are never identical, because the query languages differ.
Source record
One row as a vendor returned it, unedited.
Person
A human. Several source records may describe one person.
Reference profile
The LinkedIn record retrieved and used to check a vendor's claims. Not ground truth.
Compliance
Whether a returned row belongs in the set it was returned for. Dimensions: location, title.
Agreement
Whether a row's claim about a person survives comparison with the reference profile. Fields: name, title, company, location.
agree / disagree / not_comparable
The only three recorded outcomes. not_comparable means the comparison could not be made because one side carried no value. It is never a failure.
escalate
A fourth outcome existing only inside the rules, meaning a rule declined to decide and the model was asked. Routing, never a result, and never stored.
Valid row
A returned row that is compliant, has no contradicted field, and has both title and company comparable.
Share of union
A vendor's people as a fraction of all distinct people either vendor found. Deliberately not called recall.

Universe and inclusion

Five ICPs, chosen before any data was pulled, to vary along axes expected to stress vendor coverage differently.

The five sets and what each was chosen to stress
SetDefinitionStress axis
1 RevOps · SFRevenue operations, San Francisco, 201–1000 employeesUS tech, literal geo
2 Sales · MunichHead of sales / sales director, Munich, 201–1000non-English titles
3 Ops · OhioPlant manager / director of operations, Ohio manufacturing, 201–500industrial, non-tech
4 Marketing · SingaporeHead of marketing, Singapore, 201–1000APAC, local title conventions
5 Founders · AustinapproxFounder / co-founder, Austin SaaS, 11–20small companies, vague vertical

Subjects. Apollo.io and Clay. Both were accessed with paid API credentials held by the study author. No other vendor was evaluated.

Asymmetry disclosed. Set 5 is only approximately harmonised between the two vendors and is reported separately wherever it matters. Apollo's SaaS filter is an opaque vendor keyword tag with no Clay equivalent; three defensible Clay encodings of “SaaS” span a 180× range in result count, and the industry-enum encoding was chosen. Cost of this choice: set 5's coverage comparison confounds vendor coverage with vertical definition, and its company-agreement figure is the least trustworthy of the five.

Excluded. An earlier, non-harmonised Clay pull is retained but excluded from all reported figures. It was produced with a phrase-matching encoding not comparable to Apollo's. Cost: excluding it removes the “what a practitioner gets by typing this in naively” framing, which is a legitimate but different question.

Constants

Held fixed across both vendors and all five sets.

  • Filter intent per set, fixed before any query was issued.
  • Compliance judged against the intent, never against either vendor's query text.
  • Employee-count bands: 201–500 and 501–1000; set 3 is 201–500 and set 5 is 11–20.
  • Current employment only.
  • Reference source: LinkedIn, via one actor, one mode, one date range.
  • Agreement fields: name, title, company, location — identical for both vendors.
  • Judge model, prompt text and temperature identical across vendors and sets.
  • Identity resolution rules identical across vendors; no per-vendor tuning.
  • No vendor-specific normalisation. Every normalisation rule — legal-suffix stripping, diacritic folding, credential removal — is applied to both sides of every comparison.
  • No row was excluded after seeing its verdict.

Dataset

Retrieval · what was pulled, when, and what it cost
SourceEndpoint or actorDateCost
Apollomixed_people/api_search then people/bulk_match2026-09-101,047 enrichment credits
Claysearch/query-mode then UI export2026-09-110 — search is not credit-metered
LinkedInApify actor harvestapi/linkedin-profile-scraper, profile details without email2026-09-12$9.71

Sizes. 2,935 source records ingested — 1,047 Apollo and 1,888 Clay after de-duplication on vendor identifier. 2,428 distinct LinkedIn URLs, being the union across both vendors: each profile was purchased once regardless of how many vendors supplied it.

Sampling. None at the record level — every returned row is judged. Sampling occurs only in instrument validation, where the seed is fixed at 11.

Exclusions with counts. 26 Clay rows were dropped at ingest as duplicate vendor identifiers: the same LinkedIn URL returned twice within one set. No other exclusions.

Not a random sample of the vendors' databases. These are five deliberately chosen ICPs. Results generalise to similar segments, not to the vendors overall.

Metrics

Four axes, published separately. Each carries a known bias, and each bias is stated with the metric rather than collected in a footer.

Coverage — share of union and unique contribution

For each set, people are resolved across vendors and then counted as found by both, found only by Apollo, or found only by Clay. A vendor's share of union is its people divided by the union. A larger share means the vendor surfaced more of the people that either vendor found. It does not mean the vendor found more of the people who exist.

Known bias — this metric systematically understates overlap. Identity resolution recovers about 91% of true cross-vendor pairs, so some people found by both are counted as unique to one. Every overlap figure is a floor and every unique-contribution figure is a ceiling. It is reported anyway because the direction of the error is known and one-way, and because the alternative — joining on LinkedIn URL — assumes the correctness of a field this study exists to test.

No confidence interval is given. The population is enumerated, not sampled, so the uncertainty is systematic — matcher recall — rather than sampling error, and an interval would misrepresent its nature. Never reported as recall: neither vendor is a census and the true population is unknown.

Filter compliance

A returned row is compliant if it satisfies every stated dimension of its set. Conjunctive: the right city with the wrong title is a wrong row, not a partial match. Location is matched on a structured vocabulary after diacritic folding and punctuation stripping. Titles are matched on a token set after diacritic folding, stopword removal, light plural stemming, abbreviation expansion (VPvice president), and joining of co-founder/cofounder.

Known bias — compliance is generous by construction. A rule may conclude agreement when required tokens are present, but may never conclude a disagreement from string evidence alone, because any rule able to reject Vertriebsleiter for head of sales would reject every translation, abbreviation and synonym. Unconfirmable titles escalate to the model, which may fail them; the rules alone can only pass. Compliance figures therefore run high for both vendors and discriminate weakly. They are reported to show that neither vendor returns obviously off-target rows, not to separate the two.

Field agreement with the reference profile

For each returned row with a retrievable reference profile, four fields are compared. Each yields agree, disagree, or not comparable.

The denominator rule. not_comparable is excluded from every rate — a field absent on either side is not evidence of error. The count of excluded comparisons is published beside every rate, because a 95% rate over 40 comparisons is not the same claim as 95% over 400.

Known bias — this measures agreement with LinkedIn, not correctness. LinkedIn is self-reported, frequently stale, and in the title field is often marketing copy rather than a job title. Where a vendor and LinkedIn disagree, the defensible statement is that they disagree. The metric is directional evidence about vendor freshness, not a verdict on truth.

A second, asymmetric bias. A vendor whose LinkedIn URLs are staler has more profiles fail to retrieve, so its remaining agreement is computed on a healthier subset. Reference-retrieval success per vendor is therefore published as a first-class figure on the same screen as the agreement rates, not as a footnote.

Missing is not wrong. A field the vendor left blank is not comparable. A field the vendor populated with something contradicted is a disagreement. These are distinct states and are never collapsed, because collapsing them lets a vendor that returns less score better than one that returns more.

Valid-row rate

A row is valid if it is compliant, has zero contradicted fields, and has both title and company comparable. The third clause exists because under the denominator rule alone, a vendor returning entirely empty rows would have nothing to contradict and would score perfectly. Requiring title and company to be checkable sets a floor on usefulness; both are populated on essentially every row by both vendors, so the requirement costs a well-behaved vendor nothing.

Known bias. Name and location contribute when present but cannot sink a row alone. A vendor with sparse names is not penalised here; that shows up in field agreement instead.

No composite score. The four field rates, the compliance rate, the reference-retrieval rate and the valid-row rate are published as separate axes. No weighted index is computed, because weights would encode an assumption about a buyer we have not met.

Instrument validation

Two components of this study are themselves measuring instruments. Both were measured, and both were found imperfect.

Identity matcher

People are matched across vendors on a composite of normalised first name, surname and company. LinkedIn URL is deliberately not the key — vendor-held URLs go stale, so a URL match is not proof of sameness and a mismatch is not proof of difference. Candidate pairs are generated from the union of two blocking keys, then decided by rule, escalating to the model only where the rule declines.

Blocking keys measured against an independent probe · 507 pairs sharing an identical normalised LinkedIn URL, evidence independent of the name-and-company key being tested
KeyPrecisionRecall
first + company93.5%91.1%
first + surname94.8%93.9%
first + company token89.5%97.4%

The union blocker is used for candidate generation because blocking should optimise recall: a pair never proposed can never be recovered, while a bad pair is cheap to reject downstream.

What this validation does not establish. The probe is one-directional. Two rows sharing a URL are almost certainly the same person, but two rows with different URLs may still be the same person whose URL one vendor recorded staler. Pairs invisible to the probe are invisible to this measurement, so 91.1% is a lower bound on true recall.

Model judge

Rules settle the large majority of comparisons; the residual is sent to a fixed model at temperature 0 with a fixed prompt. The escalated set is by construction the hard cases, so the model decides few comparisons but disproportionately many contested ones. 110 comparisons were sampled with seed 11 — 25 per escalated dimension plus 5 rule-decided controls per dimension — and labelled by hand against the rubric before the model's verdicts were compared. All 20 rule-decided controls matched, indicating the rules themselves are not the weak link.

Two limitations of this validation, both material. First, the labels are not independent ground truth: they were produced by a language model validating another language model, and shared blind spots inflate agreement. This is screening, not proof, and no published claim rests on the judge alone without the underlying values being shown. Second, the acceptance floor was never set in advance. The governing decision record requires an agreed minimum before measuring; that number was not fixed, so the observed agreement is what was observed, not a threshold that was cleared.

Per-dimension agreement and the specific disagreements are published on the results page rather than summarised away.

Error and retry handling

Every reference URL resolves to exactly one of three states, and the distinction is preserved end to end.

Reference retrieval states and what each does to a result
StateMeaningEffect on results
okProfile retrievedFields compared
not_foundURL dead, private, or renamedAll fields not comparable; counted against that vendor's retrieval success
errorOur side failed — quota, timeout, transportRetried; never appears in final results

error is retried because it is our failure. not_found is never retried, because re-buying a dead link produces the same dead link.

A defect that materially affected results, caught and corrected. The scraping actor reports its own quota refusals as a dataset item with HTTP 200, in the same shape as a per-profile “profile not found”. The first implementation could not tell them apart. Reading a quota refusal as a per-profile failure wrote 2,428 fake dead links, which would have been published as a data-quality finding about both vendors when it was a fact about our billing plan. Correcting that too aggressively then read per-profile failures as batch-wide refusals and discarded 1,428 profiles already paid for. The discriminator is a field the per-profile error carries and the actor-level refusal does not; both states are now separated, 1,348 of the discarded profiles were recovered from retained datasets at no additional cost, and the remaining 80 were re-fetched and proved genuinely dead. The published dead-link figures are therefore the third measurement of that quantity. The first two were wrong in opposite directions.

Caching and re-run behaviour. Every verdict is keyed by a hash of its normalised inputs, dimension, subject, model and prompt version, with a database uniqueness constraint. Re-running the judge issues zero model calls. A prompt or rule revision changes the key, so superseded verdicts are retained alongside new ones rather than overwritten, and reported figures filter to the current revision.

Limitations

Each limitation is stated in the section it constrains; this is an index, not a substitute.

  • Overlap figures are floors and unique-contribution figures are ceilings.
  • Compliance discriminates weakly by construction.
  • Agreement measures agreement with LinkedIn, not correctness.
  • Judge validation is model-on-model and has no pre-registered floor.
  • Set 5 is only approximately harmonised.
  • Five ICPs, four countries, one English-dominant vendor pair. Set 2 is the only non-English market; a benchmark spanning more languages would likely widen the gap on title handling.
  • Single point in time. Employer accuracy decays fastest.

No right of reply was conducted. Neither Apollo nor Clay was shown these results before publication, given an opportunity to correct factual errors, or asked to comment. This is a real gap and is disclosed rather than described as a process. Scores were not, and will never be, altered in exchange for access, data, or commercial consideration.

Reproducibility and data availability

  • Rules are pure functions with no I/O; the model is an adapter behind a protocol. Layer separation is machine-enforced by import contracts.
  • The specification is a table of 70 hand-written cases, 24 of them taken from observed data rather than invented. That case table is the authoritative statement of what each rule does.
  • Two experiments are retained with the throwaway code that produced them: the identity matcher validation, and escalation volume.
  • Vendor query text per set, per vendor, is stored.
  • All verdicts, profiles and raw vendor rows are retained, including the raw JSON of every source row and every reference profile.
  • Seeds: 11 for the gold sample, 7 for escalation sampling.

Re-running ingest, resolution and reporting is free and deterministic. Re-running the reference scrape costs approximately $9.71 and would produce different results, because people change jobs.

The machine-readable aggregates behind every published figure are on the results page.

Version and revision log

Revision history · errors found after publication are recorded here with the original value preserved and visible
VersionDateChange
1.02026-09-12First publication.

Validity. Accuracy figures describe 2026-09-12. Re-verification is due 2026-12-12; past that date these figures are marked stale rather than served as current.

Corrections policy. Errors found after publication are recorded with the original value preserved. Figures are never silently amended.

Private evaluation

Run this on your segments.

This methodology is the product. If you want the same procedure run against the segments you sell into, bring your ICP definitions.

Schedule a demo