Benchmark · GTM data
Apollo and Clay disagree by segment, not overall
The aggregate gap between these two vendors is 1.9 points. By segment it runs to 13 — and it changes direction. Clay returned more valid rows on the US tech, German sales and US industrial sets; Apollo on the Singapore marketing and Austin founder sets. Both sit near 92% on employer fields, meaning about one row in twelve carries a title or company the person's own profile contradicts.
- Measured
- 2026-09-12
- Vendors
- 2
- ICP sets
- 5
- Rows returned
- 2,935
- Reference profiles
- 2,325
- Comparisons
- 17,610
MethodologyData and reproductionSearch API benchmark
Measured 2026-09-12 · re-verification due 2026-12-12
Valid rows by set
The most important figure on the page. The aggregate reverses by segment, so a buyer should read one row.
- Apollo
- Clay
- approxSet 5 is only approximately harmonised between the vendors
| Set | Apollo | Clay | Difference | Rows returned |
|---|---|---|---|---|
| 1 RevOps · SF | 84.1% (138/164) | 92.6% (237/256) | Clay +8.4 | 420 |
| 2 Sales · Munich | 75.0% (174/232) | 88.2% (464/526) | Clay +13.2 | 758 |
| 3 Ops · Ohio | 79.9% (235/294) | 83.0% (386/465) | Clay +3.1 | 759 |
| 4 Marketing · Singapore | 87.7% (107/122) | 79.8% (256/321) | Apollo +8.0 | 443 |
| 5 Founders · Austinapprox | 84.7% (199/235) | 72.2% (231/320) | Apollo +12.5 | 555 |
Read the row matching your segment, not a total. Set 5 carries the largest margin and is also the least trustworthy of the five: Apollo's SaaS filter is an opaque vendor keyword tag with no Clay equivalent, so its comparison confounds vendor coverage with vertical definition.
Coverage
Distinct people each vendor found, after resolving the same person across both.
- Apollo only
- Found by both
- Clay only
Overlap figures are floors; unique-contribution figures are ceilings. The identity matcher recovers about 91% of true cross-vendor pairs, so some people found by both are counted here as unique to one. True overlap is higher than shown and true unique contribution lower. The error is one-way and its direction is known. These are not recall figures — neither vendor is a census, so the union is a lower bound on the real population, never an estimate of it.
| Set | Union | Both | Apollo only | Clay only | Apollo share | Clay share |
|---|---|---|---|---|---|---|
| 1 RevOps · SF | 308 | 112 | 52 | 144 | 53.2% | 83.1% |
| 2 Sales · Munich | 639 | 119 | 113 | 407 | 36.3% | 82.3% |
| 3 Ops · Ohio | 621 | 133 | 161 | 327 | 47.3% | 74.1% |
| 4 Marketing · Singapore | 367 | 72 | 50 | 245 | 33.2% | 86.4% |
| 5 Founders · Austinapprox | 471 | 84 | 151 | 236 | 49.9% | 67.9% |
| Total | 2,406 | 520 | 527 | 1,359 | 43.5% | 78.1% |
These are people, not rows. The valid-row and funnel tables count returned rows; this one counts distinct people after resolution, so the two differ wherever a vendor returned the same person twice — Clay set 3 is 465 rows to 460 people, set 4 is 321 to 317. The two denominators are never mixed in one figure.
Neither vendor is a superset of the other. Apollo contributed 527 people Clay never returned — slightly more than the 520 both found. Clay's advantage is breadth, not containment, and a team using only one of them is missing a population the other sees.
Field agreement
Whether what a vendor asserts about a person survives comparison with that person's own profile.
| Vendor | Rows returned | Reference profile retrieved | Rate |
|---|---|---|---|
| Apollo | 1,047 | 1,016 | 97.0% |
| Clay | 1,888 | 1,816 | 96.2% |
| Vendor | Field | Agree | Disagree | n | Rate | Not comparable |
|---|---|---|---|---|---|---|
| Apollo — reference profile retrieved for 97.0% of returned rows | ||||||
| Apollo | Name | 1,004 | 12 | 1,016 | 98.8% | n/c 31 |
| Apollo | Location | 982 | 25 | 1,007 | 97.5% | n/c 40 |
| Apollo | Title | 939 | 75 | 1,014 | 92.6% | n/c 33 |
| Apollo | Company | 936 | 78 | 1,014 | 92.3% | n/c 33 |
| Clay — reference profile retrieved for 96.2% of returned rows | ||||||
| Clay | Name | 1,814 | 0 | 1,814 | 100.0% | n/c 74 |
| Clay | Location | 1,701 | 24 | 1,725 | 98.6% | n/c 163 |
| Clay | Title | 1,665 | 148 | 1,813 | 91.8% | n/c 75 |
| Clay | Company | 1,656 | 153 | 1,809 | 91.5% | n/c 79 |
This measures agreement with LinkedIn, not correctness. LinkedIn is self-reported, frequently stale, and in the title field is often marketing copy rather than a job title. Where a vendor and the reference profile disagree, the defensible statement is that they disagree — not that the vendor is wrong.
By set
| Set | Vendor | Name | Location | Title | Company |
|---|---|---|---|---|---|
| 1 RevOps · SF | Apollo | 98.8% n=162 n/c 2 | 98.8% n=162 n/c 2 | 90.7% n=162 n/c 2 | 93.2% n=161 n/c 3 |
| 1 RevOps · SF | Clay | 100.0% n=249 n/c 7 | 99.6% n=249 n/c 7 | 97.2% n=249 n/c 7 | 96.8% n=249 n/c 7 |
| 2 Sales · Munich | Apollo | 99.0% n=208 n/c 24 | 99.5% n=208 n/c 24 | 89.8% n=206 n/c 26 | 92.8% n=207 n/c 25 |
| 2 Sales · Munich | Clay | 100.0% n=510 n/c 16 | 99.8% n=511 n/c 15 | 93.9% n=510 n/c 16 | 94.7% n=509 n/c 17 |
| 3 Ops · Ohio | Apollo | 98.3% n=293 n/c 1 | 93.7% n=284 n/c 10 | 93.2% n=293 n/c 1 | 92.2% n=293 n/c 1 |
| 3 Ops · Ohio | Clay | 100.0% n=450 n/c 15 | 95.4% n=438 n/c 27 | 91.6% n=450 n/c 15 | 94.0% n=450 n/c 15 |
| 4 Marketing · Singapore | Apollo | 99.2% n=122 n/c 0 | 99.2% n=122 n/c 0 | 93.4% n=122 n/c 0 | 95.1% n=122 n/c 0 |
| 4 Marketing · Singapore | Clay | 100.0% n=305 n/c 16 | 99.6% n=227 n/c 94 | 86.5% n=304 n/c 17 | 90.1% n=303 n/c 18 |
| 5 Founders · Austinapprox | Apollo | 99.1% n=231 n/c 4 | 98.7% n=231 n/c 4 | 95.2% n=231 n/c 4 | 90.0% n=231 n/c 4 |
| 5 Founders · Austinapprox | Clay | 100.0% n=300 n/c 20 | 99.7% n=300 n/c 20 | 89.7% n=300 n/c 20 | 79.5% n=298 n/c 22 |
Clay's perfect name score is partly a consequence of returning less. It carries 74 non-comparable names against Apollo's 31, and abbreviates roughly 15% of surnames (Jin C.), which the rubric scores as consistent rather than wrong. It also has 163 non-comparable locations against Apollo's 40, concentrated in Singapore where 94 of 321 rows carry no city at all. That is why n/c is printed beside every rate and never folded into it.
Funnel
Returned, then compliant, then reference retrieved, then valid — per vendor.
| Stage | Apollo | Clay |
|---|---|---|
| Returned | 1,047 (100.0%) | 1,888 (100.0%) |
| Compliant with stated filter | 1,033 (98.7%) | 1,885 (99.8%) |
| Reference profile retrieved | 1,016 (97.0%) | 1,816 (96.2%) |
| Valid rows | 853 (81.5%) | 1,574 (83.4%) |
The difference between these vendors is volume, not per-row quality. Clay returned 1.8× more rows at a valid-row rate 1.9 points higher — inside the range that set-level variation moves. That aggregate reverses by set, so it is not a ranking; see valid rows by set.
The same funnel, per set
| Set | Vendor | Returned | Compliant | Reference retrieved | Valid |
|---|---|---|---|---|---|
| 1 RevOps · SF | Apollo | 164 | 161 | 162 | 138 |
| 1 RevOps · SF | Clay | 256 | 256 | 249 | 237 |
| 2 Sales · Munich | Apollo | 232 | 230 | 208 | 174 |
| 2 Sales · Munich | Clay | 526 | 523 | 512 | 464 |
| 3 Ops · Ohio | Apollo | 294 | 286 | 293 | 235 |
| 3 Ops · Ohio | Clay | 465 | 465 | 450 | 386 |
| 4 Marketing · Singapore | Apollo | 122 | 122 | 122 | 107 |
| 4 Marketing · Singapore | Clay | 321 | 321 | 305 | 256 |
| 5 Founders · Austinapprox | Apollo | 235 | 234 | 231 | 199 |
| 5 Founders · Austinapprox | Clay | 320 | 320 | 300 | 231 |
Compliance runs high for both vendors and discriminates weakly by construction. A rule may conclude that a title matches, but may never conclude a mismatch from string evidence alone — any rule able to reject Vertriebsleiter for head of sales would reject every translation and synonym. Unconfirmable titles escalate to the model, which may fail them; the rules alone can only pass. These figures show that neither vendor returns obviously off-target rows, not that one filters better.
Scope
Five ICP definitions, fixed before any data was pulled, chosen to stress vendor coverage along different axes. Full specification in the methodology.
| Set | Definition | Stress axis |
|---|---|---|
| 1 RevOps · SF | Revenue operations, San Francisco, 201–1000 employees | US tech, literal geo |
| 2 Sales · Munich | Head of sales / sales director, Munich, 201–1000 | non-English titles |
| 3 Ops · Ohio | Plant manager / director of operations, Ohio manufacturing, 201–500 | industrial, non-tech |
| 4 Marketing · Singapore | Head of marketing, Singapore, 201–1000 | APAC, local title conventions |
| 5 Founders · Austinapprox | Founder / co-founder, Austin SaaS, 11–20 | small companies, vague vertical |
Set 5 is only approximately harmonised, and is marked wherever it appears. Apollo's SaaS filter is an opaque vendor keyword tag with no Clay equivalent, and three defensible Clay encodings of “SaaS” span a 180× range in result count. The industry-enum encoding was chosen. The cost of that choice is that set 5's coverage comparison confounds vendor coverage with vertical definition, and its company-agreement figure is the least trustworthy of the five.
What the shapes show
Volume and per-row quality are close to independent. Clay returned 1.8× more rows at a valid-row rate 1.9 points higher — a gap smaller than the set-level variation inside either vendor.
Both vendors are accurate on identity and drift on employment. Name and location clear 97% for both. Title and company sit near 92% for both. The failure mode these products share is staleness about where someone works, not confusion about who they are.
The weakest cell in the study is Clay's set 5 company rate at 79.5% over 61 disagreements. Set 5 targets 11–20-person companies, where employer records change fastest and are least maintained — and it is also the set whose definitions diverge most between vendors, so the two effects cannot be separated here.
An earlier Clay pull was excluded from every figure. It used a phrase-matching encoding not comparable to Apollo's. It is retained but unreported. The cost of excluding it is that this study cannot answer “what does a practitioner get by typing this in naively” — a legitimate question, and a different one.
How verdicts were reached
Rules decide first and escalate only where they decline. The model therefore settles few comparisons and disproportionately many contested ones.
| Dimension | Rule-decided | Model-decided | Total | Escalation rate |
|---|---|---|---|---|
| Compliance · location | 2,935 | 0 | 2,935 | never escalated |
| Agreement · name | 2,920 | 15 | 2,935 | 0.5% |
| Compliance · title | 2,894 | 41 | 2,935 | 1.4% |
| Agreement · location | 2,883 | 52 | 2,935 | 1.8% |
| Agreement · company | 2,685 | 250 | 2,935 | 8.5% |
| Agreement · title | 2,651 | 284 | 2,935 | 9.7% |
| All comparisons | 16,968 (96.4%) | 642 (3.6%) | 17,610 |
Location compliance never reached the model. It is a closed structured vocabulary and is fully rule-decidable. Company and title agreement escalate hardest, which is where judge quality matters most — read instrument validation before trusting those two rates.
| Outcome | Meaning | Count | Share |
|---|---|---|---|
| Retrieved | Profile returned and fields compared | 2,325 | 95.8% |
| Dead or private | URL dead, private or renamed — a finding about the vendor that supplied it. All fields not comparable | 103 | 4.2% |
| Run failure | Our side failed — quota, timeout, transport. Retried, so it never reaches published results | 0 | 0% |
These counts are the third measurement of this quantity, not the first. The scraping actor reports its own quota refusals in the same shape as a per-profile failure. Reading a refusal as a per-profile failure wrote 2,428 fake dead links, which would have been published as a data-quality finding about both vendors when it was a fact about our billing plan. Correcting that too aggressively then discarded 1,428 profiles already paid for. The discriminator is a field the per-profile error carries and the actor-level refusal does not; 1,348 of the discarded profiles were recovered at no extra cost and the remaining 80 proved genuinely dead. The first two measurements were wrong in opposite directions.
Instrument validation
Two parts of this study are themselves measuring instruments. Both were measured, and both were found imperfect.
| Instrument | How measured | Result |
|---|---|---|
| Identity matcher | Precision / recall against an independent URL probe, 507 pairs | 93.5% / 91.1% |
| Model judge | Agreement with hand labels, seed 11 | 104/110 (94.5%) |
| Dimension | Agreement |
|---|---|
| Company | 28/30 |
| Location | 29/30 |
| Name | 18/20 |
| Title | 29/30 |
Two caveats, both material. The hand labels were produced by a language model validating another language model, so shared blind spots inflate agreement — this is screening, not independent ground truth, and no figure on this page rests on the judge alone without the underlying values being retrievable. And no acceptance floor was agreed before measuring, so 94.5% is what was observed, not a threshold that was cleared. A pre-registered floor would be a stronger claim and is not available here.
Disagreements are published rather than summarised. The model treated OPW Retail Fueling and OPW Fueling Components as different employers where the labeller read one employer with two division names, and the same for OSG USA, INC. and OSG Tap & Die, Inc. It matched Bolivar (Ohio) to Ciudad Bolívar (Venezuela) on string similarity. It accepted a headline fragment sitting in a vendor's name field as not comparable, where the labeller scored it a disagreement on the grounds that the value is present and false.
What this does not show
Named gaps, not hedges.
- Email or phone accuracy.
- Not measured — and the axis most commonly cited in this category. It would need a separate deliverability study.
- Pricing, latency, enrichment depth, intent data.
- Out of scope. Nothing here speaks to what either vendor costs or how fast it answers.
- Coverage outside these five ICPs.
- Five deliberately chosen segments in four countries. Results generalise to similar segments, not to either vendor overall.
- Either vendor's total database size.
- Nothing in this design supports a claim about how many qualifying people exist.
- Which vendor is correct where LinkedIn disagrees.
- LinkedIn is self-reported and often stale. Where a vendor and the reference profile differ, the defensible statement is that they differ.
- Languages beyond English and German.
- Set 2 is the only non-English market. A wider benchmark would likely widen the gap on title handling.
No right of reply was conducted. Neither Apollo nor Clay was shown these results before publication, given an opportunity to correct factual errors, or asked to comment. That is a real gap, and it is disclosed here rather than described as a process. Scores were not, and will never be, altered in exchange for access, data, or commercial consideration.
Data and reproduction
Published under CC BY 4.0. Attribution required; commercial use, redistribution and tabular reproduction permitted. Every figure on this page is computed from the published file at build time rather than transcribed, so the page and the file you can download cannot disagree.
| File | Contents | Status |
|---|---|---|
results_raw.json | Every figure on this page: 10 funnel rows with per-field agree/disagree/not_comparable, 5 coverage rows, the verdict breakdown and the gold-sample distribution | published |
bench.db | All 17,610 verdicts, 2,428 reference profiles and every raw vendor row, each traceable to its inputs and its decider | not yet published |
source-queries.json | The literal query text given to each vendor for each set | not yet published |
Recomputing these figures
Derived values must be computed the way the page computes them, or the numbers will not match:
- Agreement rate =
agree / (agree + disagree). Never includenot_comparablein the denominator. - Valid rate =
valid / returned. - Share of union =
by_source[vendor] / union_size. This is not recall. - Compliance rate =
compliant / returned.
Re-running ingest, resolution and reporting is free and deterministic. Re-running the reference scrape costs about $9.71 and would produce different results, because people change jobs.
Corrections. Errors found after publication are recorded with the original value preserved and visible. Figures are never silently amended.
- Version
- 1.0
- Run window
- 2026-09-10 to 2026-09-12
- Measured
- 2026-09-12
- Re-verification due
- 2026-12-12
- Published
- 2026-09-12
- Licence
- CC BY 4.0
- Attribution
- unresolved
Private evaluation
Run this on your segments.
Bring your own ICP definitions. We'll run the same comparison on the segments you actually sell into, and show you the evidence behind every row.
Schedule a demo