Benchmark · GTM data

Apollo and Clay disagree by segment, not overall

The aggregate gap between these two vendors is 1.9 points. By segment it runs to 13 — and it changes direction. Clay returned more valid rows on the US tech, German sales and US industrial sets; Apollo on the Singapore marketing and Austin founder sets. Both sit near 92% on employer fields, meaning about one row in twelve carries a title or company the person's own profile contradicts.

Measured
2026-09-12
Vendors
2
ICP sets
5
Rows returned
2,935
Reference profiles
2,325
Comparisons
17,610

Measured 2026-09-12 · re-verification due 2026-12-12

No composite score and no overall ranking is published, here or anywhere in this study. The four field rates, the compliance rate, the reference-retrieval rate and the valid-row rate are separate axes. Weighting them into one number would encode an assumption about a buyer we have not met — and the per-set table below is the reason: any single total hides a reversal.

Valid rows by set

The most important figure on the page. The aggregate reverses by segment, so a buyer should read one row.

Valid-row rate by ICP set · 2,935 returned rows · measured 2026-09-12
  • Apollo
  • Clay
  • approxSet 5 is only approximately harmonised between the vendors
0%25%50%75%100%84.192.6RevOps · SFUS tech, literal geo75.088.2Sales · Munichnon-English titles79.983.0Ops · Ohioindustrial, non-tech87.779.8Marketing · SingaporeAPAC, local title conventions84.772.2Founders · Austinapproximately harmonised
Valid-row rate by set · valid ÷ returned · a row is valid when it is compliant, has no contradicted field, and has both title and company comparable
SetApolloClayDifferenceRows returned
1 RevOps · SF84.1% (138/164)92.6% (237/256)Clay +8.4420
2 Sales · Munich75.0% (174/232)88.2% (464/526)Clay +13.2758
3 Ops · Ohio79.9% (235/294)83.0% (386/465)Clay +3.1759
4 Marketing · Singapore87.7% (107/122)79.8% (256/321)Apollo +8.0443
5 Founders · Austinapprox84.7% (199/235)72.2% (231/320)Apollo +12.5555

Read the row matching your segment, not a total. Set 5 carries the largest margin and is also the least trustworthy of the five: Apollo's SaaS filter is an opaque vendor keyword tag with no Clay equivalent, so its comparison confounds vendor coverage with vertical definition.

Coverage

Distinct people each vendor found, after resolving the same person across both.

People per set after identity resolution · union 2,406 across 5 sets · bars scaled to union size
  • Apollo only
  • Found by both
  • Clay only
RevOps · SF308Sales · Munich639Ops · Ohio621Marketing · Singapore367Founders · Austin471

Overlap figures are floors; unique-contribution figures are ceilings. The identity matcher recovers about 91% of true cross-vendor pairs, so some people found by both are counted here as unique to one. True overlap is higher than shown and true unique contribution lower. The error is one-way and its direction is known. These are not recall figures — neither vendor is a census, so the union is a lower bound on the real population, never an estimate of it.

Distinct people per set after cross-vendor identity resolution · share of union is a vendor's people divided by all people either vendor found, and is deliberately not called recall
SetUnionBothApollo onlyClay onlyApollo shareClay share
1 RevOps · SF3081125214453.2%83.1%
2 Sales · Munich63911911340736.3%82.3%
3 Ops · Ohio62113316132747.3%74.1%
4 Marketing · Singapore367725024533.2%86.4%
5 Founders · Austinapprox4718415123649.9%67.9%
Total2,4065205271,35943.5%78.1%

These are people, not rows. The valid-row and funnel tables count returned rows; this one counts distinct people after resolution, so the two differ wherever a vendor returned the same person twice — Clay set 3 is 465 rows to 460 people, set 4 is 321 to 317. The two denominators are never mixed in one figure.

Neither vendor is a superset of the other. Apollo contributed 527 people Clay never returned — slightly more than the 520 both found. Clay's advantage is breadth, not containment, and a team using only one of them is missing a population the other sees.

Field agreement

Whether what a vendor asserts about a person survives comparison with that person's own profile.

Reference-profile retrieval per vendor · conditions every rate below it, so it is published beside them rather than beneath them
VendorRows returnedReference profile retrievedRate
Apollo1,0471,01697.0%
Clay1,8881,81696.2%
Field agreement with the LinkedIn reference profile · rate is agree ÷ (agree + disagree); n/c is the count excluded as not comparable and is never counted as a disagreement
VendorFieldAgreeDisagreenRateNot comparable
Apollo — reference profile retrieved for 97.0% of returned rows
ApolloName1,004121,01698.8%n/c 31
ApolloLocation982251,00797.5%n/c 40
ApolloTitle939751,01492.6%n/c 33
ApolloCompany936781,01492.3%n/c 33
Clay — reference profile retrieved for 96.2% of returned rows
ClayName1,81401,814100.0%n/c 74
ClayLocation1,701241,72598.6%n/c 163
ClayTitle1,6651481,81391.8%n/c 75
ClayCompany1,6561531,80991.5%n/c 79

This measures agreement with LinkedIn, not correctness. LinkedIn is self-reported, frequently stale, and in the title field is often marketing copy rather than a job title. Where a vendor and the reference profile disagree, the defensible statement is that they disagree — not that the vendor is wrong.

By set

Field agreement by set and vendor · each cell carries its rate, its comparable denominator n, and the n/c count excluded from that denominator
SetVendorNameLocationTitleCompany
1 RevOps · SFApollo98.8%
n=162 n/c 2
98.8%
n=162 n/c 2
90.7%
n=162 n/c 2
93.2%
n=161 n/c 3
1 RevOps · SFClay100.0%
n=249 n/c 7
99.6%
n=249 n/c 7
97.2%
n=249 n/c 7
96.8%
n=249 n/c 7
2 Sales · MunichApollo99.0%
n=208 n/c 24
99.5%
n=208 n/c 24
89.8%
n=206 n/c 26
92.8%
n=207 n/c 25
2 Sales · MunichClay100.0%
n=510 n/c 16
99.8%
n=511 n/c 15
93.9%
n=510 n/c 16
94.7%
n=509 n/c 17
3 Ops · OhioApollo98.3%
n=293 n/c 1
93.7%
n=284 n/c 10
93.2%
n=293 n/c 1
92.2%
n=293 n/c 1
3 Ops · OhioClay100.0%
n=450 n/c 15
95.4%
n=438 n/c 27
91.6%
n=450 n/c 15
94.0%
n=450 n/c 15
4 Marketing · SingaporeApollo99.2%
n=122 n/c 0
99.2%
n=122 n/c 0
93.4%
n=122 n/c 0
95.1%
n=122 n/c 0
4 Marketing · SingaporeClay100.0%
n=305 n/c 16
99.6%
n=227 n/c 94
86.5%
n=304 n/c 17
90.1%
n=303 n/c 18
5 Founders · AustinapproxApollo99.1%
n=231 n/c 4
98.7%
n=231 n/c 4
95.2%
n=231 n/c 4
90.0%
n=231 n/c 4
5 Founders · AustinapproxClay100.0%
n=300 n/c 20
99.7%
n=300 n/c 20
89.7%
n=300 n/c 20
79.5%
n=298 n/c 22

Clay's perfect name score is partly a consequence of returning less. It carries 74 non-comparable names against Apollo's 31, and abbreviates roughly 15% of surnames (Jin C.), which the rubric scores as consistent rather than wrong. It also has 163 non-comparable locations against Apollo's 40, concentrated in Singapore where 94 of 321 rows carry no city at all. That is why n/c is printed beside every rate and never folded into it.

Funnel

Returned, then compliant, then reference retrieved, then valid — per vendor.

Row funnel per vendor · 2,935 returned rows across five ICP sets · measured 2026-09-12 · each share is of that vendor's own returned rows
StageApolloClay
Returned1,047 (100.0%)1,888 (100.0%)
Compliant with stated filter1,033 (98.7%)1,885 (99.8%)
Reference profile retrieved1,016 (97.0%)1,816 (96.2%)
Valid rows853 (81.5%)1,574 (83.4%)

The difference between these vendors is volume, not per-row quality. Clay returned 1.8× more rows at a valid-row rate 1.9 points higher — inside the range that set-level variation moves. That aggregate reverses by set, so it is not a ranking; see valid rows by set.

The same funnel, per set
Row funnel by set and vendor · counts are rows as returned, before identity resolution
SetVendorReturnedCompliantReference retrievedValid
1 RevOps · SFApollo164161162138
1 RevOps · SFClay256256249237
2 Sales · MunichApollo232230208174
2 Sales · MunichClay526523512464
3 Ops · OhioApollo294286293235
3 Ops · OhioClay465465450386
4 Marketing · SingaporeApollo122122122107
4 Marketing · SingaporeClay321321305256
5 Founders · AustinapproxApollo235234231199
5 Founders · AustinapproxClay320320300231

Compliance runs high for both vendors and discriminates weakly by construction. A rule may conclude that a title matches, but may never conclude a mismatch from string evidence alone — any rule able to reject Vertriebsleiter for head of sales would reject every translation and synonym. Unconfirmable titles escalate to the model, which may fail them; the rules alone can only pass. These figures show that neither vendor returns obviously off-target rows, not that one filters better.

Detail, method and caveats

Scope

Five ICP definitions, fixed before any data was pulled, chosen to stress vendor coverage along different axes. Full specification in the methodology.

The five sets · defined in English, independent of how either vendor was queried
SetDefinitionStress axis
1 RevOps · SFRevenue operations, San Francisco, 201–1000 employeesUS tech, literal geo
2 Sales · MunichHead of sales / sales director, Munich, 201–1000non-English titles
3 Ops · OhioPlant manager / director of operations, Ohio manufacturing, 201–500industrial, non-tech
4 Marketing · SingaporeHead of marketing, Singapore, 201–1000APAC, local title conventions
5 Founders · AustinapproxFounder / co-founder, Austin SaaS, 11–20small companies, vague vertical

Set 5 is only approximately harmonised, and is marked wherever it appears. Apollo's SaaS filter is an opaque vendor keyword tag with no Clay equivalent, and three defensible Clay encodings of “SaaS” span a 180× range in result count. The industry-enum encoding was chosen. The cost of that choice is that set 5's coverage comparison confounds vendor coverage with vertical definition, and its company-agreement figure is the least trustworthy of the five.

What the shapes show

Volume and per-row quality are close to independent. Clay returned 1.8× more rows at a valid-row rate 1.9 points higher — a gap smaller than the set-level variation inside either vendor.

Both vendors are accurate on identity and drift on employment. Name and location clear 97% for both. Title and company sit near 92% for both. The failure mode these products share is staleness about where someone works, not confusion about who they are.

The weakest cell in the study is Clay's set 5 company rate at 79.5% over 61 disagreements. Set 5 targets 11–20-person companies, where employer records change fastest and are least maintained — and it is also the set whose definitions diverge most between vendors, so the two effects cannot be separated here.

An earlier Clay pull was excluded from every figure. It used a phrase-matching encoding not comparable to Apollo's. It is retained but unreported. The cost of excluding it is that this study cannot answer “what does a practitioner get by typing this in naively” — a legitimate question, and a different one.

How verdicts were reached

Rules decide first and escalate only where they decline. The model therefore settles few comparisons and disproportionately many contested ones.

How each of 17,610 compliance and agreement comparisons was decided · rules run first and escalate only where they decline · 578 distinct model calls, 0 of which failed to connect or parse
DimensionRule-decidedModel-decidedTotalEscalation rate
Compliance · location2,93502,935never escalated
Agreement · name2,920152,9350.5%
Compliance · title2,894412,9351.4%
Agreement · location2,883522,9351.8%
Agreement · company2,6852502,9358.5%
Agreement · title2,6512842,9359.7%
All comparisons16,968 (96.4%)642 (3.6%)17,610

Location compliance never reached the model. It is a closed structured vocabulary and is fully rule-decidable. Company and title agreement escalate hardest, which is where judge quality matters most — read instrument validation before trusting those two rates.

Reference retrieval outcomes · 2,428 distinct LinkedIn URLs, the union across both vendors, each purchased once regardless of how many vendors supplied it
OutcomeMeaningCountShare
RetrievedProfile returned and fields compared2,32595.8%
Dead or privateURL dead, private or renamed — a finding about the vendor that supplied it. All fields not comparable1034.2%
Run failureOur side failed — quota, timeout, transport. Retried, so it never reaches published results00%

These counts are the third measurement of this quantity, not the first. The scraping actor reports its own quota refusals in the same shape as a per-profile failure. Reading a refusal as a per-profile failure wrote 2,428 fake dead links, which would have been published as a data-quality finding about both vendors when it was a fact about our billing plan. Correcting that too aggressively then discarded 1,428 profiles already paid for. The discriminator is a field the per-profile error carries and the actor-level refusal does not; 1,348 of the discarded profiles were recovered at no extra cost and the remaining 80 proved genuinely dead. The first two measurements were wrong in opposite directions.

Instrument validation

Two parts of this study are themselves measuring instruments. Both were measured, and both were found imperfect.

Instrument validation · full procedure in the methodology, §7
InstrumentHow measuredResult
Identity matcherPrecision / recall against an independent URL probe, 507 pairs93.5% / 91.1%
Model judgeAgreement with hand labels, seed 11104/110 (94.5%)
Model judge agreement with hand labels, by dimension · all 20 rule-decided controls matched
DimensionAgreement
Company28/30
Location29/30
Name18/20
Title29/30

Two caveats, both material. The hand labels were produced by a language model validating another language model, so shared blind spots inflate agreement — this is screening, not independent ground truth, and no figure on this page rests on the judge alone without the underlying values being retrievable. And no acceptance floor was agreed before measuring, so 94.5% is what was observed, not a threshold that was cleared. A pre-registered floor would be a stronger claim and is not available here.

Disagreements are published rather than summarised. The model treated OPW Retail Fueling and OPW Fueling Components as different employers where the labeller read one employer with two division names, and the same for OSG USA, INC. and OSG Tap & Die, Inc. It matched Bolivar (Ohio) to Ciudad Bolívar (Venezuela) on string similarity. It accepted a headline fragment sitting in a vendor's name field as not comparable, where the labeller scored it a disagreement on the grounds that the value is present and false.

What this does not show

Named gaps, not hedges.

Email or phone accuracy.
Not measured — and the axis most commonly cited in this category. It would need a separate deliverability study.
Pricing, latency, enrichment depth, intent data.
Out of scope. Nothing here speaks to what either vendor costs or how fast it answers.
Coverage outside these five ICPs.
Five deliberately chosen segments in four countries. Results generalise to similar segments, not to either vendor overall.
Either vendor's total database size.
Nothing in this design supports a claim about how many qualifying people exist.
Which vendor is correct where LinkedIn disagrees.
LinkedIn is self-reported and often stale. Where a vendor and the reference profile differ, the defensible statement is that they differ.
Languages beyond English and German.
Set 2 is the only non-English market. A wider benchmark would likely widen the gap on title handling.

No right of reply was conducted. Neither Apollo nor Clay was shown these results before publication, given an opportunity to correct factual errors, or asked to comment. That is a real gap, and it is disclosed here rather than described as a process. Scores were not, and will never be, altered in exchange for access, data, or commercial consideration.

Data and reproduction

Published under CC BY 4.0. Attribution required; commercial use, redistribution and tabular reproduction permitted. Every figure on this page is computed from the published file at build time rather than transcribed, so the page and the file you can download cannot disagree.

FileContentsStatus
results_raw.jsonEvery figure on this page: 10 funnel rows with per-field agree/disagree/not_comparable, 5 coverage rows, the verdict breakdown and the gold-sample distributionpublished
bench.dbAll 17,610 verdicts, 2,428 reference profiles and every raw vendor row, each traceable to its inputs and its decidernot yet published
source-queries.jsonThe literal query text given to each vendor for each setnot yet published

Recomputing these figures

Derived values must be computed the way the page computes them, or the numbers will not match:

  • Agreement rate = agree / (agree + disagree). Never include not_comparable in the denominator.
  • Valid rate = valid / returned.
  • Share of union = by_source[vendor] / union_size. This is not recall.
  • Compliance rate = compliant / returned.

Re-running ingest, resolution and reporting is free and deterministic. Re-running the reference scrape costs about $9.71 and would produce different results, because people change jobs.

Corrections. Errors found after publication are recorded with the original value preserved and visible. Figures are never silently amended.

Version
1.0
Run window
2026-09-10 to 2026-09-12
Measured
2026-09-12
Re-verification due
2026-12-12
Published
2026-09-12
Licence
CC BY 4.0
Attribution
unresolved

Private evaluation

Run this on your segments.

Bring your own ICP definitions. We'll run the same comparison on the segments you actually sell into, and show you the evidence behind every row.

Schedule a demo