Corrections

What we got wrong, and when we fixed it.

We publish measurements, so our own errors are part of the record. Every correction to a published benchmark is listed here with its date and the original wording, so the change can be checked rather than taken on trust.

  1. · Benchmark 01 — web search APIs

    Three methodology claims corrected against the run database

    §10 said: “Two You.com calls failed on a documented 50-word query limit. This is a real constraint on long questions, not measurement noise.” The run count and the vendor were right, but only one of the two failures was the word limit. The other was a 504 network error that consumed 19.4 seconds and exhausted all three retries — which is measurement noise. Both remain scored as misses, because a failure to return a result is a result, but the page now distinguishes the floor the study imposed from the one the product did.

    §9 claimed all settings were provider defaults except those listed. Four were neither default nor listed: country=us and search_lang=en to Brave, gl=us and hl=en to Serper, topic=general to Tavily, and includeImages=false to Linkup. All four narrow results to US English. Retrieval rates are comparable only if the requests are, so a reader reproducing the run from §9 alone would have sent different queries and had no way to tell whose numbers were wrong.

    §9 did not disclose that Keenable ran at its “pro” tier. All 400 sampled Keenable responses in the published run report it.

  2. · Benchmark 02 — GTM data providers

    Published: Apollo against Clay across five ICPs

    First publication. No corrections to date.

  3. · Benchmark 01 — web search APIs

    Published: seven web search APIs, with Serper excluded from the reliability chart

    First publication. Scores were verified the same day they were published.

    Recorded at publication rather than corrected afterwards: Serper scored at the floor on all five reliability dimensions. A single all-zero row compresses the visual range of a five-axis chart without adding information, so it was removed from that chart — not from the study. The full scored row is retained in scorecard.json so the decision can be checked, and Serper's retrieval accuracy and latency remain in the results.

    Two caveats were published with it. Removing an outlier also narrows the spread across the remaining six, which flatters them. And Serper's price carries a low-confidence flag, because its pricing page was unreachable and its credit packs sit behind signup.

How we handle errors

What happens when we find one.

We correct the page and log the date and the change here.

The original wording stays quoted inside the correction, so the change is visible rather than silent.

Figures on every page are computed from the published data files, so a corrected number cannot disagree with the file a reader downloads.

Corrections that change a published result are noted on the study page itself, not only here.