Methodology

What counts as evidence.

These rules apply to every benchmark we publish. Anything specific to one study — its sample, its instruments, its limitations — lives with that study, and each one is linked at the foot of this page.

Vocabulary

Four things are recorded about every cell.

A score on its own is an opinion. Each cell in a scorecard therefore carries what we found, how strong it is, whether a number applies, and whether we retrieved it ourselves. The definitions below are read directly from the published scorecard.json, so the wording here and the values in the file a reader downloads are the same text.

Evidence stateWhat we found when we looked for an artifact. Recorded per cell.
ValueMeans
observedThe artifact was found and is present.
observed_absentWe checked where it would be and confirmed it is not there.
not_disclosedCould not be determined from public sources. Not the same as confirmed missing.
Source gradeHow strong the thing we found is. A score is never better than the grade behind it.
ValueMeans
directly_verifiableWe retrieved the artifact ourselves and it is independently checkable.
third_party_reportedSourced from a tracker, press report, or analysis we did not produce.
vendor_claimedAsserted by the vendor with no independently checkable artifact behind it.
Score stateWhether a number applies at all.
ValueMeans
scoredA numeric score applies.
insufficient_historyThe vendor has not existed long enough to generate the evidence this dimension requires. score is null, never 0.
Verification statusWhether we retrieved the artifact ourselves, or have not yet.
ValueMeans
verifiedArtifact retrieved by us on retrieved_at.
pendingNot yet independently retrieved. Score provisional.

The scale

04, ordinal, higher is better.

The scale is ordinal, so the distance between 2 and 3 is not claimed to equal the distance between 3 and 4. Every dimension publishes a written anchor for each value, and a score is an assertion that the anchor was met — not a judgement call recorded as a number.

No composite score, and no overall rank. No weighted composite, total, or overall rank is published. Any weighting would encode our judgement about a buyer we have not met.

Missing data

Absence is not scored the same as silence.

Collapsing the two would let opacity score better than honesty. A vendor that documents a weakness therefore scores above a vendor that documents nothing, which is the opposite of what a naive scoring rule produces.

Missing-data policy, read from the published scorecard artifact.
Confirmed absent scores below undisclosedYes
Inference from non-disclosureNot permitted. We do not treat silence as evidence of a practice.
Known residual biasFavours vendors with resources to maintain public documentation surfaces, which correlates with size and funding.

Right of reply and corrections

What a vendor can expect.

Each vendor receives the rows concerning it, the anchors used, and the evidence cited before publication. Factual corrections with a citable artifact are applied and logged. Disagreements about interpretation are published alongside the score, not merged into it. No response is recorded as no response and does not change a score. Scores are never changed in exchange for access, data, or advertising.

Whether it has been conducted is recorded per study, not here. A policy is a commitment; only the study can say whether it was met for that study. Each benchmark states its own status, and the scorecard artifact carries it as a field a reader can check.

Per-study methodology

How each benchmark applies these rules.

Every published result also ships the raw data behind it. If a number here is wrong, it should be possible to prove that from the files rather than to argue about it.

Schedule a demo

Run the same test on your workflow.

Bring your requirements and your shortlist. We'll run one free evaluation and show you the evidence.

Schedule a demo