A vendor claim counts as evidence on this platform when we can show where it came from and how far our reading of it can be trusted. In practice that takes five things: the page was retrieved and stored with its address and date; the value was read from a passage or record we keep; the source is ranked by the kind of document it is; the confidence is computed from that rank and, for text read from a page, capped by how often that kind of reading was right when a person checked; and the value is published with all of it. If no value is found, the attribute is published as unknown, with the addresses checked and the date. That is a statement about our search, never about the product. On a copy of the production database from 2026-09-12, 4,400 active statements were documented and 76,550 were unknown.
What a published fact carries
A directory might store soc2 = true. We store a statement with the fields below, and the JSON API and the MCP server return every value under these names.
| Field | What it holds | Why |
|---|---|---|
value | true, a number, a date or a short text | A database constraint allows exactly one, and none for unknown. |
verification_state | verified, corroborated, vendor stated, customer reported, estimated, unknown | The first four are derived from the source tiers. The last two describe our work, not the product. |
sources | address, tier, document type, retrieval date | The database rejects a statement without a source. |
note, excerpt_kind | the passage the value was read from, and its kind | A menu entry that says "HIPAA" is not a sentence about a product. |
confidence | 5 to 99, with its parts and the cap that applied | Never 100: measured is not proven. |
observed_on, last_checked_on | when the value was read, when it was last confirmed | The formula subtracts 8 points per year of age. |
Five steps from an address to a fact
- Retrieve and store. A research task names an address and the attributes to look for there. We fetch it within the site's robots.txt and a per-host rate limit and store the text as a snapshot with its date. A hand check later judges the value against that snapshot, not against the page as it looks today.
- Find the value. A reader looks for it: a pattern for "SOC 2 Type II", a field of the page's structured markup, a number with a unit. It keeps the passage around the match. Text that repeats word for word on the pages of several products of one host is page frame, not a statement about any of them.
- Classify the excerpt. Is it running text, a navigation entry, structured markup or a quotation, and does it name the product? Both answers come from the excerpt and the product name, without another fetch.
- Rank the source. Tier 1 is primary documentation: product, API, security and pricing documentation, technical specifications, regulatory filings. Tier 3 is independent reporting, tier 5 community signal. The base confidence runs from 90 for tier 1 to 25 for tier 5; corroboration adds 6 points with two sources, 9 with three and 10 with four or more.
- Cap, then publish. For text read from a page, the number is capped at the measured precision of the reader and at that of the excerpt kind. The lower cap wins, and the evidence drawer names it. If no reader found a value, the attribute is published as unknown.
The excerpt kind is measured, not assumed
The second cap is the newer one. On 2026-09-11 a review of the live site pointed at two statements published as verified at 90: an ISO 27001 statement whose excerpt was a menu ("Trust center HIPAA GDPR ISO 27001 Resources Business blog") and a target-group statement read from a customer quotation about families looking for care for their children. Either statement may be true. The published excerpt did not carry it.
Instead of guessing what a menu entry is worth, we measured it. All 1,815 hand judgements from earlier precision samples were classified by excerpt kind after the fact, and the share of correct readings was computed per kind.
daten/messungen/2026-09-11-belegart.json.Structured markup was right in 97.4 percent of 153 cases, running text in 81.1 percent of 1,140, a navigation entry in 72.1 percent of 516. Whether the excerpt names the product matters as well: 87.6 percent correct when it does, 75.8 percent when it does not.
A rule that withdraws every statement read from navigation would have removed 372 correct readings to get rid of 144 wrong ones, so there is no such rule. The measured rate became a cap, and the excerpt kind is printed next to the excerpt. One class is not written at all: a target-group or industry attribute read from a customer quotation, because the customer speaks about the customer.
Why a tier 1 statement is rarely published at 90
All 4,400 documented statements on the 2026-09-12 copy have a tier 1 source, whose base confidence is 90. Only 891 are published at 90. A worked example with the numbers in force on 2026-09-13 shows why.
Take a statement read on 2026-09-02 from a product documentation page, by a pattern without a precision measurement of its own, from a menu entry that does not name the product.
- Source: tier 1, one source, observed less than a year ago: 90.
- Reader cap: without its own measurement, the average across all readers applies, 71.0 percent from the corpus sample of 2026-09-02: 71.
- Excerpt cap: navigation without the product's name was right in 68.6 percent of 398 hand judgements: 69.
- Published: the lower cap wins, so the statement reads verified, 69, and the evidence drawer names the excerpt kind as the reason.
3,312 of 4,400 documented statements are published below what their source alone would give them. That is the intended effect. 90 says the page is primary documentation; 69 says how often a reading like this one was right when a person checked. The number is a measured share of correct readings in a sample, not a probability that this one statement is true.
What the database looked like on 2026-09-12
These rules produce a particular shape of database. On the copy from 2026-09-12, 1,822 products and 1,787 organizations carry 80,951 active statements, and 94.6% of them are unknown.
The unknowns are deliberate. An attribute searched without result is recorded with the addresses checked, so a reader can see where we looked, when, and whose search came up empty. Nothing is overwritten either: 4,232 statements on the copy were superseded by a newer statement, and 6,102 were withdrawn, each with its reason. Both stay in the history.
The state corroborated is empty because of how states are derived, not because second sources are missing. 882 documented statements have two or more sources, but a statement with a tier 1 source is verified, and corroborated can only follow from a best source below tier 1. 692 of the documented statements cite a regulatory filing; the others cite official product, pricing, API or security documentation or technical specifications.
Why 5 of 1,822 product pages are in the index
A product page exists as soon as the product does. Whether search engines are asked to index it is decided by a gate per page type, and a page outside the index still answers. A product page needs 8 documented core facts from 2 distinct source documents, and the thresholds for every page type are listed on the methodology page. On the copy the gate is strict because the data is thin: 1,701 documented statements and 39,402 unknowns across all products, 0.9 and 21.6 per product.
The other page types follow the same logic. 51 of 1,787 vendor pages are in the index; a vendor page needs at least one product with substance. None of the 9,310 comparison pairs is; a comparison needs 6 criteria documented on both sides and 3 documented differences. The category directory and the comparison start page are in the index; the 113 individual category pages are not yet; each needs 10 products with substance and 6 filters that separate them.
The sentence we do not write
Suppose a research task checked six addresses of a vendor for a SOC 2 report and the reader found nothing. There are two ways to publish that.
Vendor X has no SOC 2 report.
No SOC 2 report found. Checked 6 addresses on 2026-09-01.
The first sentence is a claim about a company. It turns false the moment the report sits behind a login, on a trust portal we did not reach, or on a page added a week later. Section 824(1) of the German Civil Code makes whoever asserts an untrue fact that can damage another's credit, earnings or prospects liable for the damage, even without knowing it was untrue, if they should have known. The second sentence stays true in every one of those cases, because it describes what we did, and it tells the vendor which addresses to look at. On a product page an unknown reads "Checked 6 addresses on 2026-09-01" with the addresses linked, and the API returns verification_state: "unknown" with a null value, never false.
How often we are wrong
Precision is measured on samples that a person judges by hand against the stored snapshot. The corpus-wide figure on the methodology page comes from 6 successive samples of 100 documented statements drawn on 2026-09-02, with a precision of 0.74, 0.72, 0.65, 0.73, 0.71, 0.71; the last run is the one that stands. Two later samples were drawn from populations we expected to be worse, and they were.
2026-09-02-praezision-lauf6.json, 2026-09-08-nachbarn.json, 2026-09-09-reichweite.json.The neighbours sample changed how we check. A misreading often belongs to the page rather than to one attribute, for example a cloud provider's menu on the page of one of its products, so one wrong value is a reason to check the other values from the same snapshot by hand. It is not a verdict on them: 39 of the 60 neighbours were right. And a judgement stays. A statement a person marked as wrong is not restored by the next automated reading of the same passage.
Questions about this
Does "unknown" mean the feature is missing?
No. It means we searched the addresses listed next to it and found no reliable evidence on that date. A feature can sit behind a login, on a page we did not reach, or in a document sent on request. For a buyer it is a question to put to the vendor; for a vendor it is a list of pages to check.
Why is a tier 1 source 90 and never 100?
Because a current page can still be misread and an old one can be outdated. The formula stops at 99, and the caps below that are measured shares of correct readings, not opinions.
Why not withdraw everything read from a menu?
Because 72.1 percent of those readings were right. Withdrawing them would remove 372 correct statements to get rid of 144 wrong ones. The cap keeps them at the confidence they have earned, and the excerpt kind lets you judge the passage yourself.
Can a vendor pay to change a state or a confidence?
No. Nothing that ranks is for sale, and a paid vendor subscription is read by no ranking, gate or badge. A vendor can name a page on its own domain; it becomes a research task and is retrieved, stored, read and ranked like any other page. A vendor can name a page, never set a value.
How do I get these fields into my own tools?
The JSON API and the read-only MCP server described on the developers page return every value with its state, confidence and its parts, excerpt kind, dates and sources. Both call the same functions, so the same question gets the same answer on either.
Method and sources
Database figures describe a copy of the production database from 2026-09-12, the one used for that day's live audit, which re-read two statements and recomputed the page table on it. They change with every campaign run. One command reproduces all of them on any copy, using the functions the pages use: node werkzeuge/messungen/blog-zahlen.mjs; npm run seiten recomputes the page counts.
- Precision by excerpt kind:
daten/messungen/2026-09-11-belegart.json, 1,815 judgements, computed bynode werkzeuge/messungen/belegart-messen.mjs. - Corpus precision:
2026-09-02-praezision.json,2026-09-02-praezision-lauf2.json,2026-09-02-praezision-lauf3.json,2026-09-02-praezision-lauf4.json,2026-09-02-praezision-lauf5.json,2026-09-02-praezision-lauf6.json, drawn and judged withnode werkzeuge/praezision.mjs. - Pre-selected samples:
2026-09-08-nachbarn.jsonand2026-09-09-reichweite.jsonin the same directory. - Confidence formula, caps and gate thresholds: read when this page is rendered from
src/lib/konfidenz.jsandsrc/lib/seitentor.js, the files the methodology page reads. - German Civil Code (BGB), Section 824, gesetze-im-internet.de/bgb/__824.html, retrieved 2026-09-13. Subsection 2 exempts a communication whose sender does not know it is untrue, where sender or recipient has a legitimate interest in it.