Trust
Accuracy is tested, not promised.
Every AI product in this category asks you to take its word. We'd rather show the receipts. Same-repository changes that can affect model output must pass the real-model labelled gate; non-AI changes and fork contributions run a clearly labelled deterministic harness smoke. The checked-in measured run below publishes its own date and commit.
when Kestrel says a change is material, how often it's right
of the genuinely material changes, how many Kestrel catches
claims whose quoted evidence verbatim-matches the source page
Measured on 45 hand-authored benchmark cases modelled on realistic changelog and pricing language, including adversarial prompt-injection cases, labelled and checked by hand. A further 30-case interpretation-quality set is available for explicit real-model calibration runs; normal CI exercises its deterministic, report-only harness while the judge is calibrated. Small set, honestly stated — it grows with every misclassification we find, and published numbers change only when a measured artifact is deliberately refreshed.
The machinery behind the numbers.
No citation, no claim
Every assertion must quote its source verbatim. A hard gate in the pipeline checks each quoted span against the actual page text of the cited URL — claims that can't be grounded are dropped before anyone sees them. Not a policy; code.
Independent verification
Material findings are re-checked by a separate model before they can interrupt your Slack channel: claim-level entailment against the evidence, plus an independent second opinion on impact. Kestrel's confidence is never self-reported.
A blocked page is never “no change”
Bot walls, half-rendered pages and loading screens are detected and reported as crawl failures — loudly. Weekly coverage receipts tell you exactly what was checked, so you always know the difference between quiet and blind.
The gate fails closed
The eval isn't a dashboard we glance at — it's a merge gate. On a trusted change that can alter model behaviour, a missing real-model credential fails CI rather than silently substituting a smoke result. Non-AI diffs and forks are explicitly reported as harness smoke runs.
The platform underneath is audited too: Kestrel runs on Vercel, whose infrastructure holds SOC 2 Type 2 and ISO 27001 attestations.
What these numbers are not
They are measurements on our labelled set — not a universal guarantee, and not a claim that any AI is infallible. What they mean in practice: the analyst's judgment is tested the way software is tested, every miss we find becomes a new labelled case, and if a release dips below the gates you saw above, it doesn't ship. If you find a Kestrel claim that doesn't hold up against its own citation, we want it — tell us and it enters the eval set.