Trust

Accuracy is tested, not promised.

Every AI product in this category asks you to take its word. We'd rather show the receipts. Same-repository changes that can affect model output must pass the real-model labelled gate; non-AI changes and fork contributions run a clearly labelled deterministic harness smoke. The checked-in measured run below publishes its own date and commit.

Published measured run — real model, full labelled setn=45 · 2026-07-06 · 3ef65bf
Precisiongate pass
1.000
CI gate: ≥ 0.80

when Kestrel says a change is material, how often it's right

Recallgate pass
1.000
CI gate: ≥ 0.80

of the genuinely material changes, how many Kestrel catches

Citation validitygate pass
1.000
CI gate: ≥ 0.95

claims whose quoted evidence verbatim-matches the source page

confusion: tp=25 fp=0 fn=0 tn=20 · accuracy=1.000 · f1=1.000

Measured on 45 hand-authored benchmark cases modelled on realistic changelog and pricing language, including adversarial prompt-injection cases, labelled and checked by hand. A further 30-case interpretation-quality set is available for explicit real-model calibration runs; normal CI exercises its deterministic, report-only harness while the judge is calibrated. Small set, honestly stated — it grows with every misclassification we find, and published numbers change only when a measured artifact is deliberately refreshed.

The machinery behind the numbers.

No citation, no claim

Every assertion must quote its source verbatim. A hard gate in the pipeline checks each quoted span against the actual page text of the cited URL — claims that can't be grounded are dropped before anyone sees them. Not a policy; code.

Independent verification

Material findings are re-checked by a separate model before they can interrupt your Slack channel: claim-level entailment against the evidence, plus an independent second opinion on impact. Kestrel's confidence is never self-reported.

A blocked page is never “no change”

Bot walls, half-rendered pages and loading screens are detected and reported as crawl failures — loudly. Weekly coverage receipts tell you exactly what was checked, so you always know the difference between quiet and blind.

The gate fails closed

The eval isn't a dashboard we glance at — it's a merge gate. On a trusted change that can alter model behaviour, a missing real-model credential fails CI rather than silently substituting a smoke result. Non-AI diffs and forks are explicitly reported as harness smoke runs.

The platform underneath is audited too: Kestrel runs on Vercel, whose infrastructure holds SOC 2 Type 2 and ISO 27001 attestations.

What these numbers are not

They are measurements on our labelled set — not a universal guarantee, and not a claim that any AI is infallible. What they mean in practice: the analyst's judgment is tested the way software is tested, every miss we find becomes a new labelled case, and if a release dips below the gates you saw above, it doesn't ship. If you find a Kestrel claim that doesn't hold up against its own citation, we want it — tell us and it enters the eval set.