· laboratory study

Integrity assurance for contributed data, models and inference records.

One offline chain measures every artefact against a reference the team owns, then publishes each conclusion with its confidence, its evidence and the limits of what it can prove. Every number on this page is read from a committed receipt and names its denominator.

At a false-alarm budget frozen on clean models first, its strongest measured rule catches 47.9% of substitution attacks at dose 0.25 and 93.7% at dose 0.50 — while a second rule that passes the same false-alarm test catches almost none of them.

Static build over committed study results · not a live engine · no network calls in the assurance path

Overview

Six numbers, each with its denominator

The clean-null ledger and the attack ladder are different experiments on different populations. They are shown separately everywhere on this page for that reason — the first three tiles are the detection headline from the ladder, the last three are the clean-null study it was calibrated on.

Zero attacked positives in the clean ledger. That makes TPR, precision and expected loss undefined in that ledger, not zero — so detection is measured on the separate attack ladder, at thresholds the clean ledger froze first. Both facts are load-bearing; neither replaces the other.
Headline detection · frozen thresholds

At a frozen 5% false-alarm budget, this is what gets caught

Thresholds were fixed on 28,313 held-out clean models before any attacked arm was scored, so these rates are not tuned on the attacks. They are the four cells the repository's own number audit enforces — one class and one dose at a time, never pooled.

The measurement that matters most here

Bias lift is the backdoor-like case

Two rules, same false-alarm rate, different eyes

RefDiv-mean vs CTC-mean at the identical frozen threshold

Clean-null false alarms

Calibration sits near the 5% target — the intervals include it

Thresholds come from the calibration half's negatives only; the false-alarm rate is measured on a disjoint held-out half that never contributed to the threshold.

Frozen-threshold attack ladder

Detection rises with measured damage — inside one family at a time

Dose units differ by family (prune fraction, noise sigma, logit units), so the families are never pooled and never share a severity axis.

Arms caught, by dose

Hollow markers are below the declared floor of 20 and are not estimates.

Same laboratory cohort

Implemented baseline checks vs CVIAF — a tradeoff, not dominance

These comparators are written by the team. They are not a reproduced operational system, and the denominators are small enough that they are illustrative.

Model axis

2 scored positives, 4 negatives, 2 ASR-gated exclusions

Data axis

7 poisoned assets, 1 clean asset; 432 poisoned and 1,560 clean samples

The unfavourable result stays visible. At sample level the per-sample baseline reaches TPR 84.03% / precision 93.80% while CVIAF's FDR decision reaches 16.90% / 58.87%. The asset-level false-alarm reduction does not cancel that, and the two units must not be swapped.
Inference records

8 attacked and 2 genuine records — abstentions kept visible

A record checker that abstains is not the same as one that clears the record. Both are shown, along with the key assumption this run was signed under.

Corpus inventories

Two populations that overlap — never added together

The merged study and the export record share shards. Summing them would double count.

Method and limits

What this evidence does not establish

These limits are properties of the current receipts, not of the design intent.

Evidence explorer

Every number traced to a file, field and denominator

Rows marked local artifact are not tracked in git — the large corpora live outside the repository, so those rows are readable here but not linkable.

Research and references

Where these methods come from

The signals on this page are published methods, and the code names them in its own docstrings. Two repository documents carry the research and its verification log; the table below says which of those methods this build actually runs, which it only designs against, and where each lives — and what measurement has changed since the dossier was written.