rotascale authority and evidence for AI agents See the demo

Platform — Assurance

We attack our own controls, and publish where they fail.

A platform that detects personal identifiers is making a claim about an adversarial world. We test our detectors is effort, and effort is what a vendor asserts about itself. Evidence has a denominator.

red_team_report EU profile · iban, eu-vat, eu-national-id 0 defects
Grid
21 cells · identifier × mutation
Attacks evaluated
1,336 · 639 held
Cells with a defect
0 of 21
Cited by
GDPR Art. 32(1)(d)

Computed live, never stored. A saved grid keeps asserting yesterday's answer about a detector somebody has since changed — which for this artefact is the whole failure mode.

A red cell is not automatically a failure

Most of the grid is reachable only through a / or a newline — characters that separate genuinely different fields, which a detector must refuse to treat as internal. Closing those cells would buy the number with false positives, and a false positive here makes somebody erase evidence they needed.

Residual
Remove the characters the detector must refuse, ask it again, and it finds the identifier. The miss was caused by something it was right to reject.
Defect
It still misses. Something it should have handled let the identifier through — and this is the only number a fix moves, so this is the number the evidence gates on.

Two wrong gates preceded this one. Gating on every attack failed let a regime read 100% while all fourteen cells were evadable; gating on every cell evaded would have condemned a detector for refusing a character it is supposed to refuse. Both are recorded in the source, because the next person will reach for one of them.

What the grid changed

The first full run was 21 cells, 21 evaded, 2 defects: a VAT number and a national identifier both got through when written with spaces, which is how both are printed. The fix was not tolerance. It was structure.

Checked, not matched

EU VAT is now recognised per member state at that state's exact length, with the published check digit run for the twelve that have one — each algorithm confirmed against a real registration before it shipped. An algorithm we could not confirm that way is absent rather than guessed, because a wrong checksum produces confident false negatives.

observed →

Confidence is a property of the match

Twelve states publish a check digit and fifteen do not, so no single answer for the family could be true. A German number is observed and a Cypriot one asserted, in the same result, and the screen says partly checked.

observed →

The noise floor is measured

A generic eleven-digit pattern reported 100% of random strings as national identifiers — order references, phone numbers, timestamps. Three checksummed schemes report 2.1%, and a test measures it so a future scheme cannot quietly restore the old behaviour.

observed →

A pack nothing attacks is named

A jurisdiction pack with no synthetic seed is absent from the grid, and its cells are absent from the denominator with it. A denominator that shrinks in silence reads as coverage, so it is stated.

observed →

The cost, disclosed. A member state whose scheme is not implemented is no longer detected, and fifteen VAT states are recognised by length alone and reported as asserted rather than observed. Both are stated beside the result rather than absorbed. Every seed is synthetic and checksum-valid; no real identifier is used, stored or transmitted.

What a clean screen is not

The fairness screen compares outcomes across attributes a subject declared — nothing is inferred, and a decision carrying no declaration is reported as undisclosed rather than estimated. A screen that fills gaps by inference is producing a claim about individuals in order to make a claim about a population.

A clean screen is not proof of fairness and a flagged screen is not proof of discrimination. Both are grounds for a controlled analysis somebody else performs, and approval-rate parity ignores how the underlying populations differ. That belongs in the output rather than in a footnote — for the same reason every screen listing an observe grant says it refuses nothing.