Test run · July 2026 · not a client deliverable
CADi’s first end-to-end run against a public synthetic benchmark — 6,362,620 PaySim rows on SCM 3.0.0. It demonstrates the mechanism: a decision made at the authorization moment, with a named reason, and the fraud label never in scope. It is not a client deliverable and not a fraud-accuracy claim. It reports recall without the operating point that decides whether the screen could actually be run — the flag alerts on 23.9% of all transactions, roughly 190 alerts for every fraud found — states no baseline comparison, and recorded no provenance stamp while it ran. That re-instrumented run has since landed (SCM 3.2.0, 1 August 2026). It reproduced this run’s confusion matrix exactly - the same 8,012 catches and 201 misses - and added the transaction amount this run never recorded. On that measure recall is 87.5% by value against the 97.6% by count quoted below, because the 201 missed frauds average 5.7× larger than the 8,012 caught. Every recall figure in this document is by count unless it says otherwise.
Every transaction decided at authorization time. The fraud label was never in scope — it was used only afterwards, to score the decisions.
= channel where fraud occurs in this dataset. The detector concentrates in exactly those channels — and is silent on CASH IN — without ever being told where fraud lives.
332,507 flags fired in channels that carry no fraud at all. Scope the screen to the two fraud-carrying channels and the review queue shrinks by 22%, precision rises 0.53% → 0.67%, and recall is unchanged — measured on the full stream before anyone commits to the control.
Every missed fraud is a partial drain — the account was not emptied to zero. That is not a loss statistic; it is the specification for the next detector.
Perfect silence on the one channel where the mechanism cannot occur. A statistical score sprays; a mechanism tracks real money movement.