APP / PSR Regulatory Reporting Pack

APP Fraud Decision Evidence Pack

Every routing decision in this pack is the real CADi engine output over the Aizle consumer_APP_fraud-v9_1_0 portfolio of account-to-account (A2A) scam payments; nothing here is a hand-computed proxy. The engine routed all 16 scam sends and the full 494-send push-payment book they sit in (478 legitimate sends — §2). This pack records what it decided, why, and the degree to which its A2A-native detectors separate one scam typology from another in the evidence written to the case file.

Version 1.3 · 4 July 2026
Engine output · 4 episodes · 16 scam sends · 494-send book
CONFIDENTIAL
Synthetic data · no production claim

0.  Executive summary

The conclusion first, on one page, for a senior reader (CRO, Audit, Board Risk, Model Risk). Every figure is a direct engine output; the detail and basis follow in §1–§6.

Data sourceAizle aizle-consumer_APP_fraud-v9_1_0 (synthetic APP-scam portfolio)
Episodes reviewed4
Sends reviewed16
Full push-payment book (same engine, same run discipline)494 sends — 16 scam + 478 legitimate: the complete outbound push-rail population of the sample ledger (§2)
Book review surface178 of 478 legitimate sends carried ≥1 detector flag (37.2%) — a review-surface density, not a false-positive rate (§2)
Fraud events (post-run label)16 of 16
Protect decisions0
Augment decisions16
Automate decisions0
Sends carrying ≥1 detector flag4 of 16
Why zero Protect holdsNo send carried the corroborating higher-severity combination the engine requires before holding a payment; single-signal evidence is held for review, not blocked (§4).
Highest regulatory distance0.142 (send 01000862965, Romance Scam). An evidence-accumulation score, not a fraud probability (§A6).
Confidence (uniform across run)0.165 on every send: all 16 resolved to the same review posture (§3)
Commercial-cost view total£0.00: the module executed but this dataset does not populate it (post-decision, report-only; §5)
Key findingThe engine differentiates fraud typologies in the evidence it records (varying flags and regulatory distance) but did not alter the routing action: all sends routed Augment, none escalated to Protect or relaxed to Automate.
This run evidencesThis run does not evidence
Ingestion of a third-party APP-scam portfolio by the live decision engine Fraud-detection accuracy or effectiveness
A reproducible routing decision and evidence record for every send Production performance on live traffic
Detector evidence and regulatory citations written per send A causal reduction in fraud or loss (no treatment arm)
Transaction-level traceability and deterministic regeneration (§A3) Commercial ROI on this dataset (it carries no commercial signals; §5)
A corroboration bar for payment holds, observed in action (§4) Regulatory approval or endorsement
Portfolio-scale behaviour over the full push-payment book (494 sends, 478 of them legitimate) — workload split and detector surface (§2) A false-positive rate (no "should have flagged" ground truth; payee-novelty is window-truncated) or any statistical claim — 10 people, 4 fraud episodes (§2)
Separately from this run: the causal-identification gate CADi holds for future treatment-arm data is method-validated end-to-end on a real randomised 3-arm experiment (64,000 customers, public research release), with fail-closed negative controls holding on real data Any payments-domain causal effect — that validation is of the method, on non-payments (marketing) data, involves no engine routing, and is no part of this run's evidence
Reading this report

This pack evidences engine behaviour on source data: that the engine ingests a third-party APP-scam portfolio, applies four A2A-native detectors, and produces a reproducible, traceable routing decision and evidence record for every send. It is a decision-evidence pack, not a model-effectiveness or fraud-prevention assessment (§1, §A2).

Claims are marked by evidence level: Fact direct engine output; Inference interpretation of that output; Opinion author conclusion; Limitation a stated limitation. The marks let a reviewer separate what the engine produced from what the author concluded. A formal audit opinion is at §8 and an assurance statement at §A2.

1.  Scope and basis of the evidence

What this pack is, what it is not, and the exact basis of every number.

This pack reports the decisions the CADi engine produced when the four labelled Aizle APP-fraud scam episodes — and the entire 494-send outbound push-payment book they sit in — were run through the live decision engine, cut at the pre-execution authorisation point: the moment before a payment would be released. The fraud label never entered the engine; it is joined afterwards, off the decision path, only to report what each send turned out to be. Legitimate sends carry no label at all: absence is their honest ground truth.

Basis of the figures, and the boundary of the claim

Fact The routing mode, confidence, detector flags, regulatory distance, citations and Consumer-Duty scope below are direct engine output for each send, read from the engine's output record for this run (§A1).

Fact Two views. This pack reports two sides of the engine's decision: a loss-side view (§3–§4, what the engine detects) and a commercial-cost view (§5–§6, the throughput and acceptance trade-offs a payment operation cares about). Limitation The commercial-cost view is post-decision, report-only; it is never a routing input, and on this portfolio it returns £0 for the reason stated in §5.

Limitation Scope. This is the Aizle account-to-account (A2A) APP-scam portfolio, not card-acquiring / PSP traffic. Where §5–§6 use payment-operation vocabulary, they describe the kind of trade-off the engine can assess; no measured authorisation-rate or revenue number is claimed.

Limitation Not claimed. The decisions recorded are detections, not measured interventions. There is no treatment arm, so nothing here measures that the engine's action changed an outcome; no accuracy or performance claim is made; n = 4 episodes. A causal claim would require being in the live decision path with a randomised review-vs-not split, which is out of scope.

Supersedes an earlier proxy pack. A previous version of this document hand-mapped the routing and reported “2 Protect, 2 Augment, 0 Automate” at confidences of 0.60–0.75. Those figures were hand-assigned, not engine output. The real engine result below differs materially and replaces it.

2.  The portfolio

Four scam episodes, 16 sends, £34,026.82 total exposure, read directly from the synthetic ledger.

EpisodeSendsExposure FlaggedEngine routing
Impersonation: Bank1£6,916.671Augment×1
Romance Scam13£13,023.001Augment×13
Investment Scam1£10,000.001Augment×1
Impersonation: Family1£4,087.151Augment×1
Exposure by typology
Romance Scam · 38%Investment Scam · 29%Impersonation: Bank · 20%Impersonation: Family · 12%
Anatomy of the longest-running episode · Romance Scam — every send at its real calendar position and £ size
£862£1,057£836£1,248£903£1,189£950£758£830£1,449£749£1,020£1,1722025-01-292025-07-30182 days · 13 sendsdetector-flagged sendno flag recorded

Fact Grooming-pattern scams unfold as an escalating series, not a single payment. The timeline shows every send of this run's longest episode at its real calendar position and £ size; detector flags appear on 1 of the 13 sends — the opening send, where the beneficiary is still a first-time payee. Inference That is precisely the point of an episode-level evidence record: after the first payment the per-send novelty signal goes quiet, and only the arc — recurring sends to the same beneficiary, stepping up in size — carries the pattern. A single-payment view cannot see it.

The full push-payment book the episodes sit in

Fact The same engine, under the same run discipline (pre-execution decision cut, no labels on the stream), also processed the entire outbound push-rail population of the sample ledger: 494 sends — the 16 labelled scam sends plus 478 legitimate sends — drawn from 5,912 personal transactions across 4 banks (2025-01→07). Card-rail rows (POS, ATM, direct debit) are excluded by shape: they are not push payments, and an APP-scam lane cannot honestly claim to have screened them. Within the push lane the scam share is 3.2%; against the whole ledger it is the familiar sub-1% base-rate picture.

Fact Routing over the book: Augment×494 — identical review posture to the episode result (§3), for the same corroboration-bar reason (§4). The operational meaning is a workload statement: on this data, every push-payment send is routed to human-augmented review, none to straight-through processing and none to a hold.

Detector activity on the 478 legitimate sendsFired
first-time payee177
CoP mismatch31
round-sum to new payee1
Legitimate sends carrying ≥1 flag 178 / 478 (37.2%)

Limitation This is a review-surface density, not a false-positive rate. The ledger carries no ground truth for “should have been flagged”, so no precision figure exists to claim. The dominant signal — payee novelty, 177 fires — is additionally inflated by window truncation: the first observed leg of a long-running standing order reads as “first-time” inside the sample’s seven-month window. The 31 Confirmation-of-Payee mismatches on legitimate sends are genuine name-check mismatches recorded in the source data.

Limitation Sample-slice ceiling. Ten people, four fraud episodes. The book adds portfolio shape — workload split, detector surface, base-rate context — it does not add statistical power. Calibration, precision/recall and P&L claims remain out of reach at n=4 positives; they require the full corpus this sample is drawn from, or a real operator’s data (§A2).

The scam economy this portfolio samples: real-world prevalence context

Fact For scale context only: across 350,361 fraud complaint reports to the Canadian Anti-Fraud Centre (2021-01-01 → 2025-09-30), reported losses total CAD $2,687,420,244. Grouped into the UK APP vocabulary this pack uses, the mix is:

Scam family (UK APP vocabulary)Complaint reports Share of all reported lossMedian loss per losing report
Investment20,39451.4%$20,000
Impersonation29,42713.8%$6,000
Romance6,77910.9%$16,490
Purchase / invoice71,1526.4%$244
Advance fee12,2324.0%$2,979
Share of all reported CAFC loss, by scam family
Investment51.4%Impersonation13.8%Romance10.9%Purchase / invoice6.4%Advance fee4.0%

Limitation Context, not calibration. These are complaint reports (one row per report, not per transaction), self-selected and self-reported, in Canadian dollars, from a different jurisdiction; the five-family grouping is CADi's own mapping of the source categories (listed beneath). Nothing in this pack is calibrated to these figures and no UK equivalence is claimed. Their role is the opposite of a claim: they show that the episode typologies this pack reports — which fall in the Investment, Impersonation, Romance families — sit in the loss-heavy region of the real scam economy, where investment and romance/impersonation arcs dominate reported loss while high-frequency purchase scams dominate report counts at small median size.

Contains information licensed under the Open Government Licence – Canada. Source: Canadian Anti-Fraud Centre fraud-and-cybercrime reporting data, 2021-01-01 → 2025-09-30 (retrieved 2026-07-04).
Family mapping (authored): Investment ← Investments; Impersonation ← Bank Investigator, Spear Phishing, Emergency (Jail, Accident, Hospital, Help); Romance ← Romance; Purchase / invoice ← Merchandise, Counterfeit Merchandise, Vendor Fraud, Service, False Billing; Advance fee ← Prize, Loan, GRANT, Foreign Money Offer, Recovery Pitch. Enabler categories (identity theft/fraud, phishing, personal-information — ≈139k reports with a universal $0 loss by source convention) are deliberately excluded from every family above.

3.  The decisions the engine produced

Fact Observed routing split: 0 Protect, 16 Augment, 0 Automate. None of the 16 sends was routed to straight-through processing; every send was routed for human-augmented review.

Routing split · 16 sends
Protect0Augment16Automate0

Fact The engine routed 16 of 16 sends to Augment (human + Ai review) and zero to Automate (straight-through). 4 of 16 sends carried at least one detector flag (7 flags in total), and regulatory distance varies across the portfolio: from 0.000 on the unflagged sends up to 0.142 on the most-evidenced send (01000862965). Consumer-Duty scope returned UNDETERMINED (16); the engine's own scope reasons record why, and §A4 sets them out.

Fact Confidence was uniform at 0.165: every send resolved to the same review posture (Augment, human-confirms, collaborative review) and carries the same recorded confidence. Inference The confidence attaches to that resolved posture rather than varying case by case on this data. Limitation It is therefore not a per-send uncertainty estimate on this portfolio; like the £0 commercial-cost view (§5), the uniformity is information about what this dataset does not vary, and a portfolio that varies the posture-driving inputs would spread it.

Inference On this portfolio the detectors changed what the engine records (the evidence written to the case file) without changing the routing action. §4 sets out the observable reason for that separation.

The 4 flagged sends, most-evidenced first

TypologySendAmountDetector evidence (severity mix)Reg. distanceHeld at
Romance Scam01000862965£862.00CoP mismatch, first-time payee
1 higher-severity + 1 diagnostic
0.142Augment
Investment Scam02001500248£10,000.00first-time payee, round-sum to new payee
1 higher-severity + 1 diagnostic
0.117Augment
Impersonation: Bank01000707662£6,916.67balance-emptying, first-time payee
2 diagnostic
0.059Augment
Impersonation: Family04000653021£4,087.15first-time payee
1 diagnostic
0.029Augment

Fact On each flagged send the CADi detectors wrote the flags above, and their regulatory citations, to the case file; the send is held at Augment for review with interventions build evidence package, hold payout, request additional evidence. The evidence is recorded whether or not it changes the routing action. The most-evidenced send is worked end to end in §A5.

Each row below is traceable to a source record in the engine's output by its transaction ID; the detector flags and citations on each are reproduced verbatim from that record.

Send IDAmountCoPConf. Reg.distDetector flagsRouting
Impersonation: Bank · 1 send · £6,916.67 · 1 flagged by a detector
01000707662£6,916.67none0.1650.059balance-emptying, first-time payeeAugment
Romance Scam · 13 sends · £13,023.00 · 1 flagged by a detector
01000862965£862.00fail+overridden0.1650.142CoP mismatch, first-time payeeAugment
01000937323£1,057.00none0.1650.000noneAugment
01000965968£836.00none0.1650.000noneAugment
01001072691£1,248.00none0.1650.000noneAugment
01001091962£903.00none0.1650.000noneAugment
01001143065£1,189.00none0.1650.000noneAugment
01001335351£950.00none0.1650.000noneAugment
01001429129£758.00none0.1650.000noneAugment
01001600906£830.00none0.1650.000noneAugment
01001669057£1,449.00none0.1650.000noneAugment
01001850475£749.00none0.1650.000noneAugment
01001865752£1,020.00none0.1650.000noneAugment
01001898129£1,172.00none0.1650.000noneAugment
Investment Scam · 1 send · £10,000.00 · 1 flagged by a detector
02001500248£10,000.00pass0.1650.117first-time payee, round-sum to new payeeAugment
Impersonation: Family · 1 send · £4,087.15 · 1 flagged by a detector
04000653021£4,087.15none0.1650.029first-time payeeAugment

4.  Differentiation observed in the evidence, not in the action

Fact The engine separates the most-evidenced send (£862.00 Romance Scam, 01000862965) from an unflagged send by the evidence it records: 2 detector flags and regulatory distance 0.142 versus none and 0.000. Fact It does not separate them in the routing action: both route Augment.

Regulatory distance by evidence stack: the differentiation, visualised
CoP mismatch + first-time payee0.142first-time payee + round-sum0.117balance-emptying + first-time payee0.059first-time payee0.029No evidence0.000

Fact 4 of the 16 sends carry a typology-specific flag stack and a non-zero regulatory distance recorded on the dossier (§3); the remaining sends return the baseline Augment at regulatory distance 0.000. Inference The engine therefore ranks and separates these sends by the evidence it surfaces, but does not change the routing action as a result on this portfolio. The observable reason is as follows:

The same evidence under different hold bars

Inference The recorded evidence stacks make the hold policy a tunable governance choice rather than a fixed property of the engine. The table below re-reads the severity mixes published in §3 under illustrative hold bars; it is arithmetic on the recorded evidence, not a re-run.

Routing policyHold barResult on this portfolio
Review-first (the observed run) No automated hold; every A2A send held for human review 16 Augment · 0 holds (what happened)
Single-signal hold (illustrative) Any higher-severity signal 2 of 16 become hold candidates (the CoP mismatch and round-sum to new payee sends)
Evidence-stack hold (illustrative) One higher-severity signal corroborated by a supporting diagnostic 2 of 16 become hold candidates (the CoP mismatch and round-sum to new payee sends)
Corroborated-severity hold (illustrative) Two or more higher-severity signals 0 of 16 qualify

On this portfolio the two middle bars pick out the same 2 sends: every higher-severity signal recorded here arrived with a supporting diagnostic alongside it.

Limitation The illustrative bars are governance discussion aids, not the engine's configured thresholds, and a hold candidate is not a detection of additional fraud: whether any bar is appropriate is a calibration question for a powered, labelled dataset (§A2). Inference On this portfolio the strictest bar changes nothing and the loosest changes at most 2 of 16 routings: the anti-overblocking posture is visible at every bar.

What this does and does not establish

Fact Supported by the output: the engine ingests a third-party APP-scam portfolio, applies four A2A-native detectors to the facts each send carries, and writes typology-specific evidence (flags, citations, and a varying regulatory distance) to the case file. The detectors fire differentially across the 16 sends.

Limitation Not supported by the output: a change in routing action arising from that evidence (none occurred on this portfolio); and, as stated in §1, any accuracy or causal claim. The decisions recorded are detections, not measured interventions, with n = 4 episodes. Opinion Whether the supporting diagnostics should weigh more heavily toward a hold, and on what calibration, is a routing-policy question for a powered dataset and is not asserted from these four episodes.

Fact Regulatory distance is used, not decorative. The score is not display-only: the engine carries it into its internal value accounting, where it represents the regulatory cost associated with each routing option. Inference It informs that accounting and the case file rather than the routing action: the separation that allows the score to vary while the action does not.

5.  The commercial-cost view

Fact Alongside the loss-side view, the engine also assesses the commercial cost of its decisions: the throughput and acceptance trade-offs a payment operation cares about. On this portfolio the module executed on every send but the dataset does not populate it; the assessment returns £0.00. Inference The reason is the dataset, and it is stated below.

What this view is, and what it is not

Fact Post-decision, report-only. This assessment is composed after the routing decision and carried onto the case file for audit. It is not a routing input: it does not change which action the engine takes.

Limitation Not a revenue claim. This is the account-to-account APP-scam portfolio, not card-acquiring traffic; the figure is the engine's own internal arithmetic, not a measured authorisation rate, false-decline saving, or acceptance number. No such number is claimed in this pack.

Fact The zero is now shown, not asserted. The assessment decomposes into four commercial channels plus a cross-channel interaction term, and the run record serialises that decomposition for every send of the full 494-send book, with an internal consistency check passing on every one. The table reads it directly; each channel carries the method's own rigor label and the specific reason it prices £0 on this data.

ChannelResult over the bookMethod rigor Why £0 on this data
Authorization regret£0.00 · all 494 sendsanalytical approximationThe ledger carries no authorisation-stage features, so the assessment runs neutral by construction: every route prices the same.
Fraud-intervention regret£0.00 · all 494 sendsheuristicThe step-up friction inputs a live feed would carry are absent; fixed default priors stand in, and on uniform inputs the priced difference is zero.
Routing regret£0.00 · all 494 sendsanalytical approximationPriced against stand-in (synthetic) acquirer economics; an account-to-account ledger offers no real routing alternatives to price.
Retry-strategy regret£0.00 · all 494 sendsheuristicA settled-only ledger records no declines and no retry history, so the baseline retry position applies identically to every send.
Cross-domain interaction (signed)£0.00 · all 494 sendsheuristicDeclared stub in this run: the joint cross-channel computation is not performed per transaction, and the row asserts nothing beyond the per-channel values.
Four-channel commercial total £0.00 Sum of the four channels plus interaction; consistency check passes on 494/494 sends
Compliance constraint (regulatory distance) varies · non-zero on 182 of 494 heuristic Tracked separately from the commercial sum — a constraint channel, not a £ regret. This is the same evidence differentiation reported in §4, carried into the assessment.

Limitation One commercial quantity is structurally absent rather than zero: the cost of wrongly declining a legitimate payment. Measuring it requires knowing that a declined payment was in fact legitimate — a ground truth no settled-only ledger can carry, and this book records no declines at all. That quantity is not identified here, and no proxy for it is asserted.

Inference A non-zero commercial result therefore requires inputs this dataset does not carry — the authorisation, decline/retry and step-up signals a live payment feed supplies (§A2 lists what a populated run would need). The mechanism above is the part the engine brings; the missing part is data, not machinery.

6.  The two views side by side

Fact Both views on one page: the commercial-cost assessment is £0 on every send; the compliance view is the only one that varies, and it tracks the recorded evidence (§4).

Fact The commercial-cost assessment is £0 on every send — read per send from the serialised post-decision assessment, not asserted (§5 shows the channel decomposition behind it). The compliance view (regulatory distance, which the engine tracks separately) is non-zero on 4 of 16 sends and ranks them by the evidence recorded:

Reg. distanceAmount Commercial regretWhat carries the signal
0.142£862.00£0.00compliance / evidence channel
0.117£10,000.00£0.00compliance / evidence channel
0.059£6,916.67£0.00compliance / evidence channel
0.029£4,087.15£0.00compliance / evidence channel
0.000×12 sends £0.00no differentiating signal on either channel
The two-sided finding

Fact The engine carries both a loss-side view (§3–§4) and a commercial-cost view (§5). On this portfolio the loss-side detectors fire differentially while the commercial-cost view returns £0. Inference The £0 follows from the dataset: the APP-scam portfolio carries none of the commercial signals a payment gateway would supply.

Inference On this data the only view that separates one send from another is the compliance / regulatory-distance view: the same evidence-level differentiation recorded in §4. A non-zero commercial ledger is gated on gateway data not present here.

7.  The value beyond the report: counterfactual position

Fact The evidence in §1–§6 is the compliance floor: a defensible record of why each of the 16 sends was handled as it was. Inference The value CADi adds on top of that floor is commercial: the same engine that differentiates these sends by evidence also locates where an operator sits in the payments value chain, and therefore which counterfactual exposure they carry. That exposure is the figure a competing approach would have to beat.

Fact Across this portfolio the engine did not merely flag risk; it separated four distinct APP typologies (Impersonation: Bank, Impersonation: Family, Investment Scam, Romance Scam) and ranked them by the evidence each carries, from regulatory distance 0.000 on the unflagged sends up to 0.142 on the most-evidenced send (§4). Inference That ranked, typology-aware separation is the raw material of commercial direction: a generic fraud score says a payment is risky; CADi identifies which kind of risk it is, and which part of an operation owns the cost of getting it wrong.

Position in the chain decides which counterfactual applies

Inference The same payment carries a different cost depending on position in the chain. CADi reads that position from the routing trail it already produces and points to the exposure that position actually owns: not a generic risk number, but the specific regret a better decision would relieve.

Position in the chainThe exposure carried The counterfactual CADi sizes
Sending / issuing PSP Reimbursement liability on in-scope APP claims; friction cost on legitimate sends held in error. The defensible-decision record per send, and the sends where better evidence would have changed the handling: the reimbursement and false-friction regret.
Receiving PSP A share of the sending PSP's reimbursement; inbound mule and beneficiary-account exposure. The receiving-side evidence and notification trail that attributes, or defends against, that share of the cost.
Gateway / orchestrator Acceptance rate and false declines: every wrongly stopped payment is lost revenue. The acceptance-versus-friction trade-off on each decision: the revenue a tighter or looser stance moves, in either direction.

Fact On this portfolio the variation lives in the compliance / evidence channel (4 of 16 sends carry a non-zero regulatory distance; §6). Inference The commercial channel is sized the moment the engine sees the matching traffic for a given position; the mechanism below is the same on both channels.

How CADi turns this into a number: the regret ledger

Fact CADi's decision engine is built on a regret calculation: for every send it holds, alongside the action taken, the cost of the action it did not take. Inference Summed across a book of traffic, that regret is the quantity a counterfactual is measured against: the value at stake in handling the payment one way versus another.

Inference Run that calculation across a representative book for a given position in the chain, and the engine produces a ranked league table of the counterfactuals that move the most value for that position: which commercial dimension, and how much. Opinion That league table is the next layer of value CADi is built to deliver: not "these payments are risky," but "here is the decision that is costing the most, and what relieving it is worth."

What CADi gives you that ordinary reporting does not

8.  Audit opinion

A senior-reader summary of what this evidence supports, in one page.

Opinion

Opinion On the basis of the engine output examined for this run (4 episodes, 16 sends), in the author's opinion:

  1. The engine produced reproducible outputs from the source data: every routing mode, flag, citation, regulatory distance and Consumer-Duty scope in this pack is read directly from that file and can be regenerated (§A1, §A3). Fact
  2. All 16 routing decisions are traceable to a source record by transaction ID, with the detector flags and regulatory citations that drove them attached per send (§3, §A4, §A5). Fact
  3. No routing escalation occurred: all 16 sends routed Augment; none reached Protect or Automate. The threshold logic that produced this is stated and observable (§4). Fact
  4. Fraud-typology differentiation was observed at the evidence layer only (in the flags and regulatory distance recorded) and not in the routing action taken (§4). Inference

Limitation This opinion concerns engine behaviour on synthetic source data. It does not opine on model effectiveness, fraud-prevention effectiveness, or reimbursement outcomes; the basis for that exclusion is the assurance statement at §A2.

A1.  Provenance and limitations

EngineCADi decision engine with account-to-account (A2A) APP-scam detectors
Decision pointPre-execution authorisation: the moment before a payment would be released; post-decision events are excluded
Labels on decision streamNone: the fraud label is joined after the run only, off the decision path
SourceAizle aizle-consumer_APP_fraud-v9_1_0 (synthetic)
Routing basisDirect engine output: every figure read from the run's output record
Commercial-cost view (§5–§6)Post-decision, report-only; not a routing input and not a measured revenue figure
Cost coefficientsSourced from cited priors and fail-closed: they do not change any routing decision. Several have no public figure and are not asserted.
Identification methodsThe powered causal-identification gate (held for future treatment-arm data) is method-validated against a real public randomised experiment; non-payments domain, no bearing on this run's decisions. Details available on request.
Real-world context figures (§2)Canadian Anti-Fraud Centre complaint corpus, Open Government Licence – Canada; context only, calibrates nothing, authored family mapping stated in place
Causal claimNone. Detections, not measured interventions. No treatment arm; n = 4.

This pack reports decisions on fully synthetic data for demonstration; the scope and limits of the claim are stated formally in the assurance statement (§A2). It supersedes a prior version whose routing figures were hand-assigned rather than engine-derived.

A2.  Regulatory assurance statement

Scope of assurance

Opinion This report is suitable for evidencing engine behaviour: that the CADi engine ingests the stated source data and produces the reproducible, traceable routing decisions and evidence records set out herein.

Limitation This report is not suitable for evidencing model effectiveness, fraud-detection or fraud-prevention effectiveness, or APP-scam reimbursement outcomes. No accuracy, performance, or causal claim is made or implied. The dataset is synthetic; there is no treatment arm and no randomised control, so nothing in this pack measures that any engine action changed an outcome. Conclusions are drawn over n = 4 episodes / 16 sends and should not be generalised to production traffic.

What a stronger claim would require

Limitation Each claim this pack does not make is gated on a specific input this run did not have. Stated as data requirements, not promises:

Claim not made hereWhat making it would require
Detection accuracy / effectiveness A labelled portfolio at volume, with outcomes joined, and explicit false-positive / false-negative accounting
A causal reduction in fraud or loss The live decision path with a randomised review-vs-not split (a treatment arm), as stated in §1
Calibrated confidence and regulatory distance Calibration against outcome data on a labelled treatment dataset (§A6)
A non-zero commercial ledger The commercial fields a payment-gateway feed supplies (§5)
Consumer-Duty scope determination Customer-classification fields from the account-holding institution: customer location and retail status (§A4)

A3.  Reproducibility

The basis on which the figures in this pack can be regenerated from source.

Datasetaizle-consumer_APP_fraud-v9_1_0 (Aizle synthetic APP-fraud portfolio)
EngineCADi decision engine, account-to-account configuration, cut at the pre-execution authorisation point
Run scope4 episodes · 16 scam sends · full 494-send push-payment book · all routed Augment
Output recordsTwo machine-readable engine-output records — the episode run and the full-book run (the synthetic transaction data is intentionally not redistributed). The scam sends' decisions are asserted identical between the two at run time
RegenerationRe-running the engine over the named dataset reproduces both output records; this document is rendered deterministically from them

Fact Every figure in §0–§8 is a function of those two output records, which are produced by running the named dataset through the engine. An independent party with the dataset and engine access can reproduce them. The synthetic data and the run harness are available under appropriate terms.

A4.  Controls and regulation mapping

Fact The regulatory anchor behind each detector, and the evidence the engine produced against it: mapping the technical output to the obligations it speaks to.

Regulation / obligation (engine citation)Control / evidence mechanism Evidence produced in this run
Pay.UK Confirmation of Payee + PSR Specific Direction 17, Pay.UK Confirmation of Payee (account-name-checking standard for UK domestic Faster Payments / CHAPS) CoP mismatch detectorDetector output on 1 of 16 sends; flag + citation written to the case file
PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A)Round-sum to new payee detectorDetector output on 1 of 16 sends; flag + citation written to the case file
PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A)First-time payee detectorDetector output on 4 of 16 sends; flag + citation written to the case file
PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A)Balance-emptying detectorDetector output on 1 of 16 sends; flag + citation written to the case file
FCA Handbook PRIN, Principle 12 (PRIN 2.1.1)Consumer-Duty scope assessmentScope returned UNDETERMINED (16) with reasons recorded per send
PSR APP-scam mandatory reimbursement (PS23/3); FCA PRIN 2A (avoid foreseeable harm)Routing decision trailPer-send handling mode, authority mode, interventions and confidence recorded for all 16 sends

Inference The "regulation" column is the regulatory anchor the engine attaches to each detector in its own output. It records the regulatory incentive to detect, not a regulator-defined indicator: the detectors are operational signals informed by published APP-scam typologies.

Why Consumer-Duty scope returned undetermined

Fact The engine's scope assessment recorded its reasons on every send: the market is UK, the customer's location is unknown, and retail-customer status was assumed for the demonstration. Inference Scope is therefore left undetermined rather than asserted: the customer-classification fields that would determine it (customer location, retail status) are not carried by this dataset and would come from the account-holding institution's feed. Limitation UNDETERMINED is the honest state of the input data, not a detector failure; asserting scope without those fields would itself be an overclaim.

A5.  Worked case: one send, end to end

Fact The highest-regulatory-distance send in the run (01000862965, £862.00 Romance Scam) traced from transaction to final record, each step read directly from the engine output.

Transaction
£862.00 A2A send · 01000862965 · Romance Scam
Facts presented
CoP fail+overridden, first-time payee true, to-personal true
Detectors fired
CoP mismatch, first-time payee (1 higher-severity + 1 diagnostic)
Evidence generated
2 regulatory citations written to the dossier; interventions build evidence package, hold payout, request additional evidence
Regulatory distance
0.142, the highest in the run
Routing logic
The recorded evidence (1 higher-severity + 1 diagnostic) did not meet the engine's bar for a payment hold; the send is held for human-augmented review
Final record
Augment · human confirms · human-augmented review · confidence 0.165
Reading the lineage

Inference This send carried the strongest evidence in the portfolio (an overridden Confirmation-of-Payee mismatch alongside a first-time-payee diagnostic, regulatory distance 0.142), yet held at Augment for review rather than escalating to a Protect hold, because it did not carry the corroborating evidence the engine requires before holding a payment (§4). Fact The flags, citations, interventions and final handling mode shown are the engine's own output for this transaction.

A6.  Regulatory distance: methodology

What the score is, what drives it, what a higher value means, and how it is computed: enough to confirm the logic is coherent and reproducible, not a black box. Regulatory distance is an evidence-accumulation score, not a fraud probability.

Definition

Fact Regulatory distance is a composite evidence score derived from the severity and confidence of the detector outputs a send accumulates, normalised by the number of detectors run. A value of 0.000 means no differentiating evidence was observed. Higher values indicate increasing divergence from the regulator-aligned baseline expected of a lower-risk account-to-account payment, reflecting stronger evidence accumulation, not a measured fraud probability.

What drives it

Fact Three inputs, and only these three:

DriverEffect on the score
Detector severityHigher-severity signals contribute more than diagnostic signals (weights below)
Detector confidenceEach signal's contribution scales by the confidence the detector assigns it
Detector countThe total is normalised by the number of detectors run, so one signal among many does not saturate the score

The formula

Fact The score is a deterministic, rule-based composite: no trained model, no hidden state.

distance = clamp[0,1](  ( Σsignals severity_weight × confidence )  ÷  ( detector_count × severity_weightHIGH )  )

Fact Severity weights are fixed constants: higher-severity = 0.85, diagnostic = 0.40, informational = 0.15. On the account-to-account path the normalising detector count is 8. Because the denominator uses the higher-severity weight, the maximum any one detector can contribute before normalisation is its own weight, and a fully-saturated transaction (every detector firing higher-severity at full confidence) reaches 1.000.

Worked reproduction: the Romance Scam send (0.142)

Fact The highest-scoring send (01000862965) reproduces exactly from the formula:

SignalCategoryWeightConfidence Contribution
CoP mismatchhigher-severity0.850.900.765
first-time payeediagnostic0.400.500.200
Sum of contributions0.965
Normalise: 0.965 ÷ (8 detectors × 0.85) = 0.965 ÷ 6.800.142

Interpreting a value

Inference The following bands are an interpretation aid for this report: a reader's guide to magnitude, not engine thresholds (the engine consumes the raw score, not these bands).

DistanceInterpretation
0.000No differentiating evidence observed
0.001–0.050Weak evidence
0.051–0.100Moderate evidence
0.101 +Strong evidence (relative to this portfolio)

The portfolio, ranked by evidence

Fact The distinct evidence stacks observed in this run, lowest to highest:

0.000
No differentiating evidence
0.029
first-time payee
0.059
balance-emptying + first-time payee
0.117
first-time payee + round-sum to new payee
0.142
CoP mismatch + first-time payee
Maturity and honest limits

Limitation Regulatory distance is a transparent heuristic, not a calibrated or trained model. It is reproducible and auditable precisely because it is rule-based: the same inputs always yield the same score, and any reviewer can recompute it from the table above. The severity weights are fixed design choices, not fitted parameters.

Limitation It is therefore an evidence-accumulation measure, not a probability. It does not estimate the likelihood that a payment is fraudulent, and it is not calibrated against fraud outcomes: calibration against outcome data is future work, gated on a labelled treatment dataset (§A2). How a value affects the routing action is set out in §4: regulatory distance informs the engine's value accounting and the case file, and is not, by itself, the routing trigger.