Every routing decision in this pack is the real CADi engine output over the Aizle consumer_APP_fraud-v9_1_0 portfolio of account-to-account (A2A) scam payments; nothing here is a hand-computed proxy. The engine routed all 16 scam sends and the full 494-send push-payment book they sit in (478 legitimate sends — §2). This pack records what it decided, why, and the degree to which its A2A-native detectors separate one scam typology from another in the evidence written to the case file.
The conclusion first, on one page, for a senior reader (CRO, Audit, Board Risk, Model Risk). Every figure is a direct engine output; the detail and basis follow in §1–§6.
| Data source | Aizle aizle-consumer_APP_fraud-v9_1_0 (synthetic APP-scam portfolio) |
| Episodes reviewed | 4 |
| Sends reviewed | 16 |
| Full push-payment book (same engine, same run discipline) | 494 sends — 16 scam + 478 legitimate: the complete outbound push-rail population of the sample ledger (§2) |
| Book review surface | 178 of 478 legitimate sends carried ≥1 detector flag (37.2%) — a review-surface density, not a false-positive rate (§2) |
| Fraud events (post-run label) | 16 of 16 |
| Protect decisions | 0 |
| Augment decisions | 16 |
| Automate decisions | 0 |
| Sends carrying ≥1 detector flag | 4 of 16 |
| Why zero Protect holds | No send carried the corroborating higher-severity combination the engine requires before holding a payment; single-signal evidence is held for review, not blocked (§4). |
| Highest regulatory distance | 0.142 (send 01000862965, Romance Scam). An evidence-accumulation score, not a fraud probability (§A6). |
| Confidence (uniform across run) | 0.165 on every send: all 16 resolved to the same review posture (§3) |
| Commercial-cost view total | £0.00: the module executed but this dataset does not populate it (post-decision, report-only; §5) |
| Key finding | The engine differentiates fraud typologies in the evidence it records (varying flags and regulatory distance) but did not alter the routing action: all sends routed Augment, none escalated to Protect or relaxed to Automate. |
| This run evidences | This run does not evidence |
|---|---|
| Ingestion of a third-party APP-scam portfolio by the live decision engine | Fraud-detection accuracy or effectiveness |
| A reproducible routing decision and evidence record for every send | Production performance on live traffic |
| Detector evidence and regulatory citations written per send | A causal reduction in fraud or loss (no treatment arm) |
| Transaction-level traceability and deterministic regeneration (§A3) | Commercial ROI on this dataset (it carries no commercial signals; §5) |
| A corroboration bar for payment holds, observed in action (§4) | Regulatory approval or endorsement |
| Portfolio-scale behaviour over the full push-payment book (494 sends, 478 of them legitimate) — workload split and detector surface (§2) | A false-positive rate (no "should have flagged" ground truth; payee-novelty is window-truncated) or any statistical claim — 10 people, 4 fraud episodes (§2) |
| Separately from this run: the causal-identification gate CADi holds for future treatment-arm data is method-validated end-to-end on a real randomised 3-arm experiment (64,000 customers, public research release), with fail-closed negative controls holding on real data | Any payments-domain causal effect — that validation is of the method, on non-payments (marketing) data, involves no engine routing, and is no part of this run's evidence |
This pack evidences engine behaviour on source data: that the engine ingests a third-party APP-scam portfolio, applies four A2A-native detectors, and produces a reproducible, traceable routing decision and evidence record for every send. It is a decision-evidence pack, not a model-effectiveness or fraud-prevention assessment (§1, §A2).
Claims are marked by evidence level: Fact direct engine output; Inference interpretation of that output; Opinion author conclusion; Limitation a stated limitation. The marks let a reviewer separate what the engine produced from what the author concluded. A formal audit opinion is at §8 and an assurance statement at §A2.
What this pack is, what it is not, and the exact basis of every number.
This pack reports the decisions the CADi engine produced when the four labelled Aizle APP-fraud scam episodes — and the entire 494-send outbound push-payment book they sit in — were run through the live decision engine, cut at the pre-execution authorisation point: the moment before a payment would be released. The fraud label never entered the engine; it is joined afterwards, off the decision path, only to report what each send turned out to be. Legitimate sends carry no label at all: absence is their honest ground truth.
Fact The routing mode, confidence, detector flags, regulatory distance, citations and Consumer-Duty scope below are direct engine output for each send, read from the engine's output record for this run (§A1).
Fact Two views. This pack reports two sides of the engine's decision: a loss-side view (§3–§4, what the engine detects) and a commercial-cost view (§5–§6, the throughput and acceptance trade-offs a payment operation cares about). Limitation The commercial-cost view is post-decision, report-only; it is never a routing input, and on this portfolio it returns £0 for the reason stated in §5.
Limitation Scope. This is the Aizle account-to-account (A2A) APP-scam portfolio, not card-acquiring / PSP traffic. Where §5–§6 use payment-operation vocabulary, they describe the kind of trade-off the engine can assess; no measured authorisation-rate or revenue number is claimed.
Limitation Not claimed. The decisions recorded are detections, not measured interventions. There is no treatment arm, so nothing here measures that the engine's action changed an outcome; no accuracy or performance claim is made; n = 4 episodes. A causal claim would require being in the live decision path with a randomised review-vs-not split, which is out of scope.
Supersedes an earlier proxy pack. A previous version of this document hand-mapped the routing and reported “2 Protect, 2 Augment, 0 Automate” at confidences of 0.60–0.75. Those figures were hand-assigned, not engine output. The real engine result below differs materially and replaces it.
Four scam episodes, 16 sends, £34,026.82 total exposure, read directly from the synthetic ledger.
| Episode | Sends | Exposure | Flagged | Engine routing |
|---|---|---|---|---|
| Impersonation: Bank | 1 | £6,916.67 | 1 | Augment×1 |
| Romance Scam | 13 | £13,023.00 | 1 | Augment×13 |
| Investment Scam | 1 | £10,000.00 | 1 | Augment×1 |
| Impersonation: Family | 1 | £4,087.15 | 1 | Augment×1 |
Fact Grooming-pattern scams unfold as an escalating series, not a single payment. The timeline shows every send of this run's longest episode at its real calendar position and £ size; detector flags appear on 1 of the 13 sends — the opening send, where the beneficiary is still a first-time payee. Inference That is precisely the point of an episode-level evidence record: after the first payment the per-send novelty signal goes quiet, and only the arc — recurring sends to the same beneficiary, stepping up in size — carries the pattern. A single-payment view cannot see it.
Fact The same engine, under the same run discipline (pre-execution decision cut, no labels on the stream), also processed the entire outbound push-rail population of the sample ledger: 494 sends — the 16 labelled scam sends plus 478 legitimate sends — drawn from 5,912 personal transactions across 4 banks (2025-01→07). Card-rail rows (POS, ATM, direct debit) are excluded by shape: they are not push payments, and an APP-scam lane cannot honestly claim to have screened them. Within the push lane the scam share is 3.2%; against the whole ledger it is the familiar sub-1% base-rate picture.
Fact Routing over the book: Augment×494 — identical review posture to the episode result (§3), for the same corroboration-bar reason (§4). The operational meaning is a workload statement: on this data, every push-payment send is routed to human-augmented review, none to straight-through processing and none to a hold.
| Detector activity on the 478 legitimate sends | Fired |
|---|---|
| first-time payee | 177 |
| CoP mismatch | 31 |
| round-sum to new payee | 1 |
| Legitimate sends carrying ≥1 flag | 178 / 478 (37.2%) |
Limitation This is a review-surface density, not a false-positive rate. The ledger carries no ground truth for “should have been flagged”, so no precision figure exists to claim. The dominant signal — payee novelty, 177 fires — is additionally inflated by window truncation: the first observed leg of a long-running standing order reads as “first-time” inside the sample’s seven-month window. The 31 Confirmation-of-Payee mismatches on legitimate sends are genuine name-check mismatches recorded in the source data.
Limitation Sample-slice ceiling. Ten people, four fraud episodes. The book adds portfolio shape — workload split, detector surface, base-rate context — it does not add statistical power. Calibration, precision/recall and P&L claims remain out of reach at n=4 positives; they require the full corpus this sample is drawn from, or a real operator’s data (§A2).
Fact For scale context only: across 350,361 fraud complaint reports to the Canadian Anti-Fraud Centre (2021-01-01 → 2025-09-30), reported losses total CAD $2,687,420,244. Grouped into the UK APP vocabulary this pack uses, the mix is:
| Scam family (UK APP vocabulary) | Complaint reports | Share of all reported loss | Median loss per losing report |
|---|---|---|---|
| Investment | 20,394 | 51.4% | $20,000 |
| Impersonation | 29,427 | 13.8% | $6,000 |
| Romance | 6,779 | 10.9% | $16,490 |
| Purchase / invoice | 71,152 | 6.4% | $244 |
| Advance fee | 12,232 | 4.0% | $2,979 |
Limitation Context, not calibration. These are complaint reports (one row per report, not per transaction), self-selected and self-reported, in Canadian dollars, from a different jurisdiction; the five-family grouping is CADi's own mapping of the source categories (listed beneath). Nothing in this pack is calibrated to these figures and no UK equivalence is claimed. Their role is the opposite of a claim: they show that the episode typologies this pack reports — which fall in the Investment, Impersonation, Romance families — sit in the loss-heavy region of the real scam economy, where investment and romance/impersonation arcs dominate reported loss while high-frequency purchase scams dominate report counts at small median size.
Contains information licensed under the Open Government Licence – Canada. Source: Canadian Anti-Fraud Centre fraud-and-cybercrime reporting data, 2021-01-01 → 2025-09-30 (retrieved 2026-07-04).
Family mapping (authored):
Investment ← Investments; Impersonation ← Bank Investigator, Spear Phishing, Emergency (Jail, Accident, Hospital, Help); Romance ← Romance; Purchase / invoice ← Merchandise, Counterfeit Merchandise, Vendor Fraud, Service, False Billing; Advance fee ← Prize, Loan, GRANT, Foreign Money Offer, Recovery Pitch. Enabler categories (identity theft/fraud, phishing, personal-information —
≈139k reports with a universal $0 loss by source convention) are deliberately excluded from every
family above.
Fact Observed routing split: 0 Protect, 16 Augment, 0 Automate. None of the 16 sends was routed to straight-through processing; every send was routed for human-augmented review.
Fact The engine routed 16 of 16 sends to Augment (human + Ai review) and zero to Automate (straight-through). 4 of 16 sends carried at least one detector flag (7 flags in total), and regulatory distance varies across the portfolio: from 0.000 on the unflagged sends up to 0.142 on the most-evidenced send (01000862965). Consumer-Duty scope returned UNDETERMINED (16); the engine's own scope reasons record why, and §A4 sets them out.
Fact Confidence was uniform at 0.165: every send resolved to the same review posture (Augment, human-confirms, collaborative review) and carries the same recorded confidence. Inference The confidence attaches to that resolved posture rather than varying case by case on this data. Limitation It is therefore not a per-send uncertainty estimate on this portfolio; like the £0 commercial-cost view (§5), the uniformity is information about what this dataset does not vary, and a portfolio that varies the posture-driving inputs would spread it.
Inference On this portfolio the detectors changed what the engine records (the evidence written to the case file) without changing the routing action. §4 sets out the observable reason for that separation.
| Typology | Send | Amount | Detector evidence (severity mix) | Reg. distance | Held at |
|---|---|---|---|---|---|
| Romance Scam | 01000862965 | £862.00 | CoP mismatch, first-time payee 1 higher-severity + 1 diagnostic | 0.142 | Augment |
| Investment Scam | 02001500248 | £10,000.00 | first-time payee, round-sum to new payee 1 higher-severity + 1 diagnostic | 0.117 | Augment |
| Impersonation: Bank | 01000707662 | £6,916.67 | balance-emptying, first-time payee 2 diagnostic | 0.059 | Augment |
| Impersonation: Family | 04000653021 | £4,087.15 | first-time payee 1 diagnostic | 0.029 | Augment |
Fact On each flagged send the CADi detectors wrote the flags above, and their regulatory citations, to the case file; the send is held at Augment for review with interventions build evidence package, hold payout, request additional evidence. The evidence is recorded whether or not it changes the routing action. The most-evidenced send is worked end to end in §A5.
Each row below is traceable to a source record in the engine's output by its transaction ID; the detector flags and citations on each are reproduced verbatim from that record.
| Send ID | Amount | CoP | Conf. | Reg.dist | Detector flags | Routing |
|---|---|---|---|---|---|---|
| Impersonation: Bank · 1 send · £6,916.67 · 1 flagged by a detector | ||||||
| 01000707662 | £6,916.67 | none | 0.165 | 0.059 | balance-emptying, first-time payee | Augment |
| Romance Scam · 13 sends · £13,023.00 · 1 flagged by a detector | ||||||
| 01000862965 | £862.00 | fail+overridden | 0.165 | 0.142 | CoP mismatch, first-time payee | Augment |
| 01000937323 | £1,057.00 | none | 0.165 | 0.000 | none | Augment |
| 01000965968 | £836.00 | none | 0.165 | 0.000 | none | Augment |
| 01001072691 | £1,248.00 | none | 0.165 | 0.000 | none | Augment |
| 01001091962 | £903.00 | none | 0.165 | 0.000 | none | Augment |
| 01001143065 | £1,189.00 | none | 0.165 | 0.000 | none | Augment |
| 01001335351 | £950.00 | none | 0.165 | 0.000 | none | Augment |
| 01001429129 | £758.00 | none | 0.165 | 0.000 | none | Augment |
| 01001600906 | £830.00 | none | 0.165 | 0.000 | none | Augment |
| 01001669057 | £1,449.00 | none | 0.165 | 0.000 | none | Augment |
| 01001850475 | £749.00 | none | 0.165 | 0.000 | none | Augment |
| 01001865752 | £1,020.00 | none | 0.165 | 0.000 | none | Augment |
| 01001898129 | £1,172.00 | none | 0.165 | 0.000 | none | Augment |
| Investment Scam · 1 send · £10,000.00 · 1 flagged by a detector | ||||||
| 02001500248 | £10,000.00 | pass | 0.165 | 0.117 | first-time payee, round-sum to new payee | Augment |
| Impersonation: Family · 1 send · £4,087.15 · 1 flagged by a detector | ||||||
| 04000653021 | £4,087.15 | none | 0.165 | 0.029 | first-time payee | Augment |
Fact The engine separates the most-evidenced send (£862.00 Romance Scam, 01000862965) from an unflagged send by the evidence it records: 2 detector flags and regulatory distance 0.142 versus none and 0.000. Fact It does not separate them in the routing action: both route Augment.
Fact 4 of the 16 sends carry a typology-specific flag stack and a non-zero regulatory distance recorded on the dossier (§3); the remaining sends return the baseline Augment at regulatory distance 0.000. Inference The engine therefore ranks and separates these sends by the evidence it surfaces, but does not change the routing action as a result on this portfolio. The observable reason is as follows:
Inference The recorded evidence stacks make the hold policy a tunable governance choice rather than a fixed property of the engine. The table below re-reads the severity mixes published in §3 under illustrative hold bars; it is arithmetic on the recorded evidence, not a re-run.
| Routing policy | Hold bar | Result on this portfolio |
|---|---|---|
| Review-first (the observed run) | No automated hold; every A2A send held for human review | 16 Augment · 0 holds (what happened) |
| Single-signal hold (illustrative) | Any higher-severity signal | 2 of 16 become hold candidates (the CoP mismatch and round-sum to new payee sends) |
| Evidence-stack hold (illustrative) | One higher-severity signal corroborated by a supporting diagnostic | 2 of 16 become hold candidates (the CoP mismatch and round-sum to new payee sends) |
| Corroborated-severity hold (illustrative) | Two or more higher-severity signals | 0 of 16 qualify |
On this portfolio the two middle bars pick out the same 2 sends: every higher-severity signal recorded here arrived with a supporting diagnostic alongside it.
Limitation The illustrative bars are governance discussion aids, not the engine's configured thresholds, and a hold candidate is not a detection of additional fraud: whether any bar is appropriate is a calibration question for a powered, labelled dataset (§A2). Inference On this portfolio the strictest bar changes nothing and the loosest changes at most 2 of 16 routings: the anti-overblocking posture is visible at every bar.
Fact Supported by the output: the engine ingests a third-party APP-scam portfolio, applies four A2A-native detectors to the facts each send carries, and writes typology-specific evidence (flags, citations, and a varying regulatory distance) to the case file. The detectors fire differentially across the 16 sends.
Limitation Not supported by the output: a change in routing action arising from that evidence (none occurred on this portfolio); and, as stated in §1, any accuracy or causal claim. The decisions recorded are detections, not measured interventions, with n = 4 episodes. Opinion Whether the supporting diagnostics should weigh more heavily toward a hold, and on what calibration, is a routing-policy question for a powered dataset and is not asserted from these four episodes.
Fact Regulatory distance is used, not decorative. The score is not display-only: the engine carries it into its internal value accounting, where it represents the regulatory cost associated with each routing option. Inference It informs that accounting and the case file rather than the routing action: the separation that allows the score to vary while the action does not.
Fact Alongside the loss-side view, the engine also assesses the commercial cost of its decisions: the throughput and acceptance trade-offs a payment operation cares about. On this portfolio the module executed on every send but the dataset does not populate it; the assessment returns £0.00. Inference The reason is the dataset, and it is stated below.
Fact Post-decision, report-only. This assessment is composed after the routing decision and carried onto the case file for audit. It is not a routing input: it does not change which action the engine takes.
Limitation Not a revenue claim. This is the account-to-account APP-scam portfolio, not card-acquiring traffic; the figure is the engine's own internal arithmetic, not a measured authorisation rate, false-decline saving, or acceptance number. No such number is claimed in this pack.
Fact The zero is now shown, not asserted. The assessment decomposes into four commercial channels plus a cross-channel interaction term, and the run record serialises that decomposition for every send of the full 494-send book, with an internal consistency check passing on every one. The table reads it directly; each channel carries the method's own rigor label and the specific reason it prices £0 on this data.
| Channel | Result over the book | Method rigor | Why £0 on this data |
|---|---|---|---|
| Authorization regret | £0.00 · all 494 sends | analytical approximation | The ledger carries no authorisation-stage features, so the assessment runs neutral by construction: every route prices the same. |
| Fraud-intervention regret | £0.00 · all 494 sends | heuristic | The step-up friction inputs a live feed would carry are absent; fixed default priors stand in, and on uniform inputs the priced difference is zero. |
| Routing regret | £0.00 · all 494 sends | analytical approximation | Priced against stand-in (synthetic) acquirer economics; an account-to-account ledger offers no real routing alternatives to price. |
| Retry-strategy regret | £0.00 · all 494 sends | heuristic | A settled-only ledger records no declines and no retry history, so the baseline retry position applies identically to every send. |
| Cross-domain interaction (signed) | £0.00 · all 494 sends | heuristic | Declared stub in this run: the joint cross-channel computation is not performed per transaction, and the row asserts nothing beyond the per-channel values. |
| Four-channel commercial total | £0.00 | — | Sum of the four channels plus interaction; consistency check passes on 494/494 sends |
| Compliance constraint (regulatory distance) | varies · non-zero on 182 of 494 | heuristic | Tracked separately from the commercial sum — a constraint channel, not a £ regret. This is the same evidence differentiation reported in §4, carried into the assessment. |
Limitation One commercial quantity is structurally absent rather than zero: the cost of wrongly declining a legitimate payment. Measuring it requires knowing that a declined payment was in fact legitimate — a ground truth no settled-only ledger can carry, and this book records no declines at all. That quantity is not identified here, and no proxy for it is asserted.
Inference A non-zero commercial result therefore requires inputs this dataset does not carry — the authorisation, decline/retry and step-up signals a live payment feed supplies (§A2 lists what a populated run would need). The mechanism above is the part the engine brings; the missing part is data, not machinery.
Fact Both views on one page: the commercial-cost assessment is £0 on every send; the compliance view is the only one that varies, and it tracks the recorded evidence (§4).
Fact The commercial-cost assessment is £0 on every send — read per send from the serialised post-decision assessment, not asserted (§5 shows the channel decomposition behind it). The compliance view (regulatory distance, which the engine tracks separately) is non-zero on 4 of 16 sends and ranks them by the evidence recorded:
| Reg. distance | Amount | Commercial regret | What carries the signal |
|---|---|---|---|
| 0.142 | £862.00 | £0.00 | compliance / evidence channel |
| 0.117 | £10,000.00 | £0.00 | compliance / evidence channel |
| 0.059 | £6,916.67 | £0.00 | compliance / evidence channel |
| 0.029 | £4,087.15 | £0.00 | compliance / evidence channel |
| 0.000 | ×12 sends | £0.00 | no differentiating signal on either channel |
Fact The engine carries both a loss-side view (§3–§4) and a commercial-cost view (§5). On this portfolio the loss-side detectors fire differentially while the commercial-cost view returns £0. Inference The £0 follows from the dataset: the APP-scam portfolio carries none of the commercial signals a payment gateway would supply.
Inference On this data the only view that separates one send from another is the compliance / regulatory-distance view: the same evidence-level differentiation recorded in §4. A non-zero commercial ledger is gated on gateway data not present here.
Fact The evidence in §1–§6 is the compliance floor: a defensible record of why each of the 16 sends was handled as it was. Inference The value CADi adds on top of that floor is commercial: the same engine that differentiates these sends by evidence also locates where an operator sits in the payments value chain, and therefore which counterfactual exposure they carry. That exposure is the figure a competing approach would have to beat.
Fact Across this portfolio the engine did not merely flag risk; it separated four distinct APP typologies (Impersonation: Bank, Impersonation: Family, Investment Scam, Romance Scam) and ranked them by the evidence each carries, from regulatory distance 0.000 on the unflagged sends up to 0.142 on the most-evidenced send (§4). Inference That ranked, typology-aware separation is the raw material of commercial direction: a generic fraud score says a payment is risky; CADi identifies which kind of risk it is, and which part of an operation owns the cost of getting it wrong.
Inference The same payment carries a different cost depending on position in the chain. CADi reads that position from the routing trail it already produces and points to the exposure that position actually owns: not a generic risk number, but the specific regret a better decision would relieve.
| Position in the chain | The exposure carried | The counterfactual CADi sizes |
|---|---|---|
| Sending / issuing PSP | Reimbursement liability on in-scope APP claims; friction cost on legitimate sends held in error. | The defensible-decision record per send, and the sends where better evidence would have changed the handling: the reimbursement and false-friction regret. |
| Receiving PSP | A share of the sending PSP's reimbursement; inbound mule and beneficiary-account exposure. | The receiving-side evidence and notification trail that attributes, or defends against, that share of the cost. |
| Gateway / orchestrator | Acceptance rate and false declines: every wrongly stopped payment is lost revenue. | The acceptance-versus-friction trade-off on each decision: the revenue a tighter or looser stance moves, in either direction. |
Fact On this portfolio the variation lives in the compliance / evidence channel (4 of 16 sends carry a non-zero regulatory distance; §6). Inference The commercial channel is sized the moment the engine sees the matching traffic for a given position; the mechanism below is the same on both channels.
Fact CADi's decision engine is built on a regret calculation: for every send it holds, alongside the action taken, the cost of the action it did not take. Inference Summed across a book of traffic, that regret is the quantity a counterfactual is measured against: the value at stake in handling the payment one way versus another.
Inference Run that calculation across a representative book for a given position in the chain, and the engine produces a ranked league table of the counterfactuals that move the most value for that position: which commercial dimension, and how much. Opinion That league table is the next layer of value CADi is built to deliver: not "these payments are risky," but "here is the decision that is costing the most, and what relieving it is worth."
A senior-reader summary of what this evidence supports, in one page.
Opinion On the basis of the engine output examined for this run (4 episodes, 16 sends), in the author's opinion:
Limitation This opinion concerns engine behaviour on synthetic source data. It does not opine on model effectiveness, fraud-prevention effectiveness, or reimbursement outcomes; the basis for that exclusion is the assurance statement at §A2.
| Engine | CADi decision engine with account-to-account (A2A) APP-scam detectors |
| Decision point | Pre-execution authorisation: the moment before a payment would be released; post-decision events are excluded |
| Labels on decision stream | None: the fraud label is joined after the run only, off the decision path |
| Source | Aizle aizle-consumer_APP_fraud-v9_1_0 (synthetic) |
| Routing basis | Direct engine output: every figure read from the run's output record |
| Commercial-cost view (§5–§6) | Post-decision, report-only; not a routing input and not a measured revenue figure |
| Cost coefficients | Sourced from cited priors and fail-closed: they do not change any routing decision. Several have no public figure and are not asserted. |
| Identification methods | The powered causal-identification gate (held for future treatment-arm data) is method-validated against a real public randomised experiment; non-payments domain, no bearing on this run's decisions. Details available on request. |
| Real-world context figures (§2) | Canadian Anti-Fraud Centre complaint corpus, Open Government Licence – Canada; context only, calibrates nothing, authored family mapping stated in place |
| Causal claim | None. Detections, not measured interventions. No treatment arm; n = 4. |
This pack reports decisions on fully synthetic data for demonstration; the scope and limits of the claim are stated formally in the assurance statement (§A2). It supersedes a prior version whose routing figures were hand-assigned rather than engine-derived.
Opinion This report is suitable for evidencing engine behaviour: that the CADi engine ingests the stated source data and produces the reproducible, traceable routing decisions and evidence records set out herein.
Limitation This report is not suitable for evidencing model effectiveness, fraud-detection or fraud-prevention effectiveness, or APP-scam reimbursement outcomes. No accuracy, performance, or causal claim is made or implied. The dataset is synthetic; there is no treatment arm and no randomised control, so nothing in this pack measures that any engine action changed an outcome. Conclusions are drawn over n = 4 episodes / 16 sends and should not be generalised to production traffic.
Limitation Each claim this pack does not make is gated on a specific input this run did not have. Stated as data requirements, not promises:
| Claim not made here | What making it would require |
|---|---|
| Detection accuracy / effectiveness | A labelled portfolio at volume, with outcomes joined, and explicit false-positive / false-negative accounting |
| A causal reduction in fraud or loss | The live decision path with a randomised review-vs-not split (a treatment arm), as stated in §1 |
| Calibrated confidence and regulatory distance | Calibration against outcome data on a labelled treatment dataset (§A6) |
| A non-zero commercial ledger | The commercial fields a payment-gateway feed supplies (§5) |
| Consumer-Duty scope determination | Customer-classification fields from the account-holding institution: customer location and retail status (§A4) |
The basis on which the figures in this pack can be regenerated from source.
| Dataset | aizle-consumer_APP_fraud-v9_1_0 (Aizle synthetic APP-fraud portfolio) |
| Engine | CADi decision engine, account-to-account configuration, cut at the pre-execution authorisation point |
| Run scope | 4 episodes · 16 scam sends · full 494-send push-payment book · all routed Augment |
| Output records | Two machine-readable engine-output records — the episode run and the full-book run (the synthetic transaction data is intentionally not redistributed). The scam sends' decisions are asserted identical between the two at run time |
| Regeneration | Re-running the engine over the named dataset reproduces both output records; this document is rendered deterministically from them |
Fact Every figure in §0–§8 is a function of those two output records, which are produced by running the named dataset through the engine. An independent party with the dataset and engine access can reproduce them. The synthetic data and the run harness are available under appropriate terms.
Fact The regulatory anchor behind each detector, and the evidence the engine produced against it: mapping the technical output to the obligations it speaks to.
| Regulation / obligation (engine citation) | Control / evidence mechanism | Evidence produced in this run |
|---|---|---|
| Pay.UK Confirmation of Payee + PSR Specific Direction 17, Pay.UK Confirmation of Payee (account-name-checking standard for UK domestic Faster Payments / CHAPS) | CoP mismatch detector | Detector output on 1 of 16 sends; flag + citation written to the case file |
| PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A) | Round-sum to new payee detector | Detector output on 1 of 16 sends; flag + citation written to the case file |
| PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A) | First-time payee detector | Detector output on 4 of 16 sends; flag + citation written to the case file |
| PSR APP-scam mandatory reimbursement (PS23/3) + FCA Consumer Duty (PRIN 2A) | Balance-emptying detector | Detector output on 1 of 16 sends; flag + citation written to the case file |
| FCA Handbook PRIN, Principle 12 (PRIN 2.1.1) | Consumer-Duty scope assessment | Scope returned UNDETERMINED (16) with reasons recorded per send |
| PSR APP-scam mandatory reimbursement (PS23/3); FCA PRIN 2A (avoid foreseeable harm) | Routing decision trail | Per-send handling mode, authority mode, interventions and confidence recorded for all 16 sends |
Inference The "regulation" column is the regulatory anchor the engine attaches to each detector in its own output. It records the regulatory incentive to detect, not a regulator-defined indicator: the detectors are operational signals informed by published APP-scam typologies.
Fact The engine's scope assessment recorded its reasons on every send: the market is UK, the customer's location is unknown, and retail-customer status was assumed for the demonstration. Inference Scope is therefore left undetermined rather than asserted: the customer-classification fields that would determine it (customer location, retail status) are not carried by this dataset and would come from the account-holding institution's feed. Limitation UNDETERMINED is the honest state of the input data, not a detector failure; asserting scope without those fields would itself be an overclaim.
Fact The highest-regulatory-distance send in the run (01000862965, £862.00 Romance Scam) traced from transaction to final record, each step read directly from the engine output.
Inference This send carried the strongest evidence in the portfolio (an overridden Confirmation-of-Payee mismatch alongside a first-time-payee diagnostic, regulatory distance 0.142), yet held at Augment for review rather than escalating to a Protect hold, because it did not carry the corroborating evidence the engine requires before holding a payment (§4). Fact The flags, citations, interventions and final handling mode shown are the engine's own output for this transaction.
What the score is, what drives it, what a higher value means, and how it is computed: enough to confirm the logic is coherent and reproducible, not a black box. Regulatory distance is an evidence-accumulation score, not a fraud probability.
Fact Regulatory distance is a composite evidence score derived from the severity and confidence of the detector outputs a send accumulates, normalised by the number of detectors run. A value of 0.000 means no differentiating evidence was observed. Higher values indicate increasing divergence from the regulator-aligned baseline expected of a lower-risk account-to-account payment, reflecting stronger evidence accumulation, not a measured fraud probability.
Fact Three inputs, and only these three:
| Driver | Effect on the score |
|---|---|
| Detector severity | Higher-severity signals contribute more than diagnostic signals (weights below) |
| Detector confidence | Each signal's contribution scales by the confidence the detector assigns it |
| Detector count | The total is normalised by the number of detectors run, so one signal among many does not saturate the score |
Fact The score is a deterministic, rule-based composite: no trained model, no hidden state.
Fact Severity weights are fixed constants: higher-severity = 0.85, diagnostic = 0.40, informational = 0.15. On the account-to-account path the normalising detector count is 8. Because the denominator uses the higher-severity weight, the maximum any one detector can contribute before normalisation is its own weight, and a fully-saturated transaction (every detector firing higher-severity at full confidence) reaches 1.000.
Fact The highest-scoring send (01000862965) reproduces exactly from the formula:
| Signal | Category | Weight | Confidence | Contribution |
|---|---|---|---|---|
| CoP mismatch | higher-severity | 0.85 | 0.90 | 0.765 |
| first-time payee | diagnostic | 0.40 | 0.50 | 0.200 |
| Sum of contributions | 0.965 | |||
| Normalise: 0.965 ÷ (8 detectors × 0.85) = 0.965 ÷ 6.80 | 0.142 | |||
Inference The following bands are an interpretation aid for this report: a reader's guide to magnitude, not engine thresholds (the engine consumes the raw score, not these bands).
| Distance | Interpretation |
|---|---|
| 0.000 | No differentiating evidence observed |
| 0.001–0.050 | Weak evidence |
| 0.051–0.100 | Moderate evidence |
| 0.101 + | Strong evidence (relative to this portfolio) |
Fact The distinct evidence stacks observed in this run, lowest to highest:
Limitation Regulatory distance is a transparent heuristic, not a calibrated or trained model. It is reproducible and auditable precisely because it is rule-based: the same inputs always yield the same score, and any reviewer can recompute it from the table above. The severity weights are fixed design choices, not fitted parameters.
Limitation It is therefore an evidence-accumulation measure, not a probability. It does not estimate the likelihood that a payment is fraudulent, and it is not calibrated against fraud outcomes: calibration against outcome data is future work, gated on a labelled treatment dataset (§A2). How a value affects the routing action is set out in §4: regulatory distance informs the engine's value accounting and the case file, and is not, by itself, the routing trigger.