UCI ONLINE RETAIL · REPORT 03 / 03
Real data went in.
Its limits stayed visible.
Completeness is not the same as fitness for a decision. We kept the commercial view broad, the engine cohort explicit and the missing outcomes unknown. The result is an auditable ingestion rehearsal - not a payment-accuracy claim.
Real public retail data, not client data. Descriptive business intelligence, not causal inference. Positive-sale value is not settled revenue or profit. Business hypotheses and external research are labelled separately; no current-market benchmark is implied.
01 identity gap
Missing identity changes the observed mix.
Anonymous orders remain in all-country sales totals but cannot be assigned to individual customer histories.
Explore the exact figures
| Measure | Share of all-country positive totals |
|---|---|
| Anonymous share of orders | 7.2% |
| Anonymous share of sale value | 16.5% |
1,440 orders carry £1,755,276.64 in positive-sale value. Dropping them would remove a larger share of value than of orders.
At source-line level, 135,080 of 541,909 rows lack customer identity. That 24.9% row rate has a different denominator from the bars above.
02 basket mix
The unknown buyer has the larger basket.
Means do not identify who these buyers are. Missing customer identity must not be converted into an invented customer segment or a fraud label.
Explore the exact figures
| Measure | GBP per positive order |
|---|---|
| Known customer | £480 |
| Anonymous | £1,219 |
03 two populations
Full workbook coverage. Different eligible populations.
Lines, invoice IDs and compound order groups are different units. They are shown separately - not as a conversion funnel.
Commercial analysis · all countries
- 541,909 source lines examined.
- 530,104 positive sale lines retained.
- 19,960 distinct positive invoice IDs become 20,002 order groups when timestamp/customer/country conflicts are kept separate.
- 38 countries; unknown customers retained.
Engine rehearsal · UK only
- 485,123 positive UK candidate lines.
- 18,019 candidate invoice IDs.
- 39 inconsistent invoices excluded.
- 17,980 invoice-derived records processed; 17,980 dossiers and 17,980 ladder artifacts present.
04 engine handling
Processing succeeded. Authority stayed with people.
All decisions remained AUGMENT and required Human Confirms. The handling tiers below are engine outputs, not good/bad customer classifications.
Explore the exact figures
| Measure | Invoice-derived decisions |
|---|---|
| Advisory augment | 12,568 |
| Collaborative augment | 5,412 |
Every decision requested additional evidence. There were no detector flags; without outcome labels, that is not a zero-fraud finding. Decisions were not operationally actioned.
05 claim boundaries
Do not fill the blanks with certainty.
The adapter uses invoice time as a disclosed authorization-boundary proxy. The source supplies neither payment authorizations nor fraud outcomes.
- Absent outcomes are excluded from fraud, approval and outcome metrics. Adapter defaults are not observations.
- Arrival modelling is a separate, non-authoritative exercise. M2 and M3 remain uncleared.
- Public licensing permits reuse with attribution; it is not source-owner approval of CADi’s contract.
- 10,624 cancellation-or-negative-quantity rows carry £896,812.49 in absolute recorded value. They are not matched to sales: this is not a refund rate.
- Engine eligibility uses cancellation-prefix rules; its cancellation count is not interchangeable with the broader BI negative-quantity category.
06 ingestion lessons
Better inputs would unlock better questions.
A practical preparation list for a future partner dataset - not facts present in this workbook.
For stronger business intelligence
- Stable customer and order identities, with explicit missingness.
- Linked cancellations and returns, instead of unmatched negative lines.
- Gross/net/tax/discount definitions, shipping cost and margin.
- Documented timezone, recording process and trading calendar.
For stronger engine evaluation
- Actual event types and timestamps; no invoice-time substitution.
- Observed outcomes with provenance, held outside engine input where required.
- A source-owner contract and exact permission scope.
- Locked training/holdout windows and a prospective evaluation protocol.