Methodology
What each bench measures, how the cases are built, and how the work is graded.
- Environments
- KYC, Transaction Monitoring, Sanctions Screening, Fraud Prevention
- Benches
- KYC Bench, TM Bench, Sanctions Bench, Fraud Bench
- Case data
- Synthetic; no real customer data
- Sources
- Publicly available information, used for patterns only
- Sets
- Sealed and held out
- Graded on
- The decision and the evidence behind it
- Results
- Coming soon
01The environments
Each environment is one team’s queue at a fictional institution, with its own written policies and risk appetite. The agent receives an alert, a screening hit or an onboarding file, with a brief that is the same for every case. The records sit behind tools: the customer and KYC profile, accounts, transactions, counterparties, screening results and prior alerts. Nothing is pasted into the prompt. The agent decides what to pull, as an analyst would, and applies the institution’s policy rather than its own sense of what looks suspicious.
It ends the way an analyst does: a disposition, and a rationale that cites the records it relied on. Every model gets the same brief, the same tools and the same limits.
02The four benches
One bench for each environment, each with its own caseload.
KYC Bench
- Queue
- Onboarding for individuals and businesses, CDD through EDD
- Agent receives
- Application data, identity and business documents, registry extracts, screening results and the institution’s CDD and EDD policy
- Decides
- Approve, apply EDD, request information or decline, with a customer risk rating
- Graded on
- Requirements applied for the customer type, product and jurisdiction; gaps in verification found; beneficial owners identified; rating and rationale
- Results
- Coming soon
TM Bench
- Queue
- Transaction monitoring alerts, fiat and crypto
- Agent receives
- The alert, customer profile and expected activity, accounts, transactions, counterparties, wallet exposure and prior alerts
- Decides
- Close, escalate, request information, or file a SAR/STR
- Graded on
- Red flags found and ruled out, typology, the reporting decision, and a narrative that cites the records
- Results
- Coming soon
Sanctions Bench
- Queue
- Name, entity and wallet hits, adjudicated before the payment moves
- Agent receives
- The hit and the list entry, customer or counterparty identifiers, ownership records and the payment details
- Decides
- True match or false positive; release, block, reject or escalate
- Graded on
- Identifiers compared, ownership and control checked, route and goods reviewed, and the reason recorded
- Results
- Coming soon
Fraud Bench
- Queue
- Flagged logins, payments and new accounts
- Agent receives
- Device, login and behaviour signals, payment history and payee details
- Decides
- Hold, release, contact the customer or exit the relationship
- Graded on
- The pattern identified, loss weighed against turning away a genuine customer, and the rationale
- Results
- Coming soon
03How cases are built
Real case files can’t be used: SAR confidentiality and customer privacy keep them inside each institution. So every case is synthetic.
Cases start from publicly available information on how financial crime is carried out and detected. We use it for patterns, not text: the roles involved, the documents that should exist, the points where an alert would plausibly fire. No sentence of any source is carried over.
Everything else is invented: the customers, accounts, transactions, screening hits and the policies the agent works to. A case is built to be internally consistent, with enough ordinary activity around the suspicious core that the agent has to discriminate. A payroll account still pays salaries; a trading business still has seasonal peaks.
Each case keeps a private record of the pattern it draws on, held apart from the case itself. A reviewer can check that the behaviour in the case matches the behaviour described, and the source material stays out of what the agent sees.
04Every outcome, not just the guilty
Most alerts are false positives, and clearing them well is most of the job. Each bench includes cases where the right call is to act, cases where it is to clear, and cases where the file is not yet enough to decide and the right call is to request information and say what is missing. Clean cases are built with the same care as suspicious ones, and carry features that look alarming until the profile explains them.
An agent that always escalates can’t score well, and neither can one that always clears. Both will be reported as baselines next to every model.
05Review
Before any case enters a bench, it will be reviewed. The reviewer will check that the intended decision is reachable from the records provided, that it doesn’t depend on a trick of wording, and that the difficulty sits where it was designed to sit. Cases that fail go back for revision or stay out of the sealed set.
06Sealed sets
Published test sets end up in training data, and scores on them stop meaning anything. Our bench cases are sealed and held out: never used for training, never used to tune our own tools or grading, and never published.
The brief, tools and limits are fixed for a bench version. Any change starts a new version, and scores are compared only within one.
07Grading
A decision that can’t be explained won’t survive QA or an exam, so the write-up is graded as closely as the decision, the way a QA reviewer or examiner would grade it.
- The decision. A wrong disposition earns nothing, however good the write-up.
- The evidence. Every finding has to trace to a record the agent could see. Citing a record that doesn’t exist scores the case zero.
- The analysis. Each case has a private checklist: the facts that matter, the red flags present and those that should be ruled out, the typology, and what is missing. Credit comes from what the write-up shows, not from its length.
- Conduct. Stating suspicion as established fact, or wording that would tip off the customer, is recorded as a critical error.
Where a criterion needs judgement, the grader will be a model from a different family than the model under test, checked against human grading before any score is published.
08Reporting
Results will be reported as the average across cases and runs, never the best run, with error bars. Alongside each score: decision accuracy against the baselines, critical errors, cost and the kinds of mistakes made. Scores will be published with the evidence each model cited.
Results are coming soon.
09Limits
The cases and policies are synthetic. They can’t carry every quirk of a real book of business, and a score measures an agent in our environment, not in a live deployment. Automated grading is not expert sign-off. Expert review of the cases and comparisons across models are separate steps, and we will say which have been done when results are published.
Building agents for risk and compliance? We’re running private pilots with AI labs.