Benchmarks for risk and compliance agents.
Each benchmark gives an agent a caseload it has never seen and grades the work the way a QA reviewer would: the decision, and the evidence behind it.
FinCrime Bench
Four tracks, one for each environment. The cases are synthetic, sealed and held out: never used for training and never published. A correct call with no evidence behind it earns no credit, and neither does a guess. In some cases the right call is to request more information before deciding.
-
KYC
Coming soon
CDD and EDD decisions: document requirements, verification of related parties and beneficial owners, screening and risk rating; approve, request information, escalate to EDD or decline.
-
Transaction Monitoring
Coming soon
Alert dispositions across fiat and crypto: close with a rationale, request information or escalate to SAR/STR.
-
Sanctions Screening
Coming soon
Screening-hit adjudication: true match vs false positive, ownership and control; release, block, reject or escalate.
-
Fraud Prevention
Coming soon
Flagged logins, payments and new accounts: account takeover, APP scams, mule accounts and first-party fraud; hold, release, contact the customer or exit.
Results
Coming soon
Scores will be published with the evidence each model cited.
Notes
How the cases get built.
-
From a public SAR narrative to a synthetic case
How published typology patterns become fully fictional cases: provenance kept separate from invention, evidence arriving in stages, and the review every case passes before it enters a benchmark.
Building agents for risk and compliance? We’re running private pilots with AI labs.