For AI labs

RL environments to train frontier models & evaluate compliance agents.

Realistic financial crime work: onboarding customers, monitoring transactions, screening sanctions and preventing fraud.

Four environments

Built the way a financial crime team works its queue. An alert or onboarding file comes in; the agent gathers the evidence, applies policy and risk appetite, and decides whether to clear, escalate, request information or report. The rationale is judged as closely as the decision, the way a QA reviewer or examiner would.

How it works

The agent works a case the way an analyst would: gathering facts, weighing them against policy, and recording what it decided and why.

  1. Realistic synthetic customers, accounts and activity.
  2. The tools, records and policies an analyst works with.
  3. Graded on the decision and the evidence behind it.

A correct call with no evidence behind it earns no credit, and neither does a guess.

Benchmarks

FinCrime Bench tests agents against held-out cases across all four environments, scoring the decision together with the evidence that supports it. Those cases are sealed: none appear in training, so a strong score means the agent worked the case rather than memorised it. KYC Bench, focused on onboarding and due diligence decisions, is in development.

Read about the benchmarks

Research

Short notes on how the cases get built and what the benchmarks show.

Working on agents for risk and compliance? We’re running private pilots with AI labs.

Bring your own model; we supply the cases, the policies and the grading.