Benchmarks for risk and compliance agents.

Each benchmark gives an agent a caseload it has never seen and grades the work the way a QA reviewer would: the decision, and the evidence behind it.

FinCrime Bench

Four tracks, one for each environment. The cases are synthetic, sealed and held out: never used for training and never published. A correct call with no evidence behind it earns no credit, and neither does a guess. In some cases the right call is to request more information before deciding.

Results

Coming soon

Scores will be published with the evidence each model cited.

Notes

How the cases get built.

Building agents for risk and compliance? We’re running private pilots with AI labs.