Placeholder data. Every figure on this site is illustrative sample data for build purposes. No benchmark run has been published yet.
Financial Conduct Bench

Methodology

This page is the contract. If the method here is wrong, the numbers are wrong, and we would rather argue about the method in public than defend a leaderboard nobody can inspect.

What is being measured

Whether an AI system talking to a UK consumer about money behaves the way the FCA requires a firm to behave. That is broader than the advice/guidance boundary, which is why the benchmark is organised as suites rather than a single score. Suite 1 covers the perimeter. Vulnerability, consumer understanding, financial promotions, regulated-fact accuracy and consistency across customers follow.

The FCA has not written a separate rulebook for AI; it has said the existing framework applies. So every test case traces to an existing obligation — a RAO article, a CONC or MCOB provision, a Principle. No case rests on our opinion of what good looks like.

Two numbers, never one

Each run reports a pass rate and a critical breach count, and the second is never folded into the first.

A critical breach is a hard rule crossing: a personal recommendation made without the permission to make one, a merits comparison of debt solutions, a steer toward a named mortgage. These are not degrees of quality. A system that crosses the perimeter once has crossed it, and ninety good answers either side do not net it off. Averaging them together would produce a number that looks reassuring and means nothing to the person who has to sign off the deployment.

The leaderboard ordering carries the same asymmetry: any system with a critical breach sorts below every clean system regardless of pass rate.

Why refusing everything does not win

The obvious way to score well on a perimeter test is to refuse to say anything. That is also a Consumer Duty failure, so the perimeter suite scores escalation and handover as its own dimension: did the system recognise the edge of its permission and route the customer somewhere useful, or did it just stop?

The scripted baseline in the leaderboard exists to make this visible. It rarely breaches. It also rarely helps, and the dimension breakdown shows exactly that.

How cases are graded

Each case is a prompt, an expected regulatory treatment, and the rule the treatment turns on. Grading is a rubric applied to the response, and every published run includes graded examples with the rationale attached, so a reader can check our judgement on the cases we found hardest rather than only the ones we found flattering.

Where a grade is genuinely contestable — and in perimeter work a meaningful minority are — we would rather publish the case and the disagreement than quietly drop it.

The conflict of interest

Doshi builds this benchmark and also builds one of the systems it scores. That is a real conflict and pretending otherwise would be the fastest way to make the whole thing worthless.

What we do about it:

  • Every result carries a disclosure label on the author's system.
  • The test set, the harness version and the raw run files are committed to the repository. The site is a renderer over those files and cannot show a number that is not in them.
  • Runs are versioned by harness, so a result can be reproduced against the harness that produced it.
  • We publish the cases our own system fails.

What that does not fix: we choose the cases. The only real answer to that is for the case set to be public and for other people to attack it, which is why it is.

Status

Version 0.1, draft. No run has been published yet — every figure currently on the site is illustrative placeholder data, labelled as such. The method is published first so it can be criticised before there is anything at stake in defending it.