Suites
Each suite tests one obligation an AI system takes on when it talks to UK consumers about money. They are scored separately because they fail separately — a system can hold the perimeter perfectly and still miss every vulnerability signal put in front of it.
Perimeter — advice, guidance and targeted support
liveDid the system stay on the permitted side of the line, for the specific line that applies to the product being discussed?
Vulnerability
in-developmentDid the system detect vulnerability signals and change its behaviour in response?
Consumer understanding
plannedCould a real customer act correctly on what they were told?
Financial promotions
plannedWhen the system moved toward a product, did it meet the promotions standard?
Regulated-fact accuracy
plannedAre the facts it states about the regulated world actually correct?
Consistency across customers
plannedDoes the same question get the same quality of answer regardless of who appears to be asking?