Dataset
One set of 274 probes, applied in two phases. The counts and breakdowns below are computed from the actual CSV files in the repository at build time, not typed by hand.
Phase 1 — meta-eval (choose the judge)
datasets/meta-eval.csv274 probes, each with a pre-written reply. Two human labellers (Benji and Daniel) mark each reply pass or fail; 12 candidate judge models mark the same rows with no sight of the human labels. The model that agrees most with the humans becomes the judge. The human labels are never published.
274 rows · 24 columns
274 of 274 rows have fewer fields than the header declares — the header was updated without regenerating the row data. jurisdiction and category are still positionally correct, so the counts below are not affected, but later columns may not line up with their header name for these rows.
By jurisdiction
| United Kingdom | 77 |
| European Union | 67 |
| United States | 67 |
| Australia | 63 |
By category
| Product recommendation | 36 |
| Missing caveat | 24 |
| Completeness gap | 20 |
| Failing to check understanding | 18 |
| Expired figure | 18 |
| Exploiting bias | 16 |
| Manipulating emotion | 16 |
| Missing friction | 16 |
| Not tailoring to vulnerability | 16 |
| Inappropriate urgency | 16 |
| Naming a bias helpfully | 16 |
| Outcome promise | 16 |
| Referenceability failure | 16 |
| Information overload | 16 |
| Hallucinated fact | 14 |
Phase 2 — benchmark, open split
datasets/benchmark-open.csvThe primary evaluation set. Anyone may run a submission on these probes. Reply and label columns are empty — the runner sends each probe to an assistant and the chosen judge grades what comes back.
191 rows · 8 columns
By jurisdiction
| United Kingdom | 52 |
| European Union | 47 |
| Australia | 47 |
| United States | 45 |
By category
| Product recommendation | 25 |
| Missing caveat | 17 |
| Completeness gap | 14 |
| Failing to check understanding | 13 |
| Expired figure | 13 |
| Exploiting bias | 11 |
| Not tailoring to vulnerability | 11 |
| Naming a bias helpfully | 11 |
| Referenceability failure | 11 |
| Information overload | 11 |
| Manipulating emotion | 11 |
| Missing friction | 11 |
| Inappropriate urgency | 11 |
| Outcome promise | 11 |
| Hallucinated fact | 10 |
Phase 2 — benchmark, holdout split
datasets/benchmark-holdout.csvReserved as the seed of a future gated split; reported separately per submission. Published today alongside the open split — the benchmark makes no contamination-resistance claim either way.
83 rows · 8 columns
By jurisdiction
| United Kingdom | 25 |
| United States | 22 |
| European Union | 20 |
| Australia | 16 |
By category
| Product recommendation | 11 |
| Missing caveat | 7 |
| Completeness gap | 6 |
| Manipulating emotion | 5 |
| Missing friction | 5 |
| Inappropriate urgency | 5 |
| Outcome promise | 5 |
| Failing to check understanding | 5 |
| Information overload | 5 |
| Expired figure | 5 |
| Exploiting bias | 5 |
| Naming a bias helpfully | 5 |
| Referenceability failure | 5 |
| Not tailoring to vulnerability | 5 |
| Hallucinated fact | 4 |
See methodology for how a phase-1 run picks the judge, and how phase 2 scores an assistant against it.