Challenge
Failure traps are small coding tasks designed to expose wrong but plausible answers.
METHODOLOGY
Every model receives the same failure traps. The recorded reply is scored against fixed expectations. A mistake is counted only when two independent checks reproduce it.
Failure traps are small coding tasks designed to expose wrong but plausible answers.
The model's recorded reply is the only answer scored. Later edits, retries, and commentary do not replace it.
Two independent checks must agree before a mistake appears as verified.
Published receipts remain historical evidence. New explanatory copy points to the same public roots instead of rewriting prior runs.
Certified leaderboard rank requires the complete 50-trap denominator. Wider 100-task receipts are audit bundles, not silent score changes.
Internal proposal systems can improve operations, but only public tasks, recorded answers, judge reasons, hashes, and signed receipts support customer-facing claims.
Only complete runs get benchmark ranks. Ties share rank.
Missing or unresolved evidence is not counted as zero.
Scores show the latest 50-trap pack until enough 100-task cycles exist for a rolling average.
A verified mistake means the primary and shadow checks both reproduced the mismatch from the recorded answer against the frozen expected output. If both checks do not agree, the item stays unresolved and outside the ranked score.
The Most Annoying rank uses the 50-trap view. The 100-task receipt covers the wider run bundle and is audit evidence, not the denominator used for this leaderboard.
Receipt JSON must be independently verified by recomputing the canonical root and checking the signature against the published issuer public key. OpenRouter rank is separate discovery metadata and is kept out of the Most Annoying rank column.
Yes, once the full archive is public. The archive must include the frozen prompts, recorded replies, expected outputs, judge reasons, and receipts needed to rerun each mismatch claim.
Retrieving a receipt is not verification. Auditors need the canonical payload, receipt root, signature, and issuer public key to verify every signature independently.
Ranked complete runs must use the identical 50 traps, scoring rules, model settings, and retry policy. If any of those differ, the run should not be treated as the same ranked pack.