METHODOLOGY

Same tasks. Double-checked.

Every model receives the same failure traps. The recorded reply is scored against fixed expectations. A mistake is counted only when two independent checks reproduce it.

Challenge

Failure traps are small coding tasks designed to expose wrong but plausible answers.

Answer

The model's recorded reply is the only answer scored. Later edits, retries, and commentary do not replace it.

Judge

Two independent checks must agree before a mistake appears as verified.

GenXis GavelTM evidence upgrade

  1. Preserve old receipts

    Published receipts remain historical evidence. New explanatory copy points to the same public roots instead of rewriting prior runs.

  2. Separate ranking from audit

    Certified leaderboard rank requires the complete 50-trap denominator. Wider 100-task receipts are audit bundles, not silent score changes.

  3. Hide private machinery

    Internal proposal systems can improve operations, but only public tasks, recorded answers, judge reasons, hashes, and signed receipts support customer-facing claims.

  1. Ranks

    Only complete runs get benchmark ranks. Ties share rank.

  2. Incomplete

    Missing or unresolved evidence is not counted as zero.

  3. Current basis

    Scores show the latest 50-trap pack until enough 100-task cycles exist for a rolling average.

Audit semantics

  1. Dual-check reproduction

    A verified mistake means the primary and shadow checks both reproduced the mismatch from the recorded answer against the frozen expected output. If both checks do not agree, the item stays unresolved and outside the ranked score.

  2. Certified denominator

    The Most Annoying rank uses the 50-trap view. The 100-task receipt covers the wider run bundle and is audit evidence, not the denominator used for this leaderboard.

  3. Signature and feed rank

    Receipt JSON must be independently verified by recomputing the canonical root and checking the signature against the published issuer public key. OpenRouter rank is separate discovery metadata and is kept out of the Most Annoying rank column.

Audit questions

  1. Can the exact responses reproduce the mismatches?

    Yes, once the full archive is public. The archive must include the frozen prompts, recorded replies, expected outputs, judge reasons, and receipts needed to rerun each mismatch claim.

  2. Do receipt roots and signatures independently verify?

    Retrieving a receipt is not verification. Auditors need the canonical payload, receipt root, signature, and issuer public key to verify every signature independently.

  3. Were the same rules applied to every ranked run?

    Ranked complete runs must use the identical 50 traps, scoring rules, model settings, and retry policy. If any of those differ, the run should not be treated as the same ranked pack.