GENXIS GavelTM EVIDENCE Agents propose. GavelTM proves. Frozen tasks, recorded answers, dual checks, receipt roots. Expand for details.

ABOUT BENCHMARK

  1. 01

    Challenge

    Every model gets the same 50 failure traps: small coding tasks designed to expose wrong but plausible answers.

  2. 02

    Answer

    The model's recorded reply is the only thing we score.

  3. 03

    Judge

    A mistake counts only if two independent checks reproduce it.

  4. 04

    Transparency

    Scores and receipts go up now. Full-run JSON follows a week later.

Ranked claims require frozen tasks, recorded answers, deterministic checks, receipt roots, and public replay evidence. Private proposal machinery stays private; public claims stay evidence-bound.

Frozen pack
Same 50 traps for every ranked model
Recorded answer
No edits, retries, or commentary replacement
Dual check
Primary and shadow checks must agree
Receipt root
Hashes bind the released evidence archive

Certified rankings

Complete GavelTM-scored runs only

Annoying Mistakes Clean Results Unresolved Missing
Least annoyingMost annoying

Only complete GenXis GavelTM-certified models appear in the ranked chart. Incomplete or withheld evidence stays visible below without being counted as a failure.

LATEST RUN

Model results

0 models
Loading verified results

Nothing is ranked until its evidence is complete.

Ranked complete results

Shown by selected sort. Benchmark rank #1 means fewest verified mistakes; ties share rank.

    AUDITABILITY

    Unresolved runs.

    Incomplete runs

    Unresolved evidence. Not certified. Rows follow the selected sort; P-ranks still use resolved-only provisional quality.

      Not Yet Tested

      Tracked models with no benchmark attempt for this exact ID.

        Limits

        Only complete runs get a benchmark rank. Ties share it. OpenRouter's weekly order is a separate list. Missing or incomplete is not a zero. Scores are the latest 50-trap pack until enough 100-task cycles exist for a rolling average.

        AUDIT FILES

        Download the proof JSON.

        One week after each test completes, we publish one JSON file with the challenge questions, exact prompts, recorded model responses, judge reasons, and receipts. Use it to check the work yourself.

        Loading evidence archive.

          SHARE

          Use the result.