GENXIS GavelTM EVIDENCE Agents propose. GavelTM proves. Frozen tasks, recorded answers, dual checks, receipt roots. Expand for details.
ABOUT BENCHMARK
- 01
Challenge
Every model gets the same 50 failure traps: small coding tasks designed to expose wrong but plausible answers.
- 02
Answer
The model's recorded reply is the only thing we score.
- 03
Judge
A mistake counts only if two independent checks reproduce it.
- 04
Transparency
Scores and receipts go up now. Full-run JSON follows a week later.
Ranked claims require frozen tasks, recorded answers, deterministic checks, receipt roots, and public replay evidence. Private proposal machinery stays private; public claims stay evidence-bound.
- Frozen pack
- Same 50 traps for every ranked model
- Recorded answer
- No edits, retries, or commentary replacement
- Dual check
- Primary and shadow checks must agree
- Receipt root
- Hashes bind the released evidence archive
Certified rankings
Complete GavelTM-scored runs only
Only complete GenXis GavelTM-certified models appear in the ranked chart. Incomplete or withheld evidence stays visible below without being counted as a failure.
LATEST RUN
Model results
Nothing is ranked until its evidence is complete.
- Tracked
- 0
- Complete results
- 0
- Incomplete
- 0
- Not tested
- 0
Ranked complete results
Shown by selected sort. Benchmark rank #1 means fewest verified mistakes; ties share rank.
AUDITABILITY
Unresolved runs.
Incomplete runs
Unresolved evidence. Not certified. Rows follow the selected sort; P-ranks still use resolved-only provisional quality.
Showing incomplete runs.
Not Yet Tested
Tracked models with no benchmark attempt for this exact ID.
Showing not yet tested models.
Limits
Only complete runs get a benchmark rank. Ties share it. OpenRouter's weekly order is a separate list. Missing or incomplete is not a zero. Scores are the latest 50-trap pack until enough 100-task cycles exist for a rolling average.
AUDIT FILES
Download the proof JSON.
One week after each test completes, we publish one JSON file with the challenge questions, exact prompts, recorded model responses, judge reasons, and receipts. Use it to check the work yourself.
SHARE