Back to programming-model index

MODEL EVIDENCE / OpenRouter not captured

Anthropic: Claude Fable 5.1

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

anthropic/claude-fable-5.1
Official source pendingComplete result
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter not captured
Context
1000000
Input
text, image, file
Output
text
Reasoning flag
Yes
Prompt price
$10.000 / 1M tokens
Completion price
$50.000 / 1M tokens
Latest run
2026-09-02T00:26:05.052473+00:00
View denominator
50

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

MOST ANNOYING BENCHMARK

26/ 50

Complete result · 26 verified mistakes · 24 clean results · 0 unresolved

Inspect run receipt
Packs & trend

Pack 1: 6 verified mistakes, 4 clean results, 0 unresolved; Pack 2: 5 verified mistakes, 5 clean results, 0 unresolved; Pack 3: 7 verified mistakes, 3 clean results, 0 unresolved; Pack 4: 4 verified mistakes, 6 clean results, 0 unresolved; Pack 5: 4 verified mistakes, 6 clean results, 0 unresolved

2026-08-31T21:35:59.284247+00:00: not_run; 2026-09-01T22:50:04.330589+00:00: not_run; 2026-09-02T00:26:05.052473+00:00: 26

Score history

Model runs, then future slots.

  1. 01NOT RUN
  2. 02NOT RUN
  3. 0348%
  4. 04NOT RUN
  5. 05NOT RUN
  6. 06NOT RUN
  7. 07NOT RUN
  8. 08NOT RUN
  9. 09NOT RUN
  10. 10NOT RUN
  11. 11NOT RUN
  12. 12NOT RUN
  13. 13NOT RUN
  14. 14NOT RUN
  15. 15NOT RUN
  16. 16NOT RUN
  17. 17NOT RUN
  18. 18NOT RUN
  19. 19NOT RUN
  20. 20NOT RUN
  21. 21NOT RUN
  22. 22NOT RUN
  23. 23NOT RUN
  24. 24NOT RUN
Verified mistakesClean resultsUnresolved / pendingMissing / no recorded result

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. Alias TraceT001

    The model's exact answer did not match the frozen expected output.

    sha256:ff60a159a009afcbcda4611cca61dc4cc16a991d849b6cfce7ac556d0e6cd3c3
  2. Slice TraceT004

    The model's exact answer did not match the frozen expected output.

    sha256:4ddab3834c571ff929d71f587dea64bb1a3878fe529a3289817103b7c0e6a2ce
  3. Precedence TraceT006

    The model's exact answer did not match the frozen expected output.

    sha256:b961cc824002b589be2f8896381520e96a395de0e5fcde05b5ccd50e177db9e3
  4. Loop TraceT007

    The model's exact answer did not match the frozen expected output.

    sha256:08d281f1c718042151182428d1d62e7d2708713caeea4b290171c591e3d7c52e
  5. Slice TraceT008

    The model's exact answer did not match the frozen expected output.

    sha256:1473ab91ff9c39f015e773e38aad0446785012a34899e55abcbd6cf3b9af4203
  6. Precedence TraceT010

    The model's exact answer did not match the frozen expected output.

    sha256:c531cb0f4a2974dbd7947b7a54184b2143f88a66b7214f04922a1367d7d99996
  7. Precedence TraceT022

    The model's exact answer did not match the frozen expected output.

    sha256:b9c96c48feb6a0c05f06db524c3d18a039759ab9db62b8de7572232a904c33fb
  8. Slice TraceT024

    The model's exact answer did not match the frozen expected output.

    sha256:3d9f5a99cccb10477fdfbe66fe80bcbbe26a09c188641b694f2257f6e71cc423
  9. Alias TraceT025

    The model's exact answer did not match the frozen expected output.

    sha256:71ffe6d4688df01e291be11223fe1d936379c3080a4b6e1ff9e9ef9ff250fea3
  10. Slice TraceT028

    The model's exact answer did not match the frozen expected output.

    sha256:3ef512a8a15262ae2ae9fbc7e0b95d4278753bcf9c511be39cec6adfef6158e0
  11. Alias TraceT029

    The model's exact answer did not match the frozen expected output.

    sha256:bbe46f3572f023cbacdafa2579d4014d9ba3613dc6c8da6b3774db4444e79bed
  12. Alias TraceT041

    The model's exact answer did not match the frozen expected output.

    sha256:ae2984162b5d20105278f9abf48c4167af831cde2f2a29dc04ed6a5b9fbd5057
  13. Precedence TraceT042

    The model's exact answer did not match the frozen expected output.

    sha256:b4b2d68e0bbb323b977aa6fb659d8245dfa641f0e110aef032f73fc66585d540
  14. Slice TraceT044

    The model's exact answer did not match the frozen expected output.

    sha256:49854b2667770924e69acaa74732f7e4083bb8a8027b9cf98ecaa84a1ad4cb8e
  15. Alias TraceT045

    The model's exact answer did not match the frozen expected output.

    sha256:29090fd28c33b4c69be0ebf551590f33bc5df8db3aa29b63aa30ca335b80f655
  16. Precedence TraceT046

    The model's exact answer did not match the frozen expected output.

    sha256:3a7b080adeba9b28b55fa58d9341d4b1d782d0b0030d49ee4298a16d2b4f01f8
  17. Slice TraceT048

    The model's exact answer did not match the frozen expected output.

    sha256:372b89950a35212aa674e67e8b588334144db676f320b1ae28de98401fb29bd1
  18. Alias TraceT049

    The model's exact answer did not match the frozen expected output.

    sha256:d7e4a00f39181e06aca22b059ae39d08636a2092b421f463814ff9ae821625c2
  19. Slice TraceT064

    The model's exact answer did not match the frozen expected output.

    sha256:ae0a69809b865806e451692dde1b0890168f3d5ede3ee03afa1ee89beb7b55fa
  20. Precedence TraceT066

    The model's exact answer did not match the frozen expected output.

    sha256:24e49bdba3e6ef03f2a2e93ef8a10f6daf20ec58e336d1c234da8cdbcc6d39b9
  21. Slice TraceT068

    The model's exact answer did not match the frozen expected output.

    sha256:7649f5f29d6903dcaead3f12652b800d5cd244ca8249337772fc243c48e33c11
  22. Precedence TraceT070

    The model's exact answer did not match the frozen expected output.

    sha256:c48497357ad3ac678c7ee6861c0a0947a83cc63dbdb13ed955835ab99f84257a
  23. Slice TraceT084

    The model's exact answer did not match the frozen expected output.

    sha256:fbb674da202d0658192ee5e3ef0cd8a94920a9b0639b7e5bd0f2b55c92d80ba5
  24. Alias TraceT085

    The model's exact answer did not match the frozen expected output.

    sha256:2372a6eec1819dcda0f8dea88dca0b155e312e511ab6d57851bdf7cdb2895014
  25. Slice TraceT088

    The model's exact answer did not match the frozen expected output.

    sha256:cb0c4dc57023a4494ade1958f1d9f153cef3e4391198c58c2109921c57a3078d
  26. Alias TraceT089

    The model's exact answer did not match the frozen expected output.

    sha256:4e6f36a7460594f35f01078ea58d48bcd06626d52bb3ffc9832f3f46a565321c

OFFICIAL SOURCES

Provider record.

An exact primary official launch or model-card source has not yet been verified for this model ID.

STATUSOfficial source pendingVerified: PendingResearch method: official primary-web fallback
  • No exact primary launch or model card is certified yet. This page will not infer one from a similar name.

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:21279b83f5bfde51d84ee2a688158a2104ffd5ed132fcd35ea71355ff755332e
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the result with its evidence.

Anthropic: Claude Fable 5.1: 26 verified mistakes and 24 clean results across 50 sealed benchmark traps.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.