Back to programming-model index

MODEL EVIDENCE / OpenRouter #19

Anthropic: Claude Sonnet 4.6

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

anthropic/claude-sonnet-4.6
Official source pendingIncomplete run
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter #19
Context
1000000
Input
text, image, file
Output
text
Reasoning flag
Yes
Prompt price
$3.000 / 1M tokens
Completion price
$15.000 / 1M tokens
Latest run
2026-08-31T21:35:59.284247+00:00
View denominator
50

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

MOST ANNOYING BENCHMARK

25/ 50

Complete result · 25 verified mistakes · 25 clean results · 0 unresolved

Inspect run receipt
Packs & trend

Pack 1: 5 verified mistakes, 5 clean results, 0 unresolved; Pack 2: 6 verified mistakes, 4 clean results, 0 unresolved; Pack 3: 6 verified mistakes, 4 clean results, 0 unresolved; Pack 4: 4 verified mistakes, 6 clean results, 0 unresolved; Pack 5: 4 verified mistakes, 6 clean results, 0 unresolved

2026-08-28T05:52:59.444809+00:00: 25; 2026-08-31T21:35:59.284247+00:00: 25

Score history

Model runs, then future slots.

  1. 0150%
  2. 0250%
  3. 03NOT RUN
  4. 04NOT RUN
  5. 05NOT RUN
  6. 06NOT RUN
  7. 07NOT RUN
  8. 08NOT RUN
  9. 09NOT RUN
  10. 10NOT RUN
  11. 11NOT RUN
  12. 12NOT RUN
  13. 13NOT RUN
  14. 14NOT RUN
  15. 15NOT RUN
  16. 16NOT RUN
  17. 17NOT RUN
  18. 18NOT RUN
  19. 19NOT RUN
  20. 20NOT RUN
  21. 21NOT RUN
  22. 22NOT RUN
  23. 23NOT RUN
  24. 24NOT RUN
Verified mistakesClean resultsUnresolved / pendingMissing / no recorded result

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. Alias TraceT001

    The model's exact answer did not match the frozen expected output.

    sha256:1889cd17271f4f6a9401638d26ee821f94e1d8e28c80dd3e32b161f69921c0ec
  2. Slice TraceT004

    The model's exact answer did not match the frozen expected output.

    sha256:32496114399e2e1fcbd666db971e48dc41190a9c285079f17d37a70e29d0e57b
  3. Precedence TraceT006

    The model's exact answer did not match the frozen expected output.

    sha256:e86e514fccf51e34e9425caecda0cb4ccabf0f71ea5f7cb301cbba0f3031ed3c
  4. Slice TraceT008

    The model's exact answer did not match the frozen expected output.

    sha256:ee1662676af35dd9e61c847faac11f40cca82e4c9110407e37880741529b2947
  5. Precedence TraceT010

    The model's exact answer did not match the frozen expected output.

    sha256:1a1271c835aedc3ffbe297afddc85f489b8ea0c42b117ff8f44aee75dc92d626
  6. Alias TraceT021

    The model's exact answer did not match the frozen expected output.

    sha256:901ff33e4ec6b34398badb29ba7bab913e3c1870fa4e080e4064ef2a892a254d
  7. Precedence TraceT022

    The model's exact answer did not match the frozen expected output.

    sha256:a77cc3dea816f88709d99cb2d453f2006240118be14e1d295ed5ec7d2c54e009
  8. Loop TraceT023

    The model's exact answer did not match the frozen expected output.

    sha256:3521440644268b2afe31cd7eb2d3f8101bf05880254fa0f58484e2387605bd71
  9. Slice TraceT024

    The model's exact answer did not match the frozen expected output.

    sha256:c221d3dbc81994b28daf65b104ee7a8df99279d4ed1b5cc0af25578a18d4a8ab
  10. Precedence TraceT026

    The model's exact answer did not match the frozen expected output.

    sha256:a7545950812771f5c7b2f27775e38d6ac072b062060550b45d2a300a6f5ef3a3
  11. Precedence TraceT030

    The model's exact answer did not match the frozen expected output.

    sha256:2f582ccac9fcbbed6c4220957c7ee75142ba921817d482978ee093203064b037
  12. Alias TraceT041

    The model's exact answer did not match the frozen expected output.

    sha256:9b6e015bdfadda69b343710a4f270bb64b5703e1ee175eda8e84215a6321ab69
  13. Precedence TraceT042

    The model's exact answer did not match the frozen expected output.

    sha256:618132914be4bf54ce76f1c8d26ac29ec0a574fd203618b07a9da999a2420927
  14. Loop TraceT043

    The model's exact answer did not match the frozen expected output.

    sha256:823c6c3d5cdeb512f8129bbdefdd613a653bf6c975db5efe71c18819279d229f
  15. Alias TraceT045

    The model's exact answer did not match the frozen expected output.

    sha256:53f745f6aafcf5d140b7348e1824845d7321ff306024b0f77a50135f50bf9ea4
  16. Slice TraceT048

    The model's exact answer did not match the frozen expected output.

    sha256:70141726dbe842b8710bf1f05a998ab2276f363b41de60a7a447a218990f55a0
  17. Precedence TraceT050

    The model's exact answer did not match the frozen expected output.

    sha256:c9d67b9dddabfd0944282ae1c0eb5a4b4768ecfa851fecbd3e50798e20db609b
  18. Alias TraceT061

    The model's exact answer did not match the frozen expected output.

    sha256:41a599f6a7a0f316733f919851f76290c44848607db16c707d5e4e739655443d
  19. Precedence TraceT062

    The model's exact answer did not match the frozen expected output.

    sha256:446b1a4156dad428d43482a8df00f14c92bf67147e553f4efdb13768959d66a7
  20. Slice TraceT064

    The model's exact answer did not match the frozen expected output.

    sha256:02d54e096f4bccc1ec016d8adc5399bad0da988b913868287499dcbd43c202bf
  21. Precedence TraceT070

    The model's exact answer did not match the frozen expected output.

    sha256:222d37bcd0d21819d86f8331daa2f794ba068118494176afc49a70ea5e1f5953
  22. Slice TraceT084

    The model's exact answer did not match the frozen expected output.

    sha256:9d22ef6d9f67d3c53d20e4aa50409d2d657e7fc557299576ebb1f530bdcc72d7
  23. Precedence TraceT086

    The model's exact answer did not match the frozen expected output.

    sha256:8acdc7fa2f1bfae757529ea28c0daef3bd7e6cbfbe65175972ad96137eef3cc9
  24. Slice TraceT088

    The model's exact answer did not match the frozen expected output.

    sha256:e5f02433317ff385164c81054c14f4a35cbf0cea127f061f0d42ea798873fe21
  25. Precedence TraceT090

    The model's exact answer did not match the frozen expected output.

    sha256:b3e896ac38459549c0f3397f1abc27502f2fa29e8e0b98c448819e275ef5d08e

OFFICIAL SOURCES

Provider record.

An exact primary official launch or model-card source has not yet been verified for this captured model ID.

STATUSOfficial source pendingVerified: PendingResearch method: official primary-web fallback
  • No exact primary launch or model card is certified yet. This page will not infer one from a similar name.

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe only the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:6da3e3b5a85e5f626eae731ee676cf4adab9c8f8af13862968055c7dca9c3a92
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the exact page, not a screenshot without context.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.