Back to programming-model index

MODEL EVIDENCE / OpenRouter #34

Anthropic: Claude Haiku 4.5

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

anthropic/claude-haiku-4.5
Official source pendingComplete result
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter #34
Context
200000
Input
text, image, file
Output
text
Reasoning flag
Yes
Prompt price
$1.000 / 1M tokens
Completion price
$5.000 / 1M tokens
Latest run
2026-08-31T21:35:59.284247+00:00
View denominator
50

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

MOST ANNOYING BENCHMARK

45/ 50

Complete result · 45 verified mistakes · 5 clean results · 0 unresolved

Inspect run receipt
Packs & trend

Pack 1: 9 verified mistakes, 1 clean results, 0 unresolved; Pack 2: 10 verified mistakes, 0 clean results, 0 unresolved; Pack 3: 10 verified mistakes, 0 clean results, 0 unresolved; Pack 4: 8 verified mistakes, 2 clean results, 0 unresolved; Pack 5: 8 verified mistakes, 2 clean results, 0 unresolved

2026-08-28T05:52:59.444809+00:00: 44; 2026-08-31T21:35:59.284247+00:00: 45

Score history

Model runs, then future slots.

  1. 0112%
  2. 0210%
  3. 03NOT RUN
  4. 04NOT RUN
  5. 05NOT RUN
  6. 06NOT RUN
  7. 07NOT RUN
  8. 08NOT RUN
  9. 09NOT RUN
  10. 10NOT RUN
  11. 11NOT RUN
  12. 12NOT RUN
  13. 13NOT RUN
  14. 14NOT RUN
  15. 15NOT RUN
  16. 16NOT RUN
  17. 17NOT RUN
  18. 18NOT RUN
  19. 19NOT RUN
  20. 20NOT RUN
  21. 21NOT RUN
  22. 22NOT RUN
  23. 23NOT RUN
  24. 24NOT RUN
Verified mistakesClean resultsUnresolved / pendingMissing / no recorded result

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. Alias TraceT001

    The model's exact answer did not match the frozen expected output.

    sha256:54b69bd8a9f04fbc9ed3090243d488c61347d67717d7110b9318e91cac51b28f
  2. Precedence TraceT002

    The model's exact answer did not match the frozen expected output.

    sha256:f3e0dabc698678b6998e3967fcf1253f6a7aa6cefc8bd68fa4a7587b31468c80
  3. Slice TraceT004

    The model's exact answer did not match the frozen expected output.

    sha256:9084c761ed6ebab23fa676b3e78b2987b9bc47017c58f12eb6c93fbdeaf5da85
  4. Alias TraceT005

    The model's exact answer did not match the frozen expected output.

    sha256:89a0faea6f020d51dc11ffeda1986eb54ab8938f8c9f84e5f12a63d8b4eb589a
  5. Precedence TraceT006

    The model's exact answer did not match the frozen expected output.

    sha256:313c1bb67ac1d32bcbdcf02c53254083adc1baf2c3c798e608ac4f5e32bd6cec
  6. Loop TraceT007

    The model's exact answer did not match the frozen expected output.

    sha256:f8f48c599ed6f45def1184d4b18b5fc360a078ff699884d1e286c16691b673cb
  7. Slice TraceT008

    The model's exact answer did not match the frozen expected output.

    sha256:584a5f2f4dec93795acad60b05f1ed1d042db571eddb9d339628b18c40a02a9c
  8. Alias TraceT009

    The model's exact answer did not match the frozen expected output.

    sha256:4e5ceca503f62d35169970143fbcccad70d69348ebef6cb21a3fc66a42b3319e
  9. Precedence TraceT010

    The model's exact answer did not match the frozen expected output.

    sha256:5b48fcc8deb23aa6fc80fe57f6893c4b85a280384d2f07c1ff1f0ecd488e398d
  10. Alias TraceT021

    The model's exact answer did not match the frozen expected output.

    sha256:3d23cf778c62d5ebcf3519bc2a25f27f02ebb5bee8c450e7b29590a80a4edf6e
  11. Precedence TraceT022

    The model's exact answer did not match the frozen expected output.

    sha256:6da9dff510f6e58a0095d3d89277d04a987ba5c9dbb01847995dd6777a4d6d62
  12. Loop TraceT023

    The model's exact answer did not match the frozen expected output.

    sha256:f8f6a7ff040d4c8b749844bd01eb2f75dd4c2a5ee7e0f4e7d1dc0044a545ac5b
  13. Slice TraceT024

    The model's exact answer did not match the frozen expected output.

    sha256:1aba2b7e93a7cd406e7e075dac3e538b8e71bd7d82c0415d4b834f0bd2e15874
  14. Alias TraceT025

    The model's exact answer did not match the frozen expected output.

    sha256:2762bdffcff7f9ed6948fbe18f25824b91a1e6c8aefacfb35509185d1bbe9372
  15. Precedence TraceT026

    The model's exact answer did not match the frozen expected output.

    sha256:269440cbd60c48a19ec79b0837f36552614042371ba7f00f9899701ce12a83d8
  16. Loop TraceT027

    The model's exact answer did not match the frozen expected output.

    sha256:7d90b7e0ee8834a6b05cb9d3e199001b984b9f3b074a50358a883ad7862174af
  17. Slice TraceT028

    The model's exact answer did not match the frozen expected output.

    sha256:356ef83d7c12d52f2e41abbfb7e0e8c4c42faf1f9e2401863e17d360c3818fd9
  18. Alias TraceT029

    The model's exact answer did not match the frozen expected output.

    sha256:ce3e4b89e7d6920cffbd1e34d9ffd8b34340a77ff48d718eb4ae71b45d002dc9
  19. Precedence TraceT030

    The model's exact answer did not match the frozen expected output.

    sha256:416b5b4a2ee81b640d5254066cfd853e46a4dfe78ccdbb31cc1f584a22d396ca
  20. Alias TraceT041

    The model's exact answer did not match the frozen expected output.

    sha256:917937386299a07e1a3e611ca6a38b1e8e4c1f124803ffdb8d03bb3fe8de6635
  21. Precedence TraceT042

    The model's exact answer did not match the frozen expected output.

    sha256:c424c8db87d69aac20c30a17e3abde13bbbddfc041803e42618a0f113172178c
  22. Loop TraceT043

    The model's exact answer did not match the frozen expected output.

    sha256:2d28e94373324cce04588505c7f0628ca6d7705f9e90c2cd95164fbd48404637
  23. Slice TraceT044

    The model's exact answer did not match the frozen expected output.

    sha256:a88353d0ae45c50de19eb0fd5a11dbc8ee60bf98feaf5fc512f0f523b07aeaa3
  24. Alias TraceT045

    The model's exact answer did not match the frozen expected output.

    sha256:777273f0af85c2e5dca7eaf60255ad65159f22f4ad264c7c59303cb527bd4ae2
  25. Precedence TraceT046

    The model's exact answer did not match the frozen expected output.

    sha256:b43f2bc131102965bfa03f9aa467af94d1d2c1335524b49085e3c03386561aa6
  26. Loop TraceT047

    The model's exact answer did not match the frozen expected output.

    sha256:54ca958c7bcf261d8de1c5303e220841903ddd23506cbac5bf63a5dd294162e1
  27. Slice TraceT048

    The model's exact answer did not match the frozen expected output.

    sha256:aac9e9f9436b95934047659db4789e6b7eabb4335af195620fe77810ec82fcd1
  28. Alias TraceT049

    The model's exact answer did not match the frozen expected output.

    sha256:f429d0ab1dd3900528f8f73c0f3f252ff3e761ce8a26f6616020034a98f3bc0e
  29. Precedence TraceT050

    The model's exact answer did not match the frozen expected output.

    sha256:396fde104a0173fac722741f25e30ccf5f90e36be6d57e2068c4f08ffeb0d027
  30. Alias TraceT061

    The model's exact answer did not match the frozen expected output.

    sha256:f491a1bb05c7cb5979e7b99b2575f4cbab748a5907745567116869b8507fca35
  31. Precedence TraceT062

    The model's exact answer did not match the frozen expected output.

    sha256:e70d81366bb876b43604e7d33f64ce2a78ab643eae65cd8aea96ed12e5473aee
  32. Slice TraceT064

    The model's exact answer did not match the frozen expected output.

    sha256:d93a14975606b2e524bca25eff5a2be23b74266375a44570d520ad6c06ac425d
  33. Alias TraceT065

    The model's exact answer did not match the frozen expected output.

    sha256:6d82423fdda6441325e8155a0d004b9a6b3abbe4d72da5c40f573aae28afdff9
  34. Precedence TraceT066

    The model's exact answer did not match the frozen expected output.

    sha256:bfd908506c9993de059eadbfb49f06e16ea56ef538a5a980a03246ce1b28b6ec
  35. Slice TraceT068

    The model's exact answer did not match the frozen expected output.

    sha256:838e07a21d3c2776649be0eddaf1fc02dfdde79f290a09d66d007e20a30fd0e6
  36. Alias TraceT069

    The model's exact answer did not match the frozen expected output.

    sha256:20dc0b5ade7de744d5b6a6f5744a419321efd6e3d18fd5837973d53ad2e337c1
  37. Precedence TraceT070

    The model's exact answer did not match the frozen expected output.

    sha256:de7466d96a15b5222442b1f538b7f142ff39f99505a7de4e1d871cb5144a562a
  38. Alias TraceT081

    The model's exact answer did not match the frozen expected output.

    sha256:e16519e812e35df52766db9057ab8ddc38ba41a1eedd307549aa48517eea263b
  39. Precedence TraceT082

    The model's exact answer did not match the frozen expected output.

    sha256:a598f4020dbfaee0e258f433cae449d21bd7b8577d5b1d2a3ec16f6fbebbc333
  40. Loop TraceT083

    The model's exact answer did not match the frozen expected output.

    sha256:6a4fb5dbe73803e6a67f0a54ef0a1299668fc73f81655b887aff21aaa705699f
  41. Slice TraceT084

    The model's exact answer did not match the frozen expected output.

    sha256:7b0db1f2b0fb5ea8637952f989d0489cbf42cc44ea3ea077cc5f2f1e5c1032ad
  42. Alias TraceT085

    The model's exact answer did not match the frozen expected output.

    sha256:c24d228560f9028a7ec4b52492103a908d356473287d916bad68f30c2df78f68
  43. Precedence TraceT086

    The model's exact answer did not match the frozen expected output.

    sha256:ee2d16cf28b373d8b45f45d0b547e7995434b901fff4e1d7ebe1a8e4d26d4352
  44. Slice TraceT088

    The model's exact answer did not match the frozen expected output.

    sha256:4c30f34a3f695be2a5e66ee8a86dc1d7cb8a8a735af0ebaf44d4392b74087b08
  45. Alias TraceT089

    The model's exact answer did not match the frozen expected output.

    sha256:0ac9078fbdfdae11453ae42d848107d8c141a2e066dd2912f47ca40dc3e0cae5

OFFICIAL SOURCES

Provider record.

An exact primary official launch or model-card source has not yet been verified for this captured model ID.

STATUSOfficial source pendingVerified: PendingResearch method: official primary-web fallback
  • No exact primary launch or model card is certified yet. This page will not infer one from a similar name.

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe only the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:1d7e067dc7e573537b182c078285a0c8a4502904f19d269e509b778359084b64
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the exact page, not a screenshot without context.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.