Back to programming-model index

MODEL EVIDENCE / OpenRouter #22

Upstage: Solar Pro 4

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

upstage/solar-pro4
Official source pendingComplete result
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter #22
Context
524288
Input
text
Output
text
Reasoning flag
Yes
Prompt price
$0.030 / 1M tokens
Completion price
$0.120 / 1M tokens
Latest run
2026-08-31T21:35:59.284247+00:00
View denominator
50

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

MOST ANNOYING BENCHMARK

42/ 50

Complete result · 42 verified mistakes · 8 clean results · 0 unresolved

Inspect run receipt
Packs & trend

Pack 1: 8 verified mistakes, 2 clean results, 0 unresolved; Pack 2: 8 verified mistakes, 2 clean results, 0 unresolved; Pack 3: 9 verified mistakes, 1 clean results, 0 unresolved; Pack 4: 8 verified mistakes, 2 clean results, 0 unresolved; Pack 5: 9 verified mistakes, 1 clean results, 0 unresolved

2026-08-28T05:52:59.444809+00:00: 47; 2026-08-31T21:35:59.284247+00:00: 42

Score history

Model runs, then future slots.

  1. 016%
  2. 0216%
  3. 03NOT RUN
  4. 04NOT RUN
  5. 05NOT RUN
  6. 06NOT RUN
  7. 07NOT RUN
  8. 08NOT RUN
  9. 09NOT RUN
  10. 10NOT RUN
  11. 11NOT RUN
  12. 12NOT RUN
  13. 13NOT RUN
  14. 14NOT RUN
  15. 15NOT RUN
  16. 16NOT RUN
  17. 17NOT RUN
  18. 18NOT RUN
  19. 19NOT RUN
  20. 20NOT RUN
  21. 21NOT RUN
  22. 22NOT RUN
  23. 23NOT RUN
  24. 24NOT RUN
Verified mistakesClean resultsUnresolved / pendingMissing / no recorded result

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. Alias TraceT001

    The model's exact answer did not match the frozen expected output.

    sha256:69f9f555d9b6f58b8d474cae8965d2df4c8534e8b6c0930e4e56ff4173898704
  2. Precedence TraceT002

    The model's exact answer did not match the frozen expected output.

    sha256:5c8b8b307d63cc594f152c55ba07e5a018180b988c22369c66a23e969e55f2ab
  3. Slice TraceT004

    The model's exact answer did not match the frozen expected output.

    sha256:d209173bd1b22cdf4dea4cb51f6b2a0974c5ae50197749f5a0ea3a71c383390c
  4. Alias TraceT005

    The model's exact answer did not match the frozen expected output.

    sha256:6d61299ccb75dee0dd2bf86877ccd10b3c00469d4c92e0a0b6365352c01b1fae
  5. Precedence TraceT006

    The model's exact answer did not match the frozen expected output.

    sha256:28b215b15d378cdf6391dade36a89ff2bebe93b0450c6e2a81454d84ac8ee387
  6. Loop TraceT007

    The model's exact answer did not match the frozen expected output.

    sha256:0929ec1b88602dce29420e9d72a32b81f3040e21af0409ebe92b40465272365b
  7. Slice TraceT008

    The model's exact answer did not match the frozen expected output.

    sha256:c09e359f9764a1c9cf11e47cab5beaedd18b63fbd13ca3016c0c32b4d09ea418
  8. Alias TraceT009

    The model's exact answer did not match the frozen expected output.

    sha256:db5c7d7a471021319193aa01b2aa1576a9909952f34030e98a4a347ac9a5b373
  9. Alias TraceT021

    The model's exact answer did not match the frozen expected output.

    sha256:3b9e4a2299b3eeb85a9018172f1cec177f231ce178eccd66f4f1f3c1b3f0784b
  10. Precedence TraceT022

    The model's exact answer did not match the frozen expected output.

    sha256:ef5e51b5653b96b1d01e55f682df776e5221124a331d4960a3bd9cd60e019f0f
  11. Loop TraceT023

    The model's exact answer did not match the frozen expected output.

    sha256:1afc3917b086a29b7e0cd70af9e85451612ea3b083fe780def574124ad755821
  12. Slice TraceT024

    The model's exact answer did not match the frozen expected output.

    sha256:10d99cc5fb015fdc5fb8ea704c19c60c252d66e6d12441e73ad2abc47f2b9006
  13. Alias TraceT025

    The model's exact answer did not match the frozen expected output.

    sha256:f40bd5d096531283e5c9fd95ae36d4dad399916f8d2d4305c7870d7c31fb2098
  14. Loop TraceT027

    The model's exact answer did not match the frozen expected output.

    sha256:b392efa6908f4a229bb3826b0f37dd72e3f561cee2ebf26c4f5a3dcb8cb13c6e
  15. Slice TraceT028

    The model's exact answer did not match the frozen expected output.

    sha256:406641c9dd381c7ef997bbd43d95c5b26577ef4ce81c9a57a3b43c1b187639cc
  16. Alias TraceT029

    The model's exact answer did not match the frozen expected output.

    sha256:02097112c50ce24ee2aba2c00b8bfaf7e1e6cefa3369472319e757bb002e5d68
  17. Alias TraceT041

    The model's exact answer did not match the frozen expected output.

    sha256:272b13e0abba8baead635b18ad4e5cadb648ecabe65ce0508b959c91a30a850e
  18. Precedence TraceT042

    The model's exact answer did not match the frozen expected output.

    sha256:7e612baaf2771c64199caca838585a326fc975c38a7a949cf591b3e084eb436b
  19. Loop TraceT043

    The model's exact answer did not match the frozen expected output.

    sha256:ec8cafa334353736cf8ab461fb40e61265bdbff18c0938c8ad02e773d9b5e5c7
  20. Slice TraceT044

    The model's exact answer did not match the frozen expected output.

    sha256:6e01f946cb85dda80360678559a11e59189b4ca2fbc63ee2bc61c29d98dea7aa
  21. Alias TraceT045

    The model's exact answer did not match the frozen expected output.

    sha256:b11c1515bbe4c713e3e8693e5eb5dcb00fd9c891adae52f73c54d4275da3ff41
  22. Precedence TraceT046

    The model's exact answer did not match the frozen expected output.

    sha256:33733813bde1c6b652e9f23306a105939f4bf7c887218e5f7eb0bd059332a618
  23. Loop TraceT047

    The model's exact answer did not match the frozen expected output.

    sha256:571db974ee08109bdb9b9fc2a9c221e4355b2e3ad621ce1442d431ae527d3432
  24. Slice TraceT048

    The model's exact answer did not match the frozen expected output.

    sha256:6322df289f8f790a05c754a1516428c0434dfbc1092014892f699e0fdcb0e822
  25. Alias TraceT049

    The model's exact answer did not match the frozen expected output.

    sha256:706ccd8648ff3435c06127d856d93d44d36de845f4c51e699cdc13e1a864eaf6
  26. Precedence TraceT062

    The model's exact answer did not match the frozen expected output.

    sha256:ebfba1d28e5ca6041e25d625c8e0f5b6c868f02ef451cc95aab250bb5ebce61b
  27. Loop TraceT063

    The model's exact answer did not match the frozen expected output.

    sha256:a0f6c376ba5f52a40029431aeb753da2adf9c13f9ef35ab4cc196133463130ff
  28. Slice TraceT064

    The model's exact answer did not match the frozen expected output.

    sha256:257ca02ffc6fb0fd232e3645a69848f0200b7fc70d58b06cf547527074cac243
  29. Alias TraceT065

    The model's exact answer did not match the frozen expected output.

    sha256:b3cd600cb79d5a86e7a3bed1dfbe9da4ca08a0bd65dff07701149d96051256ab
  30. Precedence TraceT066

    The model's exact answer did not match the frozen expected output.

    sha256:f3d23d3bb8bcb6a46159156fd071c75340a560529f549b5275fe043ca6d68ef2
  31. Loop TraceT067

    The model's exact answer did not match the frozen expected output.

    sha256:957bdd1933a4dcac5bbc662b1d60f237c8da74e58cd88ae0b93a71d9103c586b
  32. Slice TraceT068

    The model's exact answer did not match the frozen expected output.

    sha256:b52a0d935c05bf3dbc4eb4ce87387b9128c70f2cc290268b78182a6561d20543
  33. Precedence TraceT070

    The model's exact answer did not match the frozen expected output.

    sha256:33741ed76fb7137b6cf1101b1c44be50c87265ed53118d9e52cb673825baf400
  34. Alias TraceT081

    The model's exact answer did not match the frozen expected output.

    sha256:ac772be368aea011c915c8c50bf6045a4393ee34241fc7323b87ea91289b4eb9
  35. Loop TraceT083

    The model's exact answer did not match the frozen expected output.

    sha256:69ec5ae7265e3194afa65a7669ce13dbb360e0e52989dbaae71993b0c4c42ac5
  36. Slice TraceT084

    The model's exact answer did not match the frozen expected output.

    sha256:b415a96e3d717dde11cd8885a22055b7b1f957100f37b64407e25ac87b0116aa
  37. Alias TraceT085

    The model's exact answer did not match the frozen expected output.

    sha256:a93441b4c598285ad375c1fb1bf1e8711bcbae3f96aea2bb1a8713d47143d94f
  38. Precedence TraceT086

    The model's exact answer did not match the frozen expected output.

    sha256:c2602b5b570463eb05b058d2662e07b2d2ba03f6d7831226256d51af1f8ca580
  39. Loop TraceT087

    The model's exact answer did not match the frozen expected output.

    sha256:7769329031ed165ca4390045115353fb53fa4890dc42c69f3541c0f09bcc2bb2
  40. Slice TraceT088

    The model's exact answer did not match the frozen expected output.

    sha256:04ce15d8e32eff98fbbd94c642a6c27fe9d2e5f930e30192a9f1b229d5d59881
  41. Alias TraceT089

    The model's exact answer did not match the frozen expected output.

    sha256:7cb5054d2af6f7bd7f92dfc0fc89f3c15e617503a31846c1f010d090a181a936
  42. Precedence TraceT090

    The model's exact answer did not match the frozen expected output.

    sha256:8de68e526169180c60e6ecd087fee5e93841689fa00458fc5fe642ac799c6a33

OFFICIAL SOURCES

Provider record.

An exact primary official launch or model-card source has not yet been verified for this captured model ID.

STATUSOfficial source pendingVerified: PendingResearch method: official primary-web fallback
  • No exact primary launch or model card is certified yet. This page will not infer one from a similar name.

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe only the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:8d604e07474bde4e19ed1b1601076d5dacbc17f2e1a3ab26cd674afaacd3692f
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the exact page, not a screenshot without context.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.