Back to programming-model index

MODEL EVIDENCE / OpenRouter #4

DeepSeek: DeepSeek V4 Flash 0423

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

deepseek/deepseek-v4-flash
Official source verifiedIncomplete run
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter #4
Context
1048576
Input
text
Output
text
Reasoning flag
Yes
Prompt price
$0.068 / 1M tokens
Completion price
$0.168 / 1M tokens
Latest run
2026-08-31T21:35:59.284247+00:00
View denominator
50

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

MOST ANNOYING BENCHMARK

/ 50

Incomplete run · 32 verified mistakes · 6 clean results · 12 unresolved

Inspect run receipt
Packs & trend

Pack 1: 7 verified mistakes, 1 clean results, 2 unresolved; Pack 2: 4 verified mistakes, 1 clean results, 5 unresolved; Pack 3: 6 verified mistakes, 1 clean results, 3 unresolved; Pack 4: 6 verified mistakes, 2 clean results, 2 unresolved; Pack 5: 9 verified mistakes, 1 clean results, 0 unresolved

Trend begins after a complete cycle

Score history

Model runs, then future slots.

  1. 0115%
  2. 02NOT RUN
  3. 03NOT RUN
  4. 04NOT RUN
  5. 05NOT RUN
  6. 06NOT RUN
  7. 07NOT RUN
  8. 08NOT RUN
  9. 09NOT RUN
  10. 10NOT RUN
  11. 11NOT RUN
  12. 12NOT RUN
  13. 13NOT RUN
  14. 14NOT RUN
  15. 15NOT RUN
  16. 16NOT RUN
  17. 17NOT RUN
  18. 18NOT RUN
  19. 19NOT RUN
  20. 20NOT RUN
  21. 21NOT RUN
  22. 22NOT RUN
  23. 23NOT RUN
  24. 24NOT RUN
Verified mistakesClean resultsUnresolved / pendingMissing / no recorded result

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. Alias TraceT001

    The model's exact answer did not match the frozen expected output.

    sha256:d7f7cfdcd6cb63bd14a51b93d2ebb0653cf37bb173017ec1d8c996e754c3c88a
  2. Precedence TraceT002

    The model's exact answer did not match the frozen expected output.

    sha256:03eedb9af9538e63a2c08233e538d92619007e04c45ad03550d38a2655100adf
  3. Slice TraceT004

    The model's exact answer did not match the frozen expected output.

    sha256:630982af10930b5ffe6adffc178467fcce99581d0b59169fbe6c48aad1ed44f9
  4. Alias TraceT005

    The model's exact answer did not match the frozen expected output.

    sha256:09d0dd50f6aec15272542230a3415d65eeded8ceda636fe5f59b12e748caae26
  5. Precedence TraceT006

    The model's exact answer did not match the frozen expected output.

    sha256:14b3f2cc6947d6fe24173f4170fed381be95e212e65335b6fe0ee32361a1d98a
  6. Loop TraceT007

    The model's exact answer did not match the frozen expected output.

    sha256:c1ca26408dee5a64cc47066d9ef0dac82203ed6c0bf656b2b79b1aadade0cbf3
  7. Precedence TraceT010

    The model's exact answer did not match the frozen expected output.

    sha256:1fc4d9b2ef4e2ad98a27c911f775ff9b05de1907fdf910d105317c0bc29b7abb
  8. Alias TraceT025

    The model's exact answer did not match the frozen expected output.

    sha256:424c1e992496fba88debe7cad2e4f4d4c54f503afb5dd3af3553f1881bfea75a
  9. Precedence TraceT026

    The model's exact answer did not match the frozen expected output.

    sha256:c918c334681787ba12dab58ef3070d8309d780008b095290aa070cdb018aebb4
  10. Slice TraceT028

    The model's exact answer did not match the frozen expected output.

    sha256:dbd3e79b917207d07932d596bee0312745202a52a5ed034b6ff87d6a50688393
  11. Precedence TraceT030

    The model's exact answer did not match the frozen expected output.

    sha256:21481851d39d285912aef54fbbfa1faeff6f1e513dd877470fef20a178186f10
  12. Alias TraceT041

    The model's exact answer did not match the frozen expected output.

    sha256:e90fd055892e25b8915af689cc0a87b82dd1edb488bc866455e8fbd46f2a1ced
  13. Precedence TraceT042

    The model's exact answer did not match the frozen expected output.

    sha256:41d55c1e0be27f928ae1f11d9b908cb03cd1bb7073c2fc93ee1b270986d8d2b1
  14. Slice TraceT044

    The model's exact answer did not match the frozen expected output.

    sha256:888d411080437a315979f3f257a7dcd778ed8ebb1c38eb58d504e35f0ad61e47
  15. Alias TraceT045

    The model's exact answer did not match the frozen expected output.

    sha256:3314be28f204fdbf7526dba3da8ca90f866e4df5f62eeb2b976f575921f86bda
  16. Slice TraceT048

    The model's exact answer did not match the frozen expected output.

    sha256:8deb6e870051ea7ba7d8d7e75a7630bc943d3399cdbded5b3a0f3e88a96e092e
  17. Precedence TraceT050

    The model's exact answer did not match the frozen expected output.

    sha256:ae7984bcb2b066e5319436e8b68ccbacb6d08af258a37d7f2240691d63e33d39
  18. Alias TraceT061

    The model's exact answer did not match the frozen expected output.

    sha256:59269d8b21ba7c91f92e7dd0b188b4633753cc7de2057cd6f43b5316f73dc94e
  19. Precedence TraceT062

    The model's exact answer did not match the frozen expected output.

    sha256:bb39ecd83684b0e68451ce2d30f9eb186ceafc1af01c642b0713157f501b00c4
  20. Slice TraceT064

    The model's exact answer did not match the frozen expected output.

    sha256:c38287316dd605f7c7ee5d91539f00cdadd569cfbde7c92aadaa45170dd03e75
  21. Alias TraceT065

    The model's exact answer did not match the frozen expected output.

    sha256:f7da0759b07009cfa7711719a98ec2a0cc078256c3582db48e7a3a7386150f27
  22. Loop TraceT067

    The model's exact answer did not match the frozen expected output.

    sha256:84a4987a6bf095740d6147ae7399bbd2bdd6df9bb057df50b5edcc0f46b04e03
  23. Alias TraceT069

    The model's exact answer did not match the frozen expected output.

    sha256:05db4c0aa426756cd85c3283b28224c4ad30d038bc7caeaf6a224706da4ff31f
  24. Alias TraceT081

    The model's exact answer did not match the frozen expected output.

    sha256:28db6d2da49b7dd6757f7501fd5020e66d16383c1aa42400eec41527732d9a8b
  25. Precedence TraceT082

    The model's exact answer did not match the frozen expected output.

    sha256:3708842ae95b598d3b9fc918e77cd5d2f96d3cdf2e479918e82e02559082640d
  26. Loop TraceT083

    The model's exact answer did not match the frozen expected output.

    sha256:1d6f429dfa54594f873a579b24e699a58311a19d95ab7d88737e8c9706716120
  27. Slice TraceT084

    The model's exact answer did not match the frozen expected output.

    sha256:18a325118a0f454d5cd75c176373b8d332e0a3976028e181023edd19cd4fb645
  28. Alias TraceT085

    The model's exact answer did not match the frozen expected output.

    sha256:769fe3e62c65f4ad2fa7552dd6a76f9ef9d3f4499b93591922a23471eb480453
  29. Precedence TraceT086

    The model's exact answer did not match the frozen expected output.

    sha256:1d4eec4932d4cb96af1041ca692270d00a021c95e4e7fffa209b9f0981e84786
  30. Slice TraceT088

    The model's exact answer did not match the frozen expected output.

    sha256:bdedbfabc0a2f6fc8d273128d8e4cb3f7772d2964a7a941b1f2a4ba7275b7854
  31. Alias TraceT089

    The model's exact answer did not match the frozen expected output.

    sha256:a374ad1f1a4aa2bf4b4da58a0175c005aaeadc737cf75d398987431fc6645c0c
  32. Precedence TraceT090

    The model's exact answer did not match the frozen expected output.

    sha256:55c30147b16a68dfb52f0eab6a225f7577a8653fa0b3914befae995f6b66237f

OFFICIAL SOURCES

Provider record.

DeepSeek describes this model family in its official updates and transparency material. Capabilities stated there are provider claims, separate from our benchmark.

STATUSOfficial source verifiedVerified: 2026-08-27T00:00:00+00:00Research method: official primary-web fallback

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe only the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:6be8a84b6219d15f9b0359cc3d043ca40ac7f08ff27f83227f5bf4003b954740
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the exact page, not a screenshot without context.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.