Back to programming-model index

MODEL EVIDENCE / OpenRouter #18

Google: Gemini 2.5 Flash Lite

The model identity, provider record, and independently verified benchmark results in one place. The number above is OpenRouter discovery order, not benchmark rank.

google/gemini-2.5-flash-lite
Official source pendingNot yet run
Fast numbersCaptured facts and latest status
OpenRouter feed rank
OpenRouter #18
Context
1048576
Input
text, image, file, audio, video
Output
text
Reasoning flag
Yes
Prompt price
$0.100 / 1M tokens
Completion price
$0.400 / 1M tokens
Latest run
Not run
View denominator
Not available

DEEP BENCHMARK

What the run established.

Counts stay visible for auditability. A score appears only when the Most Annoying 50-trap view satisfies the independent checks.

BENCHMARK STATUS

No complete benchmark score exists for this exact model ID. A missing run is not a zero.

Score history

Score history begins after the first public run.

Most Annoying verified mistake receipts

Only receipts from the Most Annoying view are listed here. Tolerable-view receipts are separate and are not added to this score.

  1. No individual verified mistake receipt is published for this exact model and run.

OFFICIAL SOURCES

Provider record.

An exact primary official launch or model-card source has not yet been verified for this captured model ID.

STATUSOfficial source pendingVerified: PendingResearch method: official primary-web fallback
  • No exact primary launch or model card is certified yet. This page will not infer one from a similar name.

Important: Statements in linked publisher material are provider claims. They are not findings of this benchmark.

METHOD & LIMITS

What this page can say.

Every public leaderboard score is the latest Most Annoying 50-trap view. The 100-task run receipt is the wider audit bundle, not the benchmark denominator. Benchmark rank #1 means fewest verified mistakes among complete runs; the list can still be sorted by most mistakes for readability. OpenRouter feed rank is separate discovery metadata and is kept out of the benchmark rank column. Two-week average reporting will use the last four complete 100-task cycles when enough cycles exist. Individual verified mistakes, called catches internally, require primary and shadow checks to reproduce the exact mismatch from the recorded reply against the frozen expected output. Retrieval is not signature verification; auditors should recompute receipt roots and verify signatures against the published issuer key. Routing errors, timeouts, incomplete answers, and missing evidence stay outside the ranked score. Do not judge any model solely from this benchmark. Ranked complete runs must use the same 50 traps, scoring rules, model settings, and retry policy.

  1. Dual-check reproduction

    For each listed Most Annoying mistake, primary and shadow verification must both reproduce the exact mismatch from the recorded reply against the frozen expected output. Otherwise the item remains unresolved.

  2. Certified denominator

    The rankable score is the 50-trap Most Annoying view. The run receipt may say 100 tasks because it covers the wider audit bundle; that object is evidence, not a separate leaderboard denominator.

  3. Signature and feed rank

    Independent verification means recomputing the receipt root and checking the signature against the published issuer public key. OpenRouter rank stays discovery metadata and does not affect benchmark rank.

Reproduce this result

  1. Inspect the run receipt.
  2. Inspect each Most Annoying mistake receipt.
  3. Verify each receipt root and signature.
  4. Wait for the full proof bundle if it is not released yet.
  5. Rerun the exact prompt and expected-output comparison.

Claim ceiling: OpenRouter rank and metadata describe the captured discovery feed. Provider claims are separate from this benchmark. Scores describe only the most recent public test pack and should not be used alone to judge the model generally.

Feed root
sha256:61e4106b1d08b7ef16ec9d332548214d01937a628e573fcf9a3091f601dfcb76
Model metadata root
sha256:14b8029d35483b2f3f35a3dea35e0e150d239d6c472ed21f79798e344ecf4104
Official-source registry root
sha256:e5606a5133065340a402faa6c24a6bb3fb36d82809003a75bae28e85d7ec8e68

QUOTE THE RECORD

Share the exact page, not a screenshot without context.

Discuss this result in:

GET A SECOND EXPLANATION

Ask your favorite AI what this means.

Copy a bounded prompt that includes this model page and its claim ceiling.