Pareto ·
DeepSeek V4.1 Flash
Who’s off the frontier?
● Flash━ New frontier┄ Before Flash● Beaten on both axes by Flash○ Another lab’s report
Token price, not cost per task. Different harnesses and reasoning budgets. “Beaten” means a reported score and price comparison, not statistical significance.
Change assumptions
“Across lab reports” prefers each model maker’s published result. A hollow dot marks a result found only in another lab’s comparison. “DeepSeek’s table” preserves its six priced, still-available models, without implying identical test settings.
Changing the Flash harness changes only its Terminal-Bench 2.1 and DeepSWE scores. All other models keep their published setups.
How to read this one
Scores, prices & sources
| Model | Score | $/1M | Position | Score source |
|---|
Price column follows the selected token mix. Open a source note for the raw input/output rates, harness details and conflicting reports. Download all data.
Method & judgment calls
A price frontier, not a universal model ranking
A point is dominated if another plotted point costs no more and scores no lower, with at least one strict improvement. The dashed line shows the frontier without V4.1 Flash; the green line adds it. Lines only join observed points, they do not imply achievable intermediate models. Missing scores are omitted, not zero-filled. No averages across unrelated benchmarks.
What the price means
USD per million tokens, provider-direct Standard API rates checked on 10 Sept 2026. The default mix is (3 × uncached input + output) ÷ 4, an explicit illustration, not measured usage. Output rates include billed reasoning tokens. No cache, batch, Flex, Fast mode, tool fees, tax or long-context uplifts. Cheaper tokens do not guarantee a cheaper completed task.
DeepSeek defaults to peak pricing, not its attention-grabbing off-peak rate. Peak hours are weekdays 01:00–04:00 and 06:00–10:00 UTC. Both DeepSeek models halve their rates off-peak. Gemini’s introductory rates last through 31 Dec 2026. OpenAI’s direct table lists Sol Standard at $4/$20 and Batch/Flex at $2/$10; this chart uses Standard, not the lower reseller quote.
Where launch claims disagree
Opus 5’s own launch shows 68.8% on DeepSWE; DeepSeek reports 74.0%, and OpenAI’s Astra table reports 73.7%. Kimi claims 93.5% on GPQA versus DeepSeek’s 92.9%. Sol claims 94.6% on GPQA versus DeepSeek’s 94.1%. These are source differences, not confidence intervals. Use the source switch to see their effect rather than silently choosing the highest number.
What is deliberately left out
AutomationBench results jump from 26.0% for Opus in Anthropic’s launch to 50.3% in DeepSeek’s table; versions, fixes and scoring differ. Agent’s Last Exam also has incompatible CLI and broader scores. ProgramBench “Almost@1” is not the larger score in Kimi’s launch. We do not blend these. We also keep HLE full-set/no-tools separate from text-only, HLE-Verified and tool-assisted scores.
V4 Flash has already been retired on DeepSeek’s direct API, so its old score is not paired with a made-up current price. V4 Pro is still available at this snapshot but redirects to V4.1 Flash from 14 Sept 2026. We exclude Sol Ultra’s multi-agent scores. This is a selected comparison across six labs, not an exhaustive leaderboard or an independent evaluation.
Reproducibility matters
DeepSeek’s own ablation moves Terminal-Bench 2.1 from 90.6% in DSH Minimal to 86.1% in Pi and 84.1% in Codex. DeepSWE moves from 74.2% in mini-SWE to 66.2% in Pi. That is why this chart exposes the harness choice and does not turn tiny point-score leads into blanket claims of superiority.