Skip to content
DeepSeek V4.1-Flash logo

DeepSeek V4.1-Flash: Pricing, Open Weights and Model Guide

DeepSeek

DeepSeek model released September 10, 2026, the smallest in its new architecture family and served under the API name deepseek-flash. Open weights under the MIT licence: 552B total parameters with 8B active during prefill and 16B during decode. Takes text and image input with a 1M-token context and 384K-token maximum output, thinking on by default. Peak API pricing is $0.30 input (cache miss) and $1.20 output per million tokens, half that off-peak.

Pricing & specs checkedSource: vendor documentation · checked Oct 2, 2026LMArena: 1473 (rank 40) · leaderboard · retrieved Oct 1, 2026
Pricing

$0.30 / $1.20 per 1M tokens (peak, cache miss); $0.15 / $0.60 off-peak

Context

1M tokens

LMArena Elo

1473

Overall rank

#40

Key Features

Reasoning Vision Long Context Tool Use Open Weights Low Cost Cached Input Pricing

What DeepSeek V4.1-Flash is

DeepSeek released V4.1-Flash on September 10, 2026, describing it as the smallest model in its new architecture family. Two things make it more than a minor version bump. It is open-weight, published on Hugging Face under the MIT licence, and it takes over DeepSeek's API name: requests to deepseek-flash now reach V4.1-Flash, and the older deepseek-v4-flash and deepseek-v4-flash-vision-exp names are still accepted but routed to it. If an application of yours was pinned to V4-Flash, it has been running on this model since September 10 whether you changed anything or not.

Size, and why the active count matters more than the total

DeepSeek's model card gives 552 billion parameters in total, with 8 billion active per token during prefill and 16 billion during decode. That split is what a mixture-of-experts design buys: the model carries a large pool of parameters but only switches a small slice of them on for each token, which is why it can be priced the way it is. For self-hosting, though, the total is the number that decides your hardware, because all 552B have to be loaded somewhere even if only a fraction work at once. Open weights make it possible to run; they do not make it small.

Pricing, peak and off-peak

DeepSeek prices by time of day, and off-peak rates are exactly half of peak. Per million tokens at peak: $0.30 for input on a cache miss, $0.006 on a cache hit, and $1.20 for output. Off-peak those become $0.15, $0.003 and $0.60. There is no surcharge for long prompts across the full window. The cache-hit rate is the figure to notice: at $0.006 per million, a workload that resends the same long context many times costs very little once that context is cached. As with any model, our suggestion is to budget on complete tasks rather than single calls, and to check whether your traffic falls mostly in peak or off-peak hours before comparing it with a flat-rate model.

Context, output, modes and features

The context window is 1M tokens and maximum output is 384K. The model natively reads images as well as text, which the previous V4-Pro does not, and generates text. It has a thinking mode, on by default, and a non-thinking mode for faster replies. DeepSeek lists JSON output and tool calls as supported, and fill-in-the-middle completion in beta, available in non-thinking mode only.

DeepSeek's published results, and how to read them

The model card reports 90.9 on GPQA Diamond, 74.2 on DeepSWE v1.1, 31.2 on Terminal-Bench 4.0, 74.1 on MMLU-Pro, 79.4 on HumanEval and 93.0 on GSM8K, along with vision results such as 77.9 on CVBench. These are DeepSeek's own runs under its own settings. Several of these benchmarks also appear in other vendors' launch material, but with different scaffolds, effort levels and prompt formats, so the numbers are not directly comparable across companies. LMArena has not rated V4.1-Flash yet; its Elo will appear here and on our benchmarks board once it does.

What happened to V4-Flash

DeepSeek's documentation states the models behind the older V4-Flash API names have been retired, with their requests routed to V4.1-Flash. We keep the V4-Flash page as a record of the earlier model, marked retired, and this page carries the figures for the model you actually reach today.

Key Takeaways

  • Released September 10, 2026; served as deepseek-flash, with the old V4-Flash API names routed to it.
  • Open weights under the MIT licence: 552B total parameters, 8B active in prefill and 16B in decode.
  • Peak pricing $0.30 input (cache miss), $0.006 cached and $1.20 output per million tokens; off-peak is half.
  • 1M-token context, 384K maximum output, image and text input, thinking on by default.
  • Vendor benchmarks are DeepSeek's own runs; check them against your own tasks.

Official Resources

Full Specifications

$0.30 / $1.20 per 1M tokens (peak, cache miss); $0.15 / $0.60 off-peak
Identity
DeveloperDeepSeek
ReleasedSep 2026
StatusGA
LicenceMIT
Self-hostableYes - open weights on Hugging Face
Cost
Blended $/1M tokensinput × 0.75 + output × 0.25$0.525 / 1M tokens
Input price$0.300 / 1M tokens
Output price$1.200 / 1M tokens
Cached input$0.006 / 1M tokens (peak); $0.003 off-peak
Batch discountOff-peak billing: 50% off all rates
Capacity
Context window1M tokens
Max output384k tokens
Long-context surchargeNone (flat rate across the 1M window)
Capability
Vision inYes
Audio inNo
Function callingYes
Structured outputYes - JSON output
Extended reasoningYes - thinking (default) and non-thinking modes
Web searchNo
Code executionNo
Access
APIYes
Chat appYes
Cloud marketplacesNot offered
Fine-tuningNot offered (weights are open)
Measured quality
LMArena Elo1473 (checked Oct 2026)
LMArena rankRank 40
Elo per dollarLMArena Elo ÷ blended $/1M tokens2806