Skip to content
DeepSeek V4-Pro logo

DeepSeek V4-Pro: A Low-Cost Frontier API From DeepSeek

DeepSeek

DeepSeek's V4-Pro, a paid API model with a 1M token context window and 384k max output. Reached general availability on 13 August 2026 as deepseek-v4-pro, and moved to peak/off-peak billing on 16 August: $1.32 input / $3.96 output per 1M tokens at peak, halved to $0.66 / $1.98 off-peak.

Pricing & specs verifiedSource: vendor documentation · checked Sep 13, 2026LMArena: 1457 (rank 57) · leaderboard · retrieved Sep 14, 2026
Pricing

$1.32 / $3.96 per 1M tokens (peak, cache miss); $0.66 / $1.98 off-peak

Context

1M tokens

LMArena Elo

1457

Overall rank

#57

Key Features

Large Context Low Cost Cached Input Pricing Coding Math Reasoning

What is DeepSeek V4-Pro?

DeepSeek V4-Pro is a model from the Chinese AI lab DeepSeek, offered as a paid API with a 1 million token context window. Correction (12 August 2026): this page previously described it as open-weight under a permissive MIT licence and free to self-host. That was wrong. DeepSeek charges for V4-Pro - under peak/off-peak billing introduced on 16 August 2026, standard rates are $1.32 per million input tokens and $3.96 per million output tokens, halving automatically to $0.66 / $1.98 outside peak hours - and its documentation names no open-source licence. If you need weights you can actually download and run, look at Llama 4 Scout or Mistral Large 3 instead.

Cached Input Pricing

The pricing detail worth understanding is the cache. At peak hours DeepSeek charges $1.32 per million input tokens on a cache miss but $0.044 per million on a cache hit - about 30x cheaper. Off-peak every rate halves automatically: $0.66 cache-miss input and $0.022 cached. (AIblogly analysis: for workloads that resend a large fixed prefix on every call, such as a long system prompt or a fixed document, that gap dominates the bill far more than the headline rate does. For one-off varied prompts it will not apply at all.) Output is $3.96 per million tokens at peak, $1.98 off-peak.

Coding and Math Strength

DeepSeek's pricing page does not publish coding or math benchmark scores for V4-Pro, and this guide previously claimed it "matched leading closed models" on them without a source - that claim has been removed. What is verifiable is price: at $0.435/$0.87 per million tokens it costs roughly 5-10% of Claude Opus 5 or GPT-5.6 Sol on the same API. Whether that trade is worth it for your coding workload is something to test against your own tasks, not something this guide can assert for you.

Context Window

V4-Pro offers a 1 million token context window, on par with Claude Opus 5 and Gemini 3.1 Pro and slightly ahead of GPT-5.6's 1.05M. DeepSeek's API documentation states pricing and context but does not detail architecture, so this guide does not describe one.

Correction: This Is Not a Self-Hosting Option

An earlier version of this page had a "Deployment and Cost" section describing V4-Pro as free to self-host with vLLM, SGLang or Ollama, and downloadable from Hugging Face. That was wrong and has been removed - it is a metered API product, priced per token as described above, with no published weights. If self-hosting is what you actually need, Llama 4 Scout and Mistral Large 3 on this site both genuinely offer it.

Ideal Use Cases

Based on verified pricing, DeepSeek V4-Pro fits cost-conscious, high-volume applications where paying roughly 15-30% of what Claude Opus 5 or GPT-5.6 Sol charge per API call matters more than the highest LMArena ranking. It does not fit use cases requiring self-hosting, fine-tuning on private infrastructure, or data never leaving your own systems - those need Llama 4 Scout or Mistral Large 3.

Limitations

V4-Pro is an API-only model. DeepSeek publishes no weights or open-source licence for it, so self-hosting, fine-tuning on your own infrastructure and weight inspection are not available - if those matter, Llama 4 Scout and Mistral Large 3 are the options here that genuinely offer them. It also ranks 52nd on LMArena at 1458, below Claude Opus 5, Gemini 3.1 Pro and GPT-5.6. DeepSeek publishes no multimodal capability detail, so this guide makes no claim either way. Organisations in regulated settings may separately want to consider where the API is hosted and how data is governed.

DeepSeek V4-Pro vs Competitors

On published API pricing V4-Pro is among the cheapest frontier-tier APIs on this site: $1.32 input and $3.96 output per million tokens at peak - halving to $0.66/$1.98 off-peak - against $2/$12 for Gemini 3.1 Pro and $5/$25 for Claude Opus 5. On LMArena it ranks 52nd at 1458, below all three of those. (AIblogly analysis: that is the trade the numbers actually describe - substantially lower cost, measurably lower human-preference ranking.) Note it is not a self-hosting option, despite an earlier version of this page saying so; for downloadable weights see Llama 4 Scout or Mistral Large 3.

Key Takeaways

  • A paid API model - DeepSeek publishes no open-source licence for V4-Pro
  • Priced at $1.32 / $3.96 per 1M input / output tokens at peak; $0.66 / $1.98 off-peak
  • Cached input drops to $0.044 per 1M tokens at peak ($0.022 off-peak)
  • LMArena: 1458 (+/-4), rank 52, retrieved 23 August 2026
  • 1M token context window for whole-repo and long-document work
  • Not an option where self-hosting or data sovereignty is required - see Llama 4 Scout or Mistral Large 3

Official Resources

Full Specifications

$1.32 / $3.96 per 1M tokens (peak, cache miss); $0.66 / $1.98 off-peak
Identity
DeveloperDeepSeek
ReleasedAug 2026
StatusGA
LicenceProprietary
Self-hostableNo
Cost
Blended $/1M tokensinput × 0.75 + output × 0.25$1.980 / 1M tokens
Input price$1.320 / 1M tokens
Output price$3.960 / 1M tokens
Cached input$0.044 / 1M tokens (peak); $0.022 off-peak
Batch discountOff-peak billing: 50% off all rates outside 01:00-04:00 and 06:00-10:00 UTC
Free tierNot offered
Capacity
Context window1M tokens
Max output384k tokens
Long-context surchargeNone (flat rate across the 1M window)
Capability
Vision inNo
Audio inNo
Function callingYes
Structured outputYes
Extended reasoningYes - thinking effort levels low, high and max
Web searchNo
Code executionNo
Access
APIYes
Chat appYes
Cloud marketplacesNot offered
Fine-tuningNot offered
Measured quality
LMArena Elo1457 (checked Sep 2026)
LMArena rankRank 57
Elo per dollarLMArena Elo ÷ blended $/1M tokens736