DeepSeek-V4-Pro (checkpoint 0813) and the brand-new DeepSeek V4.1 Flash bring 1M-token context, sparse-attention MoE architecture, native vision, and frontier-class agentic reasoning - open-weight, MIT-licensed, and priced from $0.22 per million input tokens off-peak.
A running log of DeepSeek's model launches, pricing changes and industry news - checked against DeepSeek's official change log and major press coverage.
DeepSeek released V4.1 Flash on 10 September 2026 (Beijing time), a day after announcing it on its developer platform. It is the smallest model in DeepSeek's new architecture series and the first with native multimodal (vision) understanding built in, rather than bolted on via a separate Vision-Exp model.
The headline engineering change is a dramatically smaller KV cache: DeepSeek says HBM requirements drop to roughly 1/4 and SSD requirements to 1/8 of the previous generation - which matters most for long-running agent workloads. DeepSeek claims V4.1 Flash beats V4 Pro on performance, cost, speed and total time, and is now routing V4 Pro API requests to V4.1 Flash at Flash prices. A full benchmark table has not yet been published on the public change log.
New Flash-series rates apply off-peak from 10 Sep 2026 12:00 Beijing time; peak-hour pricing remains 2× off-peak.
The NSA, FBI and CISA issued a joint advisory naming DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI, alleging they route distillation requests through multiple pathways to extract US models in breach of terms of use, "likely with knowledge of the Chinese government". The advisory urges US AI developers to take immediate protective action; lawmakers are weighing penalties.
Experimental multimodal model (deepseek-v4-flash-vision-exp) with V4-Flash-level text ability and vision-agent performance approaching Claude Opus 4.8. Terminal-Bench 2.1: 83.9 · NL2Repo: 57.7 · DeepSWE: 59.3.
Off-peak prices set at 50% of peak. Peak windows are 01:00–04:00 and 06:00–10:00 UTC. See the pricing table.
Rolled out across app, web and API with stronger agent capabilities, native Responses API support (Codex-compatible) and three thinking-effort levels: low, high and max. Full V4-Pro guide →
284B total / ~13B active parameters. Terminal-Bench 2.1 82.7, Arena code Elo 1577. V4 Flash guide →
Hosted deepseek-chat / deepseek-reasoner (V3.2 / R1) were switched off. Weights remain MIT-licensed on Hugging Face.
First look at the 1.6T-parameter V4 family with DeepSeek Sparse Attention and 1M context. V4 overview →
The current hosted line-up is the V4 series. Earlier models - R1, V3.2, Coder, Math, VL - are no longer served on DeepSeek's API but remain open-weight for self-hosting.
The smallest model in DeepSeek's new architecture series, with native multimodal vision understanding and a much smaller KV cache (HBM need cut to 1/4, SSD to 1/8 of the previous generation). DeepSeek says it beats V4 Pro on performance, cost, speed and total time; V4 Pro requests are now routed here and billed at Flash prices. New Flash-series pricing applies from launch day.
DeepSeek's most capable model: a 1.6T-parameter mixture-of-experts with ~49B active parameters, 61 layers of alternating CSA/HCA attention, 1M-token input and 384K output. Built for agentic workflows with native Responses API support, Codex compatibility and low / high / max thinking effort. Scores 87.9 on Terminal-Bench 2.1 and 80.6 on SWE-bench Verified.
284B total / ~13B active parameters with 1M context. Near-Pro reasoning at a fraction of the cost - the workhorse for chatbots, extraction, classification, routing and latency-sensitive apps. Terminal-Bench 2.1 82.7, NL2Repo 54.2, Cybergym 76.7, Arena code Elo 1577. Legacy deepseek-v4-flash calls are being routed to V4.1 Flash.
Experimental multimodal model: the same text, agent and reasoning ability as V4-Flash plus image understanding. Multimodal agent performance approaches Claude Opus 4.8 on vision-dependent benchmarks. Terminal-Bench 2.1 83.9, DSBench-Hard 63.6, Chartography 64.3. Call it as deepseek-v4-flash-vision-exp.
The RL-trained reasoning model that put DeepSeek on the map - chain-of-thought developed without supervised fine-tuning, 97.3% on MATH-500. No longer served on the DeepSeek API (its role is now covered by V4's thinking modes), but the MIT-licensed weights remain on Hugging Face for self-hosting.
671B-parameter MoE with 37B active per token; introduced DeepSeek Sparse Attention (DSA). Powered the old deepseek-chat and deepseek-reasoner endpoints until July 2026. Still a strong self-hosted option on 8×A100/H100-class hardware.
Purpose-built for software engineering with repository-level understanding and 82.6% on HumanEval. For hosted coding today, V4-Pro and V4.1 Flash are far stronger (Terminal-Bench 2.1, SWE-bench Verified), but Coder V2 remains a popular local model via Ollama.
Trained on competition problems (AMC, AIME, Olympiad), papers and proof corpora. Superseded for hosted use by V4-Pro's thinking modes, but still useful as a compact local model for theorem-proving research.
Full comparison of every DeepSeek model → deepseeksr1.com/deepseek-models/
Access DeepSeek through a browser, download the app, integrate via the OpenAI-compatible API, or self-host the open-weight models on your own hardware.
DeepSeek powers applications across every domain - from autonomous coding agents to scientific research, enterprise automation and creative work.
V4-Pro scores 87.9 on Terminal-Bench 2.1 and 80.6 on SWE-bench Verified. Native Responses API and Codex compatibility make it a drop-in engine for terminal agents, repo-scale refactors and CI fixers.
Literature review, hypothesis generation, experiment design and data analysis. "Max" thinking effort handles multi-step scientific reasoning in chemistry, biology and physics (HLE 60.0 with tools).
Competition math (AMC, AIME, Olympiad), university problems, theorem proving and quantitative finance - DeepSeek models earned gold-medal results at IMO and IOI 2025.
Extract, summarise and analyse PDFs, contracts, charts and screenshots. 1M-token context processes entire codebases or books in one prompt; V4.1 Flash adds native vision.
Build agents that process emails, fill forms, route tickets and automate workflows. Function calling, structured JSON output and Toolathlon-Verified 74.1 tool-use accuracy.
Write, translate and localise in 30+ languages. DeepSeek remains one of the strongest Chinese-language models available for bilingual products.
Personalised step-by-step explanations in math, science, coding and languages - and it's free on web and mobile, so students can use it without a subscription.
Brainstorm, draft, edit and polish articles, stories, marketing copy and scripts while maintaining narrative consistency across book-length projects thanks to 1M context.
Interpret datasets, write analysis scripts, generate charts via code and summarise insights. DSBench-Hard 67.2 (V4-Pro). Works with CSV, SQL and pandas workflows.
The architectural ideas behind the V4 series - and why DeepSeek can serve frontier-class models at a fraction of the usual cost.
V4-Pro has 1.6T total parameters but activates only ~49B per token; V4-Flash activates ~13B of 284B. Frontier intelligence with a small fraction of the compute per request.
Reduces long-context attention cost from O(L²) toward O(kL), making the 1M-token context window practical and affordable for whole-repository and book-length inputs.
V4's hyper-connection scheme keeps signal amplification under 2× at only ~6.7% overhead, stabilising very deep MoE training. V4-Pro alternates CSA and HCA attention across its 61 layers.
V4.1 Flash's new structure shrinks the KV cache dramatically - HBM requirements fall to about 1/4 and SSD to 1/8 of the previous generation, which is where long-running agents spend most of their cost.
V4.1 Flash understands images natively rather than via a separate vision adapter; the earlier V4-Flash-Vision-Exp already approaches Opus 4.8 on vision-dependent agent benchmarks.
Choose low, high or max reasoning effort per request. Low is fast and cheap for chat; max unlocks deep chain-of-thought for hard math, science and multi-step agent tasks.
Model weights and technical reports are published under the MIT License. Self-host, fine-tune and ship commercial products - including the retired R1 and V3.2 models.
Native function calling, structured JSON output, and a Responses API adapted for Codex-style coding agents. V4-Pro scores 74.1 on Toolathlon-Verified.
Drop-in replacement for OpenAI's Chat Completions. Change the base URL and key, swap the model name, and existing integrations keep working.
From zero to a working DeepSeek integration in minutes - whether you're a casual user or an enterprise developer.
Use the free web chat, download the iOS/Android app, or sign up for an API key at platform.deepseek.com for developer access.
For most tasks use deepseek-v4-flash (now served by V4.1 Flash) - fast, cheap and multimodal. For the hardest agentic and reasoning work use deepseek-v4-pro. For image-heavy experiments try deepseek-v4-flash-vision-exp.
Be specific. Include context, constraints and desired output format. DeepSeek excels at structured tasks - ask for JSON, tables, step-by-step solutions, or code with explanations.
Set thinking effort to low for chat and extraction, high for coding and analysis, and max for competition math, research or long agent runs. Higher effort costs more output tokens and takes longer.
Install pip install openai. Set base_url="https://api.deepseek.com" and your API key. Use the model IDs above with Chat Completions, or the native Responses API for Codex-style agents.
Keep system prompts stable to maximise cache hits - cached input on V4-Flash is $0.007/1M vs $0.22/1M uncached (off-peak). Batch non-urgent jobs outside the peak windows (01:00–04:00 and 06:00–10:00 UTC) to pay half price.
def is_prime(n): if n < 2: return False for i in range(2, int(n**0.5)+1): if n % i == 0: return False return TrueO(√n) time - much faster than checking all divisors up to n.Pay-per-token with no subscription. Off-peak hours are 50% off, and Flash-series prices were cut again on 10 September 2026. The web chat and apps stay free.
Full access to DeepSeek's chat interface. Unlimited conversations with the current V4 models, including image upload.
Free iOS and Android app with all web features plus voice input, widgets, and cross-device conversation sync.
No minimums and no monthly fee. Cached input on V4-Flash starts at $0.007 per million tokens off-peak. Top up any amount and pay only for what you use.
| Model | Context | Input (Cache Hit) | Input (Cache Miss) | Output | Status |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | — | ¥0.02/1M | ¥1.00/1M | ¥4.00/1M | New · from 10 Sep 12:00 CST |
| deepseek-v4-flash (0731) | 1M | $0.007/1M | $0.22/1M | $0.66/1M | Routed to V4.1 Flash |
| deepseek-v4-pro (0813) | 1M / 384K out | $0.022/1M | $0.66/1M | $1.98/1M | Requests routed to V4.1 Flash at Flash price |
| deepseek-v4-flash-vision-exp | 1M | Experimental - see platform for current rates | Experimental | ||
| deepseek-chat / deepseek-reasoner | 128K | Retired 24 Jul 2026 - weights still free to self-host | Retired | ||
USD figures are off-peak rates effective 16 Aug 2026. V4.1 Flash prices are DeepSeek's published CNY rates (off-peak, per 1M tokens) from 10 Sep 2026; official USD list prices had not been posted at time of writing - check the pricing page for updates.
Run smaller distilled or quantized DeepSeek models (1.5B–70B) locally with Ollama, LM Studio or llama.cpp. Needs 8–48GB VRAM depending on size.
Self-host V4-Flash (284B) or V4-Pro (1.6T) weights. V4-Flash fits a single 8-GPU node; V4-Pro needs multi-node clusters. Ideal for enterprises processing millions of requests under strict data control.
Access DeepSeek open-weight models via AWS Bedrock, Azure AI, Google Vertex AI, Together AI, Fireworks and others for enterprise compliance and SLAs.
Agentic coding and reasoning benchmarks for DeepSeek-V4-Pro-0813 against Claude Opus 4.8 and Kimi K3 - the models it is most often compared with in 2026.
● BENCHMARK SCORES (DeepSeek-V4-Pro-0813 vs competitors)
● FULL MODEL COMPARISON TABLE
| Metric | DeepSeek V4-Pro | Claude Opus 4.8 | Kimi K3 |
|---|---|---|---|
| Output price /1M | $1.98 (off-peak) | $25.00 | — |
| Context window | 1M | — | — |
| Open weights | ✓ MIT | ✗ | — |
| Terminal-Bench 2.1 | 87.9 | 85.0 | 88.3 |
| SWE-bench Verified | 80.6 | 88.6 | — |
| HLE (no tools / tools) | 42.7 / 60.0 | 49.8 / 57.9 | — |
| NL2Repo | 61.5 | 69.7 | — |
| Cybergym | 83.3 | 78.3 | — |
| Toolathlon-Verified | 74.1 | — | 76.5 |
| DeepSWE | 62.7 | — | 67.5 |
| DSBench-Hard | 67.2 | — | — |
Scores from DeepSeek's V4-Pro-0813 release notes and vendor model cards as of Aug–Sep 2026. "—" = not reported on a comparable setting. Claude Fable 5 leads SWE-bench Verified at 95.0 but at $50/1M output tokens. V4.1 Flash benchmarks will be added once DeepSeek publishes them. Deeper dives: DeepSeek vs Claude · DeepSeek vs GPT-5 · all models compared.
DeepSeek is exceptional value - but like all AI systems it has real strengths to leverage and real limitations, including some new regulatory headwinds, to be aware of.
V4-Pro output costs $1.98/1M off-peak versus $25 for Claude Opus 4.8 - and V4.1 Flash cut Flash-series prices a further 11–60% on 10 Sep 2026.
87.9 on Terminal-Bench 2.1 and 83.3 on Cybergym beat Opus 4.8; native Responses API and Codex compatibility make it easy to slot into agent frameworks.
Self-host, fine-tune and build commercial products freely. Even retired models like R1 and V3.2 remain downloadable.
Entire repositories, long contracts or book-length manuscripts fit in a single prompt, with sparse attention keeping it affordable.
Change the base URL, key and model name to migrate from ChatGPT-based code. Zero rewriting of existing integrations.
On 9 Sep 2026 the NSA, FBI and CISA accused DeepSeek and other Chinese labs of systematically distilling US models. Sanctions or procurement restrictions are possible; US government and defence-adjacent buyers should check policy before adopting the hosted API.
DeepSeek is a Chinese company; conversations on the hosted service may be stored on servers in China. Use self-hosting or a compliant cloud provider for sensitive or regulated data.
The hosted model avoids certain sensitive political topics related to China. Content filters may restrict responses that other models answer freely.
API prices double during peak windows (01:00–04:00 and 06:00–10:00 UTC), and heavy demand around launches can cause rate limits or slowdowns on the free tier.
Native vision only arrived with V4.1 Flash on 10 Sep 2026 and no benchmark table has been published yet. No audio input or image generation.
DeepSeek V4.1 Flash was released on 10 September 2026. It is the smallest model in DeepSeek's new architecture series, adds native multimodal (vision) understanding, and uses a much smaller KV cache - DeepSeek says HBM requirements fall to about 1/4 and SSD to 1/8 of the previous generation. DeepSeek claims it beats V4 Pro on performance, cost, speed and total time, and is routing V4 Pro API requests to V4.1 Flash at Flash prices. Flash-series pricing was cut the same day. A full benchmark table had not been published at the time of writing.
Yes - the web chat and the mobile apps are free, with no subscription tier required. The API is pay-per-token: off-peak, V4-Flash input costs $0.22 per million tokens ($0.007 with cache hits) and output $0.66 per million, and V4.1 Flash is cheaper still. Off-peak rates are 50% of peak. See free access options.
A token is the smallest unit of text a model processes - roughly 0.75 words (or ~4 characters) in English. The API charges separately for input tokens (what you send) and output tokens (what the model generates). With DeepSeek-V4-Pro off-peak, cache-hit input costs $0.022/1M, cache-miss input $0.66/1M and output $1.98/1M; V4-Flash is $0.007 / $0.22 / $0.66. A typical 500-word exchange costs a fraction of a cent. DeepSeek automatically caches repeated prompts such as system messages.
DeepSeek retired its hosted pre-V4 endpoints - deepseek-chat and deepseek-reasoner, built on V3.2 and R1 - on 24 July 2026. Their role is now filled by the V4 series' built-in thinking modes: choose low, high or max thinking effort instead of toggling DeepThink. The R1 and V3.2 weights remain on Hugging Face under the MIT licence, so you can still self-host them.
Start with V4.1 Flash. It is the default for chat, extraction, classification, image understanding and most coding tasks, and DeepSeek says it now matches or beats V4 Pro at a lower price - which is why V4 Pro API requests are currently routed to it. Reach for V4-Pro (1.6T parameters, 49B active) when you need the strongest possible results on long agentic runs, hard reasoning or the 384K-token output limit.
For non-sensitive work (marketing copy, code review, public data analysis) the cloud API is widely used. However, DeepSeek is a Chinese company - data processed through the public API may be stored on servers subject to Chinese law - and on 9 September 2026 US security agencies (NSA, FBI, CISA) issued an advisory accusing DeepSeek and other Chinese labs of unauthorised model distillation, raising the prospect of sanctions or procurement bans. For sensitive data (healthcare, finance, legal, government) we recommend self-hosting the open weights or using a compliant cloud provider (AWS Bedrock, Azure AI) with SOC2/HIPAA SLAs and data residency guarantees.
DeepSeek's API is OpenAI-compatible. Set base_url="https://api.deepseek.com" and your DeepSeek api_key, then change the model name to deepseek-v4-flash or deepseek-v4-pro. Streaming, function calling and structured outputs work unchanged. If you use the Codex CLI or the Responses API, DeepSeek V4 supports that natively too. Most teams migrate in minutes - see the API guide.
Yes. DeepSeek publishes weights under the MIT License. For a laptop or single GPU, the easiest route is Ollama or LM Studio with a distilled or quantized model (1.5B–70B, 8–48GB VRAM). The full V4-Flash (284B total, 13B active) fits on a single 8-GPU server; V4-Pro (1.6T) needs a multi-node cluster. See our download and self-hosting guide.
Yes. The V4 series supports OpenAI-compatible function calling, structured JSON output and multi-turn agent workflows, plus a native Responses API adapted for Codex. V4-Pro-0813 scores 74.1 on Toolathlon-Verified and 87.9 on Terminal-Bench 2.1, making it one of the strongest open-weight models for building AI agents.
DeepSeek has shipped a new checkpoint roughly every two months in 2026 (V4-Flash 0731, V4-Pro 0813, V4.1 Flash on 10 Sep). Based on that cadence, the next V4-Pro checkpoint or a V4.1 Pro is expected around mid-October 2026. We update this page as soon as DeepSeek's change log confirms a release.
DeepSeek was founded in 2023 as a subsidiary of High-Flyer, a Chinese quantitative hedge fund, and is headquartered in Hangzhou. It became globally prominent in January 2025 when DeepSeek-R1 matched frontier reasoning performance at a reported fraction of the training cost. The lab publishes its research openly and releases model weights under the MIT License. Read more in What is DeepSeek AI?
Try DeepSeek V4.1 Flash and V4-Pro free on the web, or build with an API that is open, affordable and remarkably capable.