DeepSeek V4.1 Flash launched 10 Sep 2026 - native multimodal, cheaper than V4 Flash

Explore the Unexplored
with AI That Reasons

DeepSeek-V4-Pro (checkpoint 0813) and the brand-new DeepSeek V4.1 Flash bring 1M-token context, sparse-attention MoE architecture, native vision, and frontier-class agentic reasoning - open-weight, MIT-licensed, and priced from $0.22 per million input tokens off-peak.

Start Chatting Free API Platform → Download App
1M
Token context (V4 series)
1.6T
V4-Pro params · 49B active
87.9
Terminal-Bench 2.1 (V4-Pro)
$0.22
Per 1M input (V4 Flash, off-peak)
MIT
Open-weight license
LAST UPDATED: 10 SEPTEMBER 2026
DeepSeek V4.1 Flash Released 10 Sep 2026
DeepSeek-V4-Pro-0813 1.6T Params · 49B Active
Context Window 1M Tokens In · 384K Out
Native Multimodal Vision in V4.1 Flash
Flash Pricing Cut up to 60% on 10 Sep
Terminal-Bench 2.1 87.9 (V4-Pro)
Open Weights MIT License
Off-Peak API 50% Discount
Next Checkpoint Expected ~Mid-Oct 2026
DeepSeek V4.1 Flash Released 10 Sep 2026
DeepSeek-V4-Pro-0813 1.6T Params · 49B Active
Context Window 1M Tokens In · 384K Out
Native Multimodal Vision in V4.1 Flash
Flash Pricing Cut up to 60% on 10 Sep
Terminal-Bench 2.1 87.9 (V4-Pro)
Open Weights MIT License
Off-Peak API 50% Discount
Next Checkpoint Expected ~Mid-Oct 2026
Latest Updates · September 2026

What's New at DeepSeek AI

A running log of DeepSeek's model launches, pricing changes and industry news - checked against DeepSeek's official change log and major press coverage.

10 Sep 2026 · New Model

DeepSeek V4.1 Flash is here - native vision, quarter the memory, lower prices

DeepSeek released V4.1 Flash on 10 September 2026 (Beijing time), a day after announcing it on its developer platform. It is the smallest model in DeepSeek's new architecture series and the first with native multimodal (vision) understanding built in, rather than bolted on via a separate Vision-Exp model.

The headline engineering change is a dramatically smaller KV cache: DeepSeek says HBM requirements drop to roughly 1/4 and SSD requirements to 1/8 of the previous generation - which matters most for long-running agent workloads. DeepSeek claims V4.1 Flash beats V4 Pro on performance, cost, speed and total time, and is now routing V4 Pro API requests to V4.1 Flash at Flash prices. A full benchmark table has not yet been published on the public change log.

¥0.02
Cached input / 1M (−60%)
¥1.00
Uncached input / 1M (−33%)
¥4.00
Output / 1M (−11%)
1/4 HBM
Memory vs prior gen

New Flash-series rates apply off-peak from 10 Sep 2026 12:00 Beijing time; peak-hour pricing remains 2× off-peak.

09 SEP 2026
US agencies accuse DeepSeek of "systematic" model distillation

The NSA, FBI and CISA issued a joint advisory naming DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI, alleging they route distillation requests through multiple pathways to extract US models in breach of terms of use, "likely with knowledge of the Chinese government". The advisory urges US AI developers to take immediate protective action; lawmakers are weighing penalties.

21 AUG 2026
DeepSeek-V4-Flash-Vision-Exp launches on the API

Experimental multimodal model (deepseek-v4-flash-vision-exp) with V4-Flash-level text ability and vision-agent performance approaching Claude Opus 4.8. Terminal-Bench 2.1: 83.9 · NL2Repo: 57.7 · DeepSWE: 59.3.

16 AUG 2026
Peak / off-peak API pricing introduced

Off-peak prices set at 50% of peak. Peak windows are 01:00–04:00 and 06:00–10:00 UTC. See the pricing table.

13 AUG 2026
DeepSeek-V4-Pro reaches general availability (checkpoint 0813)

Rolled out across app, web and API with stronger agent capabilities, native Responses API support (Codex-compatible) and three thinking-effort levels: low, high and max. Full V4-Pro guide →

31 JUL 2026
DeepSeek-V4-Flash GA (checkpoint 0731)

284B total / ~13B active parameters. Terminal-Bench 2.1 82.7, Arena code Elo 1577. V4 Flash guide →

24 JUL 2026
Legacy pre-V4 endpoints retired

Hosted deepseek-chat / deepseek-reasoner (V3.2 / R1) were switched off. Weights remain MIT-licensed on Hugging Face.

24 APR 2026
DeepSeek-V4 Preview announced

First look at the 1.6T-parameter V4 family with DeepSeek Sparse Attention and 1M context. V4 overview →

Model Family

Every Model, Every Use Case

The current hosted line-up is the V4 series. Earlier models - R1, V3.2, Coder, Math, VL - are no longer served on DeepSeek's API but remain open-weight for self-hosting.

FAST · 0731
DeepSeek-V4-Flash
GA 31 Jul 2026 (checkpoint 0731) · superseded by V4.1 Flash

284B total / ~13B active parameters with 1M context. Near-Pro reasoning at a fraction of the cost - the workhorse for chatbots, extraction, classification, routing and latency-sensitive apps. Terminal-Bench 2.1 82.7, NL2Repo 54.2, Cybergym 76.7, Arena code Elo 1577. Legacy deepseek-v4-flash calls are being routed to V4.1 Flash.

284B
Params (13B active)
1M
Context
$0.22
Input / 1M off-peak
👁️
VISION · EXPERIMENTAL
DeepSeek-V4-Flash-Vision-Exp
Announced 21 Aug 2026 · API only

Experimental multimodal model: the same text, agent and reasoning ability as V4-Flash plus image understanding. Multimodal agent performance approaches Claude Opus 4.8 on vision-dependent benchmarks. Terminal-Bench 2.1 83.9, DSBench-Hard 63.6, Chartography 64.3. Call it as deepseek-v4-flash-vision-exp.

83.9
Terminal-Bench 2.1
64.3
Chartography
API
Access
🧠
LEGACY · WEIGHTS ONLY
DeepSeek-R1
Released Jan 2025 · hosted endpoint retired 24 Jul 2026

The RL-trained reasoning model that put DeepSeek on the map - chain-of-thought developed without supervised fine-tuning, 97.3% on MATH-500. No longer served on the DeepSeek API (its role is now covered by V4's thinking modes), but the MIT-licensed weights remain on Hugging Face for self-hosting.

97.3%
MATH-500
CoT
Reasoning
MIT
Self-host
🗂️
LEGACY · WEIGHTS ONLY
DeepSeek-V3.2
Released Sep 2025 · hosted endpoint retired 24 Jul 2026

671B-parameter MoE with 37B active per token; introduced DeepSeek Sparse Attention (DSA). Powered the old deepseek-chat and deepseek-reasoner endpoints until July 2026. Still a strong self-hosted option on 8×A100/H100-class hardware.

671B
Parameters
128K
Context
37B
Active
💻
CODE · OPEN WEIGHTS
DeepSeek-Coder V2
MoE code model · 338 languages · self-host

Purpose-built for software engineering with repository-level understanding and 82.6% on HumanEval. For hosted coding today, V4-Pro and V4.1 Flash are far stronger (Terminal-Bench 2.1, SWE-bench Verified), but Coder V2 remains a popular local model via Ollama.

82.6%
HumanEval
338
Languages
MIT
License
📐
MATH · OPEN WEIGHTS
DeepSeek-Math
Specialised mathematics model · self-host

Trained on competition problems (AMC, AIME, Olympiad), papers and proof corpora. Superseded for hosted use by V4-Pro's thinking modes, but still useful as a compact local model for theorem-proving research.

Proof
Verification
AIME
Competition
OSS
Open weights

Full comparison of every DeepSeek model → deepseeksr1.com/deepseek-models/

Access & Download

Get DeepSeek Everywhere

Access DeepSeek through a browser, download the app, integrate via the OpenAI-compatible API, or self-host the open-weight models on your own hardware.

🌐

Web Chat - Free

No install needed. Chat with DeepSeek-V4-Pro and V4.1 Flash in your browser, with image upload and selectable thinking effort.

📱

Mobile App - iOS & Android

Full-featured iOS and Android apps with voice input, image upload, conversation history sync, and the same V4 models as the web.

⚙️

API Platform

OpenAI-compatible Chat Completions plus a native Responses API (Codex-compatible). Pay-per-token with 50% off-peak discounts and no monthly fees.

🐙

Self-Host (Open Weights)

Download MIT-licensed weights for V4-Flash, V3.2, R1 and more from Hugging Face. Run locally with sufficient GPU hardware - no API fees ever.

Python
Node.js
cURL
# pip install openai
from openai import OpenAI

# DeepSeek uses an OpenAI-compatible API
client = OpenAI(
  api_key="<your-deepseek-api-key>",
  base_url="https://api.deepseek.com"
)

response = client.chat.completions.create(
  model="deepseek-v4-flash", # or deepseek-v4-pro
  messages=[
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain MoE architecture"}
  ],
  stream=False
)

print(response.choices[0].message.content)
// npm install openai
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: '<your-deepseek-api-key>',
  baseURL: 'https://api.deepseek.com',
});

const completion = await client.chat.completions.create({
  model: 'deepseek-v4-flash',
  messages: [
    { role: 'user', content: 'Hello!' }
  ],
});

console.log(completion.choices[0].message.content);
# Replace YOUR_KEY with your API key
curl https://api.deepseek.com/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_KEY" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role":"user","content":"Hello!"}
    ]
  }'
Use Cases

What Can You Build?

DeepSeek powers applications across every domain - from autonomous coding agents to scientific research, enterprise automation and creative work.

01
Agentic Coding

V4-Pro scores 87.9 on Terminal-Bench 2.1 and 80.6 on SWE-bench Verified. Native Responses API and Codex compatibility make it a drop-in engine for terminal agents, repo-scale refactors and CI fixers.

02
Scientific Research

Literature review, hypothesis generation, experiment design and data analysis. "Max" thinking effort handles multi-step scientific reasoning in chemistry, biology and physics (HLE 60.0 with tools).

03
Mathematical Reasoning

Competition math (AMC, AIME, Olympiad), university problems, theorem proving and quantitative finance - DeepSeek models earned gold-medal results at IMO and IOI 2025.

04
Document & Image Analysis

Extract, summarise and analyse PDFs, contracts, charts and screenshots. 1M-token context processes entire codebases or books in one prompt; V4.1 Flash adds native vision.

05
Enterprise Automation

Build agents that process emails, fill forms, route tickets and automate workflows. Function calling, structured JSON output and Toolathlon-Verified 74.1 tool-use accuracy.

06
Multilingual Content

Write, translate and localise in 30+ languages. DeepSeek remains one of the strongest Chinese-language models available for bilingual products.

07
Education & Tutoring

Personalised step-by-step explanations in math, science, coding and languages - and it's free on web and mobile, so students can use it without a subscription.

08
Creative Writing

Brainstorm, draft, edit and polish articles, stories, marketing copy and scripts while maintaining narrative consistency across book-length projects thanks to 1M context.

09
Data Analysis

Interpret datasets, write analysis scripts, generate charts via code and summarise insights. DSBench-Hard 67.2 (V4-Pro). Works with CSV, SQL and pandas workflows.

Core Technology

Built Different. Built Better.

The architectural ideas behind the V4 series - and why DeepSeek can serve frontier-class models at a fraction of the usual cost.

🔀
Mixture of Experts (MoE)

V4-Pro has 1.6T total parameters but activates only ~49B per token; V4-Flash activates ~13B of 284B. Frontier intelligence with a small fraction of the compute per request.

🧩
DeepSeek Sparse Attention (DSA)

Reduces long-context attention cost from O(L²) toward O(kL), making the 1M-token context window practical and affordable for whole-repository and book-length inputs.

🔗
Manifold-Constrained Hyper-Connections (mHC)

V4's hyper-connection scheme keeps signal amplification under 2× at only ~6.7% overhead, stabilising very deep MoE training. V4-Pro alternates CSA and HCA attention across its 61 layers.

🗜️
Compact KV Cache (V4.1)

V4.1 Flash's new structure shrinks the KV cache dramatically - HBM requirements fall to about 1/4 and SSD to 1/8 of the previous generation, which is where long-running agents spend most of their cost.

👁️
Native Multimodality

V4.1 Flash understands images natively rather than via a separate vision adapter; the earlier V4-Flash-Vision-Exp already approaches Opus 4.8 on vision-dependent agent benchmarks.

🎚️
Thinking Effort Levels

Choose low, high or max reasoning effort per request. Low is fast and cheap for chat; max unlocks deep chain-of-thought for hard math, science and multi-step agent tasks.

🔓
Open Weights (MIT)

Model weights and technical reports are published under the MIT License. Self-host, fine-tune and ship commercial products - including the retired R1 and V3.2 models.

🔧
Agents, Tools & Responses API

Native function calling, structured JSON output, and a Responses API adapted for Codex-style coding agents. V4-Pro scores 74.1 on Toolathlon-Verified.

🌐
OpenAI-Compatible API

Drop-in replacement for OpenAI's Chat Completions. Change the base URL and key, swap the model name, and existing integrations keep working.

Getting Started

How to Use DeepSeek

From zero to a working DeepSeek integration in minutes - whether you're a casual user or an enterprise developer.

1

Choose Your Access Method

Use the free web chat, download the iOS/Android app, or sign up for an API key at platform.deepseek.com for developer access.

2

Select a Model

For most tasks use deepseek-v4-flash (now served by V4.1 Flash) - fast, cheap and multimodal. For the hardest agentic and reasoning work use deepseek-v4-pro. For image-heavy experiments try deepseek-v4-flash-vision-exp.

3

Write Your Prompt

Be specific. Include context, constraints and desired output format. DeepSeek excels at structured tasks - ask for JSON, tables, step-by-step solutions, or code with explanations.

4

Dial in Thinking Effort

Set thinking effort to low for chat and extraction, high for coding and analysis, and max for competition math, research or long agent runs. Higher effort costs more output tokens and takes longer.

5

Integrate via API

Install pip install openai. Set base_url="https://api.deepseek.com" and your API key. Use the model IDs above with Chat Completions, or the native Responses API for Codex-style agents.

6

Optimise Cost

Keep system prompts stable to maximise cache hits - cached input on V4-Flash is $0.007/1M vs $0.22/1M uncached (off-peak). Batch non-urgent jobs outside the peak windows (01:00–04:00 and 06:00–10:00 UTC) to pay half price.

DS
DeepSeek Assistant
● Online - V4.1 Flash
Can you write a Python function that checks if a number is prime?
Sure! Here's an efficient implementation:

def is_prime(n):
  if n < 2: return False
  for i in range(2, int(n**0.5)+1):
    if n % i == 0: return False
  return True

This runs in O(√n) time - much faster than checking all divisors up to n.
Now explain how it works step by step
Step-by-step breakdown:

1. Base case: Numbers < 2 (0, 1) are not prime by definition.
2. Loop to √n: If n has a factor > √n, there must be a corresponding factor < √n - so we only need to check up to √n.
3. Divisibility check: If any i divides n evenly, n is composite.
4. Return True: No divisors found → n is prime. ✅
Ask DeepSeek anything...
Pricing · updated 10 Sep 2026

Frontier Models at a Fraction of the Price

Pay-per-token with no subscription. Off-peak hours are 50% off, and Flash-series prices were cut again on 10 September 2026. The web chat and apps stay free.

Free Tier
API / Pay-As-You-Go
Self-Host
Web Chat
$0/month

Full access to DeepSeek's chat interface. Unlimited conversations with the current V4 models, including image upload.

DeepSeek-V4-Pro & V4.1 Flash
Thinking modes (low / high / max)
Image & file uploads
Conversation history
Web search
API access
Priority during peak hours
Start Free →
Most Popular
Mobile App
$0/month

Free iOS and Android app with all web features plus voice input, widgets, and cross-device conversation sync.

All web features included
Voice input support
Cross-device sync
Notifications
iOS 16+ / Android 8+
API access
Download App →
API Pay-As-You-Go
From $0.007/1M

No minimums and no monthly fee. Cached input on V4-Flash starts at $0.007 per million tokens off-peak. Top up any amount and pay only for what you use.

All hosted V4 models
50% off-peak discount
Automatic prompt caching
Full API documentation
Usage dashboard
Get API Key →
💡 No monthly fees. Cost = (input tokens × input price) + (output tokens × output price). Off-peak prices below are 50% of peak. Peak hours: 01:00–04:00 and 06:00–10:00 UTC. Cache hits cut input cost by ~97%.
Model Context Input (Cache Hit) Input (Cache Miss) Output Status
DeepSeek V4.1 Flash ¥0.02/1M ¥1.00/1M ¥4.00/1M New · from 10 Sep 12:00 CST
deepseek-v4-flash (0731) 1M $0.007/1M $0.22/1M $0.66/1M Routed to V4.1 Flash
deepseek-v4-pro (0813) 1M / 384K out $0.022/1M $0.66/1M $1.98/1M Requests routed to V4.1 Flash at Flash price
deepseek-v4-flash-vision-exp 1M Experimental - see platform for current rates Experimental
deepseek-chat / deepseek-reasoner 128K Retired 24 Jul 2026 - weights still free to self-host Retired

USD figures are off-peak rates effective 16 Aug 2026. V4.1 Flash prices are DeepSeek's published CNY rates (off-peak, per 1M tokens) from 10 Sep 2026; official USD list prices had not been posted at time of writing - check the pricing page for updates.

LIGHT USE / MONTH
$1–10
Personal projects
MEDIUM USE / MONTH
$10–50
Small SaaS apps
HEAVY USE / MONTH
$50–200
Production apps
Distilled / Quantized
$0/API calls

Run smaller distilled or quantized DeepSeek models (1.5B–70B) locally with Ollama, LM Studio or llama.cpp. Needs 8–48GB VRAM depending on size.

No API fees ever
Full data privacy
Works offline
Ollama / LM Studio support
Reduced capability vs full model
Local setup guide →
Full Power
Full V4-Flash / V4-Pro
GPU infrastructure

Self-host V4-Flash (284B) or V4-Pro (1.6T) weights. V4-Flash fits a single 8-GPU node; V4-Pro needs multi-node clusters. Ideal for enterprises processing millions of requests under strict data control.

Full model capability
MIT License - commercial OK
Complete data control
Fine-tune on your data
Requires ~$50K+ GPU infra
Hugging Face →
Cloud Providers
Cloud API

Access DeepSeek open-weight models via AWS Bedrock, Azure AI, Google Vertex AI, Together AI, Fireworks and others for enterprise compliance and SLAs.

Enterprise SLAs
SOC2 / HIPAA options
Regional data residency
No GPU management
Higher per-token cost
AWS Bedrock →
Benchmarks & Comparison

How DeepSeek Stacks Up

Agentic coding and reasoning benchmarks for DeepSeek-V4-Pro-0813 against Claude Opus 4.8 and Kimi K3 - the models it is most often compared with in 2026.

● BENCHMARK SCORES (DeepSeek-V4-Pro-0813 vs competitors)

DeepSeek-V4-Pro-0813
Claude Opus 4.8
Kimi K3

● FULL MODEL COMPARISON TABLE

Metric DeepSeek V4-Pro Claude Opus 4.8 Kimi K3
Output price /1M$1.98 (off-peak)$25.00
Context window1M
Open weights✓ MIT
Terminal-Bench 2.187.985.088.3
SWE-bench Verified80.688.6
HLE (no tools / tools)42.7 / 60.049.8 / 57.9
NL2Repo61.569.7
Cybergym83.378.3
Toolathlon-Verified74.176.5
DeepSWE62.767.5
DSBench-Hard67.2

Scores from DeepSeek's V4-Pro-0813 release notes and vendor model cards as of Aug–Sep 2026. "—" = not reported on a comparable setting. Claude Fable 5 leads SWE-bench Verified at 95.0 but at $50/1M output tokens. V4.1 Flash benchmarks will be added once DeepSeek publishes them. Deeper dives: DeepSeek vs Claude · DeepSeek vs GPT-5 · all models compared.

Honest Assessment

Strengths & Limitations

DeepSeek is exceptional value - but like all AI systems it has real strengths to leverage and real limitations, including some new regulatory headwinds, to be aware of.

✅ Strengths

💰
Exceptional Cost Efficiency

V4-Pro output costs $1.98/1M off-peak versus $25 for Claude Opus 4.8 - and V4.1 Flash cut Flash-series prices a further 11–60% on 10 Sep 2026.

🤖
Frontier-Class Agentic Coding

87.9 on Terminal-Bench 2.1 and 83.3 on Cybergym beat Opus 4.8; native Responses API and Codex compatibility make it easy to slot into agent frameworks.

🔓
Open Weights (MIT)

Self-host, fine-tune and build commercial products freely. Even retired models like R1 and V3.2 remain downloadable.

📏
1M-Token Context

Entire repositories, long contracts or book-length manuscripts fit in a single prompt, with sparse attention keeping it affordable.

🔌
OpenAI API Compatible

Change the base URL, key and model name to migrate from ChatGPT-based code. Zero rewriting of existing integrations.

⚠️ Limitations

⚖️
Regulatory & Geopolitical Risk

On 9 Sep 2026 the NSA, FBI and CISA accused DeepSeek and other Chinese labs of systematically distilling US models. Sanctions or procurement restrictions are possible; US government and defence-adjacent buyers should check policy before adopting the hosted API.

🔒
Data Privacy Concerns

DeepSeek is a Chinese company; conversations on the hosted service may be stored on servers in China. Use self-hosting or a compliant cloud provider for sensitive or regulated data.

🌍
Censored Topics

The hosted model avoids certain sensitive political topics related to China. Content filters may restrict responses that other models answer freely.

📈
Peak-Hour Pricing & Capacity

API prices double during peak windows (01:00–04:00 and 06:00–10:00 UTC), and heavy demand around launches can cause rate limits or slowdowns on the free tier.

🖼️
Multimodality Still Maturing

Native vision only arrived with V4.1 Flash on 10 Sep 2026 and no benchmark table has been published yet. No audio input or image generation.

FAQ

Frequently Asked Questions

What is new in DeepSeek V4.1 Flash? +

DeepSeek V4.1 Flash was released on 10 September 2026. It is the smallest model in DeepSeek's new architecture series, adds native multimodal (vision) understanding, and uses a much smaller KV cache - DeepSeek says HBM requirements fall to about 1/4 and SSD to 1/8 of the previous generation. DeepSeek claims it beats V4 Pro on performance, cost, speed and total time, and is routing V4 Pro API requests to V4.1 Flash at Flash prices. Flash-series pricing was cut the same day. A full benchmark table had not been published at the time of writing.

Is DeepSeek really free? +

Yes - the web chat and the mobile apps are free, with no subscription tier required. The API is pay-per-token: off-peak, V4-Flash input costs $0.22 per million tokens ($0.007 with cache hits) and output $0.66 per million, and V4.1 Flash is cheaper still. Off-peak rates are 50% of peak. See free access options.

What is a "token" and how much does it cost? +

A token is the smallest unit of text a model processes - roughly 0.75 words (or ~4 characters) in English. The API charges separately for input tokens (what you send) and output tokens (what the model generates). With DeepSeek-V4-Pro off-peak, cache-hit input costs $0.022/1M, cache-miss input $0.66/1M and output $1.98/1M; V4-Flash is $0.007 / $0.22 / $0.66. A typical 500-word exchange costs a fraction of a cent. DeepSeek automatically caches repeated prompts such as system messages.

What happened to DeepSeek-R1, V3.2 and DeepThink mode? +

DeepSeek retired its hosted pre-V4 endpoints - deepseek-chat and deepseek-reasoner, built on V3.2 and R1 - on 24 July 2026. Their role is now filled by the V4 series' built-in thinking modes: choose low, high or max thinking effort instead of toggling DeepThink. The R1 and V3.2 weights remain on Hugging Face under the MIT licence, so you can still self-host them.

Should I use V4-Pro or V4.1 Flash? +

Start with V4.1 Flash. It is the default for chat, extraction, classification, image understanding and most coding tasks, and DeepSeek says it now matches or beats V4 Pro at a lower price - which is why V4 Pro API requests are currently routed to it. Reach for V4-Pro (1.6T parameters, 49B active) when you need the strongest possible results on long agentic runs, hard reasoning or the 384K-token output limit.

Is DeepSeek safe to use for business / sensitive data? +

For non-sensitive work (marketing copy, code review, public data analysis) the cloud API is widely used. However, DeepSeek is a Chinese company - data processed through the public API may be stored on servers subject to Chinese law - and on 9 September 2026 US security agencies (NSA, FBI, CISA) issued an advisory accusing DeepSeek and other Chinese labs of unauthorised model distillation, raising the prospect of sanctions or procurement bans. For sensitive data (healthcare, finance, legal, government) we recommend self-hosting the open weights or using a compliant cloud provider (AWS Bedrock, Azure AI) with SOC2/HIPAA SLAs and data residency guarantees.

How do I switch from OpenAI/ChatGPT to DeepSeek? +

DeepSeek's API is OpenAI-compatible. Set base_url="https://api.deepseek.com" and your DeepSeek api_key, then change the model name to deepseek-v4-flash or deepseek-v4-pro. Streaming, function calling and structured outputs work unchanged. If you use the Codex CLI or the Responses API, DeepSeek V4 supports that natively too. Most teams migrate in minutes - see the API guide.

Can I run DeepSeek locally on my computer? +

Yes. DeepSeek publishes weights under the MIT License. For a laptop or single GPU, the easiest route is Ollama or LM Studio with a distilled or quantized model (1.5B–70B, 8–48GB VRAM). The full V4-Flash (284B total, 13B active) fits on a single 8-GPU server; V4-Pro (1.6T) needs a multi-node cluster. See our download and self-hosting guide.

Does DeepSeek support function calling / tool use? +

Yes. The V4 series supports OpenAI-compatible function calling, structured JSON output and multi-turn agent workflows, plus a native Responses API adapted for Codex. V4-Pro-0813 scores 74.1 on Toolathlon-Verified and 87.9 on Terminal-Bench 2.1, making it one of the strongest open-weight models for building AI agents.

When is the next DeepSeek model coming? +

DeepSeek has shipped a new checkpoint roughly every two months in 2026 (V4-Flash 0731, V4-Pro 0813, V4.1 Flash on 10 Sep). Based on that cadence, the next V4-Pro checkpoint or a V4.1 Pro is expected around mid-October 2026. We update this page as soon as DeepSeek's change log confirms a release.

Who made DeepSeek? Is it Chinese? +

DeepSeek was founded in 2023 as a subsidiary of High-Flyer, a Chinese quantitative hedge fund, and is headquartered in Hangzhou. It became globally prominent in January 2025 when DeepSeek-R1 matched frontier reasoning performance at a reported fraction of the training cost. The lab publishes its research openly and releases model weights under the MIT License. Read more in What is DeepSeek AI?

Get Started

Ready to go deeper?

Try DeepSeek V4.1 Flash and V4-Pro free on the web, or build with an API that is open, affordable and remarkably capable.

Start Chatting Free → Get API Key GitHub ↗