โ† Back to ConceptsVerified: 2026-07-30
DeepSeek ArchitectureDeepSeek: the open-weight lab that made frontier AI 10x cheaper and MIT-licensed.

DeepSeek Architecture

In January 2025, DeepSeek-R1 matched OpenAI o1 on reasoning benchmarks at 1/20th the training cost and triggered what Marc Andreessen called "AI's Sputnik moment." By July 2026, V4 Pro ships at 34x less than GPT-5.5 per token - with 1M context, open weights, and an MIT license. This is the full picture.

V3Nov 2024671B MoER1Jan 2025Sputnik momentV3.2Mar 2026128K contextV4 FlashApr 2026284B MoE ยท 1M ctx$0.14/$0.28V4 ProApr 20261.6T MoE ยท 1M ctx$0.44/$0.87

Five model generations in 18 months. The V4 launch on April 24, 2026 consolidated the entire lineup into two models: Flash for throughput and Pro for reasoning. Legacy aliases (deepseek-chat, deepseek-reasoner) retire July 24, 2026.

01What It Is

A Chinese AI research lab that ships open-weight models under MIT license and prices them to move.

DeepSeek is a Chinese AI lab, founded in 2023, that gives its models away and still undercuts everyone on price. Every model from V3 forward is MIT licensed - download the weights, fine-tune them, host them yourself. Its Mixture-of-Experts design activates only a fraction of total parameters per token, which is a big part of why its API pricing beats US providers by 10-34x.

02The Models

Two V4 models replaced the entire lineup.

Everything consolidated into V4 Flash and V4 Pro on April 24, 2026. R1 and V3.2 are historical references. The legacy API aliases retire July 24, 2026.

V4 Flash

The Workhorse

Params284B total / 13B active
Context1M tokens
Max Output384K tokens
Price$0.14 / $0.28 per 1M
Cache Hit$0.0028 per 1M

Default for production workloads: chat, RAG, extraction, classification, translation, tool-calling. Handles the vast majority of tasks at 4-8x less cost than Pro. Supports thinking mode for reasoning tasks.

V4 Pro

The Heavy Lifter

Params1.6T total / 49B active
Context1M tokens
Max Output384K tokens
Price$0.435 / $0.87 per 1M
Cache Hit$0.0036 per 1M

Escalate when Flash fails: multi-step reasoning, competitive math, code-golf problems, long-horizon agent planning, heavy self-debugging coding sessions. Codeforces rating 3,206 - open-source SOTA in agentic coding.

03The Cost Gap

DeepSeek is 10-34x cheaper than the US frontier.

At 10 million output tokens per month - a moderate enterprise workload - V4 Pro costs $8,700/month. GPT-5.5 costs $300,000/month for the same volume.

ModelInput / 1MOutput / 1M
DeepSeek V4 Flash$0.14$0.28
DeepSeek V4 Pro$0.435$0.87
Gemini 3.1 Pro$1.25$5.00
GPT-5.4$1.75$8.00
GPT-5.5$5.00$15.00

All prices per 1M tokens, USD. DeepSeek pricing verified July 2026. Cache hit pricing (1/10th of standard input) not shown - available on both V4 models. May 2026: V4 Pro discount made permanent at 75% off standard rates.

04Where to Run It

DeepSeek runs everywhere - API, cloud, on-prem, even your laptop.

API

Official

api.deepseek.com with OpenAI-compatible endpoints. Free chat at chat.deepseek.com. Both V4 models available. Concurrency: 2,500 (Flash), 500 (Pro). Routes through China-based infrastructure - consider latency for US/EU workloads.

Cloud

Managed

Available on Amazon Bedrock, Microsoft Azure AI Foundry, and Google Vertex AI. Also on data platforms: Snowflake Cortex AI and Databricks. OpenRouter provides multi-provider access with unified billing.

Self-Hosted

Open Weights

MIT licensed: download weights from Hugging Face, deploy on your own infrastructure. Day-zero support from vLLM, SGLang, and Hugging Face TGI. Distilled versions (1.5B to 70B) run on consumer hardware via Ollama or LM Studio.

05Tradeoffs

The price is real. So are the tradeoffs.

Consider

What you get

The same reasoning and coding quality as GPT-5.5, for a fraction of the price, with the freedom to self-host if you want it.

Watch for

What you trade

The official API routes through China, adding latency for US/EU users and raising data residency questions. The surrounding ecosystem - SDKs, plugins, native multimodal - is thinner than OpenAI's or Anthropic's. Benchmark it on your own workload before switching.

"The price gap is not a promotion. It is a structural advantage - open weights, MoE efficiency, and a research lab that trains frontier models for 1/20th the cost."

01

Default to V4 Flash. Escalate to V4 Pro only when Flash measurably fails on a specific task. Flash handles 80%+ of production workloads at 4-8x less cost.

02

Migrate off deepseek-chat and deepseek-reasoner aliases before July 24, 2026. The old names return errors after that date. One line of code: swap to deepseek-v4-flash.