Open-weight LLM comparison
DeepSeek R1 vs DeepSeek V3
Side-by-side
Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.
| Field | DeepSeek R1 | DeepSeek V3 |
|---|---|---|
| Total parameters | 671 B | 671 B |
| Active parameters (MoE) | 37 B | 37 B |
| Architecture | MoE (671B total, 37B active per token) | MoE (671B total, 37B active per token) |
| Vendor | DeepSeek | DeepSeek |
| License | MIT | MIT |
| Released | 2025-01-20 | 2024-12-26 |
| Weights @ FP16 | 1342 GB | 1342 GB |
| Weights @ INT8 | 671 GB | 671 GB |
| Weights @ INT4 | 336 GB | 336 GB |
Which is smarter (published benchmarks)
Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.
| Benchmark | DeepSeek R1 | DeepSeek V3 |
|---|---|---|
| MMLU (5-shot) | 90.8 | 88.5 |
| MATH (AIME 2024, pass@1) | 79.8 | not published |
| GPQA (Diamond) | 71.5 | not published |
| Codeforces (Elo) | 2029.0 | not published |
| HumanEval | not published | 82.6 |
| MATH (0-shot) | not published | 61.6 |
| GPQA | not published | 59.1 |
Sources: DeepSeek R1 model card · DeepSeek V3 model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.
Common questions
DeepSeek R1 vs DeepSeek V3: which is newer?
DeepSeek R1 released 2025-01-20; DeepSeek V3 released 2024-12-26.
DeepSeek R1 vs DeepSeek V3: which license is more permissive?
DeepSeek R1: MIT. DeepSeek V3: MIT. Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.
See also: every LLM comparison · DeepSeek R1 full page · DeepSeek V3 full page.