DEPLOY

Open-weight LLM comparison

DeepSeek R1 vs Qwen 2.5 7B

Side-by-side

Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.

FieldDeepSeek R1Qwen 2.5 7B
Total parameters671 B7 B
Active parameters (MoE)37 B—
ArchitectureMoE (671B total, 37B active per token)dense
VendorDeepSeekAlibaba
LicenseMITApache 2.0
Released2025-01-202024-09-19
Weights @ FP161342 GB14 GB
Weights @ INT8671 GB7 GB
Weights @ INT4336 GB4 GB

Which is smarter (published benchmarks)

Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.

BenchmarkDeepSeek R1Qwen 2.5 7B
MMLU (5-shot)90.874.2
MATH (AIME 2024, pass@1)79.8not published
GPQA (Diamond)71.5not published
Codeforces (Elo)2029.0not published
HumanEvalnot published84.8
MATH (0-shot)not published75.5

Sources: DeepSeek R1 model card · Qwen 2.5 7B model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.

Common questions

DeepSeek R1 vs Qwen 2.5 7B: which is bigger?

DeepSeek R1 has more parameters (DeepSeek R1: 671 B; Qwen 2.5 7B: 7 B). More parameters usually means higher ceiling on capability and higher memory requirement, though MoE architectures decouple total parameters from per-token compute.

DeepSeek R1 vs Qwen 2.5 7B: which is newer?

DeepSeek R1 released 2025-01-20; Qwen 2.5 7B released 2024-09-19.

DeepSeek R1 vs Qwen 2.5 7B: which needs less memory to serve?

Qwen 2.5 7B needs less HBM. Weights-only footprint at FP16: DeepSeek R1 1342 GB; Qwen 2.5 7B 14 GB. Half those numbers at INT8, quarter at INT4. Real serving adds 10-30% for KV cache.

DeepSeek R1 vs Qwen 2.5 7B: which license is more permissive?

DeepSeek R1: MIT. Qwen 2.5 7B: Apache 2.0. Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.

See also: every LLM comparison · DeepSeek R1 full page · Qwen 2.5 7B full page.