Open-weight LLM comparison
Kimi K2 vs Qwen 2.5 72B
Side-by-side
Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.
| Field | Kimi K2 | Qwen 2.5 72B |
|---|---|---|
| Total parameters | 1000 B | 72 B |
| Active parameters (MoE) | 32 B | — |
| Architecture | MoE (1T total, 32B active per token) | dense |
| Vendor | Moonshot AI | Alibaba |
| License | Modified MIT | Qwen License (72B only) |
| Released | 2025-07-11 | 2024-09-19 |
| Weights @ FP16 | 2000 GB | 144 GB |
| Weights @ INT8 | 1000 GB | 72 GB |
| Weights @ INT4 | 500 GB | 36 GB |
Which is smarter (published benchmarks)
Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.
| Benchmark | Kimi K2 | Qwen 2.5 72B |
|---|---|---|
| MMLU (5-shot) | 89.5 | 86.1 |
| HumanEval | 85.7 | 86.6 |
| MATH (0-shot) | 82.5 | 83.1 |
| GPQA | not published | 49.0 |
Sources: Kimi K2 model card · Qwen 2.5 72B model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.
Common questions
Kimi K2 vs Qwen 2.5 72B: which is bigger?
Kimi K2 has more parameters (Kimi K2: 1000 B; Qwen 2.5 72B: 72 B). More parameters usually means higher ceiling on capability and higher memory requirement, though MoE architectures decouple total parameters from per-token compute.
Kimi K2 vs Qwen 2.5 72B: which is newer?
Kimi K2 released 2025-07-11; Qwen 2.5 72B released 2024-09-19.
Kimi K2 vs Qwen 2.5 72B: which needs less memory to serve?
Qwen 2.5 72B needs less HBM. Weights-only footprint at FP16: Kimi K2 2000 GB; Qwen 2.5 72B 144 GB. Half those numbers at INT8, quarter at INT4. Real serving adds 10-30% for KV cache.
Kimi K2 vs Qwen 2.5 72B: which license is more permissive?
Kimi K2: Modified MIT. Qwen 2.5 72B: Qwen License (72B only). Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.
See also: every LLM comparison · Kimi K2 full page · Qwen 2.5 72B full page.