Open-weight LLM comparison
Command R+ vs Qwen 2.5 72B
Side-by-side
Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.
| Field | Command R+ | Qwen 2.5 72B |
|---|---|---|
| Total parameters | 104 B | 72 B |
| Architecture | dense | dense |
| Vendor | Cohere | Alibaba |
| License | CC-BY-NC-4.0 (non-commercial) | Qwen License (72B only) |
| Released | 2024-04-04 | 2024-09-19 |
| Weights @ FP16 | 208 GB | 144 GB |
| Weights @ INT8 | 104 GB | 72 GB |
| Weights @ INT4 | 52 GB | 36 GB |
Which is smarter (published benchmarks)
Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.
| Benchmark | Command R+ | Qwen 2.5 72B |
|---|---|---|
| MMLU (5-shot) | 75.7 | 86.1 |
| HumanEval | 70.7 | 86.6 |
| MATH (0-shot) | not published | 83.1 |
| GPQA | not published | 49.0 |
Sources: Command R+ model card · Qwen 2.5 72B model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.
Common questions
Command R+ vs Qwen 2.5 72B: which is bigger?
Command R+ has more parameters (Command R+: 104 B; Qwen 2.5 72B: 72 B). More parameters usually means higher ceiling on capability and higher memory requirement, though MoE architectures decouple total parameters from per-token compute.
Command R+ vs Qwen 2.5 72B: which is newer?
Qwen 2.5 72B released 2024-09-19; Command R+ released 2024-04-04.
Command R+ vs Qwen 2.5 72B: which needs less memory to serve?
Qwen 2.5 72B needs less HBM. Weights-only footprint at FP16: Command R+ 208 GB; Qwen 2.5 72B 144 GB. Half those numbers at INT8, quarter at INT4. Real serving adds 10-30% for KV cache.
Command R+ vs Qwen 2.5 72B: which license is more permissive?
Command R+: CC-BY-NC-4.0 (non-commercial). Qwen 2.5 72B: Qwen License (72B only). Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.
See also: every LLM comparison · Command R+ full page · Qwen 2.5 72B full page.