DEPLOY

Open-weight LLM comparison

Mixtral 8x22B vs Qwen 2.5 72B

Side-by-side

Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.

FieldMixtral 8x22BQwen 2.5 72B
Total parameters141 B72 B
Active parameters (MoE)39 B—
ArchitectureMoE (141B total, 39B active per token)dense
VendorMistral AIAlibaba
LicenseApache 2.0Qwen License (72B only)
Released2024-04-172024-09-19
Weights @ FP16282 GB144 GB
Weights @ INT8141 GB72 GB
Weights @ INT471 GB36 GB

Which is smarter (published benchmarks)

Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.

BenchmarkMixtral 8x22BQwen 2.5 72B
MMLU (5-shot)77.886.1
HumanEval76.286.6
MATH (maj@4)41.8not published
MATH (0-shot)not published83.1
GPQAnot published49.0

Sources: Mixtral 8x22B model card · Qwen 2.5 72B model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.

Common questions

Mixtral 8x22B vs Qwen 2.5 72B: which is bigger?

Mixtral 8x22B has more parameters (Mixtral 8x22B: 141 B; Qwen 2.5 72B: 72 B). More parameters usually means higher ceiling on capability and higher memory requirement, though MoE architectures decouple total parameters from per-token compute.

Mixtral 8x22B vs Qwen 2.5 72B: which is newer?

Qwen 2.5 72B released 2024-09-19; Mixtral 8x22B released 2024-04-17.

Mixtral 8x22B vs Qwen 2.5 72B: which needs less memory to serve?

Qwen 2.5 72B needs less HBM. Weights-only footprint at FP16: Mixtral 8x22B 282 GB; Qwen 2.5 72B 144 GB. Half those numbers at INT8, quarter at INT4. Real serving adds 10-30% for KV cache.

Mixtral 8x22B vs Qwen 2.5 72B: which license is more permissive?

Mixtral 8x22B: Apache 2.0. Qwen 2.5 72B: Qwen License (72B only). Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.

See also: every LLM comparison · Mixtral 8x22B full page · Qwen 2.5 72B full page.