DEPLOY

Open-weight LLM comparison

Llama 3.1 405B vs Mixtral 8x22B

Side-by-side

Straight from each model's release page and Hugging Face card. Memory shown is for the weights alone: add roughly 15% for real serving. Bold column marks the more parameters and the smaller memory footprint.

FieldLlama 3.1 405BMixtral 8x22B
Total parameters405 B141 B
Active parameters (MoE)—39 B
ArchitecturedenseMoE (141B total, 39B active per token)
VendorMetaMistral AI
LicenseLlama 3.1 Community LicenseApache 2.0
Released2024-07-232024-04-17
Weights @ FP16810 GB282 GB
Weights @ INT8405 GB141 GB
Weights @ INT4203 GB71 GB

Which is smarter (published benchmarks)

Vendor-reported quality scores on the standard leaderboards. Bold column marks the higher score on the same test.

BenchmarkLlama 3.1 405BMixtral 8x22B
MMLU (5-shot)88.677.8
HumanEval89.076.2
MATH (0-shot)73.8not published
GPQA51.1not published
MATH (maj@4)not published41.8

Sources: Llama 3.1 405B model card · Mixtral 8x22B model card. Benchmark methodology and prompt template can shift these numbers by several points, so treat these as relative rankings, not absolute scores.

Common questions

Llama 3.1 405B vs Mixtral 8x22B: which is bigger?

Llama 3.1 405B has more parameters (Llama 3.1 405B: 405 B; Mixtral 8x22B: 141 B). More parameters usually means higher ceiling on capability and higher memory requirement, though MoE architectures decouple total parameters from per-token compute.

Llama 3.1 405B vs Mixtral 8x22B: which is newer?

Llama 3.1 405B released 2024-07-23; Mixtral 8x22B released 2024-04-17.

Llama 3.1 405B vs Mixtral 8x22B: which needs less memory to serve?

Mixtral 8x22B needs less HBM. Weights-only footprint at FP16: Llama 3.1 405B 810 GB; Mixtral 8x22B 282 GB. Half those numbers at INT8, quarter at INT4. Real serving adds 10-30% for KV cache.

Llama 3.1 405B vs Mixtral 8x22B: which license is more permissive?

Llama 3.1 405B: Llama 3.1 Community License. Mixtral 8x22B: Apache 2.0. Apache 2.0 and MIT allow unrestricted commercial use; Llama Community License allows commercial use but restricts training larger models on outputs; CC-BY-NC and vendor-specific licenses (Qwen 72B, Gemma) have narrower terms. Check the model card for the exact clauses.

See also: every LLM comparison · Llama 3.1 405B full page · Mixtral 8x22B full page.