DEPLOY

AI chip · NVIDIA

NVIDIA B300 (Blackwell Ultra)

Blackwell Ultra data-center GPU: 288 GB HBM3e at up to 1400W, the higher-memory, higher-NVFP4 successor to the B200. It is the GPU inside the HGX B300 8-GPU board and the GB300 NVL72 rack.

Market position

Blackwell Ultra is NVIDIA's mid-cycle refresh of the B200: 288 GB HBM3e (1.5x B200's 192 GB), 1.5x dense NVFP4 compute (15 vs 10 PFLOPS) and 2x attention throughput, at up to 1400W. Ships in the HGX B300 8-GPU board and the GB300 NVL72 rack (Grace CPUs paired with B300 GPUs).

Efficiency and power

How much work you get per watt and per dollar, and how much power a full rack draws.

Perf per watt
3.57 FP8 TFLOPS/W
5,000 TFLOPS ÷ 1400 W = 3.57

What fits in 288 GB

Which open-source LLMs run on one NVIDIA B300 (Blackwell Ultra), by precision. Weights only: add roughly 15% headroom for real serving. If a model does not fit at FP16, try INT8 or INT4 (smaller quality trade-off than most people expect).

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✓70 GB ✓35 GB ✓
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✓
Llama 3.3 70B
dense
70 B140 GB ✓70 GB ✓35 GB ✓
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✓72 GB ✓36 GB ✓
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✓47 GB ✓23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✓141 GB ✓71 GB ✓
Gemma 2 27B
dense
27 B54 GB ✓27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✓104 GB ✓52 GB ✓
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Common questions

How much does NVIDIA B300 (Blackwell Ultra) cost?

NVIDIA B300 (Blackwell Ultra) doesn't have a public launch list price. Vendors like this one usually price through direct sales rather than a public sheet. See list price per TFLOP for chips that do publish.

How much memory does NVIDIA B300 (Blackwell Ultra) have?

288 GB of HBM3e, running at 8,000 GB/s. Straight from the vendor datasheet. See chips with the most memory for context.

How much power does one NVIDIA B300 (Blackwell Ultra) draw?

1400 W at the chip. A full server draws more once you add CPU, memory, networking and cooling: see the rack power number above. Compare to other chips on perf-per-watt.

Which open-source LLMs fit on one NVIDIA B300 (Blackwell Ultra)?

Of 13 open-source LLMs we track, 9 fit at FP16, 9 at INT8, and 10 at INT4 (for example Llama 3.1 8B, Llama 3.1 70B and Llama 3.1 405B). Full table above with each model's memory need.

See every answer we publish →

Key facts

Class
Compute SoC (AI accelerator)
Designer
NVIDIA
Safety-critical?
No (data-center inference / training)
Record as of
2026-09-27
Specifications (12 fields, click to expand)

Straight from the vendor datasheet. Dense throughput shown first; sparse (2:4) numbers in parentheses where the vendor publishes them. Full datasheet linked below.

Process node
TSMC 4NP
Transistors
208 B
TDP
1400 W
Memory
288 GB HBM3e
Memory bandwidth
8,000 GB/s
PCIe
Gen 6 (1,800 GB/s)
FP16 (dense)
2,500 TFLOPS (5,000 sparse)
FP8 (dense)
5,000 TFLOPS (10,000 sparse)
FP4 (dense)
15,000 TFLOPS (20,000 sparse)
Form factor
SXM6
Announced
2025-03-18
Released
2025-11-01

Source: vendor datasheet

Generation

Foundry & process

Process node
TSMC 4NP
Transistors
208 B

Sources

Adoption rows appear as we confirm each chip-in-robot pairing from a public source. See every chip we track for the full catalog or the NVIDIA page.