DEPLOY

AI chip · NVIDIA

NVIDIA H200 Tensor Core GPU

Hopper refresh with 141 GB HBM3e (from 80 GB); the mid-generation memory-bandwidth bump between H100 and Blackwell.

Deployed in 4 named data centers · Part of 1 rack design.

Market position

H200 kept H100's compute silicon and swapped HBM3 for 141 GB HBM3e at 4.8 TB/s: 1.4x memory bandwidth in the same socket. Inference shops (Together AI, Perplexity, hyperscaler tenants) upgraded here first; training shops mostly held for Blackwell.

Efficiency and power

How much work you get per watt and per dollar, and how much power a full rack draws.

Perf per watt
2.83 FP8 TFLOPS/W
1,979 TFLOPS ÷ 700 W = 2.83
Rack power
10.2 kW per NVIDIA HGX H200 (8-GPU baseboard) (~1275 W per chip)
10.2 kW ÷ 8 chips (whole-system, includes CPU, memory, NICs, cooling)
Cooling
air
Depends on the rack design

What fits in 141 GB

Which open-source LLMs run on one NVIDIA H200 Tensor Core GPU, by precision. Weights only: add roughly 15% headroom for real serving. If a model does not fit at FP16, try INT8 or INT4 (smaller quality trade-off than most people expect).

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✓70 GB ✓35 GB ✓
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✗
Llama 3.3 70B
dense
70 B140 GB ✓70 GB ✓35 GB ✓
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✗72 GB ✓36 GB ✓
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✓47 GB ✓23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✗141 GB ✓71 GB ✓
Gemma 2 27B
dense
27 B54 GB ✓27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✗104 GB ✓52 GB ✓
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Common questions

How much does NVIDIA H200 Tensor Core GPU cost?

NVIDIA H200 Tensor Core GPU doesn't have a public launch list price. Vendors like this one usually price through direct sales rather than a public sheet. See list price per TFLOP for chips that do publish.

How much memory does NVIDIA H200 Tensor Core GPU have?

141 GB of HBM3e, running at 4,800 GB/s. Straight from the vendor datasheet. See chips with the most memory for context.

How much power does one NVIDIA H200 Tensor Core GPU draw?

700 W at the chip. A full server draws more once you add CPU, memory, networking and cooling: see the rack power number above. Compare to other chips on perf-per-watt.

Which data centers use NVIDIA H200 Tensor Core GPU?

4 named data centers run them, including xAI Colossus (Memphis), Khazna QAJ1 Ajman AI-Optimized and Lake Mariner Data Campus, and 1 more. See who has the most NVIDIA H200 Tensor Core GPU for the ranked list.

Which open-source LLMs fit on one NVIDIA H200 Tensor Core GPU?

Of 13 open-source LLMs we track, 6 fit at FP16, 9 at INT8, and 9 at INT4 (for example Llama 3.1 8B, Llama 3.1 70B and Llama 3.3 70B). Full table above with each model's memory need.

See every answer we publish →

Key facts

Class
Compute SoC (AI accelerator)
Designer
NVIDIA
Safety-critical?
No (data-center inference / training)
Record as of
2026-09-27
Most recent source
2026-01-15 (across 6 sources on this page)
Specifications (17 fields, click to expand)

Straight from the vendor datasheet. Dense throughput shown first; sparse (2:4) numbers in parentheses where the vendor publishes them. Full datasheet linked below.

Process node
TSMC 4N
Transistors
80 B
Die size
814 mm²
CUDA cores
16,896
Tensor cores
528 (4th gen (Hopper))
TDP
700 W
Memory
141 GB HBM3e
Memory bandwidth
4,800 GB/s
PCIe
Gen 5 x16 (900 GB/s)
FP32
67 TFLOPS
TF32 (dense)
495 TFLOPS
FP16 (dense)
989 TFLOPS (1,979 sparse)
FP8 (dense)
1,979 TFLOPS (3,958 sparse)
INT8 (dense)
1,979 TOPS (3,958 sparse)
Form factor
SXM5
Announced
2023-11-13
Released
2024-03-01

Source: vendor datasheet

Generation

Foundry & process

Fabbed at
TSMC (all chips TSMC makes →)
Process node
TSMC 4N
Transistors
80 B on 814 mm² die

Rack designs that use it

Whole-rack systems buyers order in bulk (NVL72, HGX, TPU pods) that ship with NVIDIA H200 Tensor Core GPU inside.

See all rack designs →

Compare with

See every chip comparison →

Benchmarks

Published performance numbers by workload. Vendor datasheet figures where noted; independent measurements otherwise.

WorkloadValueUnitSourceAs of
memory bandwidth4800GB/svendor2024-01-01
memory capacity141GB HBM3evendor2024-01-01

Data centers running NVIDIA H200 Tensor Core GPU

Who has the most? →

Named data-center campuses with NVIDIA H200 Tensor Core GPU on site. Counts shown where the operator has published them; other rows are described in general terms.

  1. xAI Colossus (Memphis)as of 2026-01-15
    partially energizedreported

    Second-wave upgrade GPUs added late 2024

    Introl

  2. under constructioninferred

    H200/GB200-generation capacity for G42 / UAE AI stack

    G42 has committed multi-billion NVIDIA compute; per-campus attribution not published.

    Khazna

  3. Lake Mariner Data Campusas of 2025-01-01
    partially energizedinferred

    H200-generation HPC capacity for AI-training tenants

    AI-training-oriented HPC campus; specific chip mix not published, H200 inferred from current-gen positioning.

    TeraWulf

  4. Crusoe Childress Campusas of 2024-12-01
    announcedreported

    Recent Crusoe expansions include H200 capacity

    Crusoe

Sources

Adoption rows appear as we confirm each chip-in-robot pairing from a public source. See every chip we track for the full catalog or the NVIDIA page.