AI chip · NVIDIA
NVIDIA H200 Tensor Core GPU
Hopper refresh with 141 GB HBM3e (from 80 GB); the mid-generation memory-bandwidth bump between H100 and Blackwell.
Deployed in 4 named data centers · Part of 1 rack design.
Market position
H200 kept H100's compute silicon and swapped HBM3 for 141 GB HBM3e at 4.8 TB/s: 1.4x memory bandwidth in the same socket. Inference shops (Together AI, Perplexity, hyperscaler tenants) upgraded here first; training shops mostly held for Blackwell.
Efficiency and power
How much work you get per watt and per dollar, and how much power a full rack draws.
- Perf per watt
- 2.83 FP8 TFLOPS/W1,979 TFLOPS ÷ 700 W = 2.83
- Rack power
- 10.2 kW per NVIDIA HGX H200 (8-GPU baseboard) (~1275 W per chip)10.2 kW ÷ 8 chips (whole-system, includes CPU, memory, NICs, cooling)
- Cooling
- airDepends on the rack design
What fits in 141 GB
Which open-source LLMs run on one NVIDIA H200 Tensor Core GPU, by precision. Weights only: add roughly 15% headroom for real serving. If a model does not fit at FP16, try INT8 or INT4 (smaller quality trade-off than most people expect).
| Model | Params | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| Llama 3.1 8B dense | 8 B | 16 GB ✓ | 8 GB ✓ | 4 GB ✓ |
| Llama 3.1 70B dense | 70 B | 140 GB ✓ | 70 GB ✓ | 35 GB ✓ |
| Llama 3.1 405B dense | 405 B | 810 GB ✗ | 405 GB ✗ | 203 GB ✗ |
| Llama 3.3 70B dense | 70 B | 140 GB ✓ | 70 GB ✓ | 35 GB ✓ |
| DeepSeek V3 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| DeepSeek R1 MoE (671B total, 37B active per token) | 671 B | 1342 GB ✗ | 671 GB ✗ | 336 GB ✗ |
| Qwen 2.5 7B dense | 7 B | 14 GB ✓ | 7 GB ✓ | 4 GB ✓ |
| Qwen 2.5 72B dense | 72 B | 144 GB ✗ | 72 GB ✓ | 36 GB ✓ |
| Mixtral 8x7B MoE (46.7B total, 12.9B active per token) | 46.7 B | 93 GB ✓ | 47 GB ✓ | 23 GB ✓ |
| Mixtral 8x22B MoE (141B total, 39B active per token) | 141 B | 282 GB ✗ | 141 GB ✓ | 71 GB ✓ |
| Gemma 2 27B dense | 27 B | 54 GB ✓ | 27 GB ✓ | 14 GB ✓ |
| Command R+ dense | 104 B | 208 GB ✗ | 104 GB ✓ | 52 GB ✓ |
| Kimi K2 MoE (1T total, 32B active per token) | 1000 B | 2000 GB ✗ | 1000 GB ✗ | 500 GB ✗ |
Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.
Common questions
How much does NVIDIA H200 Tensor Core GPU cost?
NVIDIA H200 Tensor Core GPU doesn't have a public launch list price. Vendors like this one usually price through direct sales rather than a public sheet. See list price per TFLOP for chips that do publish.
How much memory does NVIDIA H200 Tensor Core GPU have?
141 GB of HBM3e, running at 4,800 GB/s. Straight from the vendor datasheet. See chips with the most memory for context.
How much power does one NVIDIA H200 Tensor Core GPU draw?
700 W at the chip. A full server draws more once you add CPU, memory, networking and cooling: see the rack power number above. Compare to other chips on perf-per-watt.
Which data centers use NVIDIA H200 Tensor Core GPU?
4 named data centers run them, including xAI Colossus (Memphis), Khazna QAJ1 Ajman AI-Optimized and Lake Mariner Data Campus, and 1 more. See who has the most NVIDIA H200 Tensor Core GPU for the ranked list.
Which open-source LLMs fit on one NVIDIA H200 Tensor Core GPU?
Of 13 open-source LLMs we track, 6 fit at FP16, 9 at INT8, and 9 at INT4 (for example Llama 3.1 8B, Llama 3.1 70B and Llama 3.3 70B). Full table above with each model's memory need.
Key facts
- Class
- Compute SoC (AI accelerator)
- Designer
- NVIDIA
- Safety-critical?
- No (data-center inference / training)
- Record as of
- 2026-09-27
- Most recent source
- 2026-01-15 (across 6 sources on this page)
Specifications (17 fields, click to expand)
Straight from the vendor datasheet. Dense throughput shown first; sparse (2:4) numbers in parentheses where the vendor publishes them. Full datasheet linked below.
- Process node
- TSMC 4N
- Transistors
- 80 B
- Die size
- 814 mm²
- CUDA cores
- 16,896
- Tensor cores
- 528 (4th gen (Hopper))
- TDP
- 700 W
- Memory
- 141 GB HBM3e
- Memory bandwidth
- 4,800 GB/s
- PCIe
- Gen 5 x16 (900 GB/s)
- FP32
- 67 TFLOPS
- TF32 (dense)
- 495 TFLOPS
- FP16 (dense)
- 989 TFLOPS (1,979 sparse)
- FP8 (dense)
- 1,979 TFLOPS (3,958 sparse)
- INT8 (dense)
- 1,979 TOPS (3,958 sparse)
- Form factor
- SXM5
- Announced
- 2023-11-13
- Released
- 2024-03-01
Source: vendor datasheet
Generation
Foundry & process
- Fabbed at
- TSMC (all chips TSMC makes →)
- Process node
- TSMC 4N
- Transistors
- 80 B on 814 mm² die
Rack designs that use it
Whole-rack systems buyers order in bulk (NVL72, HGX, TPU pods) that ship with NVIDIA H200 Tensor Core GPU inside.
- ×8NVIDIA HGX H200 (8-GPU baseboard)platform
Compare with
Benchmarks
Published performance numbers by workload. Vendor datasheet figures where noted; independent measurements otherwise.
Data centers running NVIDIA H200 Tensor Core GPU
Who has the most? →Named data-center campuses with NVIDIA H200 Tensor Core GPU on site. Counts shown where the operator has published them; other rows are described in general terms.
- xAI Colossus (Memphis)as of 2026-01-15partially energizedreported
Second-wave upgrade GPUs added late 2024
- Khazna QAJ1 Ajman AI-Optimized Data Centeras of 2025-03-01under constructioninferred
H200/GB200-generation capacity for G42 / UAE AI stack
G42 has committed multi-billion NVIDIA compute; per-campus attribution not published.
- Lake Mariner Data Campusas of 2025-01-01partially energizedinferred
H200-generation HPC capacity for AI-training tenants
AI-training-oriented HPC campus; specific chip mix not published, H200 inferred from current-gen positioning.
- Crusoe Childress Campusas of 2024-12-01announcedreported
Recent Crusoe expansions include H200 capacity
Sources
- NVIDIA H200 product pageNvidia
Adoption rows appear as we confirm each chip-in-robot pairing from a public source. See every chip we track for the full catalog or the NVIDIA page.