DEPLOY

AI chip · NVIDIA

NVIDIA H100 Tensor Core GPU

Hopper-architecture data-center AI GPU (SXM5 / PCIe). Dominant training + inference accelerator 2023-2024; 80 GB HBM3, ~700W TDP. Fabricated by TSMC on the 4N process.

Deployed in 7 named data centers · Part of 2 rack designs.

Market position

H100 was the chip the 2023-2024 AI-training build-out ran on. Meta bought 350,000. Microsoft, xAI, Anthropic, OpenAI, Oracle, CoreWeave and every neocloud filled DCs with it. Its NVLink 4 + HBM3 combo set the training-cluster template through 2024 until H200 refreshed the memory and Blackwell (B200/GB200) took the crown.

Efficiency and power

How much work you get per watt and per dollar, and how much power a full rack draws.

Perf per watt
2.83 FP8 TFLOPS/W
1,979 TFLOPS ÷ 700 W = 2.83
List $ per FP8 TFLOP
$15.16
$30,000 ÷ 1,979 TFLOPS = $15.16
Rack power
10.2 kW per NVIDIA HGX H100 (8-GPU baseboard) (~1275 W per chip)
10.2 kW ÷ 8 chips (whole-system, includes CPU, memory, NICs, cooling)
Cooling
air
Depends on the rack design

What fits in 80 GB

Which open-source LLMs run on one NVIDIA H100 Tensor Core GPU, by precision. Weights only: add roughly 15% headroom for real serving. If a model does not fit at FP16, try INT8 or INT4 (smaller quality trade-off than most people expect).

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✗
Llama 3.3 70B
dense
70 B140 GB ✗70 GB ✓35 GB ✓
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✗72 GB ✓36 GB ✓
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✗47 GB ✓23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✗141 GB ✗71 GB ✓
Gemma 2 27B
dense
27 B54 GB ✓27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✗104 GB ✗52 GB ✓
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Export controls

US Commerce Department rules that restrict where NVIDIA H100 Tensor Core GPU can be sold. Effective dates and full text below, with the official Federal Register filings and news coverage for each.

  1. The US Bureau of Industry and Security's Oct 7 2022 interim final rule restricted export of advanced AI accelerators to China, keyed to a performance threshold (>=4,800 TOPS or >=600 GB/s interconnect). Directly blocked NVIDIA A100 and H100 sales to Chinese entities without a license, triggering the A800 and H800 China-specific variants NVIDIA released within weeks.

    Sources: BIS press release, Oct 7 2022 · Federal Register, Oct 13 2022 (rule text)

  2. The Oct 17 2023 update closed the workaround NVIDIA used with A800 and H800: BIS replaced the interconnect-bandwidth threshold with a performance-density metric (total processing performance x performance density) and added a Notified Advanced Computing regime for chips that fell just below the ceiling. Direct effect: A800 and H800 could no longer be exported to China. Triggered NVIDIA's H20 as the next-generation China-specific variant.

    Sources: BIS press release, Oct 17 2023 · Federal Register, Oct 25 2023 (rule text)

  3. The Biden administration's Jan 13 2025 Framework for Artificial Intelligence Diffusion introduced a three-tier country regime for advanced GPU exports: unrestricted allies (Tier 1), per-country compute caps (Tier 2), and blocked destinations (Tier 3, including China, Russia, Iran). Set per-entity National Validated End User caps measured in compute (TFLOPs), and required licensing for large orders even to Tier 2 allies. Rescinded May 15 2025 by the Trump administration before full effect.

    Sources: Federal Register, Jan 15 2025 (framework text) · BIS AI Diffusion Framework page

Common questions

How much does NVIDIA H100 Tensor Core GPU cost?

Launch list price was $30,000 when it shipped in 2022-10-13. Street prices swing with supply and how old the generation is. Cloud rental rates vary widely by provider. See list price per TFLOP for the full ranking.

How much memory does NVIDIA H100 Tensor Core GPU have?

80 GB of HBM3, running at 3,350 GB/s. Straight from the vendor datasheet. See chips with the most memory for context.

How much power does one NVIDIA H100 Tensor Core GPU draw?

700 W at the chip. A full server draws more once you add CPU, memory, networking and cooling: see the rack power number above. Compare to other chips on perf-per-watt.

Which data centers use NVIDIA H100 Tensor Core GPU?

7 named data centers run them, including Stargate Norway (Kvandal / Narvik), Compass Datacenters Red Oak Campus (DFW III) and ByteDance Pecem, and 4 more. See who has the most NVIDIA H100 Tensor Core GPU for the ranked list.

Which open-source LLMs fit on one NVIDIA H100 Tensor Core GPU?

Of 13 open-source LLMs we track, 3 fit at FP16, 7 at INT8, and 9 at INT4 (for example Llama 3.1 8B, Llama 3.1 70B and Llama 3.3 70B). Full table above with each model's memory need.

Is NVIDIA H100 Tensor Core GPU subject to US export controls?

Yes. 3 US Commerce Department rules restrict where this chip can be sold, most recently BIS AI Diffusion Framework (Jan 13 2025) (2025-01-13). Details and effective dates in the export-control section above.

See every answer we publish →

Key facts

Class
Compute SoC (AI accelerator)
Designer
NVIDIA
Safety-critical?
No (data-center inference / training)
Record as of
2026-09-27
Most recent source
2025-07-01 (across 11 sources on this page)
Specifications (18 fields, click to expand)

Straight from the vendor datasheet. Dense throughput shown first; sparse (2:4) numbers in parentheses where the vendor publishes them. Full datasheet linked below.

Process node
TSMC 4N
Transistors
80 B
Die size
814 mm²
CUDA cores
16,896
Tensor cores
528 (4th gen (Hopper))
TDP
700 W
Memory
80 GB HBM3
Memory bandwidth
3,350 GB/s
PCIe
Gen 5 x16 (900 GB/s)
FP32
67 TFLOPS
TF32 (dense)
495 TFLOPS
FP16 (dense)
989 TFLOPS (1,979 sparse)
FP8 (dense)
1,979 TFLOPS (3,958 sparse)
INT8 (dense)
1,979 TOPS (3,958 sparse)
Form factor
SXM5
Announced
2022-03-22
Released
2022-10-13
Launch price
$30,000 (list)

Source: vendor datasheet

Generation

Foundry & process

Fabbed at
TSMC (all chips TSMC makes →)
Process node
TSMC 4N
Transistors
80 B on 814 mm² die

Rack designs that use it

Whole-rack systems buyers order in bulk (NVL72, HGX, TPU pods) that ship with NVIDIA H100 Tensor Core GPU inside.

See all rack designs →

Compare with

See every chip comparison →

Benchmarks

Published performance numbers by workload. Vendor datasheet figures where noted; independent measurements otherwise.

WorkloadValueUnitSourceAs of
memory bandwidth3350GB/svendor2023-01-01
memory capacity80GB HBM3vendor2023-01-01
peak tflops fp16 dense(SXM5 form factor)989.4TFLOPSvendor2023-01-01
peak tflops fp8 dense1978.9TFLOPSvendor2023-01-01

Data centers running NVIDIA H100 Tensor Core GPU

Who has the most? →

Named data-center campuses with NVIDIA H100 Tensor Core GPU on site. Counts shown where the operator has published them; other rows are described in general terms.

  1. under constructioninferred

    100,000 NVIDIA GPUs targeted by end of 2026 (initial phase; specific SKU may be H100/H200/Blackwell mix)

    OpenAI announcement cited 100k Nvidia GPUs without SKU. Attributed conservatively to H100 as the volume-shipping Hopper baseline; upgrade as G42/Aker discloses.

    OpenAI

  2. partially energizedinferred

    H100/H200-generation hyperscaler capacity at the Red Oak DFW III campus

    Compass DFW III hosts multiple hyperscalers with H100/H200 fleets; per-tenant chip counts not published.

    Compass Datacenters

  3. under constructioninferred

    ByteDance TikTok/Doubao inference capacity in South America (H100/H200-class)

    Chip mix not published; ByteDance's non-China DCs run predominantly NVIDIA H100/H200.

    Reuters

  4. xAI Colossus (Memphis)as of 2024-09-01
    partially energizedreported

    ~100,000 GPUs live in first cluster (Sep 2024); scaled toward 200,000 across the complex

    SemiAnalysis

  5. Crusoe Childress Campusas of 2024-06-01
    announcedreported

    Crusoe operates H100 GPU clusters at Childress

    Crusoe

  6. under constructionreported

    Nvidia complement to MTIA for training + inference at Hyperion

    Meta Engineering

  7. partially energizedreported

    Nvidia complement to MTIA for training + inference at Prometheus

    Meta Engineering

Sources

Adoption rows appear as we confirm each chip-in-robot pairing from a public source. See every chip we track for the full catalog or the NVIDIA page.