DEPLOY

Chip leaderboard

Best perf-per-watt (TFLOPS/W)

NVIDIA Jetson AGX Thor leads at 15.9 TFLOPS/W. 43 chips ranked.

How much work each chip does per watt. Straight from the vendor datasheet: best dense throughput divided by rated power. Wafer-scale and rack-scale chips look artificially strong because their datasheet numbers roll in shared overhead the per-chip numbers ignore.

#ChipValueInputs
1NVIDIA Jetson AGX Thor
NVIDIA
15.9 TFLOPS/WFP8 TFLOPS: 2,070 · TDP: 130 W
2NVIDIA DRIVE Thor
NVIDIA
15.4 TFLOPS/WFP8 TFLOPS: 2,000 · TDP: 130 W
3Google TPU v7 (Ironwood)
Google
13.2 TFLOPS/WFP8 TFLOPS: 9,228 · TDP: 700 W
4Google TPU v6 (Trillium)
Google
9.18 TFLOPS/WFP8 TFLOPS: 1,836 · TDP: 200 W
5NVIDIA B100 (Blackwell)
NVIDIA
5 TFLOPS/WFP8 TFLOPS: 3,500 · TDP: 700 W
6AMD Instinct MI350 series
AMD
4.6 TFLOPS/WFP8 TFLOPS: 4,600 · TDP: 1000 W
7NVIDIA B200 (Blackwell)
NVIDIA
4.5 TFLOPS/WFP8 TFLOPS: 4,500 · TDP: 1000 W
8AMD Instinct MI355X
AMD
3.57 TFLOPS/WFP8 TFLOPS: 5,000 · TDP: 1400 W
9NVIDIA B300 (Blackwell Ultra)
NVIDIA
3.57 TFLOPS/WFP8 TFLOPS: 5,000 · TDP: 1400 W
10NVIDIA GB300 (Blackwell Ultra)
NVIDIA
3.57 TFLOPS/WFP8 TFLOPS: 5,000 · TDP: 1400 W
11AMD Instinct MI300X
AMD
3.49 TFLOPS/WFP8 TFLOPS: 2,614 · TDP: 750 W
12NVIDIA GB200 Grace Blackwell Superchip
NVIDIA
3.33 TFLOPS/WFP8 TFLOPS: 9,000 · TDP: 2700 W
13Huawei Ascend 910D
Huawei
3 TFLOPS/WFP8 TFLOPS: 2,400 · TDP: 800 W
14NVIDIA H100 Tensor Core GPU
NVIDIA
2.83 TFLOPS/WFP8 TFLOPS: 1,979 · TDP: 700 W
15NVIDIA H200 Tensor Core GPU
NVIDIA
2.83 TFLOPS/WFP8 TFLOPS: 1,979 · TDP: 700 W
16NVIDIA H800
NVIDIA
2.83 TFLOPS/WFP8 TFLOPS: 1,979 · TDP: 700 W
17Cerebras WSE-3 (Wafer-Scale Engine 3)
Cerebras Systems
2.72 TFLOPS/WFP16 TFLOPS: 62,500 · TDP: 23000 W
18AMD Instinct MI325X
AMD
2.61 TFLOPS/WFP8 TFLOPS: 2,614 · TDP: 1000 W
19AWS Trainium 2
Amazon
2.6 TFLOPS/WFP8 TFLOPS: 1,300 · TDP: 500 W
20Tenstorrent Blackhole
Tenstorrent
2.48 TFLOPS/WFP16 TFLOPS: 745 · TDP: 300 W
21NVIDIA L40S
NVIDIA
2.09 TFLOPS/WFP8 TFLOPS: 733 · TDP: 350 W
22Intel Gaudi 3
Intel Corporation
2.04 TFLOPS/WFP8 TFLOPS: 1,835 · TDP: 900 W
23Biren BR100
Biren Technology
1.86 TFLOPS/WFP16 TFLOPS: 1,024 · TDP: 550 W
24SambaNova SN40L
SambaNova Systems
1.79 TFLOPS/WFP8 TFLOPS: 1,250 · TDP: 700 W
25Huawei Ascend 910C
Huawei
1.6 TFLOPS/WFP16 TFLOPS: 800 · TDP: 500 W
26Google TPU v5p
Google
1.53 TFLOPS/WBF16 TFLOPS: 459 · TDP: 300 W
27AWS Inferentia 2
Amazon
1.52 TFLOPS/WFP8 TFLOPS: 380 · TDP: 250 W
28Google TPU v4
Google
1.43 TFLOPS/WBF16 TFLOPS: 275 · TDP: 192 W
29Google TPU v5e
Google
1.16 TFLOPS/WBF16 TFLOPS: 197 · TDP: 170 W
30Huawei Ascend 910B
Huawei
1 TFLOPS/WFP16 TFLOPS: 400 · TDP: 400 W
31NVIDIA T4 Tensor Core GPU
NVIDIA
0.929 TFLOPS/WFP16 TFLOPS: 65 · TDP: 70 W
32Tenstorrent Wormhole
Tenstorrent
0.925 TFLOPS/WFP16 TFLOPS: 148 · TDP: 160 W
33Tesla Dojo D1
Tesla
0.905 TFLOPS/WFP8 TFLOPS: 362 · TDP: 400 W
34NVIDIA A10 Tensor Core GPU
NVIDIA
0.833 TFLOPS/WFP16 TFLOPS: 125 · TDP: 150 W
35NVIDIA A100 Tensor Core GPU
NVIDIA
0.78 TFLOPS/WFP16 TFLOPS: 312 · TDP: 400 W
36NVIDIA A800
NVIDIA
0.78 TFLOPS/WFP16 TFLOPS: 312 · TDP: 400 W
37NVIDIA H20
NVIDIA
0.74 TFLOPS/WFP8 TFLOPS: 296 · TDP: 400 W
38Intel Gaudi 2
Intel Corporation
0.72 TFLOPS/WBF16 TFLOPS: 432 · TDP: 600 W
39AMD Instinct MI250X
AMD
0.684 TFLOPS/WFP16 TFLOPS: 383 · TDP: 560 W
40Groq LPU (Language Processing Unit)
Groq
0.501 TFLOPS/WFP16 TFLOPS: 188 · TDP: 375 W
41AWS Trainium (Trn1)
Amazon
0.42 TFLOPS/WBF16 TFLOPS: 210 · TDP: 500 W
42NVIDIA V100 Tensor Core GPU
NVIDIA
0.417 TFLOPS/WFP16 TFLOPS: 125 · TDP: 300 W
43Cambricon MLU370-X8
Cambricon Technologies
0.384 TFLOPS/WFP16 TFLOPS: 96 · TDP: 250 W

43 chips ranked. 17 more are in our catalog but don't publish the numbers needed for this ranking.

Other rankings: Most HBM memory per chip · Highest memory bandwidth · Highest FP8 dense throughput · Highest FP16 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die

See also: every chip on record · head-to-head comparisons.