Chip leaderboard
Best perf-per-watt (TFLOPS/W)
NVIDIA Jetson AGX Thor leads at 15.9 TFLOPS/W. 43 chips ranked.
How much work each chip does per watt. Straight from the vendor datasheet: best dense throughput divided by rated power. Wafer-scale and rack-scale chips look artificially strong because their datasheet numbers roll in shared overhead the per-chip numbers ignore.
| # | Chip | Value | Inputs |
|---|---|---|---|
| 1 | NVIDIA Jetson AGX Thor NVIDIA | 15.9 TFLOPS/W | FP8 TFLOPS: 2,070 · TDP: 130 W |
| 2 | NVIDIA DRIVE Thor NVIDIA | 15.4 TFLOPS/W | FP8 TFLOPS: 2,000 · TDP: 130 W |
| 3 | Google TPU v7 (Ironwood) Google | 13.2 TFLOPS/W | FP8 TFLOPS: 9,228 · TDP: 700 W |
| 4 | Google TPU v6 (Trillium) Google | 9.18 TFLOPS/W | FP8 TFLOPS: 1,836 · TDP: 200 W |
| 5 | NVIDIA B100 (Blackwell) NVIDIA | 5 TFLOPS/W | FP8 TFLOPS: 3,500 · TDP: 700 W |
| 6 | AMD Instinct MI350 series AMD | 4.6 TFLOPS/W | FP8 TFLOPS: 4,600 · TDP: 1000 W |
| 7 | NVIDIA B200 (Blackwell) NVIDIA | 4.5 TFLOPS/W | FP8 TFLOPS: 4,500 · TDP: 1000 W |
| 8 | AMD Instinct MI355X AMD | 3.57 TFLOPS/W | FP8 TFLOPS: 5,000 · TDP: 1400 W |
| 9 | NVIDIA B300 (Blackwell Ultra) NVIDIA | 3.57 TFLOPS/W | FP8 TFLOPS: 5,000 · TDP: 1400 W |
| 10 | NVIDIA GB300 (Blackwell Ultra) NVIDIA | 3.57 TFLOPS/W | FP8 TFLOPS: 5,000 · TDP: 1400 W |
| 11 | AMD Instinct MI300X AMD | 3.49 TFLOPS/W | FP8 TFLOPS: 2,614 · TDP: 750 W |
| 12 | NVIDIA GB200 Grace Blackwell Superchip NVIDIA | 3.33 TFLOPS/W | FP8 TFLOPS: 9,000 · TDP: 2700 W |
| 13 | Huawei Ascend 910D Huawei | 3 TFLOPS/W | FP8 TFLOPS: 2,400 · TDP: 800 W |
| 14 | NVIDIA H100 Tensor Core GPU NVIDIA | 2.83 TFLOPS/W | FP8 TFLOPS: 1,979 · TDP: 700 W |
| 15 | NVIDIA H200 Tensor Core GPU NVIDIA | 2.83 TFLOPS/W | FP8 TFLOPS: 1,979 · TDP: 700 W |
| 16 | NVIDIA H800 NVIDIA | 2.83 TFLOPS/W | FP8 TFLOPS: 1,979 · TDP: 700 W |
| 17 | Cerebras WSE-3 (Wafer-Scale Engine 3) Cerebras Systems | 2.72 TFLOPS/W | FP16 TFLOPS: 62,500 · TDP: 23000 W |
| 18 | AMD Instinct MI325X AMD | 2.61 TFLOPS/W | FP8 TFLOPS: 2,614 · TDP: 1000 W |
| 19 | AWS Trainium 2 Amazon | 2.6 TFLOPS/W | FP8 TFLOPS: 1,300 · TDP: 500 W |
| 20 | Tenstorrent Blackhole Tenstorrent | 2.48 TFLOPS/W | FP16 TFLOPS: 745 · TDP: 300 W |
| 21 | NVIDIA L40S NVIDIA | 2.09 TFLOPS/W | FP8 TFLOPS: 733 · TDP: 350 W |
| 22 | Intel Gaudi 3 Intel Corporation | 2.04 TFLOPS/W | FP8 TFLOPS: 1,835 · TDP: 900 W |
| 23 | Biren BR100 Biren Technology | 1.86 TFLOPS/W | FP16 TFLOPS: 1,024 · TDP: 550 W |
| 24 | SambaNova SN40L SambaNova Systems | 1.79 TFLOPS/W | FP8 TFLOPS: 1,250 · TDP: 700 W |
| 25 | Huawei Ascend 910C Huawei | 1.6 TFLOPS/W | FP16 TFLOPS: 800 · TDP: 500 W |
| 26 | Google TPU v5p Google | 1.53 TFLOPS/W | BF16 TFLOPS: 459 · TDP: 300 W |
| 27 | AWS Inferentia 2 Amazon | 1.52 TFLOPS/W | FP8 TFLOPS: 380 · TDP: 250 W |
| 28 | Google TPU v4 Google | 1.43 TFLOPS/W | BF16 TFLOPS: 275 · TDP: 192 W |
| 29 | Google TPU v5e Google | 1.16 TFLOPS/W | BF16 TFLOPS: 197 · TDP: 170 W |
| 30 | Huawei Ascend 910B Huawei | 1 TFLOPS/W | FP16 TFLOPS: 400 · TDP: 400 W |
| 31 | NVIDIA T4 Tensor Core GPU NVIDIA | 0.929 TFLOPS/W | FP16 TFLOPS: 65 · TDP: 70 W |
| 32 | Tenstorrent Wormhole Tenstorrent | 0.925 TFLOPS/W | FP16 TFLOPS: 148 · TDP: 160 W |
| 33 | Tesla Dojo D1 Tesla | 0.905 TFLOPS/W | FP8 TFLOPS: 362 · TDP: 400 W |
| 34 | NVIDIA A10 Tensor Core GPU NVIDIA | 0.833 TFLOPS/W | FP16 TFLOPS: 125 · TDP: 150 W |
| 35 | NVIDIA A100 Tensor Core GPU NVIDIA | 0.78 TFLOPS/W | FP16 TFLOPS: 312 · TDP: 400 W |
| 36 | NVIDIA A800 NVIDIA | 0.78 TFLOPS/W | FP16 TFLOPS: 312 · TDP: 400 W |
| 37 | NVIDIA H20 NVIDIA | 0.74 TFLOPS/W | FP8 TFLOPS: 296 · TDP: 400 W |
| 38 | Intel Gaudi 2 Intel Corporation | 0.72 TFLOPS/W | BF16 TFLOPS: 432 · TDP: 600 W |
| 39 | AMD Instinct MI250X AMD | 0.684 TFLOPS/W | FP16 TFLOPS: 383 · TDP: 560 W |
| 40 | Groq LPU (Language Processing Unit) Groq | 0.501 TFLOPS/W | FP16 TFLOPS: 188 · TDP: 375 W |
| 41 | AWS Trainium (Trn1) Amazon | 0.42 TFLOPS/W | BF16 TFLOPS: 210 · TDP: 500 W |
| 42 | NVIDIA V100 Tensor Core GPU NVIDIA | 0.417 TFLOPS/W | FP16 TFLOPS: 125 · TDP: 300 W |
| 43 | Cambricon MLU370-X8 Cambricon Technologies | 0.384 TFLOPS/W | FP16 TFLOPS: 96 · TDP: 250 W |
43 chips ranked. 17 more are in our catalog but don't publish the numbers needed for this ranking.
Other rankings: Most HBM memory per chip · Highest memory bandwidth · Highest FP8 dense throughput · Highest FP16 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die
See also: every chip on record · head-to-head comparisons.