Chip leaderboard
Highest FP8 dense throughput
AMD Instinct MI400 leads at 20,000 TFLOPS. 25 chips ranked.
Raw compute at FP8: the precision the biggest 2024-2026 training runs use. Dense throughput, not sparse.
| # | Chip | Value | Inputs |
|---|---|---|---|
| 1 | AMD Instinct MI400 AMD | 20,000 TFLOPS | |
| 2 | Google TPU v7 (Ironwood) Google | 9,228 TFLOPS | |
| 3 | NVIDIA GB200 Grace Blackwell Superchip NVIDIA | 9,000 TFLOPS | |
| 4 | AMD Instinct MI355X AMD | 5,000 TFLOPS | |
| 5 | NVIDIA B300 (Blackwell Ultra) NVIDIA | 5,000 TFLOPS | |
| 6 | NVIDIA GB300 (Blackwell Ultra) NVIDIA | 5,000 TFLOPS | |
| 7 | AMD Instinct MI350 series AMD | 4,600 TFLOPS | |
| 8 | NVIDIA B200 (Blackwell) NVIDIA | 4,500 TFLOPS | |
| 9 | NVIDIA B100 (Blackwell) NVIDIA | 3,500 TFLOPS | |
| 10 | AMD Instinct MI300X AMD | 2,614 TFLOPS | |
| 11 | AMD Instinct MI325X AMD | 2,614 TFLOPS | |
| 12 | Huawei Ascend 910D Huawei | 2,400 TFLOPS | |
| 13 | NVIDIA Jetson AGX Thor NVIDIA | 2,070 TFLOPS | |
| 14 | NVIDIA DRIVE Thor NVIDIA | 2,000 TFLOPS | |
| 15 | NVIDIA H100 Tensor Core GPU NVIDIA | 1,979 TFLOPS | |
| 16 | NVIDIA H200 Tensor Core GPU NVIDIA | 1,979 TFLOPS | |
| 17 | NVIDIA H800 NVIDIA | 1,979 TFLOPS | |
| 18 | Google TPU v6 (Trillium) Google | 1,836 TFLOPS | |
| 19 | Intel Gaudi 3 Intel Corporation | 1,835 TFLOPS | |
| 20 | AWS Trainium 2 Amazon | 1,300 TFLOPS | |
| 21 | SambaNova SN40L SambaNova Systems | 1,250 TFLOPS | |
| 22 | NVIDIA L40S NVIDIA | 733 TFLOPS | |
| 23 | AWS Inferentia 2 Amazon | 380 TFLOPS | |
| 24 | Tesla Dojo D1 Tesla | 362 TFLOPS | |
| 25 | NVIDIA H20 NVIDIA | 296 TFLOPS |
25 chips ranked. 35 more are in our catalog but don't publish the numbers needed for this ranking.
Other rankings: Best perf-per-watt (TFLOPS/W) · Most HBM memory per chip · Highest memory bandwidth · Highest FP16 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die
See also: every chip on record · head-to-head comparisons.