Chip leaderboard
Highest FP16 dense throughput
Cerebras WSE-3 (Wafer-Scale Engine 3) leads at 62,500 TFLOPS. 38 chips ranked.
Raw compute at FP16 (the older training precision). BF16 substituted where the vendor only publishes BF16; they're effectively the same for this comparison.
| # | Chip | Value | Inputs |
|---|---|---|---|
| 1 | Cerebras WSE-3 (Wafer-Scale Engine 3) Cerebras Systems | 62,500 TFLOPS | |
| 2 | Google TPU v7 (Ironwood) Google | 4,614 TFLOPS | |
| 3 | NVIDIA B300 (Blackwell Ultra) NVIDIA | 2,500 TFLOPS | |
| 4 | AMD Instinct MI350 series AMD | 2,300 TFLOPS | |
| 5 | NVIDIA B200 (Blackwell) NVIDIA | 2,250 TFLOPS | |
| 6 | Intel Gaudi 3 Intel Corporation | 1,835 TFLOPS | |
| 7 | NVIDIA B100 (Blackwell) NVIDIA | 1,750 TFLOPS | |
| 8 | AMD Instinct MI300X AMD | 1,307 TFLOPS | |
| 9 | AMD Instinct MI325X AMD | 1,307 TFLOPS | |
| 10 | Huawei Ascend 910D Huawei | 1,200 TFLOPS | |
| 11 | Biren BR100 Biren Technology | 1,024 TFLOPS | |
| 12 | NVIDIA H100 Tensor Core GPU NVIDIA | 989 TFLOPS | |
| 13 | NVIDIA H200 Tensor Core GPU NVIDIA | 989 TFLOPS | |
| 14 | NVIDIA H800 NVIDIA | 989 TFLOPS | |
| 15 | Google TPU v6 (Trillium) Google | 918 TFLOPS | |
| 16 | Huawei Ascend 910C Huawei | 800 TFLOPS | |
| 17 | Tenstorrent Blackhole Tenstorrent | 745 TFLOPS | |
| 18 | AWS Trainium 2 Amazon | 667 TFLOPS | |
| 19 | SambaNova SN40L SambaNova Systems | 638 TFLOPS | |
| 20 | Google TPU v5p Google | 459 TFLOPS | |
| 21 | Intel Gaudi 2 Intel Corporation | 432 TFLOPS | |
| 22 | Huawei Ascend 910B Huawei | 400 TFLOPS | |
| 23 | AMD Instinct MI250X AMD | 383 TFLOPS | |
| 24 | NVIDIA L40S NVIDIA | 362 TFLOPS | |
| 25 | Tesla Dojo D1 Tesla | 362 TFLOPS | |
| 26 | NVIDIA A100 Tensor Core GPU NVIDIA | 312 TFLOPS | |
| 27 | NVIDIA A800 NVIDIA | 312 TFLOPS | |
| 28 | Google TPU v4 Google | 275 TFLOPS | |
| 29 | AWS Trainium (Trn1) Amazon | 210 TFLOPS | |
| 30 | Google TPU v5e Google | 197 TFLOPS | |
| 31 | AWS Inferentia 2 Amazon | 190 TFLOPS | |
| 32 | Groq LPU (Language Processing Unit) Groq | 188 TFLOPS | |
| 33 | NVIDIA H20 NVIDIA | 148 TFLOPS | |
| 34 | Tenstorrent Wormhole Tenstorrent | 148 TFLOPS | |
| 35 | NVIDIA A10 Tensor Core GPU NVIDIA | 125 TFLOPS | |
| 36 | NVIDIA V100 Tensor Core GPU NVIDIA | 125 TFLOPS | |
| 37 | Cambricon MLU370-X8 Cambricon Technologies | 96 TFLOPS | |
| 38 | NVIDIA T4 Tensor Core GPU NVIDIA | 65 TFLOPS |
38 chips ranked. 22 more are in our catalog but don't publish the numbers needed for this ranking.
Other rankings: Best perf-per-watt (TFLOPS/W) · Most HBM memory per chip · Highest memory bandwidth · Highest FP8 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die
See also: every chip on record · head-to-head comparisons.