DEPLOY

Chip leaderboard

Highest FP16 dense throughput

Cerebras WSE-3 (Wafer-Scale Engine 3) leads at 62,500 TFLOPS. 38 chips ranked.

Raw compute at FP16 (the older training precision). BF16 substituted where the vendor only publishes BF16; they're effectively the same for this comparison.

#ChipValueInputs
1Cerebras WSE-3 (Wafer-Scale Engine 3)
Cerebras Systems
62,500 TFLOPS
2Google TPU v7 (Ironwood)
Google
4,614 TFLOPS
3NVIDIA B300 (Blackwell Ultra)
NVIDIA
2,500 TFLOPS
4AMD Instinct MI350 series
AMD
2,300 TFLOPS
5NVIDIA B200 (Blackwell)
NVIDIA
2,250 TFLOPS
6Intel Gaudi 3
Intel Corporation
1,835 TFLOPS
7NVIDIA B100 (Blackwell)
NVIDIA
1,750 TFLOPS
8AMD Instinct MI300X
AMD
1,307 TFLOPS
9AMD Instinct MI325X
AMD
1,307 TFLOPS
10Huawei Ascend 910D
Huawei
1,200 TFLOPS
11Biren BR100
Biren Technology
1,024 TFLOPS
12NVIDIA H100 Tensor Core GPU
NVIDIA
989 TFLOPS
13NVIDIA H200 Tensor Core GPU
NVIDIA
989 TFLOPS
14NVIDIA H800
NVIDIA
989 TFLOPS
15Google TPU v6 (Trillium)
Google
918 TFLOPS
16Huawei Ascend 910C
Huawei
800 TFLOPS
17Tenstorrent Blackhole
Tenstorrent
745 TFLOPS
18AWS Trainium 2
Amazon
667 TFLOPS
19SambaNova SN40L
SambaNova Systems
638 TFLOPS
20Google TPU v5p
Google
459 TFLOPS
21Intel Gaudi 2
Intel Corporation
432 TFLOPS
22Huawei Ascend 910B
Huawei
400 TFLOPS
23AMD Instinct MI250X
AMD
383 TFLOPS
24NVIDIA L40S
NVIDIA
362 TFLOPS
25Tesla Dojo D1
Tesla
362 TFLOPS
26NVIDIA A100 Tensor Core GPU
NVIDIA
312 TFLOPS
27NVIDIA A800
NVIDIA
312 TFLOPS
28Google TPU v4
Google
275 TFLOPS
29AWS Trainium (Trn1)
Amazon
210 TFLOPS
30Google TPU v5e
Google
197 TFLOPS
31AWS Inferentia 2
Amazon
190 TFLOPS
32Groq LPU (Language Processing Unit)
Groq
188 TFLOPS
33NVIDIA H20
NVIDIA
148 TFLOPS
34Tenstorrent Wormhole
Tenstorrent
148 TFLOPS
35NVIDIA A10 Tensor Core GPU
NVIDIA
125 TFLOPS
36NVIDIA V100 Tensor Core GPU
NVIDIA
125 TFLOPS
37Cambricon MLU370-X8
Cambricon Technologies
96 TFLOPS
38NVIDIA T4 Tensor Core GPU
NVIDIA
65 TFLOPS

38 chips ranked. 22 more are in our catalog but don't publish the numbers needed for this ranking.

Other rankings: Best perf-per-watt (TFLOPS/W) · Most HBM memory per chip · Highest memory bandwidth · Highest FP8 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die

See also: every chip on record · head-to-head comparisons.