Chip leaderboard
Highest memory bandwidth
Cerebras WSE-3 (Wafer-Scale Engine 3) leads at 21,000,000 GB/s. 43 chips ranked.
How fast each chip can read and write its own memory. This is the ceiling for LLM inference speed, where every generated token pulls the full model out of memory again. Higher is faster.
| # | Chip | Value | Inputs |
|---|---|---|---|
| 1 | Cerebras WSE-3 (Wafer-Scale Engine 3) Cerebras Systems | 21,000,000 GB/s | Type: on-die SRAM |
| 2 | Groq LPU (Language Processing Unit) Groq | 80,000 GB/s | Type: SRAM (230 MB on-die) |
| 3 | AMD Instinct MI400 AMD | 19,600 GB/s | Type: HBM4 |
| 4 | NVIDIA GB200 Grace Blackwell Superchip NVIDIA | 16,000 GB/s | Type: HBM3e |
| 5 | AMD Instinct MI350 series AMD | 8,000 GB/s | Type: HBM3e |
| 6 | AMD Instinct MI355X AMD | 8,000 GB/s | Type: HBM3e |
| 7 | NVIDIA B100 (Blackwell) NVIDIA | 8,000 GB/s | Type: HBM3e |
| 8 | NVIDIA B200 (Blackwell) NVIDIA | 8,000 GB/s | Type: HBM3e |
| 9 | NVIDIA B300 (Blackwell Ultra) NVIDIA | 8,000 GB/s | Type: HBM3e |
| 10 | NVIDIA GB300 (Blackwell Ultra) NVIDIA | 8,000 GB/s | Type: HBM3e |
| 11 | Google TPU v7 (Ironwood) Google | 7,400 GB/s | Type: HBM3e |
| 12 | AMD Instinct MI325X AMD | 6,000 GB/s | Type: HBM3e |
| 13 | AMD Instinct MI300X AMD | 5,300 GB/s | Type: HBM3 |
| 14 | NVIDIA GH200 Grace Hopper Superchip NVIDIA | 4,900 GB/s | Type: HBM3e |
| 15 | NVIDIA H200 Tensor Core GPU NVIDIA | 4,800 GB/s | Type: HBM3e |
| 16 | Huawei Ascend 910D Huawei | 4,000 GB/s | Type: HBM3 |
| 17 | NVIDIA H20 NVIDIA | 4,000 GB/s | Type: HBM3 |
| 18 | Intel Gaudi 3 Intel Corporation | 3,675 GB/s | Type: HBM2e |
| 19 | NVIDIA H100 Tensor Core GPU NVIDIA | 3,350 GB/s | Type: HBM3 |
| 20 | NVIDIA H800 NVIDIA | 3,350 GB/s | Type: HBM3 |
| 21 | AMD Instinct MI250X AMD | 3,200 GB/s | Type: HBM2e |
| 22 | Huawei Ascend 910C Huawei | 3,200 GB/s | Type: HBM2e |
| 23 | AWS Trainium 2 Amazon | 2,900 GB/s | Type: HBM3 |
| 24 | Google TPU v5p Google | 2,765 GB/s | Type: HBM2e |
| 25 | Intel Gaudi 2 Intel Corporation | 2,450 GB/s | Type: HBM2e |
| 26 | Biren BR100 Biren Technology | 2,300 GB/s | Type: HBM2e |
| 27 | NVIDIA A100 Tensor Core GPU NVIDIA | 2,039 GB/s | Type: HBM2e |
| 28 | NVIDIA A800 NVIDIA | 2,039 GB/s | Type: HBM2e |
| 29 | Google TPU v6 (Trillium) Google | 1,600 GB/s | Type: HBM3 |
| 30 | Huawei Ascend 910B Huawei | 1,600 GB/s | Type: HBM2e |
| 31 | Google TPU v4 Google | 1,200 GB/s | Type: HBM2 |
| 32 | NVIDIA V100 Tensor Core GPU NVIDIA | 900 GB/s | Type: HBM2 |
| 33 | NVIDIA L40S NVIDIA | 864 GB/s | Type: GDDR6 |
| 34 | AWS Inferentia 2 Amazon | 820 GB/s | Type: HBM3 |
| 35 | AWS Trainium (Trn1) Amazon | 820 GB/s | Type: HBM2e |
| 36 | Google TPU v5e Google | 819 GB/s | Type: HBM2 |
| 37 | NVIDIA A10 Tensor Core GPU NVIDIA | 600 GB/s | Type: GDDR6 |
| 38 | Tenstorrent Blackhole Tenstorrent | 512 GB/s | Type: GDDR6 |
| 39 | NVIDIA T4 Tensor Core GPU NVIDIA | 320 GB/s | Type: GDDR6 |
| 40 | Tenstorrent Wormhole Tenstorrent | 288 GB/s | Type: GDDR6 |
| 41 | NVIDIA Jetson AGX Thor NVIDIA | 273 GB/s | Type: LPDDR5X |
| 42 | Meta MTIA v2 Meta | 205 GB/s | Type: LPDDR5 |
| 43 | NVIDIA Jetson AGX Orin NVIDIA | 205 GB/s | Type: LPDDR5 |
43 chips ranked. 17 more are in our catalog but don't publish the numbers needed for this ranking.
Other rankings: Best perf-per-watt (TFLOPS/W) · Most HBM memory per chip · Highest FP8 dense throughput · Highest FP16 dense throughput · Highest INT8 dense throughput · Cheapest per TFLOP at launch (list price) · Most transistors per die
See also: every chip on record · head-to-head comparisons.