Chip registry
Chip leaderboards
Every AI chip we track, ranked head-to-head on the same measure. Numbers straight from each vendor's datasheet, sorted and divided with the math visible on every page.
Best perf-per-watt (TFLOPS/W)
How much work each chip does per watt. Straight from the vendor datasheet: best dense throughput divided by rated power. Wafer-scale and rack-scale chips look artificially strong because their datasheet numbers roll in shared overhead the per-chip numbers ignore.
Most HBM memory per chip
How much high-bandwidth memory each chip has. This is the ceiling on the largest model that fits on one chip without splitting. Cerebras WSE-3 lists on-die SRAM, not HBM, but it's included since the question is capacity.
Highest memory bandwidth
How fast each chip can read and write its own memory. This is the ceiling for LLM inference speed, where every generated token pulls the full model out of memory again. Higher is faster.
Highest FP8 dense throughput
Raw compute at FP8: the precision the biggest 2024-2026 training runs use. Dense throughput, not sparse.
Highest FP16 dense throughput
Raw compute at FP16 (the older training precision). BF16 substituted where the vendor only publishes BF16; they're effectively the same for this comparison.
Highest INT8 dense throughput
Raw compute at INT8: the standard precision for serving inference. Dense throughput, not sparse.
Cheapest per TFLOP at launch (list price)
The list price at launch divided by peak throughput. Lower is better value. Only chips with a public launch price appear; street prices and cloud rental rates can differ a lot.
Most transistors per die
How many transistors are on each chip's die. Wafer-scale designs like Cerebras WSE-3 dominate by definition; chiplet packages like GB200 and MI300X come next.
See also: every chip on record · head-to-head comparisons · reference-design clusters.