DEPLOY

AI chip · Google

Google TPU v6 (Trillium)

Sixth-generation TPU announced May 2024; ~4.7x peak compute over TPU v5e. Backbone of Gemini 2.0-era training runs.

Deployed in 1 named data center · Part of 1 rack design.

Market position

Google's inference-optimised sixth generation (v6e). 4.7x TPU v5e per-chip FP16, 67% more memory bandwidth. Serves Gemini 1.5 Flash and Gemini 2 Nano-class workloads inside Google Cloud. v6p (Ironwood) is the training-optimised sibling.

Efficiency and power

How much work you get per watt and per dollar, and how much power a full rack draws.

Perf per watt
9.18 FP8 TFLOPS/W
1,836 TFLOPS ÷ 200 W = 9.18
Cooling
water
Depends on the rack design

What fits in 32 GB

Which open-source LLMs run on one Google TPU v6 (Trillium), by precision. Weights only: add roughly 15% headroom for real serving. If a model does not fit at FP16, try INT8 or INT4 (smaller quality trade-off than most people expect).

ModelParamsFP16INT8INT4
Llama 3.1 8B
dense
8 B16 GB ✓8 GB ✓4 GB ✓
Llama 3.1 70B
dense
70 B140 GB ✗70 GB ✗35 GB ✗
Llama 3.1 405B
dense
405 B810 GB ✗405 GB ✗203 GB ✗
Llama 3.3 70B
dense
70 B140 GB ✗70 GB ✗35 GB ✗
DeepSeek V3
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
DeepSeek R1
MoE (671B total, 37B active per token)
671 B1342 GB ✗671 GB ✗336 GB ✗
Qwen 2.5 7B
dense
7 B14 GB ✓7 GB ✓4 GB ✓
Qwen 2.5 72B
dense
72 B144 GB ✗72 GB ✗36 GB ✗
Mixtral 8x7B
MoE (46.7B total, 12.9B active per token)
46.7 B93 GB ✗47 GB ✗23 GB ✓
Mixtral 8x22B
MoE (141B total, 39B active per token)
141 B282 GB ✗141 GB ✗71 GB ✗
Gemma 2 27B
dense
27 B54 GB ✗27 GB ✓14 GB ✓
Command R+
dense
104 B208 GB ✗104 GB ✗52 GB ✗
Kimi K2
MoE (1T total, 32B active per token)
1000 B2000 GB ✗1000 GB ✗500 GB ✗

Math: FP16 = params × 2 bytes; INT8 = params × 1 byte; INT4 = params × 0.5 bytes. MoE models sum every expert (full weights on disk), not the per-token active subset.

Common questions

How much does Google TPU v6 (Trillium) cost?

Google TPU v6 (Trillium) doesn't have a public launch list price. Vendors like this one usually price through direct sales rather than a public sheet. See list price per TFLOP for chips that do publish.

How much memory does Google TPU v6 (Trillium) have?

32 GB of HBM3, running at 1,600 GB/s. Straight from the vendor datasheet. See chips with the most memory for context.

How much power does one Google TPU v6 (Trillium) draw?

200 W at the chip. A full server draws more once you add CPU, memory, networking and cooling: see the rack power number above. Compare to other chips on perf-per-watt.

Which data centers use Google TPU v6 (Trillium)?

1 named data center runs them, including Google Fort Wayne / New Haven. See who has the most Google TPU v6 (Trillium) for the ranked list.

Which open-source LLMs fit on one Google TPU v6 (Trillium)?

Of 13 open-source LLMs we track, 2 fit at FP16, 3 at INT8, and 4 at INT4 (for example Llama 3.1 8B, Qwen 2.5 7B and Mixtral 8x7B). Full table above with each model's memory need.

See every answer we publish →

Key facts

Class
Compute SoC (AI accelerator)
Designer
Google
Safety-critical?
No (data-center inference / training)
Record as of
2026-09-27
Most recent source
2024-05-01 (across 2 sources on this page)
Specifications (9 fields, click to expand)

Straight from the vendor datasheet. Dense throughput shown first; sparse (2:4) numbers in parentheses where the vendor publishes them. Full datasheet linked below.

TDP
200 W
Memory
32 GB HBM3
Memory bandwidth
1,600 GB/s
BF16 (dense)
918 TFLOPS
FP8 (dense)
1,836 TFLOPS
INT8 (dense)
1,836 TOPS
Form factor
OAM (Trillium pod, up to 256 chips)
Announced
2024-05-14
Released
2024-12-01

Source: vendor datasheet

Generation

Foundry & process

Fabbed at
TSMC (all chips TSMC makes →)

Rack designs that use it

Whole-rack systems buyers order in bulk (NVL72, HGX, TPU pods) that ship with Google TPU v6 (Trillium) inside.

See all rack designs →

Compare with

See every chip comparison →

Benchmarks

Published performance numbers by workload. Vendor datasheet figures where noted; independent measurements otherwise.

WorkloadValueUnitSourceAs of
peak tflops bf16(~4.7x TPU v5e)918TFLOPSvendor2024-05-01

Data centers running Google TPU v6 (Trillium)

Who has the most? →

Named data-center campuses with Google TPU v6 (Trillium) on site. Counts shown where the operator has published them; other rows are described in general terms.

  1. under constructioninferred

    Google TPU capacity in Fort Wayne / New Haven campus (Trillium-generation)

    Chip mix not published per-campus; inferred from Google's current-generation TPU deployment posture.

    Google

Sources

Adoption rows appear as we confirm each chip-in-robot pairing from a public source. See every chip we track for the full catalog or the Google page.