Case study · Gate-level simulation and power activity
A 247-million-gate transformer inference chip, on one laptop GPU
To measure at the size real inference chips reach, we wrote one: an integer transformer datapath in Verilog, synthesized to gates, with a C program that defines every correct answer. At 246.8 million gates it runs 64 independent data sets through 385,186 clock cycles in 6 hours 38 minutes on an 8 GB RTX 4060 laptop GPU, and every output of every data set matches the C reference byte for byte.
34.4 million flip-flops, 32 transformer stages, one netlist.
64 data sets × 385,186 cycles, each with its own weights and inputs.
data sets matching the C reference, every output byte.
simulation time for SAIF switching activity on all 246.8M nets.
What the chip computes
Each expert is a complete transformer block over a 64 × 64 block of int8 activations:
M = LayerNorm( CausalMultiHeadAttention(X) + X ) Y = LayerNorm( W2 · GELU(W1 · M) + M )
A layer holds several experts side by side, each with its own weights. A router scores every block against each expert and the highest-scoring ones run it; the layer's output is their average. Layers are pipelined, one block per layer period. All arithmetic is exact integer arithmetic, so the gates either match the reference byte for byte or they do not.
- Attention, completeQ, K and V projections, rotary position embedding, causal multi-head attention, softmax with the maximum subtracted, output projection, residual and LayerNorm.
- Tokens in, tokens outAn embedding table on the way in, a 256-way unembedding and argmax on the way out, and a decode mode that generates one token at a time against a KV cache.
- The control of a real partToken and descriptor queues, credit-based flow control between layers, top-K expert routing, and clock gating on idle experts.
- Deliberately mixed logicSystolic arrays, dividers and square roots, table-driven softmax and GELU, address-decoded weight banks and a crossbar.
Weights are generated, not trained, and weights and caches live in flip-flops rather than SRAM macros.
How it is built
One stage is written in Verilog and synthesized once. A tool then wires copies of that netlist into a grid of any size, so a quarter-billion-gate chip costs the same synthesis as a single stage, and building the full grid takes about two minutes. One C program generates every weight and parameter and computes the expected output, and each level is checked against it: every unit in RTL, the whole grid in RTL, the gate netlist on a CPU reference simulator, and the full chip on the GPU.
The diagrams show the 4 × 8 grid at width 64 as it was run for the figures on this page: four experts per layer, all running every block, eight layers.
1 · The grid
8 layers × 4 experts2 · One stage
the unit replicated 32 times3 · The attention head
one systolic array, six products4 · The feed-forward half
a second array of the same type5 · The run, to scale
385,186 clock cyclesThe runs
One netlist per row, each data set with its own weights, parameters, inputs and expected output. Every data set of every row matched the C reference.
RTX 4060 laptop GPU, 8 GB, 2-state
| Grid (experts × layers, width) | Gates | Flip-flops | Data sets | Cycles per data set | Simulation |
|---|---|---|---|---|---|
| 1 × 4, width 16 | 2,579,229 | 289,992 | 64 | 51,266 | 53 s |
| 1 × 2, width 64 | 15,426,333 | 2,152,544 | 16 | 259,186 | 14 min 8 s |
| 1 × 2, width 64 | 15,426,333 | 2,152,544 | 64 | 259,186 | 14 min 44 s |
| 2 × 16, width 16 | 20,633,605 | 2,319,936 | 64 | 74,370 | 5 min 42 s |
| 2 × 4, width 64 | 61,703,757 | 8,610,176 | 64 | 225,378 | 54 min 24 s |
| 4 × 8, width 64 | 246,812,997 | 34,440,704 | 64 | 385,186 | 6 h 38 min |
- 64 data sets cost little more than 16On the 15.4M-gate grid, 64 data sets took 884 s against 848 s for 16: four times the verification for 4% more time.
- Shape does not matter247.6M gates built as 384 narrow stages and 246.8M as 32 wide ones ran within 5% of each other.
- Switching activity for powerA SAIF file for all 246.8M nets, 14.8 GB, costs 37% more simulation time. Every net's time at 0, 1 and X adds up to the run's duration exactly.
- A quarter billion gates in 4 GBThe 246.8M-gate run used 4.1 GB of the 8 GB card, about twelve bytes a gate plus the state of all 64 data sets.
The runs in the table are on the dense configuration, before the router, tokens in and out and clock gating were added. The final chip, with all of them, has been checked on the GPU against the C reference on every data set, from 1.2M gates up to a 259M-gate grid.
Next
The same chip at 2 billion gates on a single datacentre GPU, and switching activity from the final chip's routed runs, where idle experts are clock-gated and power depends on the data.