XEROLANE

Case study · Gate-level simulation and power activity

A 247-million-gate transformer inference chip, on one laptop GPU

To measure at the size real inference chips reach, we wrote one: an integer transformer datapath in Verilog, synthesized to gates, with a C program that defines every correct answer. At 246.8 million gates it runs 64 independent data sets through 385,186 clock cycles in 6 hours 38 minutes on an 8 GB RTX 4060 laptop GPU, and every output of every data set matches the C reference byte for byte.

246.8M gates

34.4 million flip-flops, 32 transformer stages, one netlist.

6:38h:min

64 data sets × 385,186 cycles, each with its own weights and inputs.

64of 64

data sets matching the C reference, every output byte.

+37%

simulation time for SAIF switching activity on all 246.8M nets.

What the chip computes

Each expert is a complete transformer block over a 64 × 64 block of int8 activations:

M = LayerNorm( CausalMultiHeadAttention(X) + X )
Y = LayerNorm( W2 · GELU(W1 · M) + M )

A layer holds several experts side by side, each with its own weights. A router scores every block against each expert and the highest-scoring ones run it; the layer's output is their average. Layers are pipelined, one block per layer period. All arithmetic is exact integer arithmetic, so the gates either match the reference byte for byte or they do not.

  • Attention, completeQ, K and V projections, rotary position embedding, causal multi-head attention, softmax with the maximum subtracted, output projection, residual and LayerNorm.
  • Tokens in, tokens outAn embedding table on the way in, a 256-way unembedding and argmax on the way out, and a decode mode that generates one token at a time against a KV cache.
  • The control of a real partToken and descriptor queues, credit-based flow control between layers, top-K expert routing, and clock gating on idle experts.
  • Deliberately mixed logicSystolic arrays, dividers and square roots, table-driven softmax and GELU, address-decoded weight banks and a crossbar.

Weights are generated, not trained, and weights and caches live in flip-flops rather than SRAM macros.

How it is built

One stage is written in Verilog and synthesized once. A tool then wires copies of that netlist into a grid of any size, so a quarter-billion-gate chip costs the same synthesis as a single stage, and building the full grid takes about two minutes. One C program generates every weight and parameter and computes the expected output, and each level is checked against it: every unit in RTL, the whole grid in RTL, the gate netlist on a CPU reference simulator, and the full chip on the GPU.

The diagrams show the 4 × 8 grid at width 64 as it was run for the figures on this page: four experts per layer, all running every block, eight layers.

1 · The grid

8 layers × 4 experts
665 inputs clk · rst 5 shifts [25 b] 3 scales [48 b] → every stage, fixed x_byte [8] x_valid wchain_in [512] w_shift cfg_in [40] ln_in [24] 4 load strobes 4,096 bytes, one per clock layer 0 expert 0 stage 0 expert 1 stage 1 expert 2 stage 2 expert 3 stage 3 (ΣY+2)>>2 layer 1 expert 0 stage 4 expert 1 stage 5 expert 2 stage 6 expert 3 stage 7 (ΣY+2)>>2 layer 2 expert 0 stage 8 expert 1 stage 9 expert 2 stage 10 expert 3 stage 11 (ΣY+2)>>2 layer 3 expert 0 stage 12 expert 1 stage 13 expert 2 stage 14 expert 3 stage 15 (ΣY+2)>>2 layer 4 expert 0 stage 16 expert 1 stage 17 expert 2 stage 18 expert 3 stage 19 (ΣY+2)>>2 layer 5 expert 0 stage 20 expert 1 stage 21 expert 2 stage 22 expert 3 stage 23 (ΣY+2)>>2 layer 6 expert 0 stage 24 expert 1 stage 25 expert 2 stage 26 expert 3 stage 27 (ΣY+2)>>2 layer 7 expert 0 stage 28 expert 1 stage 29 expert 2 stage 30 expert 3 stage 31 (ΣY+2)>>2 10 outputs y_out [8] y_valid · busy configuration chains, stage 0 → stage 31: weights 512 b · requantiser {bias, mult} 40 b · LayerNorm {β, γ} 24 b layer 0 layer 1 layer 2 layer 3 layer 4 layer 5 layer 6 layer 7 next stage
All four experts of a layer receive the same byte stream at the same clocks, and the layer outputs their average, (Y₀+Y₁+Y₂+Y₃+2) >> 2, clamped to int8. Weights and per-channel parameters reach the stages through three daisy chains that thread all 32 stages, so the pin count does not grow with the grid.

2 · One stage

the unit replicated 32 times
layer_stage ×32 · 7,712,827 cells · 1,076,272 flip-flops read, written back and advanced: one full turn per block, loaded through the chain banked weight file: 6 banks × 64 words × 512 b = 196,608 bits one pointer; address decode, read multiplexers, a six-way select weights from previous stage to next stage Wq Wk Wv Wo W1 W2 bytes in staging buffer 64 × 64 bytes fills while the layer is busy 64 clk attn_block Q, K, V · scores · causal softmax Pr·V · Wo · +X residual LayerNorm → M gather 64-byte row M ffn_block M·W1 → GELU → A A·W2 → +M residual → B LayerNorm → Y Y, 1 byte/clk → combiner requantiser 40 b · LayerNorm 24 b chains a_cfg · a_ln f_cfg · f_ln to next stage
A stage is one transformer block plus what a pipeline needs. The staging buffer collects the next block from the previous layer while this one is still busy. The weight file holds all six weight matrices in six banks behind one pointer, and turns once per block.

3 · The attention head

one systolic array, six products
from the weight store: Wq · Wk · Wv · Wo weight input 64 words, w_load held N clocks nn_core: array + 64 requantisers + sequencer systolic MAC array 64 × 64 = 4,096 cells int8 × int8 → 24-bit acc weight-stationary sums leave 2N clocks later 64 requant ×64 + bias × mult >> shift clamp int8 elementwise ×64 ×scale >>sh + residual scratchpads, 64 × 64 B each K keys V values S scores Q queries Pr probabilities O attention out R residual sum X input block Kᵀ: column gather (crossbar) · V rows rows X Q Pr O X row, added to U on the last product result rows written back: Q · K · V · S · O · R softmax max · exp tables · 1/Σ, causal LayerNorm mean · var · √ · 1/σ · γ, β M, 1 byte/clk → FFN
One 64 × 64 systolic array computes every matrix product of attention in sequence, loading a different weight matrix each time. Four matrices come from the weight file. V is a row read from a scratchpad and Kᵀ a column read, one byte from each of K’s 64 rows: a crossbar, unlike anything else in the design.

4 · The feed-forward half

a second array of the same type
M rows from gather array × W1 requant ×64 GELU 17-pt table A scratchpad array × W2 requant ×64 + M residual LayerNorm γ, β Y 1 byte/clk M stays in the FFN scratchpad and is added back after W2
W1’s products pass through GELU, a 17-point table with linear interpolation. W2’s products get M added back as a residual, and a second LayerNorm gives the stage output.

5 · The run, to scale

385,186 clock cycles
setup, 16,386 clocks: reset 2 · weight chain 12,288 · requantiser and LayerNorm chains 4,096 1 2 3 4 5 6 7 8 9 10 SAIF runs: first 58,000 clocks (116,000 rows) 0 100,000 200,000 300,000 385,186 design clocks (two stimulus rows each)
Setup loads every stage’s weights and parameters through the chains. Then ten stage periods of 36,880 clocks carry the block through all eight layers and out. Each of the 64 data sets has its own weights, parameters and input block, all through the same netlist.

The runs

One netlist per row, each data set with its own weights, parameters, inputs and expected output. Every data set of every row matched the C reference.

RTX 4060 laptop GPU, 8 GB, 2-state

Grid (experts × layers, width)GatesFlip-flopsData setsCycles per data setSimulation
1 × 4, width 162,579,229289,9926451,26653 s
1 × 2, width 6415,426,3332,152,54416259,18614 min 8 s
1 × 2, width 6415,426,3332,152,54464259,18614 min 44 s
2 × 16, width 1620,633,6052,319,9366474,3705 min 42 s
2 × 4, width 6461,703,7578,610,17664225,37854 min 24 s
4 × 8, width 64246,812,99734,440,70464385,1866 h 38 min
  • 64 data sets cost little more than 16On the 15.4M-gate grid, 64 data sets took 884 s against 848 s for 16: four times the verification for 4% more time.
  • Shape does not matter247.6M gates built as 384 narrow stages and 246.8M as 32 wide ones ran within 5% of each other.
  • Switching activity for powerA SAIF file for all 246.8M nets, 14.8 GB, costs 37% more simulation time. Every net's time at 0, 1 and X adds up to the run's duration exactly.
  • A quarter billion gates in 4 GBThe 246.8M-gate run used 4.1 GB of the 8 GB card, about twelve bytes a gate plus the state of all 64 data sets.

The runs in the table are on the dense configuration, before the router, tokens in and out and clock gating were added. The final chip, with all of them, has been checked on the GPU against the C reference on every data set, from 1.2M gates up to a 259M-gate grid.

Next

The same chip at 2 billion gates on a single datacentre GPU, and switching activity from the final chip's routed runs, where idle experts are clock-gated and power depends on the data.