Transformer Block Architecture Explained
The Transformer Block
What is a Transformer Block?
A transformer block is the repeating unit that makes up the entire model. Each block contains two sub-layers: a multi-head self-attention layer (which lets tokens communicate) and a feed-forward network (which processes each token independently). Both sub-layers use residual connections (adding the input back to the output) and layer normalization to keep training stable. GPT-3 stacks 96 of these blocks; LLaMA 2 70B uses 80. The depth — number of stacked blocks — determines how many rounds of "thinking" the model can do.
The Transformer Block
A transformer is a stack of identical blocks. GPT-2 has 12. GPT-3 has 96. Llama-2 has 32. Each block sees the same shape of data and refines it — every single one running the same pattern.
GPT-2 small — 12 blocks, d_model 768
Same shape in and out, every block — which is exactly why they stack.
One Block, Two Sub-Layers
Every transformer block contains exactly two sub-layers:
- Multi-Head Attention (MHA) — lets each token gather information from all other tokens
- Feed-Forward Network (FFN) — transforms each token independently, applying learned knowledge
But attention and FFN don't work alone. Each one has two helpers:
- Layer Normalization — as data passes through 96 blocks, numbers can grow huge or shrink to near zero. LayerNorm rescales them to a healthy range before each sub-layer — like adjusting the volume so the signal stays clear.
- Residual connection — data enters a sub-layer (e.g., attention), but a copy of the data before the sub-layer is kept aside. After the sub-layer finishes, its output is added to that saved copy. So the result is: original + what the sub-layer learned. Nothing from the original is lost — the sub-layer only adds new information on top. Why? Without this, a 96-layer model would gradually lose the original signal — each layer would overwrite the previous one. The residual connection guarantees that information can flow straight from the first layer to the last, untouched.
Look at the simulation. The two dashed panels are the two sub-layers — everything inside a panel belongs to one of them, and the block is nothing but those two panels stacked. Data flows top to bottom. The dashed line down the left of each panel is its residual: the input bypasses the sub-layer entirely and is added back at the "+" box. Click any box to read what it computes.
The 2-Line Block
Andrej Karpathy's nanoGPT captures the entire block in two lines:
x = x + attn(ln1(x))
x = x + ffn(ln2(x))
That's it. Every major LLM — GPT-4, Claude, Gemini, Llama — runs this pattern billions of times per forward pass. The elegance is in the simplicity: normalize, transform, add back.
Stacking Blocks
A single block captures limited patterns. Stack many blocks and the model builds up understanding layer by layer — each layer sees the output of all layers before it:
Each block inherits everything from the blocks before it — the "add back" pattern (residual connections) means original information is never lost, while new understanding accumulates on top with each layer.
Understanding one block means understanding all of them. The same two lines — attention residual, FFN residual — repeat unchanged from the first block to the last, whether a model has 12 layers or 96.
Layer Normalization
Layer Normalization
Before each sub-layer runs, the input passes through a normalization step. Without it, activations accumulate errors across 96 blocks — numbers grow too large or shrink toward zero, and gradients become untrainable.
How Layer Normalization Works
For each token vector independently, Layer Normalization:
- Computes the mean μ across all
d_modeldimensions - Computes the standard deviation σ
- Subtracts the mean and divides by σ — centering at zero, unit variance
- Applies learned scale γ (gamma) and shift β (beta) to restore representational capacity
LN(x) = γ · (x − μ) / σ + β
The key detail: normalization happens per token, not per batch. Each token's d_model-dimensional vector is normalized independently. This is what distinguishes LayerNorm from BatchNorm — it works correctly with any sequence length and batch size, including batch size 1.
Analogy: Auto-Adjusting Brightness
Think of each token's vector as a photo. Layer Normalization is auto-adjust — it standardizes brightness and contrast regardless of how dark or washed-out the original was. The learned γ and β parameters let the model re-apply its preferred "exposure" after normalization.
RMSNorm — The Modern Variant
Modern LLMs — Llama, Mistral, Falcon — replace LayerNorm with RMSNorm:
RMSNorm(x) = γ · x / RMS(x) where RMS(x) = √(mean(x²))
Two changes from LayerNorm:
- No mean subtraction — skip centering entirely
- No β shift — only one set of learned parameters (γ)
RMSNorm is faster and uses less memory. Empirically it matches LayerNorm quality while cutting the normalization compute nearly in half.
RMSNorm saves compute per layer. Multiply that saving across 96 layers and billions of tokens per training run, and the difference is substantial. This is why every frontier model since Llama adopted it.
Residual Connections
Residual Connections
Look at the + in x = x + attn(ln1(x)). That addition is a residual connection — and it is arguably the most important architectural decision in the transformer.
Adding a Delta, Not Replacing
Without the residual:
x = attn(ln1(x)) # sub-layer replaces x entirely
With the residual:
x = x + attn(ln1(x)) # sub-layer adds a change to x
The sub-layer only needs to learn what to change — the delta Δ. The original information flows through untouched. This is a much easier learning problem than reconstructing x from scratch.
Gradient Flow: The Highway
Deep networks have a vanishing gradient problem. Backpropagation multiplies gradients through each layer — after 96 multiplications, the gradient reaching the first block can be nearly zero, and the early layers learn nothing.
Residual connections solve this by creating a gradient highway. The gradient can flow directly through the addition operation without any transformation — it passes through every block unchanged. Deep training becomes feasible.
This insight came from ResNets in 2015 (He et al.) — a computer vision architecture. The transformer borrowed it directly.
Analogy: Tracked Changes
Writing a document in "tracked changes" mode: every edit is layered on top of the original text, not replacing it. You can always see what changed and roll back. Each transformer block makes tracked changes to the representation — the original signal is never lost.
Without residual connections, training 96 layers would be nearly impossible — gradients would vanish long before reaching the early blocks. The residual connection is the reason that "more layers = more capable" actually works at scale.
Further reading
- Deep Residual Learning (He et al., 2015) — the original ResNet paper that introduced residual connections
- Understanding the Difficulty of Training Deep Feedforward Neural Networks (Glorot & Bengio, 2010) — the vanishing gradient problem explained
- The Vanishing Gradient Problem (Wikipedia) — accessible overview of why deep networks struggle without skip connections
The Feed-Forward Network
The Feed-Forward Network
Attention and FFN do two very different jobs:
- Attention = a group meeting. Every token talks to every other token and gathers information: "who is relevant to me?" After attention, the token for "sat" now knows about "cat" and "mat."
- FFN = individual thinking. Each token goes back to its own desk and processes what it just learned — alone, with no further communication. The token for "sat" takes its gathered context and thinks: "given that I know about cat and mat, what should my updated representation be?"
Attention is communication. FFN is computation.
Attention — the group meeting
Every token talks to every other token. Lines are the exchange; sat is asking "who here is relevant to me?"
Same four tokens both times — watch them move rather than reappear. And the desks are identical on purpose: the FFN applies the same weights at every desk. What differs is only what each token brought back from the meeting.
How do you read [batch, seq, 768]?
Shapes like this are about to appear in every paragraph below, so it is worth thirty seconds. It is one array with three axes, and the plainest way to picture it is a stack of paper:
A stack of batch sheets. Each sheet has seq rows. Each row is 768 numbers, and one row is one token.
batch— how many separate conversations the server is running through the model at the same time. Your request is one of them: on a quiet server this is 1, under load it might be 64. Why stack them at all? Because the expensive part of a forward pass is hauling the weights out of GPU memory, not the arithmetic — once a weight matrix has been loaded, applying it to 64 sequences costs barely more than applying it to one. Batching is throughput bought almost for free. None of this belongs to the model; it is a serving decision, remade fresh every step.seq— how many tokens are in the sequence: your prompt plus everything generated so far. It grows by one with each new token, and its ceiling is the context window.768— how many numbers it takes to represent one token. This isd_model, the model's hidden width, and unlike the other two it is fixed — 768 for GPT-2, 4096 for Llama-2 7B, 12288 for GPT-3. It was chosen when the model was designed and is baked into the shape of every weight matrix.
sequences at once — the server picks this
tokens in each — your input picks this
numbers per token — the model was built this wide
[2, 5, 768]
sequence 1 of 2 — 5 tokens
One row — the token sat — opened up:
Change batch and seq all you like — every request does. 768 has no button because it is welded into every weight matrix in the model; changing it would mean a different model.
That split — two axes set at runtime, one welded to the architecture — is why only the last one gets a name (d_model) and the other two are just counts.
It also explains exactly who is allowed to touch what:
| Axis | What mixes along it |
|---|---|
batch | Nothing, ever. Sequence 3 cannot influence sequence 1 — that would be one user's data leaking into another's answer. Batching is purely a hardware trick: the GPU loads the weights once and applies them to everyone present. |
seq | Attention, and only attention. This is the group meeting from above — the one axis along which tokens are allowed to see each other. |
768 | LayerNorm, the FFN, and attention's own projections. All of them work on one token's 768 numbers at a time. |
What does each one actually output?
Both return a tensor of exactly the same shape they were given — [batch, seq, 768], one vector per token. Neither changes the shape. What differs is what the output depends on: attention's result for a token is built from every token in the sequence, the FFN's from that token alone.
But the shape is the less interesting half. The important part is that neither sub-layer returns the token's new representation. Look again at where the call sits:
x = x + attn(ln1(x))
x = x + ffn(ln2(x))
attn(...) is on the right-hand side of a +. What it returns is an update — the contribution from context — which is then added to what was already there. The token itself is the running sum in the residual stream, and each sub-layer only ever adds to it.
So when we say "after attention, sat knows about cat and mat," that describes the stream after the addition. Attention's own output does not contain sat at all; it contains only what the other tokens contributed. Same for the FFN: what comes back is retrieved knowledge to add, not a replacement token.
Updating sat — reads all four tokens
x = x + attn(LN(x))
Every token feeds sat's update. What comes out is only the contribution from context — sat itself is not in it.
8 of 768 dimensions. The middle row is what the sub-layer returns — same shape as its input, and added to it, never replacing it.
Toggle between the two and watch the middle row — that row is the sub-layer's actual output. Notice that switching to the FFN carries the previous result forward as its input: the two sub-layers chain through the residual stream, each adding its own increment.
Expand → Activate → Contract
The FFN is a two-layer MLP with a specific shape pattern:
Input: [batch, seq, 768] # d_model
Expand: [batch, seq, 3072] # 4 × d_model
Activate: [batch, seq, 3072] # non-linearity
Contract: [batch, seq, 768] # back to d_model
Why Expand Then Contract?
The expanded space is a temporary workspace. With 3072 dimensions, the model can form rich intermediate combinations that couldn't be expressed in 768 dimensions. Then it contracts back — keeping only what's useful, discarding the rest.
Contraction back to d_model is not optional. It's what keeps the transformer shape-preserving: every block outputs the same dimensions it takes as input, so blocks can stack indefinitely.
FFN as Knowledge Storage
Research from Geva et al. (2021) showed that FFN layers function as key-value memory banks. The first weight matrix W₁ acts as keys — patterns to match. The second matrix W₂ acts as values — what information to retrieve when a pattern matches.
Factual associations like "Paris is the capital of France" are stored in FFN neurons. Attention routes information; FFN stores and retrieves it.
Activation Functions: ReLU → GELU → SwiGLU
The activation function sits between the expand and contract steps. It decides which signals pass through and which get suppressed. Toggle the curves below to compare:
Each generation solved a problem with the previous one:
-
ReLU (2012) — the breakthrough that made deep learning work. Everything negative becomes exactly zero:
max(0, x). Fast and simple. But it has a flaw: once a neuron outputs negative, its gradient is exactly zero — the neuron "dies" and stops learning forever. In a 96-layer model, many neurons die. -
GELU (2016, used in GPT/BERT) — fixes the "dying neuron" problem. Instead of a hard cutoff at zero, it smoothly curves — small negative values still get through slightly. Unlike ReLU, whose gradient is zero across the entire negative range, GELU's gradient is zero at only one point — so a neuron is never permanently stuck the way a dead ReLU neuron is. Toggle ReLU off and GELU on in the graph above to see the smooth transition near x=0.
-
SwiGLU (2020, used in Llama/Mistral/PaLM) — fixes a different problem: in ReLU and GELU, the activation function is fixed — it applies the same transformation regardless of input. SwiGLU adds a learned gate: a second branch of weights decides which dimensions to keep and which to suppress, adapting per input. The model can learn "for this token, dimensions 5-10 matter; for that token, dimensions 50-60 matter." More parameters, but measurably better quality.
The FFN contains roughly 2/3 of a transformer's total parameters. In GPT-3, the four FFN weight matrices per block dominate the 175B parameter count. This layer is where the model's "knowledge" is stored — the attention mechanism retrieves context, but FFN is the library.
Pre-Norm vs Post-Norm
Pre-Norm vs Post-Norm
There are two ways to place Layer Normalization inside a residual block. This ordering looks like a minor implementation detail — at 96 layers deep, it determines whether training converges at all.
Two Orderings
Post-Norm (original 2017 paper):
x = LN(x + sublayer(x)) # normalize after residual add
Pre-Norm (modern standard):
x = x + sublayer(LN(x)) # normalize before sublayer
The difference: in Post-Norm, the normalization sits outside the residual addition and sees the combined signal. In Pre-Norm, normalization sits inside — the residual path bypasses it entirely.
Why Pre-Norm Won
In Pre-Norm, the residual connection carries the unmodified input directly to the addition — LN never touches the highway. Gradient magnitude is preserved across every block.
In Post-Norm, every gradient must pass through the normalization operation at each layer boundary. This destabilizes training at depth, requiring careful learning rate warm-up schedules and lower initial rates. At 96 layers, the instability compounds.
Pre-Norm trains reliably without special warm-up schedules, which is why it's the dominant choice in modern LLMs — GPT-2, GPT-3, Llama, and Mistral all use it. Post-Norm hasn't disappeared, though: depth-scaled variants like DeepNorm (used to train 1,000-layer transformers) and sandwich/double-norm hybrids keep it in play where stability at extreme depth or long context matters.
Karpathy's Code
In nanoGPT, Karpathy writes it directly:
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
ln_1 and ln_2 are called inside the sub-layer calls — LN runs first, then the sub-layer, then the residual add. That's Pre-Norm. The names ln_1 and ln_2 reflect that there are two norms per block, one per sub-layer.
Toggle the Pre-Norm / Post-Norm switch in the simulation to see where the normalization boxes appear in the data flow.
A trivial ordering difference — normalize before or after — but at 96 layers deep it determines whether the model trains at all. Pre-Norm is now the dominant standard; Post-Norm survives mainly in depth-scaled variants like DeepNorm rather than in mainstream stacks.
Modern Variants & Scale
Modern Variants & Scale
The core transformer block has not changed since "Attention Is All You Need" in 2017. What has changed are the components inside it — each swap improving quality, speed, or both.
What Changed: 2017 → Today
Five components swapped out, same overall structure:
- Normalization: LayerNorm (γ, β) → RMSNorm (γ only) — simpler, faster, same quality
- FFN activation: ReLU → SwiGLU (gated) — learned gating, better quality
- Position encoding: Sinusoidal (absolute) → RoPE (rotary, relative) — handles longer sequences
- Attention: Standard (O(n²) memory) → Flash Attention (IO-aware) — 2-5× faster, less memory
- Norm placement: Post-Norm → Pre-Norm — trains stably at 96+ layers
RoPE (Rotary Position Embedding) encodes relative position by rotating query and key vectors — it naturally extrapolates to longer sequences than seen during training. Covered in the Embeddings module.
Flash Attention is not a new algorithm — it computes the same attention scores — but it reorganizes memory access to tile computations within GPU SRAM instead of writing intermediate results to HBM. The result: 2–5× faster, far less memory. Every production LLM uses it.
Weight tying shares the embedding matrix and the output projection matrix (they are both d_vocab × d_model). The output projection maps the final hidden state back to logits over the vocabulary. Using the same weights as the input embedding eliminates one full copy of that large matrix — saving hundreds of millions of parameters in large models.
Parameter Counting
A single transformer block contains approximately 12d² parameters (where d is d_model):
- Attention: Q, K, V, O projection matrices → ~4d²
- FFN: two weight matrices with 4× expansion → ~8d²
For GPT-3 (d = 12288):
12 × 12288² ≈ 1.8B parameters per block
1.8B × 96 blocks ≈ 172B parameters
The remaining ~3B come from embeddings and layer norms — the blocks dominate.
The transformer block's architecture has been stable since 2017. What changed were the components within: RMSNorm for speed, SwiGLU for quality, RoPE for length generalization, Flash Attention for efficiency. The two-line structure remains unchanged.
Further Reading
- Andrej Karpathy — nanoGPT — minimal GPT implementation; the
model.pyBlock class is 20 lines - Jay Alammar — The Illustrated GPT-2 — visual walk-through of GPT-2's block structure and parameter shapes
- Lilian Weng — The Transformer Family v2 — comprehensive survey of architectural variants through 2023