← All writing
Paper Breakdown

Attention Is All You Need: what the paper actually says

Rereading Vaswani et al. (Google Brain, 2017) while tracing why a pre-norm model trained faster than its post-norm counterpart.

I was debugging a training instability in a medium-scale transformer. The model diverged around step 8,000, consistently, regardless of batch size or gradient clipping threshold. The engineer who wrote it had followed "the original transformer architecture." When I actually pulled up "Attention Is All You Need" and compared, I found the original uses post-layer normalization — LayerNorm after the residual addition, not before. Every modern LLM (GPT-2 onward, LLaMA, Mistral) uses pre-norm, which trains more stably. The "original transformer" engineers implement from memory is already a hybrid of 2017 and 2019 improvements.

This made me reread the paper carefully. What I found: the architecture is more specific than the textbook version, the design choices are more reasoned than they look, and several decisions in the paper have been quietly overturned by subsequent work in ways that aren't always acknowledged.

The problem Vaswani et al. were actually solving

The paper's framing is often reduced to "transformers are better than RNNs." The actual problem is more specific: sequence transduction with long-range dependencies, trained at scale on parallel hardware.

In 2017, sequence-to-sequence models for machine translation were dominated by LSTM encoder-decoders with attention mechanisms (Bahdanau et al., 2014). These worked, but had two coupled problems:

Recurrent computation is sequential per timestep. To process a sequence of length N on a GPU with thousands of parallel compute units, you'd execute N sequential steps regardless. Training a deep LSTM on WMT English-German with 4.5 million sentence pairs took weeks and couldn't be meaningfully accelerated by adding more hardware.

Long-range dependencies were partially addressed by attention, but the attention was added on top of the RNN — the RNN still compressed the sequence into hidden states that carried all context through a fixed-size bottleneck. You were attending over compressed representations, not over raw input positions.

Vaswani et al.'s core claim, stated plainly in the abstract: "the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely." The operative word is "solely." What they replaced recurrence with, and how they made it work, is what the paper actually specifies.

Scaled dot-product attention: the specific design

The attention mechanism computes:

Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V

This looks simple, but each part is a specific design choice.

Q, K, V as projections of the same sequence. For encoder self-attention, Q, K, and V are computed by applying three separate learned linear projections to the same input sequence X:

Q = X W_Q,  K = X W_K,  V = X W_V

where W_Q, W_K, W_V ∈ R^(d_model × d_k). Every position can directly "query" every other position by computing a similarity score between its Q vector and the other position's K vector.

The 1/sqrt(d_k) scaling factor. This is specific and motivated. The paper explains: for large d_k, dot products grow large in magnitude, pushing softmax into regions with very small gradients. If Q and K are vectors of independent unit-variance random variables, the dot product Q·K has variance d_k. Dividing by sqrt(d_k) normalizes the variance back to 1, keeping softmax out of saturation.

The practical consequence: without this scaling, attention weights collapse toward near-one-hot distributions (one key dominates), gradients vanish, and training stalls. The paper includes an ablation showing that removing the scaling degrades performance, particularly at higher d_k values.

Softmax normalization over all positions. The output is a weighted sum of V vectors where weights are softmaxed similarity scores. Weights sum to 1 and are non-negative — a soft lookup over the full input. This is what gives attention its "soft lookup table" character: a query retrieves a blend of values weighted by similarity to each key.

Multi-head attention: the design the paper actually uses

The paper doesn't use a single attention function. It uses multi-head attention:

MultiHead(Q, K, V) = Concat(head_1, ..., head_h) W_O
head_i = Attention(Q W_Qi, K W_Ki, V W_Vi)

The hyperparameters in the base model: h = 8 heads, d_model = 512, d_k = d_v = d_model/h = 64.

Why multiple heads? The paper's reasoning: a single attention function would produce one similarity metric over d_model-dimensional keys and values. Different aspects of similarity may need different projections. An 8-head attention with d_k = 64 uses 8 learned projections onto 64-dimensional subspaces, computes attention in each, then concatenates the results and projects back.

What the heads actually learn isn't specified in the paper — that's empirical. Subsequent work (Voita et al., 2019; Clark et al., 2019) found that heads specialize: some attend to syntactic relations, some to positional patterns, some to coreference chains. Voita et al. also found that 8 of 16 heads in a 6-layer encoder could be pruned with minimal quality loss. The paper designed multi-head for representational capacity; specialization emerged.

Total computation cost: d_model² parameters per attention layer — identical to a single-head attention with full d_model dimensionality. Multi-head doesn't increase FLOPs; it restructures them.

Positional encoding: the sinusoidal choice and why

Self-attention is permutation-invariant. Shuffle the input sequence and the attention output changes only in corresponding positions — the operation has no built-in notion of order. Without positional information, the encoder can't distinguish "the dog bit the man" from "the man bit the dog."

The paper's solution: add a positional encoding to each token embedding before the first layer. The specific encoding uses sinusoids:

PE(pos, 2i)   = sin(pos / 10000^(2i/d_model))
PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model))

where pos is the position and i is the dimension index.

The paper's reasoning: for any fixed offset k, PE(pos + k) can be expressed as a linear transformation of PE(pos). This means the model can potentially learn to attend to relative positions rather than just absolute ones, and encodings can generalize to sequence lengths not seen during training.

The paper compares sinusoidal to learned positional embeddings and finds "nearly identical results," choosing sinusoidal partly for the extrapolation property and partly because it requires no learned parameters.

What the paper doesn't say: sinusoidal encodings don't extrapolate as cleanly in practice as the math suggests. Models trained on 512-token sequences perform poorly on 1,024-token sequences even with sinusoidal encodings, because attention patterns at long range are undertrained regardless of how the position is encoded. This is why RoPE (Su et al., 2021) and ALiBi (Press et al., 2022) replaced sinusoidal encodings in modern LLMs. The original choice was reasonable given 2017's knowledge; it's just not what anyone uses now.

The feed-forward sublayer: the part people underestimate

Each transformer layer has two sublayers: attention, and a position-wise feed-forward network:

FFN(x) = max(0, x W_1 + b_1) W_2 + b_2

Two linear layers with a ReLU activation. The inner dimension is 4 × d_model — for the base model with d_model = 512, the inner dimension is 2,048. For the large model (d_model = 1,024), it's 4,096.

The 4× expansion factor is a specific design choice that has persisted through nearly every transformer variant. The paper doesn't deeply justify it; it was a hyperparameter that worked well. Subsequent scaling law research found that over-parameterized FFN layers are beneficial — you'd often prefer a wider FFN to a deeper stack given the same FLOPs budget. The 4× ratio turns out to be a reasonable prior.

What the FFN is doing: the attention layer aggregates information across positions; the FFN applies a position-wise nonlinear transformation after aggregation. It doesn't mix positions — each position's representation is updated independently. This "memory" framing (attention reads context, FFN applies stored computations) is speculative, but ROME (Meng et al., 2022) found that factual associations in GPT models are stored primarily in FFN weights, which is consistent with it.

Modern LLMs replace ReLU with SwiGLU or GELU and change the inner dimension — LLaMA uses ~2.7× rather than 4×, adjusting for the three-matrix SwiGLU expansion. The structure is the same; the specific recipe changed.

Post-norm vs. pre-norm: the detail that breaks training

The paper specifies post-norm: apply LayerNorm after adding the residual:

x = LayerNorm(x + Sublayer(x))

Modern transformers use pre-norm:

x = x + Sublayer(LayerNorm(x))

Post-norm transformers train less stably because the input to each sublayer has variable scale: the residual path adds unnormalized values, and the early training signal is sensitive to initialization. Pre-norm ensures each sublayer receives normalized input, making gradient flow more predictable.

This is load-bearing. The paper requires a specific warmup schedule to make post-norm work:

lr = d_model^(-0.5) * min(step^(-0.5), step * warmup_steps^(-1.5))

with warmup_steps = 4,000. Linear warmup for the first 4,000 steps, then inverse-square-root decay. This schedule is necessary for the original post-norm architecture — without it, training diverges. Pre-norm transformers are far less sensitive to warmup and can often be trained with a cosine schedule without elaborate warmup.

If you implement the paper's post-norm and see training diverge around steps 5,000-10,000, check whether your warmup schedule precisely follows this formula. The formula isn't a suggestion; it was tuned specifically for post-norm stability.

What the paper actually proved in ablations

Section 6 has an ablation study that's often skipped. The key results:

Head count matters, but less than you'd expect. Reducing from 8 to 1 head (same d_model) drops BLEU by 0.9 on EN-DE. Going from 8 to 16 heads shows no consistent gain. The empirical sweet spot was approximately d_model / 64 heads.

Key dimension matters more than value dimension. Reducing d_k hurts more than reducing d_v. The interpretation: the quality of the similarity lookup (determined by key projections) is a harder bottleneck than the richness of retrieved information.

Label smoothing hurts perplexity but helps BLEU. The paper uses ε = 0.1 — target distribution puts 0.1 probability mass on non-target tokens. Perplexity goes up (the model is less confident) but BLEU improves. Smooth targets prevent over-concentration on the argmax prediction and improve calibration on the real evaluation metric. This is a general lesson: training loss and evaluation metric are not always aligned.

Big model results: 28.4 BLEU on WMT 2014 EN-DE, outperforming all prior single models and ensembles. Base model: 27.3 BLEU at roughly 1/8 the training cost.

The encoder-decoder structure

The paper's architecture is a full encoder-decoder. The encoder has N = 6 identical layers of self-attention + FFN. The decoder has N = 6 layers of masked self-attention + cross-attention + FFN.

Masked self-attention in the decoder: the decoder processes output autoregressively — each token attends only to previous output tokens. During training, teacher forcing provides the full output sequence at once, but the mask ensures position i can only attend to positions < i. Without masking, the model would attend to future tokens during training and cheat.

Cross-attention: the decoder's second sublayer attends to the encoder output. Q comes from the decoder; K and V come from the encoder. This is how the decoder incorporates source information — it can query any part of the encoder's representation at each decoding step, without an LSTM-style bottleneck. This mechanism was introduced in Bahdanau et al. (2014); the paper applies it on top of transformer representations rather than LSTM hidden states, and the encoder produces richer representations through full-sequence self-attention first.

For decoder-only LLMs (GPT family, LLaMA), the encoder-decoder structure is absent. These models use only the decoder stack with cross-attention removed. The paper's translation framing made encoder-decoder natural; for generative prediction over a single sequence, decoder-only is sufficient and architecturally simpler. The paper's full encoder-decoder is rarely used in modern deployments — most production LLMs are decoder-only or encoder-only.

Production failure modes from the original design

Post-norm divergence beyond 12 layers. Scale depth with the original post-norm and training diverges without significant care. The warmup schedule is load-bearing. Using Adam with default hyperparameters and no warmup often produces a model that trains without error but optimizes into a degenerate minimum early on. Switching to pre-norm removes this failure mode.

Quadratic memory in sequence length. Attention materializes an N × N score matrix. At N = 512, this is trivial. At N = 32,768, it's enormous. The paper doesn't discuss this because 512 tokens was state-of-the-art for translation in 2017. Every long-context optimization — FlashAttention, ring attention, sliding window attention — is a response to this quadratic scaling that the original paper didn't need to address.

Sinusoidal PE doesn't extrapolate cleanly. Models trained at max length 512 don't generalize to 1,024 tokens despite the paper's mathematical argument. The issue is undertrained attention patterns at long range, not the encoding itself. Extrapolation beyond training length requires explicit handling (RoPE, ALiBi, or careful training on progressively longer sequences).

Softmax attention cannot produce null results. Every query attends somewhere — softmax always produces a positive-weight distribution over all keys. If no key is relevant to a given query, the model still produces a noisy weighted average. For tasks where "I don't know" is the right answer for certain positions, this is a structural problem. Sparse attention or gated alternatives address it but aren't in the original paper.

When NOT to use self-attention

Short sequences with local structure. For sequences under ~64 tokens where dependencies are primarily local, CNNs or even simple MLPs often match transformers while being significantly faster. Attention computes O(N²) pairwise interactions; if you only need O(N) local interactions, you're doing unnecessary work.

Memory-constrained inference on variable-length sequences. Quadratic memory scaling means you can't naively serve very long sequences without FlashAttention or equivalent. On CPU or edge hardware where memory is the constraint, an LSTM of comparable quality often has a better memory profile despite the FLOPs disadvantage. The transformer's memory advantage only appears at the system level when batching many requests.

Very long sequences with purely local dependencies. Document classification where only a few key sentences matter, time-series forecasting with strong periodicity, genomic sequence analysis with purely local motifs. For these, full-sequence attention is expensive computation providing marginal benefit over a sliding-window or hierarchical approach.

When you need O(N) time and space. State space models (Mamba, S4) achieve O(N) time and space complexity with comparable quality on certain tasks. For sequences longer than ~64K tokens on memory-limited hardware, SSMs can be preferable. The original transformer architecture assumes you can materialize the N×N matrix or run a tiling scheme like FlashAttention.

What the paper actually changed

The paper didn't invent attention, positional encodings, residual connections, or layer normalization. Each came from prior work.

What it did: demonstrate that these components, assembled without recurrence, were sufficient — and superior — for the top sequence-to-sequence tasks of the era. The "all you need" in the title is specific: you need attention (and the rest of the transformer machinery), but you don't need recurrence. That was the empirical claim, and they proved it on WMT.

The specific recipe — 8 heads, d_k = 64, FFN inner dimension = 4 × d_model, post-norm, sinusoidal PE, N = 6 layers — was reasonable for 2017, validated by ablations on WMT. Most of it has been superseded: pre-norm replaced post-norm, RoPE replaced sinusoidal PE, GQA or MQA replaced full multi-head for serving efficiency, SwiGLU/GELU replaced ReLU in the FFN.

The core insight — every position attending directly to every other position, in parallel, with learned projections — is what persisted. That's what every LLM in production is built on. The specific recipe in the paper is a first-generation prototype.


If you're debugging training instability in a transformer, first check whether you're using pre-norm or post-norm — the paper uses post-norm, but modern frameworks default to pre-norm, and they behave very differently under the same optimizer settings. If you're reading about a modern LLM architecture and wondering how it differs from the original, the RoPE post and the GQA post cover the two most significant architectural changes that happened between 2017 and now.