A ground-up construction, one piece at a time
The task this document builds around is translation: given the English sentence "the black cat sleeps," produce its Spanish equivalent, "el gato negro duerme." We use translation as our running lens throughout because it exercises every part of the architecture naturally — encoder, decoder, and the connection between them. The ideas generalize well beyond it.
Notice something immediately: the word order is not the same. In English, the adjective comes before the noun — black cat. In Spanish, it comes after — gato negro. A system that simply translates word by word, left to right, will get this wrong. A good translation model needs to understand the structure of the whole source sentence before it can produce the target. That structural understanding is exactly what the transformer is built to provide.
The transformer was introduced in 2017 in a paper titled "Attention Is All You Need" (Vaswani et al.) and rapidly became the dominant architecture for language tasks. Before it, the standard approach was recurrent neural networks — architectures that processed text one word at a time, left to right, passing a hidden state forward at each step. The transformer replaced that sequential processing with a fully parallel mechanism called attention, which lets every token in a sequence directly interact with every other token in a single step. This made training dramatically faster and gave models a much easier time capturing relationships between distant words.
Today, virtually every large language model you interact with — GPT-4, Claude, Gemini, LLaMA, and others — is a transformer or a close descendant of one. The architecture has also spread well beyond text: transformers are now used in image recognition (Vision Transformers), protein structure prediction (AlphaFold), audio generation, and more. Learning how a transformer works is, in large part, learning the foundation of modern AI.
Rather than presenting the full transformer architecture and then explaining what each piece does, we build it up from scratch — one component at a time. Each version of the model will have multiple shortcomings. Rather than fixing them all at once, we identify the most important one, add the minimal change needed to address it, and move on. The next version then fixes the next most significant remaining problem, and so on. Every addition is motivated by one concrete, demonstrable failure in the version before it.
This progression has eight stages, called MVPs (minimum viable products). Every section follows the same structure:
Before we build anything, there is one distinction worth establishing now, because it runs through every section: the difference between training and inference. Training is the process of adjusting a model's parameters — its internal numerical weights — so that it gets progressively better at its task, given a large set of labeled examples. In our case those examples are paired English and Spanish sentences; the model sees both sides, measures how wrong its predictions are, and updates its weights accordingly. Inference is the process of passing a new input through a trained, fixed model to produce the desired output. The model's parameters do not change during inference — only the input changes.
We have a large dataset of paired sentences — English alongside correct Spanish translations. The model sees both sides, makes a prediction, measures how wrong it is against the known answer, and adjusts its internal numbers to do better next time. This happens millions of times.
Both sides known. Error measured. Weights updated.
We have only the source sentence. The model must produce the translation from scratch, with no reference answer to compare against. There is no error signal, no weight update — just a forward pass through the learned model.
Source only known. Target must be generated.
The mechanical consequences of this distinction — especially in how the decoder behaves differently in each mode — are one of the most important and least-explained aspects of the transformer. We will return to it with full technical detail in Part 4, when the decoder is introduced.
Before we build the transformer, we need to answer a more basic question: how do you feed a sentence into a neural network at all? Networks operate on numbers. Strings of text are not numbers. The answer is embeddings — and they are the foundation everything else is built on.
An embedding is simply a list of numbers — a vector — assigned to each word in the vocabulary. We maintain a large table, one row per vocabulary word, where each row is that word's vector. Looking up a word's embedding means finding its row in that table.
These vectors are learned: they start as random numbers and are updated during training until words with similar meanings end up with similar vectors. "Cat" and "kitten" end up nearby; "cat" and "airplane" end up far apart. But at the very start of training, every word has a random vector.
In this document we use 3-dimensional embeddings to keep the numbers readable. Real models use 512 or more dimensions. Here is what the embeddings for our four source tokens might look like early in training:
We now have four vectors — one per token. But to produce an output we need a single fixed-size vector. Predicting the next word means scoring every word in the Spanish vocabulary and picking the most likely one — a form of classification over V classes, one per vocabulary word. A variable-length sequence of token vectors cannot be fed directly into that scoring layer; it needs one fixed-size vector. The simplest possible answer: average them. Add up all four embedding vectors and divide by four.
| Token | dim 0 | dim 1 | dim 2 |
|---|---|---|---|
| the | 0.10 | 0.30 | −0.20 |
| black | 0.80 | −0.40 | 0.60 |
| cat | 0.50 | 0.90 | 0.10 |
| sleeps | −0.30 | 0.20 | 0.70 |
| average → | 0.28 | 0.25 | 0.30 |
This averaged vector — the context vector — is then passed through two operations to produce a prediction. First, a linear projection: a learned matrix \(W\) maps the context vector from the embedding space (\(\mathbb{R}^3\) in our example, or \(\mathbb{R}^d\) in general) into the vocabulary space (\(\mathbb{R}^V\), where \(V\) is the number of words in the Spanish vocabulary — potentially tens of thousands). "Linear" here means the transformation is a matrix multiplication with no nonlinear activation; "projection" means we are mapping between two different vector spaces. The result is a \(V\)-dimensional vector of raw scores, one per Spanish word: a higher score means the model thinks that word is more likely given this context. Second, a softmax converts those raw scores into a proper probability distribution summing to 1. The word with the highest probability is the model's prediction for the first Spanish token.
The context vector is the elementwise average of the \(n\) token embeddings:
\[ c = \frac{1}{n} \sum_{i=1}^{n} e_i \]The prediction is produced by a linear projection followed by softmax:
\[ p(y) = \text{softmax}(W c + b) \]where \( W \in \mathbb{R}^{V \times d} \) is a weight matrix and \( b \in \mathbb{R}^{V} \) is a bias vector — both are learnable parameters updated during training. Row \(i\) of \(W\) is a learned \(d\)-dimensional vector associated with vocabulary word \(i\); the dot product of row \(i\) with \(c\) scores how well the context matches that word. Softmax then converts the full vector of \(V\) raw scores into probabilities:
\[ \text{softmax}(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{V} e^{z_j}} \]We have the paired example. We compute context vector
[0.28, 0.25, 0.30], pass it through the linear layer, and get a
probability over all Spanish vocabulary words. We compare against the known
first word — el — measure the error, and backpropagate. The
embedding table and linear weights are both updated. Repeat for millions of
sentence pairs.
We have only "the black cat sleeps." We compute the same
context vector [0.28, 0.25, 0.30] and pass it through the same
linear layer — but now there is no known answer. We take the
highest-probability Spanish word as our prediction. There is no weight update;
everything is frozen.
Averaging embeddings is a functional starting point — the model can technically learn something — but it has three fundamental problems. Each one will be fixed by a specific, identifiable change in the sections that follow.
| Sentence | dim 0 | dim 1 | dim 2 |
|---|---|---|---|
| the black cat sleeps | 0.28 | 0.25 | 0.30 |
| sleeps cat black the | 0.28 | 0.25 | 0.30 |
In Part 1 we identified the root problem: averaging treats every word as equally relevant. Self-attention replaces that flat average with a learned, content-dependent weighting. Each token looks at every other token, scores how relevant it is, and builds a new representation by blending the ones that matter most to it.
Self-attention happens in three steps, applied simultaneously to every token.
Step 1: Score every other token. For token i, compute a dot product against every token j in the sequence — including itself. The dot product measures similarity: a high score means the two tokens are pointing in a similar direction in embedding space, and therefore that j is likely relevant to updating i's representation.
Step 2: Normalize into a probability distribution. Pass all scores for token i through softmax. This converts the raw scores into attention weights that sum to 1. A high weight on token j means token i will draw heavily from token j's embedding in the next step.
Step 3: Take the weighted sum. Compute a weighted sum of every token's embedding, weighted by the attention weights from step 2. The result, zi, is token i's new contextualized representation — it has been updated to reflect the content of the tokens around it, weighted by relevance.
For each token \(i\), compute raw scores against all other positions \(j\):
\[ \text{score}_{ij} = e_i \cdot e_j \]Normalize into attention weights via softmax (over all \(j\) for fixed \(i\)):
\[ \alpha_{ij} = \text{softmax}_j(\text{score}_{ij}) = \frac{e^{e_i \cdot e_j}}{\sum_{k} e^{e_i \cdot e_k}} \]Compute the contextualized representation as a weighted sum of all embeddings:
\[ z_i = \sum_j \alpha_{ij}\, e_j \]Note that \(j\) ranges over every position including \(i\) itself — a token is always a candidate to attend to itself. Also note that \(\alpha_{ij} \neq \alpha_{ji}\) in general: how much token i attends to token j is not the same as how much j attends to i.
We use the same four embeddings from Part 1. The attention computation produces a 4×4 score matrix — one row per query token, one column per key token — then softmax weights, then the final contextualized vectors.
Step 1: Score matrix (each cell is the dot product of the row token's embedding with the column token's embedding):
| the | black | cat | sleeps | |
|---|---|---|---|---|
| the | +0.14 | −0.16 | +0.30 | −0.11 |
| black | −0.16 | +1.16 | +0.10 | +0.10 |
| cat | +0.30 | +0.10 | +1.07 | +0.10 |
| sleeps | −0.11 | +0.10 | +0.10 | +0.62 |
The diagonal (highlighted) shows each token's self-similarity — how strongly it matches itself. "black" and "cat" have the highest self-scores because their embeddings have the largest magnitudes. "the", a low-information function word, has the smallest self-score.
Step 2: Attention weights (softmax applied row by row; each row sums to 1; bolder cell = highest weight in that row):
| the | black | cat | sleeps | |
|---|---|---|---|---|
| the | 0.27 | 0.20 | 0.32 | 0.21 |
| black | 0.14 | 0.51 | 0.18 | 0.18 |
| cat | 0.21 | 0.17 | 0.45 | 0.17 |
| sleeps | 0.18 | 0.22 | 0.22 | 0.37 |
A few things to notice. "black" attends most to itself (0.51) — its high self-score is so dominant that it draws more than half the weight. "sleeps" also attends most to itself (0.37), but much less exclusively — the remaining weight is spread across the other three tokens. "the" is the most distributed: it attends most to "cat" (0.32), slightly more than to itself (0.27). This is the content-dependent weighting that was impossible in Part 1 — no two rows look the same.
Step 3: Contextualized representations (weighted sum of embeddings using the weights above; compare original vs. updated):
| original embedding | after attention z | ||||||
|---|---|---|---|---|---|---|---|
| d0 | d1 | d2 | d0 | d1 | d2 | ||
| the | +0.10 | +0.30 | −0.20 | → | +0.28 | +0.33 | +0.25 |
| black | +0.80 | −0.40 | +0.60 | → | +0.46 | +0.03 | +0.42 |
| cat | +0.50 | +0.90 | +0.10 | → | +0.33 | +0.43 | +0.23 |
| sleeps | −0.30 | +0.20 | +0.70 | → | +0.20 | +0.24 | +0.38 |
Every token's representation has moved. "the" moved the most dramatically: it started at [+0.10, +0.30, −0.20] and landed at [+0.28, +0.33, +0.25] — the −0.20 in dimension 2 has flipped positive, pulled by the strong positive values of "cat" and "sleeps" in that dimension. "black", which attends mostly to itself, moved the least: its representation stayed close to its original, diluted only slightly by the other tokens.
This is the core of what contextualization means: each token's final representation is no longer just its own embedding — it is a blend of the whole sequence, weighted by how relevant each token decided the others were to it.
Self-attention fixes failure 1 from Part 1 (uniform weighting) and failure 3 (dilution). But failure 2 — permutation invariance — is still present. Consider a longer example:
The reason is structural, not incidental: without positional encoding, "flies" in position 1 uses the same query embedding as "flies" in position 6, and both compute dot products against the same global set of key embeddings, weighted the same way. Position simply does not exist as information anywhere in the computation. This will not be fixed until Part 5, when we add sinusoidal positional encodings that inject position information directly into the embeddings before attention runs.
The encoder has a clean property worth naming now because it contrasts sharply with the decoder: it behaves identically at training time and inference time. The source sentence is always fully known in both cases — whether you are training on a labeled translation pair or translating a new sentence from scratch. So the encoder always runs exactly once, in a single forward pass over all source tokens in parallel, with no sequential dependency.
Source sentence is known. All four tokens are embedded and processed in parallel. The 4×4 attention matrix is computed in one matrix multiplication; all four z vectors are produced simultaneously. No sequential steps — this is what makes the transformer trainable at scale.
Exactly the same computation. Source sentence is still fully known. The encoder produces z1…z4 in one forward pass, identically to training. The interesting train/inference asymmetry lives entirely in the decoder — which we will introduce in Part 4.
In Part 2, every token played three roles simultaneously using the same single vector: it queried ("what am I looking for?"), it offered itself as a key ("what do I advertise to others' queries?"), and it provided a value ("what do I actually contribute if someone attends to me?"). Using one vector for all three means the model cannot optimize them independently. A word cannot learn to ask a different kind of question than it answers.
The fix is one of the cleanest ideas in the transformer: give each role its own learned linear projection. Instead of one embedding per token, we maintain three separate weight matrices — \(W_Q\), \(W_K\), and \(W_V\) — each of which transforms the embedding into a specialized representation for its role.
The three projections work as follows for each token \(i\):
The attention score between token \(i\) and token \(j\) is now the dot product of \(i\)'s query with \(j\)'s key, divided by \(\sqrt{d}\):
The three projections for token \(i\):
\[ q_i = W_Q e_i, \qquad k_j = W_K e_j, \qquad v_j = W_V e_j \]Scaled dot-product attention:
\[ \alpha_{ij} = \text{softmax}_j\!\left(\frac{q_i \cdot k_j}{\sqrt{d}}\right), \qquad z_i = \sum_j \alpha_{ij}\, v_j \]Why divide by \(\sqrt{d}\)? If the components of \(q_i\) and \(k_j\) each have variance \(\sigma^2\), the dot product \(q_i \cdot k_j = \sum_{m=1}^{d} q_{im} k_{jm}\) has variance \(d\sigma^4\) — it grows with \(d\). As \(d\) increases, the scores become so large that softmax saturates: one score dominates, the weights collapse toward a one-hot vector, and gradients vanish. Dividing by \(\sqrt{d}\) normalises the variance back to \(\sigma^4\), independent of dimension size.
Note also that the output \(z_i\) is now a weighted sum of value vectors \(v_j\), not of the original embeddings \(e_j\). The value projection separates "what information I use to decide who to attend to" (keys and queries) from "what information I actually receive from attending" (values).
We apply three simple projection matrices to our four token embeddings. Each produces a different view of the same underlying embedding — the same sentence, seen through three different lenses:
| query q = WQe | key k = WKe | value v = WVe | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| d0 | d1 | d2 | d0 | d1 | d2 | d0 | d1 | d2 | |||
| the | +0.10 | −0.20 | +0.30 | +0.30 | +0.10 | −0.20 | +0.10 | +0.30 | +0.20 | ||
| black | +0.80 | +0.60 | −0.40 | −0.40 | +0.80 | +0.60 | +0.80 | −0.40 | −0.60 | ||
| cat | +0.50 | +0.10 | +0.90 | +0.90 | +0.50 | +0.10 | +0.50 | +0.90 | −0.10 | ||
| sleeps | −0.30 | +0.70 | +0.20 | +0.20 | −0.30 | +0.70 | −0.30 | +0.20 | −0.70 | ||
Each token now has three distinct representations, none of which is simply its original embedding. The query, key, and value for "black" are all different from each other — and different from "black"'s original embedding [+0.80, −0.40, +0.60]. Each projection has rotated and rescaled the information in a different direction, giving the model three separate handles to tune independently during training.
The most striking effect of projections is visible in the attention weights. Compare the two patterns directly:
| Part 2 — no projections (e·e) | ||||
|---|---|---|---|---|
| the | black | cat | sleeps | |
| the | 0.27 | 0.20 | 0.32 | 0.21 |
| black | 0.14 | 0.51 | 0.18 | 0.18 |
| cat | 0.21 | 0.17 | 0.45 | 0.17 |
| sleeps | 0.18 | 0.22 | 0.22 | 0.37 |
| Part 3 — with projections (q·k / √d) | ||||
|---|---|---|---|---|
| the | black | cat | sleeps | |
| the | 0.23 | 0.24 | 0.24 | 0.28 |
| black | 0.26 | 0.20 | 0.37 | 0.18 |
| cat | 0.19 | 0.25 | 0.27 | 0.29 |
| sleeps | 0.21 | 0.35 | 0.23 | 0.20 |
The change for "black" is the most telling. In Part 2, "black" attended to itself with weight 0.51 — more than half its attention was self-directed, because its embedding was the most similar to itself. In Part 3, "black" now attends most to "cat" (0.37), with its self-attention reduced to 0.20. The projections have broken the self-similarity dominance, allowing the attention pattern to express relationships that the raw embedding geometry could not.
Similarly, "sleeps" now attends most to "black" (0.35) rather than to itself (0.20). Whether these particular patterns are semantically meaningful is beside the point — the weights here are random and untrained. The key insight is structural: the model now has the capacity to learn patterns like these through training, which it did not have before.
Finally, the output \(z_i\) is now a weighted sum of value vectors, not the original embeddings. The z vectors for Part 3 look nothing like the original embeddings — the value projection has transformed the information space entirely:
| d0 | d1 | d2 | |
|---|---|---|---|
| z_the | +0.25 | +0.25 | −0.32 |
| z_black | +0.32 | +0.36 | −0.23 |
| z_cat | +0.27 | +0.26 | −0.34 |
| z_sleeps | +0.36 | +0.17 | −0.33 |