Understanding Transformers
from First Principles

A ground-up construction, one piece at a time

Our running example: the black cat sleeps → el gato negro duerme
Work in progress
This document is being written in order, and it currently runs through MVP 2: Learned Projections (Q, K, V). Everything up to that point is complete and self-contained — you can read it end to end and come away understanding how self-attention works and why each piece of it is there. The remaining stages — MVP 3 through MVP 7, and the closing section on training and inference — are listed below so you can see where the construction is headed, but they have not been written yet.

Contents

  1. Orientation — what problem are we solving?
  2. The Baseline: Bag of Embeddings — the simplest possible start
  3. MVP 1: Single-Head Self-Attention
  4. MVP 2: Learned Projections (Q, K, V)
  5. MVP 3: Positional Encodingin progress
  6. MVP 4: Feed-Forward Networkin progress
  7. MVP 5: Residual Connections & Layer Normin progress
  8. MVP 6: Multi-Head Attentionin progress
  9. MVP 7: Stacked Layers — the Full Architecturein progress
  10. Training and Inference in Fullin progress
Part 0

Orientation

The task this document builds around is translation: given the English sentence "the black cat sleeps," produce its Spanish equivalent, "el gato negro duerme." We use translation as our running lens throughout because it exercises every part of the architecture naturally — encoder, decoder, and the connection between them. The ideas generalize well beyond it.

Notice something immediately: the word order is not the same. In English, the adjective comes before the noun — black cat. In Spanish, it comes after — gato negro. A system that simply translates word by word, left to right, will get this wrong. A good translation model needs to understand the structure of the whole source sentence before it can produce the target. That structural understanding is exactly what the transformer is built to provide.

Transformers in context

The transformer was introduced in 2017 in a paper titled "Attention Is All You Need" (Vaswani et al.) and rapidly became the dominant architecture for language tasks. Before it, the standard approach was recurrent neural networks — architectures that processed text one word at a time, left to right, passing a hidden state forward at each step. The transformer replaced that sequential processing with a fully parallel mechanism called attention, which lets every token in a sequence directly interact with every other token in a single step. This made training dramatically faster and gave models a much easier time capturing relationships between distant words.

Today, virtually every large language model you interact with — GPT-4, Claude, Gemini, LLaMA, and others — is a transformer or a close descendant of one. The architecture has also spread well beyond text: transformers are now used in image recognition (Vision Transformers), protein structure prediction (AlphaFold), audio generation, and more. Learning how a transformer works is, in large part, learning the foundation of modern AI.

How this document works

Rather than presenting the full transformer architecture and then explaining what each piece does, we build it up from scratch — one component at a time. Each version of the model will have multiple shortcomings. Rather than fixing them all at once, we identify the most important one, add the minimal change needed to address it, and move on. The next version then fixes the next most significant remaining problem, and so on. Every addition is motivated by one concrete, demonstrable failure in the version before it.

This progression has eight stages, called MVPs (minimum viable products). Every section follows the same structure:

On the math. Every section contains mathematical notation in clearly marked callout boxes. The math makes the ideas precise, but you do not need it to follow the argument. If you are comfortable with it, the boxes add rigor; if not, skip them — the surrounding text carries the full explanation on its own.
Optional: the math
Math callout boxes look like this throughout the document. They are always self-contained and always skippable.

Training and inference — an early distinction

Before we build anything, there is one distinction worth establishing now, because it runs through every section: the difference between training and inference. Training is the process of adjusting a model's parameters — its internal numerical weights — so that it gets progressively better at its task, given a large set of labeled examples. In our case those examples are paired English and Spanish sentences; the model sees both sides, measures how wrong its predictions are, and updates its weights accordingly. Inference is the process of passing a new input through a trained, fixed model to produce the desired output. The model's parameters do not change during inference — only the input changes.

Training

We have a large dataset of paired sentences — English alongside correct Spanish translations. The model sees both sides, makes a prediction, measures how wrong it is against the known answer, and adjusts its internal numbers to do better next time. This happens millions of times.

the black cat sleeps → el gato negro duerme

Both sides known. Error measured. Weights updated.

Inference

We have only the source sentence. The model must produce the translation from scratch, with no reference answer to compare against. There is no error signal, no weight update — just a forward pass through the learned model.

the black cat sleeps → ???

Source only known. Target must be generated.

The mechanical consequences of this distinction — especially in how the decoder behaves differently in each mode — are one of the most important and least-explained aspects of the transformer. We will return to it with full technical detail in Part 4, when the decoder is introduced.

Part 1 — The Baseline

Bag of Embeddings

Before we build the transformer, we need to answer a more basic question: how do you feed a sentence into a neural network at all? Networks operate on numbers. Strings of text are not numbers. The answer is embeddings — and they are the foundation everything else is built on.

Step 1: Turning words into vectors

An embedding is simply a list of numbers — a vector — assigned to each word in the vocabulary. We maintain a large table, one row per vocabulary word, where each row is that word's vector. Looking up a word's embedding means finding its row in that table.

These vectors are learned: they start as random numbers and are updated during training until words with similar meanings end up with similar vectors. "Cat" and "kitten" end up nearby; "cat" and "airplane" end up far apart. But at the very start of training, every word has a random vector.

In this document we use 3-dimensional embeddings to keep the numbers readable. Real models use 512 or more dimensions. Here is what the embeddings for our four source tokens might look like early in training:

Tokendimension 0dimension 1dimension 2
the0.100.30−0.20
black0.80−0.400.60
cat0.500.900.10
sleeps−0.300.200.70
Optional: the math
The embedding table is a matrix \( E \in \mathbb{R}^{V \times d} \), where \(V\) is the vocabulary size and \(d\) is the embedding dimension. Looking up word \(x_i\) is written \( e_i = E[x_i] \) — "take row \(x_i\) of \(E\)." The entire table is a learnable parameter, updated by gradient descent during training.

Step 2: Collapsing the sequence into a single vector

We now have four vectors — one per token. But to produce an output we need a single fixed-size vector. Predicting the next word means scoring every word in the Spanish vocabulary and picking the most likely one — a form of classification over V classes, one per vocabulary word. A variable-length sequence of token vectors cannot be fed directly into that scoring layer; it needs one fixed-size vector. The simplest possible answer: average them. Add up all four embedding vectors and divide by four.

Tokendim 0dim 1dim 2
the0.100.30−0.20
black0.80−0.400.60
cat0.500.900.10
sleeps−0.300.200.70
average → 0.280.250.30

This averaged vector — the context vector — is then passed through two operations to produce a prediction. First, a linear projection: a learned matrix \(W\) maps the context vector from the embedding space (\(\mathbb{R}^3\) in our example, or \(\mathbb{R}^d\) in general) into the vocabulary space (\(\mathbb{R}^V\), where \(V\) is the number of words in the Spanish vocabulary — potentially tens of thousands). "Linear" here means the transformation is a matrix multiplication with no nonlinear activation; "projection" means we are mapping between two different vector spaces. The result is a \(V\)-dimensional vector of raw scores, one per Spanish word: a higher score means the model thinks that word is more likely given this context. Second, a softmax converts those raw scores into a proper probability distribution summing to 1. The word with the highest probability is the model's prediction for the first Spanish token.

Optional: the math

The context vector is the elementwise average of the \(n\) token embeddings:

\[ c = \frac{1}{n} \sum_{i=1}^{n} e_i \]

The prediction is produced by a linear projection followed by softmax:

\[ p(y) = \text{softmax}(W c + b) \]

where \( W \in \mathbb{R}^{V \times d} \) is a weight matrix and \( b \in \mathbb{R}^{V} \) is a bias vector — both are learnable parameters updated during training. Row \(i\) of \(W\) is a learned \(d\)-dimensional vector associated with vocabulary word \(i\); the dot product of row \(i\) with \(c\) scores how well the context matches that word. Softmax then converts the full vector of \(V\) raw scores into probabilities:

\[ \text{softmax}(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{V} e^{z_j}} \]

Training and inference at this stage

During training

We have the paired example. We compute context vector [0.28, 0.25, 0.30], pass it through the linear layer, and get a probability over all Spanish vocabulary words. We compare against the known first word — el — measure the error, and backpropagate. The embedding table and linear weights are both updated. Repeat for millions of sentence pairs.

During inference

We have only "the black cat sleeps." We compute the same context vector [0.28, 0.25, 0.30] and pass it through the same linear layer — but now there is no known answer. We take the highest-probability Spanish word as our prediction. There is no weight update; everything is frozen.

Three things that are wrong with this

Averaging embeddings is a functional starting point — the model can technically learn something — but it has three fundamental problems. Each one will be fixed by a specific, identifiable change in the sections that follow.

Root cause
All three failures trace back to the same underlying problem: there is no mechanism for words to interact with each other. Each embedding is looked up independently and blended with equal weight. The embedding for "black" has no way to influence the representation of "cat." A model that cannot let words influence each other cannot understand the structure of the sentence — it sees only a bag of words, not a sequence.
Coming up in Part 2
The fix is self-attention: instead of averaging with equal weight, each token computes a score against every other token, uses those scores to weight the blend, and produces a new representation that has been contextualized by what is around it.
Part 2 — MVP 1

Single-Head Self-Attention

Added this section
✓ Self-attention
Still to come
1/√d scaling Q, K, V projections Positional encoding Feed-forward network Residual connections Multi-head attention Stacked layers

In Part 1 we identified the root problem: averaging treats every word as equally relevant. Self-attention replaces that flat average with a learned, content-dependent weighting. Each token looks at every other token, scores how relevant it is, and builds a new representation by blending the ones that matter most to it.

The mechanism

Self-attention happens in three steps, applied simultaneously to every token.

Step 1: Score every other token. For token i, compute a dot product against every token j in the sequence — including itself. The dot product measures similarity: a high score means the two tokens are pointing in a similar direction in embedding space, and therefore that j is likely relevant to updating i's representation.

Step 2: Normalize into a probability distribution. Pass all scores for token i through softmax. This converts the raw scores into attention weights that sum to 1. A high weight on token j means token i will draw heavily from token j's embedding in the next step.

Step 3: Take the weighted sum. Compute a weighted sum of every token's embedding, weighted by the attention weights from step 2. The result, zi, is token i's new contextualized representation — it has been updated to reflect the content of the tokens around it, weighted by relevance.

Optional: the math

For each token \(i\), compute raw scores against all other positions \(j\):

\[ \text{score}_{ij} = e_i \cdot e_j \]

Normalize into attention weights via softmax (over all \(j\) for fixed \(i\)):

\[ \alpha_{ij} = \text{softmax}_j(\text{score}_{ij}) = \frac{e^{e_i \cdot e_j}}{\sum_{k} e^{e_i \cdot e_k}} \]

Compute the contextualized representation as a weighted sum of all embeddings:

\[ z_i = \sum_j \alpha_{ij}\, e_j \]

Note that \(j\) ranges over every position including \(i\) itself — a token is always a candidate to attend to itself. Also note that \(\alpha_{ij} \neq \alpha_{ji}\) in general: how much token i attends to token j is not the same as how much j attends to i.

Traced through "the black cat sleeps"

We use the same four embeddings from Part 1. The attention computation produces a 4×4 score matrix — one row per query token, one column per key token — then softmax weights, then the final contextualized vectors.

Step 1: Score matrix (each cell is the dot product of the row token's embedding with the column token's embedding):

theblackcatsleeps
the+0.14−0.16+0.30−0.11
black−0.16+1.16+0.10+0.10
cat+0.30+0.10+1.07+0.10
sleeps−0.11+0.10+0.10+0.62

The diagonal (highlighted) shows each token's self-similarity — how strongly it matches itself. "black" and "cat" have the highest self-scores because their embeddings have the largest magnitudes. "the", a low-information function word, has the smallest self-score.

Step 2: Attention weights (softmax applied row by row; each row sums to 1; bolder cell = highest weight in that row):

theblackcatsleeps
the0.270.200.320.21
black0.140.510.180.18
cat0.210.170.450.17
sleeps0.180.220.220.37

A few things to notice. "black" attends most to itself (0.51) — its high self-score is so dominant that it draws more than half the weight. "sleeps" also attends most to itself (0.37), but much less exclusively — the remaining weight is spread across the other three tokens. "the" is the most distributed: it attends most to "cat" (0.32), slightly more than to itself (0.27). This is the content-dependent weighting that was impossible in Part 1 — no two rows look the same.

Step 3: Contextualized representations (weighted sum of embeddings using the weights above; compare original vs. updated):

original embeddingafter attention z
d0d1d2d0d1d2
the+0.10+0.30−0.20→+0.28+0.33+0.25
black+0.80−0.40+0.60→+0.46+0.03+0.42
cat+0.50+0.90+0.10→+0.33+0.43+0.23
sleeps−0.30+0.20+0.70→+0.20+0.24+0.38

Every token's representation has moved. "the" moved the most dramatically: it started at [+0.10, +0.30, −0.20] and landed at [+0.28, +0.33, +0.25] — the −0.20 in dimension 2 has flipped positive, pulled by the strong positive values of "cat" and "sleeps" in that dimension. "black", which attends mostly to itself, moved the least: its representation stayed close to its original, diluted only slightly by the other tokens.

This is the core of what contextualization means: each token's final representation is no longer just its own embedding — it is a blend of the whole sequence, weighted by how relevant each token decided the others were to it.

What this version still cannot do

Self-attention fixes failure 1 from Part 1 (uniform weighting) and failure 3 (dilution). But failure 2 — permutation invariance — is still present. Consider a longer example:

The permutation problem
"Time flies like an arrow. Fruit flies like a banana."

There are two instances of "flies" in this sentence — one a verb ("time flies"), one a noun ("fruit flies"). A reader has no trouble distinguishing them because of the words around each instance. But in this MVP, both instances of "flies" have the same embedding, and without any position information, they also attend over exactly the same set of keys with exactly the same weights. The result: \(z_{\text{flies}_1} = z_{\text{flies}_2}\) — identical contextualized representations, despite the words meaning different things in each context.

The reason is structural, not incidental: without positional encoding, "flies" in position 1 uses the same query embedding as "flies" in position 6, and both compute dot products against the same global set of key embeddings, weighted the same way. Position simply does not exist as information anywhere in the computation. This will not be fixed until Part 5, when we add sinusoidal positional encodings that inject position information directly into the embeddings before attention runs.

Training and inference at this stage

The encoder has a clean property worth naming now because it contrasts sharply with the decoder: it behaves identically at training time and inference time. The source sentence is always fully known in both cases — whether you are training on a labeled translation pair or translating a new sentence from scratch. So the encoder always runs exactly once, in a single forward pass over all source tokens in parallel, with no sequential dependency.

Training — encoder

Source sentence is known. All four tokens are embedded and processed in parallel. The 4×4 attention matrix is computed in one matrix multiplication; all four z vectors are produced simultaneously. No sequential steps — this is what makes the transformer trainable at scale.

Inference — encoder

Exactly the same computation. Source sentence is still fully known. The encoder produces z1…z4 in one forward pass, identically to training. The interesting train/inference asymmetry lives entirely in the decoder — which we will introduce in Part 4.

Coming up in Part 3
Self-attention as described here forces every token to use the same embedding vector for three conceptually distinct jobs: forming a query ("what am I looking for?"), offering a key ("what do I advertise to others?"), and providing a value ("what do I actually contribute if attended to?"). Collapsing three roles into one vector means the model has no freedom to optimize them independently — it cannot learn to ask one kind of question while offering a different kind of information. Part 3 fixes this by introducing three separate learned projections, one per role.
Part 3 — MVP 2

Learned Projections & Scaling

Added this section
✓ Q, K, V projections ✓ 1/√d scaling
Still to come
Positional encoding Feed-forward network Residual connections Multi-head attention Stacked layers

In Part 2, every token played three roles simultaneously using the same single vector: it queried ("what am I looking for?"), it offered itself as a key ("what do I advertise to others' queries?"), and it provided a value ("what do I actually contribute if someone attends to me?"). Using one vector for all three means the model cannot optimize them independently. A word cannot learn to ask a different kind of question than it answers.

The fix is one of the cleanest ideas in the transformer: give each role its own learned linear projection. Instead of one embedding per token, we maintain three separate weight matrices — \(W_Q\), \(W_K\), and \(W_V\) — each of which transforms the embedding into a specialized representation for its role.

Three roles, three matrices

The three projections work as follows for each token \(i\):

The attention score between token \(i\) and token \(j\) is now the dot product of \(i\)'s query with \(j\)'s key, divided by \(\sqrt{d}\):

Optional: the math

The three projections for token \(i\):

\[ q_i = W_Q e_i, \qquad k_j = W_K e_j, \qquad v_j = W_V e_j \]

Scaled dot-product attention:

\[ \alpha_{ij} = \text{softmax}_j\!\left(\frac{q_i \cdot k_j}{\sqrt{d}}\right), \qquad z_i = \sum_j \alpha_{ij}\, v_j \]

Why divide by \(\sqrt{d}\)? If the components of \(q_i\) and \(k_j\) each have variance \(\sigma^2\), the dot product \(q_i \cdot k_j = \sum_{m=1}^{d} q_{im} k_{jm}\) has variance \(d\sigma^4\) — it grows with \(d\). As \(d\) increases, the scores become so large that softmax saturates: one score dominates, the weights collapse toward a one-hot vector, and gradients vanish. Dividing by \(\sqrt{d}\) normalises the variance back to \(\sigma^4\), independent of dimension size.

Note also that the output \(z_i\) is now a weighted sum of value vectors \(v_j\), not of the original embeddings \(e_j\). The value projection separates "what information I use to decide who to attend to" (keys and queries) from "what information I actually receive from attending" (values).

Traced through "the black cat sleeps"

We apply three simple projection matrices to our four token embeddings. Each produces a different view of the same underlying embedding — the same sentence, seen through three different lenses:

query q = WQe key k = WKe value v = WVe
d0d1d2 d0d1d2 d0d1d2
the +0.10−0.20+0.30 +0.30+0.10−0.20 +0.10+0.30+0.20
black +0.80+0.60−0.40 −0.40+0.80+0.60 +0.80−0.40−0.60
cat +0.50+0.10+0.90 +0.90+0.50+0.10 +0.50+0.90−0.10
sleeps −0.30+0.70+0.20 +0.20−0.30+0.70 −0.30+0.20−0.70

Each token now has three distinct representations, none of which is simply its original embedding. The query, key, and value for "black" are all different from each other — and different from "black"'s original embedding [+0.80, −0.40, +0.60]. Each projection has rotated and rescaled the information in a different direction, giving the model three separate handles to tune independently during training.

How the attention pattern changes

The most striking effect of projections is visible in the attention weights. Compare the two patterns directly:

Part 2 — no projections (e·e)
theblackcatsleeps
the0.270.200.320.21
black0.140.510.180.18
cat0.210.170.450.17
sleeps0.180.220.220.37
Part 3 — with projections (q·k / √d)
theblackcatsleeps
the0.230.240.240.28
black0.260.200.370.18
cat0.190.250.270.29
sleeps0.210.350.230.20

The change for "black" is the most telling. In Part 2, "black" attended to itself with weight 0.51 — more than half its attention was self-directed, because its embedding was the most similar to itself. In Part 3, "black" now attends most to "cat" (0.37), with its self-attention reduced to 0.20. The projections have broken the self-similarity dominance, allowing the attention pattern to express relationships that the raw embedding geometry could not.

Similarly, "sleeps" now attends most to "black" (0.35) rather than to itself (0.20). Whether these particular patterns are semantically meaningful is beside the point — the weights here are random and untrained. The key insight is structural: the model now has the capacity to learn patterns like these through training, which it did not have before.

Finally, the output \(z_i\) is now a weighted sum of value vectors, not the original embeddings. The z vectors for Part 3 look nothing like the original embeddings — the value projection has transformed the information space entirely:

d0d1d2
z_the+0.25+0.25−0.32
z_black+0.32+0.36−0.23
z_cat+0.27+0.26−0.34
z_sleeps+0.36+0.17−0.33
Coming up in Part 4
We have now fixed both the uniform-weighting problem (Part 2) and given the model the expressive power to learn arbitrary attention patterns (Part 3). But one of the original three failures remains untouched: the model is still completely blind to word order. "The black cat sleeps" and "sleeps cat black the" still produce the same output. Part 4 adds positional encoding — the only mechanism in the entire architecture that injects position information into the computation.