Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers
Stanford Online
Course Introduction and Instructors 0:06
Afin opens the class alongside his twin brother Shervin, explaining that both studied at CentraleSupélec in France before attending MIT and Stanford respectively, then working together at Uber, Google, and now Netflix. They have taught versions of this material since 2021, and this is the third year it runs as an official Stanford class. The course aims to explain how large language models work, covering the transformer architecture, how these models are trained, how they are used, and where their strengths and weaknesses lie.
Audience, Prerequisites, and Logistics 7:31
The class suits future research scientists, people building personal projects with coding agents, and really anyone wanting basic AI literacy. Helpful background includes linear algebra and basic machine learning concepts like loss functions, neural networks, and embeddings, though these will also be covered in class. Sessions run Fridays from 3:30 to 5:20, are recorded and posted within two days, and the two-unit course has no homework but includes a midterm on October 23rd covering the first four lectures and a final covering the last five, each worth 50 percent. Materials include a textbook, a regularly updated cheat sheet available in about 15 languages, slides posted a day before each lecture, and past exams from 2025 for practice, with communication happening through Canvas, ED, and email.
From RNNs to the Transformer 12:31
Ten to fifteen years ago, in the 2010s, the field relied on recurrent neural networks, with a separate dedicated model built for each task such as sentiment analysis, translation, or name recognition. The turning point came in 2017 with the paper Attention Is All You Need, which introduced the transformer architecture, a design that scales well, meaning that adding more data, compute, and parameters keeps improving performance. This scaling led to modern large language models, and the release of ChatGPT in November 2022 marked another turning point by introducing everyday chatbot interaction to the public. The course will trace how this evolved into today's agentic coding tools.
What Tokenization Means 17:01
Since models understand numbers rather than raw text, the first step is tokenization, splitting text into indivisible units called tokens, each represented as a vector of numbers. The set of all allowed tokens is called the vocabulary, and the chosen vocabulary affects the resulting sequence length, the number of tokens a piece of text breaks into. One simple approach splits text by words using spaces as delimiters, which is easy to interpret in English but runs into trouble with word variations like singular, plural, or gendered forms, which can inflate the vocabulary size.
Downsides of Word Level Tokenization 21:03
Splitting text into whole words creates a large vocabulary, since every word needs its own learned representation. It also fails to capture shared roots, so bear and bears end up as unrelated tokens even though they mean the same thing. Finally, any word not seen during training becomes an out of vocabulary problem at inference time, meaning the model has no representation for it at all.
Character and Subword Alternatives 23:30
Tokenizing by character avoids the out of vocabulary problem entirely, since you only need the fixed set of letters in a language, and it also handles misspellings gracefully. The tradeoff is that sequences become much longer, which slows the model down, and single-letter representations are far less interpretable than word representations. The middle ground is subword tokenization, which breaks words into common chunks learned from a training corpus, so reading becomes read plus ing, letting bear and bears share a token while keeping sequences short. The main cost is that building this vocabulary takes extra work and depends heavily on matching your training data to your intended use, especially for multilingual applications. Byte pair encoding, or BPE, is the dominant method today: starting from individual letters, it repeatedly merges the most frequent adjacent pairs into new tokens until the vocabulary reaches a target size, and modern large language models typically use vocabularies in the hundreds of thousands of tokens, often covering many languages at once.
Special Tokens for Structure 31:30
Beyond tokens for ordinary text, models rely on special tokens to mark structure. An unknown token stands in for any text the vocabulary can't represent, such as an unfamiliar spelling. Beginning of sequence and end of sequence tokens, often written BOS and EOS, tell the model when to start and stop generating text. A padding token forces sequences to a uniform length, which helps hardware process batches efficiently, and newer conventions add tokens like user and assistant markers, though none of these naming conventions are fully standardized across models.
From Tokens to Representations 37:33
Once text is split into tokens, the next question is how to represent them numerically. The simplest approach, one hot encoding, gives each token a unique vector, but this makes every token equally distant from every other, telling you nothing about meaning. What you actually want is representations where related tokens, like teddy bear and soft, end up close together. An early method called word2vec learns such representations through a proxy task, a stand-in objective like predicting a missing word from its surrounding context. Using a simple neural network, a one hot input token is passed through a matrix multiplication into a smaller hidden layer, then projected back out to vocabulary size and normalized into probabilities, with a loss function like cross entropy pushing the predicted probability toward the actual next word.
Walking Through the Word2Vec Training Loop 43:00
You start with a one-hot encoding of a token, which is a vector as long as your vocabulary size, V. This gets compressed into a hidden state, or latent representation, of size D, a smaller number you choose yourself. The network then expands this back out into predicted word probabilities, again of size V, one probability per possible word. You compare this prediction against the actual next word, such as predicting cute after teddy bear, compute the loss, and update the weights. Repeating this across a whole corpus, the authors found something remarkable: the learned latent representations captured real meaning, so that similarity comparisons reveal relationships like teddy bear is to soft as Persian is to poetry, or Paris is to France as Berlin is to Germany. This method dates back to around 2013.
Limitations of Simple Word Embeddings 45:30
This approach struggles with words never seen during training. It also scales awkwardly, since the representation size grows with vocabulary size, though that is really a limitation of the tokenizer rather than the method itself. A bigger problem is that words can carry multiple meanings, so cute in a sincere sentence and cute said sarcastically get identical representations, much like river bank versus a bank you visit. Word order is ignored too: the child is hugging the teddy bear and the teddy bear is hugging the child produce the same embeddings despite meaning different things.
RNNs and Their Long Range Problem 48:30
RNNs addressed word order by keeping a hidden state that accumulates the meaning of the sequence as it processes one word at a time, feeding forward into the next prediction. This let them handle tasks like sentiment classification or labeling words as nouns or verbs. But RNNs struggled to remember information from far back in a sequence, the long range dependency problem, since everything gets compressed into a single evolving vector. Saying my teddy bear is so cute, it is 3 feet tall requires connecting it back to teddy bear across a gap. Variants like LSTM, with an added cell state, tried to help but still fell short. RNNs were also slow to train, since predictions must be made one token at a time in sequence.
Attention as a Direct Connection 55:01
Attention, developed in the 2010s, lets a model connect directly to earlier tokens instead of relying only on a compressed hidden state. In translation, knowing the source word directly helps far more than inferring it through an intermediate state. The 2017 paper Attention Is All You Need turned this into self-attention, where a token's representation is computed as a function of every other token in the sequence, though this raw form initially loses track of word order, a problem addressed later.
Queries Keys and Values Explained 59:30
The mechanism uses three learned quantities: query, key, and value, each obtained by projecting a token's representation through separate projection matrices. The query for a token like teddy bear asks what best describes it, then measures similarity, often via dot product, against the keys of all other tokens. High similarity, say between teddy bear's query and cute's key, means cute's value contributes heavily to teddy bear's new representation, a weighted sum across all tokens. Mathematically this whole operation is expressed as a softmax of the scaled query-key product, times the values, with softmax ensuring the weights sum to one.
Matrix form of self-attention 1:06:00
You can arrange all the query, key, and value vectors for a sequence into matrices, and multiplying the query matrix by the transposed key matrix gives you every dot product between queries and keys at once. Adding the value matrix into this multiplication produces, for each position, a weighted sum of the values based on those similarity scores. This matrix formulation is the compact way self-attention gets written and computed in practice, and it is flagged as a formula worth remembering.
Why attention replaces recurrence 1:10:01
The transformer architecture comes from the 2017 paper Attention Is All You Need, built for machine translation, such as English to French. Earlier models were recurrent, processing tokens one by one while carrying a hidden state forward. The transformer instead lets every token interact directly with every other token through self-attention, removing the need for recurrence entirely. The architecture has two halves: an encoder that builds meaningful representations of the source text, and a decoder that uses those representations plus its own internal embeddings to generate the translation.
Tokens, positions, and the encoder 1:12:31
Each input token is first converted into a learned embedding, and since that embedding alone carries no sense of word order, a position embedding is added on top. Positions can be learned outright, which fails once a sequence runs longer than anything seen in training, or built from fixed sine and cosine functions at varying frequencies, similar to how watch hands moving at different speeds together indicate an exact time. Inside the encoder, a self-attention layer lets each token's representation be shaped by every other token via the query, key, and value projections, and a feed-forward layer then projects these representations into a higher-dimensional space with a nonlinearity.
How the decoder generates translations 1:18:31
Generation starts with a beginning-of-sequence token fed into the decoder, which first applies self-attention over only the tokens generated so far, then applies a second attention layer where those tokens' queries interact with keys and values coming from the encoder's output. A feed-forward layer follows, and the final representation is projected onto the vocabulary to get probabilities for the next token, which is chosen and fed back in autoregressively to repeat the process.
Supporting tricks: residuals, normalization, masking 1:22:00
Residual connections add a layer's input back to its output, so a layer refines the input rather than replacing it, which helps deep networks train through backpropagation. Layer normalization stabilizes activations to aid convergence, and masking in the decoder's self-attention ensures a token can only interact with tokens already decoded, since future tokens do not yet exist.
Multi-head attention, dropout, and label smoothing 1:24:02
Multi-head attention runs several different query-key-value projections in parallel, much like using multiple filters in a convolutional network, so the model learns several ways of relating tokens at once. Dropout randomly disables some units during training so the model does not over-rely on specific features, improving generalization. Label smoothing softens the target for next-word prediction, for example treating the word day as only about ninety percent certain rather than absolute, spreading the rest across other plausible words, which the original authors found improved translation quality as measured by metrics like BLEU.
Setting up the worked example 1:27:32
Shervin then takes over to walk through one full worked example combining everything just described. The original transformer was built for machine translation, using French and German as target languages, and the example sentence chosen is a cute teddy bear is reading. The first step is tokenization, splitting the sentence into arbitrary pieces rather than the optimal subword units mentioned earlier, and adding beginning-of-sequence and end-of-sequence special tokens to mark where the sequence starts and ends.
Embedding tokens and positions 1:29:00
Each token is turned into a vector using a lookup table of learnable weights called embeddings, and since attention needs to know where tokens sit relative to each other, a position embedding built from sine and cosine waves of different frequencies is added to it. The result is a matrix representation for the whole sequence that is at least partly learnable, and this matrix is what feeds into the encoder.
Queries keys and values dimensions 1:31:34
Inside the encoder, the input embedding of size D model is projected through learned matrices into queries, keys, and values, each living in its own space of size DQ, DK, and DV, with DQ forced to equal DK so the dot product between queries and keys works. Multiplying queries by the transpose of keys gives an n by n matrix of similarity scores, which is scaled by the square root of DK to keep the variance in check, passed through softmax to get a probability distribution, and then used to weight the values. This whole process runs eight times in parallel, called multi-headed attention, and the results are concatenated and projected back to D model size, keeping dimensions consistent as the data moves into a feed-forward network with a larger hidden layer, all repeated six times as in the original paper.
Decoder and cross attention 1:37:30
The decoder starts from a placeholder beginning-of-sequence token, adds its embedding and position information, and passes it through a masked self-attention layer that only lets each token see itself and earlier tokens, since translation must proceed causally. A second attention layer, cross attention, then lets the decoder's queries attend to keys and values coming from the encoder's output. After a feed-forward layer, a final linear and softmax layer turns this into a probability distribution over the vocabulary, the predicted word is fed back in as the next input, and this repeats until an end-of-sequence token appears.
What really makes it work 1:40:01
Asked what the secret sauce is, the presenter points to the attention layer as the key mechanism, and to the feed-forward network as where most of the model's parameters and learning capacity live, with everything else being smaller tricks plus the quality of the training data. Masking is described as a computational trick using a triangular matrix and softmax of negative infinity to block forbidden connections, and the different softmax uses are clarified: attention softmax projects a token over other tokens, while the final softmax predicts the next word. The lecture closes by noting this 2017 architecture is just the starting point, with training methods and agent-based applications to come in later lectures.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

