Stanford Online

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models: summary

YouTube summary16 sectionsWatch on YouTube ↗

This is an AI-generated summary of the YouTube video "Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models" (Stanford Online), made with Samuraize and published by Samuraize. It condenses the YouTube video into 16 titled sections you can read in a couple of minutes, each linking to the moment in the video it covers.

1
Filed under💻 Technology0 comments🍱 Add to trayReport
Study this
Export

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models

Stanford Online

切

Transformer recap 0:05

The lecture starts by revisiting tokenization, embeddings, RNNs, and the 2017 attention paper. It recalls that transformers use query, key, and value to decide what each token should attend to. It also restates that the original encoder-decoder transformer was built for machine translation.

切

BERT and GPT 5:01

The next step is to split the transformer into two main variants. BERT keeps only the encoder and uses self-attention across all tokens, with a CLS token whose embedding can stand for the whole input. It is trained first with masked language modeling and next sentence prediction, then fine-tuned for tasks like sentiment classification. GPT keeps only the decoder and predicts the next token one step at a time. It treats many tasks as text to text and is simpler because it does not need special input handling.

切

Scaling decoder models 19:31

The final topic here is the move toward larger decoder-only models. Early 2020s work showed that performance improves with more data, more parameters, and more compute, and that bigger models become more token efficient. The slide described test loss against training tokens and showed that larger models reach better results sooner, setting up the question of what to do when compute is limited.

切

Compute and training size 22:01

Flops means floating point operations, and here it is a measure of compute. A capital FLOPS means compute speed, so it is easy to mix the two up. The question is how to split limited compute between model size and the number of tokens you train on. A 2022 paper found a sweet spot and showed that many models then were too large for their training data, so a smaller model trained on more tokens did better under the same compute budget.

切

Decoder models and experts 26:00

For next-token prediction, cross entropy is the usual loss. In pre-training, data size matters more than data quality, though quality matters more later in training. LLM means large language model because the models are large, often with billions or more parameters and huge token counts. Modern systems usually use a decoder-only transformer, and sometimes a mixture of experts in the feed-forward network. A router picks which expert to activate for each token, and a sparse setup uses only the top k experts instead of all of them.

切

Routing and balance 35:01

The router is learned by backpropagation, but it can collapse onto a few experts if they get picked early. To stop that, training adds a load-balancing loss that pushes token routing toward a more even spread across experts. Routing happens at the token level, and stop-grad is used so the auxiliary loss treats the routed fractions as constants while still updating the routing probabilities. Newer models often add shared experts and use hundreds of experts, with only a small set activated each pass. Recent frontier models are huge, with hundreds of billions of parameters or more.

切

Position Embeddings Revisited 46:02

Learned position embeddings add a vector for each token index to the token embedding. That lets the model know where a token sits in a sequence. But it breaks when you see longer inputs at test time, because positions beyond training have no learned vectors. The fixed formula avoids that problem and gives similar results.

切

RoPE and Rotation 57:02

The common modern choice is RoPE, short for rotary position embedding. It does not add a bias. It rotates the query and key vectors by an angle tied to position. The point is that the attention score then depends on the difference between two positions. The slide works through the 2D rotation math and shows that multiplying by a rotation matrix really does rotate the vector.

切

RoPE and local attention 1:07:00

Rotary position embeddings, or RoPE, turn position into a rotation that depends on the index. The same idea extends to higher dimensions in 2D blocks. It also makes the dot product between query and key tend to shrink as tokens move farther apart. After that comes sliding window attention, also called local attention. A token only looks a limited distance back, but stacked decoder blocks still let information travel farther, much like receptive fields in convolutional networks.

切

Memory and normalization 1:14:30

Decoder models keep the keys and values from past tokens so they do not have to recompute them at every step. That saves work, but it costs memory, so methods like grouped query attention, or GQA, share key and value projections across heads or groups. The architecture has also shifted from postnorm, where normalization comes after the sublayer, to prenorm, where it comes before it. Layer norm has often been replaced by RMS norm, which rescales without recentring and uses one learned factor.

切

Choosing next tokens 1:21:00

Next token prediction uses the decoder stack and a final linear layer over the vocabulary. A simple choice is greedy decoding, where you always pick the most probable token, but that can lead to a fixed and robotic result. Beam search keeps the best k paths, but it is still deterministic. Sampling is the other main route. You can sample from the full distribution, or restrict yourself to top k tokens or to top p, where the cumulative probability stays under a set cutoff. Determinism is useful for tasks like summarization and for consistent LLM-as-judge ratings.

切

Softmax temperature 1:28:30

The final projection produces vocabulary scores, and softmax turns them into a probability distribution. Temperature T changes how sharp that distribution is. With T equal to 1, you get the regular softmax. The next step is to think about what happens when you raise or lower T.

切

Temperature And Sampling 1:29:31

Small temperature makes one token stand out and pushes the others toward zero. Large temperature flattens the distribution so the next token choice becomes more even. The same decoder can feel deterministic, but batching and floating point addition can break exact repeatability. The order of operations in tensors can change results, so the implementation matters.

切

Structured Output And Prompts 1:33:30

When you need JSON or another fixed format, asking for it directly often works but is not fully safe. Guided decoding narrows the next token choices to match the syntax you want. Prompting then uses a prompt as input, with the context window covering both your input and the tokens the model has already generated.

切

Context Limits And Retrieval 1:37:00

Context length is capped for computation, and more input is not always better. A million tokens is now a common input scale, while output can reach about 100k tokens or more. A token is about three quarters of a word. As the window grows, context rot can set in, and the model may have more trouble finding one needed fact buried in the middle, so compacting context and starting a new chat can help.

切

In Context Learning 1:40:00

You can improve performance without changing weights by giving examples in the prompt. Few shot learning teaches the format and sometimes the reasoning, but it can also overfit to the examples and add cost and latency. People often turn examples into instructions instead. Chain of thought pushes this further by showing the model how to reason, and self-consistency runs several reasoning paths and uses majority voting to steady the answer.

AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

Summarize your own YouTube video

Paste a YouTube link, article, PDF, ebook or slide deck and get a summary like this in seconds. Free to try, no sign-up needed.

⚔️ Try the YouTube summarizer

Discussion

Sign in to join the discussion. Sign in

More from the Bento Box

Browse the Bento Box →

We use Microsoft Clarity and Google Analytics to see what breaks and where visitors come from. They set cookies and send data to the US. Product events are counted without cookies either way. Cookie details