Education

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models

by Stanford Online

Share:

📚 Main topics

  • Transformer variantsThe lecture contrasts encoder-only models such as BERT, which produce contextual representations for downstream tasks, with decoder-only GPT models, which generate text through next-token prediction. It explains why decoder-only models became the dominant general-purpose approach. 4:50
  • Pre-training and fine-tuningBERT is pre-trained using masked language modeling and next-sentence prediction, then adapted to tasks such as sentiment classification. Decoder-only models can often handle tasks by framing them as text-to-text prompts, without requiring task-specific fine-tuning. 7:30
  • Scaling and computeThe lecture reviews scaling trends and compute-optimal training, including the finding that models can be undertrained when their parameter count is too large relative to the training tokens available. It distinguishes total floating-point operations from hardware throughput. 19:33
  • Mixture of expertsSparse mixture-of-experts models route each token to a subset of feed-forward experts, allowing a model to have many parameters without activating all of them for every token. The lecture also covers routing collapse and load-balancing losses. 29:28
  • Position information and attentionThe lecture explains why transformers need positional information, introduces rotary position embeddings (RoPE), and discusses local or sliding-window attention as a way to limit computation while stacked layers still pass information across longer distances. 45:18
  • Efficient transformer componentsThe lecture covers grouped-query attention, which reduces key/value memory use during decoding, and describes the shift from post-normalization and LayerNorm toward pre-normalization and RMSNorm. 1:13:39
  • Decoding and promptingThe lecture compares greedy decoding and beam search with sampling, explains temperature and guided decoding for structured outputs, and introduces context windows, in-context learning, chain-of-thought prompting, and self-consistency. 1:21:34

✨ Key takeaways

  • LLMs predict tokensA language model assigns probabilities to token sequences; decoder-only LLMs generate by repeatedly predicting one token at a time. Their scale involves not just parameters, but also training data and compute. 26:49
  • Scale has trade-offsBigger models may perform better and use training tokens more efficiently, but limited compute calls for balancing model size against the amount of training data. Serving costs also matter in production. 20:06
  • Sparse models need careful routingActivating only selected experts can reduce computation per input while supporting a larger overall model, but effective training requires preventing the router from overusing a small set of experts. 30:32
  • Modern attention is optimized in several waysRoPE adds positional information directly to queries and keys, sliding-window attention limits each layer’s local interactions, and grouped-query attention reduces the number of key/value representations that must be stored. 1:02:00
  • More context is not always betterLong contexts can make it harder for a model to retrieve a particular detail, and generated tokens also occupy context-window capacity. 1:35:40

🧠 Lessons learned

  • Choose decoding to fit the taskSampling can make responses less repetitive, while deterministic decoding can be useful when outputs need to be reproducible. Temperature adjusts how concentrated or broad the sampling distribution is. 1:24:40
  • Constrain outputs when format mattersPrompting a model to return JSON may work, but guided decoding can restrict the next-token choices to those allowed by the required syntax. 1:33:39
  • Improve behavior through promptsIn-context examples can demonstrate a task or format without changing model weights, though they add input cost and may lead the model to overfit to the examples. Instructions and chain-of-thought prompting offer related approaches. 1:39:50
  • Use multiple reasoning paths selectivelySelf-consistency generates several candidate reasoning paths and can stabilize an answer through majority voting, though it requires additional generations. 1:41:55

🏁 Conclusion/next steps

  • Course directionThe lecture connects the original transformer to today’s LLMs and summarizes the architectural and inference techniques used to make them more capable and efficient. Later lectures will return to topics including encoder-only text generation, training stages, evaluation, and reasoning models. 13:14

Transcript excerpt

0:05 Hello everyone and welcome to lecture two of CME295. So today is an exciting day because we'll be talking about large language models and in particular how they relate to the transformer architecture that we saw last lecture. But before we start, we will recap what we saw last time and we will do that at every lecture to see how this connect to what we'll see today. So if you remember last time we were

0:38 wondering how models could process text. So models they don't understand text as it is they understand numbers. So if you remember one thing that we looked at was first of all how to divide our text into indivisible units and this process is called tokenization. And we saw a few kinds of tokenization algorithms. And we said that the most common tokenization algorithm these days is the subword tokenization because it

1:10 has a nice tradeoff between the vocabulary that it was building and the sequence length that input text would have. But this was not it because once you divide a text into tokens the next step is to compute embeddings. So if you remember we saw Word2vec as a way to compute embeddings but they were not aware of the context. So then we talked about another class of models that were

🔒 The full, searchable transcript is available with Pro.

🔒 Unlock Premium Features

This is a premium feature. Upgrade to unlock unlimited Q&A, transcripts, mindmaps, and translations.

Questions & Answers

Common questions about this video

¿En qué se diferencian los modelos basados solo en el encoder, como BERT, de los modelos basados solo en el decoder, como GPT?

BERT usa el encoder para generar representaciones bidireccionales de los tokens y suele adaptarse a tareas específicas mediante fine-tuning. GPT usa solo el decoder y predice el siguiente token de forma autoregresiva, por lo que puede plantear muchas tareas como generación de texto sin requerir necesariamente fine-tuning. 4:50

¿Qué es el enmascaramiento de lenguaje (MLM) en BERT y para qué sirve?

En el preentrenamiento, se seleccionan y corrompen algunos tokens de la secuencia. El modelo usa el contexto para predecir los tokens originales, lo que le ayuda a aprender cómo se escribe el lenguaje. 8:02

¿Qué es una mezcla dispersa de expertos y cómo se evita que el modelo dependa siempre de los mismos expertos?

Una mezcla dispersa de expertos emplea un enrutador para seleccionar solo algunos expertos —normalmente los top-k— para cada token, en vez de activar todos. Una pérdida auxiliar de balanceo de carga incentiva que los tokens se distribuyan entre los expertos y ayuda a prevenir el colapso del enrutamiento. 30:32

¿Por qué los transformers modernos suelen usar RoPE en lugar de sumar embeddings posicionales a los embeddings de tokens?

Los embeddings posicionales aditivos afectan las representaciones que pasan por todas las capas, aunque la información de posición resulta especialmente relevante en la atención. RoPE modifica directamente las consultas y las claves según sus posiciones; su producto interno incorpora la diferencia entre ellas y se generaliza a distintas posiciones. 1:02:00

¿Qué hacen la atención local y GQA, y por qué son útiles?

La atención local o de ventana deslizante limita las posiciones a las que atiende cada token, reduciendo el alcance de la atención en cada capa; al apilar capas, la información puede propagarse más lejos. GQA comparte proyecciones de claves y valores entre grupos de cabezas, reduciendo la memoria necesaria para almacenarlos durante la generación. 1:09:24

¿Cómo se puede controlar la variabilidad de las respuestas al generar texto?

Elegir siempre el token más probable o usar beam search produce resultados deterministas. Para generar respuestas más variadas, se puede muestrear de la distribución de probabilidades, por ejemplo con top-k o top-p. La temperatura también influye: una temperatura baja hace la distribución más concentrada, mientras que una alta la vuelve más plana. 1:26:46

¿Qué es el context rot y por qué una ventana de contexto más larga no siempre mejora el rendimiento?

El context rot es la dificultad creciente del modelo para recuperar un dato concreto cuando aumenta la longitud del contexto. En pruebas como “needle in a haystack”, se observa que más información alrededor puede dificultar encontrar el dato relevante. 1:38:14

¿Qué es el aprendizaje en contexto y cómo se relaciona con chain of thought?

El aprendizaje en contexto consiste en proporcionar ejemplos o instrucciones en el prompt para orientar al modelo sin cambiar sus pesos. Chain of thought extiende esta idea al incluir ejemplos que muestran cómo llegar a una respuesta, lo que puede mejorar el rendimiento al enseñar un recorrido de razonamiento. 1:39:50

🔒 Unlock Premium Features

Access to Chat is a premium feature. Upgrade now to unlock unlimited studying tools.

🔒 Unlock Premium Features

Access to Mindmap is a premium feature. Upgrade now to unlock unlimited studying tools.

🔒 Unlock Premium Features

Access to Translation is a premium feature. Upgrade now to unlock unlimited studying tools.

Get unlimited summaries, Q&A, transcripts and more with Pro

Upgrade to Pro

Suggestions

🔒 Unlock Premium Features

Access to AI Suggestions is a premium feature. Upgrade now to unlock unlimited studying tools.