Stanford Online

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers

Published 2026-09-28 · Duration 1:44:28

Summary

This lecture provides a deep dive into the mechanics of Transformers and Large Language Models (LLMs), tracing the evolution of NLP from early models like RNNs to modern, scalable architectures. The core focus is on the self-attention mechanism, which allows tokens to compute representations based on all other tokens in the sequence. The lecture details the structure of the Transformer (Encoder and Decoder), the role of specialized tokens (e.g., `[BOS]`, `[EOS]`), and critical engineering components like position embeddings, residual connections, and multi-head attention.

Download summary

Key takeaways

  1. The Transformer Architecture 18:13

    The Transformer, introduced in 2017, replaced recurrent models (RNNs) by relying entirely on attention mechanisms, enabling massive scalability and parallel processing. It consists of an Encoder (processes source text) and a Decoder (generates the target sequence).

  2. Self-Attention Mechanism 22:13

    The representation of a token is computed as a weighted average of all other tokens, determined by the Query (Q), Key (K), and Value (V) vectors. The core calculation involves the softmax of (Q * K^T) / $\sqrt{d_k}$ multiplied by V.

  3. Tokenization Techniques 3:36

    Text must be converted to numbers. Subword tokenization, specifically Byte Pair Encoding (BPE), is the most popular method as it balances vocabulary size and sequence length, improving robustness and efficiency.

  4. Model Enhancements 26:13

    Modern implementations use techniques like residual connections (to aid gradient flow in deep networks), layer normalization (for stable convergence), and masked self-attention (in the decoder) to ensure causal generation.

Technical details

  • Tokenization 216s

    Text is divided into tokens (e.g., using BPE) and mapped to a vocabulary. Special tokens like `[UNK]` (unknown), `[BOS]` (beginning of sequence), `[EOS]` (end of sequence), and `[PAD]` (padding) are essential for defining sequence boundaries and dimensions.

  • Word Representation 3727s

    Initial methods like Word2Vec (CBOW, Skip-gram) learned representations through proxy tasks (e.g., predicting surrounding words). Limitations included ignoring word order and failing to capture polysemy (multiple meanings).

  • Attention Mechanism 595s

    Attention allows a token to directly connect to and weigh the importance of all other tokens in the sequence. The mechanism uses three learned projections: Query (Q), Key (K), and Value (V).

  • Transformer Encoder 722s

    The Encoder computes context-aware representations for the input sequence using a self-attention layer followed by a feed-forward neural network. It processes the entire input sequence simultaneously.

  • Transformer Decoder 1222s

    The Decoder generates the output autoregressively. It uses two attention layers: 1) Masked self-attention (to prevent looking at future tokens) and 2) Cross-attention (to relate the generated tokens to the source input from the Encoder).

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.