Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers
Summary
This lecture provides a deep dive into the mechanics of Transformers and Large Language Models (LLMs), tracing the evolution of NLP from early models like RNNs to modern, scalable architectures. The core focus is on the self-attention mechanism, which allows tokens to compute representations based on all other tokens in the sequence. The lecture details the structure of the Transformer (Encoder and Decoder), the role of specialized tokens (e.g., `[BOS]`, `[EOS]`), and critical engineering components like position embeddings, residual connections, and multi-head attention.
Key takeaways
-
The Transformer Architecture
18:13
The Transformer, introduced in 2017, replaced recurrent models (RNNs) by relying entirely on attention mechanisms, enabling massive scalability and parallel processing. It consists of an Encoder (processes source text) and a Decoder (generates the target sequence).
-
Self-Attention Mechanism
22:13
The representation of a token is computed as a weighted average of all other tokens, determined by the Query (Q), Key (K), and Value (V) vectors. The core calculation involves the softmax of (Q * K^T) / $\sqrt{d_k}$ multiplied by V.
-
Tokenization Techniques
3:36
Text must be converted to numbers. Subword tokenization, specifically Byte Pair Encoding (BPE), is the most popular method as it balances vocabulary size and sequence length, improving robustness and efficiency.
-
Model Enhancements
26:13
Modern implementations use techniques like residual connections (to aid gradient flow in deep networks), layer normalization (for stable convergence), and masked self-attention (in the decoder) to ensure causal generation.
Technical details
-
Tokenization
216s
Text is divided into tokens (e.g., using BPE) and mapped to a vocabulary. Special tokens like `[UNK]` (unknown), `[BOS]` (beginning of sequence), `[EOS]` (end of sequence), and `[PAD]` (padding) are essential for defining sequence boundaries and dimensions.
-
Word Representation
3727s
Initial methods like Word2Vec (CBOW, Skip-gram) learned representations through proxy tasks (e.g., predicting surrounding words). Limitations included ignoring word order and failing to capture polysemy (multiple meanings).
-
Attention Mechanism
595s
Attention allows a token to directly connect to and weigh the importance of all other tokens in the sequence. The mechanism uses three learned projections: Query (Q), Key (K), and Value (V).
-
Transformer Encoder
722s
The Encoder computes context-aware representations for the input sequence using a self-attention layer followed by a feed-forward neural network. It processes the entire input sequence simultaneously.
-
Transformer Decoder
1222s
The Decoder generates the output autoregressively. It uses two attention layers: 1) Masked self-attention (to prevent looking at future tokens) and 2) Cross-attention (to relate the generated tokens to the source input from the Encoder).
Mentioned resources
- Stanford CME295 Syllabus
- Super Study Guide Transformers and LLMs
- Stanford Online Graduate Education
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.