# Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers

## Executive summary

This lecture provides a deep dive into the mechanics of Transformers and Large Language Models (LLMs), tracing the evolution of NLP from early models like RNNs to modern, scalable architectures. The core focus is on the self-attention mechanism, which allows tokens to compute representations based on all other tokens in the sequence. The lecture details the structure of the Transformer (Encoder and Decoder), the role of specialized tokens (e.g., `[BOS]`, `[EOS]`), and critical engineering components like position embeddings, residual connections, and multi-head attention.

## Key takeaways

- The Transformer Architecture: The Transformer, introduced in 2017, replaced recurrent models (RNNs) by relying entirely on attention mechanisms, enabling massive scalability and parallel processing. It consists of an Encoder (processes source text) and a Decoder (generates the target sequence).
- Self-Attention Mechanism: The representation of a token is computed as a weighted average of all other tokens, determined by the Query (Q), Key (K), and Value (V) vectors. The core calculation involves the softmax of (Q * K^T) / $\sqrt{d_k}$ multiplied by V.
- Tokenization Techniques: Text must be converted to numbers. Subword tokenization, specifically Byte Pair Encoding (BPE), is the most popular method as it balances vocabulary size and sequence length, improving robustness and efficiency.
- Model Enhancements: Modern implementations use techniques like residual connections (to aid gradient flow in deep networks), layer normalization (for stable convergence), and masked self-attention (in the decoder) to ensure causal generation.

## Technical details

- Tokenization: Text is divided into tokens (e.g., using BPE) and mapped to a vocabulary. Special tokens like `[UNK]` (unknown), `[BOS]` (beginning of sequence), `[EOS]` (end of sequence), and `[PAD]` (padding) are essential for defining sequence boundaries and dimensions.
- Word Representation: Initial methods like Word2Vec (CBOW, Skip-gram) learned representations through proxy tasks (e.g., predicting surrounding words). Limitations included ignoring word order and failing to capture polysemy (multiple meanings).
- Attention Mechanism: Attention allows a token to directly connect to and weigh the importance of all other tokens in the sequence. The mechanism uses three learned projections: Query (Q), Key (K), and Value (V).
- Transformer Encoder: The Encoder computes context-aware representations for the input sequence using a self-attention layer followed by a feed-forward neural network. It processes the entire input sequence simultaneously.
- Transformer Decoder: The Decoder generates the output autoregressively. It uses two attention layers: 1) Masked self-attention (to prevent looking at future tokens) and 2) Cross-attention (to relate the generated tokens to the source input from the Encoder).

## Practical implications

- Understanding LLM mechanisms is crucial for achieving 'AI literacy,' regardless of professional role.
- Knowledge of these architectures is vital for building personal projects and understanding the capabilities and limitations of coding agents.
- The principles of attention and Transformers are foundational to modern AI systems, including advanced chatbots and agentic behavior.

## Topics

Large Language Models (LLMs), Transformer Architecture, Natural Language Processing (NLP), Self-Attention, Byte Pair Encoding (BPE), Recurrent Neural Networks (RNNs), Stanford CME295 Syllabus, Super Study Guide Transformers and LLMs, Stanford Online Graduate Education

Source: https://www.youtube.com/watch?v=114i2Kz-LZA
