Lesson 1.1 • Video
1.1 The Evolution of Transformer Architectures
Lecture Synopsis
In this comprehensive opening session, Dr. Jonathan Vance explores the historical progression from Recurrent Neural Networks (RNNs) and LSTMs to the seminal 2017 paper "Attention Is All You Need". We break down Scaled Dot-Product Attention, Multi-Head projections, Positional Encoding vectors, and GPU parallelization advantages.
Key Technical Concepts
- Scaled Dot-Product Attention: Softmax(QK^T / √d_k)V
- Sinusoidal vs Rotary Positional Embeddings (RoPE)
- Causal Attention Masking for Autoregressive Decoders
Instructor Credentials
Dr. Jonathan Vance
Chief AI Architect & Ex-Google DeepMind