arrow_back
AI

AI & Digital Business Masterclass

Student
Lesson 1.1 • Video

1.1 The Evolution of Transformer Architectures

Lecture Synopsis

In this comprehensive opening session, Dr. Jonathan Vance explores the historical progression from Recurrent Neural Networks (RNNs) and LSTMs to the seminal 2017 paper "Attention Is All You Need". We break down Scaled Dot-Product Attention, Multi-Head projections, Positional Encoding vectors, and GPU parallelization advantages.

Key Technical Concepts

  • Scaled Dot-Product Attention: Softmax(QK^T / √d_k)V
  • Sinusoidal vs Rotary Positional Embeddings (RoPE)
  • Causal Attention Masking for Autoregressive Decoders

Instructor Credentials

Instructor

Dr. Jonathan Vance

Chief AI Architect & Ex-Google DeepMind