Estimated Time to Complete
Only available after login
What You'll Learn
Concepts:
Transformer Architecture
Single-Node End-to-End Replication of a GPT-2-Grade Model
Perplexity as a Language Model Evaluation Metric
Instruction Tuning of Large Language Models
Autoregressive Language Modeling
Muon and Orthogonalised-Momentum Optimisers for Pretraining
Tokenizer Artefacts Behind Arithmetic and Spelling Failures
Causal Masking That Prevents a Token Attending to Its Future
Byte Pair Encoding
FlashAttention
Neural Scaling Laws
Encoder and Vision Models That Still Ship Absolute Positional Embeddings
Learned Reward Models Where No Verifier Exists
Positional Encoding and Length Extrapolation
Dense Models Where Memory Footprint Beats Parameter Count
The Unembedding Projection from Hidden State to Vocabulary Logits
Group-Relative Policy Optimization and Verifiable Rewards
Beam Search Decoding for Sequence Generation
Feedforward Networks (MLPs) in Transformers
Tokenization
Full Multi-Head Attention Where the KV Cache Is Not the Constraint
Grouped-Query and Multi-Query Attention
Rotary Position Embeddings (RoPE)
Self Attention Mechanism in Transformers
Quadratic Cost of Self-Attention in Sequence Length
Token Embeddings
Multi-Head Attention
Multi-Head Latent Attention and Low-Rank KV Compression
Temperature and Sampling in Language Model Decoding
Top-p and Top-k Truncation of the Sampling Tail
Residual Connections and Layer Normalization
Explicit Attention Matrices for Interpretability and Custom Masks
Reinforcement Learning from Human Feedback (RLHF)
Pretraining of Large Language Models at Scale
Agent-Driven Overnight Search Over Training Code
Mixture of Experts in Deep Learning
Cross-Entropy as an Optimisation Objective
Decoding as a Memory-Bandwidth-Bound Workload
Large Language Model Pretraining and Training Corpus Quality
Greedy Decoding and Its Degeneration into Repetition
Materialising the Full Attention Matrix in High-Bandwidth Memory
Softmax Function in Deep Learning
Stochastic Gradient Descent and the Adam Optimizer
Direct Preference Optimization Without a Separate Reward Model
Tokenization and the Vocabulary Trade-Off
Dense Parameter Scaling Where Every Token Activates Every Weight
KV Caching and Grouped-Query Attention
Positional Encoding
What you will learn
No introduction video available
About Johnny Five
J
Guide profile coming soon.