Conceptual
Login

About How to Build LLMs: Text Generation

Copied

How token generation actually works inside a foundation model, from tokenization through attention to sampling — then how the model gets trained and served. Follows the path Karpathy's nanochat takes to a GPT-2-grade model for around $100.

Estimated Time to Complete

Only available after login

What You'll Learn

Concepts:
Transformer Architecture Single-Node End-to-End Replication of a GPT-2-Grade Model Perplexity as a Language Model Evaluation Metric Instruction Tuning of Large Language Models Autoregressive Language Modeling Muon and Orthogonalised-Momentum Optimisers for Pretraining Tokenizer Artefacts Behind Arithmetic and Spelling Failures Causal Masking That Prevents a Token Attending to Its Future Byte Pair Encoding FlashAttention Neural Scaling Laws Encoder and Vision Models That Still Ship Absolute Positional Embeddings Learned Reward Models Where No Verifier Exists Positional Encoding and Length Extrapolation Dense Models Where Memory Footprint Beats Parameter Count The Unembedding Projection from Hidden State to Vocabulary Logits Group-Relative Policy Optimization and Verifiable Rewards Beam Search Decoding for Sequence Generation Feedforward Networks (MLPs) in Transformers Tokenization Full Multi-Head Attention Where the KV Cache Is Not the Constraint Grouped-Query and Multi-Query Attention Rotary Position Embeddings (RoPE) Self Attention Mechanism in Transformers Quadratic Cost of Self-Attention in Sequence Length Token Embeddings Multi-Head Attention Multi-Head Latent Attention and Low-Rank KV Compression Temperature and Sampling in Language Model Decoding Top-p and Top-k Truncation of the Sampling Tail Residual Connections and Layer Normalization Explicit Attention Matrices for Interpretability and Custom Masks Reinforcement Learning from Human Feedback (RLHF) Pretraining of Large Language Models at Scale Agent-Driven Overnight Search Over Training Code Mixture of Experts in Deep Learning Cross-Entropy as an Optimisation Objective Decoding as a Memory-Bandwidth-Bound Workload Large Language Model Pretraining and Training Corpus Quality Greedy Decoding and Its Degeneration into Repetition Materialising the Full Attention Matrix in High-Bandwidth Memory Softmax Function in Deep Learning Stochastic Gradient Descent and the Adam Optimizer Direct Preference Optimization Without a Separate Reward Model Tokenization and the Vocabulary Trade-Off Dense Parameter Scaling Where Every Token Activates Every Weight KV Caching and Grouped-Query Attention Positional Encoding

What you will learn

No introduction video available

About Johnny Five

J

Guide profile coming soon.