Conceptual

Tokenization

Splitting raw text into discrete units — words, subwords, or bytes — that become the model's vocabulary. Subword schemes like BPE bound vocabulary size while keeping rare words representable.