Conceptual

Self-Supervised Pretraining: Learning Representations by Masking

Hiding part of an input and training a model to reconstruct it turns unlabelled sequence into a supervision signal. This is how biological foundation models exploit the billions of sequences that have no experimental labels attached.