Conceptual

TinyHelen: A Leaner-Data Pipeline for Training and Evaluating Tiny Language Models

A data-refinement pipeline that uses large language models to construct a simplified ('leaner') language environment for small language models: it eliminates dataset noise, minimizes vocabulary, and preserves genre-specific linguistic patterns while keeping the text distribution aligned with conventional large-LM corpora. The pipeline yields matched pretraining, instruction-tuning, and evaluation datasets on which tiny language models achieve higher learning efficiency and stronger instruction-following than when trained on the original, unrefined data, enabling resource-efficient study of learning objectives, architectures, and training techniques.