Conceptual

Data Augmentation for Chinese Disease Name Normalization in Clinical NLP

A collection of data-augmentation techniques, with supporting modules, that synthesize additional labeled examples for mapping free-text disease-name variants onto standardized names or codes. Designed for the severe data scarcity of Chinese clinical text, where most diseases appear in few-shot or zero-shot form, the approach delivers consistent accuracy gains across multiple baseline models and training objectives, especially in low-resource regimes.