Conceptual

Orthographic Variation Between Dialog and Dialogue in NLP Literature

A quantitative meta-study of how the NLP and AI research community spells the word dialog(ue). Using a Semantic Scholar corpus of tens of thousands of papers, it measures the distribution of 'dialogue', 'dialog', and mixed use across venues, disciplines, and two decades, and tests whether author nationality or linguistic context predicts the choice using dependency parses, masked-language-model embeddings, and logistic models. Students learn how to design a corpus-based scientometric study of orthographic variation and how to weigh its statistical evidence, including bootstrap intervals and false-discovery-rate-corrected significance.