Conceptual

NLP for Low-Resource South Asian Languages: A Breadth-First Survey

A breadth-first review of natural-language, multimodal, and speech-processing research published between January 2022 and October 2024 for South Asian languages, with a spotlight on 21 low-resource languages (e.g. Assamese, Bhojpuri, Bodo, Dhivehi, Kashmiri, Khasi, Meitei, Odia, Sindhi). It maps the sub-fields - language models and their adaptation to South Asian scripts, code-mixing, datasets and benchmarks, multimodal machine translation, image captioning, multimodal sentiment and hate-speech detection, and speech modeling and recognition under low-resource conditions - and distills recurring trends, challenges, and research gaps. Methodologically it demonstrates an LLM-driven literature-curation pipeline: GPT-4o performs in-context relevance classification over candidate papers and an O1 reasoning model clusters the relevant set into coherent themes, which outperformed conventional topic modeling.