Conceptual

Multi-Dimensional Attack-Defense Data Construction for LLM Safety Alignment

A data-centric technique for improving the safety of large language models. Safety-alignment training data is enriched by expanding the diversity of adversarial attack instructions across many intent categories and by regenerating higher-quality safe responses, then filtering both with a perplexity- and LLM-score-based safety reward model before supervised fine-tuning.