Conceptual

Weak-to-Strong Trustworthiness Generalization in Language Model Fine-Tuning

A study of whether trustworthiness properties, not just task accuracy, transfer under weak-to-strong generalization, where a large model fine-tuned only on a smaller model's outputs surpasses it. The learner sees two training strategies: Weak Trustworthiness Fine-tuning adds a fairness, privacy, or robustness regularizer to the weak model's loss, and Weak+WTS Fine-tuning also regularizes the strong model during transfer. Key finding: fairness and adversarial/out-of-distribution robustness transfer when both models are regularized, whereas differential-privacy protection does not reliably carry over.