2501.00418
This paper studies whether trustworthiness properties transfer under weak-to-strong generalization, the phenomenon where a large 'strong' language model fine-tuned only on a smaller 'weak' model's ou…
A study of whether trustworthiness properties, not just task accuracy, transfer under weak-to-strong generalization, where a large model fine-tuned only on a smaller model's outputs surpasses it. The learner sees two training strategies: Weak Trustworthiness Fine-tuning adds a fairness, privacy, or robustness regularizer to the weak model's loss, and Weak+WTS Fine-tuning also regularizes the strong model during transfer. Key finding: fairness and adversarial/out-of-distribution robustness transfer when both models are regularized, whereas differential-privacy protection does not reliably carry over.
This paper studies whether trustworthiness properties transfer under weak-to-strong generalization, the phenomenon where a large 'strong' language model fine-tuned only on a smaller 'weak' model's ou…