Feature Redundancy and Training-Window Effects in Financial Time-Series Prediction
An empirical finding that, for time-series financial risk models such as mortgage default prediction, adding more historical training data and more non-critical features can degrade rather than improve accuracy. Students learn why longer windows can inject noise from outdated market regimes and why disciplined feature selection with shorter, recent windows often yields better-generalizing predictions, challenging the default assumption that more data and more features are always better.
2501.00034
This empirical study of mortgage default prediction documents a feature redundancy paradox: contrary to the assumption that longer training histories and more features improve machine-learning models…