Conceptual

Pre-Trained Transformers Learning Spectral Estimation Algorithms

A theoretical and empirical study showing that multi-layer Transformers, trained across many instances, can internalize spectral (eigenvector-based) estimation algorithms rather than only learning in context. A constructive proof maps stacked Transformer layers onto iterative eigenvector-recovery steps, using two-component Gaussian-mixture classification as the example, and experiments confirm the network performs PCA and clustering.