Conceptual

Superposition of Base and Fine-Tuned Model Representations in Transformers

An architecture that mitigates catastrophic forgetting by superimposing the hidden representations of a base Transformer and a fine-tuned copy within one shared parameter space, instead of overwriting weights during fine-tuning. Students learn how autoencoders adaptively reconstruct the appropriate hidden state from the input distribution and how B-spline blending coefficients allow dynamic switching between the base and specialized model behaviors at inference.