Conceptual

Trivial Baselines as the Gate on Interpretability Claims from Attention Weights

An interpretability claim about attention weights is only meaningful if the attention-derived signal survives comparison against features that require no model at all. Auditing single-cell transformer foundation models against CRISPR perturbation outcomes shows attention encoding genuine, layer-organised biological structure while adding zero incremental value over univariate gene-level features such as expression variance, and causal ablation of the heads ranked most regulatory degrading the model less than ablating random heads. The pattern generalises past its biological setting: an apparent pairwise relational signal can be a re-expression of per-item properties, so a trivial-baseline test, a conditional incremental-value test under hard splits, residualisation or propensity matching, and ablation paired with intervention-fidelity diagnostics are what separate a real relational structure from a confounded one.