Indexed metadata

No Equivariant Architecture Covers All Equivariant Attention

Tīkun Ông

Source record

Source: arXiv

Published: Aug 31, 2026

arXiv: 2608.30417

Open original source ↗

Source abstract

We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group GG, then GG can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: the equivariance locus of unconstrained MHSA forms a union of extremely many Zariski-irreducible components in a reduced parameter space, and any single architecture covers at most one. For G=D4G=D_4 acting on CC copies of the regular representation as the token feature space, we show that there are Ω(C64)Ω(C^{64}) components for eight attention heads.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.