Indexed metadata

Towards a mathematical theory of superposition

Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon

Source record

Source: arXiv

Published: Aug 27, 2026

arXiv: 2608.27540

Open original source ↗

Source abstract

We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector xx of active features is encoded through an overcomplete dictionary WW, and feature recovery is performed by applying ReLU(WWx+b)\operatorname{ReLU}(W^\top W x+b) with an appropriate bias vector bb. We prove several recovery theorems for this model. In the random-support setting, we establish high-probability support recovery for nearly tight, low-coherence dictionaries, with guarantees when the expected sparsity is up to order d/lognd/\log n. In the worst-case support setting, we give a sharp and computable criterion for which sparsity levels permit support recovery. We apply this criterion to Gaussian random matrices and equiangular tight frames. For real equiangular tight frames with n>d+1n>d+1, we determine the exact recovery threshold in terms of the coherence. The proof of this result for real equiangular tight frames relies on a novel characterization---which should be of independent interest to frame theorists---of the distribution of signs in the Gram matrix.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.