Indexed metadata

Gaussian Equivalence for Multi-Head Self-Attention

Tomohiro Hayase, Ryo Karakida

Source record

Source: arXiv

Published: Oct 7, 2026

arXiv: 2610.10033

Open original source ↗

Source abstract

A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.

Gaussian Equivalence for Multi-Head Self-Attention — Mathematical Frontier Network