Indexed metadata

An equality condition for the Dobrushin bound on attention rollout and how often it holds in trained transformers

Przemysław Rola

Source record

Source: arXiv

Published: Oct 4, 2026

arXiv: 2610.05558

Open original source ↗

Source abstract

The Dobrushin coefficient of each attention-rollout factor satisfies κ(12(I+A))≤12(1+κ(A))κ(\frac12(I+A))\le\frac12(1+κ(A)), and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is tight: equality holds if and only if some token pair attaining κ(A)κ(A) is mutually self-dominant - each of the two attends to itself at least as strongly as the other attends to it. The condition is far from automatic: uniformly random stochastic matrices satisfy it only 24-30% of the time. When tested on the head-averaged attention of each individual input and restricted to content tokens - image patches, words or tabular features, excluding cls, register and separator tokens - the condition holds for essentially every input at every layer of DINOv2 (three model sizes), RoBERTa and DistilBERT. In the supervised models DeiT-B and ViT-B/16 it holds for 91% and 64% of input-layer pairs respectively, with all failures occurring late in depth. In FT-Transformer trained on two standard tabular benchmarks it holds for only 11-44% of input-layer pairs. The special tokens account for almost all failures in DINOv2 and the language models: when they are included, the condition holds for only 82-97% of input-layer pairs.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.

An equality condition for the Dobrushin bound on attention rollout and how often it holds in trained transformers — Mathematical Frontier Network