Indexed metadata

Spectral Fingerprinting: Mathematical Identification of Generative Audio via Deconvolution Artifacts

Yaniv Proselkov, Aneta Kahleova, Matej Štágl

Source record

Source: Crossref

Published: Jan 1, 2026

DOI: 10.2139/ssrn.7578198

Open original source ↗

Source abstract

The rapid proliferation of high-fidelity generative music models, in particular Suno and Udio, has created a need for reliable forensic tools that can distinguish synthetic from human-produced audio. This paper presents the mathematical foundations and empirical evaluation of ai-music-detector, our open-source framework for detecting AI-generated music. Building on the observation that the transposed-convolution (deconvolution) layers used for upsampling in neural vocoders inject deterministic, architecture-dependent periodic artifacts into the spectrum, we formalise these artifacts as the result of a periodic overlap-count modulation and of incompletely suppressed spectral images. From this analysis we derive a compact forensic feature-a normalised spectral-residual "fakeprint" computed in the 1-8 kHz band-and classify it with a regularised linear model. A complementary convolutional model, operating on a cepstral representation that is unaffected by pitch shifting, provides robustness to common audio manipulations. On a held-out test set of 17,866 tracks drawn from FMA, a corpus of human-made music, and SONICS, a corpus of Suno and Udio songs, the linear fakeprint classifier attains 99.88% accuracy (F 1 = 0.9991), confirming that the generative pipelines represented in that corpus remain readily detectable in the frequency domain. A controlled robustness comparison further shows that the linear and convolutional models are complementary: a pitch shift defeats the former while leaving the latter intact, whereas band-limiting defeats the latter while leaving the former intact. A preliminary cross-generator probe qualifies that figure; on generators absent from the training pool the linear model detects almost nothing, and the convolutional model transfers only to generators sharing the same transposed-convolution decoder, indicating that what is being detected is a property of the decoder architecture rather than of AI-generated music in general.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.

Spectral Fingerprinting: Mathematical Identification of Generative Audio via Deconvolution Artifacts — Mathematical Frontier Network