BézierFormer: Affine-Invariant Shape Classification via Control Point Attention
Xiao Liu, Jean-Michel Morel, Roy Y. He
Source record
Source: Crossref
Published: Jun 27, 2026
DOI: 10.1007/s10851-026-01322-9
Open original source ↗Source abstract
Abstract We address the resolution paradox in modern deep learning, where networks receive far more spatial information than they demonstrably utilize for shape classification. We show that it is possible to train networks directly on spatially sparse and structurally compressed shape representations rather than dense pixel grids. Specifically, we extract vector graphic representations of shapes from raster images and train on control points of curves, which naturally encode the sparse, localized features (corners, curvature extrema) that both human vision and interpretability studies identify as critical for recognition. To effectively process these sparse geometric primitives, we propose an attention-based architecture, BézierFormer, that processes each parametric curve independently through shared-weight transformations and then synthesizes global shape understanding through tailored attention mechanisms. This combination of sparse vector graphic training data and segment-wise processing with attention-based synthesis achieves computational efficiency while maintaining high discriminative power, demonstrating that classification can be performed with dramatically fewer geometric primitives than pixels in conventional approaches.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.