Indexed metadata

From Gradients to Combinatorics: Sparse Combinatorial Training of All-ReLU Neural Networks

Carlo Lucheroni

Source record

Source: Crossref

Published: Jan 1, 2026

DOI: 10.2139/ssrn.7422314

Open original source ↗

Source abstract

DiBa, a gradient-free training framework for all-ReLU neural network architectures is introduced. It is based on a random dictionary of KK hinge functions acting as neurons, a sparse parametrically controllable combinatorial selection of these neurons by a hyperparameter MKM\le K, on linear optimization, and on a L1L_1 loss. Within this setting, training can be formulated as a mixed-integer linear program, which goes to replace nonconvex gradient optimization by exact combinatorial-linear optimization. Given MM, once a dictionary is randomly drawn, DiBa can optimally select on training data a subset of neurons while at once it tunes their weights. Hence MM acts as a model-complexity parameter. Error performance on the test set for different values MM leads to the identification of the best DiBa model, that is, the best network structure. The framework can thus be also used in a procedure of model regularization by validation. The framework is evaluated on electricity-market forecasting tasks involving price and volume time series. Compared with dense all-ReLU Extreme Learning Machines, and gradient-trained all-ReLU feedforward neural networks under L1L_1, DiBa achieves much higher flexibility, better forecasting performance, and lower sensitivity to randomness across an experimental set of 50 independent random initializations of DiBa and compared models. The framework also admits a natural constrained-quadratic reformulation suitable for hybrid quantum-classical optimization. These results suggest that sparse combinatorial training can provide a viable architecture-optimization alternative to gradient-based learning for a class of all-ReLU L1L_1 models.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.