From Gradients to Combinatorics: Sparse Combinatorial Training of All-ReLU Neural Networks
Carlo Lucheroni
Source abstract
DiBa, a gradient-free training framework for all-ReLU neural network architectures is introduced. It is based on a random dictionary of hinge functions acting as neurons, a sparse parametrically controllable combinatorial selection of these neurons by a hyperparameter , on linear optimization, and on a loss. Within this setting, training can be formulated as a mixed-integer linear program, which goes to replace nonconvex gradient optimization by exact combinatorial-linear optimization. Given , once a dictionary is randomly drawn, DiBa can optimally select on training data a subset of neurons while at once it tunes their weights. Hence acts as a model-complexity parameter. Error performance on the test set for different values leads to the identification of the best DiBa model, that is, the best network structure. The framework can thus be also used in a procedure of model regularization by validation. The framework is evaluated on electricity-market forecasting tasks involving price and volume time series. Compared with dense all-ReLU Extreme Learning Machines, and gradient-trained all-ReLU feedforward neural networks under , DiBa achieves much higher flexibility, better forecasting performance, and lower sensitivity to randomness across an experimental set of 50 independent random initializations of DiBa and compared models. The framework also admits a natural constrained-quadratic reformulation suitable for hybrid quantum-classical optimization. These results suggest that sparse combinatorial training can provide a viable architecture-optimization alternative to gradient-based learning for a class of all-ReLU models.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.