Optimizing Regularized Logistic Regression Pipelines on Cleaned Biomarker Datasets: A Leakage-Proof Benchmarking Protocol for Diagnostic and Prognostic Modeling
Stanley Nwakamma, Serif Oyindamola Oyesiji, Harouna Wendpanga Yann Christian Sankara, Toussida Fatah Tanguy Minoungou
Source record
Source: Crossref
Published: Jan 1, 2026
DOI: 10.54660/ijaiet.2026.7.2.64-72
Open original source ↗Source abstract
Logistic regression remains the workhorse of biomarker-based diagnostic and prognostic modeling, and for principled reasons: it yields calibrated probabilities when fitted and validated correctly, its coefficients support biological and clinical interpretation, and its statistical behavior under limited events is thoroughly characterized. Yet the gap between the method's textbook properties and its published practice is wide and documented, with prediction research reviews finding pervasive deficiencies in validation, calibration assessment, sample size justification, and handling of missing data, deficiencies that persist even in high-stakes settings. The availability of cleaned, curated biomarker datasets removes the data-quality alibi and exposes the remaining variance for what it is: modeling protocol. This paper develops a research concept for a systematic optimization and benchmarking protocol for logistic regression pipelines on cleaned biomarker datasets. The protocol fixes a leakage-proof evaluation harness, nested resampling with all preprocessing inside the resampling loop and grouping that respects study and batch structure, and then treats the pipeline's design space as the experimental object: regularization family and strength, feature selection strategy including stability-based selection, treatment of continuous predictors, class imbalance handling, missing data strategy, rare-event corrections, and post-hoc calibration methods, each varied under prespecified contrasts. Performance is reported as a profile, discrimination with uncertainty, calibration curves and indices, and decision-analytic net benefit across clinically relevant thresholds, never as a single headline number, with optimism-corrected internal validation and temporal or cohort-external validation where data permit. Prespecified hypotheses target the conditions under which elastic net dominates unpenalized fitting, when stability selection changes selected panels, and how far calibration methods repair miscalibration without refitting. Reporting follows prediction-model standards throughout.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.