Indexed metadata

KGEN: A Knowledge-Guided Evolutionary Framework for Feature Selection in Breast Cancer Risk Assessment

Saadia Humayun, Tariq Mahmood, Munira Moosajee, Muhammad Shoaib Siddiqui

Source record

Source: Crossref

Published: Oct 2, 2026

DOI: 10.3390/math14193591

Open original source ↗

Source abstract

Choosing which features go into a clinical risk model is not only a statistical problem. A model that is accurate but whose feature list makes no clinical sense will not be trusted by the physicians it is meant to help, and most feature-selection methods still include a clinician’s judgment only as a filter applied after the search, not as part of it. This paper introduces KGEN, a genetic algorithm that builds an oncologist’s feature ratings directly into both the starting population and the fitness function, alongside data-driven importance scores, through one tunable weight. We tested this on 67,858 women from the PLCO Cancer Screening Trial, searching over 75 candidate features, against a matched no-expert control run under an identical search budget. Across four independent runs, expert guidance raised clinical relevance substantially and consistently (0.929–1.000 versus 0.682 for the no-expert control and 0.669–0.720 for four of the five conventional baselines, the exception being BCRAT, whose fixed feature set happens to reach a perfect 1.000; pooled p=0.0019 across sub-populations), while discrimination stayed within the same narrow range as every conventional baseline throughout, and, in three of the four sub-populations, was even modestly but significantly higher than the no-expert control’s: a clinically sensible feature list did not come at the cost of accuracy. The same wide-margin clinical-relevance pattern held on an independent hospital-readmission dataset of 99,343 encounters and an independent breast-cancer gene-expression cohort, each with a small, measurable accuracy trade-off against the strongest conventional baselines. Feature-selection stability across runs was also higher under expert guidance for three of four sub-populations. This suggests that a clinician’s judgment can be built into a search algorithm as a working input, not just a check performed afterward.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.