A Disease-Guided Representative Gene Selection Framework for High-Dimensional Gene Expression Analysis
Cihan Kuzudisli, Bahjat F. Qaqish, Burcu Bakir-Gungor, Malik Yousef
Source abstract
Gene expression datasets provide valuable information for disease classification and biomarker discovery; however, their high dimensionality and limited sample size may limit classification performance and reduce biological interpretability. This study proposes GeDiRep, a prior knowledge-guided framework for identifying compact and informative gene subsets. The method first organizes filtered genes into disease-associated groups using curated gene–disease associations from DisGeNET. Each group is then scored according to its predictive contribution, and representative genes are selected using Random Forest-based feature importance. Representative genes from the top-ranked groups are progressively accumulated, and the resulting gene subsets are used to assess classification performance on the test set. Experiments on eight microarray datasets showed that GeDiRep reduced the average number of selected features from 40.3 to 8.3 compared with G-S-M while improving the average AUC from 0.83 to 0.87. In comparison with traditional feature selection methods using the same number of genes, GeDiRep also achieved competitive AUC values. Biological analyses, including term–gene network, hub gene, and heatmap analyses, supported the functional relevance and stability of several selected genes. Overall, GeDiRep provides a structured and interpretable framework for high-dimensional gene expression analysis by selecting reduced yet discriminative and biologically meaningful gene subsets.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.