CUES: A Multiplicative Composite Metric for Evaluating Clinical Prediction Models Theory, Inference, and Properties
Ali Mohammad Alqudah, Zahra Moussavi
Source abstract
Evaluating artificial intelligence (AI) models in clinical medicine requires more than conventional metrics such as accuracy, Area Under the Receiver Operating Characteristic (AUROC), or F1-score, which often overlook key considerations such as fairness, reliability, and real-world utility. We introduce CUES as a multiplicative composite score for clinical prediction models; it is defined as CUES=(C⋅U⋅E⋅S)1/4, where C represents calibration, U integrated clinical utility, E equity across patient subpopulations, and S sampling stability. We formally establish boundedness, monotonicity, and differentiability on the domain (0,1]4, derive first-order sensitivity relations, and provide asymptotic approximations for its sampling distribution via the delta method. To facilitate inference, we propose bootstrap procedures for constructing confidence intervals and for comparative model evaluation. Analytic examples illustrate how CUES can diverge from traditional metrics, capturing dimensions of predictive performance that are essential for clinical reliability but often missed by AUROC or F1-score alone. By integrating multiple facets of clinical utility and robustness, CUES provides a comprehensive tool for model evaluation, comparison, and selection in real-world medical applications.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.