Evaluation of the F1* Score Across Twelve Text Datasets with Prevalence Sensitivity Analysis
Yang Kyu Lim, Hyeon Gyu Kim
Source abstract
The F1 score depends on the positive-class prevalence (π), which complicates comparisons across datasets. Our previous work proposed F1* as an equal-weight composite of F1 and accuracy and derived auxiliary estimation functions. The present study evaluates that previously defined index using 12 text datasets and four models under RepeatedKFold with five folds and five repeats (25 held-out cross-validation evaluations per dataset–model pair). All 1200 runs used identical persisted splits across models, fold-local recurrent preprocessing, and raw-prediction audit. Mean F1 values across datasets were 0.9571 for BERT, 0.9263 for LSTM, 0.9244 for GRU, and 0.9235 for RNN. Mean absolute differences for the auxiliary reconstructions were 0.0012 for F1 and 0.0012 for accuracy. The evaluation also includes MCC, a controlled-prevalence analysis, and an evaluation-sample-size sensitivity analysis using fixed out-of-fold (OOF) predictions from all four models. The evidence remains bounded to binary text datasets and does not establish universal metric superiority.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.