Indexed metadata

Evaluation of the F1* Score Across Twelve Text Datasets with Prevalence Sensitivity Analysis

Yang Kyu Lim, Hyeon Gyu Kim

Source record

Source: Crossref

Published: Sep 4, 2026

DOI: 10.3390/math14173194

Open original source ↗

Source abstract

The F1 score depends on the positive-class prevalence (π), which complicates comparisons across datasets. Our previous work proposed F1* as an equal-weight composite of F1 and accuracy and derived auxiliary estimation functions. The present study evaluates that previously defined index using 12 text datasets and four models under RepeatedKFold with five folds and five repeats (25 held-out cross-validation evaluations per dataset–model pair). All 1200 runs used identical persisted splits across models, fold-local recurrent preprocessing, and raw-prediction audit. Mean F1 values across datasets were 0.9571 for BERT, 0.9263 for LSTM, 0.9244 for GRU, and 0.9235 for RNN. Mean absolute differences for the auxiliary reconstructions were 0.0012 for F1 and 0.0012 for accuracy. The evaluation also includes MCC, a controlled-prevalence analysis, and an evaluation-sample-size sensitivity analysis using fixed out-of-fold (OOF) predictions from all four models. The evidence remains bounded to binary text datasets and does not establish universal metric superiority.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.