Fault Diagnosis of Coal-Fired Power Plants Based on Multi-Scale Spatiotemporal Features and TabPFN
Xilong Ye, Chenglong Miao, Weiwei Jia, Xinyi Huang, Maofa Wang, Jun Tan
Source abstract
The safe and stable operation of coal-fired generating units is of critical strategic importance for ensuring the reliable supply of power systems. However, the fault evolution of industrial thermal systems exhibits the characteristics of strong nonlinearity and a long incubation period, coupled with the extreme scarcity of key fault samples (Few-shot) in actual production, which severely limits the engineering application of traditional data-driven diagnostic methods. Existing deep learning models, which are highly dependent on massive and balanced labeled data, not only struggle to overcome the overfitting bottleneck in scenarios with scarce fault samples, but also frequently introduce severe label noise (Label Noise) by ignoring the physical incubation period of faults, resulting in the degradation of the model’s decision boundary. To address the above challenges, this paper proposes a novel fault diagnosis framework integrating multi-scale spatiotemporal feature engineering and the Tabular Prior-Data Fitted Network (TabPFN). Starting from the physical mechanism of the system, this paper develops a dynamic label cleaning strategy based on multivariate statistical deviation, which accurately defines the fault divergence point to eliminate the noise in the incubation period. The constructed multi-scale spatiotemporal feature engineering integrating first-order difference and sliding window statistics can effectively map the transient mutation and steady-state evolution trend of the system. The introduced pre-trained TabPFN model based on the Transformer architecture, relying on its Bayesian inference capability and in-context learning (In-Context Learning) mechanism, can realize parameter-tuning-free and efficient classification for scarce samples. Experiments based on high-fidelity dynamic simulation data from GE Steam Power show that under the strict setting of limiting the training set to only 2000 samples, the proposed method achieves a comprehensive diagnostic accuracy of up to 99.29% and an F1-score of 0.9929 for seven typical operating conditions. Multi-dimensional comparative experiments and ablation studies confirm that the proposed framework comprehensively outperforms six mainstream baseline models, including XGBoost and SVM, in terms of precision, recall, and anti-interference robustness, and also delivers outstanding performance when benchmarked against deep learning models. This provides a brand-new theoretical perspective and technical paradigm for equipment health management in the context of industrial big data.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.