Improving Relative Humidity Forecasting Accuracy Using a Hybrid CEEMDAN‐Based Deep Learning and Machine Learning Approach: A Comparative Analysis
John Kamwele Mutinda, Tecla Mutave Kyalo, Jackson Ndoto Munyao, Amos Kipkorir Langat
Source abstract
Forecasting relative humidity remains challenging due to its nonlinear, nonstationary characteristics and long‐memory dependence. This study proposes a hybrid decomposition‐ensemble model, Complete Ensemble Empirical Mode Decomposition with Adaptive Noise–Sample Entropy–Gated Recurrent Unit–Ridge Regression (CEEMDAN‐SE‐GRU‐Ridge), for one‐step‐ahead humidity forecasting using daily data from Nairobi, Kenya (2017–2024). The methodology systematically addresses multiscale dynamics: A rolling window CEEMDAN decomposition is applied to the raw series to extract intrinsic mode functions (IMFs), preventing look‐ahead bias by using only past observations at each decomposition step. Sample entropy quantifies the complexity of each IMF, and k‐means clustering regroups them into three frequency components (high, medium, and low frequency) while keeping the residue separate. A GRU network models the three frequency components, capturing temporal dependencies in volatile patterns, and ridge regression forecasts the smooth residue, providing stable regularized predictions. The final forecast is the sum of all component predictions. On an 80/20 train‐test split, the proposed model achieves superior performance (RMSE = 4.1953, MAE = 3.2808, MAPE = 5.1135 % , and R 2 = 0.8286), outperforming 15 benchmark models including standalone machine learning, deep learning architectures, and six hybrid approaches. Diebold–Mariano tests confirm statistical significance for all comparisons, and Model Confidence Set analysis retains the proposed model as the only member of the Superior Set of Models. Ablation studies verify the positive contribution of each component: CEEMDAN decomposition, sample‐entropy clustering, and residue‐specific ridge regression, with the largest improvement observed from the clustering step. Robustness checks with 60/40, 70/30, and 90/10 splits demonstrate consistent performance across varying training data sizes. The results show that rolling window decomposition‐based, component‐specific modeling substantially enhances humidity forecasting accuracy while avoiding data leakage, offering practical value for climate‐sensitive applications in agriculture, water resource management, and early warning systems.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.