Recovering a physics-based generating function with ANFIS and particle swarm optimisation: a synthetic-data proof of concept for arsenic removal by iron-based adsorbents
Sheikh Hayath Mahmud, Anupam Chowdhury
Source abstract
ABSTRACT Arsenic in groundwater is still one of the largest chronic exposures in public health, reaching an estimated 140 million people in some 50 countries. Iron-based adsorbents are among the cheapest treatment options available, and a reliable way of predicting how they behave under changing water chemistry would be genuinely useful for design. This paper reports a deliberately limited step towards that goal. It builds a synthetic dataset of 380 observations for five iron phases, zerovalent iron, goethite (FeOOH), haematite (Fe2O3), magnetite (Fe3O4) and ferrihydrite (Fe(OH)3), not by extracting measurements from the literature but by writing down a multiplicative algebraic generator (Gaussian pH term, Langmuir-type dose term, pseudo-second-order time term, mild Arrhenius temperature term and a concentration-saturation term), parameterised from ranges reported in 16 experimental studies and perturbed with 2.5% Gaussian noise. An Adaptive Neuro-Fuzzy Inference System trained by a hybrid least-squares/Adam scheme and refined by Particle Swarm Optimisation was then asked to learn that generator and was benchmarked against XGBoost, Random Forest, Support Vector Regression and an Multilayer Perceptron (MLP) under identical group-aware splits. On the held-out test set, the model reached R2 = 0.9340 [root mean square error (RMSE) 5.99%] for removal efficiency and R2 = 0.9963 (RMSE 0.66 mg/g) for adsorption capacity; a bagged ensemble raised removal to R2 = 0.9369 but pushed capacity down to 0.8924. These numbers should not be read as predictive skill. They measure how well each architecture recovers a known algebraic function, and the capacity figure is further inflated because q = C0 × (η/100)/D is computed from variables that are themselves model inputs, a form of leakage it quantifies rather than hides. Even with that limitation, two findings still hold up, and these, it says, are the real contribution. First, the Takagi-Sugeno-Kang consequent structure recovers a separable multiplicative function far more accurately than tree- or kernel-based learners (ΔR2 for capacity of +0.40 to +0.60), which is a statement about architecture, not about arsenic. Second, the PSO stage bought almost nothing (removal R2 0.9338 → 0.9340), and group k-fold cross-validation collapsed to removal R2 = 0.699 ± 0.12 and capacity R2 = 0.619 ± 0.37, with onefold at −0.06 a textbook instance of the benchmark inflation described by Kapoor & Narayanan 2023. The framework is offered as a proof of concept and a cautionary benchmark; it is not yet a design tool, and validation against real experimental and field data is a precondition for any practical use.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.