ExplainBench: A Reproducibility-First Explainability Benchmark for Trustworthy On-Device Agentic AI in Energy Systems
Rakesh Kumar Agrawal, Wasim Mohammed Amin Tambe, Nihar Karra
Source record
Source: Crossref
Published: Sep 2, 2026
DOI: 10.20944/preprints202609.0201.v1
Open original source ↗Source abstract
Explainable AI (XAI) evaluation has matured substantially for static classifiers, yet, to the best of our knowledge, no reproducible benchmark has combined explanation-quality evaluation with the on-device, energy-system-specific setting this paper addresses — smart-grid edge controllers, home and building energy management agents, industrial energy IoT, electric-vehicle charging and vehicle-to-grid systems, and renewable microgrid controllers. ExplainBench evaluates explanation quality within agentic decision workflows; it does not evaluate agent capability, autonomy, or intelligence, and provides no runtime assurance or control function. This paper introduces ExplainBench, the fourth module of the BQEB (BIO-Quantum Energy Brain) research program. Consistent with the architecture established across Papers 1–3, ExplainBench introduces no new architectural layer: it registers as the fourth module of the existing Benchmark Layer, alongside ForecastBench, SecBench, and FoundationBench, and records the reproducibility metadata, attribution provenance, and human-rating protocols this evaluation requires within the existing Software and Governance Layer registries, rather than defining parallel mechanisms. The framework specifies nine component metrics — faithfulness, fidelity, stability, consistency, actionability, uncertainty awareness, human interpretability, computational efficiency, and deployment suitability — formally aggregated into six composite indices: the Explainability Quality Index (EQI), Explanation Stability Score (ESS), Counterfactual Consistency Score (CCS), Energy Decision Transparency Score (EDTS), Human Interpretability Rating (HIR), and Explanation Efficiency Index (EEI), assembled into an Explainability Reliability Matrix across a registry of eight forecasting architectures and eight explanation methods. Six algorithms specify benchmark execution, per-instance metric computation, counterfactual scoring, composite aggregation, cross-model statistical comparison, and leaderboard generation, each with explicit complexity bounds. This manuscript defines benchmark infrastructure; it reports no executed evaluations, and every planned-output table in this manuscript is explicitly marked reserved for future experimental validation, consistent with the evidentiary discipline established in Papers 1–3.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.