A Leakage-Aware Evaluation of a Multi-Scale Attention Temporal Convolutional Network for Binary Network Intrusion Detection
Yu Yang, Jinliang Yuan, Minna Gao, Mingmei Chen, Dan Gong, Le Gao
Source abstract
Reliable evaluation is as important as model design in benchmark-based network intrusion detection. This study evaluates a Multi-Scale Attention Temporal Convolutional Network (MS-ATCN) for binary intrusion detection under a leakage-aware and reproducibility-oriented protocol on CIC-IDS2017, NSL-KDD, and UNSW-NB15. MS-ATCN combines multi-scale temporal convolution, channel and temporal attention, and class-weighted focal loss over windows of flow-level records. Because these components are established techniques, the study focuses on their integrated empirical behavior rather than proposing a new learning paradigm. The evaluation includes repeated-seed experiments, seed-aligned Wilcoxon signed-rank tests with Holm correction, component ablations, an exploratory single-run record-order check, a single-seed CIC-IDS2017 day-level holdout, an UNSW-NB15 identifier-restoration stress test, exploratory one-at-a-time hyperparameter perturbations, and computational-efficiency measurements. The repeated-run means show that MS-ATCN is not consistently the best neural model: MLP has higher mean accuracy and Macro-F1 on CIC-IDS2017, XGBoost and LightGBM have the strongest repeated-run results on NSL-KDD and UNSW-NB15, and CNN has higher mean accuracy and Macro-F1 than MS-ATCN on UNSW-NB15. No seed-aligned main-model or ablation comparison reached the Holm-adjusted 0.05 threshold. With five paired seeds, however, the exact two-sided Wilcoxon test cannot attain an unadjusted p-value below 0.0625, so adjusted nonsignificance does not establish equivalence. The exploratory diagnostics indicate sensitivity to split policy and record ordering but do not establish a general benefit from local temporal context. The findings support cautious multi-metric interpretation, strong tabular baselines, repeated-seed reporting, and explicit separation of model-only efficiency from end-to-end deployment performance.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.