Indexed metadata

A Leakage-Aware Evaluation of a Multi-Scale Attention Temporal Convolutional Network for Binary Network Intrusion Detection

Yu Yang, Jinliang Yuan, Minna Gao, Mingmei Chen, Dan Gong, Le Gao

Source record

Source: Crossref

Published: Sep 16, 2026

DOI: 10.3390/math14183357

Open original source ↗

Source abstract

Reliable evaluation is as important as model design in benchmark-based network intrusion detection. This study evaluates a Multi-Scale Attention Temporal Convolutional Network (MS-ATCN) for binary intrusion detection under a leakage-aware and reproducibility-oriented protocol on CIC-IDS2017, NSL-KDD, and UNSW-NB15. MS-ATCN combines multi-scale temporal convolution, channel and temporal attention, and class-weighted focal loss over windows of flow-level records. Because these components are established techniques, the study focuses on their integrated empirical behavior rather than proposing a new learning paradigm. The evaluation includes repeated-seed experiments, seed-aligned Wilcoxon signed-rank tests with Holm correction, component ablations, an exploratory single-run record-order check, a single-seed CIC-IDS2017 day-level holdout, an UNSW-NB15 identifier-restoration stress test, exploratory one-at-a-time hyperparameter perturbations, and computational-efficiency measurements. The repeated-run means show that MS-ATCN is not consistently the best neural model: MLP has higher mean accuracy and Macro-F1 on CIC-IDS2017, XGBoost and LightGBM have the strongest repeated-run results on NSL-KDD and UNSW-NB15, and CNN has higher mean accuracy and Macro-F1 than MS-ATCN on UNSW-NB15. No seed-aligned main-model or ablation comparison reached the Holm-adjusted 0.05 threshold. With five paired seeds, however, the exact two-sided Wilcoxon test cannot attain an unadjusted p-value below 0.0625, so adjusted nonsignificance does not establish equivalence. The exploratory diagnostics indicate sensitivity to split policy and record ordering but do not establish a general benefit from local temporal context. The findings support cautious multi-metric interpretation, strong tabular baselines, repeated-seed reporting, and explicit separation of model-only efficiency from end-to-end deployment performance.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.