Explainable XGBoost and LSTM Models for Cyberattack Detection in Resource-Constrained IoT Devices
Keywords:
Cyberattack detection, explainable artificial intelligence, Internet of Things, LSTM, resource-constrained devices, XGBoostAbstract
Internet of Things (IoT) devices are frequently deployed with limited processing capacity, restricted memory, and limited security visibility, which increases their exposure to botnet compromise. This study evaluates explainable XGBoost and LSTM models for binary cyberattack detection using public N-BaIoT traffic data. The proposed workflow starts with public CSV collection, deterministic source-balanced sampling, cleaning of missing and infinite values, device-grouped validation, model training, metric evaluation, SHAP-based explanation analysis, feature ablation, and local CPU resource measurement. A reproducible subset of 101,000 cleaned rows was constructed from 89 CSV sources across nine commercial IoT devices. The experiment used 5-fold StratifiedGroupKFold with IoT device identity as the grouping variable to reduce train-test device leakage. XGBoost achieved 99.52 ± 0.71% accuracy, 99.58 ± 0.61% F1-score, and 0.9995 ± 0.0011 ROC-AUC. LSTM achieved 99.46 ± 1.10% accuracy, 99.54 ± 0.93% F1-score, and 0.9999 ± 0.0002 ROC-AUC. Histogram Gradient Boosting achieved the highest internal-baseline F1-score (99.82 ± 0.29%), confirming the strength of tree-based tabular learners on N-BaIoT. XGBoost remains the main explainable model because it provides direct TreeSHAP support, compact training cost, and competitive cross-device performance, whereas LSTM is retained as a sequential comparator for short source-local windows. SHAP stability analysis showed partially stable feature attribution, with a mean top-10 Jaccard similarity of 0.547. Resource results are interpreted as local CPU proxy measurements rather than embedded-hardware deployment validation.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Internetworking Indonesia Journal

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.