A Deep Learning NLP Framework with Directional Augmentation for Cybersecurity Compliance Monitoring

Authors

DOI:

https://doi.org/10.29407/intensif.v10i2.28533

Keywords:

BERT, Cybersecurity Compliance, Deep Learning, Directional Augmentation, Indonesia Regulatory Framework, Natural Language Processing

Abstract

Background: The increasing complexity of cybersecurity mandates, such as BSSN and Kominfo regulations, presents a significant challenge for automated compliance auditing in Indonesia. Traditional NLP models often struggle with the semantic gap between formal regulatory language and raw technical system telemetry, leading to high false-negative rates in security monitoring. Objective: The purpose of this research to develop a robust deep learning framework to automate cybersecurity compliance assessments while addressing the linguistic challenges of the Indonesian regulatory landscape. The primary goal is to enhance the detection of non-compliant system behaviors by bridging the gap between documentation and real-time logs. Methods: The proposed framework utilized a Transformer-based BERT architecture integrated with a novel Directional Augmentation (DA) mechanism. The methodology follows a four-phase process: (1) data collection of 8,240 labeled points, (2) an Anti-Leak Grouping Strategy to prevent data memorization, (3) implementation of DA through Semantic Polarization and Technical Jargon Injection, and (4) model training and evaluation.  Result: The findings of this research are indicate that the proposed framework significantly outperformed the baseline BERT model. Test Accuracy rose from 90.15% to 95.72%, while Validation Accuracy improved from 91.20% to 96.88%. The final model achieved an Overall Accuracy of 94.39%, maintaining a balanced F1-score and effectively reducing False Negatives to only 32 cases in the detection of security violations, Conlussion : Integrating Directional Augmentation into BERT optimizes Indonesian cybersecurity auditing by synchronizing BSSN/Kominfo regulatory language with technical telemetry through semantic polarization and jargon injection.

Downloads

Download data is not yet available.
Abstract views: 19 , PDF downloads: 8

References

[1] J. Chen, D. Tam, and C. Raffel, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” vol. 11, pp. 191–211, 2023. doi : 10.1162/tacl_a_00543

[2] M. Misiek and T. Hyla, “XGBoost-Based URL Phishing Detection Method With Cross-Dataset Validation,” IEEE Access, vol. 14, no. March, pp. 43200–43212, 2026, doi: 10.1109/ACCESS.2026.3672690.

[3] C. D. Manning, “P Re - Training T Ext E Ncoders As D Iscriminators R Ather T Han G Enerators,” pp. 1–18, 2020.doi : 10.48550/arXiv.2003.10555

[4] A. R. & A. R. Romil Rawat, Hitesh Rawat and We, “Scientific Reports Article in Press An entropy-guided hybrid framework for real-time phishing detection in digital communication systems IN IN,” Scientific Reports, vol. 1, no. 1, pp. 1–28, 2026.doi : 10.1038/s41598-026-08142-w

[5] Y. Yang, J. Wang, C. Bhagavatula, Y. Choi, and D. Downey, “Generative Data Augmentation for Commonsense Reasoning,” pp. 1008–1025, 2020. doi : 10.18653/v1/2020.emnlp-main.811

[6] V. Baladari, “Adaptive Cybersecurity Strategies : Mitigating Cyber Threats and Protecting Data Privacy,” vol. 7, pp. 279–288, 2020. doi : 10.35940/ijitee.G5873.059720

[7] F. Anis et al., “Analisis keamanan sistem informasi perguruan tinggi berbasis indeks kami,” pp. 437–444, 2021.doi : 10.30865/mib.v5i2.2831

[8] A. Halbouni and G. S. Member, “Machine Learning and Deep Learning Approaches for CyberSecurity : A Review,” IEEE Access, vol. 10, no. Ml, pp. 19572–19585, 2022, doi: 10.1109/ACCESS.2022.3151248.

[9] Z. Yang, Z. Dai, Y. Yang, and J. Carbonell, “XLNet : Generalized Autoregressive Pretraining for Language Understanding,” no. NeurIPS, pp. 1–11, 2019. doi : 10.48550/arXiv.1906.08237

[10] L. Elluri, S. A. I. Sree, L. Chukkapalli, K. P. Joshi, T. I. M. Finin, and A. Joshi, “A BERT Based Approach to Measure Web Services Policies Compliance With GDPR,” vol. 9, 2021. doi : 10.1109/ACCESS.2021.3090022

[11] M. L. Dam, D. U. Y. H. Ha, T. T. Huynh, M. Voznak, and C. H. Tran, “Adaptive Cloud – Edge Coordination for Real-Time Phishing URL Detection With Distributed Caching and ONNX-Based Inference,” IEEE Access, vol. 14, no. April, pp. 67492–67517, 2026. doi : 10.1109/ACCESS.2026.3681492

[12] T. He, C. Li, J. Wang, M. Wang, Z. Wang, and C. Jiao, “An emotion analysis in learning environment based on theme-specified drawing by convolutional neural network”.

[13] A. Odetunde, B. I. Adekunle, and J. C. Ogeawuchi, “Using Predictive Analytics and Automation Tools for Real-Time Regulatory Reporting and Compliance Monitoring,” pp. 650–661, 2022.doi : 10.5121/csit.2022.121350

[14] Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” no. 1, 2019.

[15] N. Reimers and I. Gurevych, “Sentence-BERT : Sentence Embeddings using Siamese BERT-Networks,” pp. 3982–3992, 2019. doi : 10.18653/v1/D19-1410

[16] M. Bayer, “A Survey on Data Augmentation for Text Classification,” vol. 55, no. 7, 2026, doi: 10.1145/3544558.

[17] Y. Li, X. Li, Y. Yang, and R. Dong, “A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation,” pp. 1–12, 2020. doi : 10.1145/3340531.3411981

[18] J. Thapa and G. Chahal, “Phishing Detection in the Gen-AI Era : Quantized LLMs vs Classical Models,” Computer & Security, vol. 3, no. 1, pp. 1–8, 2025. doi : 10.1016/j.cose.2024.104125

[19] M. Dandotiya, N. K. Goyal, A. Khunteta, and B. Tiwari, “Real time identification of phishing attacks through machine learning enhanced browser extensions IN IN,” Scientific Reports, vol. 1, no. 1, pp. 0–33, 2026. doi : 10.1038/s41598-026-09211-4

[20] E. K. Dawson, S. L. Cartwright, and J. R. Whitfield, “Multi-Modal Phishing Website Detection with Real-Time Rendering Signals and TLS Fingerprints,” Research Square, vol. 2, no. 1, pp. 1–9, 2026. doi : 10.21203/rs.3.rs-3982145/v1

[21] Y. Wang, “Intelligent Compliance Risk Detection in the Pharmaceutical Industry via Transformer-Driven Semantic Discrimination,” no. 07, 2024.doi : 10.1016/j.asoc.2024.111812

[22] P. Kumar, C. Kushwaha, and A. K. Rajpoot, “Advances in Phishing Detection : A Comprehensive Review of Machine Learning , Deep Learning , and Transformer-based Approaches,” Deep Learning, vol. 4, no. 8, pp. 1–8, 2025.

[23] A. A. Alani and A. Al-azzawia, “PhishFusionNet : A Wide and Deep Phishing Detection with a Hybrid Learning Approach,” CYBERNETICS AND INFORMATION TECHNOLOGIES, vol. 26, no. 1, pp. 140–158, 2026, doi: 10.2478/cait-2026-0008.

[24] D. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, “Albert: A Lite Bert For Self-Supervised Learning Of Language Representations,” conference paper at ICLR 2020 ALBERT:, pp. 1–17, 2020.doi : 10.48550/arXiv.1909.11942

[25] Y. Chai, H. Xie, and J. S. Qin, “Text data augmentation for large language models : a comprehensive survey of methods , challenges , and,” 2026.

[26] C. Lee, B. Kim, and H. Kim, “Computers & Security The silence of the phishers : Early-stage voice phishing detection with runtime permission requests,” Computers & Security, vol. 152, no. February, p. 104364, 2025, doi: 10.1016/j.cose.2025.104364.

[27] C. Shorten, T. M. Khoshgoftaar, and B. Furht, Text Data Augmentation for Deep Learning. Springer International Publishing, 2021. doi: 10.1186/s40537-021-00492-0.

[28] D. Pengcheng He, Xiaodong Liu, “Deberta: Decoding-Enhanced Bert With Dis- Entangled Attention,” conference paper at ICLR, 2021.doi : 10.48550/arXiv.2006.03654

[29] J. Wei and K. Zou, “EDA : Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks,” pp. 6382–6388, 2019. doi : 10.18653/v1/D19-1670

[30] A. Fathurohman and R. W. Witjaksono, “Analysis and Design of Information Security Management System Based on ISO 27001 : 2013 Using ANNEX Control ( Case Study : District of Government of Bandung City ),” vol. 1, no. 1, 2020, doi: 10.25008/bcsee.v1i1.2.

[31] C. S. Governance, I. C. Diplomacy, and E. Agency, “Tata Kelola Keamanan Siber dan Diplomasi Siber Indonesia di Bawah Kelembagaan Badan Siber dan Sandi Negara,” vol. 10, no. 2, pp. 113–128, 2019. doi : 10.22146/globalsouth.55011

[32] M. D. Deepa and A. Tamilarasi, “Bidirectional Encoder Representations from Transformers ( BERT ) Language Model for Sentiment Analysis task : Review,” vol. 12, no. 7, pp. 1708–1721, 2021. doi : 10.17762/de.v2021i7.7126

Downloads

PlumX Metrics

Published

2026-08-22

Issue

Section

Article

How to Cite

[1]
T. G. Laksana, P. D. Atika, A. R. Mahbub, A. R. Iskandar, and W. N. W. Ahmad, “A Deep Learning NLP Framework with Directional Augmentation for Cybersecurity Compliance Monitoring”, INTENSIF: J. Ilm. Penelit. dan Penerap. Tek. Sist. Inf., vol. 10, no. 2, pp. 236–249, Aug. 2026, doi: 10.29407/intensif.v10i2.28533.