A Deep Learning NLP Framework with Directional Augmentation for Cybersecurity Compliance Monitoring
DOI:
https://doi.org/10.29407/intensif.v10i2.28533Keywords:
BERT, Cybersecurity Compliance, Deep Learning, Directional Augmentation, Indonesia Regulatory Framework, Natural Language ProcessingAbstract
Background: The increasing complexity of cybersecurity mandates, such as BSSN and Kominfo regulations, presents a significant challenge for automated compliance auditing in Indonesia. Traditional NLP models often struggle with the semantic gap between formal regulatory language and raw technical system telemetry, leading to high false-negative rates in security monitoring. Objective: The purpose of this research to develop a robust deep learning framework to automate cybersecurity compliance assessments while addressing the linguistic challenges of the Indonesian regulatory landscape. The primary goal is to enhance the detection of non-compliant system behaviors by bridging the gap between documentation and real-time logs. Methods: The proposed framework utilized a Transformer-based BERT architecture integrated with a novel Directional Augmentation (DA) mechanism. The methodology follows a four-phase process: (1) data collection of 8,240 labeled points, (2) an Anti-Leak Grouping Strategy to prevent data memorization, (3) implementation of DA through Semantic Polarization and Technical Jargon Injection, and (4) model training and evaluation. Result: The findings of this research are indicate that the proposed framework significantly outperformed the baseline BERT model. Test Accuracy rose from 90.15% to 95.72%, while Validation Accuracy improved from 91.20% to 96.88%. The final model achieved an Overall Accuracy of 94.39%, maintaining a balanced F1-score and effectively reducing False Negatives to only 32 cases in the detection of security violations, Conlussion : Integrating Directional Augmentation into BERT optimizes Indonesian cybersecurity auditing by synchronizing BSSN/Kominfo regulatory language with technical telemetry through semantic polarization and jargon injection.
Downloads
References
[1] J. Chen, D. Tam, and C. Raffel, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” vol. 11, pp. 191–211, 2023. doi : 10.1162/tacl_a_00543
[2] M. Misiek and T. Hyla, “XGBoost-Based URL Phishing Detection Method With Cross-Dataset Validation,” IEEE Access, vol. 14, no. March, pp. 43200–43212, 2026, doi: 10.1109/ACCESS.2026.3672690.
[3] C. D. Manning, “P Re - Training T Ext E Ncoders As D Iscriminators R Ather T Han G Enerators,” pp. 1–18, 2020.doi : 10.48550/arXiv.2003.10555
[4] A. R. & A. R. Romil Rawat, Hitesh Rawat and We, “Scientific Reports Article in Press An entropy-guided hybrid framework for real-time phishing detection in digital communication systems IN IN,” Scientific Reports, vol. 1, no. 1, pp. 1–28, 2026.doi : 10.1038/s41598-026-08142-w
[5] Y. Yang, J. Wang, C. Bhagavatula, Y. Choi, and D. Downey, “Generative Data Augmentation for Commonsense Reasoning,” pp. 1008–1025, 2020. doi : 10.18653/v1/2020.emnlp-main.811
[6] V. Baladari, “Adaptive Cybersecurity Strategies : Mitigating Cyber Threats and Protecting Data Privacy,” vol. 7, pp. 279–288, 2020. doi : 10.35940/ijitee.G5873.059720
[7] F. Anis et al., “Analisis keamanan sistem informasi perguruan tinggi berbasis indeks kami,” pp. 437–444, 2021.doi : 10.30865/mib.v5i2.2831
[8] A. Halbouni and G. S. Member, “Machine Learning and Deep Learning Approaches for CyberSecurity : A Review,” IEEE Access, vol. 10, no. Ml, pp. 19572–19585, 2022, doi: 10.1109/ACCESS.2022.3151248.
[9] Z. Yang, Z. Dai, Y. Yang, and J. Carbonell, “XLNet : Generalized Autoregressive Pretraining for Language Understanding,” no. NeurIPS, pp. 1–11, 2019. doi : 10.48550/arXiv.1906.08237
[10] L. Elluri, S. A. I. Sree, L. Chukkapalli, K. P. Joshi, T. I. M. Finin, and A. Joshi, “A BERT Based Approach to Measure Web Services Policies Compliance With GDPR,” vol. 9, 2021. doi : 10.1109/ACCESS.2021.3090022
[11] M. L. Dam, D. U. Y. H. Ha, T. T. Huynh, M. Voznak, and C. H. Tran, “Adaptive Cloud – Edge Coordination for Real-Time Phishing URL Detection With Distributed Caching and ONNX-Based Inference,” IEEE Access, vol. 14, no. April, pp. 67492–67517, 2026. doi : 10.1109/ACCESS.2026.3681492
[12] T. He, C. Li, J. Wang, M. Wang, Z. Wang, and C. Jiao, “An emotion analysis in learning environment based on theme-specified drawing by convolutional neural network”.
[13] A. Odetunde, B. I. Adekunle, and J. C. Ogeawuchi, “Using Predictive Analytics and Automation Tools for Real-Time Regulatory Reporting and Compliance Monitoring,” pp. 650–661, 2022.doi : 10.5121/csit.2022.121350
[14] Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” no. 1, 2019.
[15] N. Reimers and I. Gurevych, “Sentence-BERT : Sentence Embeddings using Siamese BERT-Networks,” pp. 3982–3992, 2019. doi : 10.18653/v1/D19-1410
[16] M. Bayer, “A Survey on Data Augmentation for Text Classification,” vol. 55, no. 7, 2026, doi: 10.1145/3544558.
[17] Y. Li, X. Li, Y. Yang, and R. Dong, “A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation,” pp. 1–12, 2020. doi : 10.1145/3340531.3411981
[18] J. Thapa and G. Chahal, “Phishing Detection in the Gen-AI Era : Quantized LLMs vs Classical Models,” Computer & Security, vol. 3, no. 1, pp. 1–8, 2025. doi : 10.1016/j.cose.2024.104125
[19] M. Dandotiya, N. K. Goyal, A. Khunteta, and B. Tiwari, “Real time identification of phishing attacks through machine learning enhanced browser extensions IN IN,” Scientific Reports, vol. 1, no. 1, pp. 0–33, 2026. doi : 10.1038/s41598-026-09211-4
[20] E. K. Dawson, S. L. Cartwright, and J. R. Whitfield, “Multi-Modal Phishing Website Detection with Real-Time Rendering Signals and TLS Fingerprints,” Research Square, vol. 2, no. 1, pp. 1–9, 2026. doi : 10.21203/rs.3.rs-3982145/v1
[21] Y. Wang, “Intelligent Compliance Risk Detection in the Pharmaceutical Industry via Transformer-Driven Semantic Discrimination,” no. 07, 2024.doi : 10.1016/j.asoc.2024.111812
[22] P. Kumar, C. Kushwaha, and A. K. Rajpoot, “Advances in Phishing Detection : A Comprehensive Review of Machine Learning , Deep Learning , and Transformer-based Approaches,” Deep Learning, vol. 4, no. 8, pp. 1–8, 2025.
[23] A. A. Alani and A. Al-azzawia, “PhishFusionNet : A Wide and Deep Phishing Detection with a Hybrid Learning Approach,” CYBERNETICS AND INFORMATION TECHNOLOGIES, vol. 26, no. 1, pp. 140–158, 2026, doi: 10.2478/cait-2026-0008.
[24] D. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, “Albert: A Lite Bert For Self-Supervised Learning Of Language Representations,” conference paper at ICLR 2020 ALBERT:, pp. 1–17, 2020.doi : 10.48550/arXiv.1909.11942
[25] Y. Chai, H. Xie, and J. S. Qin, “Text data augmentation for large language models : a comprehensive survey of methods , challenges , and,” 2026.
[26] C. Lee, B. Kim, and H. Kim, “Computers & Security The silence of the phishers : Early-stage voice phishing detection with runtime permission requests,” Computers & Security, vol. 152, no. February, p. 104364, 2025, doi: 10.1016/j.cose.2025.104364.
[27] C. Shorten, T. M. Khoshgoftaar, and B. Furht, Text Data Augmentation for Deep Learning. Springer International Publishing, 2021. doi: 10.1186/s40537-021-00492-0.
[28] D. Pengcheng He, Xiaodong Liu, “Deberta: Decoding-Enhanced Bert With Dis- Entangled Attention,” conference paper at ICLR, 2021.doi : 10.48550/arXiv.2006.03654
[29] J. Wei and K. Zou, “EDA : Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks,” pp. 6382–6388, 2019. doi : 10.18653/v1/D19-1670
[30] A. Fathurohman and R. W. Witjaksono, “Analysis and Design of Information Security Management System Based on ISO 27001 : 2013 Using ANNEX Control ( Case Study : District of Government of Bandung City ),” vol. 1, no. 1, 2020, doi: 10.25008/bcsee.v1i1.2.
[31] C. S. Governance, I. C. Diplomacy, and E. Agency, “Tata Kelola Keamanan Siber dan Diplomasi Siber Indonesia di Bawah Kelembagaan Badan Siber dan Sandi Negara,” vol. 10, no. 2, pp. 113–128, 2019. doi : 10.22146/globalsouth.55011
[32] M. D. Deepa and A. Tamilarasi, “Bidirectional Encoder Representations from Transformers ( BERT ) Language Model for Sentiment Analysis task : Review,” vol. 12, no. 7, pp. 1708–1721, 2021. doi : 10.17762/de.v2021i7.7126
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Tri Ginanjar Laksana, Prima Dina Atika, Asep Ramdhani Mahbub, Ade Rahmat Iskandar, Wan Nooraishy Wan Ahmad

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Copyright on any article is retained by the author(s).
- The author grants the journal, the right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work’s authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal’s published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
- The article and any associated published material is distributed under the Creative Commons Attribution-ShareAlike 4.0 International License


