Generative AI-Based Website Security Testing: Case Study of Educational Websites in Indonesia

Authors

DOI:

https://doi.org/10.21512/commit.v20i2.15010

Keywords:

Generative AI, Hybrid Security Testing, Vulnerability, Educational Website, OWASP, STRIDE

Abstract

The integration of Generative Artificial Intelligence (Generative AI) into educational websites has expanded digital service capabilities through chatbots and dynamic content systems. However, it has simultaneously introduced novel security vulnerabilities, such as prompt injection, data leakage, and insecure output handling, that remain insufficiently examined, particularly in the Indonesian context where rapid AI adoption outpaces security readiness. The research aims to develop and validate a hybrid security testing framework capable of detecting and mitigating vulnerabilities specific to Generative AI-based websites. The research adopts a Design Science Research (DSR) approach with an experimental mixed-methods strategy, combining quantitative security testing through penetration testing and adversarial red teaming with qualitative analysis based on threat modeling. The evaluation is conducted on Generative AI-enabled educational websites and simulated prototypes using Open Web Application Security Project (OWASP) Top 10 for Large Language Models (LLM) risk mapping, the Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege (STRIDE) framework, and adversarial benchmarks. The results indicate that dominant vulnerabilities are associated with prompt injection, information leakage, and insecure model output handling. The proposed framework significantly improves detection effectiveness with an accuracy rate of 92.6% while reducing attack success rates following layered remediation. These findings confirm that conventional web security mechanisms are insufficient for LLM-based systems and require adaptive, AI-specific testing approaches. The research contributes by strengthening the integration of AI security frameworks in the form of a new OWASP-STRIDE hybrid model and practically by providing actionable security testing guidance for developers and educational institutions seeking to deploy Generative AI more securely.

Dimensions

Author Biographies

Fransiskus Mario Hartono Tjiptabudi, STIKOM Uyelindo Kupang

Information System Study Program

Ricky Imanuel Ndaumanu, Universitas Widya Dharma Pontianak

Informatics Study Program, Faculty of Information Technology

Donna Setiawati, Universitas Boyolali

Management Study Program, Faculty of Economics and Business

References

[1] D. Russo, “Navigating the complexity of Generative AI adoption in software engineering,” ACM Transactions on Software Engineering and Methodology, vol. 33, no. 5, pp. 1–50, 2024.

[2] F. M. H. Tjiptabudi, “Integrated information and communication media based on Organization Goal-Oriented Requirement Engineering (OGORE),” Jurnal Sistem Informasi, vol. 19, no. 1, pp. 28–42, 2023.

[3] X. Zhang, Z. Zhang, Q. Zhong, X. Zheng, Y. Zhang, S. Hu, and L. Y. Zhang, “Masked language model based textual adversarial example detection,” in Proceedings of the 2023 ACM Asia Conference on Computer and Communications Security. New York, United States: Association for Computing Machinery, July 10–14, 2023, pp. 925–937.

[4] I. Jada and T. O. Mayayise, “The impact of artificial intelligence on organisational cyber security: An outcome of a systematic literature review,” Data and Information Management, vol. 8, no. 2, pp. 1–15, 2024.

[5] E. A. Altulaihan, A. Alismail, and M. Frikha, “A survey on web application penetration testing,” Electronics, vol. 12, no. 5, pp. 1–23, 2023.

[6] M. A. Ferrag, F. Alwahedi, A. Battah, B. Cherif, A. Mechri, N. Tihanyi, T. Bisztray, and M. Debbah, “Generative AI in cybersecurity: A comprehensive review of LLM applications and vulnerabilities,” Internet of Things and Cyber-Physical Systems, vol. 5, pp. 1–46, 2025.

[7] D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath et al., “Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,” 2022. [Online]. Available: https://arxiv.org/abs/2209.07858

[8] F. P. Utama and R. M. H. Nurhadi, “Uncovering the risk of academic information system vulnerability through PTES and OWASP method,” CommIT (Communication and Information Technology) Journal, vol. 18, no. 1, pp. 39–51, 2024.

[9] M. Podpora, M. Baranowski, M. Chopcian, L. Kwasniewicz, and W. Radziewicz, “LLM firewall using validator agent for prevention against prompt injection attacks,” Applied Sciences, vol. 16, no. 1, pp. 1–24, 2025.

[10] L. Mauri and E. Damiani, “Modeling threats to AI-ML systems using STRIDE,” Sensors, vol. 22, no. 17, pp. 1–21, 2022.

[11] M. A. Angganegara, I. Y. Mukti, and M. Fathinnuddin, “Integration of STRIDE and MITRE ATT&CK frameworks for enhanced cyber threat modeling: A case study of digital merchant banking application,” in 2025 International Conference on Advancement in Data Science, Elearning and Information System (ICADEIS). Bandung, Indonesia: IEEE, Feb. 3–4, 2025, pp. 1–6.

[12] P. M. J. Delport, R. Von Solms, and M. Gerber, “Methodological guidelines for design science research,” Procedia Computer Science, vol. 237, pp. 195–203, 2024.

[13] J. da Assunc¸ ˜ao Moutinho, G. Fernandes, and R. Rabechini Jr, “Evaluation in design science: A framework to support project studies in the context of university research centres,” Evaluation and Program Planning, vol. 102, 2024.

[14] S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022, pp. 3214–3252.

[15] M. S. Jabbar, S. Al-Azani, A. Alotaibi, and M. Ahmed, “Red teaming large language models: A comprehensive review and critical analysis,” Information Processing & Management, vol. 62, no. 6, 2025.

[16] X. Liu and B. Zhong, “Integrating generative artificial intelligence into student learning: A systematic review from a TPACK perspective,” Educational Research Review, vol. 49, pp. 1–34, 2025.

[17] M. M. A. Parambil, J. Rustamov, S. G. Ahmed, Z. Rustamov, A. I. Awad, N. Zaki, and F. Alnajjar, “Integrating AI-based and conventional cybersecurity measures into online higher education settings: Challenges, opportunities, and prospects,” Computers and Education: Artificial Intelligence, vol. 7, pp. 1–30, 2024.

[18] S. Shrestha, C. Banda, A. K. Mishra, F. Djebbar, and D. Puthal, “Investigation of cybersecurity bottlenecks of AI agents in industrial automation,” Computers, vol. 14, no. 11, pp. 1–43, 2025.

[19] M. Uddin, M. S. Irshad, I. A. Kandhro, F. Alanazi, F. Ahmed, M. Maaz et al., “Generative AI revolution in cybersecurity: A comprehensive review of threat intelligence and operations,” Artificial Intelligence Review, vol. 58, pp. 1–39, 2025.

[20] Z. Azam, M. M. Islam, and M. N. Huda, “Comparative analysis of intrusion detection systems and machine learning-based model analysis through decision tree,” IEEE Access, vol. 11, pp. 80 348–80 391, 2023.

[21] S. Potla, “Threat modeling in application security: A practical approach,” European Journal of Computer Science and Information Technology, vol. 13, pp. 10–19, 4 2025.

[22] S. F. Wen, A. Shukla, and B. Katt, “Artificial intelligence for system security assurance: A systematic literature review,” International Journal of Information Security, vol. 24, pp. 1–42, 2025.

[23] S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “RealToxicityPrompts: Evaluating neural toxic degeneration in language models,” in Findings of the Association for Computational Linguistics: EMNLP 2020. Online: Association for Computational Linguistics, 2020, pp. 3356–3369.

[24] A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023. [Online]. Available: https://arxiv.org/abs/2307.15043

[25] K. F. Hasan, H. H. Shajeeb, C. Abeydeera, B. Turnbull, and M. Warren, “ISADM: An integrated STRIDE, ATT&CK, and D3FEND model for threat modeling against real-world adversaries,” IEEE Access, vol. 13, pp. 217 316–217 348, 2025.

[26] D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto, “Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks,” in 2024 IEEE security and privacy workshops (SPW). San Francisco, CA, USA: IEEE, May 23, 2024, pp. 132–143.

[27] AWS, “AI security scoping matrix.” [Online]. Available: https://aws.amazon.com/id/ai/security/generative-ai-scoping-matrix/

[28] S. Gregor and A. R. Hevner, “Positioning and presenting design science research for maximum impact,” MIS Quarterly, vol. 37, no. 2, pp. 337–355, 2013.

[29] T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung et al., “Model evaluation for extreme risks,” 2023. [Online]. Available: https://arxiv.org/abs/2305.15324

[30] A. Shostack, Threat modeling: Designing for security. John Wiley & Sons, 2014.

[31] Z. Liao, K. Chen, Y. Lin, K. Li, Y. Liu, H. Chen et al., “Attack and defense techniques in large language models: A survey and new perspectives,” Neural Networks, vol. 196, 2026.

[32] G. Recupito, F. Pecorelli, G. Catolino, V. Lenarduzzi, D. Taibi, D. Di Nucci, and F. Palomba, “Technical debt in AI-enabled systems: On the prevalence, severity, impact, and management strategies for code and architecture,” Journal of Systems and Software, vol. 216, pp. 1–22, 2024.

Downloads

Published

2026-07-27

How to Cite

[1]
F. M. H. Tjiptabudi, R. I. Ndaumanu, and D. Setiawati, “Generative AI-Based Website Security Testing: Case Study of Educational Websites in Indonesia”, CommIT (Communication and Information Technology) Journal, vol. 20, no. 2, pp. 233–245, Jul. 2026.
Abstract 84  .
PDF downloaded 43  .