LARGE LANGUAGE MODELS IN PENETRATION TESTING AUTOMATION: CURRENT STATE, LIMITATIONS AND PROSPECTS FOR APPLICATION IN PROFESSIONAL TRAINING OF INFORMATION SECURITY SPECIALISTS
Abstract
The article analyzes the current state of research on the application of large language models (LLMs) for penetration testing automation. The reasons for the rapid growth of interest in this area are examined, driven by the increasing complexity of IT infrastructures, the high cost of traditional security audit methods, and the development of the DevSecOps concept. Based on the systematization of review and applied research from 2023–2025, four key areas of LLM application in cybersecurity are identified: decision support and strategic attack planning, automation of penetration testing stages, security of language models themselves (AI Red Teaming), and intelligent source code analysis. A critical analysis is conducted of multi-agent system architectures using Retrieval-Augmented Generation (RAG) approaches, testing and evaluation systems, and commercial solutions. It is established that even the most advanced LLM-based penetration testing systems operate at intermediate autonomy levels (3–4 out of 5), unable to independently formulate goals and evaluate the consequences of their decisions. The thesis is substantiated that as of the end of 2025, LLMs serve primarily as an analyst tool rather than an autonomous penetration tester. The pedagogical aspects of integrating LLM tools into professional training of information security specialists are examined, and the necessity of developing critical thinking among students regarding the capabilities and limitations of these technologies is justified.
Online viewer
References
- Novikov, A. M. (1997). Professional education in Russia: Development prospects. ICP NPO RAO. (In Russian)
- Leont'ev, A. N. (1975). Activity. Consciousness. Personality. Politizdat. (In Russian)
- Lerner, I. Ya. (1974). Problem-based learning. Znanie. (In Russian)
- Tikhomirov, V. P., & Tikhomirova, N. V. (2019). Digital transformation of education: Challenges and opportunities. Open Education, (4), 4–14. (In Russian)
- Alaryani, M., et al. (2025). Towards supporting penetration testing education with large language models: An evaluation and comparison. arXiv preprint arXiv:2501.17539.
- Happe, A., & Cito, J. (2023). Getting pwn'd by AI: Penetration testing with large language models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 2082–2086).
- Deng, G., et al. (2023). PentestGPT: An LLM-empowered automatic penetration testing tool. arXiv preprint arXiv:2308.06782.
- Xu, Z., et al. (2024). AutoAttacker: A large language model guided system to implement automatic cyber-attacks. arXiv preprint arXiv:2403.01038.
- Fang, R., et al. (2024). LLM agents can autonomously exploit one-day vulnerabilities. arXiv preprint arXiv:2404.08144.
- Shao, R., et al. (2023). Empirical analysis of large language models in automated vulnerability discovery. arXiv preprint arXiv:2312.02337.
- Bhatt, M., et al. (2023). Purple Llama CyberSecEval: A secure coding benchmark for language models. arXiv preprint arXiv:2312.04724.
- Wan, Y., et al. (2024). Cybersecurity issues and challenges: In brief. IEEE Access, 12, 1–15.
- Moskal, S., et al. (2023). LLM in the shell: Generative AI-powered pentesting. arXiv preprint arXiv:2312.09552.
- Charan, P. V. S., et al. (2023). From text to MITRE techniques: Examining the layered approach to explainability with a large language model (LLM). arXiv preprint arXiv:2308.09197.
- Tihanyi, N., et al. (2024). CyberMetric: A benchmark dataset for evaluating large language models knowledge in cybersecurity. arXiv preprint arXiv:2402.07688.
- Peng, B., et al. (2024). PenHeal: A two-stage LLM framework for automated pentesting and optimal remediation. arXiv preprint arXiv:2407.17788.
- Wan, Y., et al. (2023). HackMentor: Fine-tuning large language models for cybersecurity. arXiv preprint arXiv:2312.01234.
- Yang, J., et al. (2023). SWE-bench: Can language models resolve real-world GitHub issues? arXiv preprint arXiv:2310.06770.
- Kraevskiy, V. V. (2006). Methodology of pedagogy: A new stage. Akademiya. (In Russian)
- Batyshev, S. Ya. (2010). Professional pedagogy (S. Ya. Batyshev & A. M. Novikov, Eds.; 3rd ed.). EGVES. (In Russian)
- Verbitskiy, A. A. (1991). Active learning in higher education: A contextual approach. Vysshaya Shkola. (In Russian)
- Slastenin, V. A. (2002). Pedagogy. Akademiya. (In Russian)
- Zeer, E. F. (2009). Psychology of professional education. Akademiya. (In Russian)
- Alias Robotics. (2025). Cybersecurity AI (CAI): An open, bug bounty-ready agentic framework. arXiv preprint arXiv:2504.06017.
License
Copyright (c) 2025 V. S. Grekov (Author)