Analysis of the Precision of Large Language Models in the Identification of Security Vulnerabilities and Weaknesses in Generated Code

  • Federico Muñoz-Babiano
    UNIR (International University of La Rioja), Av. de la Paz, 137, Logroño, 37008, La Rioja, Spain federico.munoz[at]unir.net
  • Paula Lamo
  • Ricardo S. Alonso
    AIR Institute, HUB de Innovación Tecnológico La Aldehuela, Zamora, 49022, Castilla y León, Spain

Abstract

Large Language Models (LLMs) have accelerated code generation, yet their security implications remain underexplored. This work evaluates the capability of eight general-purpose and code-specialized LLMs to detect weaknesses and vulnerabilities in generated code using CVE and CWE references. Stratified sampling of the CVEfixes dataset across five programming languages is assessed using precision, recall, F1-score, and relevance-oriented metrics (REINP, REINR, and REINF1). Results show limited detection of specific vulnerabilities (CVE), but better performance for general weaknesses (CWE), especially in Python and Ruby. We discuss limitations in model training and prompt conditioning, and outline improvements through dataset diversification, prompt engineering and hybrid human-in-the-loop approaches. The study highlights the current potential and limitations of LLMs for practical vulnerability detection in generated code.
  • Referencias
  • Cómo citar
  • Del mismo autor
  • Métricas
Ahmad, B., Thakur, S., Tan, B., Karri, R., & Pearce, H. (2024). On hardware security bug code fixes by prompting large language models. IEEE Transactions on Information Forensics and Security, 19, 4043-4057. https://doi.org/10.1109/TIFS.2024.3374558

Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., & Sanghai, S. (2023). GQA: Training generalized multi-query transformers from multi-head checkpoints. arXiv. https://doi.org/10.18653/v1/2023.emnlp-main.298

Albanese, M., Adebiyi, O., & Onovae, F. (2024). CVE2CWE: Automated mapping of software vulnerabilities. In Proceedings of the International Conference on Security and Cryptography (pp. 500-507). SCITEPRESS. https://doi.org/10.5220/0012770400003767

Alian, A. A., Sobhy, B., Nasser, M., & Hani, L. (2023). Backslash map: An automated vulnerability scanner. In Proceedings of the International Conference on Intelligent Computing and Information Systems (pp. 476-482). IEEE. https://doi.org/10.1109/ICICIS58388.2023.10391153

Allamanis, M., Jackson-Flux, H., & Brockschmidt, M. (2021). Self-supervised bug detection and repair. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.

Alon, U., Brody, S., Levy, O., & Yahav, E. (2018). code2seq: Sequences from structured code representations. arXiv. https://arxiv.org/abs/1808.01400

Alon, U., Zilberstein, M., Levy, O., & Yahav, E. (2019). code2vec: Learning distributed representations of code. In Proceedings of the Association for Computing Machinery. ACM. https://doi.org/10.1145/3290353

Backslash Security. (2025, abril). Can AI Vibe Coding Be Trusted? https://www.backslash.security/blog/can-ai-vibe-coding-be-trusted

Bhandari, G., Naseer, A., & Moonen, L. (2021). CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software. In Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (pp. 38-39). ACM. https://doi.org/10.1145/3475960.3475985

Bo, Y., Zhang, N., Li, S. P., & Xia, X. (2020). Survey of intelligent code completion. Journal of Software, 31(5), 1435-1453.

Cassano, F., Gouwar, J., Nguyen, D. P., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., Guha, A., Greenberg, M., & Jangda, A. (2023). Multipl-E: A scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7), 3675-3691. https://doi.org/10.1109/TSE.2023.3267446

Chakraborty, S., Krishna, R., Ding, Y., & Ray, B. (2021). Deep learning-based vulnerability detection. IEEE Transactions on Software Engineering, 48(9), 3280-3296. https://doi.org/10.1109/TSE.2021.3087402

Chen, J., Huang, H., Lyu, Y., An, J., Shi, J., Yang, C., Zhang, T., Tian, H., Li, Y., Li, Z., Zhou, X., Hu, X., & Lo, D. (2025). SecureAgentBench: Benchmarking secure code generation under realistic vulnerability scenarios. arXiv. https://arxiv.org/abs/2509.22097

Chen, X., Liu, C., & Song, D. (2018). Tree-to-tree neural networks for program translation. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.

Chen, Y., Ding, Z., Alowain, L., Chen, X., Deepmind, G., & Wagner, D. (2023). DiverseVUL: A vulnerable source code dataset. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses (pp. 654-668). ACM. https://doi.org/10.1145/3607199.3607242

Choi, Y.-D., Na, C. W., Kim, H., & Lee, J.-H. (2023). ReadSum: Retrieval-augmented transformer. IEEE Access, 11, 51155-51165. https://doi.org/10.1109/ACCESS.2023.3271992

Croft, R., Babar, M. A., & Kholoosi, M. M. (2023). Data quality for software vulnerability datasets. In Proceedings of the 45th International Conference on Software Engineering (pp. 121-133). IEEE. https://doi.org/10.1109/ICSE48619.2023.00022

Dai, S.-C., Xu, J., & Tao, G. (2025). Rethinking the evaluation of secure code generation. arXiv. https://arxiv.org/abs/2503.15554

Du, X., Liu, M., Wang, K., Wang, H., Liu, J., Chen, Y., Feng, J., Sha, C., & Peng, X. (2024). Evaluating LLMs in class-level code generation. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (pp. 1-13). IEEE/ACM. https://doi.org/10.1145/3597503.3639219

Fan, J., Li, Y., Wang, S., & Nguyen, T. N. (2020). AC/C++ code vulnerability dataset. In Proceedings of the 17th International Conference on Mining Software Repositories (pp. 508-512). ACM. https://doi.org/10.1145/3379597.3387501

Feng, Y., Wang, F., Wong, K. K., Wang, S., Yu-hong, L., Zhu, M., Wang, B., & Chen, W. (2023). PromptMagician: Interactive prompt engineering for image creation. IEEE Transactions on Visualization and Computer Graphics, 30, 295-305. https://doi.org/10.1109/TVCG.2023.3327168

Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W., Zettlemoyer, L., & Lewis, M. (2023). InCoder: A generative model for code infilling. arXiv. https://arxiv.org/abs/2204.05999

Ganti, M., Orr, L., & Wu, S. (2024). Evaluating text-to-SQL model failures on Real-World data. In Proceedings of the IEEE 38th International Conference on Data Engineering (pp. 1-1). IEEE. https://doi.org/10.1109/ICDE60146.2024.00456

Gokcimen, T., & Das, B. (2025). A novel system for strengthening security in LLMs. Alexandria Engineering Journal, 123, 71-90. https://doi.org/10.1016/j.aej.2025.03.030

Gu, X., Zhang, H., & Kim, S. (2018). Deep code search. In Proceedings of the 40th International Conference on Software Engineering (pp. 933-944). ACM. https://doi.org/10.1145/3180155.3180167

Guo, L. (2022). Using metacognitive prompts to enhance self-regulated learning. Journal of Computer Assisted Learning, 38(3), 811-832. https://doi.org/10.1111/jcal.12650

Hindle, A., Barr, E. T., Gabel, M., Su, Z., & Devanbu, P. (2016). On the naturalness of software. Communications of the ACM, 59(5), 122-131. https://doi.org/10.1145/2902362

Hliš, T., Četina, L., Beranič, T., & Pavlič, L. (2023). Evaluating usability of intelligent code assistants. Applied Sciences, 13(24), 13061. https://doi.org/10.3390/app132413061

Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., & Vinyals, O. (2022). An empirical analysis of compute-optimal large language model training. In Proceedings of the Advances in Neural Information Processing Systems (pp. 30016-30030). Curran Associates.

Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.

Jaoua, I., Ben, S. O., & Sahraoui, H. (2025). Combining large language models with static analyzers for code review generation. arXiv. https://arxiv.org/abs/2502.06633

Joshi, H., Sanchez, J. C., Gulwani, S., Le, V., Verbruggen, G., & Radiček, I. (2023). Repair is nearly generation: Multilingual program repair with LLMs. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 5131-5140). AAAI Press. https://doi.org/10.1609/aaai.v37i4.25642

Kalia, A. K., Xiao, J., Krishna, R., Sinha, S., Vuković, M., & Banerjee, D. (2021). Mono2Micro: a practical and effective tool for decomposing monolithic Java applications to microservices. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1214-1224). ACM. https://doi.org/10.1145/3468264.3473915

Kulal, S., Pasupat, P., Chandra, K., Lee, M., Padon, O., Aiken, A., & Liang, P. S. (2019). SPOC: Search-based pseudocode-to-code. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates.

Kumar, A., Jindal, K., Sharma, H., & Chaudhary, A. (2024). Enhancing syndicate lending through AI-Powered Fine-Tuning: Leveraging LLAMA2 for dynamic decision support. In Proceedings of the 2024 International Conference on Communication, Computer Sciences and Engineering (pp. 342-347). IEEE. https://doi.org/10.1109/IC3SE62002.2024.10593463

Lajkó, M., Csuvik, V., & Vidács, L. (2022). Towards JavaScript program repair with generative pre-trained transformer (GPT-2). In Proceedings of the Third International Workshop on Automated Program Repair (pp. 61-68). ACM. https://doi.org/10.1145/3524459.3527350

Latibari, B. S., Nazari, N., Alam Chowdhury, M., Immanuel Gubbi, K., Fang, C., Ghimire, S., Hosseini, E., Sayadi, H., Homayoun, H., Salehi, S., & Sasan, A. (2024). Transformers: A security perspective. IEEE Access, 12, 181071-181105. https://doi.org/10.1109/ACCESS.2024.3509372

Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. arXiv. https://arxiv.org/abs/2207.01780

LeClair, A., Jiang, S., & McMillan, C. (2019). Generating summaries of code. In Proceedings of the International Conference on Software Engineering (pp. 795-806). IEEE. https://doi.org/10.1109/ICSE.2019.00087

Li, Y., Shi, J., & Zhang, Z (2024). An approach for rapid source code development based on ChatGPT and prompt engineering. IEEE Access, 12, 53074-53087. https://doi.org/10.1109/ACCESS.2024.3385682

Li, Y., Wang, S., Nguyen, T. N., & Van Nguyen, S. (2019). Improving bug detection via context-based representation learning. Proceedings of the ACM on Programming Languages, 3(OOPSLA), 1-30. https://doi.org/10.1145/3360588

Lin, R., Fu, Y., Yi, W., Yang, J., Cao, J., Dong, Z., Xie, F., & Li, H. (2024). Vulnerabilities and security patch detection in OSS: A survey. ACM Computing Surveys, 57, 1-37. https://doi.org/10.1145/3694782

Lin, X. V., Wang, C., Zettlemoyer, L., & Ernst, M. D. (2018). NL2Bash: A corpus and semantic parser for natural language interface to the linux operating system. arXiv. https://arxiv.org/abs/1802.08979

Liu, F., Li, J., & Zhang, L. (2023). Syntax and domain aware model for unsupervised program translation. arXiv. https://arxiv.org/abs/2302.03908

Liu, S., Gao, C., Chen, S., Nie, L. Y., & Liu, Y. (2020). ATOM: Commit message generation based on abstract syntax tree and hybrid ranking. IEEE Transactions on Software Engineering, 48(5), 1800-1817. https://doi.org/10.1109/TSE.2020.3038681

Nikitopoulos, G., Dritsa, K., Louridas, P., & Mitropoulos, D. (2021). CrossVul dataset. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1565-1569). ACM. https://doi.org/10.1145/3468264.3473122

Nitin, V., Asthana, S., Ray, B., & Krishna, R. (2022). CARGO: AI-Guided dependency analysis for migrating monolithic applications to microservices architecture. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (pp. 1-12). IEEE/ACM. https://doi.org/10.1145/3551349.3556960

Obreja, D. M., & Rughiniș, R. (2023). The moral status of artificial intelligence: Exploring users’ anticipatory ethics in the controversy regarding LaMDA’s sentience. In Proceedings of the International Conference on Control Systems and Computer Science (pp. 411-417). IEEE. https://doi.org/10.1109/CSCS59211.2023.00071

Omar, M., Sorin, V., Collins, J. D., Reich, D., Freeman, R., Gavin, N., Charney, A., Stump, L., Bragazzi, N. L., Nadkarni, G. N., & Klang, E. (2025). Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Communications Medicine, 5(1). https://doi.org/10.1038/s43856-025-01021-3

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training LLMs with human feedback. In Proceedings of the Advances in Neural Information Processing Systems (pp. 27730-27744). Curran Associates.

Park, D., An, G. T., Kamyod, C., & Kim, C. G. (2023). A study on performance improvement of prompt engineering for generative AI with a large language model. Journal of Web Engineering, 22(8), 1187-1206. https://doi.org/10.13052/jwe1540-9589.2285

Samsi, S., Zhao, D., McDonald, J., Li, B., Michaleas, A., Jones, M., Bergeron, W., Kepner, J., Tiwari, D., & Gadepally, V. (2023). From words to watts: Benchmarking the energy costs of large language model inference. arXiv. https://doi.org/10.1109/HPEC58863.2023.10363447

SecureIT Project. (2021). CVEfixes Dataset. GitHub. https://github.com/secureIT-project/CVEfixes

Sv, S., Sunil, S., AS, P. A., & Satish, G. (2024). Democratizing data science using LLMs. In Proceedings of the 4th International Conference on Pervasive Computing and Social Networking (pp. 1065-1069). IEEE. https://doi.org/10.1109/ICPCSN62568.2024.00177

Teja, N. S., Kumar, K., & Malarvel, M. (2024). Multilingual text enhancer with Llama2. In Proceedings of the 3rd International Conference on Applied Artificial Intelligence and Computing (pp. 895-900). IEEE. https://doi.org/10.1109/ICAAIC60222.2024.10575122

Tóth, R., Bisztray, T., & Erdődi, L. (2024). LLMs in web development: Evaluating LLM-Generated PHP code unveiling vulnerabilities and limitations. In Proceedings of the Lecture Notes in Computer Science (pp. 425-437). Springer. https://doi.org/10.1007/978-3-031-68738-9_34

Vasiliniuc, M.-S., & Groza, A. (2023). AI-assisted mobile code generation. arXiv. https://arxiv.org/abs/2308.04736

Verbert, K., Manouselis, N., Ochoa, X., Wolpers, M., Drachsler, H., Bosnic, I., & Duval, E. (2012). Context-aware recommender systems for learning: A survey and future challenges. IEEE Transactions on Learning Technologies, 5(4), 318-335. https://doi.org/10.1109/TLT.2012.11

Wang, W., Li, G., Ma, B., Xia, X., & Jin, Z. (2020). Detecting code clones with GNNs. In Proceedings of the 27th International Conference on Software Analysis, Evolution and Reengineering (pp. 261-271). IEEE. https://doi.org/10.1109/SANER48275.2020.9054857

Wang, Y., Le, H., Gotmare, A. D., Bui, N. D. Q., Li, J., & Hoi, S. C. H. (2023). CodeT5+: A large code language model. arXiv. https://arxiv.org/abs/2305.07922

Wang, Y., Wang, W., Joty, S., & Hoi, S. C. H. (2021). CodeT5: Identifier-aware pretrained models. arXiv. https://arxiv.org/abs/2109.00859

Xu, S., Yao, Y., Feng, X., Gu, T., Tong, H., & Jian L. (2019). Commit Message Generation for Source Code Changes. https://doi.org/10.24963/ijcai.2019/552

Xu, X., Su, Z., Guo, J., Zhang, K., Wang, Z., & Zhang, X. (2024). ProSec: Security alignment for code LLMs. arXiv. https://arxiv.org/abs/2411.12882

Yao, D., Zhang, J., Harris, I. G., & Carlsson, M. (2024). FuzzLLM: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. arXiv. https://doi.org/10.1109/ICASSP48485.2024.10448041

Yin, P. (2021). Learning Structured Neural Semantic Parsers [Tesis doctoral, Carnegie Mellon University].

Yin, P., & Neubig, G. (2018). TranX: Neural abstract syntax parser. arXiv. https://arxiv.org/abs/1810.02720

Yu, L., Zhang, J., Wang, X., Ma, J., Yang, L., & Zhang, F. (2025). Towards secure and explainable smart contract generation with Security-Aware group relative policy optimization. arXiv. https://arxiv.org/abs/2509.09942

Yu, T., Zhang, R., Er, H. Y., Li, S., Xue, E., Pang, B., Lin, X. V., Tan, Y. C., Shi, T., Li, Z., Jiang, Y., Yasunaga, M., Shim, S., Chen, T., Fabbri, A., Li, Z., Chen, L., Zhang, Y., Dixit, S., & Zhang, V. (2019). COSQL: Conversational text-to-SQL. arXiv. https://arxiv.org/abs/1909.05378

Zhang, J., Wang, X., Zhang, H., Sun, H., Wang, K., & Liu, X. (2019). A novel neural source code representation based on abstract syntax tree. In Proceedings of the 41st International Conference on Software Engineering (pp. 783-794). IEEE. https://doi.org/10.1109/ICSE.2019.00086

Zhang, Y., Qiu, Z., Stol, K. J., Zhu, W., Zhu, J., Tian, Y., & Liu, H. (2024). Automatic commit message generation. IEEE Transactions on Software Engineering, 50, 816-835. https://doi.org/10.1109/TSE.2024.3364675

Zheng, Q., Xiao, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Shen, L., Wang, Z., Wang, A., Li, Y., Su, T., Yang, Z., & Tang, J. (2023). CodeGeeX: Pre-trained model for code generation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5673-5684). ACM. https://doi.org/10.1145/3580305.3599790

Zhou, Y., Liu, S., Siow, J., Du, X., & Liu, Y. (2019). Devign: Vulnerability identification via GNNs. In Proceedings of the Neural Information Processing Systems. Curran Associates.
Muñoz-Babiano, F., Lamo, P., & Alonso, R. S. (2026). Analysis of the Precision of Large Language Models in the Identification of Security Vulnerabilities and Weaknesses in Generated Code. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 14, e32926. https://doi.org/10.14201/adcaij.32926

Downloads

Download data is not yet available.
+ −