Analysis of the Precision of Large Language Models in the Identification of Security Vulnerabilities and Weaknesses in Generated Code
Abstract Large Language Models (LLMs) have accelerated code generation, yet their security implications remain underexplored. This work evaluates the capability of eight general-purpose and code-specialized LLMs to detect weaknesses and vulnerabilities in generated code using CVE and CWE references. Stratified sampling of the CVEfixes dataset across five programming languages is assessed using precision, recall, F1-score, and relevance-oriented metrics (REINP, REINR, and REINF1). Results show limited detection of specific vulnerabilities (CVE), but better performance for general weaknesses (CWE), especially in Python and Ruby. We discuss limitations in model training and prompt conditioning, and outline improvements through dataset diversification, prompt engineering and hybrid human-in-the-loop approaches. The study highlights the current potential and limitations of LLMs for practical vulnerability detection in generated code.
- Referencias
- Cómo citar
- Del mismo autor
- Métricas
Ahmad, B., Thakur, S., Tan, B., Karri, R., & Pearce, H. (2024). On hardware security bug code fixes by prompting large language models. IEEE Transactions on Information Forensics and Security, 19, 4043-4057. https://doi.org/10.1109/TIFS.2024.3374558
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., & Sanghai, S. (2023). GQA: Training generalized multi-query transformers from multi-head checkpoints. arXiv. https://doi.org/10.18653/v1/2023.emnlp-main.298
Albanese, M., Adebiyi, O., & Onovae, F. (2024). CVE2CWE: Automated mapping of software vulnerabilities. In Proceedings of the International Conference on Security and Cryptography (pp. 500-507). SCITEPRESS. https://doi.org/10.5220/0012770400003767
Alian, A. A., Sobhy, B., Nasser, M., & Hani, L. (2023). Backslash map: An automated vulnerability scanner. In Proceedings of the International Conference on Intelligent Computing and Information Systems (pp. 476-482). IEEE. https://doi.org/10.1109/ICICIS58388.2023.10391153
Allamanis, M., Jackson-Flux, H., & Brockschmidt, M. (2021). Self-supervised bug detection and repair. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.
Alon, U., Brody, S., Levy, O., & Yahav, E. (2018). code2seq: Sequences from structured code representations. arXiv. https://arxiv.org/abs/1808.01400
Alon, U., Zilberstein, M., Levy, O., & Yahav, E. (2019). code2vec: Learning distributed representations of code. In Proceedings of the Association for Computing Machinery. ACM. https://doi.org/10.1145/3290353
Backslash Security. (2025, abril). Can AI Vibe Coding Be Trusted? https://www.backslash.security/blog/can-ai-vibe-coding-be-trusted
Bhandari, G., Naseer, A., & Moonen, L. (2021). CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software. In Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (pp. 38-39). ACM. https://doi.org/10.1145/3475960.3475985
Bo, Y., Zhang, N., Li, S. P., & Xia, X. (2020). Survey of intelligent code completion. Journal of Software, 31(5), 1435-1453.
Cassano, F., Gouwar, J., Nguyen, D. P., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., Guha, A., Greenberg, M., & Jangda, A. (2023). Multipl-E: A scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7), 3675-3691. https://doi.org/10.1109/TSE.2023.3267446
Chakraborty, S., Krishna, R., Ding, Y., & Ray, B. (2021). Deep learning-based vulnerability detection. IEEE Transactions on Software Engineering, 48(9), 3280-3296. https://doi.org/10.1109/TSE.2021.3087402
Chen, J., Huang, H., Lyu, Y., An, J., Shi, J., Yang, C., Zhang, T., Tian, H., Li, Y., Li, Z., Zhou, X., Hu, X., & Lo, D. (2025). SecureAgentBench: Benchmarking secure code generation under realistic vulnerability scenarios. arXiv. https://arxiv.org/abs/2509.22097
Chen, X., Liu, C., & Song, D. (2018). Tree-to-tree neural networks for program translation. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.
Chen, Y., Ding, Z., Alowain, L., Chen, X., Deepmind, G., & Wagner, D. (2023). DiverseVUL: A vulnerable source code dataset. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses (pp. 654-668). ACM. https://doi.org/10.1145/3607199.3607242
Choi, Y.-D., Na, C. W., Kim, H., & Lee, J.-H. (2023). ReadSum: Retrieval-augmented transformer. IEEE Access, 11, 51155-51165. https://doi.org/10.1109/ACCESS.2023.3271992
Croft, R., Babar, M. A., & Kholoosi, M. M. (2023). Data quality for software vulnerability datasets. In Proceedings of the 45th International Conference on Software Engineering (pp. 121-133). IEEE. https://doi.org/10.1109/ICSE48619.2023.00022
Dai, S.-C., Xu, J., & Tao, G. (2025). Rethinking the evaluation of secure code generation. arXiv. https://arxiv.org/abs/2503.15554
Du, X., Liu, M., Wang, K., Wang, H., Liu, J., Chen, Y., Feng, J., Sha, C., & Peng, X. (2024). Evaluating LLMs in class-level code generation. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (pp. 1-13). IEEE/ACM. https://doi.org/10.1145/3597503.3639219
Fan, J., Li, Y., Wang, S., & Nguyen, T. N. (2020). AC/C++ code vulnerability dataset. In Proceedings of the 17th International Conference on Mining Software Repositories (pp. 508-512). ACM. https://doi.org/10.1145/3379597.3387501
Feng, Y., Wang, F., Wong, K. K., Wang, S., Yu-hong, L., Zhu, M., Wang, B., & Chen, W. (2023). PromptMagician: Interactive prompt engineering for image creation. IEEE Transactions on Visualization and Computer Graphics, 30, 295-305. https://doi.org/10.1109/TVCG.2023.3327168
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W., Zettlemoyer, L., & Lewis, M. (2023). InCoder: A generative model for code infilling. arXiv. https://arxiv.org/abs/2204.05999
Ganti, M., Orr, L., & Wu, S. (2024). Evaluating text-to-SQL model failures on Real-World data. In Proceedings of the IEEE 38th International Conference on Data Engineering (pp. 1-1). IEEE. https://doi.org/10.1109/ICDE60146.2024.00456
Gokcimen, T., & Das, B. (2025). A novel system for strengthening security in LLMs. Alexandria Engineering Journal, 123, 71-90. https://doi.org/10.1016/j.aej.2025.03.030
Gu, X., Zhang, H., & Kim, S. (2018). Deep code search. In Proceedings of the 40th International Conference on Software Engineering (pp. 933-944). ACM. https://doi.org/10.1145/3180155.3180167
Guo, L. (2022). Using metacognitive prompts to enhance self-regulated learning. Journal of Computer Assisted Learning, 38(3), 811-832. https://doi.org/10.1111/jcal.12650
Hindle, A., Barr, E. T., Gabel, M., Su, Z., & Devanbu, P. (2016). On the naturalness of software. Communications of the ACM, 59(5), 122-131. https://doi.org/10.1145/2902362
Hliš, T., Četina, L., Beranič, T., & Pavlič, L. (2023). Evaluating usability of intelligent code assistants. Applied Sciences, 13(24), 13061. https://doi.org/10.3390/app132413061
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., & Vinyals, O. (2022). An empirical analysis of compute-optimal large language model training. In Proceedings of the Advances in Neural Information Processing Systems (pp. 30016-30030). Curran Associates.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.
Jaoua, I., Ben, S. O., & Sahraoui, H. (2025). Combining large language models with static analyzers for code review generation. arXiv. https://arxiv.org/abs/2502.06633
Joshi, H., Sanchez, J. C., Gulwani, S., Le, V., Verbruggen, G., & Radiček, I. (2023). Repair is nearly generation: Multilingual program repair with LLMs. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 5131-5140). AAAI Press. https://doi.org/10.1609/aaai.v37i4.25642
Kalia, A. K., Xiao, J., Krishna, R., Sinha, S., Vuković, M., & Banerjee, D. (2021). Mono2Micro: a practical and effective tool for decomposing monolithic Java applications to microservices. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1214-1224). ACM. https://doi.org/10.1145/3468264.3473915
Kulal, S., Pasupat, P., Chandra, K., Lee, M., Padon, O., Aiken, A., & Liang, P. S. (2019). SPOC: Search-based pseudocode-to-code. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates.
Kumar, A., Jindal, K., Sharma, H., & Chaudhary, A. (2024). Enhancing syndicate lending through AI-Powered Fine-Tuning: Leveraging LLAMA2 for dynamic decision support. In Proceedings of the 2024 International Conference on Communication, Computer Sciences and Engineering (pp. 342-347). IEEE. https://doi.org/10.1109/IC3SE62002.2024.10593463
Lajkó, M., Csuvik, V., & Vidács, L. (2022). Towards JavaScript program repair with generative pre-trained transformer (GPT-2). In Proceedings of the Third International Workshop on Automated Program Repair (pp. 61-68). ACM. https://doi.org/10.1145/3524459.3527350
Latibari, B. S., Nazari, N., Alam Chowdhury, M., Immanuel Gubbi, K., Fang, C., Ghimire, S., Hosseini, E., Sayadi, H., Homayoun, H., Salehi, S., & Sasan, A. (2024). Transformers: A security perspective. IEEE Access, 12, 181071-181105. https://doi.org/10.1109/ACCESS.2024.3509372
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. arXiv. https://arxiv.org/abs/2207.01780
LeClair, A., Jiang, S., & McMillan, C. (2019). Generating summaries of code. In Proceedings of the International Conference on Software Engineering (pp. 795-806). IEEE. https://doi.org/10.1109/ICSE.2019.00087
Li, Y., Shi, J., & Zhang, Z (2024). An approach for rapid source code development based on ChatGPT and prompt engineering. IEEE Access, 12, 53074-53087. https://doi.org/10.1109/ACCESS.2024.3385682
Li, Y., Wang, S., Nguyen, T. N., & Van Nguyen, S. (2019). Improving bug detection via context-based representation learning. Proceedings of the ACM on Programming Languages, 3(OOPSLA), 1-30. https://doi.org/10.1145/3360588
Lin, R., Fu, Y., Yi, W., Yang, J., Cao, J., Dong, Z., Xie, F., & Li, H. (2024). Vulnerabilities and security patch detection in OSS: A survey. ACM Computing Surveys, 57, 1-37. https://doi.org/10.1145/3694782
Lin, X. V., Wang, C., Zettlemoyer, L., & Ernst, M. D. (2018). NL2Bash: A corpus and semantic parser for natural language interface to the linux operating system. arXiv. https://arxiv.org/abs/1802.08979
Liu, F., Li, J., & Zhang, L. (2023). Syntax and domain aware model for unsupervised program translation. arXiv. https://arxiv.org/abs/2302.03908
Liu, S., Gao, C., Chen, S., Nie, L. Y., & Liu, Y. (2020). ATOM: Commit message generation based on abstract syntax tree and hybrid ranking. IEEE Transactions on Software Engineering, 48(5), 1800-1817. https://doi.org/10.1109/TSE.2020.3038681
Nikitopoulos, G., Dritsa, K., Louridas, P., & Mitropoulos, D. (2021). CrossVul dataset. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1565-1569). ACM. https://doi.org/10.1145/3468264.3473122
Nitin, V., Asthana, S., Ray, B., & Krishna, R. (2022). CARGO: AI-Guided dependency analysis for migrating monolithic applications to microservices architecture. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (pp. 1-12). IEEE/ACM. https://doi.org/10.1145/3551349.3556960
Obreja, D. M., & Rughiniș, R. (2023). The moral status of artificial intelligence: Exploring users’ anticipatory ethics in the controversy regarding LaMDA’s sentience. In Proceedings of the International Conference on Control Systems and Computer Science (pp. 411-417). IEEE. https://doi.org/10.1109/CSCS59211.2023.00071
Omar, M., Sorin, V., Collins, J. D., Reich, D., Freeman, R., Gavin, N., Charney, A., Stump, L., Bragazzi, N. L., Nadkarni, G. N., & Klang, E. (2025). Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Communications Medicine, 5(1). https://doi.org/10.1038/s43856-025-01021-3
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training LLMs with human feedback. In Proceedings of the Advances in Neural Information Processing Systems (pp. 27730-27744). Curran Associates.
Park, D., An, G. T., Kamyod, C., & Kim, C. G. (2023). A study on performance improvement of prompt engineering for generative AI with a large language model. Journal of Web Engineering, 22(8), 1187-1206. https://doi.org/10.13052/jwe1540-9589.2285
Samsi, S., Zhao, D., McDonald, J., Li, B., Michaleas, A., Jones, M., Bergeron, W., Kepner, J., Tiwari, D., & Gadepally, V. (2023). From words to watts: Benchmarking the energy costs of large language model inference. arXiv. https://doi.org/10.1109/HPEC58863.2023.10363447
SecureIT Project. (2021). CVEfixes Dataset. GitHub. https://github.com/secureIT-project/CVEfixes
Sv, S., Sunil, S., AS, P. A., & Satish, G. (2024). Democratizing data science using LLMs. In Proceedings of the 4th International Conference on Pervasive Computing and Social Networking (pp. 1065-1069). IEEE. https://doi.org/10.1109/ICPCSN62568.2024.00177
Teja, N. S., Kumar, K., & Malarvel, M. (2024). Multilingual text enhancer with Llama2. In Proceedings of the 3rd International Conference on Applied Artificial Intelligence and Computing (pp. 895-900). IEEE. https://doi.org/10.1109/ICAAIC60222.2024.10575122
Tóth, R., Bisztray, T., & Erdődi, L. (2024). LLMs in web development: Evaluating LLM-Generated PHP code unveiling vulnerabilities and limitations. In Proceedings of the Lecture Notes in Computer Science (pp. 425-437). Springer. https://doi.org/10.1007/978-3-031-68738-9_34
Vasiliniuc, M.-S., & Groza, A. (2023). AI-assisted mobile code generation. arXiv. https://arxiv.org/abs/2308.04736
Verbert, K., Manouselis, N., Ochoa, X., Wolpers, M., Drachsler, H., Bosnic, I., & Duval, E. (2012). Context-aware recommender systems for learning: A survey and future challenges. IEEE Transactions on Learning Technologies, 5(4), 318-335. https://doi.org/10.1109/TLT.2012.11
Wang, W., Li, G., Ma, B., Xia, X., & Jin, Z. (2020). Detecting code clones with GNNs. In Proceedings of the 27th International Conference on Software Analysis, Evolution and Reengineering (pp. 261-271). IEEE. https://doi.org/10.1109/SANER48275.2020.9054857
Wang, Y., Le, H., Gotmare, A. D., Bui, N. D. Q., Li, J., & Hoi, S. C. H. (2023). CodeT5+: A large code language model. arXiv. https://arxiv.org/abs/2305.07922
Wang, Y., Wang, W., Joty, S., & Hoi, S. C. H. (2021). CodeT5: Identifier-aware pretrained models. arXiv. https://arxiv.org/abs/2109.00859
Xu, S., Yao, Y., Feng, X., Gu, T., Tong, H., & Jian L. (2019). Commit Message Generation for Source Code Changes. https://doi.org/10.24963/ijcai.2019/552
Xu, X., Su, Z., Guo, J., Zhang, K., Wang, Z., & Zhang, X. (2024). ProSec: Security alignment for code LLMs. arXiv. https://arxiv.org/abs/2411.12882
Yao, D., Zhang, J., Harris, I. G., & Carlsson, M. (2024). FuzzLLM: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. arXiv. https://doi.org/10.1109/ICASSP48485.2024.10448041
Yin, P. (2021). Learning Structured Neural Semantic Parsers [Tesis doctoral, Carnegie Mellon University].
Yin, P., & Neubig, G. (2018). TranX: Neural abstract syntax parser. arXiv. https://arxiv.org/abs/1810.02720
Yu, L., Zhang, J., Wang, X., Ma, J., Yang, L., & Zhang, F. (2025). Towards secure and explainable smart contract generation with Security-Aware group relative policy optimization. arXiv. https://arxiv.org/abs/2509.09942
Yu, T., Zhang, R., Er, H. Y., Li, S., Xue, E., Pang, B., Lin, X. V., Tan, Y. C., Shi, T., Li, Z., Jiang, Y., Yasunaga, M., Shim, S., Chen, T., Fabbri, A., Li, Z., Chen, L., Zhang, Y., Dixit, S., & Zhang, V. (2019). COSQL: Conversational text-to-SQL. arXiv. https://arxiv.org/abs/1909.05378
Zhang, J., Wang, X., Zhang, H., Sun, H., Wang, K., & Liu, X. (2019). A novel neural source code representation based on abstract syntax tree. In Proceedings of the 41st International Conference on Software Engineering (pp. 783-794). IEEE. https://doi.org/10.1109/ICSE.2019.00086
Zhang, Y., Qiu, Z., Stol, K. J., Zhu, W., Zhu, J., Tian, Y., & Liu, H. (2024). Automatic commit message generation. IEEE Transactions on Software Engineering, 50, 816-835. https://doi.org/10.1109/TSE.2024.3364675
Zheng, Q., Xiao, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Shen, L., Wang, Z., Wang, A., Li, Y., Su, T., Yang, Z., & Tang, J. (2023). CodeGeeX: Pre-trained model for code generation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5673-5684). ACM. https://doi.org/10.1145/3580305.3599790
Zhou, Y., Liu, S., Siow, J., Du, X., & Liu, Y. (2019). Devign: Vulnerability identification via GNNs. In Proceedings of the Neural Information Processing Systems. Curran Associates.
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., & Sanghai, S. (2023). GQA: Training generalized multi-query transformers from multi-head checkpoints. arXiv. https://doi.org/10.18653/v1/2023.emnlp-main.298
Albanese, M., Adebiyi, O., & Onovae, F. (2024). CVE2CWE: Automated mapping of software vulnerabilities. In Proceedings of the International Conference on Security and Cryptography (pp. 500-507). SCITEPRESS. https://doi.org/10.5220/0012770400003767
Alian, A. A., Sobhy, B., Nasser, M., & Hani, L. (2023). Backslash map: An automated vulnerability scanner. In Proceedings of the International Conference on Intelligent Computing and Information Systems (pp. 476-482). IEEE. https://doi.org/10.1109/ICICIS58388.2023.10391153
Allamanis, M., Jackson-Flux, H., & Brockschmidt, M. (2021). Self-supervised bug detection and repair. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.
Alon, U., Brody, S., Levy, O., & Yahav, E. (2018). code2seq: Sequences from structured code representations. arXiv. https://arxiv.org/abs/1808.01400
Alon, U., Zilberstein, M., Levy, O., & Yahav, E. (2019). code2vec: Learning distributed representations of code. In Proceedings of the Association for Computing Machinery. ACM. https://doi.org/10.1145/3290353
Backslash Security. (2025, abril). Can AI Vibe Coding Be Trusted? https://www.backslash.security/blog/can-ai-vibe-coding-be-trusted
Bhandari, G., Naseer, A., & Moonen, L. (2021). CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software. In Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (pp. 38-39). ACM. https://doi.org/10.1145/3475960.3475985
Bo, Y., Zhang, N., Li, S. P., & Xia, X. (2020). Survey of intelligent code completion. Journal of Software, 31(5), 1435-1453.
Cassano, F., Gouwar, J., Nguyen, D. P., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M.-H., Zi, Y., Anderson, C. J., Feldman, M. Q., Guha, A., Greenberg, M., & Jangda, A. (2023). Multipl-E: A scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7), 3675-3691. https://doi.org/10.1109/TSE.2023.3267446
Chakraborty, S., Krishna, R., Ding, Y., & Ray, B. (2021). Deep learning-based vulnerability detection. IEEE Transactions on Software Engineering, 48(9), 3280-3296. https://doi.org/10.1109/TSE.2021.3087402
Chen, J., Huang, H., Lyu, Y., An, J., Shi, J., Yang, C., Zhang, T., Tian, H., Li, Y., Li, Z., Zhou, X., Hu, X., & Lo, D. (2025). SecureAgentBench: Benchmarking secure code generation under realistic vulnerability scenarios. arXiv. https://arxiv.org/abs/2509.22097
Chen, X., Liu, C., & Song, D. (2018). Tree-to-tree neural networks for program translation. In Proceedings of the Neural Information Processing Systems (pp. 27865-27876). Curran Associates.
Chen, Y., Ding, Z., Alowain, L., Chen, X., Deepmind, G., & Wagner, D. (2023). DiverseVUL: A vulnerable source code dataset. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses (pp. 654-668). ACM. https://doi.org/10.1145/3607199.3607242
Choi, Y.-D., Na, C. W., Kim, H., & Lee, J.-H. (2023). ReadSum: Retrieval-augmented transformer. IEEE Access, 11, 51155-51165. https://doi.org/10.1109/ACCESS.2023.3271992
Croft, R., Babar, M. A., & Kholoosi, M. M. (2023). Data quality for software vulnerability datasets. In Proceedings of the 45th International Conference on Software Engineering (pp. 121-133). IEEE. https://doi.org/10.1109/ICSE48619.2023.00022
Dai, S.-C., Xu, J., & Tao, G. (2025). Rethinking the evaluation of secure code generation. arXiv. https://arxiv.org/abs/2503.15554
Du, X., Liu, M., Wang, K., Wang, H., Liu, J., Chen, Y., Feng, J., Sha, C., & Peng, X. (2024). Evaluating LLMs in class-level code generation. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (pp. 1-13). IEEE/ACM. https://doi.org/10.1145/3597503.3639219
Fan, J., Li, Y., Wang, S., & Nguyen, T. N. (2020). AC/C++ code vulnerability dataset. In Proceedings of the 17th International Conference on Mining Software Repositories (pp. 508-512). ACM. https://doi.org/10.1145/3379597.3387501
Feng, Y., Wang, F., Wong, K. K., Wang, S., Yu-hong, L., Zhu, M., Wang, B., & Chen, W. (2023). PromptMagician: Interactive prompt engineering for image creation. IEEE Transactions on Visualization and Computer Graphics, 30, 295-305. https://doi.org/10.1109/TVCG.2023.3327168
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W., Zettlemoyer, L., & Lewis, M. (2023). InCoder: A generative model for code infilling. arXiv. https://arxiv.org/abs/2204.05999
Ganti, M., Orr, L., & Wu, S. (2024). Evaluating text-to-SQL model failures on Real-World data. In Proceedings of the IEEE 38th International Conference on Data Engineering (pp. 1-1). IEEE. https://doi.org/10.1109/ICDE60146.2024.00456
Gokcimen, T., & Das, B. (2025). A novel system for strengthening security in LLMs. Alexandria Engineering Journal, 123, 71-90. https://doi.org/10.1016/j.aej.2025.03.030
Gu, X., Zhang, H., & Kim, S. (2018). Deep code search. In Proceedings of the 40th International Conference on Software Engineering (pp. 933-944). ACM. https://doi.org/10.1145/3180155.3180167
Guo, L. (2022). Using metacognitive prompts to enhance self-regulated learning. Journal of Computer Assisted Learning, 38(3), 811-832. https://doi.org/10.1111/jcal.12650
Hindle, A., Barr, E. T., Gabel, M., Su, Z., & Devanbu, P. (2016). On the naturalness of software. Communications of the ACM, 59(5), 122-131. https://doi.org/10.1145/2902362
Hliš, T., Četina, L., Beranič, T., & Pavlič, L. (2023). Evaluating usability of intelligent code assistants. Applied Sciences, 13(24), 13061. https://doi.org/10.3390/app132413061
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., & Vinyals, O. (2022). An empirical analysis of compute-optimal large language model training. In Proceedings of the Advances in Neural Information Processing Systems (pp. 30016-30030). Curran Associates.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.
Jaoua, I., Ben, S. O., & Sahraoui, H. (2025). Combining large language models with static analyzers for code review generation. arXiv. https://arxiv.org/abs/2502.06633
Joshi, H., Sanchez, J. C., Gulwani, S., Le, V., Verbruggen, G., & Radiček, I. (2023). Repair is nearly generation: Multilingual program repair with LLMs. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 5131-5140). AAAI Press. https://doi.org/10.1609/aaai.v37i4.25642
Kalia, A. K., Xiao, J., Krishna, R., Sinha, S., Vuković, M., & Banerjee, D. (2021). Mono2Micro: a practical and effective tool for decomposing monolithic Java applications to microservices. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1214-1224). ACM. https://doi.org/10.1145/3468264.3473915
Kulal, S., Pasupat, P., Chandra, K., Lee, M., Padon, O., Aiken, A., & Liang, P. S. (2019). SPOC: Search-based pseudocode-to-code. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates.
Kumar, A., Jindal, K., Sharma, H., & Chaudhary, A. (2024). Enhancing syndicate lending through AI-Powered Fine-Tuning: Leveraging LLAMA2 for dynamic decision support. In Proceedings of the 2024 International Conference on Communication, Computer Sciences and Engineering (pp. 342-347). IEEE. https://doi.org/10.1109/IC3SE62002.2024.10593463
Lajkó, M., Csuvik, V., & Vidács, L. (2022). Towards JavaScript program repair with generative pre-trained transformer (GPT-2). In Proceedings of the Third International Workshop on Automated Program Repair (pp. 61-68). ACM. https://doi.org/10.1145/3524459.3527350
Latibari, B. S., Nazari, N., Alam Chowdhury, M., Immanuel Gubbi, K., Fang, C., Ghimire, S., Hosseini, E., Sayadi, H., Homayoun, H., Salehi, S., & Sasan, A. (2024). Transformers: A security perspective. IEEE Access, 12, 181071-181105. https://doi.org/10.1109/ACCESS.2024.3509372
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. arXiv. https://arxiv.org/abs/2207.01780
LeClair, A., Jiang, S., & McMillan, C. (2019). Generating summaries of code. In Proceedings of the International Conference on Software Engineering (pp. 795-806). IEEE. https://doi.org/10.1109/ICSE.2019.00087
Li, Y., Shi, J., & Zhang, Z (2024). An approach for rapid source code development based on ChatGPT and prompt engineering. IEEE Access, 12, 53074-53087. https://doi.org/10.1109/ACCESS.2024.3385682
Li, Y., Wang, S., Nguyen, T. N., & Van Nguyen, S. (2019). Improving bug detection via context-based representation learning. Proceedings of the ACM on Programming Languages, 3(OOPSLA), 1-30. https://doi.org/10.1145/3360588
Lin, R., Fu, Y., Yi, W., Yang, J., Cao, J., Dong, Z., Xie, F., & Li, H. (2024). Vulnerabilities and security patch detection in OSS: A survey. ACM Computing Surveys, 57, 1-37. https://doi.org/10.1145/3694782
Lin, X. V., Wang, C., Zettlemoyer, L., & Ernst, M. D. (2018). NL2Bash: A corpus and semantic parser for natural language interface to the linux operating system. arXiv. https://arxiv.org/abs/1802.08979
Liu, F., Li, J., & Zhang, L. (2023). Syntax and domain aware model for unsupervised program translation. arXiv. https://arxiv.org/abs/2302.03908
Liu, S., Gao, C., Chen, S., Nie, L. Y., & Liu, Y. (2020). ATOM: Commit message generation based on abstract syntax tree and hybrid ranking. IEEE Transactions on Software Engineering, 48(5), 1800-1817. https://doi.org/10.1109/TSE.2020.3038681
Nikitopoulos, G., Dritsa, K., Louridas, P., & Mitropoulos, D. (2021). CrossVul dataset. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 1565-1569). ACM. https://doi.org/10.1145/3468264.3473122
Nitin, V., Asthana, S., Ray, B., & Krishna, R. (2022). CARGO: AI-Guided dependency analysis for migrating monolithic applications to microservices architecture. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (pp. 1-12). IEEE/ACM. https://doi.org/10.1145/3551349.3556960
Obreja, D. M., & Rughiniș, R. (2023). The moral status of artificial intelligence: Exploring users’ anticipatory ethics in the controversy regarding LaMDA’s sentience. In Proceedings of the International Conference on Control Systems and Computer Science (pp. 411-417). IEEE. https://doi.org/10.1109/CSCS59211.2023.00071
Omar, M., Sorin, V., Collins, J. D., Reich, D., Freeman, R., Gavin, N., Charney, A., Stump, L., Bragazzi, N. L., Nadkarni, G. N., & Klang, E. (2025). Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Communications Medicine, 5(1). https://doi.org/10.1038/s43856-025-01021-3
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training LLMs with human feedback. In Proceedings of the Advances in Neural Information Processing Systems (pp. 27730-27744). Curran Associates.
Park, D., An, G. T., Kamyod, C., & Kim, C. G. (2023). A study on performance improvement of prompt engineering for generative AI with a large language model. Journal of Web Engineering, 22(8), 1187-1206. https://doi.org/10.13052/jwe1540-9589.2285
Samsi, S., Zhao, D., McDonald, J., Li, B., Michaleas, A., Jones, M., Bergeron, W., Kepner, J., Tiwari, D., & Gadepally, V. (2023). From words to watts: Benchmarking the energy costs of large language model inference. arXiv. https://doi.org/10.1109/HPEC58863.2023.10363447
SecureIT Project. (2021). CVEfixes Dataset. GitHub. https://github.com/secureIT-project/CVEfixes
Sv, S., Sunil, S., AS, P. A., & Satish, G. (2024). Democratizing data science using LLMs. In Proceedings of the 4th International Conference on Pervasive Computing and Social Networking (pp. 1065-1069). IEEE. https://doi.org/10.1109/ICPCSN62568.2024.00177
Teja, N. S., Kumar, K., & Malarvel, M. (2024). Multilingual text enhancer with Llama2. In Proceedings of the 3rd International Conference on Applied Artificial Intelligence and Computing (pp. 895-900). IEEE. https://doi.org/10.1109/ICAAIC60222.2024.10575122
Tóth, R., Bisztray, T., & Erdődi, L. (2024). LLMs in web development: Evaluating LLM-Generated PHP code unveiling vulnerabilities and limitations. In Proceedings of the Lecture Notes in Computer Science (pp. 425-437). Springer. https://doi.org/10.1007/978-3-031-68738-9_34
Vasiliniuc, M.-S., & Groza, A. (2023). AI-assisted mobile code generation. arXiv. https://arxiv.org/abs/2308.04736
Verbert, K., Manouselis, N., Ochoa, X., Wolpers, M., Drachsler, H., Bosnic, I., & Duval, E. (2012). Context-aware recommender systems for learning: A survey and future challenges. IEEE Transactions on Learning Technologies, 5(4), 318-335. https://doi.org/10.1109/TLT.2012.11
Wang, W., Li, G., Ma, B., Xia, X., & Jin, Z. (2020). Detecting code clones with GNNs. In Proceedings of the 27th International Conference on Software Analysis, Evolution and Reengineering (pp. 261-271). IEEE. https://doi.org/10.1109/SANER48275.2020.9054857
Wang, Y., Le, H., Gotmare, A. D., Bui, N. D. Q., Li, J., & Hoi, S. C. H. (2023). CodeT5+: A large code language model. arXiv. https://arxiv.org/abs/2305.07922
Wang, Y., Wang, W., Joty, S., & Hoi, S. C. H. (2021). CodeT5: Identifier-aware pretrained models. arXiv. https://arxiv.org/abs/2109.00859
Xu, S., Yao, Y., Feng, X., Gu, T., Tong, H., & Jian L. (2019). Commit Message Generation for Source Code Changes. https://doi.org/10.24963/ijcai.2019/552
Xu, X., Su, Z., Guo, J., Zhang, K., Wang, Z., & Zhang, X. (2024). ProSec: Security alignment for code LLMs. arXiv. https://arxiv.org/abs/2411.12882
Yao, D., Zhang, J., Harris, I. G., & Carlsson, M. (2024). FuzzLLM: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. arXiv. https://doi.org/10.1109/ICASSP48485.2024.10448041
Yin, P. (2021). Learning Structured Neural Semantic Parsers [Tesis doctoral, Carnegie Mellon University].
Yin, P., & Neubig, G. (2018). TranX: Neural abstract syntax parser. arXiv. https://arxiv.org/abs/1810.02720
Yu, L., Zhang, J., Wang, X., Ma, J., Yang, L., & Zhang, F. (2025). Towards secure and explainable smart contract generation with Security-Aware group relative policy optimization. arXiv. https://arxiv.org/abs/2509.09942
Yu, T., Zhang, R., Er, H. Y., Li, S., Xue, E., Pang, B., Lin, X. V., Tan, Y. C., Shi, T., Li, Z., Jiang, Y., Yasunaga, M., Shim, S., Chen, T., Fabbri, A., Li, Z., Chen, L., Zhang, Y., Dixit, S., & Zhang, V. (2019). COSQL: Conversational text-to-SQL. arXiv. https://arxiv.org/abs/1909.05378
Zhang, J., Wang, X., Zhang, H., Sun, H., Wang, K., & Liu, X. (2019). A novel neural source code representation based on abstract syntax tree. In Proceedings of the 41st International Conference on Software Engineering (pp. 783-794). IEEE. https://doi.org/10.1109/ICSE.2019.00086
Zhang, Y., Qiu, Z., Stol, K. J., Zhu, W., Zhu, J., Tian, Y., & Liu, H. (2024). Automatic commit message generation. IEEE Transactions on Software Engineering, 50, 816-835. https://doi.org/10.1109/TSE.2024.3364675
Zheng, Q., Xiao, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Shen, L., Wang, Z., Wang, A., Li, Y., Su, T., Yang, Z., & Tang, J. (2023). CodeGeeX: Pre-trained model for code generation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5673-5684). ACM. https://doi.org/10.1145/3580305.3599790
Zhou, Y., Liu, S., Siow, J., Du, X., & Liu, Y. (2019). Devign: Vulnerability identification via GNNs. In Proceedings of the Neural Information Processing Systems. Curran Associates.
Muñoz-Babiano, F., Lamo, P., & Alonso, R. S. (2026). Analysis of the Precision of Large Language Models in the Identification of Security Vulnerabilities and Weaknesses in Generated Code. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 14, e32926. https://doi.org/10.14201/adcaij.32926
Downloads
Download data is not yet available.
+
−