AI-Assisted Assessment of Programming Assignments: A Comparative Study of Automated and Human Evaluation Methods
Author Affiliations
- 1Department of Computer Science, Nrupathunga University (Formerly Government Science College), Bangalore, Karnataka, India
- 2Department of CSE, BGS College of Engineering and Technology, Bangalore, Karnataka, India
Res. J. Recent Sci., Volume 15, Issue (3), Pages 91-95, July,2 (2026)
Abstract
Assessment of programming assignments is an essential part of teaching computer science. Conventional manual grading takes a lot of time and is frequently inconsistent. Using automated code analysis and artificial intelligence methods, the study suggests an Artificial Intelligence (AI) -assisted assessment system for programming assignments. Logical correctness, language syntax, readability, documentation, and code standards are all assessed by the proposed framework. The findings show that while retaining a high degree of agreement with instructor ratings, AI-assisted assessment can drastically cut down on grading time. According to the results, AI can be a useful tool for decision-making to assess a large number of programming assignments to save time.
References
- Bernik, A., Radošević, D., & Čep, A. (2025)., A comparative study of large language models in programming education: Accuracy, efficiency, and feedback in student assignment grading., Applied Sciences, 15(18), 10055.
- Jukiewicz, M. (2026)., A systematic comparison of large language models for automated assignment assessment in programming education: Exploring the importance of architecture and vendor., Computers and Education Open, 10, 100364.
- Mohamed, K., Yousef, M., Medhat, W., Mohamed, E. H., Khoriba, G., & Arafa, T. (2025)., Hands-on analysis of using large language models for the auto evaluation of programming assignments., Information Systems, 128, 102473.
- Ala-Mutka, K. M. (2005)., A survey of automated assessment approaches for programming assignments., Computer Science Education, 15(2), 83–102.
- Kiesler, N., & Schiffner, D. (2023)., Large language models in introductory programming education: ChatGPT, Cornell University.
- Yousef, M., Mohamed, K., Medhat, W., Mohamed, E. H., Khoriba, G., & Arafa, T. (2025)., BeGrading: large language models for enhanced feedback in programming education., Neural Computing and Applications, 37(2), 1027-1040.
- Raihan, N., Goswami, D., Puspo, S. S. C., Siddiq, M. L., Newman, C., Ranasinghe, T., ... & Zampieri, M. (2026)., On the performance of large language models on introductory programming assignments., Journal of Intelligent Information Systems, 64(1), 239-263.
- Fan, G., Liu, D., Zhang, R., & Pan, L. (2025)., The impact of AI-assisted pair programming on student motivation, programming anxiety, collaborative learning, and programming performance: A comparative study with traditional pair programming and individual approaches., International Journal of STEM Education, 12(1), 16.
- Cohen, J. (1960)., A coefficient of agreement for nominal scales., Educational and Psychological Measurement, 20(1), 37–46.
- Zhang, D. W., Boey, M., Tan, Y. Y., & Jia, A. H. S. (2024)., Evaluating large language models for criterion-based grading from agreement to consistency., npj Science of Learning, 9(1), 79.
