본문 바로가기
  • Home

Effectiveness of a GPT-OSS-Based Automated Essay Scoring System and the Educational Value of LLM-Generated Feedback: A Case Study of Descriptive Writing Tasks

  • The Journal of General Education
  • 2026, (35), pp.127~157
  • DOI : 10.24173/jge.2026.04.30.4
  • Publisher : Da Vinci Mirae Institute of General Education
  • Research Area : Social Science > Education > Field of Education > General Education
  • Received : March 22, 2026
  • Accepted : April 20, 2026
  • Published : April 30, 2026

Jo Meounggun 1 Lee, Ju Hyun 2

1호서대학교
2한국과학기술원

Accredited

ABSTRACT

This study aims to examine whether an automated essay scoring system based on a large language model (LLM) can function as a practical assessment tool in university writing education, and to investigate the educational value of the feedback it generates. To this end, an automated scoring system was implemented using the GPT-OSS 20B model in a local environment, and three scoring strategies—persona-based, chain-of-thought, and pairwise comparison—were applied to the same dataset and comparatively analyzed. In addition to measuring score agreement, the quality of the generated feedback was also evaluated by experts to examine its qualitative validity. The results show that the chain-of-thought approach demonstrated the highest agreement with human scoring, while the persona-based approach showed a conservative scoring tendency and the pairwise comparison approach showed a lenient tendency. In addition, the LLM showed high consistency in identifying low-quality texts but had limitations in distinguishing subtle qualitative differences among high-performing texts. The generated feedback focused on content organization and expression improvement, indicating its educational applicability. Meanwhile, differences in computational cost were observed depending on the scoring strategy, and in particular, the pairwise comparison approach showed high computational burden, suggesting limitations for practical application. These findings suggest that LLM-based automated scoring systems are more effective when used as collaborative tools supporting feedback provision rather than as replacements for human evaluators.

Citation status

* References for papers published after 2025 are currently being built.