본문 바로가기
  • Home

A Methodology for LLM-Based Assessment of Literature Essay Examinations : From the Examiner’s Perspective through a Case Study on Camus’s L’Étranger and Saint-Exupéry’s Le Petit Prince

  • The Journal of General Education
  • 2026, (36), pp.159~194
  • DOI : 10.24173/jge.2026.07.31.5
  • Publisher : Da Vinci Mirae Institute of General Education
  • Research Area : Social Science > Education > Field of Education > General Education
  • Received : June 21, 2026
  • Accepted : July 20, 2026
  • Published : July 31, 2026

Hwa-Jin Ahn 1

1한국과학기술원

Accredited

ABSTRACT

This study proposes an LLM-based methodology for evaluating French literature essay examinations and validates its effectiveness using 30 mid-term exam scripts evaluated by the instructor who both designed and graded the assessment. Unlike conventional validation studies employing independent external raters, this study examines the extent to which an LLM can reproduce the evaluation standards of that instructor. Focusing on comparative essays on Camus’s L’étranger and Saint-Exupéry’s Le Petit Prince, we developed an LLM scoring pipeline (utilizing Claude Sonnet 4.6) comprising: (1) question-specific credit/penalty matrices, (2) Dual-Reference Fusion (DRF) merging analytical and summary model answers, and (3) multi-axis scoring combining textual fidelity and stance clarity, augmented by a Stance Coherence Check. Quantitative analysis revealed a strong correlation between the AI and the human grader (Pearson r=0.786, Spearman ρ=0.651). Qualitatively, while the LLM exhibited relative leniency toward lower-tier responses, the human instructor more actively recognized and rewarded the intellectual potential in top-tier essays. These findings demonstrate that a multi-stage prompt design aligned with pedagogical intent enables LLMs to stably track the nuances of humanistic reasoning. Ultimately, this study underscores the practical viability of positioning LLMs not as a replacement, but as a first-pass "co-rater" in humanities assessment. While this methodology may extend to other humanities essays involving dual-text comparative structures, the empirical validation reported here is limited to the two works and single course examined.

Citation status

* References for papers published after 2025 are currently being built.