본문 바로가기
  • Home

Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)

  • International Journal of Glocal Language and Literary Studies(약칭: IGLL)
  • Abbr : IGLL
  • 2026, 23(23), pp.373~387
  • Publisher : Glocal Institute of Language and Literary Studies(GILLS)
  • Research Area : Humanities > Other Humanities
  • Received : July 20, 2026
  • Accepted : August 20, 2026
  • Published : August 30, 2026

이일권 1 Seungrak Choi 2

1인사혁신처
2한림대학교

Accredited

ABSTRACT

Critical thinking is widely regarded as a core competency in the age of artificial intelligence, and cultivating it requires the ability to evaluate arguments properly. This study assessed the argument- evaluation capabilities of commercially available AI systems powered by large language models (LLMs) using argument-strengthening and argument-weakening items from Korea’s Public Service Aptitude Test ( PSAT). First, it discusses the educational and practical education. Based on this theoretical discussion, 20 previously administered items were selected from the verbal reasoning section of the PSAT to compare the performance of ChatGPT 5.2 Pro, Claude Opus 4.5, and Gemini 3 Ultra. Each model was presented with the same items ten times. ChatGPT achieved the highest accuracy rate at 99.0%, followed by Claude at 87.0% and Gemini at 83.0%. In particular, repeated testing revealed systematic error patterns specific to each model, as well as persistent errors that would be difficult to detect from a single response alone. Given that existing performance evaluations of LLM-based systems have focused primarily on mathematical and coding abilities, this study highlights the value of strengthening and weakening items as a means of comprehensively assessing the core components of critical thinking. It further argues that review and evaluation by instructors remain essential when AI is used for educational purposes.

Journal Copyright Policy

No CCL information provided

Citation status

* References for papers published after 2025 are currently being built.

This paper was written with support from the National Research Foundation of Korea.