@article{ART003371807},
author={이일권 and Seungrak Choi},
title={Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)},
journal={International Journal of Glocal Language and Literary Studies(약칭: IGLL)},
year={2026},
volume={23},
number={23},
pages={373-387}
TY - JOUR
AU - 이일권
AU - Seungrak Choi
TI - Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)
JO - International Journal of Glocal Language and Literary Studies(약칭: IGLL)
PY - 2026
VL - 23
IS - 23
PB - Glocal Institute of Language and Literary Studies(GILLS)
SP - 373
EP - 387
AB - Critical thinking is widely regarded as a core competency in the age of artificial intelligence, and cultivating it requires the ability to evaluate arguments properly. This study assessed the argument- evaluation capabilities of commercially available AI systems powered by large language models (LLMs) using argument-strengthening and argument-weakening items from Korea’s Public Service Aptitude Test ( PSAT). First, it discusses the educational and practical education. Based on this theoretical discussion, 20 previously administered items were selected from the verbal reasoning section of the PSAT to compare the performance of ChatGPT 5.2 Pro, Claude Opus 4.5, and Gemini 3 Ultra. Each model was presented with the same items ten times. ChatGPT achieved the highest accuracy rate at 99.0%, followed by Claude at 87.0% and Gemini at 83.0%. In particular, repeated testing revealed systematic error patterns specific to each model, as well as persistent errors that would be difficult to detect from a single response alone. Given that existing performance evaluations of LLM-based systems have focused primarily on mathematical and coding abilities, this study highlights the value of strengthening and weakening items as a means of comprehensively assessing the core components of critical thinking. It further argues that review and evaluation by instructors remain essential when AI is used for educational purposes.
KW - Critical Thinking;Argument Evaluation;Strengthen/Weaken items;Large Language Model(LLM);PSAT(Public Service Aptitude Test)
DO -
UR -
ER -
이일권 and Seungrak Choi. (2026). Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT). International Journal of Glocal Language and Literary Studies(약칭: IGLL), 23(23), 373-387.
이일권 and Seungrak Choi. 2026, "Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)", International Journal of Glocal Language and Literary Studies(약칭: IGLL), vol.23, no.23 pp.373-387.
이일권, Seungrak Choi "Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)" International Journal of Glocal Language and Literary Studies(약칭: IGLL) 23.23 pp.373-387 (2026) : 373.
이일권, Seungrak Choi. Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT). 2026; 23(23), 373-387.
이일권 and Seungrak Choi. "Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)" International Journal of Glocal Language and Literary Studies(약칭: IGLL) 23, no.23 (2026) : 373-387.
이일권; Seungrak Choi. Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT). International Journal of Glocal Language and Literary Studies(약칭: IGLL), 23(23), 373-387.
이일권; Seungrak Choi. Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT). International Journal of Glocal Language and Literary Studies(약칭: IGLL). 2026; 23(23) 373-387.
이일권, Seungrak Choi. Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT). 2026; 23(23), 373-387.
이일권 and Seungrak Choi. "Argument-Evaluation Capabilities of Commercial AI: A Focus on Strengthening and Weakening Items in the Public Service Aptitude Test (PSAT)" International Journal of Glocal Language and Literary Studies(약칭: IGLL) 23, no.23 (2026) : 373-387.