This study investigated whether and to what extent EFL college learners and a trained bilingual rater differed in evaluating the adequacy of machine translation (MT) across four text types: English for Specific Purposes (ESP), argumentative (Opinion), expository (Information), and literary (Literature). A mixed-methods design was adopted to examine translation adequacy ratings and learner perceptions of MTassisted reading. At the macro level, the two groups showed close agreement: both rated Opinion and ESP as highly adequate (Rater M = 4.63, 4.55; Student M = 4.29, 4.41) and Literature as the most vulnerable domain (Rater M = 3.64; Student M = 4.06).
Despite this shared rank order, a notable disparity emerged at the micro level. The rater scored Literature more critically owing to its stylistic complexity, whereas the learners rated it more leniently, a pattern that may reflect the surface-level readability of the text rather than the quality of its translation. Survey results also revealed a mismatch between perceived accuracy and actual preference: although ChatGPT was rated as the most accurate tool by 62% of participants, more learners preferred Papago (57.1%) over ChatGPT (42.9%) due to its L1-optimized interface. Furthermore, significant affective trade-offs emerged as the learners viewed MT as a helpful scaffold for reading, while simultaneously holding mixed feelings of trust and skepticism toward MT output. The study concludes with pedagogical implications for developing critical MT literacy in EFL curricula.