본문 바로가기
  • Home

A comparable-corpus study of linguistic features in Korean-to-Chinese translations produced for AI training and published in Korean newspapers

  • The Journal of Translation Studies
  • Abbr : JTS
  • 2026, 27(3), pp.293~333
  • DOI : 10.15749/jts.2026.27.3.010
  • Publisher : The Korean Association for Translation Studies
  • Research Area : Humanities > Interpretation and Translation Studies
  • Received : August 15, 2026
  • Accepted : September 15, 2026
  • Published : September 30, 2026

Hwang Eun Ha ORD ID 1,  Fei, Li 2

1배재대학교
2숙명여자대학교

Accredited

ABSTRACT

This study compares Korean-to-Chinese translations compiled for AI training and published in Korean newspaper Chinese editions with non-translated Chinese news. A three-way comparable corpus of 720,000 sentences, comprising 240,000 sentences per subcorpus, was constructed from the AI Hub Korean–Chinese parallel corpus, four Korean newspaper Chinese editions, and THUCNews. The subcorpora were balanced for news genre and the social, economic, and cultural domains; the AI Hub subset was restricted to files sourced exclusively from Korean news outlets where source information was available. Both translation sets showed lower size-adjusted lexical diversity than the reference corpus and a shared preference for expanded noun phrases over clause chaining. AI-training translations showed still lower lexical diversity, particularly in particles, idioms, adverbs, and conjunctions, as well as more frequent right-adjunct relations and use of the attributive marker de (的). Their consistent use of half-width punctuation resulted from the corpus-construction pipeline rather than from translation. The shared tendencies are consistent with simplification and normalization. Differences between the translation sets may be associated with intended readership, editorial unit, quality-control procedures, and translation production method, although their individual effects cannot be distinguished in the present design. Quality control for AI training corpora should therefore extend beyond source-target correspondence to target-language linguistic profiles and corpus-construction procedures.

Journal Copyright Policy

No CCL information provided

Citation status

* References for papers published after 2025 are currently being built.