@article{ART002881708},
author={Gyu-min Lee and Sanghoun Song},
title={The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills},
journal={The Sociolinguistic Journal of Korea},
issn={1226-4822},
year={2022},
volume={30},
number={3},
pages={213-243},
doi={10.14353/sjk.2022.30.3.08}
TY - JOUR
AU - Gyu-min Lee
AU - Sanghoun Song
TI - The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills
JO - The Sociolinguistic Journal of Korea
PY - 2022
VL - 30
IS - 3
PB - The Sociolinguistic Society Of Korea
SP - 213
EP - 243
SN - 1226-4822
AB - Despite the massive impact of COVID-19 on society, beyond the numbers of confirmed cases and deaths, there remains a lack of large-scale data depicting the effects of the virus on the society of the Republic of Korea. To fill this gap, we collected 1.822 million news articles with more than 1 billion morphemes from January 2020 to June 2022, creating a Korean version of the Coronavirus Corpus. This corpus is introduced in the current study. In addition, to demonstrate how such massive corpus can be utilized, we conducted information theoretical analyses to see how the stance of the press media on topics such as vaccines and social distancing affected the COVID-19 situation in the Republic of Korea. Specifically, we utilized several computational linguistic skills including concordance building and sentiment analysis through both traditional and machine learning techniques and measured the transfer entropy to estimate the impact with information theory. The results suggest that the overall impact of the press media on the society was minimal to non-existent.
KW - COVID-19;media;Republic of Korea;corpus;computational linguistics;sentiment analysis;diachronic analysis
DO - 10.14353/sjk.2022.30.3.08
ER -
Gyu-min Lee and Sanghoun Song. (2022). The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills. The Sociolinguistic Journal of Korea, 30(3), 213-243.
Gyu-min Lee and Sanghoun Song. 2022, "The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills", The Sociolinguistic Journal of Korea, vol.30, no.3 pp.213-243. Available from: doi:10.14353/sjk.2022.30.3.08
Gyu-min Lee, Sanghoun Song "The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills" The Sociolinguistic Journal of Korea 30.3 pp.213-243 (2022) : 213.
Gyu-min Lee, Sanghoun Song. The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills. 2022; 30(3), 213-243. Available from: doi:10.14353/sjk.2022.30.3.08
Gyu-min Lee and Sanghoun Song. "The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills" The Sociolinguistic Journal of Korea 30, no.3 (2022) : 213-243.doi: 10.14353/sjk.2022.30.3.08
Gyu-min Lee; Sanghoun Song. The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills. The Sociolinguistic Journal of Korea, 30(3), 213-243. doi: 10.14353/sjk.2022.30.3.08
Gyu-min Lee; Sanghoun Song. The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills. The Sociolinguistic Journal of Korea. 2022; 30(3) 213-243. doi: 10.14353/sjk.2022.30.3.08
Gyu-min Lee, Sanghoun Song. The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills. 2022; 30(3), 213-243. Available from: doi:10.14353/sjk.2022.30.3.08
Gyu-min Lee and Sanghoun Song. "The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills" The Sociolinguistic Journal of Korea 30, no.3 (2022) : 213-243.doi: 10.14353/sjk.2022.30.3.08