@article{ART002102136},
author={Bokkeun Sun},
title={Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System},
journal={Journal of The Korea Society of Computer and Information},
issn={1598-849X},
year={2016},
volume={21},
number={4},
pages={55-62}
TY - JOUR
AU - Bokkeun Sun
TI - Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System
JO - Journal of The Korea Society of Computer and Information
PY - 2016
VL - 21
IS - 4
PB - The Korean Society Of Computer And Information
SP - 55
EP - 62
SN - 1598-849X
AB - In this paper, we propose the INDEM(Internet News Data Extraction Middleware) system for the removal of the unnecessary data in internet news. Although data on the internet can be used in various fields such as source of data of IR(Information Retrieval), Data mining and knowledge information service, it contains a lot of unnecessary information. The removal of the unnecessary data is a problem to be solved prior to the study of the knowledge-based information service that is based on the data of the web page. The INDEM system parses html and explores the XPath, and it is to perform the analysis. The user simply utilize INDEM by implementing an abstract class that provides INDEM, and can obtain the analysis information. INDEM System through this process delivers the analysis information including the main contents of news site to the users. In this paper, the INDEM system was adapted in a stand-alone and web service system and it was evaluated on the basis of 16 news site. As a result, performance of the INDEM system is affected in html source data size and complexity of used html grammar than the main news data size.
KW - Information Extraction;Web news page;Middleware;XPath Grouping;Data Mining
DO -
UR -
ER -
Bokkeun Sun. (2016). Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System. Journal of The Korea Society of Computer and Information, 21(4), 55-62.
Bokkeun Sun. 2016, "Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System", Journal of The Korea Society of Computer and Information, vol.21, no.4 pp.55-62.
Bokkeun Sun "Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System" Journal of The Korea Society of Computer and Information 21.4 pp.55-62 (2016) : 55.
Bokkeun Sun. Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System. 2016; 21(4), 55-62.
Bokkeun Sun. "Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System" Journal of The Korea Society of Computer and Information 21, no.4 (2016) : 55-62.
Bokkeun Sun. Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System. Journal of The Korea Society of Computer and Information, 21(4), 55-62.
Bokkeun Sun. Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System. Journal of The Korea Society of Computer and Information. 2016; 21(4) 55-62.
Bokkeun Sun. Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System. 2016; 21(4), 55-62.
Bokkeun Sun. "Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System" Journal of The Korea Society of Computer and Information 21, no.4 (2016) : 55-62.