SeonahLee
|
LIM, BYUNG HWA
| 2026, 14(1)
| pp.17~44
| number of Cited : 0
This study constructs large language model (LLM)-based dynamic business networks for KOSPI-listed firms and examines whether the output language of LLM-generated business descriptions affects network informativeness. Using annual reports from 2010 to 2024, we extract the Business Overview and MD&A sections and generate standardized business descriptions in both Korean and English from the same underlying disclosures. To reduce look-ahead bias, firm and product identifiers are masked before LLM generation. We then embed the generated descriptions with multiple language and finance-specific models and construct annual peer networks based on cosine similarity. The results show that the proposed networks identify firm-level economic peers not fully captured by the Korean Standard Industrial Classification (KSIC). More importantly, Korean summaries and English summaries generated from Korean disclosures produce different similarity structures and lead-lag return signals. Asset-pricing tests indicate that network-based peer-return factors generate significant alphas beyond standard risk factors and a KSIC-based benchmark. LLM-based English business descriptions tend to deliver more stable and broadly significant lead-lag performance in several specifications, while Korean descriptions also provide meaningful signals in selected cases. These findings suggest that output-language choice is not a neutral preprocessing step, but a substantive modeling decision in LLM-based financial text analysis.