본문 바로가기
  • Home

Predicting the Size of Consonant Inventories through Machine Learning

  • Journal of Humanities, Seoul National University
  • 2026, 83(3), pp.297~325
  • Publisher : Institute of Humanities, Seoul National University
  • Research Area : Humanities > Other Humanities
  • Received : July 11, 2026
  • Accepted : August 10, 2026
  • Published : August 31, 2026

Lee Jin-Ho 1

1서울대학교

Accredited

ABSTRACT

Thi s study investigates how accurately machine learning can estimate consonant inventory sizes, or the number of consonant phonemes, across the world’s languages. A typologically balanced dataset of 3,035 languages was divided into training (2,124, 70%), validation (455, 15%), and test (456, 15%) sets. Predictive models were built using 47 binary consonant variables alongside genealogical and geographical data. Among five algorithms compared, LightGBM performed best. On the test set, this model achieved a 13.7% error rate and an R 2 of 0.770. Furthermore, 24.8% of the languages had an error rate below 5.0%, and the predictions showed 61.2% agreement with the W ALS-based five-level classification. Compared to baseline models, LightGBM attained high accuracy using only binary segmental features and limited extra-linguistic variables. Analysis revealed that prediction accuracy varied significantly across several factors but could not be explained by isolated consonants or simple co-occurrence patterns. The results suggest that the model’s predictions were more likely driven by complex, multidimensional relationships among variables than by isolated predictors.

Citation status

* References for papers published after 2025 are currently being built.