본문 바로가기
  • Home

Design and Implementation of a PostgreSQL-Milvus Data Pipeline with a PostgreSQL-Based Korean Hybrid Retrieval Module: An Evaluation of Retrieval Performance

  • Journal of The Korea Society of Computer and Information
  • Abbr : JKSCI
  • 2026, 31(9), pp.237~246
  • Publisher : The Korean Society Of Computer And Information
  • Research Area : Engineering > Computer Science
  • Received : July 29, 2026
  • Accepted : August 20, 2026
  • Published : September 30, 2026

Min-Ho Park 1,  Hyunchul Ahn 1

1국민대학교

Accredited

ABSTRACT

This paper presents the design and implementation of a PostgreSQL–Milvus data pipeline for Korean semantic search applications, including ontology-oriented services. PostgreSQL 18 is used as the source of truth for managing structured data and dense embeddings stored through the pgvector extension. INSERT, UPDATE, and DELETE events are captured by Debezium and delivered through Kafka to a dedicated Milvus consumer, allowing changes in PostgreSQL to be propagated to Milvus. For Korean information retrieval, the pipeline incorporates the BGE-M3 multilingual embedding model, which generates 1,024-dimensional dense vectors, and the textsearch_ko extension based on the mecab-ko morphological analyzer. A PostgreSQL-based hybrid retrieval function combines dense and lexical rankings using weighted reciprocal rank fusion within a single SQL statement. As a component-level evaluation, the retrieval effectiveness of the PostgreSQL-based hybrid search module was examined using an evaluation subset constructed from MIRACL-ko, consisting of 1,500 passages and 100 queries. Dense vector retrieval achieved Recall@10 of 0.9466, Recall@100 of 1.0000, MRR of 0.7898, and nDCG@10 of 0.8077. The weighted-RRF hybrid retrieval maintained comparable recall while improving MRR to 0.8219 and nDCG@10 to 0.8284. These results indicate that the proposed hybrid retrieval module improves ranking quality over dense retrieval alone under the given experimental setting. The end-to-end performance of the full PostgreSQL–Milvus pipeline, including propagation latency, failure recovery, cross-store consistency, and scalability, remains to be evaluated in future work.

Journal Copyright Policy

No CCL information provided

Citation status

* References for papers published after 2025 are currently being built.