AI Briefing
KO
Pick

Open-Source Leaderboard for Korean Embedding Models Released

·2026.07.26 09:30

Key point

An open-source embedding leaderboard has launched, enabling objective comparison of Korean RAG and retrieval performance.

Details

The ko-embedding-leaderboard project has been released to fairly compare Korean embedding models, which are essential when building RAG (Retrieval-Augmented Generation) systems.

This project customizes MTEB (Massive Text Embedding Benchmark) for the Korean-language environment, and has the following features.

  • Evaluation method: Rankings are calculated using the average of NDCG@5 and NDCG@10, leveraging 7 Korean retrieval datasets (Ko-StrategyQA, LawIRKo, etc.).
  • Categorization: Rankings are provided separately for Dense embeddings and Sparse embeddings based on model characteristics.
  • Key results (Dense): perplexity-ai/pplx-embed-v1-4b currently ranks 1st, with Korean-specialized models and multilingual general-purpose models evenly represented in the top ranks.
  • Key results (Sparse): telepix/PIXIE-Splade-v1.5 holds a top ranking.

This leaderboard is an open-source project that is continuously updated as new models are added.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.