Introducing IndQA
Key point
IndQA, a new benchmark for evaluating AI models' understanding of Indian culture and languages, has been released.
Details
Existing multilingual benchmarks such as MMMLU have already reached saturated scores, making it difficult to measure models' actual progress. In addition, evaluation methods centered on simple translation or multiple-choice questions have limitations in capturing genuine understanding of a language's context, culture, and history.
IndQA, developed to address this, evaluates how well AI models understand and reason about questions posed in Indian languages. India is a massive market with about 1 billion non-English-speaking users and 22 official languages, and it is ChatGPT's second-largest market.
The key features of IndQA are as follows:
- Covers 12 languages (including Hinglish) and 10 cultural domains
- 2,278 questions, with participation from 261 domain experts
- Covers a wide range of subjects including architecture, art, food, history, law, and religion
The evaluation method is based on expert-written rubrics. Each answer is scored by a model-based grader according to criteria set by experts.
To improve the quality of the questions, an adversarial filtering process was applied. Only high-difficulty questions that the latest models—GPT-4o, OpenAI o3, and GPT-4.5—failed to solve were selected, ensuring room to measure models' future progress.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.