NVIDIA Nemotron-Personas-Korea - A 1 Million-Record Synthetic Persona Dataset Based on South Korea's Actual Population Distribution
Key point
NVIDIA has released a synthetic persona dataset of 1 million records reflecting South Korea's population distribution.
Details
NVIDIA has released Nemotron-Personas-Korea. It is a synthetic persona dataset reflecting South Korea's actual demographic, geographic, and disposition distributions, and is notable for being structured to be easily readable in Korean.
- Composed of 1 million records, 7 million personas, and 26 fields.
- Contains various attributes including name, gender, age, marital status, education level, occupation, and region of residence.
- Built based on public data from KOSIS, the Supreme Court, the National Health Insurance Service, and the Korea Rural Economic Institute, among others.
- Synthesized using NeMo Data Designer and google/gemma-4-31B-it, and released under CC BY 4.0.
It also pointed out that Korean personas generated by existing LLMs deviate significantly from reality in aspects such as food preferences, occupation distribution, and regional distribution. Examples cited included biases such as excessive preference for salads and extreme skewing toward certain occupations.
This dataset was designed to reduce such biases. In particular, it was structured to more faithfully reflect elderly populations, rural areas, and education/occupation distributions, and it was explained that the dataset can be used as synthetic data for training and evaluating sovereign AI models.
In detail:
- Includes 17 metropolitan/provincial regions (si-do) and 252 city/county/district regions (si-gun-gu).
- Reflects over 209,000 unique name combinations.
- Age, marital status, household type, education, and occupation distributions were introduced as being similar to actual statistics.
- Also reflects characteristics of the Korean labor market, such as online shopping salespersons, building security guards, and building cleaners.
Additionally, this dataset claims to have higher realism than LLM-based synthetic personas in terms of city, household, housing, and food distributions. It stated that the name distribution reflects naming trends by generation since the 1940s, reducing anachronistic combinations.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.