Kakao Reveals Behind-the-Scenes of 'AI TOP 100' Exam Design to Verify AI Utilization Capabilities
Key point
The article introduces an exam design strategy that ensures discrimination and blocks AI's 'click-to-solve' approach through Human-in-the-loop collaboration and dataset personalization.
Details
The exam design process for the AI TOP 100 competition hosted by Kakao has been revealed. This competition focused on measuring 'human capability' to solve real-world problems using AI, rather than the performance of AI models themselves.
Exam Principles: Evaluating Process and Collaboration
The exam designers adopted Human-in-the-loop as a core principle, rather than focusing on simple deliverables. They verified whether the process of human analysis, AI problem-solving, and human verification was functioning, and intentionally embedded technical limitations to avoid the 'paradox of clarity'—problems that AI could solve instantly. Additionally, they aimed for an open structure that allowed diverse approaches without forcing specific solutions.
Ensuring Fairness and Difficulty Design
To fundamentally prevent cheating among participants, a dataset personalization strategy was adopted. By providing data with different numerical values or proper nouns for each participant, answer sharing was rendered ineffective. To adjust difficulty, a multi-item structure of Easy, Medium, and Hard was applied per question, with fine-tuning achieved through randomizing multiple-choice options and providing verification points for subjective questions.
Verification Process and Response to Latest Models
The exam design process went through a 4-stage pipeline from idea to final confirmation, repeating a 3-stage validation procedure including internal testing by examiners, alpha testing, and private beta testing. Notably, although Gemini 3 was released three days before the finals, it was confirmed that the designed killer questions still required human analysis and verification despite the performance improvements of the latest models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.