Toss Builds 'Toss Benchmark' to Bridge Gap Between Public Leaderboards and Real-World Performance
Key point
Toss analyzed the discrepancy between public benchmarks and actual service performance, establishing Toss Benchmark as a domain-specific evaluation standard.
Details
The Toss LLM Modeling team built Toss Benchmark to select models suitable for real-world work environments. Scores on existing public leaderboards were skewed toward English proficiency or high scores in Reasoning mode, creating a gap with performance in operational environments requiring Korean policy understanding or fast responses.
Ranking Reversals Based on Korean Language and Reasoning Settings
Changing the prompt language from English to Korean or disabling the Reasoning feature caused significant fluctuations in rankings between models. In math and coding evaluations, rankings reversed when comparing Korean translations to English originals, and the extent of performance drop when Reasoning was Off varied by model. Notably, GLM-5.1 dropped by 30.6 points when Reasoning was Off, whereas gemma-4-26B-A4B-it dropped by only 6.0 points, revealing differences in stability for operational environments.
Domain-Specific Evaluation: Toss IFBench and Knowledge
To evaluate actual work processing capabilities, Toss IFBench (policy compliance) and Toss Knowledge (domain knowledge) were introduced. Results from Toss IFBench showed that format compliance and judgment flexibility are distinct capabilities, while Toss Knowledge measures commerce domain brands, terminology, and product specifications through multiple-choice questions.
Correlation Between Benchmarks and Actual Operational Tasks
Analysis of 11 models revealed a strong positive correlation (Pearson +0.87) between the format instruction scores of Toss IFBench and performance on actual operational tasks. Additionally, higher Toss Knowledge scores correlated with better performance on detailed commerce tasks such as product classification, grouping, and inspection. Toss plans to use this data as the basis for model recommendations in its internal GenAI Portal and as a criterion for verifying improvements before and after training for Toss Foundation model development.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.