AI Briefing
KO

Translation Benchmark #1

·2026.04.14 19:36

Key point

TranslateGemma-12b ranked #1 overall across 6 LLM translation benchmarks.

Details

English subtitles were translated into Spanish, Japanese, Korean, Thai, Chinese (Simplified/Traditional), evaluating 167 segments per language pair.

Evaluation used reference-free QE metrics such as MetricX-24 (lower is better) and COMETKiwi (higher is better), along with a combined TQI = COMETKiwi × exp(−MetricX / 10).

  • Overall TQI #1: TranslateGemma-12b (0.6335)
  • #2: gemini-3.1-flash-lite-preview (0.5981)
  • #3: deepseek-v3.2 (0.5946)
  • #4: claude-sonnet-4-6 (0.5811)
  • #5: gpt-5.4-mini (0.5785)
  • #6: gpt-5.4-nano (0.5562)

All models' COMETKiwi scores were similar, at 0.75–0.79, but rankings diverged significantly due to large gaps in MetricX-24 fidelity scores.

However, the author noted that since MetricX-24 is a Google metric and TranslateGemma is also a Google model, the #1 gap may be somewhat overestimated.

Additionally, Claude Sonnet ranked lowest in Japanese, yet its COMETKiwi score was decent, revealing a pattern of good fluency but shaky meaning fidelity.

Gemini Flash Lite stayed in the upper ranks even compared to full-size frontier models, demonstrating the competitiveness of lightweight models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.