AI Briefing
KO

OCR Cost Reversal

·2026.04.23 14:40

Key point

In a benchmark of 42 documents and 7,560 calls, cheaper/older models outperformed on OCR performance and cost efficiency.

Details

To verify the pattern of overspending on the latest, large models in OCR and document extraction workflows, 18 LLMs were compared under identical conditions.

  • 42 standard documents were selected, and each model was run 10 times each, for a total of 7,560 calls.
  • The evaluation criteria were pass^n (large-scale reliability), cost-per-success, latency, and critical field accuracy.
  • The key finding was that for standard OCR, smaller, older models achieved accuracy comparable to premium models while being much cheaper.

The entire dataset and framework were open-sourced, along with a separate mini-bench and leaderboard.

  • GitHub: ocr-mini-bench
  • Leaderboard: arbitrhq.ai/leaderboards/

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.