AI Briefing
KO

Interfaze: A New Model Architecture Delivering High Accuracy at Scale

·2026.05.12 01:22

Key point

Interfaze has unveiled a new model architecture and API for OCR, STT, and structured output.

1 / 2

Details

Interfaze presents a hybrid architecture combining DNN/CNN + transformer decoder, explaining that it is tailored for deterministic tasks such as OCR, object/GUI detection, web extraction/search, STT, and translation.

  • Specs: context window 1M tokens, max output 32k tokens, input supports text/image/audio/file, with reasoning disabled by default.
  • Own benchmarks: it published comparison results against Gemini-3-Flash, Claude-Sonnet-4.6, GPT-5.4-Mini, and Grok-4.3 on OCRBench V2, olmOCR, RefCOCO, VoxPopuli-Cleaned-AA, Spider 2.0-Lite, GPQA Diamond, MMMLU, MMMU-Pro, and SOB.
  • Pricing: set at input $1.50 / 1M tokens, output $3.50 / 1M tokens.
  • OCR: it explains that it returns pixel coordinates for both text and graphical objects together from complex PDFs and images, and even exposes raw OCR precontext and confidence.
  • STT: it presented WER 2.4% on VoxPopuli-Cleaned-AA, and stated it processes 209 seconds of audio per second.
  • Integration: it can be used by connecting an OpenAI Chat Completions-compatible API at https://api.interfaze.ai/v1.

SOB is explained as a benchmark that measures the model's ability to accurately fill in values in a JSON structure when the correct answer is already given in the context.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.