Confucius4-R2T2: Real-time speech recognition with 200ms latency for confirmed text output
netease-youdao/Confucius4-R2T2
About the project
It adopts an 'append-only' approach where already output text is not modified during speech recognition. Through adjustable decoding chunks ranging from 80ms to 2 seconds, it guarantees low latency averaging 200–600ms.

A fine-tuned model based on Qwen3-ASR, it applies the Longest Stable Prefix learning paradigm to immediately emit only stable prefixes. This enables immediate data processing without text flickering in real-time subtitle generation or LLM agent pipelines.
It maintains offline recognition accuracy even with added streaming capabilities, achieving SOTA-level performance in English and Chinese. It supports the vLLM backend, running stably in environments requiring high throughput.
netease-youdao/Confucius4-R2T2
The original page has no description.
automatic-speech-recognition
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.