Audio8-ASR-Infinite: 24-hour uninterrupted streaming ASR with adjustable latency
Edge0/Audio8-ASR-Infinite
About the project
This streaming model allows you to freely adjust speech recognition latency from 240ms to 560ms. You can select an audio clock of 80ms, 120ms, or 160ms to balance recognition speed and resource consumption.
By applying Rolling KV Cache technology, it maintains constant memory usage and latency. Despite a default context length of 30 seconds, it performs transcription continuously without interruption for 24 hours.
It features Semantic VAD, which distinguishes between pauses during speech, stuttering, and actual end-of-speech, cases that were difficult for existing acoustic VADs to differentiate. It supports Chinese and English, and enables real-time inference via vLLM.
Performance is ensured by combining the Voxtral Realtime architecture with the Qwen2.5-3B-Instruct decoder. It is distributed under the Apache 2.0 license and is currently a preview version that includes transcription functionality.
Edge0/Audio8-ASR-Infinite
The original page has no description.
automatic-speech-recognition
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.