Researchers Release 'Turnbench', a Turn-Taking Evaluation Benchmark by Dialogue Type
Key point
Researchers have released 'Turnbench', a turn-taking evaluation benchmark by dialogue type featuring 30 hours of high-quality data and a standardized evaluation protocol.
Details
Researchers have released 'Turnbench', a multi-domain turn-taking benchmark based on dialogue analysis. This benchmark includes 30 hours of manually labeled human dialogue corpora and 104 hours of training data, providing reproducible protocols for end-of-turn (EOT) detection and interruption detection. Performance evaluation results showed that the VAP model achieved high accuracy with an EOT recall of 0.845 and an interruption recall of 0.945, with an EOT latency of 368ms and an interruption latency of 994ms. In contrast, Gemini recorded the lowest FPR (0.01–0.03) but demonstrated conservative performance with a recall of 0.61–0.71.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.