DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Key point
Apple researchers released DiscoSign, a sign language translation framework that incorporates context and spatial consistency.
Details
Existing sign language processing systems operate only at the sentence level, ignoring discourse phenomena essential for sign language understanding. Apple researchers introduced DiscoSign, a framework based on linguistic research, to introduce context awareness when translating text into sign language gloss.
This approach uses an LLM (Large Language Model)-based modular translation framework and addresses the following three core phenomena:
- Spatial Deixis Resolution: Processing entities to maintain consistent spatial positions throughout the discourse
- Question-Answer Chunks (QAC): Handling quasi-parallel structures that perform specific discourse functions
- Concept-Gloss Consistency: Ensuring stable mapping between English concepts and American Sign Language (ASL) notation
Since existing translation metrics do not evaluate discourse-level quality, the researchers introduced a new set of evaluation metrics to assess discourse cohesion across each dimension. Experimental results show that DiscoSign significantly improves spatial consistency and entity tracking performance compared to sentence-level translation, while maintaining competitive single-sentence gloss translation quality. This study presents the first systematic framework and evaluation method for discourse-level text-to-sign language gloss translation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.