AI Briefing
KO

Automatic Generation of Sign Language Annotations Using Sign Language Models

·2026.04.30 09:00

Key point

A pseudo-annotation pipeline was proposed that automatically generates annotation candidates based on sign language video and English.

Details

A pseudo-annotation pipeline was proposed that takes sign language video and English as input and ranks gloss, fingerspelled word, and sign classifier candidates along with their time spans.

Large-scale sign language datasets such as ASL STEM Wiki and FLEURS-ASL span hundreds of hours but contain only partial annotations, leaving them underutilized. The pipeline combines the following signals to estimate annotation candidates:

  • Sparse predictions from a fingerspelling recognizer
  • Predictions from an isolated sign recognizer (ISR)
  • Calibration using a K-shot LLM

Baseline models were also presented. They achieved 6.7% CER on FSBoard and 74% top-1 accuracy on ASL Citizen, and professional interpreters added sequence-level gloss labels to about 500 videos from ASL STEM Wiki to create a gold-standard benchmark including gloss, classifier, and fingerspelling. The human annotations and over 300 hours of pseudo-annotations will be released as supplementary material.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.