DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Apple Machine Learning Research
DiscoSign introduces a discourse-aware text-to–sign-language gloss translation framework that targets discourse phenomena ignored by sentence-level systems. It improves spatial consistency and entity tracking on experiments using sentence-level and discourse-level datasets. The method adds modules for spatial coreference, question-answer clause handling, and concept-gloss consistency plus discourse-level evaluation metrics that replace sentence-only quality measures.
Why it matters
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving…