Author ORCID Identifier
Document Type
Thesis
Date of Award
2026
Degree Name
Master of Science (MS)
Department
Computer Science
First Advisor
Lina S Chato
Abstract
Automatic sign language translation remains challenging because models must recover linguistic structure from complex visual motion while relying on limited annotated data. This thesis investigates whether compact skeletal key-points can serve as an effective alternative to RGB video for continuous sign language recognition and translation, motivated by the privacy, computational efficiency, and transferability of skeletal representations. To address this problem, a segmented key-point-based framework is evaluated in continuous sign language recognition (Sign2Gloss), gloss-mediated translation (Sign2Gloss2Text), gloss-free translation (Sign2Text), and transfer learning within a common modeling pipeline. The proposed framework is a motion-aware multi-stream transformer encoder. Rather than processing the signer as a single feature sequence, the model separates pose, left-hand, right-hand, and facial landmarks into anatomically meaningful streams, augments each stream with velocity and acceleration features, and integrates them through cross-stream attention. This enables the encoder to capture complementary articulatory information while preserving the compactness and privacy advantages of skeletal representations. Experiments conducted on PHOENIX-2014-T and How2Sign show that the proposed architecture consistently outperforms a flat key-point encoder, reducing sign recognition word error rate from 28.17% to 26.20%, improving gloss-mediated translation from 18.27 to 20.36 BLEU-4, improving gloss-free translation from 13.82 to 14.38 BLEU-4 on PHOENIX-2014-T, and increasing How2Sign translation performance from 2.17 to 2.98 BLEU-4 when trained from scratch, while maintaining real-time inference performance. Transfer learning experiments were conducted to investigate the transferability of key-point-based sign language representations across datasets and languages. Cross-lingual transfer from PHOENIX-2014-T to How2Sign, reduced Keyword Word Error Rate (KW-WER) from 0.68 to 0.66 on How2Sign, suggesting that the encoder learned articulatory and motion patterns that generalize across sign languages. This improvement did not translate into meaningful gains in translation performance, with BLEU-4 remaining essentially unchanged, indicating that transferable low-level motion representations alone are insufficient to overcome domain-specific differences. In contrast, large-scale in-domain pre-training on YouTube-ASL followed by fine-tuning on How2Sign reduced KW-WER to 0.55 and improved BLEU-4 from 2.98 to 5.63. These findings demonstrate that anatomically structured key-point representations can learn transferable sign language features and improve performance in sign language recognition and translation systems.
Subject Categories
Computer Sciences
Keywords
key-point representation, motion modeling, multi-stream transformer, sign language, recognition, sign language translation, transfer learning
Number of Pages
101
Publisher
University of South Dakota
Recommended Citation
Kagozi, Alex Muchiri, "KEY-POINT-BASED SIGN LANGUAGE TRANSLATION WITH MULTI-STREAM MOTION MODELING AND CROSS-LINGUAL TRANSFER" (2026). Dissertations and Theses. 450.
https://red.library.usd.edu/diss-thesis/450