Author ORCID Identifier

https://orcid.org//0009-0006-5481-1636

Document Type

Thesis

Date of Award

2026

Degree Name

Master of Science (MS)

Department

Computer Science

First Advisor

Lina S Chato

Abstract

Automatic sign language translation remains challenging because models must recover linguistic structure from complex visual motion while relying on limited annotated data. This thesis investigates whether compact skeletal key-points can serve as an effective alternative to RGB video for continuous sign language recognition and translation, motivated by the privacy, computational efficiency, and transferability of skeletal representations. To address this problem, a segmented key-point-based framework is evaluated in continuous sign language recognition (Sign2Gloss), gloss-mediated translation (Sign2Gloss2Text), gloss-free translation (Sign2Text), and transfer learning within a common modeling pipeline. The proposed framework is a motion-aware multi-stream transformer encoder. Rather than processing the signer as a single feature sequence, the model separates pose, left-hand, right-hand, and facial landmarks into anatomically meaningful streams, augments each stream with velocity and acceleration features, and integrates them through cross-stream attention. This enables the encoder to capture complementary articulatory information while preserving the compactness and privacy advantages of skeletal representations. Experiments conducted on PHOENIX-2014-T and How2Sign show that the proposed architecture consistently outperforms a flat key-point encoder, reducing sign recognition word error rate from 28.17% to 26.20%, improving gloss-mediated translation from 18.27 to 20.36 BLEU-4, improving gloss-free translation from 13.82 to 14.38 BLEU-4 on PHOENIX-2014-T, and increasing How2Sign translation performance from 2.17 to 2.98 BLEU-4 when trained from scratch, while maintaining real-time inference performance. Transfer learning experiments were conducted to investigate the transferability of key-point-based sign language representations across datasets and languages. Cross-lingual transfer from PHOENIX-2014-T to How2Sign, reduced Keyword Word Error Rate (KW-WER) from 0.68 to 0.66 on How2Sign, suggesting that the encoder learned articulatory and motion patterns that generalize across sign languages. This improvement did not translate into meaningful gains in translation performance, with BLEU-4 remaining essentially unchanged, indicating that transferable low-level motion representations alone are insufficient to overcome domain-specific differences. In contrast, large-scale in-domain pre-training on YouTube-ASL followed by fine-tuning on How2Sign reduced KW-WER to 0.55 and improved BLEU-4 from 2.98 to 5.63. These findings demonstrate that anatomically structured key-point representations can learn transferable sign language features and improve performance in sign language recognition and translation systems.

Subject Categories

Computer Sciences

Keywords

key-point representation, motion modeling, multi-stream transformer, sign language, recognition, sign language translation, transfer learning

Number of Pages

101

Publisher

University of South Dakota

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.