PLOS ONE 2025 Sign LanguageGraph NetworksTransformers

SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks

Safaeid Hossain Arib1, Rabeya Akter1, Sejuti Rahman1, Shafin Rahman2

1Dept. of Robotics and Mechatronics Engineering, University of Dhaka   2Dept. of Electrical and Computer Engineering, North South University

PLOS ONE 20(2): e0316298, 2025 · Accepted at the WiML Workshop, NeurIPS 2025

TL;DR

Transformers over RGB capture context but miss the skeleton's graph structure. SignFormer-GCN adds a spatio-temporal graph stream over keypoints, so the model sees both what the scene looks like and how the body moves.

  • 19.75BLEU-4 on RWTH-PHOENIX-2014T (gloss-free)
  • 8.53BLEU-4 on How2Sign test, best among compared methods
  • 9.43Mparameters, vs. 115.41M for GFSLT-VLP
Overview figure for SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks
SignFormer-GCN. (A) I3D features and keypoint features are encoded by a transformer encoder (B) and an STGCN-LSTM encoder (C), fused, and decoded into a spoken-language sentence.

Abstract

Sign language is a complex visual language system that uses hand gestures, facial expressions, and body movements to convey meaning. It is the primary means of communication for millions of deaf and hard-of-hearing individuals worldwide. Tracking physical actions, such as hand movements and arm orientation, alongside expressive actions, including facial expressions, mouth movements, eye movements, eyebrow gestures, head movements, and body postures, using only RGB features can be limiting due to discrepancies in backgrounds and signers across different datasets. Despite this limitation, most Sign Language Translation (SLT) research relies solely on RGB features. We used keypoint features, and RGB features to capture better the pose and configuration of body parts involved in sign language actions and complement the RGB features. Similarly, most works on SLT research have used transformers, which are good at capturing broader, high-level context and focusing on the most relevant video frames. Still, the inherent graph structure associated with sign language is neglected and fails to capture low-level details. To solve this, we used a joint encoding technique using a transformer and STGCN architecture to capture the context of sign language expressions and spatial and temporal dependencies on skeleton graphs. Our method, SignFormer-GCN, achieves competitive performance in RWTH-PHOENIX-2014T, How2Sign, and BornilDB v1.0 datasets experimentally, showcasing its effectiveness in enhancing translation accuracy through different sign languages.

Results

RWTH-PHOENIX-2014T · gloss-free · higher is better

MethodBLEU-1BLEU-2BLEU-3BLEU-4
Conv2d-RNN27.1015.6110.828.35
Joint-SLT30.8818.5713.1210.19
Tokenization-SLT37.2223.8817.0813.25
TSPNet-Joint36.1023.1216.8813.41
GASLT39.0726.7421.8615.74
GFSLT-VLP (115.41M params)43.7133.1826.1121.44
SignFormer-GCN (9.43M params)41.1930.8924.2319.75

SignFormer-GCN is competitive with the much larger GFSLT-VLP at roughly one-twelfth of the parameters.

How2Sign (American Sign Language)

MethodVal rBLEUVal BLEU-4Test rBLEUTest BLEU-4
slt_how2sign2.798.892.218.03
asl_video2text3.299.392.567.95
SignFormer-GCN3.979.902.968.53

What matters

  • The graph stream helps: adding the STGCN-LSTM encoder improves test BLEU-4 on PHOENIX-2014T from 18.80 to 19.75.
  • Simple fusion works best: fusing the two streams by summation outperformed a linear-layer or LSTM fusion.

Citation

@article{arib2025signformer,
  title   = {{SignFormer-GCN}: Continuous Sign Language Translation Using
             Spatio-Temporal Graph Convolutional Networks},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Rahman, Sejuti and Rahman, Shafin},
  journal = {PLOS ONE},
  volume  = {20},
  number  = {2},
  pages   = {e0316298},
  year    = {2025},
  doi     = {10.1371/journal.pone.0316298}
}