PLOS ONE 2025 Sign LanguageGraph NetworksTransformers

SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks

Safaeid Hossain Arib1, Rabeya Akter1, Sejuti Rahman1, Shafin Rahman2

1Dept. of Robotics and Mechatronics Engineering, University of Dhaka   2Dept. of Electrical and Computer Engineering, North South University

PLOS ONE 20(2): e0316298, 2025 · Accepted at the WiML Workshop, NeurIPS 2025

TL;DR

Transformers over RGB capture context but miss the skeleton's graph structure. SignFormer-GCN adds a spatio-temporal graph stream over keypoints, so the model sees both what the scene looks like and how the body moves.

  • 19.75BLEU-4 on RWTH-PHOENIX-2014T, gloss-free
  • 8.53BLEU-4 on How2Sign test, best among compared methods
  • 9.43Mparameters, vs. 115.41M for GFSLT-VLP
  • 3sign languages: German, American and Bangla
Overview figure for SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks
SignFormer-GCN. (A) I3D features and keypoint features are encoded by a transformer encoder (B) and an STGCN-LSTM encoder (C), fused, and decoded into a spoken-language sentence.

Abstract

Sign language is a complex visual language system that uses hand gestures, facial expressions, and body movements to convey meaning. It is the primary means of communication for millions of deaf and hard-of-hearing individuals worldwide. Tracking physical actions, such as hand movements and arm orientation, alongside expressive actions, including facial expressions, mouth movements, eye movements, eyebrow gestures, head movements, and body postures, using only RGB features can be limiting due to discrepancies in backgrounds and signers across different datasets. Despite this limitation, most Sign Language Translation (SLT) research relies solely on RGB features. We used keypoint features, and RGB features to capture better the pose and configuration of body parts involved in sign language actions and complement the RGB features. Similarly, most works on SLT research have used transformers, which are good at capturing broader, high-level context and focusing on the most relevant video frames. Still, the inherent graph structure associated with sign language is neglected and fails to capture low-level details. To solve this, we used a joint encoding technique using a transformer and STGCN architecture to capture the context of sign language expressions and spatial and temporal dependencies on skeleton graphs. Our method, SignFormer-GCN, achieves competitive performance in RWTH-PHOENIX-2014T, How2Sign, and BornilDB v1.0 datasets experimentally, showcasing its effectiveness in enhancing translation accuracy through different sign languages.

Context and skeleton

Motivation

Sign language translation maps a video of signing directly to a spoken-language sentence. It is hard for two reasons. Signers move at different speeds, so videos vary in length. And frames do not line up one-to-one with words, because sign and spoken languages order meaning differently.

Most translators use only RGB video and a transformer. Transformers are good at broad, high-level context, but RGB features are sensitive to backgrounds and signer appearance, and they ignore the topology of the human body. Signing is, at its core, the motion of joints. A spatio-temporal skeleton graph captures exactly that, so SignFormer-GCN reads both.

Method

Two streams, one decoder

  1. Appearance streamAn I3D network embeds 16-frame clips. The embeddings get positional encodings and pass through a transformer encoder that models context across the whole video.
  2. Skeleton streamMediaPipe keypoints form a graph over joints and time. Stacked STGCN blocks model joint relationships and an LSTM models how they evolve.
  3. Fusion and decodingThe two encodings are fused by summation, and a transformer decoder generates the sentence directly, with no gloss supervision.

Given a video \(V = \{f_1, \dots, f_t\}\), the model learns \(P(S \mid V)\) for a sentence \(S = \{w_1, \dots, w_n\}\). The appearance stream embeds each clip and adds temporal position.

\[ E_t = P\big(F_{\mathrm{temporal}}(F_{\mathrm{I3D}}(V_{t:t+15}))\big), \qquad \hat{E}_t = E_t + E_{\mathrm{pos}}(t) \]

The whole model has 9.43M parameters and trains in about 2.5 minutes per epoch on a single RTX 3090.

Datasets

DatasetLanguageSignersHours (train)Vocabulary (train)Domain
RWTH-PHOENIX-2014TGerman (DGS)99.22KWeather forecasts
How2SignAmerican (ASL)1169.615.6KInstructional videos
BornilDB v1.0Bangla (BdSL)349.8614.1KNot specified

Results

Gloss-free translation · BLEU · higher is better

RWTH-PHOENIX-2014T (German)

MethodBLEU-1BLEU-2BLEU-3BLEU-4
Conv2d-RNN27.1015.6110.828.35
+ Luong attention29.8617.5211.969.00
+ Bahdanau attention32.2419.0312.839.58
Joint-SLT30.8818.5713.1210.19
Tokenization-SLT37.2223.8817.0813.25
TSPNet-Sequential35.6522.8016.6012.97
TSPNet-Joint36.1023.1216.8813.41
GASLT39.0726.7421.8615.74
GFSLT-VLP (115.41M params)43.7133.1826.1121.44
SignFormer-GCN (9.43M params)41.1930.8924.2319.75

SignFormer-GCN is competitive with the much larger GFSLT-VLP at about one-twelfth of the parameters, and ahead of every other gloss-free method.

How2Sign (American)

MethodVal rBLEUVal BLEU-1Val BLEU-4Test rBLEUTest BLEU-1Test BLEU-4
slt_how2sign2.7935.208.892.2134.018.03
asl_video2text3.2935.259.392.5633.207.95
SignFormer-GCN3.9737.379.902.9634.918.53

BornilDB v1.0 (Bangla)

SplitBLEU-1BLEU-2BLEU-3BLEU-4
Validation7.623.051.370.72
Test7.372.891.180.58

BornilDB v1.0 is a new, low-resource Bangla benchmark with only three signers, which makes it far harder than the other two datasets.

Example translations

DatasetReferencePrediction
How2Signthe other thing i did was i had a little bit of line out.another thing i would do is i have a little bit of line out.
How2Signjust enough to build up some chest strength.enough to build some chest strength.
PHOENIXguten abend liebe zuschauer (good evening dear viewers)guten abend liebe zuschauer (good evening dear viewers)
BornilDB(Is everything okay?)(Is everything fine?)
BornilDB(Didn't he know?)(Does he not know?)

Bangla examples are shown with the paper's English translations.

What matters

Ablations on RWTH-PHOENIX-2014T · test BLEU-4

</tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tbody> </table> </div>
  • The graph stream adds real signal. Removing the STGCN-LSTM encoder costs about a full BLEU-4 point.
  • Moderate depth works best. Three STGCN blocks and a single LSTM layer beat deeper variants.
  • Simple fusion wins. Summing the two streams outperformed fusing them with a linear layer or an LSTM in validation BLEU-4.

Future work

The next step is to narrow the semantic gap between the three representations involved (video, keypoints and text) so the model can learn even richer translations from sign language video.

VariantBLEU-4
Skeleton stream
Without STGCN-LSTM18.80
With STGCN-LSTM19.75
STGCN layers
1 / 218.87 / 19.26
319.75
4 / 519.06 / 19.16
Encoder–decoder layers
2–218.27
3–319.10
6–319.75
LSTM layers
119.75
2 / 319.31 / 18.89

Citation

@article{arib2025signformer,
  title   = {{SignFormer-GCN}: Continuous Sign Language Translation Using
             Spatio-Temporal Graph Convolutional Networks},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Rahman, Sejuti and Rahman, Shafin},
  journal = {PLOS ONE},
  volume  = {20},
  number  = {2},
  pages   = {e0316298},
  year    = {2025},
  doi     = {10.1371/journal.pone.0316298}
}