SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks
1Dept. of Robotics and Mechatronics Engineering, University of Dhaka 2Dept. of Electrical and Computer Engineering, North South University
PLOS ONE 20(2): e0316298, 2025 · Accepted at the WiML Workshop, NeurIPS 2025
TL;DR
Transformers over RGB capture context but miss the skeleton's graph structure. SignFormer-GCN adds a spatio-temporal graph stream over keypoints, so the model sees both what the scene looks like and how the body moves.
- 19.75BLEU-4 on RWTH-PHOENIX-2014T, gloss-free
- 8.53BLEU-4 on How2Sign test, best among compared methods
- 9.43Mparameters, vs. 115.41M for GFSLT-VLP
- 3sign languages: German, American and Bangla

Abstract
Sign language is a complex visual language system that uses hand gestures, facial expressions, and body movements to convey meaning. It is the primary means of communication for millions of deaf and hard-of-hearing individuals worldwide. Tracking physical actions, such as hand movements and arm orientation, alongside expressive actions, including facial expressions, mouth movements, eye movements, eyebrow gestures, head movements, and body postures, using only RGB features can be limiting due to discrepancies in backgrounds and signers across different datasets. Despite this limitation, most Sign Language Translation (SLT) research relies solely on RGB features. We used keypoint features, and RGB features to capture better the pose and configuration of body parts involved in sign language actions and complement the RGB features. Similarly, most works on SLT research have used transformers, which are good at capturing broader, high-level context and focusing on the most relevant video frames. Still, the inherent graph structure associated with sign language is neglected and fails to capture low-level details. To solve this, we used a joint encoding technique using a transformer and STGCN architecture to capture the context of sign language expressions and spatial and temporal dependencies on skeleton graphs. Our method, SignFormer-GCN, achieves competitive performance in RWTH-PHOENIX-2014T, How2Sign, and BornilDB v1.0 datasets experimentally, showcasing its effectiveness in enhancing translation accuracy through different sign languages.
Context and skeleton
Motivation
Sign language translation maps a video of signing directly to a spoken-language sentence. It is hard for two reasons. Signers move at different speeds, so videos vary in length. And frames do not line up one-to-one with words, because sign and spoken languages order meaning differently.
Most translators use only RGB video and a transformer. Transformers are good at broad, high-level context, but RGB features are sensitive to backgrounds and signer appearance, and they ignore the topology of the human body. Signing is, at its core, the motion of joints. A spatio-temporal skeleton graph captures exactly that, so SignFormer-GCN reads both.
Method
Two streams, one decoder
- Appearance streamAn I3D network embeds 16-frame clips. The embeddings get positional encodings and pass through a transformer encoder that models context across the whole video.
- Skeleton streamMediaPipe keypoints form a graph over joints and time. Stacked STGCN blocks model joint relationships and an LSTM models how they evolve.
- Fusion and decodingThe two encodings are fused by summation, and a transformer decoder generates the sentence directly, with no gloss supervision.
Given a video \(V = \{f_1, \dots, f_t\}\), the model learns \(P(S \mid V)\) for a sentence \(S = \{w_1, \dots, w_n\}\). The appearance stream embeds each clip and adds temporal position.
The whole model has 9.43M parameters and trains in about 2.5 minutes per epoch on a single RTX 3090.
Datasets
| Dataset | Language | Signers | Hours (train) | Vocabulary (train) | Domain |
|---|---|---|---|---|---|
| RWTH-PHOENIX-2014T | German (DGS) | 9 | 9.2 | 2K | Weather forecasts |
| How2Sign | American (ASL) | 11 | 69.6 | 15.6K | Instructional videos |
| BornilDB v1.0 | Bangla (BdSL) | 3 | 49.86 | 14.1K | Not specified |
Results
Gloss-free translation · BLEU · higher is better
RWTH-PHOENIX-2014T (German)
| Method | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 |
|---|---|---|---|---|
| Conv2d-RNN | 27.10 | 15.61 | 10.82 | 8.35 |
| + Luong attention | 29.86 | 17.52 | 11.96 | 9.00 |
| + Bahdanau attention | 32.24 | 19.03 | 12.83 | 9.58 |
| Joint-SLT | 30.88 | 18.57 | 13.12 | 10.19 |
| Tokenization-SLT | 37.22 | 23.88 | 17.08 | 13.25 |
| TSPNet-Sequential | 35.65 | 22.80 | 16.60 | 12.97 |
| TSPNet-Joint | 36.10 | 23.12 | 16.88 | 13.41 |
| GASLT | 39.07 | 26.74 | 21.86 | 15.74 |
| GFSLT-VLP (115.41M params) | 43.71 | 33.18 | 26.11 | 21.44 |
| SignFormer-GCN (9.43M params) | 41.19 | 30.89 | 24.23 | 19.75 |
SignFormer-GCN is competitive with the much larger GFSLT-VLP at about one-twelfth of the parameters, and ahead of every other gloss-free method.
How2Sign (American)
| Method | Val rBLEU | Val BLEU-1 | Val BLEU-4 | Test rBLEU | Test BLEU-1 | Test BLEU-4 |
|---|---|---|---|---|---|---|
| slt_how2sign | 2.79 | 35.20 | 8.89 | 2.21 | 34.01 | 8.03 |
| asl_video2text | 3.29 | 35.25 | 9.39 | 2.56 | 33.20 | 7.95 |
| SignFormer-GCN | 3.97 | 37.37 | 9.90 | 2.96 | 34.91 | 8.53 |
BornilDB v1.0 (Bangla)
| Split | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 |
|---|---|---|---|---|
| Validation | 7.62 | 3.05 | 1.37 | 0.72 |
| Test | 7.37 | 2.89 | 1.18 | 0.58 |
BornilDB v1.0 is a new, low-resource Bangla benchmark with only three signers, which makes it far harder than the other two datasets.
Example translations
| Dataset | Reference | Prediction |
|---|---|---|
| How2Sign | the other thing i did was i had a little bit of line out. | another thing i would do is i have a little bit of line out. |
| How2Sign | just enough to build up some chest strength. | enough to build some chest strength. |
| PHOENIX | guten abend liebe zuschauer (good evening dear viewers) | guten abend liebe zuschauer (good evening dear viewers) |
| BornilDB | (Is everything okay?) | (Is everything fine?) |
| BornilDB | (Didn't he know?) | (Does he not know?) |
Bangla examples are shown with the paper's English translations.
What matters
Ablations on RWTH-PHOENIX-2014T · test BLEU-4
| Variant | BLEU-4 | |
|---|---|---|
| Skeleton stream | ||
| Without STGCN-LSTM | 18.80 | </tr>|
| With STGCN-LSTM | 19.75 | </tr>|
| STGCN layers | ||
| 1 / 2 | 18.87 / 19.26 | </tr>|
| 3 | 19.75 | </tr>|
| 4 / 5 | 19.06 / 19.16 | </tr>|
| Encoder–decoder layers | ||
| 2–2 | 18.27 | </tr>|
| 3–3 | 19.10 | </tr>|
| 6–3 | 19.75 | </tr>|
| LSTM layers | ||
| 1 | 19.75 | </tr>|
| 2 / 3 | 19.31 / 18.89 | </tr> </tbody> </table> </div>|
Citation
@article{arib2025signformer,
title = {{SignFormer-GCN}: Continuous Sign Language Translation Using
Spatio-Temporal Graph Convolutional Networks},
author = {Arib, Safaeid Hossain and Akter, Rabeya and Rahman, Sejuti and Rahman, Shafin},
journal = {PLOS ONE},
volume = {20},
number = {2},
pages = {e0316298},
year = {2025},
doi = {10.1371/journal.pone.0316298}
}