Paper
Title: A Multimodal Spatio-Temporal GCN Model with Enhancements for Isolated Sign Recognition
Authors: Yang Zhou, Zhaoyang Xia, Yuxiao Chen, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.45/
Summary
A multimodal Spatio-Temporal GCN for isolated sign recognition (ISR). Builds on Yan et al. (2018) ST-GCN with 27 body/hand keypoints from AlphaPose, late-fuses forward and backward joint and bone streams, adds a per-channel Gating module (multilayer convolutions + temperature softmax) to weight informative frames, and pairs it with a transformer-based encoder branch consuming dominant/non-dominant start and end handshapes for temporal localization. Trained in PyTorch 1.7.0 on an NVIDIA Quadro RTX8000 with SGD + Nesterov momentum, batch 64, 200 epochs.
Results / Conclusions
- Combined isolated datasets: 79.98% Top-1 / 95.04% Top-5.
- Combined isolated + pre-segmented: 80.76% Top-1 / 95.18% Top-5.
- Pre-segmented only: 80.39% Top-1 / 92.96% Top-5.
- Beats Dafnis et al. 2022b on WLASL (79.59 vs 77.43 Top-1) and on combined WLASL+ASLLVD (81.56 vs 78.70 Top-1).
Code & Data
Suggested Reproduction
- Implement ST-GCN with the described 27-keypoint topology and the per-channel Gating module.
- Add the transformer branch over start/end handshapes.
- Train on WLASL / ASLLVD / RIT / DSP combinations as in the paper.
- Reproduce Top-1 / Top-5 numbers.
Bibtex
@inproceedings{zhou-etal-2024-multimodal,
title = "A Multimodal Spatio-Temporal {GCN} Model with Enhancements for Isolated Sign Recognition",
author = "Zhou, Yang and
Xia, Zhaoyang and
Chen, Yuxiao and
Neidle, Carol and
Metaxas, Dimitris N.",
editor = "Efthimiou, Eleni and
Fotinea, Stavroula-Evita and
Hanke, Thomas and
Hochgesang, Julie A. and
Mesch, Johanna and
Schulder, Marc",
booktitle = "Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.signlang-1.45/",
pages = "408--419"
}
Paper
Title: A Multimodal Spatio-Temporal GCN Model with Enhancements for Isolated Sign Recognition
Authors: Yang Zhou, Zhaoyang Xia, Yuxiao Chen, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.45/
Summary
A multimodal Spatio-Temporal GCN for isolated sign recognition (ISR). Builds on Yan et al. (2018) ST-GCN with 27 body/hand keypoints from AlphaPose, late-fuses forward and backward joint and bone streams, adds a per-channel Gating module (multilayer convolutions + temperature softmax) to weight informative frames, and pairs it with a transformer-based encoder branch consuming dominant/non-dominant start and end handshapes for temporal localization. Trained in PyTorch 1.7.0 on an NVIDIA Quadro RTX8000 with SGD + Nesterov momentum, batch 64, 200 epochs.
Results / Conclusions
Code & Data
Suggested Reproduction
Bibtex