Skip to content

Reproduce: Zhou et al. 2024 — Multimodal ST-GCN for Isolated Sign Recognition #27

Description

@AmitMY

Paper

Title: A Multimodal Spatio-Temporal GCN Model with Enhancements for Isolated Sign Recognition
Authors: Yang Zhou, Zhaoyang Xia, Yuxiao Chen, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.45/

Summary

A multimodal Spatio-Temporal GCN for isolated sign recognition (ISR). Builds on Yan et al. (2018) ST-GCN with 27 body/hand keypoints from AlphaPose, late-fuses forward and backward joint and bone streams, adds a per-channel Gating module (multilayer convolutions + temperature softmax) to weight informative frames, and pairs it with a transformer-based encoder branch consuming dominant/non-dominant start and end handshapes for temporal localization. Trained in PyTorch 1.7.0 on an NVIDIA Quadro RTX8000 with SGD + Nesterov momentum, batch 64, 200 epochs.

Results / Conclusions

  • Combined isolated datasets: 79.98% Top-1 / 95.04% Top-5.
  • Combined isolated + pre-segmented: 80.76% Top-1 / 95.18% Top-5.
  • Pre-segmented only: 80.39% Top-1 / 92.96% Top-5.
  • Beats Dafnis et al. 2022b on WLASL (79.59 vs 77.43 Top-1) and on combined WLASL+ASLLVD (81.56 vs 78.70 Top-1).

Code & Data

Suggested Reproduction

  1. Implement ST-GCN with the described 27-keypoint topology and the per-channel Gating module.
  2. Add the transformer branch over start/end handshapes.
  3. Train on WLASL / ASLLVD / RIT / DSP combinations as in the paper.
  4. Reproduce Top-1 / Top-5 numbers.

Bibtex

@inproceedings{zhou-etal-2024-multimodal,
    title = "A Multimodal Spatio-Temporal {GCN} Model with Enhancements for Isolated Sign Recognition",
    author = "Zhou, Yang  and
      Xia, Zhaoyang  and
      Chen, Yuxiao  and
      Neidle, Carol  and
      Metaxas, Dimitris N.",
    editor = "Efthimiou, Eleni  and
      Fotinea, Stavroula-Evita  and
      Hanke, Thomas  and
      Hochgesang, Julie A.  and
      Mesch, Johanna  and
      Schulder, Marc",
    booktitle = "Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.signlang-1.45/",
    pages = "408--419"
}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions