Paper
Title: Preprocessing Mediapipe Keypoints with Keypoint Reconstruction and Anchors for Isolated Sign Language Recognition
Authors: Kyunggeun Roh, Huije Lee, Eui Jun Hwang, Sukmin Cho, Jong C. Park
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.36/
Summary
Trains a 4-layer Transformer encoder (and a SPOTER encoder-decoder for comparison) for isolated sign language recognition (ISLR) over MediaPipe pose features. Targets two failure modes: body-centric normalization that under-weights hand shape, and MediaPipe's >50% hand-undetection rate on WLASL. Three preprocessing components: (1) palm-anchor-based separate normalization for body and hands, (2) bilinear-interpolation hand keypoint reconstruction with average-pose initialization for empty first/last frames, (3) frame duplication to a fixed length of 512. Trained with Adam, lr 1e-5, 200 epochs on WLASL-100 and AUTSL.
Results / Conclusions
- WLASL-100: +6.05% accuracy, 83.26% top-1 — best among pose-based methods at the time.
- Ablations show each component contributes; gains are largest for handshape-dependent signs.
Code & Data
- No public code repository disclosed in the paper.
Suggested Reproduction
- Implement the three preprocessing components on top of MediaPipe pose extraction.
- Train a 4-layer Transformer encoder with the reported hyperparameters on WLASL-100 and AUTSL.
- Reproduce the +6.05% gain and 83.26% top-1 numbers, and the per-component ablation deltas.
Bibtex
@inproceedings{roh-etal-2024-preprocessing,
title = "Preprocessing Mediapipe Keypoints with Keypoint Reconstruction and Anchors for Isolated Sign Language Recognition",
author = "Roh, Kyunggeun and
Lee, Huije and
Hwang, Eui Jun and
Cho, Sukmin and
Park, Jong C.",
editor = "Efthimiou, Eleni and
Fotinea, Stavroula-Evita and
Hanke, Thomas and
Hochgesang, Julie A. and
Mesch, Johanna and
Schulder, Marc",
booktitle = "Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.signlang-1.36/",
pages = "323--334"
}
Paper
Title: Preprocessing Mediapipe Keypoints with Keypoint Reconstruction and Anchors for Isolated Sign Language Recognition
Authors: Kyunggeun Roh, Huije Lee, Eui Jun Hwang, Sukmin Cho, Jong C. Park
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.36/
Summary
Trains a 4-layer Transformer encoder (and a SPOTER encoder-decoder for comparison) for isolated sign language recognition (ISLR) over MediaPipe pose features. Targets two failure modes: body-centric normalization that under-weights hand shape, and MediaPipe's >50% hand-undetection rate on WLASL. Three preprocessing components: (1) palm-anchor-based separate normalization for body and hands, (2) bilinear-interpolation hand keypoint reconstruction with average-pose initialization for empty first/last frames, (3) frame duplication to a fixed length of 512. Trained with Adam, lr 1e-5, 200 epochs on WLASL-100 and AUTSL.
Results / Conclusions
Code & Data
Suggested Reproduction
Bibtex