Paper
Title: Diffusion Models for Sign Language Video Anonymization
Authors: Zhaoyang Xia, Yang Zhou, Ligong Han, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.44/
Summary
Introduces DiffSLVA, text-guided sign-language video anonymization. Given a source signing video and a text prompt describing a target identity/style, DiffSLVA generates a new video where the original signer's appearance (body, clothing, ethnicity, gender) is replaced while preserving the linguistic content. Uses pretrained large-scale text-guided latent diffusion models with ControlNet conditioned on Holistically-Nested Edge (HED) maps to circumvent the pose-estimation bottleneck of prior anonymization methods. Adds a dedicated facial expression enhancement module to preserve linguistically essential non-manual features. Fine-tunes a state-of-the-art image animation model (Zhao & Zhang, 2022) on a mixed dataset for the facial enhancement module. Uses cross-frame attention and two-stage optical-flow-guided latent fusion (following Yang et al. 2023, Rerender-A-Video) for temporal consistency.
Results / Conclusions
- Demonstrates text-prompted appearance transfer with preserved sign content (qualitative results in the paper).
- Avoids accurate pose estimation by using HED edges, robust to motion blur and occlusion.
Code & Data
Suggested Reproduction
- Clone DiffSLVA repo and install dependencies.
- Run end-to-end on a sample sign-language video with a text prompt.
- Compare anonymized output qualitatively against AnonySign and SLVA baselines.
- Validate with sign-recognition back-translation to check linguistic preservation.
Bibtex
@inproceedings{xia-etal-2024-diffusion,
title = "Diffusion Models for Sign Language Video Anonymization",
author = "Xia, Zhaoyang and
Zhou, Yang and
Han, Ligong and
Neidle, Carol and
Metaxas, Dimitris N.",
editor = "Efthimiou, Eleni and
Fotinea, Stavroula-Evita and
Hanke, Thomas and
Hochgesang, Julie A. and
Mesch, Johanna and
Schulder, Marc",
booktitle = "Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.signlang-1.44/",
pages = "395--407"
}
Paper
Title: Diffusion Models for Sign Language Video Anonymization
Authors: Zhaoyang Xia, Yang Zhou, Ligong Han, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.44/
Summary
Introduces DiffSLVA, text-guided sign-language video anonymization. Given a source signing video and a text prompt describing a target identity/style, DiffSLVA generates a new video where the original signer's appearance (body, clothing, ethnicity, gender) is replaced while preserving the linguistic content. Uses pretrained large-scale text-guided latent diffusion models with ControlNet conditioned on Holistically-Nested Edge (HED) maps to circumvent the pose-estimation bottleneck of prior anonymization methods. Adds a dedicated facial expression enhancement module to preserve linguistically essential non-manual features. Fine-tunes a state-of-the-art image animation model (Zhao & Zhang, 2022) on a mixed dataset for the facial enhancement module. Uses cross-frame attention and two-stage optical-flow-guided latent fusion (following Yang et al. 2023, Rerender-A-Video) for temporal consistency.
Results / Conclusions
Code & Data
Suggested Reproduction
Bibtex