Skip to content

Reproduce: Xia et al. 2024 — DiffSLVA Diffusion-Based Sign Language Anonymization #26

Description

@AmitMY

Paper

Title: Diffusion Models for Sign Language Video Anonymization
Authors: Zhaoyang Xia, Yang Zhou, Ligong Han, Carol Neidle, Dimitris N. Metaxas
Venue: SignLang 2024 (LREC-COLING 2024 workshop)
ACL Anthology: https://aclanthology.org/2024.signlang-1.44/

Summary

Introduces DiffSLVA, text-guided sign-language video anonymization. Given a source signing video and a text prompt describing a target identity/style, DiffSLVA generates a new video where the original signer's appearance (body, clothing, ethnicity, gender) is replaced while preserving the linguistic content. Uses pretrained large-scale text-guided latent diffusion models with ControlNet conditioned on Holistically-Nested Edge (HED) maps to circumvent the pose-estimation bottleneck of prior anonymization methods. Adds a dedicated facial expression enhancement module to preserve linguistically essential non-manual features. Fine-tunes a state-of-the-art image animation model (Zhao & Zhang, 2022) on a mixed dataset for the facial enhancement module. Uses cross-frame attention and two-stage optical-flow-guided latent fusion (following Yang et al. 2023, Rerender-A-Video) for temporal consistency.

Results / Conclusions

  • Demonstrates text-prompted appearance transfer with preserved sign content (qualitative results in the paper).
  • Avoids accurate pose estimation by using HED edges, robust to motion blur and occlusion.

Code & Data

Suggested Reproduction

  1. Clone DiffSLVA repo and install dependencies.
  2. Run end-to-end on a sample sign-language video with a text prompt.
  3. Compare anonymized output qualitatively against AnonySign and SLVA baselines.
  4. Validate with sign-recognition back-translation to check linguistic preservation.

Bibtex

@inproceedings{xia-etal-2024-diffusion,
    title = "Diffusion Models for Sign Language Video Anonymization",
    author = "Xia, Zhaoyang  and
      Zhou, Yang  and
      Han, Ligong  and
      Neidle, Carol  and
      Metaxas, Dimitris N.",
    editor = "Efthimiou, Eleni  and
      Fotinea, Stavroula-Evita  and
      Hanke, Thomas  and
      Hochgesang, Julie A.  and
      Mesch, Johanna  and
      Schulder, Marc",
    booktitle = "Proceedings of the LREC-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.signlang-1.44/",
    pages = "395--407"
}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions