Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,8 @@ For more information, demos, and examples, please visit our [Project Page](https
| VibeVoice-ASR-7B | [HF Link](https://huggingface.co/microsoft/VibeVoice-ASR) | [Playground](https://aka.ms/vibevoice-asr) |
| VibeVoice-ASR-Streaming | [HF Link](https://huggingface.co/collections/microsoft/vibevoice-68a2ef24a875c44be47b034f) | [Documentation](docs/vibevoice-asr-streaming.md) |
| VibeVoice-ASR-BitNet (CPU) | [HF Link](https://huggingface.co/microsoft/VibeVoice-ASR-BitNet) | [VibeASR.cpp](https://github.com/microsoft/VibeASR.cpp) |
| VibeVoice-TTS-1.5B | [HF Link](https://huggingface.co/microsoft/VibeVoice-1.5B) | Disabled |
| VibeVoice-TTS-1.5B (legacy source) | [HF Link](https://huggingface.co/microsoft/VibeVoice-1.5B) | Disabled |
| VibeVoice-TTS-1.5B (Transformers-native) | [HF Link](https://huggingface.co/microsoft/VibeVoice-1.5B-HF) | [Publication workflow](docs/publish-vibevoice-1.5b-hf.md) |
| VibeVoice-Realtime-0.5B | [HF Link](https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B) | [Colab](https://colab.research.google.com/github/microsoft/VibeVoice/blob/main/demo/vibevoice_realtime_colab.ipynb) |

</div>
Expand Down
82 changes: 82 additions & 0 deletions docs/publish-vibevoice-1.5b-hf.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
# Publishing VibeVoice 1.5B for Transformers

This release path creates `microsoft/VibeVoice-1.5B-HF`, the official
Transformers-native artifact for the original TTS checkpoint. The original
[`microsoft/VibeVoice-1.5B`](https://huggingface.co/microsoft/VibeVoice-1.5B)
remains the legacy source/provenance repository; link it to `-HF` after the
owner publication rather than replacing its main branch.

## Pinned inputs

| Input | Pinned revision | Use |
| --- | --- | --- |
| [`microsoft/VibeVoice-1.5B`](https://huggingface.co/microsoft/VibeVoice-1.5B) | `c00898d257e6b46004e3e2866a47534085fb685a` | Official weight source |
| [Transformers](https://github.com/huggingface/transformers) | `640a08a597034221ca1c4fc0c129cf0118179225` | Canonical VibeVoice converter and native implementation |
| [`Qwen/Qwen2.5-1.5B`](https://huggingface.co/Qwen/Qwen2.5-1.5B) | `8faed761d45a263340a0528343f099c05c9a4323` | Canonical tokenizer input |
| [`vibevoice/VibeVoice-1.5B-hf`](https://huggingface.co/vibevoice/VibeVoice-1.5B-hf) | `edc39f80f5cae656da37baf8faa8f5502bf7081f` | Independent 1,204-key layout evidence only |

The Transformers commit contains `VibeVoiceForConditionalGeneration`,
`VibeVoiceProcessor`, `AutoModelForTextToWaveform` registration, and the
canonical `convert_vibevoice_to_hf.py` converter. It postdates the v5.16.1
release, so use the pinned source commit rather than a released package.

The sidecar reference above is never downloaded by the converter or used at
runtime. The converter downloads only the official Microsoft checkpoint and
the pinned Qwen tokenizer, invokes the canonical Transformers converter, and
fails unless all 1,204 source-to-native keys, shapes, and dtypes agree.

## Owner conversion and publication

Run this in a clean Python environment with sufficient disk and memory for the
multi-GB checkpoint:

```bash
git clone https://github.com/huggingface/transformers.git /tmp/transformers
git -C /tmp/transformers checkout 640a08a597034221ca1c4fc0c129cf0118179225
python -m pip install -e "/tmp/transformers[torch]"

python tools/release/convert_vibevoice_1_5b_hf.py \
--transformers-source /tmp/transformers \
--output-dir /tmp/VibeVoice-1.5B-HF
```

The output includes native safetensor shards, config, generation config,
tokenizer, processor, chat template, model card, and
`conversion-manifest.json`. The manifest records every pinned input and the
strict tensor-alignment digest, including the clean VibeVoice release-tool
commit. The script has no upload option.

Only an owner authorized for the Microsoft Hugging Face namespace should
publish the reviewed local artifact:

```bash
huggingface-cli upload microsoft/VibeVoice-1.5B-HF /tmp/VibeVoice-1.5B-HF . \
--commit-message "Publish Transformers-native VibeVoice 1.5B"
```

## Acceptance

The unit checks do not download model weights:

```bash
python -m unittest discover -s tests -p 'test_publish_vibevoice_1_5b_hf.py' -v
```

The conversion command performs the optional real-weight validation. After the
owner publishes, this native-only load confirms the public artifact has no
remote code or sidecar dependency:

```python
from transformers import AutoModelForTextToWaveform, AutoProcessor

model_id = "microsoft/VibeVoice-1.5B-HF"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=False)
model = AutoModelForTextToWaveform.from_pretrained(
model_id,
dtype="auto",
trust_remote_code=False,
)

assert processor.__class__.__name__ == "VibeVoiceProcessor"
assert model.__class__.__name__ == "VibeVoiceForConditionalGeneration"
```
1 change: 1 addition & 0 deletions docs/vibevoice-tts.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@

**Model:** [VibeVoice-1.5B](https://huggingface.co/microsoft/VibeVoice-1.5B)<br>
**Report:** [Technical Report](https://arxiv.org/pdf/2508.19205)<br>
**Transformers-native publication:** [VibeVoice-1.5B-HF](https://huggingface.co/microsoft/VibeVoice-1.5B-HF) ([owner workflow](publish-vibevoice-1.5b-hf.md))<br>


<div align="center">
Expand Down
256 changes: 256 additions & 0 deletions tests/test_publish_vibevoice_1_5b_hf.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,256 @@
from __future__ import annotations

import importlib.util
import json
import sys
import tempfile
import unittest
from pathlib import Path
from unittest.mock import patch


ROOT = Path(__file__).resolve().parents[1]
SCRIPT = ROOT / "tools" / "release" / "convert_vibevoice_1_5b_hf.py"
SPEC = importlib.util.spec_from_file_location("publish_vibevoice", SCRIPT)
assert SPEC is not None and SPEC.loader is not None
PUBLISH = importlib.util.module_from_spec(SPEC)
sys.modules[SPEC.name] = PUBLISH
SPEC.loader.exec_module(PUBLISH)


def write_safetensors_header(path: Path, tensors: dict[str, dict[str, object]]) -> None:
header = json.dumps(tensors, separators=(",", ":")).encode()
path.write_bytes(len(header).to_bytes(8, "little") + header)


def write_checkpoint(directory: Path, tensors: dict[str, dict[str, object]]) -> None:
shard = "model-00001-of-00001.safetensors"
write_safetensors_header(directory / shard, tensors)
(directory / "model.safetensors.index.json").write_text(
json.dumps({"weight_map": {name: shard for name in tensors}}),
encoding="utf-8",
)


class PublishVibeVoiceTests(unittest.TestCase):
def test_pinned_revisions_are_full_git_hashes(self) -> None:
for revision in (
PUBLISH.OFFICIAL_SOURCE_REVISION,
PUBLISH.TRANSFORMERS_REVISION,
PUBLISH.QWEN_TOKENIZER_REVISION,
PUBLISH.NATIVE_REFERENCE_REVISION,
):
self.assertRegex(revision, r"^[0-9a-f]{40}$")

def test_rejects_dirty_transformers_checkout(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
checkout = Path(temporary_directory)
converter = (
checkout
/ "src"
/ "transformers"
/ "models"
/ "vibevoice"
/ "convert_vibevoice_to_hf.py"
)
converter.parent.mkdir(parents=True)
converter.touch()
with patch.object(
PUBLISH.subprocess,
"run",
side_effect=[
PUBLISH.subprocess.CompletedProcess(
args=[],
returncode=0,
stdout=f"{PUBLISH.TRANSFORMERS_REVISION}\n",
),
PUBLISH.subprocess.CompletedProcess(
args=[],
returncode=0,
stdout=" M src/transformers/models/vibevoice/modular_vibevoice.py\n",
),
],
):
with self.assertRaisesRegex(PUBLISH.ConversionError, "uncommitted changes"):
PUBLISH.assert_transformers_revision(checkout)

def test_rejects_untracked_transformers_file(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
checkout = Path(temporary_directory)
converter = (
checkout
/ "src"
/ "transformers"
/ "models"
/ "vibevoice"
/ "convert_vibevoice_to_hf.py"
)
converter.parent.mkdir(parents=True)
converter.touch()
with patch.object(
PUBLISH.subprocess,
"run",
side_effect=[
PUBLISH.subprocess.CompletedProcess(
args=[],
returncode=0,
stdout=f"{PUBLISH.TRANSFORMERS_REVISION}\n",
),
PUBLISH.subprocess.CompletedProcess(
args=[],
returncode=0,
stdout="?? src/transformers/sidecar.py\n",
),
],
):
with self.assertRaisesRegex(PUBLISH.ConversionError, "uncommitted changes"):
PUBLISH.assert_transformers_revision(checkout)

def test_records_clean_release_tool_checkout(self) -> None:
expected_revision = "a" * 40
with patch.object(
PUBLISH.subprocess,
"run",
side_effect=[
PUBLISH.subprocess.CompletedProcess(
args=[],
returncode=0,
stdout=f"{expected_revision}\n",
),
PUBLISH.subprocess.CompletedProcess(args=[], returncode=0, stdout=""),
],
):
self.assertEqual(PUBLISH.release_tool_revision(), expected_revision)

def test_reads_metadata_from_indexed_safetensors_headers(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
checkpoint = Path(temporary_directory)
write_checkpoint(
checkpoint,
{
"tensor_a": {"dtype": "BF16", "shape": [2, 3], "data_offsets": [0, 12]},
"tensor_b": {"dtype": "F32", "shape": [], "data_offsets": [12, 16]},
},
)

metadata = PUBLISH.checkpoint_tensor_metadata(checkpoint, expected_tensor_count=2)

self.assertEqual(metadata["tensor_a"].shape, (2, 3))
self.assertEqual(metadata["tensor_a"].dtype, "BF16")
self.assertEqual(metadata["tensor_b"].shape, ())

def test_rejects_index_header_key_mismatch(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
checkpoint = Path(temporary_directory)
write_safetensors_header(
checkpoint / "model-00001-of-00001.safetensors",
{"actual": {"dtype": "BF16", "shape": [1], "data_offsets": [0, 2]}},
)
(checkpoint / "model.safetensors.index.json").write_text(
json.dumps({"weight_map": {"indexed": "model-00001-of-00001.safetensors"}}),
encoding="utf-8",
)

with self.assertRaisesRegex(PUBLISH.ConversionError, "index/header mismatch"):
PUBLISH.checkpoint_tensor_metadata(checkpoint, expected_tensor_count=1)

def test_rejects_dtype_or_shape_mismatch(self) -> None:
source = {"original": PUBLISH.TensorMetadata(shape=(2, 3), dtype="BF16")}
native = {"native": PUBLISH.TensorMetadata(shape=(2, 3), dtype="F32")}

with self.assertRaisesRegex(PUBLISH.ConversionError, "metadata mismatch"):
PUBLISH.assert_tensor_metadata_matches(
source,
native,
lambda _: "native",
expected_tensor_count=1,
)

def test_rejects_key_mismatch(self) -> None:
source = {"original": PUBLISH.TensorMetadata(shape=(2, 3), dtype="BF16")}
native = {"other": PUBLISH.TensorMetadata(shape=(2, 3), dtype="BF16")}

with self.assertRaisesRegex(PUBLISH.ConversionError, "Tensor key mismatch"):
PUBLISH.assert_tensor_metadata_matches(
source,
native,
lambda _: "native",
expected_tensor_count=1,
)

def test_rejects_remote_code_metadata(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
output = Path(temporary_directory)
for name in PUBLISH.REQUIRED_OUTPUT_FILES:
(output / name).write_text("{}", encoding="utf-8")
(output / "config.json").write_text(
json.dumps(
{
"model_type": "vibevoice",
"architectures": ["VibeVoiceForConditionalGeneration"],
}
),
encoding="utf-8",
)
(output / "processor_config.json").write_text(
json.dumps({"processor_class": "VibeVoiceProcessor"}),
encoding="utf-8",
)
(output / "tokenizer_config.json").write_text(
json.dumps({"auto_map": {"AutoTokenizer": "untrusted.Module"}}),
encoding="utf-8",
)

with self.assertRaisesRegex(PUBLISH.ConversionError, "remote custom code"):
PUBLISH.assert_native_assets(output)

def test_alignment_digest_is_deterministic(self) -> None:
source = {
"first": PUBLISH.TensorMetadata(shape=(1,), dtype="BF16"),
"second": PUBLISH.TensorMetadata(shape=(2,), dtype="F32"),
}
native = {
"native.first": source["first"],
"native.second": source["second"],
}

first = PUBLISH.assert_tensor_metadata_matches(
source,
native,
lambda name: f"native.{name}",
expected_tensor_count=2,
)
second = PUBLISH.assert_tensor_metadata_matches(
dict(reversed(list(source.items()))),
native,
lambda name: f"native.{name}",
expected_tensor_count=2,
)

self.assertEqual(first, second)

def test_manifest_records_pinned_provenance(self) -> None:
with tempfile.TemporaryDirectory() as temporary_directory:
output = Path(temporary_directory)
(output / "config.json").write_text("{}", encoding="utf-8")
PUBLISH.write_manifest(output, "f" * 64, "a" * 40)
manifest = json.loads((output / "conversion-manifest.json").read_text(encoding="utf-8"))

self.assertEqual(
manifest["source"]["revision"],
PUBLISH.OFFICIAL_SOURCE_REVISION,
)
self.assertEqual(
manifest["canonical_converter"]["revision"],
PUBLISH.TRANSFORMERS_REVISION,
)
self.assertEqual(manifest["release_tool"]["revision"], "a" * 40)
self.assertEqual(manifest["tensor_alignment"]["source_tensor_count"], 1204)
self.assertEqual(
manifest["native_reference"]["revision"],
PUBLISH.NATIVE_REFERENCE_REVISION,
)


if __name__ == "__main__":
unittest.main()
Loading