Skip to content

Commit ec8addd

Browse files
authored
asr: default to mmx speech transcribe (mmx-cli >= 1.0.26); keep --provider api REST fallback (1.9.0)
1 parent 443ca01 commit ec8addd

8 files changed

Lines changed: 201 additions & 39 deletions

File tree

‎plugins/Wzdhehe/html2video-for-mcode/.claude-plugin/plugin.json‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
{
22
"name": "html2video-for-mcode",
3-
"version": "1.8.1",
3+
"version": "1.9.0",
44
"description": "Turn a topic, outline, or script into a narrated MP4: HTML slides with staged entrance animations, TTS voiceover, ffmpeg assembly, and ASR verification.",
55
"skills": [
66
"./skills/html2video-for-mcode/SKILL.md"

‎plugins/Wzdhehe/html2video-for-mcode/CHANGELOG.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,15 @@
11
# Changelog
22

3+
## 1.9.0 — 2026-09-20
4+
5+
**ASR toolchain simplification: `mmx speech transcribe` is now the default provider (mmx-cli ≥ 1.0.26, merged upstream in MiniMax-AI/cli#262)**
6+
7+
- `scripts/asr.mjs` now picks its provider automatically: with mmx-cli ≥ 1.0.26 on `PATH` it shells out to `mmx speech transcribe` — the same login as TTS, so the API key never passes through this script at all. Results are staged through a dedicated `os.tmpdir()` scratch directory before landing in jailed paths; the external CLI never touches user paths directly. `--provider api` forces the previous direct-REST path for environments without mmx-cli, and `--from` (offline compare) no longer trips the missing-key gate.
8+
- The capability probe reads `mmx --version` — `--help` exits 0 even for unknown subcommands, so it cannot gate. On Windows the CLI is invoked through `cmd.exe /c` because Node cannot spawn npm-global commands directly (bare name → ENOENT, `.cmd` → EINVAL).
9+
- The endpoint allowlist is unchanged and provider-independent: a non-official `--base-url` / `MINIMAX_BASE_URL` is still rejected before any provider resolution.
10+
- Docs corrected where they claimed "mmx-cli has no ASR subcommand" (false since 1.0.26): SKILL.md (scripts table, runtime comparison table, first-time setup, Phase 5), `references/tts-and-timing.md` (ASR paths B/C), and both READMEs' requirements/disclosure bullets — the mmx provider spends the same paid, quota-metered MiniMax account as REST.
11+
- Tests +4 (invalid `--provider` guidance; keyless `--provider api` guidance; an mmx-shim integration pair asserting the exact argv mapping — `speech transcribe --model asr-1.0 --response-format … --timestamp-level … --language …` — tmpdir staging/cleanup, and failure passthrough with the upgrade hint). REST-stub tests now pin `--provider api` so they stay deterministic on machines that do have mmx. **248 tests / 14 files**: verified locally with tools (245 pass / 0 fail / 3 capability skips) and end-to-end against real mmx 1.0.26 (keyless run, transcript exact including the number).
12+
313
## 1.8.1 — 2026-09-19
414

515
**Thirteenth-review residue (found by the independent sweep over the published `e845cbd8` tree): `SKILL.md` was the one install surface 1.8.0 left behind**

‎plugins/Wzdhehe/html2video-for-mcode/plugin.json‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
11
{
22
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
33
"name": "html2video-for-mcode",
4-
"version": "1.8.1",
4+
"version": "1.9.0",
55
"description": "Turn a topic, outline, or script into a narrated MP4: HTML slides with staged entrance animations, TTS voiceover, ffmpeg assembly, and ASR verification.",
66
"author": {
77
"name": "Wzdhehe",

‎plugins/Wzdhehe/html2video-for-mcode/skills/html2video-for-mcode/SKILL.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -47,7 +47,7 @@ The skill ships with **12 command-line scripts + 8 internal modules** (call them
4747
| `scripts/preview-page.mjs <project> [--open] [--no-script]` | Generates the **play page** `preview/play/index.html` (single file, zero dependencies, double-click to view over file://): **with animations on, ←→/touch swipe step through the entrance level by level** — each press of → reveals the next level (a sub-heading or small chart animates in on the spot), and only after the last level does it turn the page; with animations off it just turns pages. X toggles animations on/off, P toggles narration on/off (all toggles are buttons in the bottom bar, and the label text literally reads "动效开/动效关" "口播开/口播关" — animations on/off, narration on/off; the top bar keeps only the page number), O overview, F fullscreen. **Frame switching is double-buffered with no white flash; UI text adapts between Chinese and English per the `lang` in script.json**. **It does only one thing: "show the frame"** — no timer/progress bar/read-along highlight/replay button (to see timing, watch the finished video). Snapshots inject the measured delays from `timings.json` (the level-by-level animation shares its source with the finished video); the narration panel decides automatically from the data whether to load at all, collapses into a bottom drawer with default-collapsed state in narrow windows/on phones, and the canvas adapts to landscape or vertical |
4848
| `scripts/build-video.mjs <project> [--asr] [--dry-run] [--transition cut|xfade]` | Encodes each slide → concatenates → aligns the audio track → muxes → self-checks + emits `out/subs.srt`; `--asr` splits the audio **by sentence** + generates a checklist; `--dry-run` only prints the ffmpeg commands that would run (for troubleshooting). **Transitions default to a hard cut** (no black frame between segments; the first segment dissolves in from the cover and the last fades out), and `--transition xfade` or `transition` in `script.json` switches to a 0.4s cross-dissolve (the dissolve eats the extra frames kept at the end of each segment, so the total duration is unchanged). When subtitle stills exist it splices as "frame-sequence segments + subtitle segments" (the segment lengths must sum exactly to the slide's duration) |
4949
| `scripts/grab-frames.mjs <project> [--ids 05,11] [--at 0.5] [--both]` | **Finished-video frame extraction check**: compute each slide's absolute start from `timings.json` and pull frames from `out/final.mp4` into `build/introspect/` — problems like a caption covered by subtitles, the last-level element missing from the frame, or too little still time at the end are "invisible in a still, only visible in the finished video", so go through them slide by slide before publishing |
50-
| `scripts/asr.mjs <project> [--api-key K] [--verify-timing] [--from <transcript>] [--allow-any-endpoint]` | Calls ASR to transcribe and verifies sentence by sentence whether the audio says what the script says (mismatched numbers/traditional characters count as ✗); `--verify-timing` uses character-level timestamps to measure sentence onset. The key is only ever sent to official domains |
50+
| `scripts/asr.mjs <project> [--provider mmx\|api] [--api-key K] [--verify-timing] [--from <transcript>] [--allow-any-endpoint]` | Calls ASR to transcribe and verifies sentence by sentence whether the audio says what the script says (mismatched numbers/traditional characters count as ✗); `--verify-timing` uses character-level timestamps to measure sentence onset. Defaults to `mmx speech transcribe` (mmx-cli ≥ 1.0.26, same login as TTS — no key handling here); `--provider api` calls REST directly instead, and then the key is only ever sent to official domains |
5151

5252
Internal modules (imported by the scripts above, never run on their own): `tools.mjs` (ffmpeg/ffprobe probing + path jail `safeId/safeRel/inside`), `url-policy.mjs` (ASR endpoint allowlist + SSRF/redirect policy + download filenames), `nofx-css.mjs` (single source of truth for the `no-fx` rules), `chart-css.mjs` (single source of truth for chart animations and chart primitives), `table-css.mjs` (table primitives: `.tbl/.kv/.matrix/.rank`), `css-kit.mjs` (the **managed-region mechanism**: rev delimiter comments + `findAllBlocks` + in-place replacement + the status vocabulary `ok`/`stale`/`duplicate`/`broken`/`legacy-outside`/`missing`, shared by `--upgrade-css`/`--check-css`/check-slides/preview-page/capture/build-video), `tokens-template.mjs` (the single source of the generated tokens.css body + its content hash `TOKENS_REV`), `limits.mjs` (the 2 MB scan cap that applies to both tokens.css and slide HTML). The whole generated body is wrapped in a `tokens` managed region **with the three toolkit blocks nested inside**, so after any source module changes, `--upgrade-css` propagates exactly into old projects; `BLOCK_IDS` in css-kit is the single list of region ids.
5353

@@ -84,11 +84,11 @@ The script layer (screenshots / rendering / compositing / validation / images) i
8484
| TTS synthesis | `mcode-tools connector call connector__matrix__batch_text_to_audio --args '{...}'` (≤10 items per batch, the main path); for a single voice test use `connector__matrix__synthesize_speech` | `mmx speech synthesize --text "第一句口播。" --voice <voice_id> --speed 1.1 --out audio/01.mp3`; voice list via `mmx speech voices` |
8585
| Writing results to disk | `get_asset_url <node_id>` → download to `audio/<id>.mp3` | `--out` writes straight to disk |
8686
| BGM music | `connector__matrix__batch_text_to_music` (≤5 items per batch) | ⚠ **mmx-cli has no music generation** → have the user provide a music file (confirm licensing, then register it in MANIFEST), or skip BGM |
87-
| Reverse ASR verification | `mcode-tools upload_temp_url` + `connector__matrix__listen_audio` | **`node scripts/asr.mjs <project>`** — calls REST (`/v1/speech_to_text`) directly with the same API Key, with no dependency on mcode and no need to install whisper; **it automatically compares against the expected text in the checklist and back-fills it, and mismatched numbers/traditional characters (Cantonese) are marked ✗ outright**. To measure sentence onset with character-level timestamps: `--verify-timing` |
87+
| Reverse ASR verification | `mcode-tools upload_temp_url` + `connector__matrix__listen_audio` | **`node scripts/asr.mjs <project>`** — transcribes via `mmx speech transcribe` (mmx-cli ≥ 1.0.26; same mmx login as TTS, no separate key juggling), with no dependency on mcode and no need to install whisper; `--provider api` + `MINIMAX_API_KEY` is the direct-REST fallback for environments without mmx-cli. **It automatically compares against the expected text in the checklist and back-fills it, and mismatched numbers/traditional characters (Cantonese) are marked ✗ outright**. To measure sentence onset with character-level timestamps: `--verify-timing` |
8888
| Image assets | Built-in browser inspecting the official site's DOM (preferred; take the image URL to disk with `scripts/fetch-official-images.mjs --url`) / official brand kit | Same image paths as on the left (in this environment use `fetch-official-images`, which ships its own Playwright headless browser for rendering pages); for abstract illustrations use `mmx image generate --prompt "..." --aspect-ratio 16:9 --n 3` (use `9:16` for vertical projects; **abstract concept images only — generating logos / screenshots / real people's faces is forbidden**), then register it per `image-sources.md` |
8989
| Research | `web_search` / `web_fetch` (built in); for SPA/JS-rendered pages open them in the **built-in browser** to read the body text (usage discipline in references/research.md) | `mmx search "<keyword>"` / `mmx text chat`; for SPA pages render the body text with the project's Playwright **headless browser** (minimal command in research.md) |
9090

91-
**First-time mmx-cli setup** (non-mcode environments): `npm install -g mmx-cli` → `mmx auth login --api-key sk-xxx` → verify with `mmx quota`. A 401 is usually a region mismatch: `mmx config set --key region --value cn|global`. On the script side, use `MINIMAX_API_KEY` (required) and `MINIMAX_REGION=cn|global` (optional) to line up with the same identity.
91+
**First-time mmx-cli setup** (non-mcode environments): `npm install -g mmx-cli` (TTS works on any version; ASR needs ≥ 1.0.26 for `speech transcribe`) → `mmx auth login --api-key sk-xxx` → verify with `mmx quota`. A 401 is usually a region mismatch: `mmx config set --key region --value cn|global`. On the script side, ASR reuses that same mmx login; set `MINIMAX_API_KEY` (+ `MINIMAX_REGION=cn|global`) only if you force the direct-REST path (`--provider api`).
9292

9393
**Discipline does not change with the environment**: durations still come from ffprobe measurements, subtitles still come from `clauses[]`, voices still have to be auditioned and language-verified (via asr.mjs above), and no Gate is skipped. **Search and body-text extraction are not limited to the tools in the table above either**: if other search skills/plugins are installed locally (web search, web reading, etc.), or any headless browser (the project's Playwright counts), use whichever one can get the body text — only the tool changes; source grading, two-source cross-checking and scope annotation are not negotiable.
9494

@@ -229,7 +229,7 @@ node <skill>/scripts/build-video.mjs <project> --asr
229229
- build-video does this automatically: encode each slide (uniform parameters) → concat (automatic fallback re-encode on duration drift) → align the audio track with `apad` using each slide's measured duration → **bed the BGM underneath (only when bgm is configured: loop to fill, fade in/out, voice first; falls back to voice-only if mixing fails)** → mux → ffprobe duration check + full decode self-check, non-zero exit code if it fails; it also outputs `out/subs.srt` (same source and same windows as the burned-in subtitles; bilingual subtitles are generated automatically as two lines from each clause's text2, for platform upload).
230230
- `--asr`: splits out `asr/part-<id>-<k>.mp3` **by sentence** + generates `asr/checklist.md`. The transcription is compared with the expected text: numbers, years and product names must match; homophones are tolerable. **If a segment's transcription bleeds in the start of the previous sentence, that sentence actually started later than estimated — run check-timing to calibrate.**
231231
- mcode: upload with `mcode-tools upload_temp_url`, then hand off to `connector__matrix__listen_audio`.
232-
- **Other environments**: `MINIMAX_API_KEY=sk-xxx node scripts/asr.mjs <project>` — calls REST directly (same Key), compares and back-fills the checklist automatically; failing items (mismatched numbers/traditional characters) are reported with a non-zero exit code. For more accurate onset times: `--verify-timing`.
232+
- **Other environments**: `node scripts/asr.mjs <project>` — transcribes via `mmx speech transcribe` (mmx-cli ≥ 1.0.26, same login as TTS), compares and back-fills the checklist automatically; failing items (mismatched numbers/traditional characters) are reported with a non-zero exit code. Without mmx-cli: `MINIMAX_API_KEY=sk-xxx node scripts/asr.mjs <project> --provider api` calls REST directly with the same Key. For more accurate onset times: `--verify-timing`.
233233
- For a slide that fails: change the narration or redo that segment's TTS → re-run plan-timings → re-render that slide (just delete the corresponding slide's frame directory).
234234

235235
**Gate 5**: the user reviews the finished video `out/final.mp4` + the ASR verification table, including a subtitle readability check (play it once muted — can the subtitles carry the meaning) and the BGM level (is the voice always clear). **Before publishing, pull real frames from the finished video and look at them slide by slide**: `node <技能>/scripts/grab-frames.mjs <project>` → `build/introspect/` — whether captions/credits are covered by subtitles, whether all the last-level elements made it into the frame, and whether the ending holds still long enough; these three are invisible in both stills and the play page.

‎plugins/Wzdhehe/html2video-for-mcode/skills/html2video-for-mcode/references/tts-and-timing.md‎

Lines changed: 10 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -146,16 +146,23 @@ build-video `--asr` cuts `asr/part-<id>-<k>.mp3` **by sentence** (sentence k of
146146
2. `mcode-tools connector call connector__matrix__listen_audio --args '{"audio_info": {"url": "<URL>"}}'`.
147147
3. Compare the transcript against the "expected narration" in the checklist: numbers, years and product names must match exactly; homophone/punctuation differences are acceptable.
148148

149-
**Path B · direct REST (any environment, the same single API Key; mmx-cli has no ASR subcommand)**
149+
**Path B · mmx-cli (any environment, the same login as TTS — mmx-cli ≥ 1.0.26 ships `speech transcribe`)**
150+
151+
```bash
152+
node scripts/asr.mjs <项目目录> # 默认走 mmx speech transcribe, 不需要配 Key → 逐句转写 → 自动比对 → 回填 asr/checklist.md
153+
```
154+
155+
**Path C · direct REST (fallback for environments without mmx-cli, the same single API Key)**
150156

151157
```bash
152158
export MINIMAX_API_KEY=sk-xxx # 与 mmx-cli 同一把; 海外套餐加 MINIMAX_REGION=global
153-
node scripts/asr.mjs <项目目录> # 逐句转写 → 自动比对 → 回填 asr/checklist.md
159+
node scripts/asr.mjs <项目目录> --provider api
154160
```
155161

162+
- Both paths hit the same backend, share the same limits, and differ only in who holds the credentials; `--provider mmx|api` forces one explicitly (the default auto-picks mmx when ≥ 1.0.26 is on PATH).
156163
- The script performs the **comparison verdict** automatically: numbers mismatched, or traditional characters appearing (suspected Cantonese) → ✗ and the process exits non-zero; low similarity → ⚠ pending review.
157164
- If the audio exceeds the API limits (>500s or >50MB) it is automatically transcoded with ffmpeg to mono 16k mp3 before upload; no manual handling needed.
158-
- If you already have transcripts (e.g. produced with whisper in another environment) you can feed them straight in for comparison: `node scripts/asr.mjs <项目> --from transcripts.json` (JSON shaped like `{"part-01-1": "文本"}`).
165+
- If you already have transcripts (e.g. produced with whisper in another environment) you can feed them straight in for comparison (no network needed): `node scripts/asr.mjs <项目> --from transcripts.json` (JSON shaped like `{"part-01-1": "文本"}`).
159166
- If you want **word-level timestamps to measure each sentence's onset** (more accurate than silence detection; the API supports `timestamp_level: word`): `node scripts/asr.mjs <项目> --verify-timing`, which outputs an "estimate vs ASR-measured" comparison table; if the deviation is large, adjust `clauses[].start` in `timings.json` per the results and re-render.
160167

161168
For slides that don't pass: change the narration or redo that segment's TTS → re-run plan-timings → delete `build/frames/<id>/` and `out/slide-<id>.mp4` → re-run capture (that slide) and build-video. Don't redo the whole video.

0 commit comments

Comments
 (0)