Skip to content

DM: optimistic shard DDL conflict on single source has no recovery path (unlock-ddl-lock rejects skip-and-wait-for-redirect; force-remove ignored) #12840

Description

@wjhuang2016

What did you do?

Run an optimistic shard-merge task on a single source (two databases of one MySQL instance merged into one downstream table). Execute a conflicting DDL on one branch, e.g. ALTER TABLE db1.tb1 CHARACTER SET utf8mb4 COLLATE utf8mb4_bin vs ALTER TABLE db2.tb1 CHARACTER SET utf8mb4 COLLATE utf8mb4_general_ci (different collations).

What did you expect to see?

The conflict can be resolved through the documented manual path (unlock-ddl-lock, correcting the upstream, or restarting the task).

What did you see instead?

The lock enters skip and wait for redirect and every recovery path fails:

  • unlock-ddl-lock (skip/exec) → 38020 "lock ... is in skip and wait for redirect status, not conflicted" (Optimist.UnlockLock only accepts ConflictDetected operations)
  • unlock-ddl-lock --force-remove → same 38020 (the optimistic branch ignores req.ForceRemove; dm/master/server.go:959-965 only passes req.Op)
  • upstream correction DDL never arrives: the same binlog stream is blocked at the conflicting DDL (checkpoint frozen)
  • stop-task + start-task: the lock is recreated from etcd and replay conflicts again

The task is permanently stuck; the only way out is remove-meta/recreate, which loses the checkpoint.

Root cause

  • Optimist.UnlockLock rejects any operation whose conflict stage is not ConflictDetected (dm/master/shardddl/optimist.go:270-278)
  • Server.UnlockDDLLock optimistic branch ignores req.ForceRemove (dm/master/server.go:959-965)
  • single source: both branches share one binlog stream, so no other source can drive the redirect (multi-source recovery was not verified in this environment)

Versions

  • tiflow master 52c8e14; also present on release-8.5

Note

The conflict detection itself is correct (schemacmp reports "incompatible collation"); the broken part is the recovery path for the ConflictSkipWaitRedirect state. Single-source optimistic shard merge has no code check or test coverage (all optimistic tests use double sources).

Production reachability

  • Trigger prerequisites: shard-mode: optimistic (explicit user config) + a conflicting DDL across branches (different collation, column-type change, ADD NOT NULL, etc. — the conflict set DM documents as unsupported) + single-source topology (two databases of one MySQL instance sharing one binlog stream).
  • Frequency: medium-low (requires the optimistic-mode + conflict-DDL combination). Single-source optimistic shard merge is a legal configuration (no code check rejects it; no test covers it — all optimistic tests use double sources).
  • Consequence: task permanently stuck (checkpoint frozen at the conflicting DDL), sync interrupted; recovery requires remove-meta/recreate, losing the checkpoint. The documented manual recovery path (unlock-ddl-lock) is unusable for the skip and wait for redirect state.
  • Reachability: medium (config-combination gate). Multi-source optimistic recovery was not verified in this environment — it may be recoverable via another source driving the redirect; the single-source deadlock is reproduced deterministically.

Suggested labels

area/dm, type/bug, affects-8.5, severity/major, impact/func-failure, subject/replication-interruption (label add requires repo admin/triage rights, which the reporter does not have).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions