Skip to content

syncer(dm): resize worker resources after updating worker-count - #12861

Open
joechenrh wants to merge 1 commit into
pingcap:masterfrom
joechenrh:fix/dm-syncer-resize-worker-count
Open

joechenrh wants to merge 1 commit into
pingcap:masterfrom
joechenrh:fix/dm-syncer-resize-worker-count

Conversation

@joechenrh

@joechenrh joechenrh commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

What problem does this PR solve?

Issue Number: close #12813

Syncer.CheckCanUpdateCfg and Syncer.Update accept a new worker-count when foreign key causality is not in use, which is the default. After the task resumes, the DML workers use the new count, but s.workerJobTSArray and s.toDBConns keep the size they had when the syncer was created. Raising worker-count makes queues beyond the old count index out of range, and dm-worker panics.

What is changed and how it works?

worker-count is one of the fields the update path is intended to support, so this PR makes the change take effect instead of rejecting it:

  • Syncer.reset resizes workerJobTSArray to WorkerCount + workerJobTSArrayInitSize when the size differs. reset runs after all goroutines of the previous Run have exited, which are the only readers of the slice.
  • resetDBs resets the DML connections through a new resetDMLDBs. When the number of connections differs from WorkerCount, it creates a new downstream pool with WorkerCount connections and closes the old pool only after the new connections are ready. If creating them fails, the old pool is kept and resuming can retry.

The existing guard that rejects changing worker-count when foreign key causality is in use is not changed.

Check List

Tests

  • Unit test
    • TestResetAfterWorkerCountUpdateResizesJobTSArray: increases, decreases, then increases worker-count on the same syncer, and checks that every new DML queue can update its timestamp slot. Fails on master.
    • TestResetDMLDBsAfterWorkerCountUpdate: checks that the DML connections are recreated with the new count, and that a failed recreation keeps the old pool. Fails on master.
  • go test ./dm/syncer/... passes with failpoints enabled.

Questions

Will it cause performance regression or break compatibility?

No. Connections are recreated only when worker-count changed.

Do you need to update user documentation, design documentation or monitoring documentation?

No.

Release note

Fix the issue that DM-worker might panic after increasing `worker-count` of a task through OpenAPI.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Refactor

    • Improved handling of worker-count changes while the syncer is paused, including resizing internal processing capacity.
    • Downstream database connections are refreshed to match the current worker count.
    • Existing connections remain available if a replacement attempt fails, and old connections are closed after successful replacement.
  • Tests

    • Added coverage for worker-count updates, connection pool replacement, cleanup, and recovery from reconnection failures.

@ti-chi-bot ti-chi-bot Bot added release-note Denotes a PR that will be considered when it comes time to generate release notes. do-not-merge/needs-triage-completed area/dm Issues or PRs related to DM. size/L Denotes a PR that changes 100-499 lines, ignoring generated files. labels Sep 17, 2026
@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: bc41d96e-6ae1-4513-bf29-d2f5d3d22ab4

📥 Commits

Reviewing files that changed from the base of the PR and between 1824057 and 27332e7.

📒 Files selected for processing (1)
  • dm/syncer/syncer.go

Included review availability: Your plan provides up to 2 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The syncer reset path now resizes worker timestamp storage and recreates downstream DML connections when WorkerCount changes while paused. Tests cover repeated count changes, failed connection creation, and successful pool replacement.

Changes

Worker-count resume handling

Layer / File(s) Summary
Resize worker timestamp storage
dm/syncer/syncer.go, dm/syncer/syncer_test.go
reset() reallocates workerJobTSArray to WorkerCount + workerJobTSArrayInitSize. Tests cover worker counts 4, 2, and 5 and verify timestamp updates do not panic.
Rebuild downstream DML connections
dm/syncer/syncer.go, dm/syncer/syncer_test.go
createDMLDBs() centralizes downstream pool creation. resetDMLDBs() creates a replacement pool before closing the old pool. Tests verify failed creation preserves the old pool and successful creation replaces it.

Priority: ⬆️ High

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix · Severity of issue fixed: High

Sequence Diagram(s)

sequenceDiagram
  participant Syncer
  participant createDMLDBs
  participant DownstreamDB
  Syncer->>Syncer: reset worker timestamp array
  Syncer->>createDMLDBs: create replacement pool
  createDMLDBs->>DownstreamDB: create WorkerCount connections
  DownstreamDB-->>createDMLDBs: pool or error
  createDMLDBs-->>Syncer: new pool
  Syncer->>DownstreamDB: close old pool after success
Loading

Suggested reviewers: gmhdbjd, olivers929

Merge Risk: ⚪ Minimal · up to 27332

The worker-count resume path safely resizes worker state and replaces downstream connections without establishing a concrete merge-blocking risk.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes resizing worker resources after a worker-count update, which is the main change.
Description check ✅ Passed The description follows the required template. It includes the issue number, problem, implementation details, tests, compatibility assessment, documentation assessment, and release note.
Linked Issues check ✅ Passed The PR meets the coding requirements for issue #12813. reset resizes and zero-initializes workerJobTSArray to the current WorkerCount. resetDMLDBs recreates the DML connection pool when its si…
Out of Scope Changes check ✅ Passed The changes stay within issue #12813. Production changes update reset-time worker storage and DML connection handling. Tests verify the required worker-count update and failure behavior. No unrelated …
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

A rabbit sees the workers grow
New timestamp slots now neatly show
The old pool waits while new links start
Failed links leave the old ones part
Then fresh connections hop in line
And every queue runs fine

Comment @coderabbitai help to get the list of available commands.

@ti-chi-bot

ti-chi-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

@coderabbitai[bot]: adding LGTM is restricted to approvers and reviewers in OWNERS files.

Details

In response to this:

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@ti-chi-bot

ti-chi-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: coderabbitai[bot]
Once this PR has been reviewed and has the lgtm label, please assign gmhdbjd for approval. For more information see the Code Review Process.
Please ensure that each of them provides their approval before proceeding.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@codecov

codecov Bot commented Sep 17, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 57.4519%. Comparing base (03ed243) to head (3410208).

❌ Your project check has failed because the head coverage (57.4519%) is below the target coverage (60.0000%). You can increase the head coverage or adjust the target coverage.

Additional details and impacted files
Components Coverage Δ
cdc 57.4519% <ø> (∅)
dm ∅ <ø> (∅)
engine ∅ <ø> (∅)
Flag Coverage Δ
cdc 57.4519% <ø> (?)
unit 57.4519% <ø> (+6.7973%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

@@               Coverage Diff                @@
##             master     #12861        +/-   ##
================================================
+ Coverage   50.6546%   57.4519%   +6.7973%     
================================================
  Files           213        525       +312     
  Lines         17720      69700     +51980     
================================================
+ Hits           8976      40044     +31068     
- Misses         8178      26742     +18564     
- Partials        566       2914      +2348     
🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Syncer.Update accepts a new worker-count, and the resumed DML workers
use it to index the job timestamp slots and the downstream connections.
Both were sized only when the syncer was created, so raising
worker-count made dm-worker panic with index out of range after the
task resumed.

Resize the job timestamp slots when resetting the syncer, and recreate
the DML connections when their number differs from worker-count. The
old connections are kept if creating the new ones fails, so resuming
can retry. The connection setup is shared with createDBs through
createDMLDBs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@joechenrh
joechenrh force-pushed the fix/dm-syncer-resize-worker-count branch from 27332e7 to 0f762a6 Compare September 18, 2026 01:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/dm Issues or PRs related to DM. release-note Denotes a PR that will be considered when it comes time to generate release notes. size/L Denotes a PR that changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[DM] Increasing syncer worker-count and resuming panics with index out of range

1 participant