Watch DCP fields per GPU model to fix startup on mixed-model nodes - #736
CodeBuildder wants to merge 2 commits into
Conversation
On a node with more than one GPU model (issue NVIDIA#657 reports an A30 with MIG plus an RTX2000), dcgm-exporter puts every GPU into one shared DCGM group and watches one field group against it. DCP/profiling fields are validated for the whole group at registration time, so if any GPU in the group doesn't support a requested profiling field, the whole watch call fails and the pod never comes up. Ordinary fields don't have this problem: DCGM just reports those as a per-entity NOT_SUPPORTED value at scrape time instead of failing the watch. The eligibility check for DCP fields already asks DCGM which profiling metric groups are supported, but only for GPU index 0, and applies that one answer to the whole node. That's a reasonable shortcut given nodes are almost always provisioned with a single GPU model, but it breaks down exactly here. This partitions the GPU watch by model instead: each model gets its own DCGM group, and each group only gets the DCP fields DCGM reports that model as supporting. Non-DCP fields are unaffected. Nodes with a single GPU model, the overwhelming majority of real deployments, are unchanged and keep using the original single-group path. I considered probing every individual GPU rather than one representative GPU per model, but DCGM reports profiling support per model, not per unit, so that would just repeat identical queries for no benefit. Verified with a reproduction: mocked an A30 + RTX2000 node and confirmed the pre-fix code fails registration with the same "Feature not supported" error from the report, then confirmed the fix succeeds and splits the field list correctly per model. This is validated against a mocked DCGM client, not real hardware, so it should still get a look from someone who can try it on an actual mixed-model node. Signed-off-by: kaushik-kumaran <kaushik.kumaran@ibm.com>
8c2b003 to
22d4422
Compare
|
I tested the PR on my single node cluster: Here is the config: The pod is no longer crashing, but here are the logs: |
|
I queried the metrics but don't see any DCGM_FI_PROF_ metrics. |
Reported by @gfrankliu testing this branch on real mixed hardware (A30 MIG + RTX2000 Ada): the pod no longer crashes, but the RTX2000 entity repeatedly logs "Repairing stale DCGM profiling watch" for the FP64 field and never recovers. Watch registration was already scoped per GPU model, but the read path wasn't: fieldsToScrape/latestValues asks DCGM for every configured field on every entity regardless of which model's group actually requested it. For an ordinary field that's harmless, DCGM reports one it doesn't support as a per-entity blank value. For a DCP field it isn't: asking for a field that was deliberately never watched for that entity returns DCGM_ST_NOT_WATCHED, and the existing stale-watch repair logic then retries forever trying to fix a watch that can't exist. Added ModelSupportsDCPField, caching the per-model answer the same way the watch-registration path already does, and filter the scrape-time field list with it before reading an entity's values. The cache clears on every DCGM reinit via the same queryDCPMetrics call that already refreshes DCP capability on that cadence. Signed-off-by: kaushik-kumaran <kaushik.kumaran@ibm.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
@gfrankliu thank you for actually testing this on your mixed A30 MIG + RTX2000 hardware, that's exactly the setup this fix targets and I didn't have access to it myself. Good news: the core fix works, the pod comes up and stays up. Your logs caught a real follow-up gap though: the RTX2000's repeated "Repairing stale DCGM profiling watch" for the FP64 field. What was happening: watch registration is now correctly scoped per GPU model, but the read path wasn't. It was still asking DCGM for every configured field on every entity regardless of which model's group actually watched it. For an ordinary field that's harmless, but for a DCP field, asking for one that was deliberately never watched for that entity gets you Just pushed a fix for that: the scrape-time field list is now filtered per entity the same way the watch registration already is, so the RTX2000 will stop being asked for a field it was never watched for. Also added test coverage for this specific gap so it can't regress silently. Would appreciate it if you could pull the latest commit and confirm the repair-loop warnings are gone on your setup when you get a chance. No pressure on timing, just flagging it's ready whenever you want to retest. |
|
Thanks @CodeBuildder ! The new commit fixed the startup errors and I can see the DCGM_FI_PROF_* metrics now. A few observations:
Again, here are the GPUs on the node: |
|
Thanks for confirming, glad the follow-up fix actually cleared the repair-loop on your setup. On your two observations:
If it turns out to be new, I'd want to dig into it properly rather than guess further. |
|
|
The NVRM message is unrelated. This PR looks fine in my test and can be merged. |
|
@CodeBuildder thanks again. This has been running fine on our mixed A30 (MIG) + RTX2000 nodes. Since 4.8.4 landed, the branch now conflicts with While reading the diff, I noticed two things that could make profiling metrics quietly disappear in production: 1. The watch path and scrape path handle a
So if that query fails for a model, the scrape reads DCP fields that were never watched and gets 2. GPU 0 still decides DCP support for the whole node
On mixed-model nodes the result also depends on which card enumerates as GPU 0. Would it make sense to take the union of supported fields across one representative GPU per model here, since the watch path already filters per model? Neither of these blocks our current testing, but we'd like to run an official release of this in prod, so I wanted to raise them before merge. |
Fixes #657.
The problem
On a node with more than one GPU model, dcgm-exporter can fail to start entirely. The reporter's case is an A30 running in MIG mode alongside an RTX2000. The pod logs end with:
Root cause: all GPUs of a given entity type share one DCGM group, watched with one field group. DCP/profiling fields (
DCGM_FI_PROF_*) are validated for the whole group at registration time, so if any single GPU in that group doesn't support a requested profiling field, the entire watch call fails and the exporter exits. Ordinary fields don't hit this: DCGM just reports those as a per-entityNOT_SUPPORTEDvalue at scrape time instead of failing the watch.The DCP eligibility check already exists (
queryDCPMetricscallsGetSupportedMetricGroups), but it only queries GPU index 0 and applies that single answer to the whole node. That's a reasonable shortcut since nodes are almost always provisioned with one GPU model, but it's exactly what breaks on a mixed node: whichever GPU happens to be at index 0 decides what the whole node is allowed to watch, and if a different model in the same node doesn't support a field that GPU 0 does, the shared watch call fails for everyone.The fix
WatchDeviceFieldGroupsnow partitions GPUs by model when the field list includes a DCP field:GetSupportedMetricGroups, queried once per model, not per GPU).I considered querying every individual GPU instead of one representative GPU per model, but DCGM reports profiling support per model, not per unit, so that would just repeat identical queries without changing behavior.
Testing
Added
internal/pkg/devicewatcher/dcp_model_partition_test.go, covering the model-partitioning logic directly and a full mocked run ofWatchDeviceFieldGroupswith an A30 + RTX2000 node. As part of verifying this, I reproduced the reported failure against the pre-fix code with the same mocked scenario (confirmed it fails with the same "Feature not supported" error from the report), then confirmed this change resolves it. Also added a same-model control test to confirm the unmodified path is unaffected.This is verified against a mocked DCGM client, not real hardware, so it would be good to get eyes from someone who can try it against an actual mixed-model node.