Skip to content

Watch DCP fields per GPU model to fix startup on mixed-model nodes - #736

Open
CodeBuildder wants to merge 2 commits into
NVIDIA:mainfrom
CodeBuildder:fix/mixed-gpu-model-dcp-watch-657
Open

CodeBuildder wants to merge 2 commits into
NVIDIA:mainfrom
CodeBuildder:fix/mixed-gpu-model-dcp-watch-657

Conversation

@CodeBuildder

@CodeBuildder CodeBuildder commented Sep 13, 2026 •

Copy link
Copy Markdown

Fixes #657.

The problem

On a node with more than one GPU model, dcgm-exporter can fail to start entirely. The reporter's case is an A30 running in MIG mode alongside an RTX2000. The pod logs end with:

error watching fields: Feature not supported

Root cause: all GPUs of a given entity type share one DCGM group, watched with one field group. DCP/profiling fields (DCGM_FI_PROF_*) are validated for the whole group at registration time, so if any single GPU in that group doesn't support a requested profiling field, the entire watch call fails and the exporter exits. Ordinary fields don't hit this: DCGM just reports those as a per-entity NOT_SUPPORTED value at scrape time instead of failing the watch.

The DCP eligibility check already exists (queryDCPMetrics calls GetSupportedMetricGroups), but it only queries GPU index 0 and applies that single answer to the whole node. That's a reasonable shortcut since nodes are almost always provisioned with one GPU model, but it's exactly what breaks on a mixed node: whichever GPU happens to be at index 0 decides what the whole node is allowed to watch, and if a different model in the same node doesn't support a field that GPU 0 does, the shared watch call fails for everyone.

The fix

WatchDeviceFieldGroups now partitions GPUs by model when the field list includes a DCP field:

  • Each distinct GPU model gets its own DCGM group.
  • Each group's field list is filtered to the DCP fields DCGM actually reports that model as supporting (GetSupportedMetricGroups, queried once per model, not per GPU).
  • Non-DCP fields are untouched and go into every group as before.
  • A node with a single GPU model, the common case, never enters this path and keeps using the original single-group logic exactly as it is today.

I considered querying every individual GPU instead of one representative GPU per model, but DCGM reports profiling support per model, not per unit, so that would just repeat identical queries without changing behavior.

Testing

Added internal/pkg/devicewatcher/dcp_model_partition_test.go, covering the model-partitioning logic directly and a full mocked run of WatchDeviceFieldGroups with an A30 + RTX2000 node. As part of verifying this, I reproduced the reported failure against the pre-fix code with the same mocked scenario (confirmed it fails with the same "Feature not supported" error from the report), then confirmed this change resolves it. Also added a same-model control test to confirm the unmodified path is unaffected.

This is verified against a mocked DCGM client, not real hardware, so it would be good to get eyes from someone who can try it against an actual mixed-model node.

On a node with more than one GPU model (issue NVIDIA#657 reports an A30 with
MIG plus an RTX2000), dcgm-exporter puts every GPU into one shared DCGM
group and watches one field group against it. DCP/profiling fields are
validated for the whole group at registration time, so if any GPU in
the group doesn't support a requested profiling field, the whole watch
call fails and the pod never comes up. Ordinary fields don't have this
problem: DCGM just reports those as a per-entity NOT_SUPPORTED value at
scrape time instead of failing the watch.

The eligibility check for DCP fields already asks DCGM which profiling
metric groups are supported, but only for GPU index 0, and applies that
one answer to the whole node. That's a reasonable shortcut given nodes
are almost always provisioned with a single GPU model, but it breaks
down exactly here.

This partitions the GPU watch by model instead: each model gets its own
DCGM group, and each group only gets the DCP fields DCGM reports that
model as supporting. Non-DCP fields are unaffected. Nodes with a single
GPU model, the overwhelming majority of real deployments, are unchanged
and keep using the original single-group path.

I considered probing every individual GPU rather than one representative
GPU per model, but DCGM reports profiling support per model, not per
unit, so that would just repeat identical queries for no benefit.

Verified with a reproduction: mocked an A30 + RTX2000 node and confirmed
the pre-fix code fails registration with the same "Feature not
supported" error from the report, then confirmed the fix succeeds and
splits the field list correctly per model. This is validated against a
mocked DCGM client, not real hardware, so it should still get a look
from someone who can try it on an actual mixed-model node.

Signed-off-by: kaushik-kumaran <kaushik.kumaran@ibm.com>
@CodeBuildder
CodeBuildder force-pushed the fix/mixed-gpu-model-dcp-watch-657 branch from 8c2b003 to 22d4422 Compare September 13, 2026 00:49
@gfrankliu

gfrankliu commented Sep 15, 2026 •

Copy link
Copy Markdown

I tested the PR on my single node cluster:

$ nvidia-smi -L
GPU 0: NVIDIA A30 (UUID: GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3)
  MIG 1g.6gb      Device  0: (UUID: MIG-c2e1f9b6-a5fd-5698-b5ea-00ebbdd34806)
  MIG 1g.6gb      Device  1: (UUID: MIG-84339efb-1018-5c35-97da-8ba001802bca)
  MIG 1g.6gb      Device  2: (UUID: MIG-0dbd45d4-a7c2-55cf-8586-752fac76e78f)
  MIG 1g.6gb      Device  3: (UUID: MIG-fb5eb194-36a0-55a0-9570-bacbf9f33106)
GPU 1: NVIDIA RTX 2000 Ada Generation (UUID: GPU-fc60126e-514d-3522-2828-c5daa8e99d85)
GPU 2: NVIDIA A30 (UUID: GPU-63bc0b91-2265-1b94-831c-15c41fe92675)
  MIG 1g.6gb      Device  0: (UUID: MIG-b2bc7070-2d13-5b7e-8e95-9ed9e560ba92)
  MIG 1g.6gb      Device  1: (UUID: MIG-9dede469-0ac1-5836-812a-63a42122485b)
  MIG 1g.6gb      Device  2: (UUID: MIG-e6b2eb9c-9cae-515d-9c1c-860272d09193)
  MIG 1g.6gb      Device  3: (UUID: MIG-85d711d6-474a-5856-8cc4-3657be932c31)

Here is the config:

apiVersion: v1
kind: ConfigMap
metadata:
  name: dcgm-custom-metrics
  labels:
    name: dcgm-custom-metrics
data:
  # prettier-ignore
  dcgm-custom-metrics.csv: |
    # Format
    # If line starts with a '#' it is considered a comment
    # DCGM FIELD, Prometheus metric type, help message
    # https://docs.nvidia.com/datacenter/dcgm/latest/dcgm-api/dcgm-api-field-ids.html

    # Memory usage
    DCGM_FI_DEV_FB_FREE, gauge, Framebuffer memory free (in MiB).
    DCGM_FI_DEV_FB_USED, gauge, Framebuffer memory used (in MiB).

    # DCP metrics (DCGM_FI_PROF_*) - testing on RTX2000/A30 mix
    DCGM_FI_PROF_GR_ENGINE_ACTIVE,   gauge, Ratio of time the graphics engine is active (in %).
    DCGM_FI_PROF_SM_ACTIVE,          gauge, The ratio of cycles an SM has at least 1 warp assigned (in %).
    DCGM_FI_PROF_SM_OCCUPANCY,       gauge, The ratio of number of warps resident on an SM (in %).
    DCGM_FI_PROF_PIPE_TENSOR_ACTIVE, gauge, Ratio of cycles the tensor (HMMA) pipe is active (in %).
    DCGM_FI_PROF_DRAM_ACTIVE,        gauge, Ratio of cycles the device memory interface is active sending or receiving data (in %).
    DCGM_FI_PROF_PIPE_FP64_ACTIVE,   gauge, Ratio of cycles the fp64 pipes are active (in %).
    DCGM_FI_PROF_PIPE_FP32_ACTIVE,   gauge, Ratio of cycles the fp32 pipes are active (in %).
    DCGM_FI_PROF_PIPE_FP16_ACTIVE,   gauge, Ratio of cycles the fp16 pipes are active (in %).

    # Power & Thermal
    DCGM_FI_DEV_POWER_USAGE, gauge, Power usage (in W).
    DCGM_FI_DEV_GPU_TEMP,    gauge, GPU temperature (in °C).
    DCGM_FI_DEV_MEMORY_TEMP, gauge, Memory temperature (in °C).

    # Utilization
    DCGM_FI_DEV_GPU_UTIL,      gauge, Ratio of time the GPU is active (in %).
    DCGM_FI_DEV_MEM_COPY_UTIL, gauge, Memory utilization (in %).
    DCGM_FI_DEV_ENC_UTIL,      gauge, Ratio of time the encoder is active (in %).
    DCGM_FI_DEV_DEC_UTIL,      gauge, Ratio of time the decoder is active (in %).

The pod is no longer crashing, but here are the logs:

time=2026-09-15T05:31:19.701Z level=INFO msg="Starting dcgm-exporter" Version=4.6.0-4.8.3
time=2026-09-15T05:31:19.706Z level=INFO msg="Attempting to initialize DCGM."
time=2026-09-15T05:31:20.322Z level=INFO msg="Initialized DCGM Fields module."
time=2026-09-15T05:31:20.324Z level=INFO msg="Attempting to initialize NVML library."
time=2026-09-15T05:31:20.324Z level=INFO msg="NVML provider successfully initialized for Kubernetes MIG support"
time=2026-09-15T05:31:20.324Z level=INFO msg="DCGM successfully initialized!"
time=2026-09-15T05:31:20.905Z level=INFO msg="Successfully queried DCGM profiling metric groups" reload_id=0 count=7 gpu_model="NVIDIA A30"
time=2026-09-15T05:31:20.905Z level=INFO msg="Building registry for current GPU topology"
time=2026-09-15T05:31:20.905Z level=INFO msg="Using metric file '/etc/dcgm-exporter/dcgm-custom-metrics.csv'"
time=2026-09-15T05:31:20.905Z level=INFO msg="Initializing system entities of type 'GPU'"
time=2026-09-15T05:31:21.020Z level=INFO msg="Initializing system entities of type 'NvSwitch'"
time=2026-09-15T05:31:21.021Z level=INFO msg="Not collecting NvSwitch metrics; no switches to monitor"
time=2026-09-15T05:31:21.021Z level=INFO msg="Initializing system entities of type 'NvLink'"
time=2026-09-15T05:31:21.021Z level=WARN msg="Failed to initialize NvSwitch/NvLink info" error="no switches to monitor"
time=2026-09-15T05:31:21.080Z level=INFO msg="Initializing system entities of type 'CPU'"
time=2026-09-15T05:31:21.878Z level=INFO msg="Not collecting CPU metrics; error retrieving DCGM CPU hierarchy v2: This request is serviced by a module of DCGM that is not currently loaded"
time=2026-09-15T05:31:21.878Z level=INFO msg="Initializing system entities of type 'CPU Core'"
time=2026-09-15T05:31:21.878Z level=INFO msg="Not collecting CPU Core metrics; error retrieving DCGM CPU hierarchy v2: This request is serviced by a module of DCGM that is not currently loaded"
time=2026-09-15T05:31:22.398Z level=INFO msg="Registry built successfully" collector_count=1
time=2026-09-15T05:31:22.399Z level=INFO msg="Kubernetes metrics collection enabled!"
time=2026-09-15T05:31:22.399Z level=INFO msg="HTTP server started - ready to serve metrics"
time=2026-09-15T05:31:22.399Z level=INFO msg="Watching for changes in file" file=/etc/dcgm-exporter/dcgm-custom-metrics.csv debounce=200ms
time=2026-09-15T05:31:22.399Z level=INFO msg="Starting webserver"
time=2026-09-15T05:31:22.400Z level=INFO msg="Listening on" address=[::]:9400
time=2026-09-15T05:31:22.400Z level=INFO msg="TLS is disabled." http2=false address=[::]:9400
time=2026-09-15T05:32:16.045Z level=WARN msg="Repairing stale DCGM profiling watch" reason="profiling field 1006 for entity GPU:1 returned status -16" repairBackoff=1m0s
time=2026-09-15T05:32:16.392Z level=ERROR msg="DCGM profiling watch remains stale after repair" reason="profiling field 1006 for entity GPU:1 returned status -16"
time=2026-09-15T05:33:16.011Z level=ERROR msg="DCGM profiling watch repair is rate limited" reason="profiling field 1006 for entity GPU:1 returned status -16" retryAfter=34.323334ms
time=2026-09-15T05:34:16.003Z level=WARN msg="Repairing stale DCGM profiling watch" reason="profiling field 1006 for entity GPU:1 returned status -16" repairBackoff=1m0s
time=2026-09-15T05:34:16.327Z level=ERROR msg="DCGM profiling watch remains stale after repair" reason="profiling field 1006 for entity GPU:1 returned status -16"
time=2026-09-15T05:35:16.014Z level=WARN msg="Repairing stale DCGM profiling watch" reason="profiling field 1006 for entity GPU:1 returned status -16" repairBackoff=1m0s
time=2026-09-15T05:35:16.394Z level=ERROR msg="DCGM profiling watch remains stale after repair" reason="profiling field 1006 for entity GPU:1 returned status -16"
time=2026-09-15T05:36:15.991Z level=ERROR msg="DCGM profiling watch repair is rate limited" reason="profiling field 1006 for entity GPU:1 returned status -16" retryAfter=23.556578ms
time=2026-09-15T05:37:15.987Z level=WARN msg="Repairing stale DCGM profiling watch" reason="profiling field 1006 for entity GPU:1 returned status -16" repairBackoff=1m0s
time=2026-09-15T05:37:16.297Z level=ERROR msg="DCGM profiling watch remains stale after repair" reason="profiling field 1006 for entity GPU:1 returned status -16"

@gfrankliu

Copy link
Copy Markdown

I queried the metrics but don't see any DCGM_FI_PROF_ metrics.

$ curl dcgm_pod_ip:9400/metrics
# HELP DCGM_FI_DEV_DEC_UTIL Ratio of time the decoder is active (in %).
# TYPE DCGM_FI_DEV_DEC_UTIL gauge
DCGM_FI_DEV_DEC_UTIL{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 46
# HELP DCGM_FI_DEV_ENC_UTIL Ratio of time the encoder is active (in %).
# TYPE DCGM_FI_DEV_ENC_UTIL gauge
DCGM_FI_DEV_ENC_UTIL{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 9
# HELP DCGM_FI_DEV_FB_FREE Framebuffer memory free (in MiB).
# TYPE DCGM_FI_DEV_FB_FREE gauge
DCGM_FI_DEV_FB_FREE{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 11275
DCGM_FI_DEV_FB_FREE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-95xcf"} 2207
DCGM_FI_DEV_FB_FREE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-wwk9g"} 2207
DCGM_FI_DEV_FB_FREE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-9dtcv"} 2207
DCGM_FI_DEV_FB_FREE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-t7grm"} 2207
DCGM_FI_DEV_FB_FREE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-94fhb"} 2207
DCGM_FI_DEV_FB_FREE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-4s2c4"} 2207
DCGM_FI_DEV_FB_FREE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-px2gm"} 2207
DCGM_FI_DEV_FB_FREE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-2w5b2"} 2207
# HELP DCGM_FI_DEV_FB_USED Framebuffer memory used (in MiB).
# TYPE DCGM_FI_DEV_FB_USED gauge
DCGM_FI_DEV_FB_USED{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 4674
DCGM_FI_DEV_FB_USED{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-95xcf"} 3744
DCGM_FI_DEV_FB_USED{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-wwk9g"} 3744
DCGM_FI_DEV_FB_USED{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-9dtcv"} 3744
DCGM_FI_DEV_FB_USED{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-t7grm"} 3744
DCGM_FI_DEV_FB_USED{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-94fhb"} 3744
DCGM_FI_DEV_FB_USED{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-4s2c4"} 3744
DCGM_FI_DEV_FB_USED{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-px2gm"} 3744
DCGM_FI_DEV_FB_USED{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-2w5b2"} 3744
# HELP DCGM_FI_DEV_GPU_TEMP GPU temperature (in °C).
# TYPE DCGM_FI_DEV_GPU_TEMP gauge
DCGM_FI_DEV_GPU_TEMP{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 44
DCGM_FI_DEV_GPU_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-95xcf"} 45
DCGM_FI_DEV_GPU_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-wwk9g"} 50
DCGM_FI_DEV_GPU_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-9dtcv"} 44
DCGM_FI_DEV_GPU_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-t7grm"} 50
DCGM_FI_DEV_GPU_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-94fhb"} 45
DCGM_FI_DEV_GPU_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-4s2c4"} 50
DCGM_FI_DEV_GPU_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-px2gm"} 45
DCGM_FI_DEV_GPU_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-2w5b2"} 50
# HELP DCGM_FI_DEV_GPU_UTIL Ratio of time the GPU is active (in %).
# TYPE DCGM_FI_DEV_GPU_UTIL gauge
DCGM_FI_DEV_GPU_UTIL{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 70
# HELP DCGM_FI_DEV_MEMORY_TEMP Memory temperature (in °C).
# TYPE DCGM_FI_DEV_MEMORY_TEMP gauge
DCGM_FI_DEV_MEMORY_TEMP{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 0
DCGM_FI_DEV_MEMORY_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-95xcf"} 48
DCGM_FI_DEV_MEMORY_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-wwk9g"} 51
DCGM_FI_DEV_MEMORY_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-9dtcv"} 46
DCGM_FI_DEV_MEMORY_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-t7grm"} 50
DCGM_FI_DEV_MEMORY_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-94fhb"} 46
DCGM_FI_DEV_MEMORY_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-4s2c4"} 51
DCGM_FI_DEV_MEMORY_TEMP{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-px2gm"} 46
DCGM_FI_DEV_MEMORY_TEMP{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-2w5b2"} 53
# HELP DCGM_FI_DEV_MEM_COPY_UTIL Memory utilization (in %).
# TYPE DCGM_FI_DEV_MEM_COPY_UTIL gauge
DCGM_FI_DEV_MEM_COPY_UTIL{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 14
# HELP DCGM_FI_DEV_POWER_USAGE Power usage (in W).
# TYPE DCGM_FI_DEV_POWER_USAGE gauge
DCGM_FI_DEV_POWER_USAGE{gpu="1",UUID="GPU-fc60126e-514d-3522-2828-c5daa8e99d85",pci_bus_id="00000000:B1:00.0",device="nvidia1",modelName="NVIDIA RTX 2000 Ada Generation",hostname="dcgm-exporter-wwwjh"} 34.959
DCGM_FI_DEV_POWER_USAGE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-95xcf"} 173.625
DCGM_FI_DEV_POWER_USAGE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="3",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-wwk9g"} 177.291
DCGM_FI_DEV_POWER_USAGE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-9dtcv"} 173.625
DCGM_FI_DEV_POWER_USAGE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="4",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-t7grm"} 177.685
DCGM_FI_DEV_POWER_USAGE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-94fhb"} 165.644
DCGM_FI_DEV_POWER_USAGE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="5",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-4s2c4"} 126.788
DCGM_FI_DEV_POWER_USAGE{gpu="2",UUID="GPU-63bc0b91-2265-1b94-831c-15c41fe92675",pci_bus_id="00000000:CA:00.0",device="nvidia2",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-px2gm"} 165.644
DCGM_FI_DEV_POWER_USAGE{gpu="0",UUID="GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3",pci_bus_id="00000000:65:00.0",device="nvidia0",modelName="NVIDIA A30",GPU_I_PROFILE="1g.6gb",GPU_I_ID="6",hostname="dcgm-exporter-wwwjh",container="triton-server",namespace="default",pod="triton-server-f7cb5d77-2w5b2"} 183.657

Reported by @gfrankliu testing this branch on real mixed hardware (A30
MIG + RTX2000 Ada): the pod no longer crashes, but the RTX2000 entity
repeatedly logs "Repairing stale DCGM profiling watch" for the FP64
field and never recovers.

Watch registration was already scoped per GPU model, but the read path
wasn't: fieldsToScrape/latestValues asks DCGM for every configured
field on every entity regardless of which model's group actually
requested it. For an ordinary field that's harmless, DCGM reports one
it doesn't support as a per-entity blank value. For a DCP field it
isn't: asking for a field that was deliberately never watched for that
entity returns DCGM_ST_NOT_WATCHED, and the existing stale-watch repair
logic then retries forever trying to fix a watch that can't exist.

Added ModelSupportsDCPField, caching the per-model answer the same way
the watch-registration path already does, and filter the scrape-time
field list with it before reading an entity's values. The cache clears
on every DCGM reinit via the same queryDCPMetrics call that already
refreshes DCP capability on that cadence.

Signed-off-by: kaushik-kumaran <kaushik.kumaran@ibm.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@CodeBuildder

Copy link
Copy Markdown
Author

@gfrankliu thank you for actually testing this on your mixed A30 MIG + RTX2000 hardware, that's exactly the setup this fix targets and I didn't have access to it myself.

Good news: the core fix works, the pod comes up and stays up. Your logs caught a real follow-up gap though: the RTX2000's repeated "Repairing stale DCGM profiling watch" for the FP64 field.

What was happening: watch registration is now correctly scoped per GPU model, but the read path wasn't. It was still asking DCGM for every configured field on every entity regardless of which model's group actually watched it. For an ordinary field that's harmless, but for a DCP field, asking for one that was deliberately never watched for that entity gets you DCGM_ST_NOT_WATCHED (that's the -16 in your logs), and the existing self-healing logic then retries forever trying to repair a watch that was never meant to exist.

Just pushed a fix for that: the scrape-time field list is now filtered per entity the same way the watch registration already is, so the RTX2000 will stop being asked for a field it was never watched for. Also added test coverage for this specific gap so it can't regress silently.

Would appreciate it if you could pull the latest commit and confirm the repair-loop warnings are gone on your setup when you get a chance. No pressure on timing, just flagging it's ready whenever you want to retest.

@gfrankliu

gfrankliu commented Sep 16, 2026 •

Copy link
Copy Markdown

Thanks @CodeBuildder !

The new commit fixed the startup errors and I can see the DCGM_FI_PROF_* metrics now.

A few observations:

  1. The old version of dcgm-exporter on a A30 only node (running dcgm version 4.5.2-4.8.1-ubuntu22.04) starts up log has one extra line, not sure if it is just because a version change now that we are running 4.6.0-4.8.3-distroless:
time=2026-09-08T05:17:23.610Z level=INFO msg="Profiling endpoints enabled at /debug/pprof/"
  1. sudo dmesg -T prints below after running the new version:
[Wed Sep 16 05:23:27 2026] NVRM: free_os_event: failed to find OS event:
[Wed Sep 16 05:23:27 2026] NVRM: free_os_event:    hParent: 0xc1d775c9
[Wed Sep 16 05:23:27 2026] NVRM: free_os_event:    fd: 17
[Wed Sep 16 05:36:00 2026] NVRM: GPU0 nvCheckOkFailedNoLog: Check failed: Requested object not found [NV_ERR_OBJECT_NOT_FOUND] (0x00000057) returned from gisubscriptionGetGPUInstanceSubscription(pRsClient, RES_GET_HANDLE(pSubdevice), &pGPUInstanceSubscription) @ kernel_mig_manager.c:3105
[Wed Sep 16 05:36:00 2026] NVRM: GPU2 nvCheckOkFailedNoLog: Check failed: Requested object not found [NV_ERR_OBJECT_NOT_FOUND] (0x00000057) returned from gisubscriptionGetGPUInstanceSubscription(pRsClient, RES_GET_HANDLE(pSubdevice), &pGPUInstanceSubscription) @ kernel_mig_manager.c:3105

Again, here are the GPUs on the node:

$ nvidia-smi -L
GPU 0: NVIDIA A30 (UUID: GPU-9a63752e-af8d-fe52-7f0d-968bca6d52e3)
  MIG 1g.6gb      Device  0: (UUID: MIG-c2e1f9b6-a5fd-5698-b5ea-00ebbdd34806)
  MIG 1g.6gb      Device  1: (UUID: MIG-84339efb-1018-5c35-97da-8ba001802bca)
  MIG 1g.6gb      Device  2: (UUID: MIG-0dbd45d4-a7c2-55cf-8586-752fac76e78f)
  MIG 1g.6gb      Device  3: (UUID: MIG-fb5eb194-36a0-55a0-9570-bacbf9f33106)
GPU 1: NVIDIA RTX 2000 Ada Generation (UUID: GPU-fc60126e-514d-3522-2828-c5daa8e99d85)
GPU 2: NVIDIA A30 (UUID: GPU-63bc0b91-2265-1b94-831c-15c41fe92675)
  MIG 1g.6gb      Device  0: (UUID: MIG-b2bc7070-2d13-5b7e-8e95-9ed9e560ba92)
  MIG 1g.6gb      Device  1: (UUID: MIG-9dede469-0ac1-5836-812a-63a42122485b)
  MIG 1g.6gb      Device  2: (UUID: MIG-e6b2eb9c-9cae-515d-9c1c-860272d09193)
  MIG 1g.6gb      Device  3: (UUID: MIG-85d711d6-474a-5856-8cc4-3657be932c31)
$ nvidia-smi --version
NVIDIA-SMI version  : 595.91.07
NVML version        : 595.91
DRIVER version      : 595.91.07
CUDA Version        : 13.2

@CodeBuildder

Copy link
Copy Markdown
Author

Thanks for confirming, glad the follow-up fix actually cleared the repair-loop on your setup.

On your two observations:

  1. The missing "Profiling endpoints enabled at /debug/pprof/" line isn't related to this PR. That's printed only when --enable-pprof is explicitly set (see internal/pkg/server/server.go), so it's a config difference between your two deployments, not a behavior change from this fix.

  2. The dmesg MIG subscription messages (gisubscriptionGetGPUInstanceSubscription, NV_ERR_OBJECT_NOT_FOUND) I can't verify myself, I don't have MIG hardware to reproduce kernel-level driver behavior against. One honest thing worth checking on your end: did you see the same dmesg output on the old version too, or only after switching to this branch? This PR does change how GPU groups get created, specifically, your two A30s now get their own dedicated DCGM group instead of sharing one group with the RTX2000, so it's plausible that changes the GPU-instance subscription pattern DCGM triggers. But I don't want to claim that's the cause without knowing whether this is new behavior or something your node already logs regardless of exporter version.

If it turns out to be new, I'd want to dig into it properly rather than guess further.

@gfrankliu

Copy link
Copy Markdown
  1. Regarding missing "Profiling endpoints enabled at /debug/pprof/" line, thanks for the pointer, and I see the "if condition" was added in 4.5.2-4.8.2 https://github.com/NVIDIA/dcgm-exporter/blob/main/internal/pkg/server/server.go#L116 Prior to that, the message was always printed. That explains why my existing production system running 4.5.2-4.8.1 has the line.

  2. Regarding dmesg, old version dcgm-exporter doesn't have DCGM_FI_PROF_ metrics enabled so not an apple to apple comparison. I couldn't reproduce anymore today, even running the new image built from this branch. Maybe that is just one-off message, depending on other GPU workload running on the same node.

@gfrankliu

Copy link
Copy Markdown

The NVRM message is unrelated. This PR looks fine in my test and can be merged.

@gfrankliu

Copy link
Copy Markdown

@CodeBuildder thanks again. This has been running fine on our mixed A30 (MIG) + RTX2000 nodes. Since 4.8.4 landed, the branch now conflicts with main in internal/pkg/collector/gpu_collector.go and internal/pkg/devicewatcher/device_watcher.go (main added the 127-field group chunking in field_group_split.go). Could you rebase onto 4.8.4? Happy to retest the rebased build on the same hardware.

While reading the diff, I noticed two things that could make profiling metrics quietly disappear in production:

1. The watch path and scrape path handle a GetSupportedMetricGroups error in opposite ways

  • In watchFieldGroupsPartitionedByModel, an error drops all DCP fields for that model (only non-profiling fields get watched).
  • In ModelSupportsDCPField, the same error keeps the field in the scrape list.

So if that query fails for a model, the scrape reads DCP fields that were never watched and gets DCGM_ST_NOT_WATCHED, which kicks off the repair loop this PR is meant to prevent. While the repair is pending, repairProfilingWatch falls back to nonProfilingMetrics, which removes PROF metrics for every GPU on the node, not only the affected one. The error result also isn't cached, so every scrape makes one DCGM call and logs one warning per DCP field per entity. Could the scrape path fail closed on error (or reuse the result from watch time) so the two paths always agree?

2. GPU 0 still decides DCP support for the whole node
queryDCPMetrics still calls GetSupportedMetricGroups(0), and its answer sets CollectDCP and the set of DCP counters (fieldIsSupported) for the whole node. So:

  • Fields GPU 0's model doesn't support are never offered to the other models, even though they'd now be watched per model.
  • If that single query fails (at startup or after any DCGM reinit), DCP collection is turned off node-wide until the next reinit.

On mixed-model nodes the result also depends on which card enumerates as GPU 0. Would it make sense to take the union of supported fields across one representative GPU per model here, since the watch path already filters per model?

Neither of these blocks our current testing, but we'd like to run an official release of this in prod, so I wanted to raise them before merge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4.5.2-4.8.1 fails to start on node with A30 (MIG enabled) and RTX2000

2 participants