Skip to content

docs(helm): document nvidia.com/gpu toleration override - #733

Closed
KR-Ravindra wants to merge 1 commit into
NVIDIA:mainfrom
KR-Ravindra:docs/gpu-taint-toleration
Closed

KR-Ravindra wants to merge 1 commit into
NVIDIA:mainfrom
KR-Ravindra:docs/gpu-taint-toleration

Conversation

@KR-Ravindra

@KR-Ravindra KR-Ravindra commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Installing the chart with default values on a cluster whose GPU nodes are tainted nvidia.com/gpu=true:NoSchedule (EKS/GKE GPU node groups, Karpenter/Cluster Autoscaler GPU pools, and the taint the NVIDIA device plugin chart tolerates by default) gives a DaemonSet with DESIRED 0: the release succeeds, no exporter pods run, and neither deployment/values.yaml nor the README mentions taints. #732 proposed adding the toleration to the defaults; the maintainer feedback there was that this is a breaking change and that the override should be documented instead, so this PR only documents it.

Fix

  • deployment/values.yaml: a comment block above tolerations explaining that a user-supplied list replaces the default rather than extending it, with an example that adds the nvidia.com/gpu toleration. The stray #- operator: Exists example line below the list is folded into the same comment. The default list itself is unchanged.
  • README.md: a "Scheduling on tainted GPU nodes" subsection at the end of the Kubernetes quickstart with the same values snippet and one sentence on why the DaemonSet otherwise schedules no pods on those nodes.
  • deployment/README.md: a "Scheduling on Tainted GPU Nodes" section under Configuration with the same snippet, so the chart's own README stays aligned with values.yaml.

No template, raw manifest, test, or default value changes, and no chart version bump.

How tested

  • helm lint deployment/: 1 chart linted, 0 failed.
  • helm template test deployment/ output is byte-identical before and after this change (only comments changed in values.yaml).
  • helm template test deployment/ -f values.yaml with the documented override renders both the control-plane and nvidia.com/gpu tolerations on the DaemonSet, so the snippet does what the text says.

Links

This change was prepared with an AI agent operated by KR-Ravindra, who reviewed it.

@KR-Ravindra
KR-Ravindra marked this pull request as ready for review September 9, 2026 18:30
@KR-Ravindra

Copy link
Copy Markdown
Contributor Author

@nccurry this is the documentation-only follow-up from #732: defaults untouched, the README quickstart and the values.yaml comment now show the nvidia.com/gpu toleration override and note that a set list replaces the default. helm template output is byte-identical to main.

@iacker

iacker commented Sep 16, 2026

Copy link
Copy Markdown

The values.yaml comment reads well and the replace-not-merge behaviour matches the template.

One gap: deployment/README.md is the chart's own doc and has no tolerations section. deployment/AGENTS.md asks to keep values.yaml and README snippets aligned, so that file looks like the natural place for this too.

The chart default tolerations only cover the control-plane taint, so on
clusters that taint GPU nodes with nvidia.com/gpu the DaemonSet schedules
no pods and nothing in the chart or README says why. Add a values.yaml
comment and a README subsection showing the override. Defaults unchanged.

Follow-up to NVIDIA#732. Related to NVIDIA#731.

Signed-off-by: KR Ravindra <42912207+KR-Ravindra@users.noreply.github.com>
@KR-Ravindra
KR-Ravindra force-pushed the docs/gpu-taint-toleration branch from 03c0441 to eed7817 Compare September 16, 2026 09:43
@KR-Ravindra

Copy link
Copy Markdown
Contributor Author

@iacker done in eed7817: added a "Scheduling on Tainted GPU Nodes" section to deployment/README.md (under Configuration) with the same override snippet, so the chart README, values.yaml comment and top-level README now say the same thing. helm lint passes and helm template output is still byte-identical to main. Thanks for catching it.

@nccurry

nccurry commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Set Helm tolerations for nvidia.com/gpu with operator: Exists and effect: NoSchedule. This is documented in 4.8.4. Closing.

@nccurry nccurry closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants