docs(helm): document nvidia.com/gpu toleration override - #733
KR-Ravindra wants to merge 1 commit into
Conversation
|
The values.yaml comment reads well and the replace-not-merge behaviour matches the template. One gap: deployment/README.md is the chart's own doc and has no tolerations section. deployment/AGENTS.md asks to keep values.yaml and README snippets aligned, so that file looks like the natural place for this too. |
The chart default tolerations only cover the control-plane taint, so on clusters that taint GPU nodes with nvidia.com/gpu the DaemonSet schedules no pods and nothing in the chart or README says why. Add a values.yaml comment and a README subsection showing the override. Defaults unchanged. Follow-up to NVIDIA#732. Related to NVIDIA#731. Signed-off-by: KR Ravindra <42912207+KR-Ravindra@users.noreply.github.com>
03c0441 to
eed7817
Compare
|
@iacker done in eed7817: added a "Scheduling on Tainted GPU Nodes" section to deployment/README.md (under Configuration) with the same override snippet, so the chart README, values.yaml comment and top-level README now say the same thing. helm lint passes and helm template output is still byte-identical to main. Thanks for catching it. |
|
Set Helm |
Problem
Installing the chart with default values on a cluster whose GPU nodes are tainted
nvidia.com/gpu=true:NoSchedule(EKS/GKE GPU node groups, Karpenter/Cluster Autoscaler GPU pools, and the taint the NVIDIA device plugin chart tolerates by default) gives a DaemonSet withDESIRED 0: the release succeeds, no exporter pods run, and neitherdeployment/values.yamlnor the README mentions taints. #732 proposed adding the toleration to the defaults; the maintainer feedback there was that this is a breaking change and that the override should be documented instead, so this PR only documents it.Fix
deployment/values.yaml: a comment block abovetolerationsexplaining that a user-supplied list replaces the default rather than extending it, with an example that adds thenvidia.com/gputoleration. The stray#- operator: Existsexample line below the list is folded into the same comment. The default list itself is unchanged.README.md: a "Scheduling on tainted GPU nodes" subsection at the end of the Kubernetes quickstart with the same values snippet and one sentence on why the DaemonSet otherwise schedules no pods on those nodes.deployment/README.md: a "Scheduling on Tainted GPU Nodes" section under Configuration with the same snippet, so the chart's own README stays aligned withvalues.yaml.No template, raw manifest, test, or default value changes, and no chart version bump.
How tested
helm lint deployment/: 1 chart linted, 0 failed.helm template test deployment/output is byte-identical before and after this change (only comments changed invalues.yaml).helm template test deployment/ -f values.yamlwith the documented override renders both the control-plane andnvidia.com/gputolerations on the DaemonSet, so the snippet does what the text says.Links
nvidia.com/gpu, so the DaemonSet schedules zero pods on tainted GPU nodes #731This change was prepared with an AI agent operated by KR-Ravindra, who reviewed it.