Repository navigation
fix(apps/tencentcloud/cache/brc): raise brc requests and pin memory to Guaranteed QoS - #2424
Conversation
…o Guaranteed QoS The two brc replicas were repeatedly evicted with NodeHasInsufficientMemory: the pod's real footprint (~7 CPU / ~12.3Gi) far exceeded its requests (4 CPU / 8Gi), so Karpenter sized the node for the request (8 vCPU / 16GB, ~12.5Gi allocatable) and the pod then overshot the node. Raise the requests to match observed usage and set memory requests == limits so brc gets a Guaranteed memory QoS (kubelet eviction is driven by memory pressure, not CPU). CPU limit stays 16 to allow bursting.
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
There was a problem hiding this comment.
I have already done a preliminary review for you, and I hope to help you do a better job.
Summary
This PR addresses memory eviction issues for the brc pods in Tencent Cloud by increasing the CPU and memory resource requests to match the memory limits, thus enabling Guaranteed QoS for memory and preventing kubelet memory-pressure evictions. The approach is a straightforward update in the release.yaml resource requests, aligning requests to limits for memory and increasing CPU requests to ensure node sizing by Karpenter. The change is minimal, focused, and well-reasoned with a clear explanation of the cause and effect.
Code Improvements
-
Resource Request & Limit Consistency (apps/tencentcloud/cache/brc/release.yaml, lines ~71-78)
While settingrequests.memoryequal tolimits.memoryis a good practice for Guaranteed QoS, the CPU request is increased to 8 while the limit remains 16. This is reasonable based on the explanation, but it would be helpful to add a comment in the YAML manifest to clarify this choice for future maintainers, e.g.:# CPU requests increased to 8 for node sizing; limits kept at 16 for bursting requests: cpu: "8" memory: 16Gi limits: cpu: "16" memory: 16Gi
-
Potential Edge Case: Pod Scheduling on Nodes with Exact Resource Match
By increasing requests to 8 CPU and 16Gi memory, the pod may only schedule on nodes that can provide these resources. If the cluster has limited such nodes, consider documenting or handling fallback scenarios if pod scheduling is delayed or blocked.
Best Practices
-
Documentation in YAML (apps/tencentcloud/cache/brc/release.yaml, lines ~71-78)
Adding inline comments in the resource section explaining the rationale behind equal memory requests and limits will improve readability and maintainability, especially since this is a deliberate fix for evictions. -
Testing / Validation Coverage
The PR description states verification steps but does not mention any automated tests or monitoring changes. Consider adding or mentioning tests (e.g., integration tests, pod stability checks) or updated monitoring alerts to detect if evictions reoccur. -
Consistency in Resource Quantities Formatting
The YAML mixes quoted and unquoted CPU values ("4"→"8"for CPU requests and"16"for limits). For clarity and consistency, use either quoted or unquoted consistently, preferably unquoted for numeric CPU values, e.g.:requests: cpu: 8 memory: 16Gi limits: cpu: 16 memory: 16Gi
Critical Issues
- None identified. The change is minimal and aligns with Kubernetes best practices for resource requests and limits to influence scheduling and evictions.
Summary of actionable suggestions:
- Add comments in
release.yamlnear resource requests/limits explaining the rationale for future maintainers. - Use consistent formatting for CPU resource quantities (prefer unquoted numerics).
- Consider adding or documenting automated tests or monitoring to detect recurrence of evictions or scheduling issues.
- Document any potential scheduling impact due to increased resource requests in cluster capacity planning.
Summary
Fix the repeated
NodeHasInsufficientMemoryevictions of thebrc(bazel-remote) replicas on tencentcloud.The pod's real footprint (~7 CPU / ~12.3Gi) was far above its requests (
4 CPU / 8Gi). Karpenter sizes nodes from the request, so it placedbrcon 8 vCPU / 16GB nodes (~12.5Gi allocatable) and the pod then overshot the node, triggering kubelet memory-pressure eviction:Changes
apps/tencentcloud/cache/brc/release.yaml:requests.cpu:4->8requests.memory:8Gi->16Gilimitsunchanged:cpu: 16,memory: 16GiSetting
requests.memory == limits.memorygives the container a Guaranteed memory QoS, which strongly protects it from node memory-pressure eviction. Kubelet eviction is driven by memory (and disk/PID) pressure, not CPU; CPU oversubscription only causes throttling, so the CPU limit stays at 16 to allow bursting whilerequests.cpu: 8forces Karpenter to provision a node large enough.No change to
replicaCount,sessionAffinity,max_size, or the PVCs.Verification
git diffis limited to the two resource lines.kustomize build apps/tencentcloud/cache/brcrenders the new requests/limits correctly.kubectl -n cache get eventsshould stop showing evictions forbrc.