Skip to content

fix(apps/tencentcloud/cache/brc): raise brc requests and pin memory to Guaranteed QoS - #2424

Merged
wuhuizuo merged 1 commit into
mainfrom
fix/tencentcloud-brc-memory-guaranteed-qos
Sep 24, 2026
Merged

wuhuizuo merged 1 commit into
mainfrom
fix/tencentcloud-brc-memory-guaranteed-qos

Conversation

@wuhuizuo

Copy link
Copy Markdown
Collaborator

Summary

Fix the repeated NodeHasInsufficientMemory evictions of the brc (bazel-remote) replicas on tencentcloud.

The pod's real footprint (~7 CPU / ~12.3Gi) was far above its requests (4 CPU / 8Gi). Karpenter sizes nodes from the request, so it placed brc on 8 vCPU / 16GB nodes (~12.5Gi allocatable) and the pod then overshot the node, triggering kubelet memory-pressure eviction:

Warning  Evicted  pod/brc-1  The node was low on resource: memory.
Container bazel-remote was using 12.4Gi, request is 8Gi

Changes

apps/tencentcloud/cache/brc/release.yaml:

  • requests.cpu: 4 -> 8
  • requests.memory: 8Gi -> 16Gi
  • limits unchanged: cpu: 16, memory: 16Gi

Setting requests.memory == limits.memory gives the container a Guaranteed memory QoS, which strongly protects it from node memory-pressure eviction. Kubelet eviction is driven by memory (and disk/PID) pressure, not CPU; CPU oversubscription only causes throttling, so the CPU limit stays at 16 to allow bursting while requests.cpu: 8 forces Karpenter to provision a node large enough.

No change to replicaCount, sessionAffinity, max_size, or the PVCs.

Verification

  • git diff is limited to the two resource lines.
  • kustomize build apps/tencentcloud/cache/brc renders the new requests/limits correctly.
  • After Flux reconciles: pods should land on larger nodes and kubectl -n cache get events should stop showing evictions for brc.

…o Guaranteed QoS

The two brc replicas were repeatedly evicted with NodeHasInsufficientMemory:
the pod's real footprint (~7 CPU / ~12.3Gi) far exceeded its requests
(4 CPU / 8Gi), so Karpenter sized the node for the request (8 vCPU / 16GB,
~12.5Gi allocatable) and the pod then overshot the node.

Raise the requests to match observed usage and set memory requests == limits
so brc gets a Guaranteed memory QoS (kubelet eviction is driven by memory
pressure, not CPU). CPU limit stays 16 to allow bursting.
@ti-chi-bot

ti-chi-bot Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign dillon-zheng for approval. For more information see the Code Review Process.
Please ensure that each of them provides their approval before proceeding.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@ti-chi-bot ti-chi-bot Bot added the area/apps label Sep 24, 2026

@ti-chi-bot ti-chi-bot Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have already done a preliminary review for you, and I hope to help you do a better job.

Summary
This PR addresses memory eviction issues for the brc pods in Tencent Cloud by increasing the CPU and memory resource requests to match the memory limits, thus enabling Guaranteed QoS for memory and preventing kubelet memory-pressure evictions. The approach is a straightforward update in the release.yaml resource requests, aligning requests to limits for memory and increasing CPU requests to ensure node sizing by Karpenter. The change is minimal, focused, and well-reasoned with a clear explanation of the cause and effect.


Code Improvements

  • Resource Request & Limit Consistency (apps/tencentcloud/cache/brc/release.yaml, lines ~71-78)
    While setting requests.memory equal to limits.memory is a good practice for Guaranteed QoS, the CPU request is increased to 8 while the limit remains 16. This is reasonable based on the explanation, but it would be helpful to add a comment in the YAML manifest to clarify this choice for future maintainers, e.g.:

    # CPU requests increased to 8 for node sizing; limits kept at 16 for bursting
    requests:
      cpu: "8"
      memory: 16Gi
    limits:
      cpu: "16"
      memory: 16Gi
  • Potential Edge Case: Pod Scheduling on Nodes with Exact Resource Match
    By increasing requests to 8 CPU and 16Gi memory, the pod may only schedule on nodes that can provide these resources. If the cluster has limited such nodes, consider documenting or handling fallback scenarios if pod scheduling is delayed or blocked.


Best Practices

  • Documentation in YAML (apps/tencentcloud/cache/brc/release.yaml, lines ~71-78)
    Adding inline comments in the resource section explaining the rationale behind equal memory requests and limits will improve readability and maintainability, especially since this is a deliberate fix for evictions.

  • Testing / Validation Coverage
    The PR description states verification steps but does not mention any automated tests or monitoring changes. Consider adding or mentioning tests (e.g., integration tests, pod stability checks) or updated monitoring alerts to detect if evictions reoccur.

  • Consistency in Resource Quantities Formatting
    The YAML mixes quoted and unquoted CPU values ("4" → "8" for CPU requests and "16" for limits). For clarity and consistency, use either quoted or unquoted consistently, preferably unquoted for numeric CPU values, e.g.:

    requests:
      cpu: 8
      memory: 16Gi
    limits:
      cpu: 16
      memory: 16Gi

Critical Issues

  • None identified. The change is minimal and aligns with Kubernetes best practices for resource requests and limits to influence scheduling and evictions.

Summary of actionable suggestions:

  • Add comments in release.yaml near resource requests/limits explaining the rationale for future maintainers.
  • Use consistent formatting for CPU resource quantities (prefer unquoted numerics).
  • Consider adding or documenting automated tests or monitoring to detect recurrence of evictions or scheduling issues.
  • Document any potential scheduling impact due to increased resource requests in cluster capacity planning.

@wuhuizuo
wuhuizuo merged commit d25c401 into main Sep 24, 2026
5 checks passed
@wuhuizuo
wuhuizuo deleted the fix/tencentcloud-brc-memory-guaranteed-qos branch September 24, 2026 09:44
@ti-chi-bot ti-chi-bot Bot added the size/XS label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant