Skip to content

Alive check - #4600

Open
CarensirA-MF wants to merge 5 commits into
mainfrom
alive-check
Open

CarensirA-MF wants to merge 5 commits into
mainfrom
alive-check

Conversation

@CarensirA-MF

@CarensirA-MF CarensirA-MF commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

My understanding web liveness probe was kind of answering the wrong question. It hits /landing/, which needs Redis, the database pool and a template render, so it reports "is Postgres fast right now?" when the cluster aganet is asking "would restarting this process help?".

Based on the opservastions (via k9s) every pod was blocked on the same exhausted pool, every probe failed within the same minute, and restarted the web pods in a wave. Each restart dropped in-flight requests and took up to 200 s to return, while the remaining pods absorbed the load.

What changes

/alive on Root, with tools.sessions, tools.oidc and tools.reset_threadlocal off for the path, the same exemptions the static paths use. No template render either way.

liveness_checks_dependencies, default on: /alive pings Redis through the session client and runs SELECT 1 on a pooled connection. Either failing returns 503. Under pool exhaustion the connection checkout waits the pool timeout, which exceeds the probe timeout, so a stuck pod fails the probe and is restarted. That matches the current behavior minus the page render.

Off: /alive returns 200 as long as CherryPy can hand the request to a thread.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant