fix(postgres): give probes room on an I/O-bound single node
renovate-ci / validate-renovate (push) Canceled after 21s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 10s
ci / lint-yaml (push) Successful in 4s
ci / lint-dockerfiles (push) Successful in 4s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 21s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 22s

pg_isready with the 1s default times out under I/O stall and kubelet kills a healthy postgres mid-recovery; each kill restarts a multi-minute fsync from zero and loops forever. readiness/liveness timeout 5s, liveness threshold 5, startup budget 15min.
This commit is contained in:
forust committed 2026-09-28 12:35:16 +02:00
1 parent cda0022d81
commit 2f891a5d31
1 file changed
+11 -1
+11 -1
View File
@@ -60,16 +60,26 @@ spec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 10
periodSeconds: 10
# Generous timeout: on an I/O-bound single node even exec can take
# seconds, and a 1s default kills a healthy postgres mid-recovery.
timeoutSeconds: 5
startupProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
failureThreshold: 30
# Crash recovery on an I/O-starved single node can fsync for 10+
# minutes; killing postgres mid-recovery restarts the fsync from
# zero and loops forever. 90x10s = 15 minutes of grace.
failureThreshold: 90
periodSeconds: 10
livenessProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 30
periodSeconds: 20
# Same I/O reasoning as readiness, plus more misses before a kill:
# restarting postgres on a loaded node only makes recovery longer.
timeoutSeconds: 5
failureThreshold: 5
resources:
requests:
memory: "512Mi"