fix(immich): size postgres probes for HDD stalls
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 13s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 13s
Postmaster was SIGKILLed in a loop: 70s fsync stalls on the loaded rotational disk outlasted the 5-minute startup budget and the 60s liveness tolerance, and every kill bought another full WAL replay. Startup budget 15min, liveness 5x60s. Already applied live with kubectl; this keeps git in sync.
This commit is contained in:
1 parent
a2af854867
commit
b6e1dc0362
1 file changed
+10
-3
@@ -90,10 +90,16 @@ spec:
|
|||||||
# rotational disk on a loaded single node, and pg_isready can take
|
# rotational disk on a loaded single node, and pg_isready can take
|
||||||
# seconds during WAL recovery. A 1s timeout kills the container
|
# seconds during WAL recovery. A 1s timeout kills the container
|
||||||
# mid-recovery and restarts the spiral.
|
# mid-recovery and restarts the spiral.
|
||||||
|
#
|
||||||
|
# Budgets are sized for HDD stalls, not for a healthy disk: fsync of
|
||||||
|
# a single file was observed taking 70s under node IO pressure, so
|
||||||
|
# the startup budget is 15 minutes and liveness tolerates 5 minutes
|
||||||
|
# of unresponsiveness. Killing a stalled-but-healthy postmaster only
|
||||||
|
# buys another full WAL replay, which is more IO, not less.
|
||||||
startupProbe:
|
startupProbe:
|
||||||
exec:
|
exec:
|
||||||
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
||||||
failureThreshold: 60
|
failureThreshold: 180
|
||||||
periodSeconds: 5
|
periodSeconds: 5
|
||||||
timeoutSeconds: 5
|
timeoutSeconds: 5
|
||||||
readinessProbe:
|
readinessProbe:
|
||||||
@@ -105,8 +111,9 @@ spec:
|
|||||||
exec:
|
exec:
|
||||||
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
||||||
initialDelaySeconds: 30
|
initialDelaySeconds: 30
|
||||||
periodSeconds: 20
|
periodSeconds: 60
|
||||||
timeoutSeconds: 5
|
timeoutSeconds: 10
|
||||||
|
failureThreshold: 5
|
||||||
# The image template sets shared_buffers to 512MB, and the vchord and
|
# The image template sets shared_buffers to 512MB, and the vchord and
|
||||||
# vectors workers are Rust binaries with a real RSS footprint on top
|
# vectors workers are Rust binaries with a real RSS footprint on top
|
||||||
# of postmaster, checkpointer and friends. 1Gi was enough to start
|
# of postmaster, checkpointer and friends. 1Gi was enough to start
|
||||||
|
|||||||
Reference in new issue
Block a user