ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 13s
Postmaster was SIGKILLed in a loop: 70s fsync stalls on the loaded rotational disk outlasted the 5-minute startup budget and the 60s liveness tolerance, and every kill bought another full WAL replay. Startup budget 15min, liveness 5x60s. Already applied live with kubectl; this keeps git in sync.
146 lines
5.6 KiB
YAML
146 lines
5.6 KiB
YAML
# Immich's own database, separate from the shared postgres in the database
|
|
# namespace. It has to be separate: v3 checks the VectorChord version at startup
|
|
# and refuses to boot without it, VectorChord needs its .so in
|
|
# shared_preload_libraries, and that can only be read when postmaster starts.
|
|
# So the shared instance would have to be rebuilt on a custom image carrying
|
|
# vchord and restarted - for every consumer of it (authentik, gitea, netbox,
|
|
# netronome, penpot, statuspage). Not worth it for one photo library.
|
|
apiVersion: v1
|
|
kind: Service
|
|
metadata:
|
|
name: immich-postgres
|
|
namespace: immich
|
|
labels:
|
|
app: immich-postgres
|
|
spec:
|
|
selector:
|
|
app: immich-postgres
|
|
ports:
|
|
- name: postgres
|
|
port: 5432
|
|
targetPort: postgres
|
|
---
|
|
apiVersion: apps/v1
|
|
kind: StatefulSet
|
|
metadata:
|
|
name: immich-postgres
|
|
namespace: immich
|
|
labels:
|
|
app: immich-postgres
|
|
spec:
|
|
serviceName: immich-postgres
|
|
replicas: 1
|
|
selector:
|
|
matchLabels:
|
|
app: immich-postgres
|
|
template:
|
|
metadata:
|
|
labels:
|
|
app: immich-postgres
|
|
spec:
|
|
containers:
|
|
- name: postgres
|
|
# v3.x expects vchord for its vector work and vectors (pgvecto.rs)
|
|
# for some index types. This image ships both and preloads them, plus
|
|
# its own shared_buffers and wal settings, through
|
|
# /etc/postgresql/postgresql.conf - which its entrypoint reaches via
|
|
# `postgres -c config_file=...` in the image CMD.
|
|
#
|
|
# So there is deliberately no `command:` here. Overriding it replaces
|
|
# that config_file, and it also loses the step where the entrypoint
|
|
# drops from root to the postgres user: postmaster then starts as
|
|
# root and refuses to run.
|
|
image: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0@sha256:1a078b237c1d9b420b0ee59147386b4aa60d3a07a8e6a402fc84a57e41b043a4
|
|
env:
|
|
- name: POSTGRES_USER
|
|
value: immich
|
|
- name: POSTGRES_DB
|
|
value: immich
|
|
- name: POSTGRES_PASSWORD
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: immich-secrets
|
|
key: DB_PASSWORD
|
|
# Only read when the data directory is empty, so the checksums are
|
|
# decided here and never again.
|
|
- name: POSTGRES_INITDB_ARGS
|
|
value: --data-checksums
|
|
# The postgres-data volume lives on sdc, which is rotational. The
|
|
# SSD template is the default; HDD only changes the planner costs
|
|
# (effective_io_concurrency, random_page_cost), nothing structural.
|
|
- name: DB_STORAGE_TYPE
|
|
value: HDD
|
|
- name: TZ
|
|
valueFrom:
|
|
configMapKeyRef:
|
|
name: immich-config
|
|
key: TZ
|
|
ports:
|
|
- name: postgres
|
|
containerPort: 5432
|
|
volumeMounts:
|
|
- name: postgres-data
|
|
mountPath: /var/lib/postgresql/data
|
|
# The upstream compose file asks docker for 128mb of shm. Kubernetes
|
|
# gives every container 64mb, which is not what postmaster expects
|
|
# for parallel query workers and the WAL writer.
|
|
- name: shm
|
|
mountPath: /dev/shm
|
|
# Probes use a generous timeout on purpose: the data lives on a
|
|
# rotational disk on a loaded single node, and pg_isready can take
|
|
# seconds during WAL recovery. A 1s timeout kills the container
|
|
# mid-recovery and restarts the spiral.
|
|
#
|
|
# Budgets are sized for HDD stalls, not for a healthy disk: fsync of
|
|
# a single file was observed taking 70s under node IO pressure, so
|
|
# the startup budget is 15 minutes and liveness tolerates 5 minutes
|
|
# of unresponsiveness. Killing a stalled-but-healthy postmaster only
|
|
# buys another full WAL replay, which is more IO, not less.
|
|
startupProbe:
|
|
exec:
|
|
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
|
failureThreshold: 180
|
|
periodSeconds: 5
|
|
timeoutSeconds: 5
|
|
readinessProbe:
|
|
exec:
|
|
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
|
periodSeconds: 10
|
|
timeoutSeconds: 5
|
|
livenessProbe:
|
|
exec:
|
|
command: ["sh", "-c", "pg_isready -U immich -d immich"]
|
|
initialDelaySeconds: 30
|
|
periodSeconds: 60
|
|
timeoutSeconds: 10
|
|
failureThreshold: 5
|
|
# The image template sets shared_buffers to 512MB, and the vchord and
|
|
# vectors workers are Rust binaries with a real RSS footprint on top
|
|
# of postmaster, checkpointer and friends. 1Gi was enough to start
|
|
# the server but the vectors worker kept dying in it, so the limit
|
|
# sits at 2Gi. The request stays at the idle cost.
|
|
resources:
|
|
requests:
|
|
cpu: "50m"
|
|
memory: "256Mi"
|
|
limits:
|
|
cpu: "1000m"
|
|
memory: "2Gi"
|
|
volumes:
|
|
- name: shm
|
|
emptyDir:
|
|
medium: Memory
|
|
sizeLimit: 128Mi
|
|
volumeClaimTemplates:
|
|
- metadata:
|
|
name: postgres-data
|
|
spec:
|
|
accessModes: ["ReadWriteOnce"]
|
|
# Retain: this is the metadata for a library that only exists in one
|
|
# place, and local-path cannot expand a bound volume, so this size has
|
|
# to hold until the library is rebuilt or dumped elsewhere.
|
|
storageClassName: local-path-retain
|
|
resources:
|
|
requests:
|
|
storage: 32Gi
|