fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works.
This commit is contained in:
1 parent
f49d91b63d
commit
f54589a05c
2 files changed
+7
-2
No files matched your search
@@ -11,7 +11,12 @@ APPLY_PRUNE="${APPLY_PRUNE:-false}"
|
||||
# Commit CI validated. Empty for a manual workflow_dispatch, which falls back to
|
||||
# the current origin/main.
|
||||
DEPLOY_SHA="${DEPLOY_SHA:-}"
|
||||
DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-/var/backups/homelab-deploy}"
|
||||
# Handoff point between the apply stage (writes) and the verify stage (reads).
|
||||
# Under the deploy user's own XDG state directory rather than /var/backups: the
|
||||
# deploy is unprivileged, /var/backups does not exist on a minimal Arch host, and
|
||||
# creating it would need root — which is why the first real deploy died here with
|
||||
# "is not writable" before touching a single workload. $HOME comes from sshd.
|
||||
DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state}/homelab-deploy}"
|
||||
# Per-workload rollout budget and how many workloads to watch at once. The whole
|
||||
# apply job has its own timeout-minutes as a backstop.
|
||||
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
|
||||
|
||||
@@ -24,7 +24,7 @@ chmod 600 "$ssh_key"
|
||||
ssh -i "$ssh_key" -p "$deploy_port" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=accept-new \
|
||||
"${DEPLOY_USER}@${DEPLOY_HOST}" \
|
||||
"REPO=$deploy_path APPLY_PRUNE=${APPLY_PRUNE:-false} DEPLOY_SHA=${DEPLOY_SHA:-} STAGE=$1 bash -se" <<'EOF'
|
||||
"REPO=$deploy_path APPLY_PRUNE=${APPLY_PRUNE:-false} DEPLOY_SHA=${DEPLOY_SHA:-} DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-} STAGE=$1 bash -se" <<'EOF'
|
||||
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
||||
run_stage "$STAGE"
|
||||
EOF
|
||||
Reference in new issue
Block a user