verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite of what its own contract says. A non-empty file means something failed, so the function returned success exactly when a workload never came up, and failure when everything was fine. Every rollback was therefore skipped, and every deploy that changed anything ended red with an empty failure list and a bogus "Rolled back successfully". Worse, the check only ever ran at the end of stage_apply_k8s, inside the same process as the apply. A job killed by timeout-minutes, cancelled by a new push, or cut off by a dropped SSH connection never reached it, which is precisely when a rollback matters. The three helm upgrades alone can consume the whole 30-minute job budget, so that path was reachable. Verification now lives in its own job, gated on always(), so it runs whatever happened to the apply. The apply stage publishes its pre-apply snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything, and the verify stage picks it up from there. A snapshot whose recorded commit does not match the deploy is refused rather than trusted, so a stale pointer from an earlier run cannot make the rollback revert the wrong workloads. An unwritable snapshot directory now fails the deploy up front instead of silently continuing without a way back. cancel-in-progress becomes false for the same reason: cancelling a run kills the apply job and takes the verify job with it, which is the failure this change exists to prevent. Both applies are idempotent, so queueing costs little. The SSH key moves to a per-run directory removed on exit, and the deploy is pinned to the exact commit CI validated.
31 lines
1.1 KiB
Bash
Executable File
31 lines
1.1 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# usage: ssh-run.sh <stage>
|
|
# Runs one deploy-lib.sh stage on the workstation over SSH.
|
|
set -euo pipefail
|
|
|
|
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
|
|
: "${DEPLOY_USER:?missing DEPLOY_USER}"
|
|
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
|
|
|
|
deploy_port="${DEPLOY_PORT:-22}"
|
|
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
|
|
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
|
|
|
|
# The private key is written to a per-run directory that is removed on exit, so a
|
|
# failed or cancelled job cannot leave deploy credentials in the runner's temp
|
|
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
|
|
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
|
|
trap 'rm -rf "$key_dir"' EXIT INT TERM
|
|
|
|
ssh_key="$key_dir/deploy_key"
|
|
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
|
|
chmod 600 "$ssh_key"
|
|
|
|
ssh -i "$ssh_key" -p "$deploy_port" \
|
|
-o BatchMode=yes -o StrictHostKeyChecking=accept-new \
|
|
"${DEPLOY_USER}@${DEPLOY_HOST}" \
|
|
"REPO=$deploy_path APPLY_PRUNE=${APPLY_PRUNE:-false} DEPLOY_SHA=${DEPLOY_SHA:-} STAGE=$1 bash -se" <<'EOF'
|
|
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
|
run_stage "$STAGE"
|
|
EOF
|