A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation.
The file is empty.