`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.
Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
is a Prometheus endpoint, not a readiness signal.
The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.
Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
Keep the existing Redis PVC and data while migrating the edu-master workload from Deployment to StatefulSet.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- restore-seed-job.yaml.example: translate runbook to English
- secrets.yaml.example: translate section headers to English
- webinar-checker.yaml: translate initContainer dependency-order comments to English
Co-authored-by: assistant