Compare commits

..
Author SHA1 Message Date
forust 5c8bc15e60 docs(reloader): describe workload opt-in and reload policy
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 9s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 6s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 10s
ci / lint-actionlint (pull_request) Successful in 5s
ci / lint-shellcheck (pull_request) Successful in 8s
ci / lint-prettier (pull_request) Successful in 16s
ci / lint-ruff (pull_request) Successful in 7s
ci / lint-yaml (pull_request) Successful in 10s
ci / lint-dockerfiles (pull_request) Successful in 7s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
2026-10-06 16:14:04 +02:00
forust 3c4732ae20 docs: document homelab services, deployment, and repository review 2026-10-06 16:09:50 +02:00
125 changed files with 2617 additions and 3452 deletions

No files matched your search

+88
View File
@@ -0,0 +1,88 @@
# Build and deployment workflows
Gitea Actions checks this repository, builds its custom images, and deploys
selected services to the workstation. Workflows use the self-hosted runner labels
`linux`, `arch`, and `homelab`; deployment jobs also require `prod`.
## Checks
`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff,
yamllint, hadolint, and kubeconform. Tool versions are pinned in
`workflows/tool-versions.env` and installed by `install-ci-tools.sh`.
Compose CI checks structure without resolving local environment files or paths.
On the reviewed main commit it only discovers standard filenames; the
`fix/deploy-validation` branch adds the manual Compose entry points too.
Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main,
CI also attempts server-side dry-runs for marked services; these require an
existing namespace and contact the cluster's admission webhooks. A cluster that
is unreachable produces a warning and skips that CI pass. Deploy validation has
its own dry-run stage.
`renovate-ci.yaml` validates Renovate settings and checks that its generated
ConfigMap matches `renovate/renovate.json`.
## Image builds
CI builds changed custom images for `errorpages`, both `homepages` variants, and
the two `edu_master` Python services. Main builds publish `main`, `prod`, and a
commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest
checks.
Kubernetes deployment resolves the lab's own registry images to digests, preferring
commit-specific tags. Third-party image versions remain declared in the manifests.
## Deploy selection
`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the
workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection
uses an `active` file beside a standard `compose.yaml` or `compose.yml`.
Kustomize overlays are supported, although the current tree primarily contains
plain manifests.
Secret files, examples, Helm values, and patch files are excluded from plain
manifest selection. Create local Kubernetes Secrets separately in their target
namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each
release controlled by its configured marker. Other charts need separate setup.
## Trigger and required settings
Automatic deployment follows a successful main CI run when the repository Actions
variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch
and targets the fetched main branch when no validated commit SHA is provided.
A manual dispatch does not prove that this commit passed CI.
Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and,
where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses
`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git,
Docker, kubectl, Helm, jq, and the state directory used for snapshots.
Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop
is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied.
Preflight fetches and resets the remote checkout. It refuses when tracked files
have local changes; ignored local env and Secret files stay in place. Do not use
a development checkout with uncommitted tracked changes as the deployment target.
## Stages and recovery
1. Preflight fetches the target commit and checks the remote working tree.
2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and
checks referenced Secrets.
3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases,
applies resources, and refreshes owned custom images.
4. Apply Compose recreates marked stacks and checks container state.
5. Verify Kubernetes checks changed workloads and attempts rollback for failures.
6. Smoke probes public routes after verification.
The two apply jobs share a remote lock. Workflow concurrency queues deployments
rather than interrupting an older apply. Snapshots live under
`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default.
They contain the pre-apply workload data and commit identifier.
Rollback uses workload revisions. It does not restore ConfigMaps, Secrets,
database schemas, or data. Helm-owned workloads are handled through the Helm
upgrade's rollback path; the generic rollback skips them. Compose has no automatic
rollback. See the [review](../docs/repository-review.md) for remaining recovery
limitations, including SSH retries and serial rollback timing.
-3
View File
@@ -1,3 +0,0 @@
{
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
}
-122
View File
@@ -1,122 +0,0 @@
# Homelab CI/CD
The native Gitea runner runs on **vps**; production runs on **workstation**.
Jobs run on `homelab:host`, one at a time. No job images or Kubernetes credentials
are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
From this checkout on the VPS:
```sh
sudo bash .gitea/runner/setup-runner.sh
```
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
For a new host, install the same runner binary and register as `gitea-runner`
using the registration token interactively, label `homelab:host`, and working
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
## Workstation setup
As the existing SSH deploy user on workstation:
```sh
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
```
The controller uses `/srv/homelab` as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
`~/.config/homelab-deploy/environment`. Check these before installing.
Configure Gitea Actions Variables:
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
## Releases and deployment
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. EDU images remain pinned to the
digests released by their application repository. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA:
- `full`: required for the first baseline; reconcile all active components.
- `changed`: compare with the last fully successful production deploy.
- `plan`: validate configuration and show selection without changing production resources.
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
## Status and recovery
`--retry` repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
```sh
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
```
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using `compose-before/<stack>.json`, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures.
Update the workstation dispatcher only when no deploy is running.
## Validation and migration rollback
```sh
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
```
Test on a separate namespace before the initial production `full` run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
-11
View File
@@ -1,11 +0,0 @@
[worker.oci]
gc = true
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
[[worker.oci.gcpolicy]]
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
all = true
-10
View File
@@ -1,10 +0,0 @@
runner:
file: /var/lib/gitea-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab:host
cache:
enabled: false
container:
docker_host: unix:///var/run/docker.sock
-18
View File
@@ -1,18 +0,0 @@
[Unit]
Description=Gitea Actions runner
After=network-online.target docker.service
Wants=network-online.target
[Service]
User=gitea-runner
Group=gitea-runner
SupplementaryGroups=docker
WorkingDirectory=/var/lib/gitea-runner
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
Restart=on-failure
RestartSec=5
UMask=0077
[Install]
WantedBy=multi-user.target
-12
View File
@@ -1,12 +0,0 @@
[Unit]
Description=Homelab deploy %i
[Service]
Type=exec
EnvironmentFile=%h/.config/homelab-deploy/environment
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
RuntimeMaxSec=5h
TimeoutStopSec=135min
KillMode=control-group
UMask=0077
-48
View File
@@ -1,48 +0,0 @@
#!/usr/bin/env bash
# Native host runner, with pinned user-space tools and no extra CI images.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in docker curl python3 git tar xz flock runuser systemctl; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
docker info >/dev/null
docker compose version >/dev/null
docker buildx version >/dev/null
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
# Reuse the established service account and runner registration.
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
mkdir -p /etc/gitea-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
chmod 755 "$scratch"
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
python3 - <<'PYLABELS'
import json
from pathlib import Path
registration = Path('/var/lib/gitea-runner/.runner')
if registration.exists():
labels = json.loads(registration.read_text()).get('labels', [])
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
labels.append('homelab:host')
config = Path('/etc/gitea-runner/config.yaml')
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
PYLABELS
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
if [ ! -f /var/lib/gitea-runner/.runner ]; then
echo 'Register once as gitea-runner with homelab:host before starting the service.'
exit 0
fi
systemctl daemon-reload
systemctl enable --now gitea-runner.service
systemctl restart gitea-runner.service
echo "Runner ready. Configuration backups: *.before-$stamp"
-32
View File
@@ -1,32 +0,0 @@
#!/usr/bin/env bash
# Run as the existing deploy user on workstation. Never resets the working tree.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="${HOMELAB_REPO:-/srv/homelab}"
for tool in python3 git kubectl helm docker flock timeout; do
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
done
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
echo "Run once: sudo loginctl enable-linger $USER" >&2
exit 1
fi
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
chmod 700 "$config"
if [ ! -f "$config/environment" ]; then
context="$(kubectl config current-context)"
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
chmod 600 "$config/environment"
fi
# Do not replace a dispatcher while an existing deploy uses it.
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
echo 'An existing deploy is running; wait before updating the controller' >&2
exit 1
fi
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
systemctl --user daemon-reload
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
-94
View File
@@ -1,94 +0,0 @@
#!/usr/bin/env bash
# Local regressions only: kubectl is mocked and Docker is used for config parsing.
set -euo pipefail
repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate"
git -C "$scratch/repo" init -q
for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do
touch "$scratch/repo/$file"
done
git -C "$scratch/repo" add .
# shellcheck source=../workflows/compose-lint.sh
source "$repo/.gitea/workflows/compose-lint.sh"
actual="$(cd "$scratch/repo" && compose_files)"
expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml'
[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; }
cat >"$scratch/compose.yaml" <<'YAML'
services:
example:
image: busybox:1.37.0
environment:
REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression}
YAML
unset HOMELAB_TEST_REQUIRED
if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then
echo 'Full Compose validation accepted a missing variable' >&2
exit 1
fi
grep -q 'required for this regression' "$scratch/config.log"
HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml"
cat >"$scratch/resources.json" <<'JSON'
{"kind":"List","items":[
{"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{
"containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}],
"initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}],
"imagePullSecrets":[{"name":"registry"}],
"volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}]
}}}},
{"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}},
{"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}}
]}
JSON
actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)"
expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron'
[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; }
REPO="$repo"
# shellcheck source=../workflows/deploy-lib.sh
source "$repo/.gitea/workflows/deploy-lib.sh"
K8S_MANIFESTS=("$scratch/resources.json")
KUSTOMIZE_APPS=()
# No live cluster access. Reject credentials in app even if they exist elsewhere.
kubectl() {
case "$1" in
create) cat "$scratch/resources.json" ;;
get)
if [ "$3" = credentials ] && [ "$5" = app ]; then
return 1
fi
return 0
;;
*) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;;
esac
}
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Namespace-scoped Secret check accepted a missing Secret' >&2
exit 1
fi
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
# API/rendering errors must not produce an empty reference list and pass.
kubectl() { return 1; }
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
exit 1
fi
kubectl() { return 0; }
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight skipped an installed CRD' >&2
exit 1
fi
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
echo 'VMAgent preflight skipped an unrelated manifest' >&2
exit 1
fi
kubectl() { return 1; }
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2
exit 1
fi
printf '%s\n' 'Deploy validation regressions passed.'
+410 -33
View File
@@ -1,29 +1,38 @@
name: ci
"on":
on:
push:
branches:
- main
pull_request: null
workflow_dispatch: null
- "**"
pull_request:
workflow_dispatch:
# Every job here is checkout plus local tools. The token needs to read the tree
# and nothing else, and saying so keeps a future step that reaches for the API
# from quietly holding a token that can write to the repository.
permissions:
contents: read
actions: read
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
REGISTRY: gcr.forust.xyz
jobs:
checks:
runs-on: homelab
timeout-minutes: 30
lint-compose:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
echo "$tools_dir" >> "$GITHUB_PATH"
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# Structure check for every committed Compose file, active or not.
# Interpolation, env-file and bind-mount resolution are all switched off,
# because inactive stacks have no .env here and would only fail on their
# ${VAR:?} guards. Active stacks get the full check with interpolation in
# the deploy workflow, where the real .env files live.
- name: Validate Compose files
shell: bash
run: |
@@ -52,15 +61,35 @@ jobs:
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
lint-actionlint:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
export PATH="$tools_dir:$PATH"
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
lint-shellcheck:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)"
export PATH="$tools_dir:$PATH"
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash'
)
@@ -69,11 +98,20 @@ jobs:
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
bash .gitea/tests/deploy-validation.sh
lint-prettier:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Check formatting with Prettier
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
export PATH="$tools_dir:$PATH"
mapfile -t prettier_files < <(
git ls-files \
@@ -87,17 +125,36 @@ jobs:
fi
prettier --check --ignore-unknown "${prettier_files[@]}"
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint and format-check Python with Ruff
shell: bash
run: |
set -euo pipefail
ruff check . .gitea/workflows
ruff format --check . .gitea/workflows
python3 -m unittest discover -s tests -v
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint YAML syntax
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
export PATH="$tools_dir:$PATH"
mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \
@@ -111,10 +168,20 @@ jobs:
fi
yamllint -c .yamllint "${yaml_files[@]}"
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Dockerfiles
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
export PATH="$tools_dir:$PATH"
mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
@@ -126,10 +193,20 @@ jobs:
fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
validate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Kubernetes manifests against JSON schemas
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
@@ -146,27 +223,327 @@ jobs:
-ignore-missing-schemas \
-summary \
"${manifests[@]}"
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
# skipped above. The live API server knows the real CRD schemas (and runs the
# cert-manager / Traefik admission webhooks), so validate there too.
#
# Only services marked with a k8s/active marker are checked: server-side
# dry-run needs the target namespace to exist, and inactive services are not
# deployed. Services being enabled for the first time are still covered by
# the JSON-schema pass above.
#
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
# the admission webhooks of the production API server, so anyone able to open
# a pull request would be able to run arbitrary manifest content through
# cert-manager and Traefik. A pull request has nothing to gain from it either:
# only main is ever deployed, and this job runs to completion before the
# deploy workflow is allowed to start, so a bad CRD is still caught before
# anything reaches the cluster -- just on the push rather than on the PR.
- name: Note the server-side check is not running here
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
shell: bash
run: |
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
"Traefik admission webhooks against the production API server, so it is limited" \
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
"server-side pass still runs on main before the deploy."
- name: Validate active manifests against the live API server
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
shell: bash
run: |
set -euo pipefail
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
exit 0
fi
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
manifests=()
kustomize_apps=()
for dir in "${k8s_dirs[@]}"; do
if [ ! -f "${dir}/active" ]; then
echo "skip (no k8s/active): ${dir}"
continue
fi
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/overlays/prod")
elif [ -f "${dir}/base/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/base")
else
while IFS= read -r f; do
[ -n "$f" ] && manifests+=("$f")
done < <(
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
fi
done
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
failed=0
for m in ${manifests[@]+"${manifests[@]}"}; do
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
failed=1
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
fi
done
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
failed=1
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
exit 1
fi
echo "server-side dry-run: all active manifests accepted by the API server"
build:
needs:
- checks
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
runs-on: homelab
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
# userbot moved to its own repo; upstream's code is upstream's gate now.
# The rule is unchanged: publishing and passing the checks are the same
# gate, so a commit that fails any of these still cannot move :prod.
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
outputs:
services: ${{ steps.services.outputs.services }}
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
fetch-depth: 0
- name: Build changed images and write release
- name: Detect changed docker-built services
id: services
shell: bash
run: |
set -euo pipefail
base="${{ github.event.before }}"
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
base="$(git rev-list --max-parents=0 HEAD)"
fi
# A failed diff used to leave changed_files empty, which reads exactly
# like "nothing to build": the job went green having built nothing and
# the tag never moved. The status is checked, not assumed.
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
exit 1
fi
mapfile -t changed_files <<<"$changed"
services=()
add_service() {
local name="$1"
local seen=0
for existing in "${services[@]}"; do
if [ "$existing" = "$name" ]; then
seen=1
break
fi
done
if [ "$seen" -eq 0 ]; then
services+=("$name")
fi
}
for file in "${changed_files[@]}"; do
case "$file" in
errorpages/*)
add_service errorpages
;;
homepages/*)
add_service homepages
;;
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
add_service edu_master
;;
esac
done
if [ "${#services[@]}" -eq 0 ]; then
echo "No docker-built services changed."
echo "services=" >> "$GITHUB_OUTPUT"
exit 0
fi
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
# The pin step below also writes (manifest PUTs), and it runs on every
# main push — including manifest-only ones where services is empty. A
# stale persistent login on the old runner used to mask this; a clean
# runner pushes anonymously and gets 401.
if: steps.services.outputs.services != '' || github.ref_name == 'main'
shell: bash
# Through env, not by substitution into the script. A secret written
# into a run: block is pasted into the shell source before bash parses
# it, so a password containing a quote, a backtick or $(...) becomes
# code that runs. Masking the value in the log does not prevent that.
env:
GITEA_TOKEN: ${{ github.token }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: python3 .gitea/workflows/release.py build
- name: Store commit release
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
with:
name: release-${{ github.sha }}
path: release.json
if-no-files-found: error
retention-days: 30
run: |
set -euo pipefail
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
-u "$REGISTRY_USERNAME" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
# This step was the one run: block in the workflow without it, and it
# is the one that cannot afford it: a docker push that failed partway
# through the loop used to be followed by more pushes, the loop's exit
# status came from the last one, and the job went green with half the
# images missing from the registry.
set -euo pipefail
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
# Tags for this push. The commit-pinned name is the point of this
# step: the deploy resolves it in preference to :prod, so a deploy
# that sat in the queue behind a later push still gets the build of
# the commit CI validated, instead of whatever :prod points at by the
# time it runs. See render_pinned in deploy-lib.sh.
commit_tag=""
if [ "${GITHUB_REF_NAME}" = "main" ]; then
commit_tag="sha-${GITHUB_SHA:0:12}"
fi
set_tags() {
tags=()
case "${GITHUB_REF_NAME}" in
main) tags+=("main" "prod") ;;
dev) tags+=("dev") ;;
esac
if [ -n "$commit_tag" ]; then
tags+=("$commit_tag")
fi
}
for service in "${services[@]}"; do
case "$service" in
errorpages)
image="${REGISTRY}/forust/error-pages"
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
homepages)
for variant in forust xdfnx; do
case "$variant" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for variant in session-keeper webinar-checker; do
case "$variant" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
# Every image the tree names has to carry the commit-pinned name, not only
# the ones this push rebuilt. A push that touches nothing but manifests
# builds nothing, and its deploy would then find no commit-pinned tag to
# resolve and quietly fall back to the moving :prod - which is the whole
# failure the commit-pinned name exists to remove.
#
# Re-tagging copies the manifest list and transfers no layers, so pinning
# six images that already exist costs six registry writes.
#
# The list is derived from the tree rather than written out here, so an
# image added to a manifest is covered without a second place to update.
- name: Pin the commit name on the images this push did not rebuild
if: github.ref_name == 'main'
shell: bash
run: |
set -euo pipefail
commit_tag="sha-${GITHUB_SHA:0:12}"
mapfile -t repos < <(
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
| sort -u
)
if [ "${#repos[@]}" -eq 0 ]; then
echo "No own images referenced by the tree."
exit 0
fi
echo "pinning ${#repos[@]} image(s) to $commit_tag"
for repo in "${repos[@]}"; do
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
echo " already built by this push: ${repo##*/}"
continue
fi
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
continue
fi
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
echo " pinned ${repo##*/}"
done
+2 -1
View File
@@ -21,7 +21,8 @@
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*compose.yaml' '*compose.yml'
'*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \
'*/docker-compose.yaml' '*/docker-compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
-93
View File
@@ -1,93 +0,0 @@
#!/usr/bin/env python3
"""Resolve Compose images without changing project names or local bind paths."""
import json
import os
import re
import subprocess
import sys
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def resolve(reference):
if '@sha256:' in reference:
return reference
descriptor = json.loads(
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
)
digest = descriptor['digest']
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError(f'Invalid registry digest for {reference}')
# Strip tag only from the final path segment (registry ports are preserved).
repository = reference.rsplit('/', 1)
repository[-1] = repository[-1].split(':')[0]
return '/'.join(repository) + '@' + digest
def prepare(source_file):
config_repo = Path(os.environ['CONFIG_REPO'])
source_repo = Path(os.environ['REPO'])
directory = Path(os.environ['RUN_DIR'])
relative = source_file.relative_to(source_repo)
project_dir = config_repo / relative.parent
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
project = config['name']
previous_file = directory / 'previous.json'
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text())
before = json.loads(json.dumps(config))
for service, settings in config['services'].items():
reference = settings.get('image')
if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
if image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference]
else:
pinned = resolve(reference)
settings['image'] = pinned
locks[reference] = pinned
# Capture what is running, not the current value of its mutable tag.
ids = output(
'docker',
'ps',
'-aq',
'--filter',
f'label=com.docker.compose.project={project}',
'--filter',
f'label=com.docker.compose.service={service}',
).splitlines()
actual = set()
for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True)
destination = folder / f'{relative.parent.name}.json'
destination.write_text(json.dumps(data, indent=2) + '\n')
destination.chmod(0o600)
images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved')
print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
)
if __name__ == '__main__':
prepare(Path(sys.argv[1]))
-323
View File
@@ -1,323 +0,0 @@
#!/usr/bin/env python3
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
import argparse
import contextlib
import fcntl
import importlib.util
import json
import math
import os
import re
import shutil
import subprocess
import sys
import time
from pathlib import Path
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
def command(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def atomic_json(path, data):
temporary = path.with_suffix('.tmp')
temporary.write_text(json.dumps(data, indent=2) + '\n')
temporary.chmod(0o600)
temporary.replace(path)
@contextlib.contextmanager
def lock(name):
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
with (STATE / name).open('a') as stream:
fcntl.flock(stream, fcntl.LOCK_EX)
yield
def load_module(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def run_directory(run_id):
if not RUN_ID.fullmatch(run_id):
raise ValueError('Run ID must be numeric workflow-id and attempt')
return STATE / 'runs' / run_id
def start(run_id):
payload = sys.stdin.buffer.read(256 * 1024 + 1)
if len(payload) > 256 * 1024:
raise ValueError('Deploy request exceeds 256 KiB')
request = json.loads(payload)
sha = request['release']['sha']
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
raise ValueError('Invalid deploy SHA or mode')
if not isinstance(request['refresh_images'], bool):
raise ValueError('refresh_images must be boolean')
directory = run_directory(run_id)
with lock('prepare.lock'):
if (directory / 'request.json').exists():
if json.loads((directory / 'request.json').read_text()) != request:
raise ValueError('Run ID already belongs to a different request')
else:
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
if not (directory / 'source').exists():
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
raise ValueError('Prepared source does not match deploy SHA')
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
release_module.validate_release(request['release'], sha)
atomic_json(directory / 'release.json', request['release'])
atomic_json(directory / 'request.json', request)
if not (directory / 'status.json').exists():
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
# Starting an existing active or finished ID is idempotent; never re-apply it.
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
print(f'Accepted deploy {run_id} ({sha})')
def environment(directory):
request = json.loads((directory / 'request.json').read_text())
return {
**os.environ,
'REPO': str(directory / 'source'),
'CONFIG_REPO': str(CONFIG_REPO),
'RUN_DIR': str(directory),
'DEPLOY_SHA': request['release']['sha'],
'RELEASE_FILE': str(directory / 'release.json'),
'DEPLOY_PLAN': str(directory / 'plan.json'),
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
'ROLLOUT_PARALLELISM': '4',
}
def stage(directory, name, budget):
status = json.loads((directory / 'status.json').read_text())
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
return status['stages'][name]['result'] == 'success'
started = time.time()
status['stages'][name] = {'result': 'running', 'started': started}
atomic_json(directory / 'status.json', status)
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
with (directory / f'{name}.log').open('a') as log:
# timeout kills the whole stage process group, including children, before recovery.
result = subprocess.run( # noqa: S603, S607
[
shutil.which('timeout') or '/usr/bin/timeout',
'--signal=TERM',
'--kill-after=30s',
str(budget),
'bash',
str(script),
name,
],
env=environment(directory),
stdout=log,
stderr=subprocess.STDOUT,
check=False,
).returncode
status = json.loads((directory / 'status.json').read_text())
status['stages'][name].update(
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
)
atomic_json(directory / 'status.json', status)
return result == 0
def make_plan(directory):
source = directory / 'source'
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
helm = json.loads(command('helm', 'list', '--all', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose']
atomic_json(directory / 'plan.json', plan)
if previous:
atomic_json(directory / 'previous.json', previous)
# Local config is deliberately separate from the immutable Git source.
return plan
def finish_success(directory, plan):
# Repeating finalization after a crash is safe while holding deploy.lock.
plan['run_id'] = directory.name
path = directory / 'compose-images.json'
previous = directory / 'previous.json'
plan['compose-images'] = (
json.loads(path.read_text())
if path.exists()
else json.loads(previous.read_text()).get('compose-images', {})
if previous.exists()
else {}
)
atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success'
atomic_json(directory / 'status.json', status)
try:
retain_completed(directory)
except (OSError, subprocess.CalledProcessError) as error:
print(f'Retention deferred: {error}', flush=True)
def recover(directory, retry=False):
status = json.loads((directory / 'status.json').read_text())
if status['state'] in ('success', 'planned'):
return
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
return
if retry:
for name in ('verify-k8s', 'smoke'):
if status['stages'].get(name, {}).get('result') == 'failure':
del status['stages'][name]
atomic_json(directory / 'status.json', status)
snapshot = directory / 'snapshot/current'
if snapshot.exists():
stage(directory, 'verify-k8s', 7200)
stage(directory, 'smoke', 600)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'failure'
atomic_json(directory / 'status.json', status)
def execute(run_id):
directory = run_directory(run_id)
with lock('deploy.lock'):
status = json.loads((directory / 'status.json').read_text())
if status['state'] != 'queued':
return
# A crashed predecessor must be recovered before another apply begins.
for other in (STATE / 'runs').iterdir():
if (
other != directory
and (other / 'status.json').exists()
and json.loads((other / 'status.json').read_text())['state'] == 'running'
):
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running'
atomic_json(directory / 'status.json', status)
try:
plan = make_plan(directory)
print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True,
)
if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed')
if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'planned'
atomic_json(directory / 'status.json', status)
return
# Budget includes both rollout checks and rollback waves, plus API overhead.
count = int(
command(
'bash',
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
'workload-count',
env=environment(directory),
)
)
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy')
k8s_ok = stage(directory, 'apply-k8s', 2700)
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
verify_ok = stage(directory, 'verify-k8s', verify_budget)
smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
finish_success(directory, plan)
except Exception as error:
with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n')
recover(directory)
raise
def retain_completed(current):
finished = []
for directory in (STATE / 'runs').iterdir():
status_file = directory / 'status.json'
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
finished.append(directory)
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
if directory == current:
continue
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
shutil.rmtree(directory)
def follow(run_id, phase):
directory = run_directory(run_id)
groups = {
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
'verify': ('verify-k8s',),
'smoke': ('smoke',),
}
names = groups[phase]
offsets = {}
while True:
status = json.loads((directory / 'status.json').read_text())
for name in (*names, 'controller'):
path = directory / f'{name}.log'
if path.exists():
with path.open() as stream:
stream.seek(offsets.get(name, 0))
content = stream.read()
if content:
print(content, end='', flush=True)
offsets[name] = stream.tell()
stages = status['stages']
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
return all(stages[name]['result'] == 'success' for name in names)
if status['state'] in ('success', 'failure', 'planned'):
return status['state'] in ('success', 'planned')
time.sleep(3)
def main():
os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow'))
parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
args = parser.parse_args()
directory = run_directory(args.run_id)
if args.action == 'start':
start(args.run_id)
elif args.action == 'execute':
execute(args.run_id)
elif args.action == 'recover':
with lock('deploy.lock'):
recover(directory, retry=args.retry)
elif args.action == 'status':
print((directory / 'status.json').read_text())
if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif not follow(args.run_id, args.phase):
sys.exit(1)
if __name__ == '__main__':
main()
+445 -287
View File
@@ -1,5 +1,5 @@
#!/usr/bin/env bash
# Workstation deploy stages; invoked by the durable controller against pinned source.
# Shared stages for the deploy workflow. Runs on the workstation, invoked as:
# REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF'
# source "$REPO/.gitea/workflows/deploy-lib.sh"
# run_stage "$STAGE"
@@ -8,8 +8,8 @@ set -euo pipefail
: "${REPO:?REPO must be set}"
APPLY_PRUNE="${APPLY_PRUNE:-false}"
CONFIG_REPO="${CONFIG_REPO:-$REPO}"
# Exact SHA accepted by the CI gate for both manual and automatic deploys.
# Commit CI validated. Empty for a manual workflow_dispatch, which falls back to
# the current origin/main.
DEPLOY_SHA="${DEPLOY_SHA:-}"
# Handoff point between the apply stage (writes) and the verify stage (reads).
# Under the deploy user's own XDG state directory rather than /var/backups: the
@@ -20,7 +20,7 @@ DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state
# Per-workload rollout budget and how many workloads to watch at once. The whole
# apply job has its own timeout-minutes as a backstop.
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-4}"
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-8}"
WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps"
log() {
@@ -31,16 +31,6 @@ warn() {
echo "WARNING: $*" >&2
}
# Prune needs the complete desired set in one invocation. Per-file pruning
# treats resources from the other files as absent and can delete them.
check_prune_mode() {
if [ "$APPLY_PRUNE" = "true" ]; then
echo "ERROR: APPLY_PRUNE=true is unsupported by the per-file deploy loop." >&2
echo "Disable it; remove obsolete resources explicitly after review." >&2
return 1
fi
}
collect_k8s() {
git -C "$REPO" ls-files -- "$1" \
| grep -E '\.ya?ml$' \
@@ -60,25 +50,6 @@ kustomize_overlay() {
fi
}
selected_service() {
local kind="$1" service="$2" section=selected
[ -n "${DEPLOY_PLAN:-}" ] || return 0
if [ "${DEPLOY_SMOKE_ALL:-false}" = true ]; then section=active; fi
jq -e --arg kind "$kind" --arg service "$service" --arg section "$section" \
'.[$section][$kind] | index($service) != null' "$DEPLOY_PLAN" >/dev/null
}
# Resolve .env and relative binds on the persistent workstation tree. Locked
# JSON configs keep the same Compose project name and volume names.
compose() {
local cf="$1" locked project_dir
project_dir="$CONFIG_REPO/$(basename "$(dirname "$cf")")"
shift
locked="${RUN_DIR:-/nonexistent}/compose/$(basename "$(dirname "$cf")").json"
if [ -f "$locked" ]; then cf="$locked"; fi
(cd "$CONFIG_REPO" && docker compose --project-directory "$project_dir" -f "$cf" "$@")
}
select_manifests() {
K8S_MANIFESTS=()
KUSTOMIZE_APPS=()
@@ -86,7 +57,6 @@ select_manifests() {
local kd_rel kd overlay cf_rel cf f
while IFS= read -r kd_rel; do
kd="$REPO/$kd_rel"
selected_service k8s "${kd_rel%/k8s}" || continue
if [ ! -f "$kd/active" ]; then
echo "skip (no k8s/active): $kd_rel"
continue
@@ -108,7 +78,6 @@ select_manifests() {
)
while IFS= read -r cf_rel; do
cf="$REPO/$cf_rel"
selected_service compose "$(dirname "$cf_rel")" || continue
if [ -f "$(dirname "$cf")/active" ]; then
echo "compose: $cf_rel"
COMPOSE_STACKS+=("$cf")
@@ -125,10 +94,16 @@ select_manifests() {
# on failure roll them back to the revision that was running before, so a bad
# push to main cannot leave a service crash-looping.
#
# The workstation controller runs apply and verification as separate durable
# stages. Runner jobs only follow their logs. ExecStopPost recovers interrupted
# runs using the per-run snapshot, even when the SSH connection has gone away.
# Verification lives in its own workflow job, not at the end of the apply stage.
# Inside a single process it is worthless exactly when it is needed most: a job
# killed by timeout-minutes or cancelled mid-apply never reaches the rollback
# code, and leaves a half-applied cluster behind. Split out, the apply job can
# die in any way and the verify job still runs.
#
# That split needs a handoff point on the workstation, because the two stages are
# separate processes on separate runner jobs: DEPLOY_SNAPSHOT_DIR/current, written
# before anything is applied, read by the verify stage afterwards.
# Creates this run's snapshot directory and publishes it as the handoff point for
# the verify stage. Fails hard by design: a deploy that cannot record what it is
# about to change must not start, because then nothing can be rolled back for it
@@ -158,35 +133,18 @@ snapshot_dir() {
}
save_snapshot() {
local dir="$1" releases revision status
local dir="$1"
log "Saving pre-apply snapshot to $dir"
workload_generations >"$dir/generations.before" || return 1
kubectl get "$WORKLOAD_KINDS" -A -o json >"$dir/workloads.json" || return 1
kubectl get controllerrevisions.apps -A -o json >"$dir/controller-revisions.json" || return 1
jq --slurpfile revisions "$dir/controller-revisions.json" '
[.items[] | . as $w | {
kind: (.kind | ascii_downcase), namespace: .metadata.namespace, name: .metadata.name, uid: .metadata.uid,
revision: (if .kind == "Deployment" then (.metadata.annotations["deployment.kubernetes.io/revision"] // "0" | tonumber)
else ([$revisions[0].items[] | select(.metadata.namespace == $w.metadata.namespace)
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
releases="$(helm list --all -A -o json)" || return 1
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
if ! jq -e --arg r "$release" --arg n "$namespace" \
'any(.[]; .name == $r and .namespace == $n)' <<<"$releases" >/dev/null; then
continue
fi
helm status "$release" -n "$namespace" -o json >"$dir/helm-$release.json" || return 1
status="$(jq -r '.info.status' "$dir/helm-$release.json")"
if [ "$status" != deployed ]; then
# Never capture a pending/failed revision as the recovery target.
helm history "$release" -n "$namespace" -o json >"$dir/helm-$release.history.json" || return 1
revision="$(jq '[.[] | select(.status == "deployed" or .status == "superseded") | .revision] | max // 0' \
"$dir/helm-$release.history.json")"
jq --argjson revision "$revision" '.version = $revision' "$dir/helm-$release.json" >"$dir/helm-$release.tmp"
mv "$dir/helm-$release.tmp" "$dir/helm-$release.json"
workload_generations >"$dir/generations.before" 2>/dev/null \
|| warn "could not snapshot workload generations"
kubectl get "$WORKLOAD_KINDS" -A -o yaml >"$dir/workloads.yaml" 2>/dev/null \
|| warn "could not snapshot workloads"
for release in prometheus-stack loki alloy; do
if helm status "$release" -n prometheus >/dev/null 2>&1; then
{
echo "revision: $(helm history "$release" -n prometheus -o json 2>/dev/null)"
helm get values "$release" -n prometheus --all 2>/dev/null
} >"$dir/helm-$release.txt"
fi
done
# The verify stage compares this against the commit it is deploying, to refuse
@@ -210,24 +168,294 @@ workload_generations() {
# moved since the snapshot, i.e. the ones this apply actually touched.
changed_workloads() {
local before="$1"
local ns name kind gen old current
current="$(workload_generations)" || return 1
local ns name kind gen old
while read -r ns name kind gen; do
[ -n "${gen:-}" ] || continue
old="$(awk -v want_ns="$ns" -v want_name="$name" -v want_kind="$kind" \
'$1 == want_ns && $2 == want_name && $3 == want_kind { print $4; exit }' "$before" 2>/dev/null || true)"
if [ -n "${RUN_DIR:-}" ] && ! grep -qxF "$kind $ns $name" "$RUN_DIR/workload-refs"; then
continue
fi
old="$(awk -v want_ns="$ns" -v want_name="$name" \
'$1 == want_ns && $2 == want_name { print $4; exit }' "$before" 2>/dev/null || true)"
if [ "$old" != "$gen" ]; then
printf '%s %s %s\n' "$kind" "$ns" "$name"
fi
done <<<"$current"
done < <(workload_generations)
}
# Resolve owned image references exclusively from the checked CI artifact.
# Prints "<ns> <kind>/<name> <image>" for every workload this repository owns that
# runs an image from our own registry.
#
# The repository is the scope, deliberately. The cluster also holds workloads on
# our registry that no manifest here declares (they are applied out of band), and
# those are somebody else's to deploy. Walking the manifests rather than the
# cluster means those can never be restarted by this pipeline, now or later.
owned_registry_workloads() {
local kd_rel f
while IFS= read -r kd_rel; do
[ -f "$REPO/$kd_rel/active" ] || continue
while IFS= read -r f; do
[ -n "$f" ] || continue
# A file that does not mention the registry cannot declare a workload on it,
# and parsing costs ~2.5s per file against a millisecond for the grep. The
# filter keeps this at a handful of parses instead of one per manifest.
grep -q 'gcr\.forust\.xyz/forust/' "$REPO/$f" 2>/dev/null || continue
# kubectl prints a bare object for a single-document file and a List for a
# multi-document one, so normalise both shapes before filtering.
kubectl apply --dry-run=client -f "$REPO/$f" -o json 2>/dev/null \
| jq -r '
(if .items then .items[] else . end)
| select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| select(any((.spec.template.spec.containers // [])[]?;
(.image // "") | test("^gcr\\.forust\\.xyz/forust/")))
| (.metadata.namespace // "default") as $ns
| ([.spec.template.spec.containers[].image
| select(test("^gcr\\.forust\\.xyz/forust/"))][0]) as $img
| "\($ns) \(.kind | ascii_downcase)/\(.metadata.name) \($img)"
' 2>/dev/null || true
done < <(collect_k8s "$kd_rel" || true)
done < <(
git -C "$REPO" ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
}
# Prints the digest an image tag resolves to for this cluster's architecture, or
# nothing when it cannot be resolved.
#
# Only the manifest entry matching the node architecture counts. A multi-arch tag
# also carries `unknown/unknown` entries for the build attestation, and a pod's
# imageID is always the per-platform digest, so comparing the wrong entry would
# mark every workload stale forever and restart the whole cluster on every deploy.
registry_digest() {
local arch
arch="$(kubectl get nodes -o jsonpath='{.items[0].status.nodeInfo.architecture}' 2>/dev/null || true)"
[ -n "$arch" ] || arch=amd64
# The || true is load-bearing. Every caller runs under set -euo pipefail, and
# pipefail reports the rightmost non-zero stage, so a ref the registry does not
# have would abort the caller at the assignment instead of yielding an empty
# string. The callers check for empty themselves and report it by name.
#
# Retried with a hard timeout because the registry has a known hang mode (and
# a known blink mode: a single failed lookup aborts the whole apply file in
# render_pinned). A short sleep between attempts lets a restarting registry
# come back instead of failing the deploy on one bad second.
local attempt=0 digest=""
while [ "$attempt" -lt 3 ]; do
digest="$(timeout 25s docker manifest inspect "$1" 2>/dev/null \
| jq -r --arg arch "$arch" '
.manifests[]?
| select(.platform.os == "linux" and .platform.architecture == $arch)
| .digest
' 2>/dev/null \
| head -1 || true)"
[ -n "$digest" ] && break
attempt=$((attempt + 1))
if [ "$attempt" -lt 3 ]; then
echo "WARNING: registry lookup of $1 failed (attempt $attempt/3), retrying in 5s" >&2
sleep 5
fi
done
printf '%s' "$digest"
}
# The commit this deploy is for: what CI validated, or - on a manual dispatch,
# whatever stage_preflight just checked out.
deploy_commit() {
local c="${DEPLOY_SHA:-}"
[ -n "$c" ] || c="$(git -C "$REPO" rev-parse HEAD 2>/dev/null || true)"
printf '%.12s' "${c:-}"
}
# Resolves one of our image refs to the digest THIS commit's build produced.
#
# A manifest naming `:prod` names a pointer, not a version, and the deploy
# resolves it when the apply runs - which is not when CI ran it. Deploy runs are
# queued rather than cancelled (see deploy.yaml), so two pushes in a row leave
# the first deploy resolving the second push's build: the right manifests with
# the wrong code, and nothing anywhere reports it. ci therefore publishes every
# image it ships under `sha-<commit12>`, a name that cannot move, and that is
# the name resolved here.
#
# The fallback to the plain tag is for an image this pipeline never built. It
# reports itself, because a fallback nobody sees is the failure this removes.
pinned_digest() {
local ref="$1" commit pinned
commit="$(deploy_commit)"
if [ -n "$commit" ]; then
pinned="$(registry_digest "${ref%:*}:sha-$commit")"
if [ -n "$pinned" ]; then
printf '%s' "$pinned"
return 0
fi
fi
pinned="$(registry_digest "$ref")"
if [ -n "$pinned" ]; then
echo "WARNING: ${ref} carries no sha-${commit:-<unknown>} tag; resolved the moving tag instead" >&2
fi
printf '%s' "$pinned"
}
# Rewrites our own images to immutable digests on the way into the cluster.
# Reads a manifest stream on stdin, writes the pinned stream to stdout.
#
# A digest is not knowable when a manifest is written, so it is never committed:
# git keeps a readable `:prod` tag and the exact bytes are chosen here, at apply
# time, from the tag ci published for the commit being deployed. That is what
# makes rollback mean something. `kubectl rollout undo` restores the previous
# ReplicaSet's pod template verbatim, and a template naming a digest restores the
# exact bytes that were serving before. A template naming a moving tag does not —
# the tag has already moved by the time the rollback runs, so the "rollback"
# re-pulls the very image that just failed and the cluster stays broken.
#
# imagePullPolicy is deliberately left alone. The manifests no longer set it, and a
# reference that is not `:latest` defaults to IfNotPresent, which is what the
# Kubernetes docs ask for alongside a digest: the bytes under a digest cannot
# change, so pulling again buys nothing.
#
# An image that cannot be resolved is fatal. Carrying on would quietly apply a
# mutable tag again, which is the exact failure this function exists to remove.
render_pinned() {
python3 "$REPO/.gitea/workflows/release.py" render
local src refs map ref digest missing=0
src="$(mktemp)"
refs="$(mktemp)"
map="$(mktemp)"
cat >"$src"
grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+' "$src" | sort -u >"$refs" || true
while read -r ref; do
[ -n "$ref" ] || continue
digest="$(pinned_digest "$ref")"
if [ -z "$digest" ]; then
echo "ERROR: cannot resolve ${ref} in the registry; applying nothing." >&2
echo " The build job has to push that tag before the deploy resolves it." >&2
missing=$((missing + 1))
continue
fi
printf '%s\t%s\n' "$ref" "$digest" >>"$map"
done <"$refs"
if [ "$missing" -gt 0 ]; then
rm -f "$src" "$refs" "$map"
return 1
fi
awk -v mapfile="$map" '
BEGIN {
while ((getline line < mapfile) > 0) {
i = index(line, "\t")
d[substr(line, 1, i - 1)] = substr(line, i + 1)
}
}
{
if (match($0, /^[[:space:]]*image:[[:space:]]*gcr\.forust\.xyz\/forust\/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+[[:space:]]*$/)) {
name = $0
sub(/^[[:space:]]*image:[[:space:]]*/, "", name)
sub(/[[:space:]]*$/, "", name)
if (name in d) {
pad = $0
sub(/image:.*/, "", pad)
# Drop the tag: the canonical form used in the docs is repo@sha256:...,
# and leaving :prod next to the digest reads like it still matters.
repo = name
sub(/:[A-Za-z0-9._-]+$/, "", repo)
print pad "image: " repo "@" d[name]
next
}
}
print
}
' "$src"
rm -f "$src" "$refs" "$map"
}
# Restarts every owned workload whose running image is not the one its tag
# resolves to now.
#
# This used to be how a rebuild reached the cluster at all: the manifests pinned
# `:latest`, so a rebuild left the pod template byte-identical, `kubectl apply`
# decided there was nothing to do, and the cluster served the previous build
# indefinitely. The apply now pins digests via render_pinned, so a rebuild moves
# the pod template and rolls out on its own.
#
# What is left is the drift check: a hand-run `kubectl set image`, or anything
# else that edits a live workload behind the deploy's back, is the only way to end
# up serving a digest the tag has moved past. It stays idempotent, so a redeploy
# that changed no image still does not bounce healthy services.
#
# The container is matched on its repository rather than on the exact reference:
# once render_pinned has run, a pod's status reports `repo@sha256:...` while this
# still reads the repository's `:prod` tag out of the manifest.
restart_stale_images() {
local ns target image want selector running entry one
local unchecked=0
local -A digests=()
local -a stale=()
while read -r ns target image; do
[ -n "${target:-}" ] || continue
if [ -z "${digests[$image]:-}" ]; then
digests[$image]="$(pinned_digest "$image")"
fi
want="${digests[$image]}"
if [ -z "$want" ]; then
warn "cannot resolve ${image##*/} in the registry, leaving $target alone"
unchecked=$((unchecked + 1))
continue
fi
selector="$(kubectl get "$target" -n "$ns" -o jsonpath='{.spec.selector.matchLabels}' 2>/dev/null \
| jq -r 'to_entries | map("\(.key)=\(.value)") | join(",")' 2>/dev/null)"
if [ -z "$selector" ]; then
warn "cannot read the pod selector of $target, skipping"
unchecked=$((unchecked + 1))
continue
fi
running="$(kubectl get pods -n "$ns" -l "$selector" -o json 2>/dev/null \
| jq -r --arg repo "${image%%:*}" '
.items[] | .status.containerStatuses[]?
| select(.image == $repo
or (.image | startswith($repo + ":"))
or (.image | startswith($repo + "@")))
| .imageID
' 2>/dev/null)"
if [ -z "$running" ]; then
# Scaled to zero. Nothing is serving stale code, and imagePullPolicy
# resolves the tag when it is scaled back up.
continue
fi
entry=""
while IFS= read -r one; do
[ -n "$one" ] || continue
entry="${one##*@}"
if [ "$entry" != "$want" ]; then
stale+=("$ns $target")
break
fi
done <<<"$running"
done < <(owned_registry_workloads)
if [ "${#stale[@]}" -eq 0 ]; then
if [ "$unchecked" -gt 0 ]; then
# Say so plainly. Reporting "everything is current" after checking nothing
# would tell the operator the deploy is fine when it may not be.
warn "No workload needed a restart, but $unchecked could not be checked"
else
log "All owned workloads already run the image their tag points at"
fi
return 0
fi
log "Restarting ${#stale[@]} workload(s) running an image their tag has moved past"
for ref in "${stale[@]}"; do
log " $ref"
done
local failed=()
for ref in "${stale[@]}"; do
ns="${ref%% *}"
target="${ref#* }"
if ! kubectl rollout restart "$target" -n "$ns" >/dev/null 2>&1; then
failed+=("$ref")
fi
done
if [ "${#failed[@]}" -gt 0 ]; then
warn "could not restart: ${failed[*]}"
return 1
fi
}
# verify_workloads <failed-file> <kind> <ns> <name> ...
@@ -271,44 +499,31 @@ verify_workloads() {
# settle. Prints a report and returns non-zero if any workload is still unhealthy,
# so the operator knows manual recovery is required.
rollback_workloads() {
local failed_file="$1" snapshot kind ns name index=0 running=0 pid revision uid
local -a pids=()
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
local failed_file="$1"
local kind ns name unrecovered=()
local -a recovered=()
while read -r kind ns name; do
[[ "$kind" =~ ^(deployment|statefulset|daemonset)$ ]] || continue
index=$((index + 1))
(
if kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.annotations}' | grep -q 'meta.helm.sh/release-name'; then
echo " skip (Helm recovery owns this workload): $kind/$ns/$name"
exit 1
fi
revision="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .revision' "$snapshot/revisions.json")"
uid="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .uid' "$snapshot/revisions.json")"
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]] || [ "$uid" != "$(kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.uid}')" ]; then
echo " no safe previous revision: $kind/$ns/$name (new or replaced workload)"
exit 1
fi
kubectl rollout undo "$kind/$name" -n "$ns" --to-revision="$revision" \
&& kubectl rollout status "$kind/$name" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s"
) >"$snapshot/rollback-$index.log" 2>&1 &
pids+=($!)
running=$((running + 1))
if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then
wait -n 2>/dev/null || true
running=$((running - 1))
[ -n "${kind:-}" ] || continue
# Helm-owned workloads are already rolled back by the release's --rollback-on-failure
# upgrade. `rollout undo` here would step back to the revision Helm just
# escaped (the failed one), so leave them for the operator instead.
if kubectl get "${kind}/${name}" -n "$ns" -o jsonpath='{.metadata.annotations}' 2>/dev/null | grep -q 'meta.helm.sh/release-name'; then
echo " skip (helm-managed, needs manual check): ${kind}/${ns}/${name}"
unrecovered+=("${kind}/${ns}/${name} (helm-managed)")
continue
fi
if kubectl rollout undo "${kind}/${name}" -n "$ns" >/dev/null 2>&1 \
&& kubectl rollout status "${kind}/${name}" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" >/dev/null 2>&1; then
echo " rolled back: ${kind}/${ns}/${name}"
recovered+=("${kind}/${ns}/${name}")
else
echo " NOT RECOVERED: ${kind}/${ns}/${name}"
unrecovered+=("${kind}/${ns}/${name}")
fi
done <"$failed_file"
local recovered=0 unrecovered=0 i=0
for pid in "${pids[@]}"; do
i=$((i + 1))
if wait "$pid"; then recovered=$((recovered + 1)); else unrecovered=$((unrecovered + 1)); fi
cat "$snapshot/rollback-$i.log"
done
echo "ROLLED_BACK=$recovered" >>"$failed_file"
echo "UNRECOVERED=$unrecovered" >>"$failed_file"
[ "$unrecovered" -eq 0 ]
echo "ROLLED_BACK=${#recovered[@]}" >>"$failed_file"
echo "UNRECOVERED=${#unrecovered[@]}" >>"$failed_file"
[ "${#unrecovered[@]}" -eq 0 ]
}
# Helm releases owned by this stage, one line each:
@@ -321,10 +536,9 @@ rollback_workloads() {
# have to be declared as custom.regex managers in renovate/renovate.json.
HELM_RELEASES=(
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
"reloader|stakater/reloader|reloader|2.2.18|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
)
# "name url" for the Helm repository hosting a chart, empty if unknown.
@@ -333,7 +547,6 @@ helm_repo_for() {
prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;;
grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;;
stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;;
victoriametrics/*) echo "victoriametrics https://victoriametrics.github.io/helm-charts" ;;
esac
}
@@ -343,12 +556,8 @@ helm_repo_for() {
helm_release_status() {
local out
if ! out="$(helm status "$1" -n "$2" 2>&1)"; then
if [[ "$out" == *"release: not found"* ]]; then
echo "not-found"
return 0
fi
printf 'ERROR: cannot read Helm status: %s\n' "$out" >&2
return 1
echo "not-found"
return 0
fi
awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]'
}
@@ -360,27 +569,16 @@ helm_release_status() {
# pending (deployed, failed, not-found). Returns non-zero when the release is
# still not recoverable, so the pipeline fails loud instead of wedging.
recover_pending_release() {
local release="$1" namespace="$2" status revision snapshot
status="$(helm_release_status "$release" "$namespace")" || return 1
local release="$1" namespace="$2" status
status="$(helm_release_status "$release" "$namespace")"
case "$status" in
pending-upgrade|pending-rollback|pending-install)
log "Release $release is $status, rolling back to the last deployed revision"
revision=""
if [ -s "$DEPLOY_SNAPSHOT_DIR/current" ]; then
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
if [ -s "$snapshot/helm-$release.json" ]; then
revision="$(jq -r '.version' "$snapshot/helm-$release.json")"
fi
fi
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]]; then
echo "ERROR: no captured Helm revision for $release; manual recovery required"
return 1
fi
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
if ! helm rollback "$release" -n "$namespace" --wait --timeout 10m >/dev/null 2>&1; then
echo "WARN: helm rollback of $release did not complete"
return 1
fi
status="$(helm_release_status "$release" "$namespace")" || return 1
status="$(helm_release_status "$release" "$namespace")"
if [ "$status" != "deployed" ]; then
echo "WARN: $release is $status after rollback"
return 1
@@ -416,16 +614,11 @@ upgrade_helm_releases() {
local entry release chart namespace version values marker repo
for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
if [ -n "${DEPLOY_PLAN:-}" ] && ! jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null; then
echo "skip (unchanged Helm release): $release"
continue
fi
if [ ! -f "$REPO/$values" ] && [ -f "$CONFIG_REPO/$values" ]; then values="$CONFIG_REPO/$values"; else values="$REPO/$values"; fi
if [ ! -f "$REPO/$marker" ]; then
echo "skip (no $marker): $release"
continue
fi
if [ ! -f "$values" ]; then
if [ ! -f "$REPO/$values" ]; then
echo "ERROR: $values is gitignored but missing on the workstation, restore it first."
return 1
fi
@@ -434,8 +627,8 @@ upgrade_helm_releases() {
echo "ERROR: no Helm repository configured for chart $chart"
return 1
fi
helm repo add "${repo%% *}" "${repo#* }" >/dev/null
helm repo update "${repo%% *}" >/dev/null
helm repo add "${repo%% *}" "${repo#* }" >/dev/null 2>&1 || true
helm repo update "${repo%% *}" >/dev/null 2>&1 || true
log "Upgrading $release ($chart $version)"
wait_for_calm "helm $release"
# A previous run with --rollback-on-failure whose own rollback never finished leaves the
@@ -451,7 +644,7 @@ upgrade_helm_releases() {
if ! helm upgrade --install "$release" "$chart" \
--namespace "$namespace" \
--version "$version" \
--values "$values" \
--values "$REPO/$values" \
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
echo "WARN: upgrade of $release failed, checking release state"
# --rollback-on-failure already attempted its own rollback; finish the job when that
@@ -467,112 +660,56 @@ upgrade_helm_releases() {
done
}
stage_doctor() {
local tool entry release chart namespace version values marker
for tool in git docker kubectl helm jq curl timeout flock python3; do
command -v "$tool" >/dev/null || { echo "Missing workstation tool: $tool"; return 1; }
done
docker compose version >/dev/null
docker buildx version >/dev/null
[ "$(kubectl config current-context)" = "${KUBE_CONTEXT:?configure KUBE_CONTEXT}" ] || { echo "Unexpected Kubernetes context"; return 1; }
[ "$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" = "${EXPECTED_CLUSTER_UID:?configure EXPECTED_CLUSTER_UID}" ] || { echo "Unexpected Kubernetes cluster"; return 1; }
kubectl get --raw=/readyz --request-timeout=10s >/dev/null
[ "$(git -C "$REPO" rev-parse HEAD)" = "$DEPLOY_SHA" ] || return 1
select_manifests
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
[ -f "$REPO/$marker" ] || continue
[ -f "$REPO/$values" ] || [ -f "$CONFIG_REPO/$values" ] || { echo "Missing values: $values"; return 1; }
done
jq '{sha, selected, helm, removed}' "$DEPLOY_PLAN"
local cf
for cf in "${COMPOSE_STACKS[@]}"; do
compose "$cf" config --quiet
while IFS= read -r network; do
docker network inspect "$network" >/dev/null || return 1
done < <(compose "$cf" config --format json | jq -r '.networks // {} | to_entries[] | select(.value.external == true) | .value.name')
python3 "$REPO/.gitea/workflows/compose-release.py" "$cf"
done
local image refs m k
refs="$(
for m in "${K8S_MANIFESTS[@]}"; do render_pinned <"$m" || return 1; done
for k in "${KUSTOMIZE_APPS[@]}"; do kubectl kustomize "$k" | render_pinned || return 1; done
)" || return 1
refs="$(grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+@sha256:[0-9a-f]{64}' <<<"$refs" | sort -u || true)"
while IFS= read -r image; do
[ -n "$image" ] || continue
timeout 60s docker buildx imagetools inspect "$image" >/dev/null
done <<<"$refs"
}
# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are
# created by cert-manager and are not prerequisites for applying a Certificate.
check_referenced_secrets() {
local m k objects refs extracted ns name
local missing=()
refs=""
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
refs+="$extracted"$'\n'
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
objects="$(kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json)" || return 1
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
refs+="$extracted"$'\n'
done
while read -r ns name; do
[ -n "${name:-}" ] || continue
if kubectl get secret "$name" -n "$ns" -o name >/dev/null 2>&1; then
echo " ok: $ns/$name"
else
echo " MISSING OR UNREADABLE: $ns/$name"
missing+=("$ns/$name")
fi
done < <(printf '%s' "$refs" | sort -u)
if [ "${#missing[@]}" -gt 0 ]; then
echo "ERROR: required pod Secrets are missing or unreadable:"
printf ' - %s\n' "${missing[@]}"
echo "Create them in the listed namespaces from the service's secret example."
return 1
stage_preflight() {
if [ ! -d "$REPO/.git" ]; then
echo "Repository not found at $REPO"
exit 1
fi
}
# The VMAgent CRD is installed by the VictoriaMetrics Operator Helm release in
# stage_apply_k8s, after this preflight stage. Skip only its dry-run until then.
skip_uninstalled_vmagent_crd() {
local manifest="$1"
if [[ "$manifest" == "$REPO/prometheus-stack/k8s/vmagent.yaml" ]] \
&& ! kubectl get crd vmagents.operator.victoriametrics.com >/dev/null 2>&1; then
echo " skip: VMAgent CRD is installed by Helm during apply: ${manifest#"$REPO"/}"
return 0
if [ -n "$DEPLOY_SHA" ]; then
log "Checking out the commit CI validated ($DEPLOY_SHA)"
git -C "$REPO" fetch origin --quiet "$DEPLOY_SHA" 2>/dev/null \
|| git -C "$REPO" fetch origin main
else
git -C "$REPO" fetch origin main
fi
return 1
target="${DEPLOY_SHA:-origin/main}"
log "Workstation state"
echo " local: $(git -C "$REPO" rev-parse --short HEAD)"
echo " target: $(git -C "$REPO" rev-parse --short "$target")"
if [ -n "$(git -C "$REPO" status --porcelain --untracked-files=no)" ]; then
echo "ERROR: workstation has local tracked modifications, refusing reset:"
git -C "$REPO" status --porcelain --untracked-files=no
git -C "$REPO" diff --stat
echo "Fix it on the workstation (commit, or 'git restore .'), then re-run the deploy."
exit 1
fi
git -C "$REPO" reset --hard "$target"
}
stage_validate() {
check_prune_mode || return 1
cd "$REPO"
select_manifests
local m k cf
# The deploy host has the local .env and secret files. Resolve them here so
# missing configuration fails before either apply job changes workloads.
# CI keeps the structure-only check for inactive stacks.
# Compose .env files and secret files are gitignored by design, so the
# workstation never has real values for the inactive stacks. This stage only
# runs the full check on active stacks; the general structure check for every
# committed Compose file (active or not) lives in the ci workflow, which has no
# .env at all.
#
# Active stacks are still validated with interpolation and env-file resolution
# off, so required-variable guards (:?) and missing local files do not fail the
# deploy. Normalization and consistency checks stay enabled.
# shellcheck source=compose-lint.sh
source "$REPO/.gitea/workflows/compose-lint.sh"
local compose_validate_flags=()
mapfile -t compose_validate_flags < <(compose_safe_flags)
log "Validate compose stacks"
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " config: $cf"
compose "$cf" config --quiet
validate_compose_file "$cf" ${compose_validate_flags[@]+"${compose_validate_flags[@]}"}
done
log "Validate k8s manifests (kubectl dry-run=client)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=client -f "$m" >/dev/null
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
@@ -580,9 +717,6 @@ stage_validate() {
done
log "Validate k8s manifests (kubectl dry-run=server)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=server -f "$m" >/dev/null
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
@@ -590,43 +724,54 @@ stage_validate() {
done
log "Checking referenced Secrets exist"
echo " (deploy never applies *secret*.yaml; create missing ones manually)"
check_referenced_secrets
}
selected_workload_refs() {
local m k
for m in "${K8S_MANIFESTS[@]}"; do
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
kubectl create --dry-run=client --validate=false -f "$m" -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
done
for k in "${KUSTOMIZE_APPS[@]}"; do
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
local ref_secrets=() missing_secrets=() all_secrets s
if [ "${#K8S_MANIFESTS[@]}" -gt 0 ]; then
while IFS= read -r s; do
[ -n "$s" ] && ref_secrets+=("$s")
done < <(
{
grep -h -A1 -E 'secretRef:|secretKeyRef:' "${K8S_MANIFESTS[@]}" 2>/dev/null || true
grep -h -E 'secretName:' "${K8S_MANIFESTS[@]}" 2>/dev/null || true
} | grep -E 'name:' | sed -E 's/.*name:[[:space:]]*//' | tr -d '"'"'"' "'"'" | sed -E 's/[[:space:]]*#.*//' | awk 'NF' | sort -u || true
)
fi
all_secrets="$(kubectl get secrets -A --no-headers -o custom-columns=:metadata.name 2>/dev/null || true)"
for s in ${ref_secrets[@]+"${ref_secrets[@]}"}; do
if printf '%s\n' "$all_secrets" | grep -qx "$s"; then
echo " ok: $s"
else
echo " MISSING: $s"
missing_secrets+=("$s")
fi
done
if [ "${#missing_secrets[@]}" -gt 0 ]; then
echo "ERROR: ${#missing_secrets[@]} referenced Secret(s) not found in the cluster:"
printf ' - %s\n' "${missing_secrets[@]}"
echo "Create them manually from the laptop, e.g.:"
echo " kubectl apply -f SERVICE/k8s/secrets.yaml # see SERVICE/k8s/secrets.yaml.example"
exit 1
fi
}
stage_apply_k8s() {
check_prune_mode || return 1
cd "$REPO"
select_manifests >/dev/null
local ns_files=() other_files=() m k
local ns_files=() other_files=() m k prune_opts=()
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
case "$m" in
*/namespace.yaml|*/namespace.yml) ns_files+=("$m") ;;
*/namespace.y?ml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
if [ "$APPLY_PRUNE" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
# Record what is about to change, and publish it for the verify job, before
# the first apply. Both are fatal on failure: see snapshot_dir.
selected_workload_refs >"$RUN_DIR/workload-refs"
local snapshot
snapshot="$(snapshot_dir)" || return 1
save_snapshot "$snapshot" || return 1
touch "$snapshot/ready"
if [ "${#ns_files[@]}" -gt 0 ]; then
log "Applying namespaces (${#ns_files[@]} files)"
@@ -634,8 +779,8 @@ stage_apply_k8s() {
kubectl apply -f "$m"
done
fi
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
if [ -f "$REPO/prometheus-stack/k8s/active" ]; then
if [ ! -f "$REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
exit 1
fi
@@ -645,8 +790,7 @@ stage_apply_k8s() {
if [ "${#other_files[@]}" -gt 0 ]; then
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
for m in "${other_files[@]}"; do
log "Applying ${m#"$REPO"/}"
if ! render_pinned <"$m" | kubectl apply -f -; then
if ! render_pinned <"$m" | kubectl apply "${prune_opts[@]}" -f -; then
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
exit 1
fi
@@ -659,6 +803,7 @@ stage_apply_k8s() {
exit 1
fi
done
restart_stale_images
# No verification here on purpose. This stage may be killed at any point by
# timeout-minutes, by the runner cancelling the job, or by a dropped SSH
@@ -673,7 +818,7 @@ stage_apply_k8s() {
# and rolls back the ones that never became healthy.
stage_verify_k8s() {
local pointer="$DEPLOY_SNAPSHOT_DIR/current"
local snapshot want have generations changed
local snapshot want have generations
local -a touched=()
if [ ! -s "$pointer" ]; then
@@ -684,7 +829,7 @@ stage_verify_k8s() {
return 1
fi
snapshot="$(head -1 "$pointer")"
if [ ! -d "$snapshot" ] || [ ! -f "$snapshot/ready" ]; then
if [ ! -d "$snapshot" ]; then
echo "ERROR: snapshot pointer refers to a missing directory: $snapshot"
return 1
fi
@@ -707,13 +852,6 @@ stage_verify_k8s() {
fi
echo " snapshot: $snapshot (commit ${have:0:12})"
local entry release chart namespace version values marker
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null || continue
recover_pending_release "$release" "$namespace" || return 1
done
generations="$snapshot/generations.before"
if [ ! -s "$generations" ]; then
# Without a baseline we cannot tell which workloads the apply touched, so
@@ -722,10 +860,9 @@ stage_verify_k8s() {
: >"$generations"
fi
changed="$(changed_workloads "$generations")" || return 1
while read -r kind ns name; do
[ -n "${kind:-}" ] && touched+=("$kind $ns $name")
done <<<"$changed"
done < <(changed_workloads "$generations")
log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)"
if [ "${#touched[@]}" -eq 0 ]; then
@@ -764,19 +901,21 @@ stage_verify_k8s() {
verify_compose_stack() {
local cf="$1"
local expected running missing=()
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
running="$(compose "$cf" ps --status running --services | sort)" || return 1
expected="$(docker compose -f "$cf" config --services 2>/dev/null | sort || true)"
running="$(docker compose -f "$cf" ps --status running --services 2>/dev/null | sort || true)"
[ -n "$expected" ] || return 0
while IFS= read -r svc; do
[ -n "$svc" ] || continue
# restart:"no" services are allowed to have exited.
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
if ! printf '%s\n' "$running" | grep -qx "$svc" \
&& ! docker compose -f "$cf" config 2>/dev/null \
| grep -A5 "^ ${svc}:" | grep -qE 'restart:\s*"?no"?'; then
missing+=("$svc")
fi
done <<<"$expected"
if [ "${#missing[@]}" -gt 0 ]; then
echo " NOT RUNNING: ${missing[*]}"
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
docker compose -f "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
return 1
fi
echo " all ${#expected} service(s) running"
@@ -860,11 +999,10 @@ traefik_routed_hosts() {
# cases, so ask Traefik which routes it built and fail on the difference.
stage_smoke() {
cd "$REPO"
if [ -n "${DEPLOY_PLAN:-}" ] && jq -e '.full_smoke' "$DEPLOY_PLAN" >/dev/null; then
DEPLOY_SMOKE_ALL=true
fi
select_manifests >/dev/null
local -a hosts=()
# Not named failed: an array of that name already exists in restart_stale_images
# above, and a scalar shadowing an array is a trap rather than a shadow.
local h code rc bad=0
while IFS= read -r h; do
[ -n "$h" ] && hosts+=("$h")
@@ -872,8 +1010,8 @@ stage_smoke() {
if [ "${#hosts[@]}" -eq 0 ]; then
# Nothing to probe means the extraction broke, not that the cluster is empty.
echo "No public routes in the selected components"
return 0
echo "ERROR: no public hostnames found in active manifests, refusing to report success"
return 1
fi
log "Probing ${#hosts[@]} public route(s)"
@@ -957,17 +1095,37 @@ stage_apply_compose() {
cd "$REPO"
select_manifests >/dev/null
local cf
for cf in "${COMPOSE_STACKS[@]}"; do
log "Applying Compose ${cf#"$REPO"/}"
compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans
verify_compose_stack "$cf"
log "Redeploying docker compose stacks (${#COMPOSE_STACKS[@]} stacks)"
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " compose: $cf"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
fi
docker compose -f "$cf" up -d --pull always --remove-orphans
done
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
local -a broken=()
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " verifying: $cf"
if ! verify_compose_stack "$cf"; then
broken+=("$cf")
fi
done
if [ "${#broken[@]}" -gt 0 ]; then
echo
echo "ERROR: ${#broken[@]} compose stack(s) did not come up:"
printf ' - %s\n' "${broken[@]}"
echo "Compose stacks are not rolled back automatically: their images use mutable"
echo "':latest' tags, so there is no previous version to return to. Check the logs"
echo "above, then re-run the deploy once the cause is fixed."
return 1
fi
}
run_stage() {
case "${1:?stage required}" in
doctor) stage_doctor ;;
preflight) stage_preflight ;;
validate) stage_validate ;;
apply-k8s) stage_apply_k8s ;;
verify-k8s) stage_verify_k8s ;;
-124
View File
@@ -1,124 +0,0 @@
#!/usr/bin/env python3
"""Calculate selected components against the last fully successful deploy."""
import hashlib
import json
import re
import subprocess
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def tracked(repo):
return output('git', '-C', str(repo), 'ls-files').splitlines()
def helm_releases(repo):
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
def inventory(repo):
files = tracked(repo)
k8s = sorted(
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
)
compose = sorted(
{
str(Path(f).parent)
for f in files
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
}
)
return {'k8s': k8s, 'compose': compose}
def file_hash(path):
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
def make_plan(repo, config_repo, release, previous, mode, live_helm):
active = inventory(repo)
all_services = set(active['k8s'] + active['compose'])
helm_inputs = {}
helm_selected = []
for name, chart, namespace, version, values, marker in helm_releases(repo):
if not (repo / marker).is_file():
continue
value_path = repo / values if (repo / values).is_file() else config_repo / values
if not value_path.is_file():
raise ValueError(f'Missing Helm values: {values}')
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
helm_inputs[name] = stamp
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
if (
mode == 'full'
or previous is None
or previous.get('helm_inputs', {}).get(name) != stamp
or live is None
or live.get('status') != 'deployed'
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
):
helm_selected.append(name)
local_inputs = {}
for service in all_services:
candidates = [config_repo / service / '.env']
if service in active['compose']:
candidates.append(config_repo / '.env')
cfg = config_repo / service / 'config'
if cfg.is_dir():
candidates.extend(
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
)
local_inputs[service] = hashlib.sha256(
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
).hexdigest()
if previous is None:
if mode == 'changed':
raise ValueError('No successful baseline; run deploy in full mode first')
changed = set(all_services)
removed = []
else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {path.split('/')[0] for path in paths}
if any(path.startswith('.gitea/') for path in paths):
changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo):
service = file.split('/')[0]
if service not in all_services or not file.endswith(('.yaml', '.yml')):
continue
text = (repo / file).read_text()
if any(
image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items()
):
changed.add(service)
removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services
)
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
if mode == 'full':
changed = set(all_services)
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
while True:
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
if expanded == changed:
break
changed = expanded
return {
'version': 1,
'sha': release['sha'],
'images': release['images'],
'active': active,
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
'helm': helm_selected,
'helm_inputs': helm_inputs,
'local_inputs': local_inputs,
'removed': sorted(set(removed)),
'full_smoke': mode == 'full' or 'traefik' in changed,
}
-10
View File
@@ -1,10 +0,0 @@
#!/usr/bin/env bash
set -euo pipefail
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
case "${1:?stage required}" in
workload-count)
select_manifests >/dev/null
selected_workload_refs | sort -u | wc -l
;;
*) run_stage "$1" ;;
esac
+161 -68
View File
@@ -1,104 +1,197 @@
name: deploy
on:
# Deploy only what CI already validated. workflow_run is used instead of
# workflow_dispatch so a red lint/validate run can never reach the cluster.
workflow_run:
workflows: [ci]
branches: [main]
types: [completed]
workflow_dispatch:
inputs:
deploy_ref:
description: "Commit already checked by successful main CI (main or SHA)"
default: main
required: true
deploy_mode:
description: "First deploy requires full; plan changes no production resources"
type: choice
options: [changed, full, plan]
default: changed
refresh_images:
description: "Explicitly refresh mutable third-party Compose tags"
type: boolean
default: false
# The deploy jobs read the tree, then reach the cluster over SSH with the
# deploy key. The Actions token itself is not part of that path, so it gets
# read-only contents and no more.
permissions:
contents: read
actions: read
concurrency:
group: deploy-main
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
# takes the verify job down with it, so a superseded deploy would leave the
# cluster half-applied and unchecked — the exact failure the verify job exists
# to catch. kubectl apply and docker compose up are both idempotent, so letting
# the older run finish and then deploying the newer commit costs little.
cancel-in-progress: false
env:
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
# finished ci run checked. Pin the exact validated commit instead, so a push
# landing mid-deploy cannot make the workstation deploy something else. Also
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
# which falls back to the current origin/main.
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
jobs:
gate:
preflight:
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
# Run workflow always bypasses the switch: dispatching it is the explicit
# intent to deploy.
if: >-
github.ref == 'refs/heads/main' &&
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
runs-on: homelab
(github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.head_branch == 'main'))
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
outputs:
sha: ${{ steps.release.outputs.sha }}
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
fetch-depth: 0
- name: Check successful CI and download the exact commit release
id: release
env:
GITEA_TOKEN: ${{ github.token }}
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
- name: Submit durable deploy to workstation
run: bash .gitea/workflows/ssh-run.sh start
apply:
needs: [gate]
runs-on: homelab
timeout-minutes: 100
- name: Fetch and reset workstation
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh preflight
validate:
needs: [preflight]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 20
steps:
- name: Checkout checked commit
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply
verify:
needs: [gate, apply]
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 130
- name: Dry-run manifests and check Secrets
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh validate
apply-k8s:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
# Apply only, no verification, so this is just the work itself: snapshot,
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop.
# Verification has its own job and its own budget.
#
# 45 is roughly four times the measured cost of the stage, which is
# deliberately not raised on a theory:
#
# helm, healthy 3 no-op upgrades ~3-5 min
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
# then rolls that one back
# apply loop ~40 manifests, 4 of which ~1 min
# resolve an image digest
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
# 9.8s to resolve their digests
#
# The helm figure is one release, not three: `set -e` aborts
# upgrade_helm_releases on the first failure, so a broken release costs
# 10m and the other two are never attempted. Multiplying 10m by three
# overstates the worst case by 20 minutes.
#
# The 45 minutes this was last raised to 45 were still not enough, and the
# job logs for those runs no longer exist, so what actually consumed the
# budget is not known - the two measurable candidates above account for
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
# known hang mode is now bounded inside registry_digest (25s timeout, 3
# attempts): a dead registry fails each owned image after ~85s instead of
# hanging the stage, and a blinking one is retried instead of failing the
# whole apply file. Still open: make the stage announce which manifest it
# is working on, so a killed run leaves a diagnosable last line.
timeout-minutes: 45
steps:
- name: Checkout checked commit
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify
- name: Apply Kubernetes manifests
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-k8s
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
# Watches the workloads this deploy changed and rolls back the ones that never
# became healthy. Runs even when the apply jobs failed, timed out or were
# cancelled — that is the whole point of splitting it out. `always()` is what
# lets it start after a failed dependency; the needs on apply-compose are a
# barrier, so verification begins only once both applies are done.
verify-k8s:
needs: [apply-k8s, apply-compose]
if: >-
always() &&
needs.apply-k8s.result != 'skipped' &&
needs.apply-compose.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
# Not raised, because the arithmetic does not close.
#
# 32 workloads are under management and the wave width is 8, so the verify
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
# every rollout times out rather than converging. That is already 20 of 30.
#
# The other 10 would have to absorb rollback, and rollback_workloads is a
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
# Any larger number is buying a bigger multiple of an unbounded term rather
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
# therefore not a bound, it is a guess with two digits.
#
# The number becomes derivable the moment rollback uses the same wave width
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
# and 45 covers verify plus rollback at full width. That change is to the
# recovery path and is not folded into a timeout edit.
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Verify workloads and roll back on failure
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh verify-k8s
# Asks the public route of every active service whether it is actually
# serving, which the rollout check above structurally cannot: a pod can
# converge and still be crash-looping, or be listening on a port no Service
# points at, or answer 500.
#
# `always()` for the same reason verify-k8s has it, and it runs after that job
# specifically because a rollback is when a route most needs re-checking. The
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
# cancelled, the probes are what say whether the cluster is serving, and
# suppressing them on a rollback would hide the one run where the answer
# matters most.
smoke:
needs: [gate, verify]
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 15
needs: [verify-k8s]
if: always() && needs.verify-k8s.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
steps:
- name: Checkout checked commit
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow public route checks
run: bash .gitea/workflows/ssh-run.sh smoke
- name: Probe the public route of every active service
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh smoke
+29 -60
View File
@@ -13,14 +13,9 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
# A runner may accept overlapping workflows even though each workflow is sequential.
exec 9>"$TOOLS_DIR/install.lock"
flock -w 300 9
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
# The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner).
@@ -53,7 +48,7 @@ esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
curl -sSLf --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
@@ -93,13 +88,10 @@ installed_version() {
# at_version <command> <expected>
at_version() {
local version expected="${2#v}"
version="$(installed_version "$1")"
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
[ "${BASH_REMATCH[2]}" = "$expected" ]
else
return 1
fi
case "$(installed_version "$1")" in
*"$2"*) return 0 ;;
*) return 1 ;;
esac
}
install_kubeconform() {
@@ -128,15 +120,6 @@ install_shellcheck() {
rm -rf "$tmp"
}
install_jq() {
if at_version jq "${JQ_VERSION}"; then
return 0
fi
fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \
"$BIN_DIR/jq"
chmod 0755 "$BIN_DIR/jq"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
@@ -184,7 +167,6 @@ install_pip_audit() {
}
install_prettier() {
install_node
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
@@ -245,41 +227,28 @@ install_actionlint() {
rm -rf "$tmp"
}
main() {
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
jq) install_jq ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
[ -d "$old" ] || continue
case "$(basename "$old")" in
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
*) rm -rf "$old" ;;
esac
done
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
printf '%s\n' "$BIN_DIR"
}
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
printf '%s\n' "$BIN_DIR"
-350
View File
@@ -1,350 +0,0 @@
#!/usr/bin/env python3
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
import argparse
import hashlib
import io
import itertools
import json
import os
import re
import shutil
import subprocess
import sys
import tempfile
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from pathlib import Path
SHA = re.compile(r'[0-9a-f]{40}')
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
IMAGES = {
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
}
# These images are released by the EDU application repository.
EXTERNAL_IMAGES = {'gcr.forust.xyz/forust/session-keeper', 'gcr.forust.xyz/forust/webinar-checker'}
def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell."""
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def validate_release(data, sha=None):
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
raise ValueError('Invalid release version or SHA')
if sha is not None and data['sha'] != sha:
raise ValueError('Release SHA does not match the checked CI commit')
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
if set(data.get('images', {})) != expected:
raise ValueError('Release must contain all owned images')
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
raise ValueError('Release has an invalid image digest')
if set(data.get('inputs', {})) != expected or not all(
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
):
raise ValueError('Release has invalid build input fingerprints')
return data
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
return None
class Gitea:
def __init__(self):
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
if urllib.parse.urlsplit(self.origin).scheme != 'https':
raise ValueError('Gitea API must use HTTPS')
self.repository = os.environ['GITHUB_REPOSITORY']
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
raise ValueError('Invalid Gitea repository')
self.token = os.environ['GITEA_TOKEN']
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
def request(self, url, *, archive=False):
if not url.startswith(self.base + '/'):
raise ValueError('Refusing to send the Actions token to another origin')
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
opener = urllib.request.build_opener(NoRedirect())
try:
response = opener.open(req, timeout=30) # noqa: S310
except urllib.error.HTTPError as error:
if not archive or error.code not in (301, 302, 303, 307, 308):
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
target = urllib.parse.urljoin(url, error.headers['Location'])
if urllib.parse.urlsplit(target).scheme != 'https':
raise ValueError('Artifact redirect must use HTTPS') from None
# Signed storage redirects must never receive the Gitea token.
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
with response:
payload = response.read(8 * 1024 * 1024 + 1)
if len(payload) > 8 * 1024 * 1024:
raise ValueError('Gitea response exceeds 8 MiB')
return payload if archive else json.loads(payload)
def pages(self, path, key, **params):
for page in range(1, 101):
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
data = self.request(f'{self.base}/{path}?{query}')
entries = data[key]
yield from entries
if len(entries) < 50:
return
raise RuntimeError('Gitea pagination limit exceeded')
def successful_runs(self, sha=None):
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
if sha:
params['head_sha'] = sha
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
if (
run.get('status') == 'completed'
and run.get('conclusion') == 'success'
and run.get('head_branch') == 'main'
and run.get('event') in ('push', 'workflow_dispatch')
and (run.get('repository') or {}).get('full_name') == self.repository
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
and (sha is None or run.get('head_sha') == sha)
):
yield run
def release(self, run):
sha = run['head_sha']
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
# A green workflow with a skipped build must not authorize a deploy.
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
raise ValueError('CI build job did not succeed')
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
if len(matching) != 1:
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
files = [entry for entry in archive.infolist() if not entry.is_dir()]
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
raise ValueError('Unexpected release archive contents')
return validate_release(json.loads(archive.read(files[0])), sha)
def fingerprint(context, dockerfile):
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
return hashlib.sha256(tree.encode()).hexdigest()
def gate(output, requested_ref, event_sha):
command('git', 'fetch', '--quiet', 'origin', 'main')
if event_sha:
if not SHA.fullmatch(event_sha):
raise ValueError('Invalid workflow_run SHA')
sha = event_sha
else:
if requested_ref == 'main':
requested_ref = 'origin/main'
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
if not SHA.fullmatch(sha):
raise ValueError('Invalid deploy SHA')
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
api = Gitea()
runs = list(api.successful_runs(sha))
if not runs:
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
release = api.release(max(runs, key=lambda run: run['id']))
output.write_text(json.dumps(release, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write(f'sha={sha}\n')
print(f'CI gate accepted {sha}')
def build(output):
sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA')
api = Gitea()
previous = None
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
continue
try:
previous = api.release(run)
break
except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
builder = 'homelab-ci'
versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
)
image = versions['BUILDKIT_IMAGE']
signature = builder_config / 'homelab-ci-image'
exists = (
subprocess.run( # noqa: S603
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
capture_output=True,
env=env,
).returncode
== 0
)
if exists and (not signature.exists() or signature.read_text().strip() != image):
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
exists = False
if not exists:
command(
'docker',
'buildx',
'create',
'--name',
builder,
'--driver',
'docker-container',
'--driver-opt',
f'image={image}',
'--buildkitd-config',
'.gitea/runner/buildkitd.toml',
env=env,
)
signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
exists = False
if old_digest and previous['inputs'].get(image) == inputs:
exists = (
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'imagetools',
'inspect',
f'{image}@{old_digest}',
],
capture_output=True,
env=env,
timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--push',
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--tag',
f'{image}:sha-{sha}',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
release['images'][image] = digest
release['inputs'][image] = inputs
validate_release(release, sha)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
# Cleanup errors must neither leak credentials nor mask the original build error.
try:
subprocess.run( # noqa: S603
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'prune',
'--builder',
'homelab-ci',
'--force',
'--max-used-space',
'1gb',
],
env=env,
timeout=60,
)
except (OSError, subprocess.TimeoutExpired):
print('CI builder cache cleanup deferred', flush=True)
finally:
shutil.rmtree(docker_config)
def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile(
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
)
rendered = []
for line in stream:
match = image_line.fullmatch(line.rstrip('\n'))
if match:
prefix, quote, image, tail = match.groups()
if image in EXTERNAL_IMAGES and f'{image}@sha256:' in line:
rendered.append(line)
continue
if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
rendered.append(line)
destination.writelines(rendered)
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('build', 'gate', 'render'))
parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='')
args = parser.parse_args()
if args.action == 'render':
render(sys.stdin, sys.stdout)
elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha)
else:
build(args.output)
if __name__ == '__main__':
main()
+1 -15
View File
@@ -2,23 +2,9 @@ name: renovate-ci
on:
pull_request:
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
push:
branches:
- main
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
workflow_dispatch:
permissions:
@@ -26,7 +12,7 @@ permissions:
jobs:
validate-renovate:
runs-on: homelab
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
+1 -1
View File
@@ -32,7 +32,7 @@ concurrency:
jobs:
run-renovate:
runs-on: homelab
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
steps:
- name: Checkout repository
-13
View File
@@ -1,13 +0,0 @@
# kubectl emits a List for files containing multiple resources.
(if .kind == "List" then .items[] else . end)
| (.metadata.namespace // "default") as $ns
| [
(.. | objects
| (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty)
| select(.optional != true)
| .name // .secretName // empty),
(.. | objects | .imagePullSecrets[]?.name)
]
| unique[]
| select(. != null and . != "")
| "\($ns) \(.)"
+63 -48
View File
@@ -1,56 +1,71 @@
#!/usr/bin/env bash
# The SSH client submits once and follows durable stages on workstation.
# usage: ssh-run.sh <stage>
# Runs one deploy-lib.sh stage on the workstation over SSH.
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
# The private key is written to a per-run directory that is removed on exit, so a
# failed or cancelled job cannot leave deploy credentials in the runner's temp
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT
chmod 700 "$key_dir"
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
chmod 600 "$key_dir/key" "$key_dir/known_hosts"
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
controller=.local/lib/homelab-deploy/controller.py
case "${1:?start, apply, verify or smoke required}" in
start)
python3 - <<'PY' >"$key_dir/request.json"
import json
import os
from pathlib import Path
release = json.loads(Path('release.json').read_text())
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
PY
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
sleep 5
done
exit "$rc"
trap 'rm -rf "$key_dir"' EXIT INT TERM
ssh_key="$key_dir/deploy_key"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
# A connection that died silently used to hang until the job timeout, and the
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
# fails on its own merits exits with the remote's status, so a real failure
# still surfaces its own log instead of burning three attempts. The stages are
# declarative applies, so re-running one that had already committed is harmless.
ssh_opts=(
-i "$ssh_key" -p "$deploy_port"
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
-o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
)
rc=0
# apply-k8s and apply-compose are separate workflow jobs so the graph stays
# intact for the verify job, but on a single node they must not run at once:
# host docker churn on top of cluster churn is what melts the node (load 40+,
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation
# with a shared lock; whoever arrives second waits.
remote_cmd=(bash -se)
case "$1" in
apply-k8s | apply-compose)
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
;;
apply|verify|smoke)
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
sleep 5
done
exit "$rc"
;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac
for attempt in 1 2 3; do
if [ "$attempt" -gt 1 ]; then
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt * 5))
fi
rc=0
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
source "$REPO/.gitea/workflows/deploy-lib.sh"
run_stage "$STAGE"
EOF
[ "$rc" -eq 0 ] && break
[ "$rc" -ne 255 ] && break
done
if [ "$rc" -ne 0 ]; then
echo ":: error::stage $1 failed over ssh (exit $rc)"
fi
exit "$rc"
-6
View File
@@ -31,9 +31,3 @@ UV_VERSION="0.12.17"
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
# Secret-reference regression tests parse rendered Kubernetes objects.
JQ_VERSION="1.8.1"
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
+163
View File
@@ -0,0 +1,163 @@
# Homelab
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
Gitea Actions that build and deploy them. Most applications have both deployment
formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes
provides their ingress through Services and EndpointSlices.
These files contain this lab's domains, IP addresses, storage paths, and private
registry names. Running them on another machine takes some editing.
## Start here
- [Service list](#services) — what each directory contains.
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
- [Repository review](docs/repository-review.md) — confirmed problems and fix branches.
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
[cert-manager](cert-manager/README.md) — common dependencies.
## What gets deployed
The `active` files are switches for the deploy workflow, not health indicators.
| File | Effect |
| ---------------------- | ----------------------------------------------------------- |
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
| Both | Run the Compose stack and apply the Kubernetes resources. |
| Neither | Keep the configuration in Git without automatic deployment. |
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
manual entry points. The deploy script does not discover them.
Kubernetes selection excludes secret files, examples, Helm values, and patches.
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
does not install their charts.
The table below describes committed configuration. It does not claim that a
service is currently healthy or running.
## Services
| Service | Configuration | Selected by markers |
| ----------------------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
## Running a Compose stack
Use the service README first. Where a service has an env example, copy it inside
that service's directory and replace the placeholders. The root `.env.example`
is an older collection of variables, not a complete configuration for every stack.
For example, from the repository root:
```sh
cd netbox
cp .env.example .env
$EDITOR .env
docker compose config --quiet
docker compose up -d
docker compose ps
```
Stacks that attach to `proxy` require an existing Docker network of that name and
an appropriate reverse proxy. Published host ports still work independently of
Traefik. Check port conflicts before starting an alternative to a Kubernetes
service: DNS, STUN, and HTTP listeners can share the same host.
`docker compose down` keeps named volumes. Adding `-v` removes them.
## Preparing Kubernetes
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
Replace the lab's hosts and addresses before using the configuration elsewhere.
Create a service's namespace, then prepare its ignored Secret from the example.
For example:
```sh
kubectl apply -f netbox/k8s/namespace.yaml
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
$EDITOR netbox/k8s/secrets.yaml
kubectl apply -f netbox/k8s/secrets.yaml
```
The deploy workflow applies the tracked resources for marked services. Avoid
applying an entire `k8s/` directory blindly: some directories contain Helm values,
examples, and alternative routes. For a manual change, apply the selected manifest
explicitly and check the resulting rollout.
Shared database passwords must agree between the `database` namespace and each
application's Secret. Updating the PostgreSQL Secret does not change an existing
role's password; see the database README.
## Local checks
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
```sh
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
.gitea/workflows/sync-renovate-configmap.sh --check
```
The [workflow README](.gitea/README.md#checks) lists the rest of the checks.
Structure checks do not establish that local Secrets, mounted files, storage,
or external services are ready.
## Data and recovery
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
configuration. Keep backups of application data and the keys needed to read it.
An image rollback does not roll back database migrations or ConfigMap contents.
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
The manifests do not provide a repository-wide backup schedule.
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
is a planning document, not evidence that NFS has been installed.
+22
View File
@@ -0,0 +1,22 @@
# AdGuard Home
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
configuration and working data, and mounts the `adguard-certs` TLS Secret.
The LoadBalancer Service exposes DNS separately from the web ingress.
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
and `certs/` before starting it. Starting both DNS deployments on the same address
can cause a port conflict.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n adguard
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -31,7 +31,7 @@ services:
- "traefik.http.routers.adguard-dev.entrypoints=websecure"
- "traefik.http.routers.adguard-dev.tls=true"
# DoH Router
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)"
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
- "traefik.http.routers.dns-over-https.entrypoints=websecure"
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
+2 -2
View File
@@ -51,8 +51,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: adguard-deployment
namespace: adguard
spec:
@@ -66,6 +64,8 @@ spec:
metadata:
labels:
app: adguard
annotations:
reloader.stakater.com/auto: "true"
spec:
containers:
- name: adguard
+22
View File
@@ -0,0 +1,22 @@
# Authentik
Identity provider with separate server and worker deployments.
Kubernetes connects to the shared PostgreSQL service in `database`. Set
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
Keep `AUTHENTIK_SECRET_KEY` with the backups.
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
Its image defaults differ from Kubernetes; check both before an upgrade.
The worker mounts the Docker socket for Docker outpost management.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n authentik
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-4
View File
@@ -27,8 +27,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-server-deployment
namespace: authentik
spec:
@@ -65,8 +63,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-worker-deployment
namespace: authentik
spec:
+18
View File
@@ -0,0 +1,18 @@
# cert-manager
Public ACME issuers and an internal certificate authority.
This directory contains chart values and issuer resources, not the controller
installation. Install the cert-manager chart with CRDs and the settings in
`k8s/cert-manager-values.yaml` before applying the issuers.
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
reachability must work for the requested names before issuance.
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
up; the tracked `.crt` is only a public certificate.
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
`kubectl apply` does not interpret the Helm values file.
See the [repository README](../README.md) for deployment selection.
+22
View File
@@ -0,0 +1,22 @@
# Cloudflare DDNS
Updates the lab DNS records when the public address changes.
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
The Compose stack also uses host networking. Configure the API token and domain
list from the relevant example; keep DNS names consistent with the ingress rules.
`config.json.example` is a separate configuration example. The current Compose
file does not mount a config.json file. Check configuration against the pinned
DDNS image when changing between environment and file-based settings.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cfddns
labels:
app: cfddns
+21
View File
@@ -0,0 +1,21 @@
# Checkmk
Checkmk Raw monitoring site with web and agent-receiver ingress.
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
volume on Compose. The agent receiver has a separate TCP route; enabling the
web route alone does not expose it.
Prepare the password in the service env or Secret example. Inspect the Checkmk
container logs during the first site creation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n checkmk
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: checkmk-deployment
namespace: checkmk
spec:
+21
View File
@@ -0,0 +1,21 @@
# Cloudflare Tunnel
A Kubernetes connector for an existing Cloudflare tunnel.
The Deployment runs in `default` and reads its token from the ignored Secret
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
in Cloudflare before starting the connector.
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
`k8s/deployment.yaml` when this tunnel is needed.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -3
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cloudflared
labels:
app: cloudflared
@@ -20,7 +18,7 @@ spec:
spec:
containers:
- name: cloudflared
image: cloudflare/cloudflared:2026.10.0
image: cloudflare/cloudflared:2026.9.3
imagePullPolicy: IfNotPresent
args:
- tunnel
+22
View File
@@ -0,0 +1,22 @@
# File converters
ConvertX for server-side conversion and BentoPDF for PDF tools.
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
Kubernetes configuration includes a local `config.yaml.example`, excluded from
normal deployment. Copy and apply the real ConfigMap separately where required.
Compose publishes ConvertX on host port 9992 as well as attaching it to the
proxy network. Replace the authentication settings from `.env.example` before
exposing it outside the lab.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n converters
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: convertx-deployment
namespace: converters
spec:
+25
View File
@@ -0,0 +1,25 @@
# CrowdSec
Helm values, dashboards, network policy, and a maintenance CronJob.
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
marker in this directory.
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
namespace Role. Review its script and schedule before enabling cleanup.
Traefik's values state that enforcement moved to a host firewall bouncer. This
repository does not install that host component.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n crowdsec
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+21
View File
@@ -0,0 +1,21 @@
# Dockmon
Docker management UI that talks to the host Docker daemon.
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
to the node hosting the pod, so this is not a cluster-wide container manager.
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
with a volume claim template. Its ServersTransport is specific to the upstream
connection; keep it with the ingress resources.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n dockmon
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+151
View File
@@ -0,0 +1,151 @@
# Repository review
Reviewed the tracked tree at `cc9c3de` and read the live workstation state on
6 October 2026. Changes are split into documentation and individual fix branches,
all based on that main commit. The original local checkout and its uncommitted
monitoring changes were preserved. No deployment was performed.
## Confirmed problems with prepared fixes
| Priority | Problem and consequence | Fix branch |
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |
| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` |
| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` |
| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` |
| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` |
| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` |
| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` |
| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` |
| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` |
| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` |
| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` |
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
The fix retains the DoH path constraint for both hostnames.
The prune fix deliberately rejects the unsafe option. It does not introduce
automatic deletion under a different implementation. Prune defaults to false,
and no tracked resource currently carries the selector label, so this is a
latent defect rather than evidence of a live deletion incident.
The session fix bounds HTTP and Redis calls, validates required credentials,
sets a cookie lifetime of two refresh intervals, and marks success only after
publishing the verified cookie. With the default ten-minute interval, an outage
longer than twenty minutes will make the existing Redis-key readiness checks fail.
That is an intentional change from indefinite apparent readiness.
The deployment fix extracts required pod Secret references from rendered JSON,
checks their namespaces, includes init containers, image-pull credentials, and
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
issued by cert-manager are not treated as pre-existing pod prerequisites.
It checks existence/access, not every key's contents or application validity.
## Live workstation observations
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
no pods were Pending or in another non-running, non-completed phase. This is a
point-in-time observation, not a complete application health test.
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
reviewed local commit. It has untracked host configuration and a separate
`userbot/` directory. It was not reset or cleaned.
| Observed difference | Implication |
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
The monitoring files already modified in the user's local tree correspond to the
live migration. They are excluded from these branches. Reconcile that work before
using this review's baseline to deploy monitoring.
## Remaining work
These need recovery design or infrastructure decisions rather than a small
configuration correction:
- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries
exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot.
If the first attempt already changed workloads, the retry snapshots that partial
state. Preserve a run-specific original baseline and verify it across retries.
- **Rollback can exceed the job budget.** Verification is parallel, but
`rollback_workloads` is serial with a five-minute limit per workload. The
thirty-minute job budget can expire before recovery finishes. Bound recovery
concurrency and account for both phases before choosing a new timeout.
- **Snapshot collection is allowed to fail.** Generation and workload snapshot
errors are warnings; verify can fall back to all workloads. A snapshot failure
must not permit unrelated workloads to be selected for automatic undo.
- **Rollback uses the previous revision, not the captured revision.** `rollout undo`
without an explicit revision cannot guarantee restoration to the snapshot after
retries or intervening rollouts. First deployments also have no previous revision.
- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the
target commit's successful CI run or document manual dispatch as an operator
override with its own required checks.
- **Direct Traefik API exposure is unauthenticated.** The latest local commit
explicitly added it for Homarr. Preserve that integration while choosing a
cluster-internal authenticated path or a verified network restriction; do not
simply disable an integration that is already in use.
- **Storage retention and backup are not reproducible as a whole.** Defaults and
several important PV policies are Delete. There is no repository-wide backup
schedule. Existing PVC StorageClass changes require migration rather than an
in-place YAML edit.
- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB
emptyDir. Decide whether pod replacement should discard files or whether it
should use persistent storage. Compose uses a host directory instead.
- **First-time activation needs a bootstrap path.** Deploy validation dry-runs
namespaced resources before the apply stage creates namespaces and installs
selected charts. On a fresh cluster, missing namespaces and CRDs need separate
preparation; activation is not a complete installer.
## Validation
Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose
files, and Kubernetes resources with available schemas. Kubeconform found 347
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
That skip count matters: passing schema validation does not validate Traefik rule
strings or other controller-specific behavior.
Fix validation covers:
- Compose discovery of manual entry points, rejection of required-variable gaps,
namespace-scoped and optional Secret references, and API/render failures.
- NetBird setup idempotence, preservation of existing keys, file permissions,
runtime rendering, and rejection of invalid trusted proxy CIDRs.
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
missing credentials, and nonpositive refresh intervals.
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
- YAML and Compose structure for the corrected router rules, compared with the
documented Traefik grammar. They were not exercised on the live proxy.
- Prune rejection before any cluster invocation.
All seven fix branches and the documentation branch merged together without
conflicts in a disposable validation worktree. The combined tree passed the
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
Compose structure checks, and 11 Python regression tests plus the shell
validation regressions. CRD server-side validation and live rollout tests were
not run.
Runtime tests use fixtures and mocks, not production credentials. Live checks read
workload metadata, storage policies, chart versions, and container state only.
They did not read Secret contents or change services.
## Reloader follow-up
`fix/reloader-integration` adds the active marker and opt-in annotations to 28
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
It corrects AdGuard's misplaced pod-template annotation. The Helm settings use
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
CronJobs. PostgreSQL workloads are excluded because their credential variables
and init scripts are only effective on an empty data directory.
The controller was already running on workstation when inspected. Its live
configuration is unchanged by the branch: merge and deploy the integration to
apply the new policy and application annotations. Configuration reload behavior
was checked against the pinned chart, with Helm rendering and manifest validation;
no production configuration was changed to provoke a test restart.
+20
View File
@@ -0,0 +1,20 @@
# Downtify
Download UI with a persistent downloads directory.
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
so check certificate and middleware availability before enabling them.
Back up downloads separately if they need to survive storage replacement.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n downtify
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+50 -12
View File
@@ -1,14 +1,52 @@
# EDU deployment ownership
# EDU session keeper and Telegram bot
Application source and release builds: `forust/edu-master`.
The homelab pipeline deploys `edu_master/k8s` and preserves explicit image digests.
The application copies in this directory are legacy and are not build inputs.
Do not publish EDU `prod` images from homelab or resolve releases from moving tags.
Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands.
For an EDU release, validate both images, select their digests in the keeper and
checker manifests, and run the existing homelab validation/apply/verification
helpers against this service. Keep the existing Secret and Redis PVC.
Coordinate Redis authentication changes with both clients and all init/probes;
keep a pre-rollout Redis backup and both previous compatible image references.
The current HTTP checker does not depend on Playwright; check other consumers
before removing the separate browser service.
`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis.
`webinar-checker/` uses that cookie through a remote Playwright browser and stores
subscribers, language preferences, and webinar history in Redis.
Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`.
`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the
checker's metrics on port 8000. Its `/health` endpoint reflects recent checks.
## Configuration
Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose
`.env.example` has stale names until `fix/session-keeper-reliability` is merged.
The code reads:
| Variable | Purpose |
| ----------------------------------------------------- | ------------------------------------------------ |
| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. |
| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. |
| `EDU_URL_BASE` | EDU site origin. |
| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. |
| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. |
| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. |
| `REDIS_HOST`, `REDIS_PORT` | Redis connection. |
| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. |
Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python
package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching
versions. The two Python images are built and published by CI.
## Bot use
Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`,
`/schedule`, and `/setclass` manage subscriptions and school views. The
administrator can manage the whitelist with `/adduser` and `/removeuser`.
Back up Redis if subscriber settings and notification history matter. Session
cookies and Telegram tokens are credentials; keep them out of shared logs.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n edu-master
kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+12 -31
View File
@@ -50,36 +50,7 @@ spec:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / http error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
- alert: WebinarCheckerNeverStarted
expr: |
(time() - edu_process_start > 120)
and (webinar_check_last_run_timestamp_seconds == 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker job has not started"
description: "The process exposes metrics but its webinar job has never started."
- alert: WebinarDeliveryPending
expr: edu_delivery_pending > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Webinar notifications await delivery"
description: "Telegram delivery has pending recipients. Check delivery failures and retry status."
- alert: EduRedisUnavailable
expr: edu_redis_connected == 0
for: 2m
labels:
severity: critical
annotations:
summary: "EDU checker cannot reach Redis"
description: "Redis health checks are failing; checker commands and delivery may be unavailable."
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
@@ -103,7 +74,7 @@ spec:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker deployment unavailable.
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
@@ -113,3 +84,13 @@ spec:
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
-22
View File
@@ -1,22 +0,0 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: redis-clients-only
namespace: edu-master
spec:
podSelector:
matchLabels:
app: edu-master-redis
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: edu-master-session-keeper
- podSelector:
matchLabels:
app: edu-master-webinar-checker
ports:
- protocol: TCP
port: 6379
-21
View File
@@ -20,27 +20,6 @@ spec:
- name: redis
image: redis:8.10.2-alpine
imagePullPolicy: IfNotPresent
env:
- name: REDIS_PASSWORD
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
- |
case "$REDIS_PASSWORD" in *[!0-9a-fA-F]*|'') echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1;; esac
[ "${#REDIS_PASSWORD}" -eq 64 ] || { echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1; }
umask 077
printf 'requirepass "%s"\n' "$REDIS_PASSWORD" > /tmp/redis-auth.conf
chown redis:redis /tmp/redis-auth.conf
exec docker-entrypoint.sh redis-server /tmp/redis-auth.conf
ports:
- containerPort: 6379
volumeMounts:
+1 -17
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: session-keeper
namespace: edu-master
labels:
@@ -16,20 +14,12 @@ spec:
type: Recreate
template:
metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels:
app: edu-master-session-keeper
spec:
initContainers:
- name: wait-redis
image: redis:8.10.2-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
@@ -43,16 +33,10 @@ spec:
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper@sha256:49285e87cc5bc4cf4ffe190813d87927916c2df8a206daac0aeb7d227c636450
image: gcr.forust.xyz/forust/session-keeper:prod
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
resources:
requests:
cpu: 25m
+8 -20
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: webinar-checker
namespace: edu-master
labels:
@@ -16,22 +14,14 @@ spec:
type: Recreate
template:
metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels:
app: edu-master-webinar-checker
spec:
# Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID)
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.2-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
@@ -49,9 +39,15 @@ spec:
sleep 2
done
echo "PHPSESSID ok"
until nc -z playwright-service 3000; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
sleep 2
done
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker@sha256:66c146f7b43cb9f0dc31ba9aa36d217e01df42ddafba5971b79c12ec215b2c01
image: gcr.forust.xyz/forust/webinar-checker:prod
ports:
- name: metrics
containerPort: 8000
@@ -64,14 +60,6 @@ spec:
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
livenessProbe:
httpGet:
path: /live
port: metrics
initialDelaySeconds: 60
periodSeconds: 15
timeoutSeconds: 3
failureThreshold: 4
envFrom:
- secretRef:
name: edu-master-secrets
+1 -1
View File
@@ -1,4 +1,4 @@
FROM python:3.14-slim
FROM python:3.11-slim
WORKDIR /app
+1 -1
View File
@@ -1,4 +1,4 @@
FROM python:3.14-slim
FROM python:3.11-slim
WORKDIR /app
+21
View File
@@ -0,0 +1,21 @@
# Error pages
Static HTTP error pages served by an Nginx image built in CI.
Edit the HTML in `html/`; the Dockerfile copies it into the image.
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
middleware. Keep the middleware's namespace and port aligned with that Service.
For a local build, run `docker build -t homelab-error-pages .` from this directory.
Compose references the private registry image rather than a build context.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n error-pages
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+25
View File
@@ -0,0 +1,25 @@
# Gitea
Git hosting with HTTP and a separate SSH route.
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
and application data. Match the Gitea database password with the shared database
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
its own database, not a second frontend for the Kubernetes instance.
Back up repositories, application configuration, and a consistent database dump
together. Gitea Actions definitions for this repository live in `../.gitea/`.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n gitea
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: gitea-deployment
namespace: gitea
spec:
+26
View File
@@ -0,0 +1,26 @@
# Glance
Dashboard pages for links, service checks, and Docker containers.
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
`user.css`. Update both copies when changing shared content.
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
no tracked Secret example; create that Secret in `glance` before starting it.
Compose expects a local `.env` with the same password.
The CSS mount points at the wrong ConfigMap on the reviewed main commit;
`fix/glance-assets` corrects it.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n glance
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -3
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: glance-deployment
namespace: glance
spec:
@@ -71,7 +69,7 @@ spec:
name: glance-config
- name: glance-assets
configMap:
name: glance-assets
name: glance-config
- name: docker-socket
hostPath:
path: /var/run/docker.sock
+18
View File
@@ -0,0 +1,18 @@
# Headscale
Headscale, Headplane, and a separate web administration UI on Docker.
Kubernetes only provides routes to the Docker host. Update the addresses in
`k8s/routing/external-service.yaml` if the host moves.
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
`config/policy.json.example` to their names without `.example`. Set the public
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
The example URLs are placeholders.
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
13000, and the other UI on 10080. The data volumes store the Headscale database,
keys, and Headplane state. The embedded DERP configuration needs reachable
addresses; Compose does not publish its UDP 3478 listener.
See the [repository README](../README.md) for deployment selection.
+24
View File
@@ -0,0 +1,24 @@
# Homarr
Dashboard with Kubernetes integration and persistent application state.
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
The committed ingress is internal. There is no `k8s/active` marker even though
manifests exist, so the workflow does not select Homarr automatically.
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
expects a local kubeconfig. Check these host ports against Traefik before use.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n homarr
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,7 +1,7 @@
services:
homarr:
container_name: homarr
image: ghcr.io/homarr-labs/homarr:v2.2.0
image: ghcr.io/homarr-labs/homarr:v2.1.2
restart: unless-stopped
volumes:
- ./appdata:/appdata
+1 -3
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: homarr-deployment
namespace: homarr
spec:
@@ -32,7 +30,7 @@ spec:
serviceAccountName: homarr
containers:
- name: homarr
image: ghcr.io/homarr-labs/homarr:v2.2.0
image: ghcr.io/homarr-labs/homarr:v2.1.2
envFrom:
- configMapRef:
name: homarr-config
+25
View File
@@ -0,0 +1,25 @@
# Homepages
Two static sites: Forust and xdfnx.
The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with
its own Dockerfile and publishes it to the private registry. Kubernetes serves
the image contents; Compose overlays the source directories as bind mounts.
Both Traefik IngressRoute and Gateway API route manifests are committed.
Keep their hostnames and backend Services aligned when changing routes.
Certificate resources cover public and internal hostnames.
Build either site locally with `docker build -f Dockerfile.forust .` or
`docker build -f Dockerfile.xdfnx .` from this directory.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n homepages
kubectl get events -n homepages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+29
View File
@@ -0,0 +1,29 @@
# Immich
Photo library with its own vector-enabled PostgreSQL and machine-learning service.
This database is separate from the shared PostgreSQL instance. Keep the server
and machine-learning versions aligned when upgrading.
Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must
already exist and contain the intended library; moving the pod to a different
node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and
the model cache has its own PVC.
Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example
uses the same library path as Kubernetes. Run one writer against that library;
do not start both deployments as independent instances over the same files.
Back up the library and a consistent database dump together. The model cache
can be rebuilt; the photo database cannot.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n immich
kubectl get events -n immich --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -14,8 +14,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-deployment
namespace: immich
labels:
-2
View File
@@ -14,8 +14,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-machine-learning-deployment
namespace: immich
labels:
-2
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1
kind: StatefulSet
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-valkey
namespace: immich
labels:
+21
View File
@@ -0,0 +1,21 @@
# Kener
Status page with Redis and persistent database and upload directories.
Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet.
Compose keeps the corresponding directories in named volumes. Set the signing
and other credentials from the env or Secret example.
The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`.
There is no active marker for either runtime.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n kener
kubectl get events -n kener --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: kener-deployment
namespace: kener
spec:
+16
View File
@@ -0,0 +1,16 @@
# Loki and Alloy
Loki log storage and Alloy collection, both deployed through Helm.
The deploy library lists separate `loki` and `alloy` releases in `prometheus`,
controlled by this directory's `k8s/active` marker. Chart versions are pinned in
`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`.
Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the
monitoring stack. Review Loki retention and storage settings before enabling
collection on a new cluster.
Check releases with `helm list -n prometheus` and inspect collector logs before
assuming that an empty Grafana query means there were no events.
See the [repository README](../README.md) for deployment selection.
+21
View File
@@ -0,0 +1,21 @@
# MeTube
Web downloader behind Traefik.
Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB
`emptyDir` for `/downloads`: completed downloads disappear when the pod is
replaced. Download files from the UI promptly if this temporary storage is intended.
Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes
would require changing the volume to a PVC and choosing a storage policy.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n metube
kubectl get events -n metube --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: metube-deployment
namespace: metube
spec:
+21
View File
@@ -0,0 +1,21 @@
# n8n
Workflow automation with persistent application and file storage.
Kubernetes keeps application state in `n8n-node-pvc` and files in
`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes.
Webhook URLs and proxy settings are committed in the application config.
There is no active marker. Review the URLs before enabling the stack, and retain
the credential encryption key with the database or application-data backup.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n n8n
kubectl get events -n n8n --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services:
n8n:
image: docker.n8n.io/n8nio/n8n:2.43.0
image: docker.n8n.io/n8nio/n8n:2.42.3
container_name: n8n
restart: unless-stopped
environment:
+1 -3
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: n8n-deployment
namespace: n8n
spec:
@@ -31,7 +29,7 @@ spec:
spec:
containers:
- name: n8n
image: docker.n8n.io/n8nio/n8n:2.43.0
image: docker.n8n.io/n8nio/n8n:2.42.3
envFrom:
- configMapRef:
name: n8n-config
+43 -102
View File
@@ -1,123 +1,64 @@
# NetBird
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
Self-hosted NetBird with the combined management, signal, relay, and STUN server.
The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
Kubernetes runs the server and dashboard in `netbird`. The server uses SQLite
in `netbird-pvc`; `k8s/config.yaml` contains the template and runtime renderer.
Prepare `netbird-secrets` from `k8s/secrets.yaml.example` before the first start.
## Files
## Routing and keys
- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
- `config.template.yaml`: non-secret server configuration rendered at startup.
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
- `.env`: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key.
- `secrets/`: ignored relay secret and datastore encryption key.
The public hostname is set in the ConfigMap and ingress rules. Keep the issuer,
dashboard endpoints, and public routes consistent. HTTP, WebSocket, and gRPC
traffic go through Traefik; the STUN route uses UDP 3478. A CDN's HTTP proxy does
not provide that UDP listener.
## First deployment
`NETBIRD_PROXY_SUBNET` controls which forwarded client addresses are trusted.
Use the actual proxy network CIDR rather than assuming another lab's subnet.
Keep the datastore encryption key with every datastore backup. Regenerating it
can make stored credentials unreadable.
Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
## Compose alternative
```bash
cd /srv/homelab/netbird
`compose.yaml` expects `entrypoint.sh`, a local `.env`, and two local files:
`secrets/relay-auth-secret` and `secrets/datastore-encryption-key`.
The reviewed main commit is missing the renderer and setup script.
`fix/netbird-compose-runtime` restores them. Merge that fix before following
these setup commands:
```sh
cd netbird
./setup.sh
$EDITOR .env
docker compose config --quiet
docker compose up -d
```
Review the values in `.env` before starting. The example public hostname is `netbird.forust.xyz`; change it if a different public domain was selected. `setup.sh` replaces `NETBIRD_PROXY_SUBNET=auto` with the first IPv4 subnet of the external Docker `proxy` network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere.
`setup.sh` is idempotent and never replaces existing secrets. Do not delete or regenerate `secrets/datastore-encryption-key` after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated.
## Network prerequisites
- Point the public hostname directly to the Docker host. Do not proxy UDP `3478` through Cloudflare or another CDN.
- Allow inbound TCP `80`, TCP `443`, and UDP `3478` through the host firewall and upstream router.
- Ensure the external `proxy` Docker network exists and Traefik uses its `websecure` entrypoint and `letsencrypt` resolver. `NETBIRD_PROXY_SUBNET` must describe that network; it is used to trust only forwarded client addresses from Traefik.
- Ensure the internal names in `.env` resolve where the local and development aliases are needed.
- Keep Traefik's `websecure` read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures `--entrypoints.websecure.transport.respondingTimeouts.readTimeout=0` in `traefik/compose.yaml`.
After startup, verify OIDC discovery through the public TLS endpoint:
```bash
curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration"
```
Open `https://${NETBIRD_DOMAIN}` immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists.
The setup script detects the IPv4 subnet of the external `proxy` network and
preserves existing secrets. Complete the initial owner setup through the public
TLS endpoint after starting the server.
## Optional host client
The client intentionally lives in a separate Compose project. Normal server deploys use `--remove-orphans`, so keeping the client in the server project would cause it to be removed.
`client.compose.yaml` runs a host-network peer in a separate Compose project.
Set `NB_SETUP_KEY` and `NETBIRD_CLIENT_HOSTNAME` in the local `.env`, then run
`docker compose -f client.compose.yaml up -d`. It needs `/dev/net/tun` and elevated
network capabilities. The normal server deployment does not start this client.
1. Create a reusable or ephemeral setup key in the NetBird dashboard.
2. Put `NB_SETUP_KEY=<key>` in the ignored `netbird/.env` file.
3. Set `NETBIRD_CLIENT_HOSTNAME` to this machine's desired peer name.
4. Start and inspect the client:
## Backup
```bash
cd /srv/homelab/netbird
docker compose -f client.compose.yaml config --quiet
docker compose -f client.compose.yaml up -d
docker compose -f client.compose.yaml exec netbird-client netbird status
Back up the SQLite data while the server is stopped, together with the encryption
key, relay secret, and local configuration. Test a restore on an isolated host.
For Compose, the datastore volume has the explicit name `netbird_data`.
Do not use `docker compose down -v` when keeping the installation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netbird
kubectl get events -n netbird --sort-by=.metadata.creationTimestamp
```
The client uses host networking and requires `NET_ADMIN`, `SYS_ADMIN`, `SYS_RESOURCE`, and `/dev/net/tun`. Remove it without affecting the server stack:
```bash
docker compose -f client.compose.yaml down
```
## Operations
Inspect status and logs:
```bash
docker compose ps
docker compose logs --tail=200 netbird-server dashboard
```
Stop or remove containers without deleting data:
```bash
docker compose down
```
Do not add `-v` to `docker compose down`; it would delete the NetBird datastore.
## Backup and restore
Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and `datastore-encryption-key` in an encrypted backup:
```bash
cd /srv/homelab/netbird
mkdir -p backups
docker compose stop netbird-server
docker run --rm \
-v netbird_data:/data:ro \
-v "$PWD/backups:/backup" \
busybox:1.37.0 \
tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" .
docker compose start netbird-server
```
Also securely back up:
- `secrets/datastore-encryption-key` — required to decrypt stored secrets.
- `secrets/relay-auth-secret` — keeps issued relay credentials valid across restoration.
- `netbird/.env` — optional, but it records the public and internal hostnames.
Test a restore in an isolated Docker host before relying on a backup.
## Upgrade
1. Take and verify a backup.
2. Review NetBird release notes for server, client, and dashboard compatibility.
3. Update the pinned tags in `compose.yaml`; update `client.compose.yaml` separately when deploying the client.
4. Pull and recreate the selected services:
```bash
docker compose pull
docker compose up -d
```
The image tags are intentionally pinned instead of using `latest`, matching this repository's pull-on-deploy policy.
See the [repository README](../README.md) for deployment selection.
-109
View File
@@ -1,109 +0,0 @@
#!/bin/sh
set -eu
umask 077
TEMPLATE_PATH=/opt/netbird/config.template.yaml
RENDERED_PATH=/run/netbird/config.yaml
RELAY_SECRET_PATH=/run/secrets/relay_auth_secret
ENCRYPTION_KEY_PATH=/run/secrets/datastore_encryption_key
is_valid_proxy_subnet() {
candidate="$1"
case "$candidate" in
0.0.0.0/0)
return 1
;;
*/*)
address="${candidate%%/*}"
prefix="${candidate#*/}"
;;
*)
return 1
;;
esac
case "$prefix" in
0|[1-9]|[1-2][0-9]|3[0-2]) ;;
*)
return 1
;;
esac
old_ifs="$IFS"
IFS=.
# shellcheck disable=SC2086
set -- $address
IFS="$old_ifs"
[ "$#" -eq 4 ] || return 1
for octet do
case "$octet" in
0|[1-9]|[1-9][0-9]|1[0-9][0-9]|2[0-4][0-9]|25[0-5]) ;;
*)
return 1
;;
esac
done
}
read_secret() {
secret_path="$1"
if [ ! -r "$secret_path" ]; then
echo "Required secret is not readable: $secret_path" >&2
exit 1
fi
secret_value="$(cat "$secret_path")"
if [ -z "$secret_value" ]; then
echo "Required secret is empty: $secret_path" >&2
exit 1
fi
printf '%s' "$secret_value"
}
if [ -z "${NETBIRD_DOMAIN:-}" ]; then
echo "NETBIRD_DOMAIN must be set" >&2
exit 1
fi
case "$NETBIRD_DOMAIN" in
*[!A-Za-z0-9.-]*)
echo "NETBIRD_DOMAIN contains unsupported characters" >&2
exit 1
;;
esac
if [ -z "${NETBIRD_PROXY_SUBNET:-}" ] || [ "$NETBIRD_PROXY_SUBNET" = "auto" ]; then
echo "NETBIRD_PROXY_SUBNET must be an explicit IPv4 CIDR; run netbird/setup.sh first" >&2
exit 1
fi
if ! is_valid_proxy_subnet "$NETBIRD_PROXY_SUBNET"; then
echo "NETBIRD_PROXY_SUBNET must be a non-default IPv4 CIDR, for example 172.20.0.0/16" >&2
exit 1
fi
if [ "$#" -ne 2 ] || [ "$1" != "--config" ] || [ "$2" != "$RENDERED_PATH" ]; then
echo "Expected: --config $RENDERED_PATH" >&2
exit 1
fi
relay_secret="$(read_secret "$RELAY_SECRET_PATH")"
encryption_key="$(read_secret "$ENCRYPTION_KEY_PATH")"
mkdir -p "$(dirname "$RENDERED_PATH")"
sed \
-e "s|__NETBIRD_DOMAIN__|${NETBIRD_DOMAIN}|g" \
-e "s|__NETBIRD_AUTH_SECRET__|${relay_secret}|g" \
-e "s|__NETBIRD_ENCRYPTION_KEY__|${encryption_key}|g" \
-e "s|__NETBIRD_PROXY_SUBNET__|${NETBIRD_PROXY_SUBNET}|g" \
"$TEMPLATE_PATH" >"$RENDERED_PATH"
if grep -q '__NETBIRD_' "$RENDERED_PATH"; then
echo "Rendered NetBird configuration still contains unresolved placeholders" >&2
exit 1
fi
exec /go/bin/netbird-server "$@"
-4
View File
@@ -32,8 +32,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbird-server-deployment
namespace: netbird
spec:
@@ -128,8 +126,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbird-dashboard-deployment
namespace: netbird
spec:
-38
View File
@@ -1,38 +0,0 @@
#!/usr/bin/env bash
# Prepare local Compose configuration without replacing existing credentials.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")"
umask 077
if [ ! -f .env ]; then
cp .env.example .env
fi
if grep -q '^NETBIRD_PROXY_SUBNET=auto$' .env; then
subnet="$(docker network inspect proxy --format '{{range .IPAM.Config}}{{println .Subnet}}{{end}}' | awk '/^[0-9]+\./ { print; exit }')"
if [ -z "$subnet" ]; then
echo "No IPv4 subnet found on the Docker proxy network. Set NETBIRD_PROXY_SUBNET in .env." >&2
exit 1
fi
# The detected value must be safe to substitute into the env file.
if [[ ! "$subnet" =~ ^[0-9.]+/[0-9]+$ ]]; then
echo "Unexpected Docker network subnet: $subnet" >&2
exit 1
fi
sed -i "s|^NETBIRD_PROXY_SUBNET=auto$|NETBIRD_PROXY_SUBNET=$subnet|" .env
fi
mkdir -p secrets
chmod 700 secrets
for name in relay-auth-secret datastore-encryption-key; do
path="secrets/$name"
if [ -e "$path" ]; then
if [ ! -s "$path" ]; then
echo "Existing secret is empty: $path. Restore it before continuing." >&2
exit 1
fi
else
openssl rand -base64 32 >"$path"
fi
chmod 600 "$path"
done
printf '%s\n' 'Local files are ready. Review .env, then run docker compose config --quiet.'
+35 -88
View File
@@ -1,96 +1,43 @@
# NetBox
NetBox for homelab documentation and visualization. Two runtimes are available:
Inventory and network documentation with a web process, worker, and Valkey.
| Runtime | Manifest | Purpose |
| ------- | -------------- | -------------------------------------------------------------- |
| Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) |
| k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) |
Kubernetes uses the shared PostgreSQL service at
`postgres.database.svc.cluster.local:5432`, database and role `netbox`.
The database and application Secrets must contain the same password.
Media, reports, scripts, and Valkey have persistent storage.
Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks
plus a second logical database for caching. The Docker stand keeps its own
PostgreSQL container, while the k8s deployment uses the shared `database` cluster
(`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey
stays a per-service StatefulSet.
Compose has its own PostgreSQL container and Valkey instances. It publishes the
web UI on `127.0.0.1:8000`; its Traefik labels can also expose it while a Docker
proxy is running. Copy `.env.example` to `.env`, replace the credentials, and run
`docker compose config --quiet` before starting it.
## Docker Compose
## First Kubernetes start
```bash
cp .env.example .env
# replace CHANGE_ME
docker compose up -d
Create the namespace and application Secret. Provision the database through the
shared database initializer on a fresh instance, or create the role and database
manually on an existing instance; see [PostgreSQL](../postgres/README.md).
The database NetworkPolicy already includes `netbox`.
Apply the selected application manifests after the database is ready. Startup
runs schema migrations, so the probes allow a longer first boot. Inspect web and
worker logs before retrying a slow migration.
## Settings and backup
`configuration/configuration.py` is the Compose settings file. Its Kubernetes
copy is embedded in `k8s/settings.yaml`; keep them aligned.
Back up the database and media together. Keep `SECRET_KEY` and
`API_TOKEN_PEPPER_1`: changing them invalidates sessions or API tokens.
A container rollback cannot undo a database migration.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netbox
kubectl get events -n netbox --sort-by=.metadata.creationTimestamp
```
The UI is available at <http://localhost:8000>. The port is bound to `127.0.0.1`
intentionally, so this stand is not exposed on the LAN or public interfaces.
The `netbox` service is also attached to the external `proxy` network and carries
Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those
labels only take effect while the Docker Traefik stack is running; it is currently
stopped, and the live ingress path in this homelab is the k8s Traefik.
Inspect startup and health with:
```bash
docker compose ps
docker compose logs -f netbox
```
Stop it with `docker compose down`; data is kept in the named volumes
`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`,
`netbox-scripts-files` and `netbox-redis-data`.
## Kubernetes
`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly
plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To
rebuild it from scratch:
```bash
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
kubectl -n database patch secret postgres-shared-secrets \
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
# 2. secrets first: the deploy workflow never applies *secret*.yaml
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
kubectl apply -f k8s/secrets.yaml
# 3. manifests
kubectl apply -f k8s/
```
The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its
NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace
or connections are dropped, and `postgres/initdb/01-create-databases.sh` already
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
StatefulSet of its own — only `netbox-valkey`.
`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS`
list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the
`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`.
Resources are permanent again now that the first-boot migrations are complete:
the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the
worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves
`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps
leave enough headroom for future schema migrations without letting one process
consume the whole node.
The first start applies ~810 migrations, each in its own transaction with DDL and
a commit; every later start is a no-op. The startup probe allows 15 minutes and
`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod
and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would
replace the probe destination with that public hostname and bypass the pod.
## Secrets
- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only
`.env.example` and `k8s/secrets.yaml.example` are committed.
- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and
the Django keys all come from the environment, so the same settings file works in
both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap
(`k8s/settings.yaml`) and must be kept in sync with the file.
- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1`
invalidates every API token.
See the [repository README](../README.md) for deployment selection.
-4
View File
@@ -14,8 +14,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-deployment
namespace: netbox
labels:
@@ -120,8 +118,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-worker-deployment
namespace: netbox
labels:
-2
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1
kind: StatefulSet
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-valkey
namespace: netbox
labels:
+22
View File
@@ -0,0 +1,22 @@
# Netronome
Network monitoring application using the shared PostgreSQL instance on Kubernetes.
Kubernetes reads application settings from its ConfigMap and Secret. Match the
Netronome role password with `NETRONOME_DB_PASSWORD` in the shared database Secret.
Its namespace is included in the PostgreSQL NetworkPolicy.
The Compose configuration is a separate deployment; review its local database
settings and env example before starting it. Keep monitoring history in the
database backup.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netronome
kubectl get events -n netronome --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services:
netronome:
image: ghcr.io/autobrr/netronome:v0.16.0
image: ghcr.io/autobrr/netronome:v0.15.0
restart: unless-stopped
container_name: netronome
ports:
+1 -3
View File
@@ -14,8 +14,6 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netronome-deployment
namespace: netronome
labels:
@@ -34,7 +32,7 @@ spec:
spec:
containers:
- name: netronome
image: ghcr.io/autobrr/netronome:v0.16.0
image: ghcr.io/autobrr/netronome:v0.15.0
ports:
- name: netronome-port
protocol: TCP
+17
View File
@@ -0,0 +1,17 @@
# Nextcloud AIO
Nextcloud All-in-One on Docker, with Kubernetes routes to the Docker host.
The master container manages its own child containers through the Docker
socket. Kubernetes does not run the Nextcloud application; the EndpointSlices
under `k8s/routing/` point to host services.
Compose publishes the AIO administration interface on 8888. The Apache frontend
uses host port 11000. `NEXTCLOUD_DATADIR` is `/mnt/nextcloud/ncdata`; prepare that
storage before first setup and do not change the path casually afterwards.
Use AIO's backup and restore tools for the managed application. Keep the master
configuration volume and the data directory with the recovery plan. Do not
remove child containers just because they do not appear as Compose services.
See the [repository README](../README.md) for deployment selection.
+12
View File
@@ -0,0 +1,12 @@
# Penpot
A Compose-only design application with frontend, backend, exporter, database, and cache.
There is no active marker or Kubernetes deployment here. Configure the public
URL and credentials from `.env.example` before starting `compose.yaml`.
Penpot has its own PostgreSQL container. The shared database initializer still
contains a Penpot role, but this Compose stack does not use it.
Back up the application assets and database together.
See the [repository README](../README.md) for deployment selection.
+21
View File
@@ -0,0 +1,21 @@
# Portainer
Container management UI backed by the host Docker socket.
Kubernetes mounts the node's Docker socket and persists application data in
`portainer-data-pvc`. This targets Docker on that node, not Kubernetes workloads.
Compose uses the `portainer_data` volume for its state.
Review initial administrator setup and route access before exposing the UI.
Neither deployment has an active marker.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n portainer
kubectl get events -n portainer --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-1
View File
@@ -1,7 +1,6 @@
POSTGRES_ADMIN_PASSWORD=
AUTHENTIK_DB_PASSWORD=
GITEA_DB_PASSWORD=
NETBOX_DB_PASSWORD=
NETRONOME_DB_PASSWORD=
PENPOT_DB_PASSWORD=
STATUSPAGE_DB_PASSWORD=
+56 -28
View File
@@ -1,35 +1,63 @@
# Shared PostgreSQL
This directory contains the shared PostgreSQL 17 deployment for Authentik,
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
per service. Per-service standalone databases were removed after the
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
not part of the shared instance).
PostgreSQL 17 for the Kubernetes deployments of Authentik, Gitea, NetBox, and Netronome.
## Compatibility baseline
The server runs in `database` as StatefulSet `postgres17`, with data in
`postgres17-data`. Applications connect to
`postgres.database.svc.cluster.local:5432`. The NetworkPolicy allows only the
listed application namespaces; add a new consumer there as well as provisioning
its database.
| Service | Current application | Shared PostgreSQL 17 |
| ---------- | ------------------- | -------------------------------------- |
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
| Statuspage | custom | Supported |
## Initialization
A major-version change must use a logical dump/restore; changing only the
image tag while keeping a data directory is not supported.
`initdb/01-create-databases.sh` creates roles and databases on an empty data
directory. The Kubernetes copy is embedded in `k8s/postgres.yaml`.
It also provisions Penpot and Statuspage roles, even though those are not active
consumers in the current Kubernetes manifests.
For Compose, copy `.env.example` to `.env`, set all passwords, and start it with
`docker compose -f shared-compose.yaml up -d`. This file is intentionally not
named `compose.yaml`, so the repository deploy workflow does not start a second
database accidentally.
Applications that use this database must also join that external network and use
`homelab-postgres:5432`.
The initializer requires every listed password. Prepare `k8s/secrets.yaml` from
the example before applying the StatefulSet. Existing application Secrets keep
copies of their own database passwords; they must match the corresponding role.
For Kubernetes, create `k8s/secrets.yaml` from the example before applying the
manifests. The `k8s/active` marker makes the normal deploy workflow include the
namespace, StatefulSet, ConfigMap, and NetworkPolicy. Applications use
`postgres.database.svc.cluster.local:5432`.
Migrate each existing database with a tested logical dump/restore before
switching an application. Do not reuse a PostgreSQL 14 or 17 data directory
with PostgreSQL 15.
The init scripts do not run again when an existing data directory is mounted.
Changing a Secret does not rotate the PostgreSQL role password. Rotate the role
with SQL and update the application Secret together.
## Compose alternative
From this directory:
```sh
cp .env.example .env
$EDITOR .env
docker compose -f shared-compose.yaml config --quiet
docker compose -f shared-compose.yaml up -d
```
Add `NETBOX_DB_PASSWORD` to `.env` as well: the reviewed env example omits it;
`fix/postgres-env-example` restores the key. Fill every required password.
This stack creates the `homelab-database` Docker network and the
`homelab-postgres` container. Compose applications need to join that network
explicitly to use it; several committed Compose stacks use their own databases.
The filename is intentional: the automatic deploy discovery does not start this
stack just because the Kubernetes database is active.
## Backup and upgrades
Keep database dumps and role definitions, including ownership and grants.
Take a logical backup before changing a major PostgreSQL version. A new image
tag over the existing data directory is not a major-version migration.
Test restores separately before changing application connection settings.
Immich uses its own vector-enabled database and is outside this shared instance.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n database
kubectl get events -n database --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+26
View File
@@ -0,0 +1,26 @@
# Monitoring stack
Prometheus, Grafana, Alertmanager, and application alert rules.
Kubernetes installs `kube-prometheus-stack` in `prometheus` through the deploy
library. Its chart version is pinned there; `k8s/grafana-values.yaml` contains the
values for the whole stack, despite the filename.
The values file is tracked. `.gitignore` also lists it, but that does not stop Git
tracking later edits. Keep local credentials in Secrets rather than treating
changes to this file as ignored.
Prepare Grafana admin and Alertmanager Secrets separately. The Alertmanager
configuration example and Telegram template are in `k8s/`; the main deploy
selection excludes the example config. Certificates and ingress expose Grafana,
and the rule files add service-specific alerts.
The values use `local-path` PVCs for Grafana, Prometheus, and Alertmanager.
Retention is limited by both time and size. Back up Grafana state and any history
that must survive a storage failure.
The Compose stack has separate Prometheus and Alertmanager configuration files.
There is no root active marker. This README describes committed main files;
local VictoriaMetrics experiments are not part of that configuration.
See the [repository README](../README.md) for deployment selection.
-17
View File
@@ -1,17 +0,0 @@
# VictoriaMetrics
The `victoria-operator` Helm release converts Prometheus Operator
`ServiceMonitor` resources into owned `VMServiceScrape` resources. The
`VMAgent` selects converted scrapes labeled `release: prometheus-stack` in all
namespaces and writes them to the existing single-node VictoriaMetrics
instance. Changes to selected `ServiceMonitor` resources are reconciled
automatically; there is no copied Prometheus scrape-config blob to regenerate.
The agent drops targets for the Prometheus server service to avoid duplicating
its self-scrape. `scraper: victoria` identifies the samples ingested by this
VMAgent.
The VictoriaMetrics Operator chart and its CRDs are installed before the
Kubernetes manifests by the normal deploy workflow. On a cluster where the
operator CRDs are not installed yet, CI skips the server-side dry-run of the
`VMAgent` resource; the deploy installs the chart before applying that resource.
-11
View File
@@ -38,8 +38,6 @@ grafana:
# One block covers both the dashboards and datasources sidecars (p95 91M / 80M).
sidecar:
datasources:
defaultDatasourceEnabled: false
resources:
requests:
memory: "96Mi"
@@ -52,18 +50,9 @@ grafana:
type: loki
url: http://loki-gateway.prometheus.svc.cluster.local
access: proxy
- name: VictoriaMetrics
type: prometheus
url: http://victoria-metrics.prometheus.svc.cluster.local:8428
access: proxy
isDefault: true
prometheus:
prometheusSpec:
# VM trial: vmagent scrapes and remote-writes to VictoriaMetrics, so the
# Prometheus server itself stands down. Encoded here (not a kubectl patch)
# so helm keeps owning spec.replicas and upgrades do not conflict on it.
replicas: 0
retention: 60d
retentionSize: 32GB
storageSpec:
-102
View File
@@ -31,105 +31,3 @@ spec:
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: prometheus-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`prom.workstation.internal`) || Host(`prom.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-prometheus
port: 9090
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alertmanager-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`am.workstation.internal`) || Host(`am.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-alertmanager
port: 9093
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: loki-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`loki.workstation.internal`) || Host(`loki.gigaforust.internal`)
kind: Rule
services:
- name: loki-gateway
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alloy-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`alloy.workstation.internal`) || Host(`alloy.gigaforust.internal`)
kind: Rule
services:
- name: alloy
port: 12345
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: victoria-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`victoria.workstation.internal`) || Host(`victoria.gigaforust.internal`)
kind: Rule
services:
- name: victoria-metrics
port: 8428
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: vmalert-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`vmalert.workstation.internal`) || Host(`vmalert.gigaforust.internal`)
kind: Rule
services:
- name: vmalert
port: 8880
tls:
secretName: internal-wildcard-tls
@@ -1,12 +0,0 @@
nameOverride: victoria-operator
operator:
enable_converter_ownership: true
resources:
requests:
cpu: 50m
memory: 96Mi
limits:
cpu: 200m
memory: 256Mi
-79
View File
@@ -1,79 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: victoria-metrics
namespace: prometheus
spec:
selector:
app: victoria-metrics
ports:
- port: 8428
targetPort: 8428
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: victoria-pvc
namespace: prometheus
spec:
resources:
requests:
storage: 10Gi
volumeMode: Filesystem
accessModes:
- ReadWriteOnce
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: victoria-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: victoria-metrics
strategy:
type: Recreate
template:
metadata:
labels:
app: victoria-metrics
spec:
containers:
- name: victoria
image: victoriametrics/victoria-metrics:v1.153.0-scratch
args:
- -storageDataPath=/vmdata
- -retentionPeriod=30d
- -httpListenAddr=:8428
ports:
- containerPort: 8428
readinessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: vmdata
mountPath: /vmdata
resources:
requests:
cpu: "100m"
memory: "256Mi"
limits:
cpu: "1000m"
memory: "1Gi"
volumes:
- name: vmdata
persistentVolumeClaim:
claimName: victoria-pvc
-29
View File
@@ -1,29 +0,0 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMAgent
metadata:
name: vmagent
namespace: prometheus
spec:
image:
tag: v1.153.0
scrapeInterval: 30s
externalLabels:
scraper: victoria
serviceScrapeNamespaceSelector: {}
serviceScrapeSelector:
matchLabels:
release: prometheus-stack
globalScrapeRelabelConfigs:
- action: drop
source_labels:
- __meta_kubernetes_service_name
regex: prometheus-stack-kube-prom-prometheus
remoteWrite:
- url: http://victoria-metrics.prometheus.svc.cluster.local:8428/api/v1/write
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1000m"
memory: 1Gi
-70
View File
@@ -1,70 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: vmalert
namespace: prometheus
spec:
selector:
app: vmalert
ports:
- port: 8880
targetPort: 8880
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: vmalert-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: vmalert
strategy:
type: Recreate
template:
metadata:
labels:
app: vmalert
spec:
containers:
- name: vmalert
image: victoriametrics/vmalert:v1.153.0
args:
- -datasource.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -remoteWrite.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -notifier.url=http://prometheus-stack-kube-prom-alertmanager.prometheus.svc.cluster.local:9093
- -rule=/etc/vm/rules/*.yaml
- -evaluationInterval=60s
- -httpListenAddr=:8880
ports:
- containerPort: 8880
readinessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: rules
mountPath: /etc/vm/rules
readOnly: true
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "256Mi"
volumes:
- name: rules
configMap:
name: prometheus-prometheus-stack-kube-prom-prometheus-rulefiles-0
+20
View File
@@ -0,0 +1,20 @@
# RackPeek
Rack inventory UI behind Traefik.
Kubernetes stores configuration in `rackpeek-pvc`. The Compose alternative uses
its own data mount. Keep rack descriptions and inventory data in the backup.
Public and internal certificates and routes are in `k8s/`. There are no tracked
Secret examples for this service.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n rackpeek
kubectl get events -n rackpeek --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+56
View File
@@ -0,0 +1,56 @@
# Reloader
Restarts opted-in workloads when the ConfigMaps or Secrets they consume change.
The deploy workflow upgrades the `reloader` Helm release in namespace `reloader`;
`k8s/active` enables it. The chart version is pinned in `deploy-lib.sh`.
## Workload integration
Put this annotation on the Deployment or StatefulSet metadata:
```yaml
metadata:
annotations:
reloader.stakater.com/auto: "true"
```
The annotation belongs to the workload, not `spec.template.metadata`.
Reloader discovers references in environment variables and mounted volumes.
This covers startup-only settings and ConfigMaps or Secrets mounted with `subPath`.
See the [upstream usage guide](https://github.com/stakater/Reloader/blob/v1.4.22/README.md#usage).
The application manifests opt in 28 workloads, including AdGuard's TLS files,
NetBird, both NetBox processes, EDU bots, and the password-protected Valkey servers.
Inactive services have the same annotations ready for later activation.
## Controller policy
The controller watches all namespaces but only restarts annotated workloads.
It uses the `annotations` reload strategy, so changes trigger a pod-template
annotation rather than injecting extra environment variables.
Jobs and CronJobs are excluded: their next execution reads current configuration.
PostgreSQL is intentionally not opted in. Its password variables and init scripts
apply to first initialization; restarting an existing database does not rotate
roles or rerun those scripts. Rotate database credentials with SQL and update the
clients' Secrets together.
Helm-managed monitoring components already have their own configuration reload
paths; Traefik watches its file-provider configuration. They are not globally
opted in. The controller does not react to files in PVCs or changes to external
services unless a watched ConfigMap or Secret changes.
## Verify
```sh
kubectl -n reloader rollout status deployment/reloader-reloader
kubectl -n reloader logs deployment/reloader-reloader --since=10m
kubectl -n netbird get deployment netbird-server-deployment \
-o jsonpath='{.metadata.annotations.reloader\.stakater\.com/auto}'
```
A changed configuration can briefly interrupt a single-replica service, especially
one using `Recreate`. Installing annotations does not validate the configuration
or migrate database data. Keep changes to shared Secrets coordinated across consumers.
See the [repository README](../README.md) for deployment selection.
View File
Whitespace-only changes.
Loaded 100 of 125 files, more files were not shown because too many files have changed in this diff. Show more