Compare commits

..
Author SHA1 Message Date
renovate-bot Bot 6a3eec0164 chore(deps): update all patch updates
ci / checks (pull_request) Successful in 41s
renovate-ci / validate-renovate (pull_request) Successful in 9s
ci / build (pull_request) Skipped
2026-10-07 10:19:32 +00:00
111 changed files with 3071 additions and 3131 deletions

No files matched your search

-50
View File
@@ -1,50 +0,0 @@
# EDU ownership handoff
## Status
The EDU ownership handoff is complete. The homelab repository no longer owns
EDU workloads, images, routes, alerts, or deployment selection. The EDU
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
build, deploy, rollback, verification, route-probe, and registry references.
It also added the serial image build matrix for the homelab services. This
handoff record is the only remaining EDU-specific file in homelab Git.
The dedicated workstation checkout is `/srv/edu-master`, at release
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
moved outside the homelab repository to
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
checkout has no EDU marker or tracked EDU application/deployment files.
`AUTODEPLOY=false` remains in place for homelab deployment.
## Release evidence
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
all validation and both image builds. Deploy run 1653 passed for the exact main
SHA above.
The workstation rollout completed for both Deployments. The deployment
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
matching expressions and healthy evaluation.
The images now run by digest:
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
runtime Secret and Fernet key were preserved during the handoff. Notification
delivery was verified before closeout, as confirmed by the operator. The
deployment did not record downtime.
The release rollback snapshot is
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
The handoff data snapshot remains at
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
Both snapshots are outside Git. Do not restore old Redis data unless recovery
requires it. Never delete or recreate the Redis PVC.
-64
View File
@@ -1,64 +0,0 @@
# CI and deployment
Gitea Actions validates changes, builds the repository's custom images, and can
deploy selected services to the workstation. CI and production deployment use
separate workflows. See the [runner and recovery guide](runner/README.md) for
installation, configuration, and operator commands.
## CI
`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit
test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs
use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool
versions are pinned in `workflows/tool-versions.env`.
Compose CI checks every committed Compose file without requiring ignored `.env`
files. Kubernetes checks validate known schemas; unknown CRDs are skipped.
On main, CI plans builds for the three owned images: `error-pages`,
`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a
digest from a successful earlier main run. The successful build job publishes a
release artifact for the exact commit SHA. Pull requests do not publish images.
## Deployment gate
`workflows/deploy.yaml` starts a deployment after successful main CI when the
`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the
requested `main` ref or commit must have successful main CI and its matching
release artifact. A manual dispatch does not bypass validation.
The workflow supports these modes:
- `changed`: select active services changed since the last successful deploy.
- `full`: select all active services; use this for the first baseline.
- `plan`: validate and show the selection without applying production resources.
`refresh_images=true` explicitly refreshes mutable third-party Compose tags.
## Selection and rollout
The active markers define automatic deployment. `<service>/active` selects a
standard Compose file; `<service>/k8s/active` selects Kubernetes resources. Helm
releases have their own markers in `workflows/deploy-lib.sh`. Service
dependencies are declared in `deploy-dependencies.json`. Removed resources are
reported for manual review; the workflow does not prune them automatically.
The workstation controller runs the checked source in a per-SHA worktree. It
validates configuration, applies Kubernetes and Compose changes in sequence,
verifies changed Kubernetes workloads, and checks public routes. A durable
systemd service continues the rollout if the Actions SSH client disconnects.
The workflow checks the exact CI release before it submits a deployment.
Kubernetes recovery uses captured workload revisions. It does not restore
ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is
manual and does not restore volume data or reverse migrations. Keep backups for
stateful services. The runner guide documents status, retry, logs, and recovery
commands.
## Settings
Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified
`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`,
`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation
also needs its existing registry authentication. Set `AUTODEPLOY=false` until
automatic production deploys are intended.
-1
View File
@@ -7,5 +7,4 @@ self-hosted-runner:
labels: labels:
- arch - arch
- homelab - homelab
- homelab-pr
- prod - prod
+5 -64
View File
@@ -1,20 +1,8 @@
# Homelab CI/CD # Homelab CI/CD
The native Gitea runners run on **vps**; production runs on **workstation**. The native Gitea runner runs on **vps**; production runs on **workstation**.
Main-branch checks and image builds use `homelab:host`. Pull request and Jobs run on `homelab:host`, one at a time. No job images or Kubernetes credentials
non-main checks use `homelab-pr:host` under a separate account without Docker are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
access. The `homelab-pr` runner is registered at User scope for `forust`, so
any repository under that account can schedule jobs that request this label.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with `setup-workstation.sh` when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation ## Runner installation
@@ -40,32 +28,6 @@ pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage. 2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes. Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
### Pull request runner
Install the unprivileged host runner on the VPS:
```sh
sudo bash .gitea/runner/setup-pr-runner.sh
```
Get a registration token from the user Actions runner settings. Run the
installer in a terminal. It asks for the token without echoing it, registers the
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
runner as User scope before merging the workflow change. An unmatched label can
fall back to the default job image.
Renovate PR validation uses `pull_request_target`, which reads the workflow from
the base branch. It checks out the PR head only after runner selection and runs
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
The PR runner has a separate home and tool cache. Do not add it to the `docker`
group or give it access to `/var/run/docker.sock`. It runs repository code from
pull requests, so keep its registration and permissions separate from the
trusted `homelab` runner. This separates users and host permissions, but both
runners still share the VPS kernel and network. Use a disposable VM if PRs from
untrusted external authors must be fully isolated.
## Workstation setup ## Workstation setup
As the existing SSH deploy user on workstation: As the existing SSH deploy user on workstation:
@@ -96,7 +58,8 @@ The deploy user's existing Docker registry authentication remains necessary.
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to successful main CI artifact, never from `:prod`. EDU images remain pinned to the
digests released by their application repository. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun. rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA: Run deploy from main with `deploy_ref=main` or a checked SHA:
@@ -157,25 +120,3 @@ run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git. reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed. state or Compose recovery files until recovery is confirmed.
### Compose configuration recovery
Successful deploys save the complete resolved Compose configuration in
`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets.
Keep them private and do not commit or upload them.
The next deploy uses this configuration for its recovery file, including old
commands, environment, mounts, ports, and removed services. The recovery command
uses `--remove-orphans` to remove services added by the failed deploy. It does
not restore volume data or reverse database migrations.
On the first run after this update, the controller can use the Compose file
from the previous successful run. If that file is absent, it reads the persistent
checkout and checks its service configuration hashes against existing containers.
A mismatch stops preflight. Restore the previous configuration before retrying.
Update the installed controller with `bash .gitea/runner/setup-workstation.sh`
from the reviewed checkout before using this change.
New namespaces are checked during preflight. Server validation of their resources
runs after namespace creation and before application resources are applied.
Plan mode does not create namespaces. A failed deferred check can leave an empty
namespace; inspect it before removing it.
-8
View File
@@ -1,8 +0,0 @@
runner:
file: /var/lib/gitea-pr-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab-pr:host
cache:
enabled: false
-27
View File
@@ -1,27 +0,0 @@
[Unit]
Description=Gitea Actions untrusted pull request runner
After=network-online.target
Wants=network-online.target
[Service]
User=gitea-pr-runner
Group=gitea-pr-runner
WorkingDirectory=/var/lib/gitea-pr-runner
Environment=HOME=/var/lib/gitea-pr-runner
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
Restart=on-failure
RestartSec=5
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=full
ProtectHome=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
UMask=0077
[Install]
WantedBy=multi-user.target
-56
View File
@@ -1,56 +0,0 @@
#!/usr/bin/env bash
# Install a native runner for untrusted PR jobs without Docker access.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in cp cut date getent id install runuser systemctl useradd; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
command -v /usr/local/bin/gitea-runner >/dev/null || {
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
exit 1
}
id gitea-pr-runner >/dev/null 2>&1 || \
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
echo 'Unexpected PR runner home; inspect the existing service first' >&2
exit 1
}
case " $(id -nG gitea-pr-runner) " in
*' docker '*)
echo 'The PR runner account must not belong to the docker group' >&2
exit 1
;;
esac
install -d -m 0755 /etc/gitea-pr-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
printf '\n'
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
unset runner_token
runuser --preserve-environment -u gitea-pr-runner -- \
/usr/local/bin/gitea-runner register \
--config /etc/gitea-pr-runner/config.yaml \
--instance https://gitea.forust.xyz \
--name homelab-pr \
--labels homelab-pr:host \
--no-interactive
unset GITEA_RUNNER_REGISTRATION_TOKEN
fi
chmod 0600 /var/lib/gitea-pr-runner/.runner
systemctl daemon-reload
systemctl enable --now gitea-pr-runner.service
systemctl restart gitea-pr-runner.service
echo "PR runner ready. Configuration backups: *.before-$stamp"
-62
View File
@@ -91,66 +91,4 @@ if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2 echo 'Secret check accepted a failed manifest render' >&2
exit 1 exit 1
fi fi
# New declared namespaces defer only their own resources during preflight.
render_selected_resources() {
cat <<'JSON'
{"apiVersion":"v1","kind":"List","items":[
{"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}}
]}
JSON
}
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources true
jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null
if validate_server_resources false 2>"$scratch/deferred.log"; then
echo 'Post-namespace validation accepted a missing namespace' >&2
exit 1
fi
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources false
jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null
render_selected_resources() {
printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}'
}
if validate_server_resources true 2>"$scratch/undeclared.log"; then
echo 'Preflight accepted an undeclared missing namespace' >&2
exit 1
fi
# Count services, not characters in the newline-separated service names.
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane web ;;
*) return 1 ;;
esac
}
verify_compose_stack example.yaml >"$scratch/compose-count.log"
grep -qF 'all 3 service(s) running' "$scratch/compose-count.log"
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane ;;
*) return 0 ;;
esac
}
if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then
echo 'Compose verification accepted a missing service' >&2
exit 1
fi
grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log"
printf '%s\n' 'Deploy validation regressions passed.' printf '%s\n' 'Deploy validation regressions passed.'
+16 -364
View File
@@ -12,14 +12,18 @@ concurrency:
group: ci-${{ github.ref }} group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }} cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
jobs: jobs:
compose: checks:
name: Compose runs-on: homelab
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} timeout-minutes: 30
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source - name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
echo "$tools_dir" >> "$GITHUB_PATH"
- name: Validate Compose files - name: Validate Compose files
shell: bash shell: bash
run: | run: |
@@ -48,73 +52,11 @@ jobs:
exit 1 exit 1
fi fi
echo "checked ${#files[@]} Compose file(s)" echo "checked ${#files[@]} Compose file(s)"
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Compose
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
&& 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
workflows:
name: Workflows
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Gitea Actions workflows with actionlint - name: Lint Gitea Actions workflows with actionlint
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Workflows
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
shell:
name: Shell
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint shell scripts with ShellCheck - name: Lint shell scripts with ShellCheck
shell: bash shell: bash
run: | run: |
@@ -128,37 +70,6 @@ jobs:
fi fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}" shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
bash .gitea/tests/deploy-validation.sh bash .gitea/tests/deploy-validation.sh
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Shell
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
formatting:
name: Formatting
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Check formatting with Prettier - name: Check formatting with Prettier
shell: bash shell: bash
run: | run: |
@@ -176,37 +87,6 @@ jobs:
fi fi
prettier --check --ignore-unknown "${prettier_files[@]}" prettier --check --ignore-unknown "${prettier_files[@]}"
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Formatting
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
python:
name: Python and tests
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint and format-check Python with Ruff - name: Lint and format-check Python with Ruff
shell: bash shell: bash
run: | run: |
@@ -214,37 +94,6 @@ jobs:
ruff check . .gitea/workflows ruff check . .gitea/workflows
ruff format --check . .gitea/workflows ruff format --check . .gitea/workflows
python3 -m unittest discover -s tests -v python3 -m unittest discover -s tests -v
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Python and tests
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
yaml:
name: YAML
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint YAML syntax - name: Lint YAML syntax
shell: bash shell: bash
run: | run: |
@@ -262,37 +111,6 @@ jobs:
fi fi
yamllint -c .yamllint "${yaml_files[@]}" yamllint -c .yamllint "${yaml_files[@]}"
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: YAML
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
dockerfiles:
name: Dockerfiles
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Dockerfiles - name: Lint Dockerfiles
shell: bash shell: bash
run: | run: |
@@ -308,37 +126,6 @@ jobs:
fi fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}" hadolint -c .hadolint.yaml "${dockerfiles[@]}"
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Dockerfiles
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
kubernetes:
name: Kubernetes
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Validate Kubernetes manifests against JSON schemas - name: Validate Kubernetes manifests against JSON schemas
shell: bash shell: bash
run: | run: |
@@ -359,162 +146,27 @@ jobs:
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
"${manifests[@]}" "${manifests[@]}"
id: check build:
- name: Write the job result needs:
if: always() - checks
env:
SUMMARY_CHECK: Kubernetes
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
image-plan:
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main' if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
runs-on: homelab runs-on: homelab
timeout-minutes: 10 timeout-minutes: 60
outputs:
matrix: ${{ steps.plan.outputs.matrix }}
steps: steps:
- name: Checkout repository - name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with: with:
fetch-depth: 0 fetch-depth: 0
- name: Detect build inputs against successful CI - name: Build changed images and write release
id: plan
env: env:
GITEA_TOKEN: ${{ github.token }} GITEA_TOKEN: ${{ github.token }}
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
- name: Store the image plan
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: build-plan
path: build-plan.json
if-no-files-found: error
retention-days: 30
- name: Write the plan result
if: always()
env:
SUMMARY_CHECK: Image plan
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
fi
images:
name: Image (${{ matrix.name }})
needs: [image-plan]
if: needs.image-plan.result == 'success'
runs-on: homelab
timeout-minutes: 60
strategy:
max-parallel: 1
fail-fast: false
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download the checked image plan
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
name: build-plan
- name: Build or reuse this image
id: check
env:
IMAGE_NAME: ${{ matrix.name }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json run: python3 .gitea/workflows/release.py build
- name: Store the image result
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: image-${{ matrix.name }}
path: image.json
if-no-files-found: error
retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image (${{ matrix.name }})
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
# Retain the build job name required by the immutable release deployment gate.
build:
needs: [image-plan, images]
runs-on: homelab
timeout-minutes: 15
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download all image results
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
path: artifacts
- name: Pin SHA tags and write the complete release
id: check
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: >-
python3 .gitea/workflows/release.py finalize
--plan artifacts/build-plan/build-plan.json
- name: Store commit release - name: Store commit release
id: artifact uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with: with:
name: release-${{ github.sha }} name: release-${{ github.sha }}
path: release.json path: release.json
if-no-files-found: error if-no-files-found: error
retention-days: 30 retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image release and SHA tags
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
+4 -75
View File
@@ -42,50 +42,15 @@ def prepare(source_file):
images_file = directory / 'compose-images.json' images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {}) locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text()) release = json.loads((directory / 'release.json').read_text())
state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy')) before = json.loads(json.dumps(config))
baseline = state / 'compose-configs' / f'{relative.parent.name}.json'
if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')):
baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name
bootstrap = not baseline.exists()
if not bootstrap:
before = json.loads(baseline.read_text())
else:
# Bootstrap from the persistent configuration, never from the new source.
persistent_file = config_repo / relative
if persistent_file.exists():
before = json.loads(
output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--format',
'json',
cwd=config_repo,
)
)
elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'):
raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy')
else:
before = {'name': project, 'services': {}}
if before['name'] != project:
raise ValueError('Compose project name changed; manual migration is required')
for service, settings in config['services'].items(): for service, settings in config['services'].items():
reference = settings.get('image') reference = settings.get('image')
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
if not reference or settings.get('build'): if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image') raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1) image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0] image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo) image_repo = '/'.join(image_repo)
# Nextcloud AIO validates the mastercontainer image reference and rejects if image_repo in release['images']:
# a digest. Keep its configured tag so AIO can start and manage its stack.
if nextcloud_aio_master:
pinned = reference
elif image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo] pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks: elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference] pinned = locks[reference]
@@ -93,12 +58,6 @@ def prepare(source_file):
pinned = resolve(reference) pinned = resolve(reference)
settings['image'] = pinned settings['image'] = pinned
locks[reference] = pinned locks[reference] = pinned
for service, settings in before['services'].items():
reference = settings['image']
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
# Capture what is running, not the current value of its mutable tag. # Capture what is running, not the current value of its mutable tag.
ids = output( ids = output(
'docker', 'docker',
@@ -110,43 +69,13 @@ def prepare(source_file):
f'label=com.docker.compose.service={service}', f'label=com.docker.compose.service={service}',
).splitlines() ).splitlines()
actual = set() actual = set()
if bootstrap and ids:
expected_hash = output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--hash',
service,
cwd=config_repo,
).split()[-1]
for container in ids:
running_hash = output(
'docker',
'inspect',
container,
'--format',
'{{ index .Config.Labels "com.docker.compose.config-hash" }}',
)
if running_hash != expected_hash:
raise ValueError(
f'{project}/{service}: persistent config differs from running config; restore the previous config'
)
for container in ids: for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}') image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}')) digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id)) actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1: if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config') raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
# AIO also rejects a digest in its recovery config. Preserve its tag in before['services'][service]['image'] = next(iter(actual)) if actual else reference
# both deploy and recovery files.
if nextcloud_aio_master:
before['services'][service]['image'] = reference
else:
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)): for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True) folder.mkdir(mode=0o700, exist_ok=True)
@@ -156,7 +85,7 @@ def prepare(source_file):
images_file.write_text(json.dumps(locks, indent=2) + '\n') images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved') print(f'Compose {project}: images pinned; local paths preserved')
print( print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans' f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
) )
+2 -106
View File
@@ -141,8 +141,7 @@ def make_plan(directory):
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py') planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text()) request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
# Helm 4 lists every release status by default and removed the --all flag. helm = json.loads(command('helm', 'list', '--all', '-A', '-o', 'json'))
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm) plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']: if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose'] plan['selected']['compose'] = plan['active']['compose']
@@ -165,10 +164,6 @@ def finish_success(directory, plan):
if previous.exists() if previous.exists()
else {} else {}
) )
configs = STATE / 'compose-configs'
configs.mkdir(mode=0o700, exist_ok=True)
for config in (directory / 'compose').glob('*.json'):
atomic_json(configs / config.name, json.loads(config.read_text()))
atomic_json(STATE / 'last-success.json', plan) atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text()) status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success' status['state'] = 'success'
@@ -217,17 +212,14 @@ def execute(run_id):
raise ValueError(f'Interrupted deploy {other.name}; run recover first') raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running' status['state'] = 'running'
atomic_json(directory / 'status.json', status) atomic_json(directory / 'status.json', status)
phase = 'plan'
try: try:
plan = make_plan(directory) plan = make_plan(directory)
print( print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}), json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True, flush=True,
) )
phase = 'doctor'
if not stage(directory, 'doctor', 600): if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed') raise RuntimeError('Preflight failed')
phase = 'validate'
if not stage(directory, 'validate', 1200): if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed') raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan': if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
@@ -236,7 +228,6 @@ def execute(run_id):
atomic_json(directory / 'status.json', status) atomic_json(directory / 'status.json', status)
return return
# Budget includes both rollout checks and rollback waves, plus API overhead. # Budget includes both rollout checks and rollback waves, plus API overhead.
phase = 'Recovery budget'
count = int( count = int(
command( command(
'bash', 'bash',
@@ -248,24 +239,14 @@ def execute(run_id):
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120) verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200: if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy') raise ValueError('More than two hours of recovery required; split this deploy')
phase = 'apply-k8s'
k8s_ok = stage(directory, 'apply-k8s', 2700) k8s_ok = stage(directory, 'apply-k8s', 2700)
phase = 'apply-compose'
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
phase = 'verify-k8s'
verify_ok = stage(directory, 'verify-k8s', verify_budget) verify_ok = stage(directory, 'verify-k8s', verify_budget)
phase = 'smoke'
smoke_ok = stage(directory, 'smoke', 600) smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)): if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report') raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
phase = 'Save the successful baseline'
finish_success(directory, plan) finish_success(directory, plan)
except Exception as error: except Exception as error:
status = json.loads((directory / 'status.json').read_text())
status['failure_stage'] = next(
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
)
atomic_json(directory / 'status.json', status)
with (directory / 'controller.log').open('a') as stream: with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n') stream.write(f'{error}\n')
recover(directory) recover(directory)
@@ -313,93 +294,10 @@ def follow(run_id, phase):
time.sleep(3) time.sleep(3)
def summary(run_id):
directory = run_directory(run_id)
request = json.loads((directory / 'request.json').read_text())
release = request['release']
plan_file = directory / 'plan.json'
lines = [
f'## Deploy `{release["sha"]}`',
'',
f'- Mode: `{request["mode"]}`',
f'- Refresh third-party images: `{request["refresh_images"]}`',
]
status = json.loads((directory / 'status.json').read_text())
if status.get('failure_stage'):
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
lines.extend(
[
'',
f'- Observed run state: **{status["state"]}**',
'',
'### Stage results',
'| Stage | Result | Exit code |',
'| --- | --- | --- |',
]
)
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
stage_result = status['stages'].get(name, {})
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
lines.extend(['', '### Apply and Helm recovery results'])
events_file = directory / 'apply-events.jsonl'
events = []
if events_file.exists():
for line in events_file.read_text().splitlines():
try:
events.append(json.loads(line))
except json.JSONDecodeError:
lines.append('- An operation record is incomplete. Check the stage log.')
latest = {(event['action'], event['target']): event['result'] for event in events}
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
if not latest:
lines.append('- No apply results were recorded.')
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
lines.extend(['', '### Kubernetes recovery'])
pointer = directory / 'snapshot/current'
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
if failed and failed.exists():
contents = failed.read_text()
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
if not contents.strip():
lines.append('- No failed workloads were recorded. See the verification result above.')
elif counts:
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
else:
lines.append('- Rollback has no recorded result yet. Check the verification log.')
else:
lines.append('- No workload rollback was recorded. This does not confirm health.')
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
if not plan_file.exists():
lines.extend(['', 'Plan was not created. Check the controller log.'])
print('\n'.join(lines))
return
plan = json.loads(plan_file.read_text())
lines.extend(['', '### Selected services'])
count = 0
for kind, services in plan['selected'].items():
for service in services:
lines.append(f'- `{kind}`: `{service}`')
count += 1
if not count:
lines.append('- None')
lines.extend(['', '### Selected Helm releases'])
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
if not plan.get('helm'):
lines.append('- None')
lines.extend(['', '### Images pinned in the checked release'])
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
lines.extend(['', '### Removed resources requiring manual review'])
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
if not plan.get('removed'):
lines.append('- None')
print('\n'.join(lines))
def main(): def main():
os.umask(0o077) os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__) parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary')) parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow'))
parser.add_argument('run_id') parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke')) parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply') parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
@@ -417,8 +315,6 @@ def main():
if (directory / 'plan.json').exists(): if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text()) plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2)) print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif args.action == 'summary':
summary(args.run_id)
elif not follow(args.run_id, args.phase): elif not follow(args.run_id, args.phase):
sys.exit(1) sys.exit(1)
+16 -86
View File
@@ -27,15 +27,6 @@ log() {
echo "== $* ==" echo "== $* =="
} }
# Store operation results without command output or local configuration values.
record_apply() {
[ -n "${RUN_DIR:-}" ] || return 0
jq -cn --arg action "$1" --arg target "$2" --arg result "$3" \
'{action: $action, target: $target, result: $result}' >>"$RUN_DIR/apply-events.jsonl" \
|| echo 'WARNING: cannot record an apply result' >&2
return 0
}
warn() { warn() {
echo "WARNING: $*" >&2 echo "WARNING: $*" >&2
} }
@@ -180,8 +171,7 @@ save_snapshot() {
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid)) | select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end) | select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1 }]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
# Helm 4 lists every release status by default and removed the --all flag. releases="$(helm list --all -A -o json)" || return 1
releases="$(helm list -A -o json)" || return 1
for entry in "${HELM_RELEASES[@]}"; do for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release _ namespace _ _ _ <<<"$entry" IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
if ! jq -e --arg r "$release" --arg n "$namespace" \ if ! jq -e --arg r "$release" --arg n "$namespace" \
@@ -330,11 +320,11 @@ rollback_workloads() {
# written straight into a `helm upgrade` command would never be updated: these # written straight into a `helm upgrade` command would never be updated: these
# have to be declared as custom.regex managers in renovate/renovate.json. # have to be declared as custom.regex managers in renovate/renovate.json.
HELM_RELEASES=( HELM_RELEASES=(
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.3.2|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active" "prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active" "victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active" "loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active" "alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
"reloader|stakater/reloader|reloader|2.2.18|reloader/k8s/reloader-values.yaml|reloader/k8s/active" "reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
) )
# "name url" for the Helm repository hosting a chart, empty if unknown. # "name url" for the Helm repository hosting a chart, empty if unknown.
@@ -386,19 +376,15 @@ recover_pending_release() {
echo "ERROR: no captured Helm revision for $release; manual recovery required" echo "ERROR: no captured Helm revision for $release; manual recovery required"
return 1 return 1
fi fi
record_apply helm-rollback "$namespace/$release" started
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
record_apply helm-rollback "$namespace/$release" failure
echo "WARN: helm rollback of $release did not complete" echo "WARN: helm rollback of $release did not complete"
return 1 return 1
fi fi
status="$(helm_release_status "$release" "$namespace")" || return 1 status="$(helm_release_status "$release" "$namespace")" || return 1
if [ "$status" != "deployed" ]; then if [ "$status" != "deployed" ]; then
record_apply helm-rollback "$namespace/$release" failure
echo "WARN: $release is $status after rollback" echo "WARN: $release is $status after rollback"
return 1 return 1
fi fi
record_apply helm-rollback "$namespace/$release" success
;; ;;
esac esac
return 0 return 0
@@ -462,13 +448,11 @@ upgrade_helm_releases() {
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade # --rollback-on-failure (+ --wait) rolls the release back when the upgrade
# times out or the workloads it touches never become ready, so a bad chart # times out or the workloads it touches never become ready, so a bad chart
# bump is not left half applied. (--atomic was this combo; deprecated.) # bump is not left half applied. (--atomic was this combo; deprecated.)
record_apply helm-upgrade "$namespace/$release" started
if ! helm upgrade --install "$release" "$chart" \ if ! helm upgrade --install "$release" "$chart" \
--namespace "$namespace" \ --namespace "$namespace" \
--version "$version" \ --version "$version" \
--values "$values" \ --values "$values" \
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then --wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
record_apply helm-upgrade "$namespace/$release" failure
echo "WARN: upgrade of $release failed, checking release state" echo "WARN: upgrade of $release failed, checking release state"
# --rollback-on-failure already attempted its own rollback; finish the job when that # --rollback-on-failure already attempted its own rollback; finish the job when that
# rollback never completed, otherwise the release stays pending-* and # rollback never completed, otherwise the release stays pending-* and
@@ -478,10 +462,8 @@ upgrade_helm_releases() {
else else
echo "ERROR: upgrade of $release failed (release is back on its previous revision)." echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
fi fi
record_apply helm-recovery-state "$namespace/$release" "$(helm_release_status "$release" "$namespace" || echo unknown)"
return 1 return 1
fi fi
record_apply helm-upgrade "$namespace/$release" success
done done
} }
@@ -571,41 +553,6 @@ skip_uninstalled_vmagent_crd() {
return 1 return 1
} }
# Render one complete resource list so new namespaces can be identified across
# files and Kustomize apps. A missing undeclared namespace remains an error.
render_selected_resources() {
local m k
{
for m in "${K8S_MANIFESTS[@]}"; do
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
kubectl create --dry-run=client --validate=false -f "$m" -o json || return 1
done
for k in "${KUSTOMIZE_APPS[@]}"; do
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json || return 1
done
} | jq -s '{apiVersion: "v1", kind: "List", items: [ .[] | if .kind == "List" then .items[] else . end ]}'
}
validate_server_resources() {
local defer_new="$1" resources existing filtered
resources="$(render_selected_resources)" || return 1
existing="$(kubectl get namespaces -o json)" || return 1
filtered="$(jq --argjson existing "$existing" --argjson defer "$defer_new" '
[.items[] | select(.kind == "Namespace") | .metadata.name] as $declared
| [$existing.items[].metadata.name] as $present
| .items |= map(
(.metadata.namespace // "default") as $ns
| if .kind == "Namespace" or ($present | index($ns)) != null then .
elif ($declared | index($ns)) == null then error("Undeclared missing namespace: " + $ns)
elif $defer then empty
else error("Namespace still missing after namespace apply: " + $ns)
end)
' <<<"$resources")" || return 1
if [ "$(jq '.items | length' <<<"$filtered")" -gt 0 ]; then
kubectl apply --dry-run=server -f - <<<"$filtered" >/dev/null
fi
}
stage_validate() { stage_validate() {
check_prune_mode || return 1 check_prune_mode || return 1
cd "$REPO" cd "$REPO"
@@ -632,7 +579,15 @@ stage_validate() {
kubectl apply -k "$k" --dry-run=client >/dev/null kubectl apply -k "$k" --dry-run=client >/dev/null
done done
log "Validate k8s manifests (kubectl dry-run=server)" log "Validate k8s manifests (kubectl dry-run=server)"
validate_server_resources true for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=server -f "$m" >/dev/null
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
kubectl apply -k "$k" --dry-run=server >/dev/null
done
log "Checking referenced Secrets exist" log "Checking referenced Secrets exist"
echo " (deploy never applies *secret*.yaml; create missing ones manually)" echo " (deploy never applies *secret*.yaml; create missing ones manually)"
check_referenced_secrets check_referenced_secrets
@@ -676,22 +631,9 @@ stage_apply_k8s() {
if [ "${#ns_files[@]}" -gt 0 ]; then if [ "${#ns_files[@]}" -gt 0 ]; then
log "Applying namespaces (${#ns_files[@]} files)" log "Applying namespaces (${#ns_files[@]} files)"
for m in "${ns_files[@]}"; do for m in "${ns_files[@]}"; do
record_apply kubectl "${m#"$REPO"/}" started kubectl apply -f "$m"
if ! kubectl apply -f "$m"; then
record_apply kubectl "${m#"$REPO"/}" failure
return 1
fi
record_apply kubectl "${m#"$REPO"/}" success
done done
fi fi
# Kustomize may declare namespaces inside its rendered resources too.
local namespace_resources
namespace_resources="$(render_selected_resources | jq '.items |= map(select(.kind == "Namespace"))')" || return 1
if [ "$(jq '.items | length' <<<"$namespace_resources")" -gt 0 ]; then
kubectl apply -f - <<<"$namespace_resources" || return 1
fi
# Complete the deferred server checks before Helm or application resources change.
validate_server_resources false || return 1
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first." echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
@@ -704,24 +646,18 @@ stage_apply_k8s() {
log "Applying resources (${#other_files[@]} files, our images pinned to digests)" log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
for m in "${other_files[@]}"; do for m in "${other_files[@]}"; do
log "Applying ${m#"$REPO"/}" log "Applying ${m#"$REPO"/}"
record_apply kubectl "${m#"$REPO"/}" started
if ! render_pinned <"$m" | kubectl apply -f -; then if ! render_pinned <"$m" | kubectl apply -f -; then
record_apply kubectl "${m#"$REPO"/}" failure
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2 echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
exit 1 exit 1
fi fi
record_apply kubectl "${m#"$REPO"/}" success
done done
fi fi
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)" log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
record_apply kustomize "${k#"$REPO"/}" started
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
record_apply kustomize "${k#"$REPO"/}" failure
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2 echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
exit 1 exit 1
fi fi
record_apply kustomize "${k#"$REPO"/}" success
done done
# No verification here on purpose. This stage may be killed at any point by # No verification here on purpose. This stage may be killed at any point by
@@ -827,13 +763,12 @@ stage_verify_k8s() {
# actually be running. # actually be running.
verify_compose_stack() { verify_compose_stack() {
local cf="$1" local cf="$1"
local expected running svc missing=() service_count=0 local expected running missing=()
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1 expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
running="$(compose "$cf" ps --status running --services | sort)" || return 1 running="$(compose "$cf" ps --status running --services | sort)" || return 1
[ -n "$expected" ] || return 0 [ -n "$expected" ] || return 0
while IFS= read -r svc; do while IFS= read -r svc; do
[ -n "$svc" ] || continue [ -n "$svc" ] || continue
service_count=$((service_count + 1))
# restart:"no" services are allowed to have exited. # restart:"no" services are allowed to have exited.
if ! printf '%s\n' "$running" | grep -qx "$svc"; then if ! printf '%s\n' "$running" | grep -qx "$svc"; then
missing+=("$svc") missing+=("$svc")
@@ -844,7 +779,7 @@ verify_compose_stack() {
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
return 1 return 1
fi fi
echo " all $service_count service(s) running" echo " all ${#expected} service(s) running"
return 0 return 0
} }
@@ -1024,12 +959,7 @@ stage_apply_compose() {
local cf local cf
for cf in "${COMPOSE_STACKS[@]}"; do for cf in "${COMPOSE_STACKS[@]}"; do
log "Applying Compose ${cf#"$REPO"/}" log "Applying Compose ${cf#"$REPO"/}"
record_apply compose "${cf#"$REPO"/}" started compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans
if ! compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans; then
record_apply compose "${cf#"$REPO"/}" failure
return 1
fi
record_apply compose "${cf#"$REPO"/}" success
verify_compose_stack "$cf" verify_compose_stack "$cf"
done done
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)" echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
+4 -4
View File
@@ -83,20 +83,20 @@ def make_plan(repo, config_repo, release, previous, mode, live_helm):
removed = [] removed = []
else: else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines() paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {service for service in all_services for path in paths if path.startswith(service + '/')} changed = {path.split('/')[0] for path in paths}
if any(path.startswith('.gitea/') for path in paths): if any(path.startswith('.gitea/') for path in paths):
changed |= all_services changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]} changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo): for file in tracked(repo):
owners = {service for service in all_services if file.startswith(service + '/')} service = file.split('/')[0]
if not owners or not file.endswith(('.yaml', '.yml')): if service not in all_services or not file.endswith(('.yaml', '.yml')):
continue continue
text = (repo / file).read_text() text = (repo / file).read_text()
if any( if any(
image in text and previous.get('images', {}).get(image) != digest image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items() for image, digest in release['images'].items()
): ):
changed |= owners changed.add(service)
removed = sorted( removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', [])) set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services - all_services
+1 -37
View File
@@ -64,23 +64,11 @@ jobs:
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA" run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
- name: Submit durable deploy to workstation - name: Submit durable deploy to workstation
run: bash .gitea/workflows/ssh-run.sh start run: bash .gitea/workflows/ssh-run.sh start
- name: Write the request result
if: always()
env:
REQUEST_RESULT: ${{ job.status }}
CHECKED_SHA: ${{ steps.release.outputs.sha }}
run: |
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
if [ "$REQUEST_RESULT" != success ]; then
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
fi
fi
apply: apply:
needs: [gate] needs: [gate]
runs-on: homelab runs-on: homelab
timeout-minutes: 120 timeout-minutes: 100
steps: steps:
- name: Checkout checked commit - name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -88,14 +76,6 @@ jobs:
ref: ${{ needs.gate.outputs.sha }} ref: ${{ needs.gate.outputs.sha }}
- name: Follow validation and sequential Kubernetes / Compose apply - name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply run: bash .gitea/workflows/ssh-run.sh apply
- name: Write the deploy result
if: always()
run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
verify: verify:
needs: [gate, apply] needs: [gate, apply]
@@ -109,14 +89,6 @@ jobs:
ref: ${{ needs.gate.outputs.sha }} ref: ${{ needs.gate.outputs.sha }}
- name: Follow workload verification and recovery - name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify run: bash .gitea/workflows/ssh-run.sh verify
- name: Write the deploy result
if: always()
run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
smoke: smoke:
needs: [gate, verify] needs: [gate, verify]
@@ -130,11 +102,3 @@ jobs:
ref: ${{ needs.gate.outputs.sha }} ref: ${{ needs.gate.outputs.sha }}
- name: Follow public route checks - name: Follow public route checks
run: bash .gitea/workflows/ssh-run.sh smoke run: bash .gitea/workflows/ssh-run.sh smoke
- name: Write the deploy result
if: always()
run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
+63 -232
View File
@@ -26,6 +26,9 @@ IMAGES = {
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'), 'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
} }
# These images are released by the EDU application repository.
EXTERNAL_IMAGES = {'gcr.forust.xyz/forust/session-keeper', 'gcr.forust.xyz/forust/webinar-checker'}
def command(*args, **kwargs): def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell.""" """Arguments are passed directly to the executable, never to a shell."""
@@ -160,7 +163,7 @@ def gate(output, requested_ref, event_sha):
print(f'CI gate accepted {sha}') print(f'CI gate accepted {sha}')
def prepare_images(output): def build(output):
sha = command('git', 'rev-parse', 'HEAD') sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha): if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA') raise ValueError('Build checkout does not match GITHUB_SHA')
@@ -175,61 +178,11 @@ def prepare_images(output):
except ValueError: except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback. # Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue continue
targets = []
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
targets.append(
{
'name': name,
'image': image,
'context': context,
'dockerfile': dockerfile,
'inputs': inputs,
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
}
)
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
def checked_plan(path):
data = json.loads(path.read_text())
sha = command('git', 'rev-parse', 'HEAD')
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Image plan does not match the checked source commit')
targets = data.get('targets', [])
if sorted(t['name'] for t in targets) != sorted(IMAGES):
raise ValueError('Image plan must contain each owned image once')
for target in targets:
name = target['name']
context, dockerfile = IMAGES[name]
if (target['context'], target['dockerfile'], target['image']) != (
context,
dockerfile,
f'gcr.forust.xyz/forust/{name}',
) or target['inputs'] != fingerprint(context, dockerfile):
raise ValueError('Image plan has invalid build inputs')
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
raise ValueError('Image plan has an invalid reuse digest')
return data
def build_images(output, report, name, plan):
data = checked_plan(plan)
sha = data['sha']
target = next(t for t in data['targets'] if t['name'] == name)
context, dockerfile = IMAGES[name]
docker_config = tempfile.mkdtemp(prefix='homelab-registry-') docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx' builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True) builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)} env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try: try:
report['phase'] = 'Registry login'
subprocess.run( # noqa: S603, S607 subprocess.run( # noqa: S603, S607
[ [
shutil.which('docker') or '/usr/bin/docker', shutil.which('docker') or '/usr/bin/docker',
@@ -244,7 +197,6 @@ def build_images(output, report, name, plan):
check=True, check=True,
env=env, env=env,
) )
report['phase'] = 'Prepare the builder'
builder = 'homelab-ci' builder = 'homelab-ci'
versions = dict( versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE) re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
@@ -279,66 +231,61 @@ def build_images(output, report, name, plan):
) )
signature.write_text(image + '\n') signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}} release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
report['images'] = release['images'] for name, (context, dockerfile) in IMAGES.items():
report['phase'] = f'Build or reuse {name}' image = f'gcr.forust.xyz/forust/{name}'
report['current'] = name inputs = fingerprint(context, dockerfile)
image = f'gcr.forust.xyz/forust/{name}' old_digest = (previous or {}).get('images', {}).get(image)
inputs = target['inputs'] exists = False
old_digest = target['reuse_digest'] if old_digest and previous['inputs'].get(image) == inputs:
exists = False exists = (
if old_digest: subprocess.run( # noqa: S603, S607
exists = ( [
subprocess.run( # noqa: S603, S607 shutil.which('docker') or '/usr/bin/docker',
[ 'buildx',
shutil.which('docker') or '/usr/bin/docker', 'imagetools',
'buildx', 'inspect',
'imagetools', f'{image}@{old_digest}',
'inspect', ],
f'{image}@{old_digest}', capture_output=True,
], env=env,
capture_output=True, timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--push',
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--tag',
f'{image}:sha-{sha}',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env, env=env,
timeout=60, )
).returncode digest = json.loads(metadata.read_text())['containerimage.digest']
== 0 release['images'][image] = digest
) release['inputs'][image] = inputs
if exists: validate_release(release, sha)
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--output',
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
if not isinstance(digest, str) or not DIGEST.fullmatch(digest):
raise ValueError('Image job returned an invalid digest')
release['images'][image] = digest
release['inputs'][image] = inputs
report['reused' if exists else 'built'].append(name)
output.write_text(json.dumps(release, indent=2) + '\n') output.write_text(json.dumps(release, indent=2) + '\n')
report['current'] = None
report['phase'] = 'Image result file saved'
finally: finally:
# Cleanup errors must neither leak credentials nor mask the original build error. # Cleanup errors must neither leak credentials nor mask the original build error.
try: try:
@@ -362,63 +309,6 @@ def build_images(output, report, name, plan):
shutil.rmtree(docker_config) shutil.rmtree(docker_config)
def write_summary(lines):
path = os.environ.get('GITHUB_STEP_SUMMARY')
if path:
try:
with Path(path).open('a') as stream:
stream.write('\n'.join(lines) + '\n\n')
except OSError:
print('WARNING: cannot write the job summary')
def check_summary():
lines = [
f'## {os.environ["SUMMARY_CHECK"]}',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
]
if os.environ.get('SUMMARY_FAILED_STEP'):
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
if os.environ['SUMMARY_RESULT'] != 'success':
lines.append('- Open the failed step log for the error details.')
write_summary(lines)
def build(output, name, plan):
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
result = 'failure'
try:
build_images(output, report, name, plan)
result = 'success'
finally:
lines = [
f'## Image build result `{name}`',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
'',
f'- Result: **{result}**',
f'- Last stage: {report["phase"]}',
]
if result == 'failure':
lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.')
if result == 'success':
lines.append('- This is one image result. The final build job must publish the complete release.')
if report['current']:
lines.append(f'- Image at the failure: `{report["current"]}`')
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
lines.extend(['', f'### {title}'])
lines.extend(f'- `{name}`' for name in report[key])
if not report[key]:
lines.append('- None')
lines.extend(['', '### Completed image digests'])
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
if not report['images']:
lines.append('- None')
write_summary(lines)
def render(stream, destination): def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA']) release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile( image_line = re.compile(
@@ -429,6 +319,9 @@ def render(stream, destination):
match = image_line.fullmatch(line.rstrip('\n')) match = image_line.fullmatch(line.rstrip('\n'))
if match: if match:
prefix, quote, image, tail = match.groups() prefix, quote, image, tail = match.groups()
if image in EXTERNAL_IMAGES and f'{image}@sha256:' in line:
rendered.append(line)
continue
if image not in release['images']: if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}') raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n' line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
@@ -438,81 +331,19 @@ def render(stream, destination):
destination.writelines(rendered) destination.writelines(rendered)
def finalize_images(output, fragments, plan):
data = checked_plan(plan)
sha = data['sha']
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name in IMAGES:
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
image = f'gcr.forust.xyz/forust/{name}'
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
raise ValueError('Image job artifact is missing or belongs to another commit')
target = next(t for t in data['targets'] if t['name'] == name)
if fragment.get('inputs') != {image: target['inputs']}:
raise ValueError('Image artifact does not match the build plan')
release['images'].update(fragment['images'])
release['inputs'].update(fragment['inputs'])
validate_release(release, sha)
# Only a complete set of successful image jobs can publish the release tags.
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
for image, digest in release['images'].items():
command(
'docker',
'buildx',
'imagetools',
'create',
'--prefer-index=false',
'--tag',
f'{image}:sha-{sha}',
f'{image}@{digest}',
env=env,
timeout=90,
)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
shutil.rmtree(docker_config)
def main(): def main():
parser = argparse.ArgumentParser(description=__doc__) parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary')) parser.add_argument('action', choices=('build', 'gate', 'render'))
parser.add_argument('--output', type=Path, default=Path('release.json')) parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main') parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='') parser.add_argument('--event-sha', default='')
parser.add_argument('--image', choices=IMAGES)
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
args = parser.parse_args() args = parser.parse_args()
if args.action == 'check-summary': if args.action == 'render':
check_summary()
elif args.action == 'render':
render(sys.stdin, sys.stdout) render(sys.stdin, sys.stdout)
elif args.action == 'gate': elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha) gate(args.output, args.ref, args.event_sha)
elif args.action == 'prepare':
prepare_images(args.output)
elif args.action == 'image':
if not args.image:
parser.error('--image is required')
build(args.output, args.image, args.plan)
else: else:
finalize_images(args.output, args.fragments, args.plan) build(args.output)
if __name__ == '__main__': if __name__ == '__main__':
+16 -26
View File
@@ -1,9 +1,7 @@
name: renovate-ci name: renovate-ci
on: on:
# Read the workflow from the trusted base branch. PR code runs only on the pull_request:
# unprivileged runner selected below.
pull_request_target:
paths: paths:
- "renovate/**" - "renovate/**"
- ".gitea/workflows/renovate-ci.yaml" - ".gitea/workflows/renovate-ci.yaml"
@@ -28,47 +26,37 @@ permissions:
jobs: jobs:
validate-renovate: validate-renovate:
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} runs-on: homelab
timeout-minutes: 20 timeout-minutes: 20
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
# renovate/k8s/cronjob.yaml is the single source of truth for the version. # renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
- name: Resolve the deployed Renovate version # so the same version that runs in the cluster is the one validated here.
- name: Resolve the deployed Renovate image
id: image id: image
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)" renovate/k8s/cronjob.yaml | head -1)"
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then if [ -z "$image" ]; then
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml" echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1 exit 1
fi fi
version="${BASH_REMATCH[1]}" echo "using $image"
echo "using Renovate $version" echo "image=$image" >> "$GITHUB_OUTPUT"
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
- name: Prepare pinned validation tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
echo "$tools_dir" >> "$GITHUB_PATH"
- name: Validate Renovate repository config - name: Validate Renovate repository config
shell: bash shell: bash
env:
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
run: | run: |
set -euo pipefail set -euo pipefail
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")" docker run --rm \
trap 'rm -rf "$npm_cache"' EXIT -v "$PWD/renovate:/opt/renovate:ro" \
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator "${{ steps.image.outputs.image }}" \
renovate-config-validator /opt/renovate/renovate.json
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml # The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches. # carries an inlined copy of the config. Fail if it no longer matches.
@@ -82,6 +70,8 @@ jobs:
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \ kubeconform \
-strict \ -strict \
-ignore-missing-schemas \ -ignore-missing-schemas \
+5 -11
View File
@@ -32,14 +32,11 @@ concurrency:
jobs: jobs:
run-renovate: run-renovate:
if: github.ref == 'refs/heads/main'
runs-on: homelab runs-on: homelab
timeout-minutes: 60 timeout-minutes: 60
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: refs/heads/main
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag. # renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version # Reading it here means this workflow validates and runs the exact version
@@ -51,23 +48,21 @@ jobs:
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)" renovate/k8s/cronjob.yaml | head -1)"
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then if [ -z "$image" ]; then
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml" echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1 exit 1
fi fi
echo "using $image" echo "using $image"
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT" echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config - name: Validate Renovate config
shell: bash shell: bash
env:
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
docker run --rm \ docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \ -v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"$RENOVATE_IMAGE" \ "${{ steps.image.outputs.image }}" \
renovate-config-validator renovate-config-validator
- name: Run Renovate - name: Run Renovate
@@ -78,7 +73,6 @@ jobs:
RENOVATE_REPOSITORIES: ${{ inputs.repositories }} RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }} RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
LOG_LEVEL: ${{ inputs.log_level }} LOG_LEVEL: ${{ inputs.log_level }}
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
@@ -95,4 +89,4 @@ jobs:
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_BASE_DIR=/tmp/renovate \ -e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \ -e LOG_LEVEL="${LOG_LEVEL:-info}" \
"$RENOVATE_IMAGE" "${{ steps.image.outputs.image }}"
+4 -18
View File
@@ -20,7 +20,7 @@ ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHo
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15 -o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4) -o ServerAliveInterval=15 -o ServerAliveCountMax=4)
controller=.local/lib/homelab-deploy/controller.py controller=.local/lib/homelab-deploy/controller.py
case "${1:?start, apply, verify, smoke or summary required}" in case "${1:?start, apply, verify or smoke required}" in
start) start)
python3 - <<'PY' >"$key_dir/request.json" python3 - <<'PY' >"$key_dir/request.json"
import json import json
@@ -41,30 +41,16 @@ PY
exit "$rc" exit "$rc"
;; ;;
apply|verify|smoke) apply|verify|smoke)
result=0
for attempt in 1 2 3; do for attempt in 1 2 3; do
rc=0 rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables. # shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$? ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && break [ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || { result="$rc"; break; } [ "$rc" -eq 255 ] || exit "$rc"
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)" echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
if [ "$attempt" -eq 3 ]; then result=255; break; fi
sleep 5 sleep 5
done done
exit "$result" exit "$rc"
;;
summary)
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
rc=0
# shellcheck disable=SC2029 # The run ID is validated above.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
if [ "$rc" -eq 0 ]; then
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
else
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
fi
fi
;; ;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;; *) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac esac
+1
View File
@@ -94,6 +94,7 @@ replacements.txt
.idea .idea
# Temp files # Temp files
edu_master/temp/
temp/* temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts) # Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/ tmp/
-163
View File
@@ -1,163 +0,0 @@
# Homelab
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
Gitea Actions that build and deploy them. Most applications have both deployment
formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes
ingress; the media stack has Compose and Kubernetes routing configuration.
These files contain this lab's domains, IP addresses, storage paths, and private
registry names. Running them on another machine takes some editing.
## Start here
- [Service list](#services) — what each directory contains.
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status.
- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository.
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
[cert-manager](cert-manager/README.md) — common dependencies.
## What gets deployed
The `active` files are switches for the deploy workflow, not health indicators.
| File | Effect |
| ---------------------- | ----------------------------------------------------------- |
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
| Both | Run the Compose stack and apply the Kubernetes resources. |
| Neither | Keep the configuration in Git without automatic deployment. |
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
manual entry points. The deploy script does not discover them.
Kubernetes selection excludes secret files, examples, Helm values, and patches.
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
does not install their charts.
The table below describes committed configuration. It does not claim that a
service is currently healthy or running.
## Services
| Service | Configuration | Selected by markers |
| ---------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
## Running a Compose stack
Use the service README first. Where a service has an env example, copy it inside
that service's directory and replace the placeholders. The root `.env.example`
is an older collection of variables, not a complete configuration for every stack.
For example, from the repository root:
```sh
cd netbox
cp .env.example .env
$EDITOR .env
docker compose config --quiet
docker compose up -d
docker compose ps
```
Stacks that attach to `proxy` require an existing Docker network of that name and
an appropriate reverse proxy. Published host ports still work independently of
Traefik. Check port conflicts before starting an alternative to a Kubernetes
service: DNS, STUN, and HTTP listeners can share the same host.
`docker compose down` keeps named volumes. Adding `-v` removes them.
## Preparing Kubernetes
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
Replace the lab's hosts and addresses before using the configuration elsewhere.
Create a service's namespace, then prepare its ignored Secret from the example.
For example:
```sh
kubectl apply -f netbox/k8s/namespace.yaml
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
$EDITOR netbox/k8s/secrets.yaml
kubectl apply -f netbox/k8s/secrets.yaml
```
The deploy workflow applies the tracked resources for marked services. Avoid
applying an entire `k8s/` directory blindly: some directories contain Helm values,
examples, and alternative routes. For a manual change, apply the selected manifest
explicitly and check the resulting rollout.
Shared database passwords must agree between the `database` namespace and each
application's Secret. Updating the PostgreSQL Secret does not change an existing
role's password; see the database README.
## Local checks
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
```fish
set tools_dir (bash .gitea/workflows/install-ci-tools.sh)
set -gx PATH $tools_dir $PATH
ruff check .
ruff format --check .
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
.gitea/workflows/sync-renovate-configmap.sh --check
```
The [workflow README](.gitea/README.md#ci) lists the rest of the checks.
Structure checks do not establish that local Secrets, mounted files, storage,
or external services are ready.
## Data and recovery
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
configuration. Keep backups of application data and the keys needed to read it.
An image rollback does not roll back database migrations or ConfigMap contents.
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
The manifests do not provide a repository-wide backup schedule.
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
is a planning document, not evidence that NFS has been installed.
-22
View File
@@ -1,22 +0,0 @@
# AdGuard Home
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
configuration and working data, and mounts the `adguard-certs` TLS Secret.
The LoadBalancer Service exposes DNS separately from the web ingress.
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
and `certs/` before starting it. Starting both DNS deployments on the same address
can cause a port conflict.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n adguard
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-22
View File
@@ -1,22 +0,0 @@
# Authentik
Identity provider with separate server and worker deployments.
Kubernetes connects to the shared PostgreSQL service in `database`. Set
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
Keep `AUTHENTIK_SECRET_KEY` with the backups.
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
Its image defaults differ from Kubernetes; check both before an upgrade.
The worker mounts the Docker socket for Docker outpost management.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n authentik
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-18
View File
@@ -1,18 +0,0 @@
# cert-manager
Public ACME issuers and an internal certificate authority.
This directory contains chart values and issuer resources, not the controller
installation. Install the cert-manager chart with CRDs and the settings in
`k8s/cert-manager-values.yaml` before applying the issuers.
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
reachability must work for the requested names before issuance.
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
up; the tracked `.crt` is only a public certificate.
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
`kubectl apply` does not interpret the Helm values file.
See the [repository README](../README.md) for deployment selection.
-22
View File
@@ -1,22 +0,0 @@
# Cloudflare DDNS
Updates the lab DNS records when the public address changes.
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
The Compose stack also uses host networking. Configure the API token and domain
list from the relevant example; keep DNS names consistent with the ingress rules.
`config.json.example` is a separate configuration example. The current Compose
file does not mount a config.json file. Check configuration against the pinned
DDNS image when changing between environment and file-based settings.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# Checkmk
Checkmk Raw monitoring site with web and agent-receiver ingress.
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
volume on Compose. The agent receiver has a separate TCP route; enabling the
web route alone does not expose it.
Prepare the password in the service env or Secret example. Inspect the Checkmk
container logs during the first site creation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n checkmk
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# Cloudflare Tunnel
A Kubernetes connector for an existing Cloudflare tunnel.
The Deployment runs in `default` and reads its token from the ignored Secret
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
in Cloudflare before starting the connector.
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
`k8s/deployment.yaml` when this tunnel is needed.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-22
View File
@@ -1,22 +0,0 @@
# File converters
ConvertX for server-side conversion and BentoPDF for PDF tools.
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
Kubernetes configuration includes a local `config.yaml.example`, excluded from
normal deployment. Copy and apply the real ConfigMap separately where required.
Compose publishes ConvertX on host port 9992 as well as attaching it to the
proxy network. Replace the authentication settings from `.env.example` before
exposing it outside the lab.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n converters
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-25
View File
@@ -1,25 +0,0 @@
# CrowdSec
Helm values, dashboards, network policy, and a maintenance CronJob.
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
marker in this directory.
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
namespace Role. Review its script and schedule before enabling cleanup.
Traefik's values state that enforcement moved to a host firewall bouncer. This
repository does not install that host component.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n crowdsec
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# Dockmon
Docker management UI that talks to the host Docker daemon.
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
to the node hosting the pod, so this is not a cluster-wide container manager.
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
with a volume claim template. Its ServersTransport is specific to the upstream
connection; keep it with the ingress resources.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n dockmon
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-134
View File
@@ -1,134 +0,0 @@
# Repository review (6 October 2026 baseline)
This records the tracked tree at `cc9c3de` and the workstation state observed on
6 October 2026. It is a historical review, not a current runtime inventory. The
listed code fixes have since merged into `main`; EDU ownership has moved to the
separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md).
See the [CI and deployment guide](../.gitea/README.md) and
[runner and recovery guide](../.gitea/runner/README.md) for the current workflow.
No deployment was performed during the original review.
## Findings at the baseline and current status
| Priority | Finding at the baseline | Current status |
| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. |
| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. |
| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. |
| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. |
| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. |
| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. |
| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. |
| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. |
| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. |
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
The fix retains the DoH path constraint for both hostnames.
The current deploy workflow deliberately rejects the unsafe prune option. It
does not introduce automatic deletion under a different implementation. The
baseline finding was a configuration risk, not evidence of a live deletion
incident.
The former session fix bounded HTTP and Redis calls, validated credentials, set
a cookie lifetime of two refresh intervals, and marked success only after
publishing the verified cookie. The service is now owned by the EDU repository;
see that repository for its current implementation.
The deployment fix extracts required pod Secret references from rendered JSON,
checks their namespaces, includes init containers, image-pull credentials, and
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
issued by cert-manager are not treated as pre-existing pod prerequisites.
It checks existence/access, not every key's contents or application validity.
## Live workstation observations
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
no pods were Pending or in another non-running, non-completed phase. This is a
point-in-time observation, not a complete application health test.
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
reviewed local commit. It has untracked host configuration and a separate
`userbot/` directory. It was not reset or cleaned.
| Observed difference | Implication |
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The
first row above records the state before that change. Read
[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current
tracked monitoring configuration; the live observations in this section remain
a snapshot from 6 October.
## Current recovery limits
The deployment controller and its recovery process changed after this review.
The current operator workflow is documented in the
[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries
are:
- Kubernetes recovery can restore captured workload revisions. It does not
restore ConfigMaps, Secrets, database schemas, or persistent data.
- Compose recovery is manual. It uses saved resolved configuration, but it does
not restore volume data or reverse database migrations.
- Removed resources require manual review and removal; the deploy workflow does
not prune them automatically.
- Plan mode does not create namespaces. During apply, server validation for new
namespaces runs after namespace creation and chart installation; a failed
check can leave an empty namespace.
- Storage policy and backup coverage remain service-specific. Check the live PV,
PVC, and backup state before changing stateful workloads.
## Validation
At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose
files, and Kubernetes resources with available schemas. Kubeconform found 347
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
That skip count matters: passing schema validation does not validate Traefik rule
strings or other controller-specific behavior.
Fix validation covers:
- Compose discovery of manual entry points, rejection of required-variable gaps,
namespace-scoped and optional Secret references, and API/render failures.
- NetBird setup idempotence, preservation of existing keys, file permissions,
runtime rendering, and rejection of invalid trusted proxy CIDRs.
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
missing credentials, and nonpositive refresh intervals.
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
- YAML and Compose structure for the corrected router rules, compared with the
documented Traefik grammar. They were not exercised on the live proxy.
- Prune rejection before any cluster invocation.
At the time of review, all seven fix branches and the documentation branch
merged together in a disposable validation worktree. That combined tree passed the
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
Compose structure checks, and 11 Python regression tests plus the shell
validation regressions. CRD server-side validation and live rollout tests were
not run.
Runtime tests use fixtures and mocks, not production credentials. Live checks read
workload metadata, storage policies, chart versions, and container state only.
They did not read Secret contents or change services.
## Reloader follow-up (baseline)
`fix/reloader-integration` added the active marker and opt-in annotations to
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
It corrected AdGuard's misplaced pod-template annotation. The Helm settings use
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
CronJobs. PostgreSQL workloads are excluded because their credential variables
and init scripts are only effective on an empty data directory.
The controller was running on the workstation when inspected. The original
review checked configuration against the pinned chart with Helm rendering and
manifest validation; it did not change production configuration to provoke a
test restart or confirm every application's live reload behavior.
-20
View File
@@ -1,20 +0,0 @@
# Downtify
Download UI with a persistent downloads directory.
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
so check certificate and middleware availability before enabling them.
Back up downloads separately if they need to survive storage replacement.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n downtify
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+14
View File
@@ -0,0 +1,14 @@
EDU_LOGIN=your_edu_login_here
EDU_PASSWORD=your_edu_password_here
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
PHPSESSID_INTERVAL=10
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
WEBINAR_CHECK_INTERVAL=60
REDIS_HOST=redis
REDIS_PORT=6379
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
TZ=Europe/Kyiv
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
WEBINAR_ADMIN_ID=123456789
+1
View File
@@ -0,0 +1 @@
1.56.0
+14
View File
@@ -0,0 +1,14 @@
# EDU deployment ownership
Application source and release builds: `forust/edu-master`.
The homelab pipeline deploys `edu_master/k8s` and preserves explicit image digests.
The application copies in this directory are legacy and are not build inputs.
Do not publish EDU `prod` images from homelab or resolve releases from moving tags.
For an EDU release, validate both images, select their digests in the keeper and
checker manifests, and run the existing homelab validation/apply/verification
helpers against this service. Keep the existing Secret and Redis PVC.
Coordinate Redis authentication changes with both clients and all init/probes;
keep a pre-rollout Redis backup and both previous compatible image references.
The current HTTP checker does not depend on Playwright; check other consumers
before removing the separate browser service.
+49
View File
@@ -0,0 +1,49 @@
services:
redis:
image: redis:8.10.2-alpine
restart: unless-stopped
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
healthcheck:
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
interval: 30s
timeout: 5s
retries: 10
start_period: 60s
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
session-keeper:
condition: service_healthy
playwright-service:
condition: service_started
volumes:
redis-data:
View File
Whitespace-only changes.
+115
View File
@@ -0,0 +1,115 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / http error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
- alert: WebinarCheckerNeverStarted
expr: |
(time() - edu_process_start > 120)
and (webinar_check_last_run_timestamp_seconds == 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker job has not started"
description: "The process exposes metrics but its webinar job has never started."
- alert: WebinarDeliveryPending
expr: edu_delivery_pending > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Webinar notifications await delivery"
description: "Telegram delivery has pending recipients. Check delivery failures and retry status."
- alert: EduRedisUnavailable
expr: edu_redis_connected == 0
for: 2m
labels:
severity: critical
annotations:
summary: "EDU checker cannot reach Redis"
description: "Redis health checks are failing; checker commands and delivery may be unavailable."
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker deployment unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: edu-master
+69
View File
@@ -0,0 +1,69 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: playwright-service
namespace: edu-master
labels:
app: edu-master-playwright
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-playwright
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-playwright
spec:
containers:
- name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command:
- npx
- -y
- playwright@1.56.0
- run-server
- --port
- "3000"
- --path
- /ws
ports:
- containerPort: 3000
readinessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 3
livenessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
---
apiVersion: v1
kind: Service
metadata:
name: playwright-service
namespace: edu-master
spec:
selector:
app: edu-master-playwright
ports:
- name: ws
port: 3000
targetPort: 3000
+22
View File
@@ -0,0 +1,22 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: redis-clients-only
namespace: edu-master
spec:
podSelector:
matchLabels:
app: edu-master-redis
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: edu-master-session-keeper
- podSelector:
matchLabels:
app: edu-master-webinar-checker
ports:
- protocol: TCP
port: 6379
+96
View File
@@ -0,0 +1,96 @@
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: redis
namespace: edu-master
labels:
app: edu-master-redis
spec:
serviceName: redis
replicas: 1
selector:
matchLabels:
app: edu-master-redis
template:
metadata:
labels:
app: edu-master-redis
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
imagePullPolicy: IfNotPresent
env:
- name: REDIS_PASSWORD
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
- |
case "$REDIS_PASSWORD" in *[!0-9a-fA-F]*|'') echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1;; esac
[ "${#REDIS_PASSWORD}" -eq 64 ] || { echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1; }
umask 077
printf 'requirepass "%s"\n' "$REDIS_PASSWORD" > /tmp/redis-auth.conf
chown redis:redis /tmp/redis-auth.conf
exec docker-entrypoint.sh redis-server /tmp/redis-auth.conf
ports:
- containerPort: 6379
volumeMounts:
- name: redis-data
mountPath: /data
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 3
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 3
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: redis-data-pvc
namespace: edu-master
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis
namespace: edu-master
spec:
selector:
app: edu-master-redis
ports:
- name: redis
port: 6379
targetPort: 6379
@@ -0,0 +1,50 @@
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
#
# Runbook:
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
# 2. docker run --rm -v edu_master_redis-data:/data \
# -v /tmp/edu-master-backup:/backup \
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
# 6. kubectl delete job redis-restore-seed -n edu-master
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
apiVersion: batch/v1
kind: Job
metadata:
name: redis-restore-seed
namespace: edu-master
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 3600
template:
spec:
restartPolicy: Never
containers:
- name: seed
image: redis:alpine
command:
- /bin/sh
- -ec
- |
ls -la /backup
cp /backup/dump.rdb /data/dump.rdb
chmod 644 /data/dump.rdb
ls -la /data
volumeMounts:
- name: redis-data
mountPath: /data
- name: backup
mountPath: /backup
readOnly: true
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
- name: backup
hostPath:
path: /tmp/edu-master-backup
type: DirectoryOrCreate
+29
View File
@@ -0,0 +1,29 @@
apiVersion: v1
kind: Secret
metadata:
name: edu-master-secrets
namespace: edu-master
type: Opaque
stringData:
# Session keeper credentials
KEEPER_LOGIN: ""
KEEPER_PASSWORD: ""
KEEPER_INTERVAL: "10"
# EDU links
EDU_URL_BASE: "https://edu.edu.vn.ua"
EDU_URL_LOGIN: "/user/login"
EDU_URL_COURSES: "/course/userlist"
EDU_URL_WEBINAR: "/webinar/useractive"
# Playwright
USER_AGENT: ""
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
# Webinar-checker
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
TZ: "Europe/Kyiv"
+15
View File
@@ -0,0 +1,15 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
@@ -1,14 +1,14 @@
apiVersion: monitoring.coreos.com/v1 apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor kind: ServiceMonitor
metadata: metadata:
name: netbird-server name: webinar-checker
namespace: netbird namespace: edu-master
labels: labels:
release: prometheus-stack release: prometheus-stack
spec: spec:
selector: selector:
matchLabels: matchLabels:
app: netbird-server app: edu-master-webinar-checker
endpoints: endpoints:
- port: metrics - port: metrics
path: /metrics path: /metrics
+69
View File
@@ -0,0 +1,69 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: session-keeper
namespace: edu-master
labels:
app: edu-master-session-keeper
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-session-keeper
strategy:
type: Recreate
template:
metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels:
app: edu-master-session-keeper
spec:
initContainers:
- name: wait-redis
image: redis:8.10.2-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper@sha256:49285e87cc5bc4cf4ffe190813d87927916c2df8a206daac0aeb7d227c636450
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 10
+87
View File
@@ -0,0 +1,87 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-webinar-checker
strategy:
type: Recreate
template:
metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels:
app: edu-master-webinar-checker
spec:
# Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID)
initContainers:
- name: wait-deps
image: redis:8.10.2-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis ok"
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
sleep 2
done
echo "PHPSESSID ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker@sha256:66c146f7b43cb9f0dc31ba9aa36d217e01df42ddafba5971b79c12ec215b2c01
ports:
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
livenessProbe:
httpGet:
path: /live
port: metrics
initialDelaySeconds: 60
periodSeconds: 15
timeoutSeconds: 3
failureThreshold: 4
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: TZ
value: "Europe/Kyiv"
resources:
requests:
cpu: "50m"
memory: "192Mi"
limits:
cpu: "600m"
memory: "384Mi"
+15
View File
@@ -0,0 +1,15 @@
FROM python:3.14-slim
WORKDIR /app
# Install system dependencies
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
# Install dependencies
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
# Copy application code
COPY . .
# Run the bot
CMD ["python", "bot.py"]
+132
View File
@@ -0,0 +1,132 @@
import logging
import os
import time
from datetime import datetime
import redis
import requests
# Configure logging
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Load configuration (adapted to .env keys)
def _env(key, default=None):
v = os.getenv(key, default)
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
return v[1:-1]
return v
LOGIN = _env('KEEPER_LOGIN')
PASSWORD = _env('KEEPER_PASSWORD')
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
USER_AGENT = _env(
'USER_AGENT',
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
)
REDIS_HOST = _env('REDIS_HOST', 'redis')
REDIS_PORT = int(_env('REDIS_PORT', 6379))
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
def touch_success_file():
"""Updates the timestamp of the success file for healthchecks."""
try:
with open(SUCCESS_FILE, 'w') as f:
f.write(str(datetime.now().timestamp()))
except Exception as e:
logger.error(f'Failed to touch success file: {e}')
def main():
logger.info('Starting Session Keeper Bot')
# Connect to Redis
try:
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
redis_client.ping()
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
except Exception as e:
logger.error(f'Failed to connect to Redis: {e}')
return
session = requests.Session()
# Set headers
headers = {
'User-Agent': USER_AGENT,
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
'Accept-Language': 'en-US,en;q=0.9',
'Cache-Control': 'max-age=0',
'Upgrade-Insecure-Requests': '1',
'Sec-Fetch-Site': 'same-origin',
'Sec-Fetch-Mode': 'navigate',
'Sec-Fetch-User': '?1',
'Sec-Fetch-Dest': 'document',
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
'Sec-Ch-Ua-Mobile': '?0',
'Sec-Ch-Ua-Platform': '"Linux"',
'Accept-Encoding': 'gzip, deflate, br',
'Priority': 'u=0, i',
}
session.headers.update(headers)
while True:
try:
logger.info('Attempting login...')
# Login payload
payload = {'login': LOGIN, 'password': PASSWORD}
# Perform Login
# Note: The user request shows a POST to /user/login with form data
# We need to make sure we handle the PHPSESSID correctly.
# If we already have a PHPSESSID, requests will send it.
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
logger.info(f'Login Response Status: {login_response.status_code}')
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
# Verify Session
logger.info('Verifying session...')
verify_response = session.get(URL_VERIFY, allow_redirects=False)
logger.info(f'Verify Response Status: {verify_response.status_code}')
if verify_response.status_code == 200:
logger.info('Session verification SUCCESS (200 OK).')
touch_success_file()
# Save PHPSESSID to Redis
phpsessid = session.cookies.get('PHPSESSID')
if phpsessid:
try:
redis_client.set('EDU_PHPSESSID', phpsessid)
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
except Exception as e:
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
elif verify_response.status_code == 302:
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
else:
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
except Exception as e:
logger.error(f'An error occurred: {e}')
logger.info(f'Sleeping for {INTERVAL} minutes...')
time.sleep(INTERVAL * 60)
if __name__ == '__main__':
main()
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.14-slim
WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
CMD ["python", "checker.py"]
File diff suppressed because it is too large. Load diff
-21
View File
@@ -1,21 +0,0 @@
# Error pages
Static HTTP error pages served by an Nginx image built in CI.
Edit the HTML in `html/`; the Dockerfile copies it into the image.
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
middleware. Keep the middleware's namespace and port aligned with that Service.
For a local build, run `docker build -t homelab-error-pages .` from this directory.
Compose references the private registry image rather than a build context.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n error-pages
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-25
View File
@@ -1,25 +0,0 @@
# Gitea
Git hosting with HTTP and a separate SSH route.
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
and application data. Match the Gitea database password with the shared database
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
its own database, not a second frontend for the Kubernetes instance.
Back up repositories, application configuration, and a consistent database dump
together. Gitea Actions definitions for this repository live in `../.gitea/`.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n gitea
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-2
View File
@@ -20,8 +20,6 @@ data:
GITEA__mailer__ENABLED: "false" GITEA__mailer__ENABLED: "false"
GITEA__metrics__ENABLED: "true"
# No code/issue search needed: bleve reindexes the whole issue index on # No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers # every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres. # the rotational disk for an hour. "db" serves issue search from postgres.
-2
View File
@@ -3,8 +3,6 @@ kind: Service
metadata: metadata:
name: gitea-service name: gitea-service
namespace: gitea namespace: gitea
labels:
app: gitea
spec: spec:
selector: selector:
app: gitea app: gitea
+1 -2
View File
@@ -7,8 +7,7 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
# Metrics are scraped directly through the cluster Service. - match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
kind: Rule kind: Rule
services: services:
- name: gitea-service - name: gitea-service
-16
View File
@@ -1,16 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: gitea
namespace: gitea
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: gitea
endpoints:
- port: http
path: /metrics
interval: 30s
scrapeTimeout: 10s
-26
View File
@@ -1,26 +0,0 @@
# Glance
Dashboard pages for links, service checks, and Docker containers.
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
`user.css`. Update both copies when changing shared content.
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
no tracked Secret example; create that Secret in `glance` before starting it.
Compose expects a local `.env` with the same password.
The pod mounts `user.css` from `glance-assets`, which is the ConfigMap that
contains that key.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n glance
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-18
View File
@@ -1,18 +0,0 @@
# Headscale
Headscale, Headplane, and a separate web administration UI on Docker.
Kubernetes only provides routes to the Docker host. Update the addresses in
`k8s/routing/external-service.yaml` if the host moves.
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
`config/policy.json.example` to their names without `.example`. Set the public
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
The example URLs are placeholders.
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
13000, and the other UI on 10080. The data volumes store the Headscale database,
keys, and Headplane state. The embedded DERP configuration needs reachable
addresses; Compose does not publish its UDP 3478 listener.
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -7,7 +7,7 @@
# # Dev server_url # # Dev server_url
# server_url: https://hs.dev_internal_domain.internal # server_url: https://hs.dev_internal_domain.internal
listen_addr: 0.0.0.0:8080 listen_addr: 0.0.0.0:8080
metrics_listen_addr: 0.0.0.0:9090 metrics_listen_addr: 127.0.0.1:9090
grpc_listen_addr: 127.0.0.1:50443 grpc_listen_addr: 127.0.0.1:50443
grpc_allow_insecure: false grpc_allow_insecure: false
noise: noise:
@@ -3,8 +3,6 @@ kind: Service
metadata: metadata:
name: headscale-server-external name: headscale-server-external
namespace: headscale namespace: headscale
labels:
app: headscale
spec: spec:
ports: ports:
- port: 8080 - port: 8080
-18
View File
@@ -1,18 +0,0 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: headscale
namespace: headscale
labels:
release: prometheus-stack
spec:
# The external Service has a manually managed EndpointSlice, not Endpoints.
discoveryRole: endpointslice
selector:
matchLabels:
app: headscale
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
-24
View File
@@ -1,24 +0,0 @@
# Homarr
Dashboard with Kubernetes integration and persistent application state.
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
The committed ingress is internal. There is no `k8s/active` marker even though
manifests exist, so the workflow does not select Homarr automatically.
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
expects a local kubeconfig. Check these host ports against Traefik before use.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n homarr
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-25
View File
@@ -1,25 +0,0 @@
# Homepages
Two static sites: Forust and xdfnx.
The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with
its own Dockerfile and publishes it to the private registry. Kubernetes serves
the image contents; Compose overlays the source directories as bind mounts.
Both Traefik IngressRoute and Gateway API route manifests are committed.
Keep their hostnames and backend Services aligned when changing routes.
Certificate resources cover public and internal hostnames.
Build either site locally with `docker build -f Dockerfile.forust .` or
`docker build -f Dockerfile.xdfnx .` from this directory.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n homepages
kubectl get events -n homepages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-29
View File
@@ -1,29 +0,0 @@
# Immich
Photo library with its own vector-enabled PostgreSQL and machine-learning service.
This database is separate from the shared PostgreSQL instance. Keep the server
and machine-learning versions aligned when upgrading.
Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must
already exist and contain the intended library; moving the pod to a different
node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and
the model cache has its own PVC.
Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example
uses the same library path as Kubernetes. Run one writer against that library;
do not start both deployments as independent instances over the same files.
Back up the library and a consistent database dump together. The model cache
can be rebuilt; the photo database cannot.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n immich
kubectl get events -n immich --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-4
View File
@@ -6,10 +6,6 @@ metadata:
data: data:
TZ: "Europe/Bratislava" TZ: "Europe/Bratislava"
IMMICH_TELEMETRY_INCLUDE: "all"
IMMICH_API_METRICS_PORT: "8081"
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
# The database in this namespace, not the shared one in the database # The database in this namespace, not the shared one in the database
# namespace: v3 needs VectorChord, and only the dedicated image carries it. # namespace: v3 needs VectorChord, and only the dedicated image carries it.
DB_HOSTNAME: "immich-postgres" DB_HOSTNAME: "immich-postgres"
-12
View File
@@ -3,8 +3,6 @@ kind: Service
metadata: metadata:
name: immich-service name: immich-service
namespace: immich namespace: immich
labels:
app: immich
spec: spec:
selector: selector:
app: immich app: immich
@@ -12,12 +10,6 @@ spec:
- name: http - name: http
port: 2283 port: 2283
targetPort: 2283 targetPort: 2283
- name: api-metrics
port: 8081
targetPort: api-metrics
- name: worker-metrics
port: 8082
targetPort: worker-metrics
--- ---
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
@@ -49,10 +41,6 @@ spec:
ports: ports:
- name: http - name: http
containerPort: 2283 containerPort: 2283
- name: api-metrics
containerPort: 8081
- name: worker-metrics
containerPort: 8082
volumeMounts: volumeMounts:
- name: immich-data - name: immich-data
mountPath: /data mountPath: /data
-20
View File
@@ -1,20 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: immich
namespace: immich
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: immich
endpoints:
- port: api-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
- port: worker-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
-21
View File
@@ -1,21 +0,0 @@
# Kener
Status page with Redis and persistent database and upload directories.
Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet.
Compose keeps the corresponding directories in named volumes. Set the signing
and other credentials from the env or Secret example.
The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`.
There is no active marker for either runtime.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n kener
kubectl get events -n kener --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
kener: kener:
image: rajnandan1/kener:v4.1.7 image: rajnandan1/kener:4.1.5
container_name: kener container_name: kener
restart: unless-stopped restart: unless-stopped
# ports: # ports:
+1 -1
View File
@@ -31,7 +31,7 @@ spec:
spec: spec:
containers: containers:
- name: kener - name: kener
image: rajnandan1/kener:v4.1.7 image: rajnandan1/kener:4.1.5
envFrom: envFrom:
- configMapRef: - configMapRef:
name: kener-config name: kener-config
-16
View File
@@ -1,16 +0,0 @@
# Loki and Alloy
Loki log storage and Alloy collection, both deployed through Helm.
The deploy library lists separate `loki` and `alloy` releases in `prometheus`,
controlled by this directory's `k8s/active` marker. Chart versions are pinned in
`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`.
Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the
monitoring stack. Review Loki retention and storage settings before enabling
collection on a new cluster.
Check releases with `helm list -n prometheus` and inspect collector logs before
assuming that an empty Grafana query means there were no events.
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# MeTube
Web downloader behind Traefik.
Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB
`emptyDir` for `/downloads`: completed downloads disappear when the pod is
replaced. Download files from the UI promptly if this temporary storage is intended.
Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes
would require changing the volume to a PVC and choosing a storage policy.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n metube
kubectl get events -n metube --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# n8n
Workflow automation with persistent application and file storage.
Kubernetes keeps application state in `n8n-node-pvc` and files in
`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes.
Webhook URLs and proxy settings are committed in the application config.
There is no active marker. Review the URLs before enabling the stack, and retain
the credential encryption key with the database or application-data backup.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n n8n
kubectl get events -n n8n --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
n8n: n8n:
image: docker.n8n.io/n8nio/n8n:2.43.2 image: docker.n8n.io/n8nio/n8n:2.43.1
container_name: n8n container_name: n8n
restart: unless-stopped restart: unless-stopped
environment: environment:
+1 -1
View File
@@ -31,7 +31,7 @@ spec:
spec: spec:
containers: containers:
- name: n8n - name: n8n
image: docker.n8n.io/n8nio/n8n:2.43.2 image: docker.n8n.io/n8nio/n8n:2.43.1
envFrom: envFrom:
- configMapRef: - configMapRef:
name: n8n-config name: n8n-config
+4 -16
View File
@@ -1,24 +1,12 @@
# NetBird # NetBird
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind Traefik. The Compose configuration uses the external Docker `proxy` network and publishes only STUN UDP `3478` directly. Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
The single-instance server uses SQLite. Back up its data and datastore encryption key together. The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
## Kubernetes
`k8s/active` selects the Kubernetes deployment. It runs the server and dashboard
in namespace `netbird`; the server stores SQLite data in `netbird-pvc`. The
configuration renderer and template are in `k8s/`. Prepare
`k8s/secrets.yaml` from `k8s/secrets.yaml.example` before the first deploy.
## Compose alternative
The Compose files are available for manual use. There is no root `active` marker,
so the automatic deploy workflow selects Kubernetes only.
## Files ## Files
- `compose.yaml`: dashboard and combined server; start it manually when using Compose. - `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
- `config.template.yaml`: non-secret server configuration rendered at startup. - `config.template.yaml`: non-secret server configuration rendered at startup.
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration. - `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key. - `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
@@ -27,7 +15,7 @@ so the automatic deploy workflow selects Kubernetes only.
## First deployment ## First deployment
Run these commands on the Docker host before the first Compose start. Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
```bash ```bash
cd /srv/homelab/netbird cd /srv/homelab/netbird
-9
View File
@@ -3,8 +3,6 @@ kind: Service
metadata: metadata:
name: netbird-server-service name: netbird-server-service
namespace: netbird namespace: netbird
labels:
app: netbird-server
spec: spec:
selector: selector:
app: netbird-server app: netbird-server
@@ -13,10 +11,6 @@ spec:
name: http name: http
targetPort: 80 targetPort: 80
protocol: TCP protocol: TCP
- port: 9090
name: metrics
targetPort: metrics
protocol: TCP
- port: 3478 - port: 3478
name: stun name: stun
targetPort: 3478 targetPort: 3478
@@ -65,9 +59,6 @@ spec:
- containerPort: 80 - containerPort: 80
name: http name: http
protocol: TCP protocol: TCP
- containerPort: 9090
name: metrics
protocol: TCP
- containerPort: 3478 - containerPort: 3478
name: stun name: stun
protocol: UDP protocol: UDP
+88 -35
View File
@@ -1,43 +1,96 @@
# NetBox # NetBox
Inventory and network documentation with a web process, worker, and Valkey. NetBox for homelab documentation and visualization. Two runtimes are available:
Kubernetes uses the shared PostgreSQL service at | Runtime | Manifest | Purpose |
`postgres.database.svc.cluster.local:5432`, database and role `netbox`. | ------- | -------------- | -------------------------------------------------------------- |
The database and application Secrets must contain the same password. | Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) |
Media, reports, scripts, and Valkey have persistent storage. | k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) |
Compose has its own PostgreSQL container and Valkey instances. It publishes the Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks
web UI on `127.0.0.1:8000`; its Traefik labels can also expose it while a Docker plus a second logical database for caching. The Docker stand keeps its own
proxy is running. Copy `.env.example` to `.env`, replace the credentials, and run PostgreSQL container, while the k8s deployment uses the shared `database` cluster
`docker compose config --quiet` before starting it. (`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey
stays a per-service StatefulSet.
## First Kubernetes start ## Docker Compose
Create the namespace and application Secret. Provision the database through the ```bash
shared database initializer on a fresh instance, or create the role and database cp .env.example .env
manually on an existing instance; see [PostgreSQL](../postgres/README.md). # replace CHANGE_ME
The database NetworkPolicy already includes `netbox`. docker compose up -d
Apply the selected application manifests after the database is ready. Startup
runs schema migrations, so the probes allow a longer first boot. Inspect web and
worker logs before retrying a slow migration.
## Settings and backup
`configuration/configuration.py` is the Compose settings file. Its Kubernetes
copy is embedded in `k8s/settings.yaml`; keep them aligned.
Back up the database and media together. Keep `SECRET_KEY` and
`API_TOKEN_PEPPER_1`: changing them invalidates sessions or API tokens.
A container rollback cannot undo a database migration.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netbox
kubectl get events -n netbox --sort-by=.metadata.creationTimestamp
``` ```
See the [repository README](../README.md) for deployment selection. The UI is available at <http://localhost:8000>. The port is bound to `127.0.0.1`
intentionally, so this stand is not exposed on the LAN or public interfaces.
The `netbox` service is also attached to the external `proxy` network and carries
Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those
labels only take effect while the Docker Traefik stack is running; it is currently
stopped, and the live ingress path in this homelab is the k8s Traefik.
Inspect startup and health with:
```bash
docker compose ps
docker compose logs -f netbox
```
Stop it with `docker compose down`; data is kept in the named volumes
`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`,
`netbox-scripts-files` and `netbox-redis-data`.
## Kubernetes
`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly
plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To
rebuild it from scratch:
```bash
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
kubectl -n database patch secret postgres-shared-secrets \
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
# 2. secrets first: the deploy workflow never applies *secret*.yaml
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
kubectl apply -f k8s/secrets.yaml
# 3. manifests
kubectl apply -f k8s/
```
The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its
NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace
or connections are dropped, and `postgres/initdb/01-create-databases.sh` already
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
StatefulSet of its own — only `netbox-valkey`.
`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS`
list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the
`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`.
Resources are permanent again now that the first-boot migrations are complete:
the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the
worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves
`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps
leave enough headroom for future schema migrations without letting one process
consume the whole node.
The first start applies ~810 migrations, each in its own transaction with DDL and
a commit; every later start is a no-op. The startup probe allows 15 minutes and
`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod
and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would
replace the probe destination with that public hostname and bypass the pod.
## Secrets
- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only
`.env.example` and `k8s/secrets.yaml.example` are committed.
- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and
the Django keys all come from the environment, so the same settings file works in
both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap
(`k8s/settings.yaml`) and must be kept in sync with the file.
- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1`
invalidates every API token.
-22
View File
@@ -1,22 +0,0 @@
# Netronome
Network monitoring application using the shared PostgreSQL instance on Kubernetes.
Kubernetes reads application settings from its ConfigMap and Secret. Match the
Netronome role password with `NETRONOME_DB_PASSWORD` in the shared database Secret.
Its namespace is included in the PostgreSQL NetworkPolicy.
The Compose configuration is a separate deployment; review its local database
settings and env example before starting it. Keep monitoring history in the
database backup.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netronome
kubectl get events -n netronome --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-17
View File
@@ -1,17 +0,0 @@
# Nextcloud AIO
Nextcloud All-in-One on Docker, with Kubernetes routes to the Docker host.
The master container manages its own child containers through the Docker
socket. Kubernetes does not run the Nextcloud application; the EndpointSlices
under `k8s/routing/` point to host services.
Compose publishes the AIO administration interface on 8888. The Apache frontend
uses host port 11000. `NEXTCLOUD_DATADIR` is `/mnt/nextcloud/ncdata`; prepare that
storage before first setup and do not change the path casually afterwards.
Use AIO's backup and restore tools for the managed application. Keep the master
configuration volume and the data directory with the recovery plan. Do not
remove child containers just because they do not appear as Compose services.
See the [repository README](../README.md) for deployment selection.
-12
View File
@@ -1,12 +0,0 @@
# Penpot
A Compose-only design application with frontend, backend, exporter, database, and cache.
There is no active marker or Kubernetes deployment here. Configure the public
URL and credentials from `.env.example` before starting `compose.yaml`.
Penpot has its own PostgreSQL container. The shared database initializer still
contains a Penpot role, but this Compose stack does not use it.
Back up the application assets and database together.
See the [repository README](../README.md) for deployment selection.
-21
View File
@@ -1,21 +0,0 @@
# Portainer
Container management UI backed by the host Docker socket.
Kubernetes mounts the node's Docker socket and persists application data in
`portainer-data-pvc`. This targets Docker on that node, not Kubernetes workloads.
Compose uses the `portainer_data` volume for its state.
Review initial administrator setup and route access before exposing the UI.
Neither deployment has an active marker.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n portainer
kubectl get events -n portainer --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
portainer: portainer:
image: portainer/portainer-ce:2.45.2 image: portainer/portainer-ce:2.45.1
container_name: portainer container_name: portainer
restart: always restart: always
volumes: volumes:
+1 -1
View File
@@ -29,7 +29,7 @@ spec:
spec: spec:
containers: containers:
- name: portainer - name: portainer
image: portainer/portainer-ce:2.45.2 image: portainer/portainer-ce:2.45.1
ports: ports:
- containerPort: 9000 - containerPort: 9000
volumeMounts: volumeMounts:
+28 -56
View File
@@ -1,63 +1,35 @@
# Shared PostgreSQL # Shared PostgreSQL
PostgreSQL 17 for the Kubernetes deployments of Authentik, Gitea, NetBox, and Netronome. This directory contains the shared PostgreSQL 17 deployment for Authentik,
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
per service. Per-service standalone databases were removed after the
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
not part of the shared instance).
The server runs in `database` as StatefulSet `postgres17`, with data in ## Compatibility baseline
`postgres17-data`. Applications connect to
`postgres.database.svc.cluster.local:5432`. The NetworkPolicy allows only the
listed application namespaces; add a new consumer there as well as provisioning
its database.
## Initialization | Service | Current application | Shared PostgreSQL 17 |
| ---------- | ------------------- | -------------------------------------- |
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
| Statuspage | custom | Supported |
`initdb/01-create-databases.sh` creates roles and databases on an empty data A major-version change must use a logical dump/restore; changing only the
directory. The Kubernetes copy is embedded in `k8s/postgres.yaml`. image tag while keeping a data directory is not supported.
It also provisions Penpot and Statuspage roles, even though those are not active
consumers in the current Kubernetes manifests.
The initializer requires every listed password. Prepare `k8s/secrets.yaml` from For Compose, copy `.env.example` to `.env`, set all passwords, and start it with
the example before applying the StatefulSet. Existing application Secrets keep `docker compose -f shared-compose.yaml up -d`. This file is intentionally not
copies of their own database passwords; they must match the corresponding role. named `compose.yaml`, so the repository deploy workflow does not start a second
database accidentally.
Applications that use this database must also join that external network and use
`homelab-postgres:5432`.
The init scripts do not run again when an existing data directory is mounted. For Kubernetes, create `k8s/secrets.yaml` from the example before applying the
Changing a Secret does not rotate the PostgreSQL role password. Rotate the role manifests. The `k8s/active` marker makes the normal deploy workflow include the
with SQL and update the application Secret together. namespace, StatefulSet, ConfigMap, and NetworkPolicy. Applications use
`postgres.database.svc.cluster.local:5432`.
## Compose alternative Migrate each existing database with a tested logical dump/restore before
switching an application. Do not reuse a PostgreSQL 14 or 17 data directory
From this directory: with PostgreSQL 15.
```sh
cp .env.example .env
$EDITOR .env
docker compose -f shared-compose.yaml config --quiet
docker compose -f shared-compose.yaml up -d
```
The example includes `NETBOX_DB_PASSWORD`; fill it and every other required
password before starting the stack.
This stack creates the `homelab-database` Docker network and the
`homelab-postgres` container. Compose applications need to join that network
explicitly to use it; several committed Compose stacks use their own databases.
The filename is intentional: the automatic deploy discovery does not start this
stack just because the Kubernetes database is active.
## Backup and upgrades
Keep database dumps and role definitions, including ownership and grants.
Take a logical backup before changing a major PostgreSQL version. A new image
tag over the existing data directory is not a major-version migration.
Test restores separately before changing application connection settings.
Immich uses its own vector-enabled database and is outside this shared instance.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n database
kubectl get events -n database --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-32
View File
@@ -1,32 +0,0 @@
# Monitoring stack
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
and vmalert. The `k8s/active` marker selects the stack. The
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
Prometheus Operator, and related components. Its Prometheus server is configured
with zero replicas while VMAgent collects metrics and writes them to the
single-node VictoriaMetrics instance.
The `victoria-operator` Helm release converts selected Prometheus Operator
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
validation commands.
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
keep their credentials out of the values file. Persistent volumes store data for
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
backups before changing storage. VictoriaMetrics currently retains 30 days of
data.
A separate Compose configuration is present for manual use. There is no root
`active` marker, so the automatic deploy workflow does not select it.
The deploy workflow does not remove resources when manifests are deleted. For a
rollback of application-metrics changes, follow the explicit cleanup steps in
the [Kubernetes monitoring notes](k8s/README.md).
See the [repository README](../README.md) for deployment selection.
-26
View File
@@ -15,29 +15,3 @@ The VictoriaMetrics Operator chart and its CRDs are installed before the
Kubernetes manifests by the normal deploy workflow. On a cluster where the Kubernetes manifests by the normal deploy workflow. On a cluster where the
operator CRDs are not installed yet, CI skips the server-side dry-run of the operator CRDs are not installed yet, CI skips the server-side dry-run of the
`VMAgent` resource; the deploy installs the chart before applying that resource. `VMAgent` resource; the deploy installs the chart before applying that resource.
## Application metrics
The application ServiceMonitors use a 30s interval and a 10s timeout:
- Headscale: the external Service points to the Compose host on port 19090.
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
The Compose configuration must bind metrics to `0.0.0.0:9090`.
- NetBird: the combined server exports `/metrics` on port 9090. The existing
`server.metricsPort` setting enables the listener.
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
ingress excludes this path. The monitor uses the internal Service directly.
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
8081 and 8082. The monitor scrapes both ports on each server replica.
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
VictoriaMetrics. All targets should report 1.
For rollback, revert the application metrics changes, run CI, and deploy the
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
workflow applies manifests and does not prune removed resources.
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
previous Compose metrics bind address, and restart only the Headscale service.
-20
View File
@@ -1,20 +0,0 @@
# RackPeek
Rack inventory UI behind Traefik.
Kubernetes stores configuration in `rackpeek-pvc`. The Compose alternative uses
its own data mount. Keep rack descriptions and inventory data in the backup.
Public and internal certificates and routes are in `k8s/`. There are no tracked
Secret examples for this service.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n rackpeek
kubectl get events -n rackpeek --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-56
View File
@@ -1,56 +0,0 @@
# Reloader
Restarts opted-in workloads when the ConfigMaps or Secrets they consume change.
The deploy workflow upgrades the `reloader` Helm release in namespace `reloader`;
`k8s/active` enables it. The chart version is pinned in `deploy-lib.sh`.
## Workload integration
Put this annotation on the Deployment or StatefulSet metadata:
```yaml
metadata:
annotations:
reloader.stakater.com/auto: "true"
```
The annotation belongs to the workload, not `spec.template.metadata`.
Reloader discovers references in environment variables and mounted volumes.
This covers startup-only settings and ConfigMaps or Secrets mounted with `subPath`.
See the [upstream usage guide](https://github.com/stakater/Reloader/blob/v1.4.22/README.md#usage).
The application manifests opt in workloads including AdGuard's TLS files,
NetBird, both NetBox processes, and the password-protected Valkey servers.
Inactive services have the same annotations ready for later activation.
## Controller policy
The controller watches all namespaces but only restarts annotated workloads.
It uses the `annotations` reload strategy, so changes trigger a pod-template
annotation rather than injecting extra environment variables.
Jobs and CronJobs are excluded: their next execution reads current configuration.
PostgreSQL is intentionally not opted in. Its password variables and init scripts
apply to first initialization; restarting an existing database does not rotate
roles or rerun those scripts. Rotate database credentials with SQL and update the
clients' Secrets together.
Helm-managed monitoring components already have their own configuration reload
paths; Traefik watches its file-provider configuration. They are not globally
opted in. The controller does not react to files in PVCs or changes to external
services unless a watched ConfigMap or Secret changes.
## Verify
```sh
kubectl -n reloader rollout status deployment/reloader-reloader
kubectl -n reloader logs deployment/reloader-reloader --since=10m
kubectl -n netbird get deployment netbird-server-deployment \
-o jsonpath='{.metadata.annotations.reloader\.stakater\.com/auto}'
```
A changed configuration can briefly interrupt a single-replica service, especially
one using `Recreate`. Installing annotations does not validate the configuration
or migrate database data. Keep changes to shared Secrets coordinated across consumers.
See the [repository README](../README.md) for deployment selection.
+88 -38
View File
@@ -1,51 +1,101 @@
# Renovate # Renovate for Gitea
Container and chart dependency updates for the Gitea repository. Renovate runs as a Kubernetes CronJob and creates container image update pull
requests in Gitea. It does not deploy changes itself.
The Kubernetes CronJob runs in `renovate` every six hours with overlapping ## Kubernetes
CronJob executions forbidden. Prepare the bot PAT from the Secret example.
Give the dedicated Gitea user access to the repositories it should update.
`renovate.json` is the source configuration. The ConfigMap is a generated copy: Create a dedicated Gitea user named `renovate-bot`, create a repository access
token, and grant it repository read/write plus issue read/write permissions.
Add `read:packages` if Renovate must inspect private Gitea registry images.
Create the ignored Secret locally; never commit the PAT:
```sh ```sh
.gitea/workflows/sync-renovate-configmap.sh cp renovate/k8s/secrets.yaml.example renovate/k8s/secrets.yaml
.gitea/workflows/sync-renovate-configmap.sh --check $EDITOR renovate/k8s/secrets.yaml
kubectl apply -f renovate/k8s/namespace.yaml
kubectl apply -f renovate/k8s/secrets.yaml
kubectl apply -f renovate/k8s/configmap.yaml
kubectl apply -f renovate/k8s/cronjob.yaml
``` ```
Run those commands from the repository root. The `renovate-ci` workflow checks The `renovate/k8s/active` marker makes the normal deployment workflow include
that the generated configuration agrees with the source. the namespace, ConfigMap, and CronJob. The Secret is intentionally excluded
from Git and must be applied separately after every new cluster.
## Run manually Run it immediately instead of waiting for the six-hour schedule.
From the repository root: Two options, both use the same `renovate/renovate.json`:
```fish
kubectl create job --from=cronjob/renovate renovate-manual-(date +%s) -n renovate
kubectl get jobs,pods -n renovate
```
Alternatively use the `renovate-run` Actions workflow. It reads the image tag
from the CronJob and accepts repository, log-level, and dry-run inputs. Actions
requires `RENOVATE_TOKEN`; `RENOVATE_GITHUB_COM_TOKEN` is optional.
The Actions concurrency group and the CronJob policy are separate, so avoid
starting both against the same repository at once.
For Compose, copy `.env.example` to `.env` in this directory and run
`docker compose -f renovate-compose.yaml run --rm renovate`. That file is a
manual entry point and is not selected by the deploy workflow.
The config also tracks chart versions in `deploy-lib.sh` and tool versions in
`.gitea/workflows/tool-versions.env`. Renovate opens pull requests; the normal CI and deploy
workflows handle changes after merge.
## Inspect
From the repository root:
```sh ```sh
kubectl get pods,svc,pvc -n renovate kubectl create job --from=cronjob/renovate renovate-manual-$(date +%s) -n renovate
kubectl get events -n renovate --sort-by=.metadata.creationTimestamp
``` ```
See the [repository README](../README.md) for deployment selection. or the `renovate-run` Actions workflow (Actions tab → `renovate-run` →
Run workflow). It runs the same image as the CronJob on the self-hosted runner
via Docker — the tag is read out of `renovate/k8s/cronjob.yaml` at run time
rather than hardcoded, so the two cannot drift apart. Required Actions secrets
(repo or org settings):
- `RENOVATE_TOKEN` — renovate-bot PAT (repository + issue read/write).
- `RENOVATE_GITHUB_COM_TOKEN` — optional, for changelogs and GitHub rate limits.
Inputs: `repositories` (default `forust/homelab`), `log_level`
(`info`/`debug`). Only one run at a time (concurrency group
`renovate-run`), same as the CronJob `Forbid` policy.
Inspect runs with:
```sh
kubectl get cronjob,jobs,pods -n renovate
kubectl logs -n renovate job/<job-name>
```
`RENOVATE_GITHUB_COM_TOKEN` is optional but recommended for changelogs and
GitHub API rate limits. Set it in the Kubernetes Secret if available.
## Compose
Copy `.env.example` to `.env`, set the PAT, and run:
```sh
docker compose -f renovate-compose.yaml run --rm renovate
```
The Compose file is intentionally named `renovate-compose.yaml`, so the
repository's automatic deployment discovery does not start it accidentally.
## Configuration
`renovate/renovate.json` is the single source of truth. The Compose file and the
`renovate-run` workflow mount that file directly.
A ConfigMap cannot read from the repository, so the CronJob needs the config
inlined. `renovate/k8s/configmap.yaml` is therefore a **generated** copy:
```sh
.gitea/workflows/sync-renovate-configmap.sh # regenerate after editing
.gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
```
The `renovate-ci` workflow runs the `--check` form on every PR and push, so a
config edit that forgets to regenerate the ConfigMap cannot be merged.
Beyond images, `customManagers` in the config track:
- Helm chart versions pinned in `.gitea/workflows/deploy-lib.sh`. The built-in
`helmv3` manager only reads `Chart.yaml` and `helm-values` only reads values
files, so neither sees a version written into a `helm upgrade` command —
these are declared as `custom.regex` managers against the `helm` datasource.
- CI linter versions in `.gitea/workflows/tool-versions.env`.
The Renovate image tag is deliberately _not_ in `tool-versions.env`:
`renovate/k8s/cronjob.yaml` owns it, and the workflows read it from there.
## How updates flow
Renovate scans both `compose.yaml` files and Kubernetes manifests, opens a
branch and PR with image tag changes, and waits for CI. After merge, the
existing deployment workflow applies Kubernetes changes or redeploys Compose
stacks. Renovate never updates running workloads directly.
+36
View File
@@ -29,6 +29,31 @@ data:
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
}, },
"customManagers": [ "customManagers": [
{
"customType": "regex",
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
"managerFilePatterns": ["edu_master/k8s/playwright.yaml", "edu_master/compose.yaml"],
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
"datasourceTemplate": "npm",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: PLAYWRIGHT_VERSION file",
"managerFilePatterns": ["edu_master/PLAYWRIGHT_VERSION"],
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n)?$"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: playwright Python client version pinned in Dockerfile ARG",
"managerFilePatterns": ["edu_master/webinar-checker/Dockerfile"],
"matchStrings": ["(?:^|\\n)ARG PLAYWRIGHT_VERSION=(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n|$)"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright",
"versioningTemplate": "pep440"
},
{ {
"customType": "regex", "customType": "regex",
"description": "kube-prometheus-stack chart version pinned in the deploy workflow", "description": "kube-prometheus-stack chart version pinned in the deploy workflow",
@@ -187,6 +212,17 @@ data:
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"], "matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
"enabled": false "enabled": false
}, },
{
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"groupName": "playwright singlesource",
"groupSlug": "playwright"
},
{
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"automerge": false
},
{ {
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file", "description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
"matchPackageNames": ["renovate/renovate"], "matchPackageNames": ["renovate/renovate"],
+1 -1
View File
@@ -19,7 +19,7 @@ spec:
restartPolicy: Never restartPolicy: Never
containers: containers:
- name: renovate - name: renovate
image: renovate/renovate:44.147.0 image: renovate/renovate:44.140.0
env: env:
- name: RENOVATE_PLATFORM - name: RENOVATE_PLATFORM
value: gitea value: gitea
+1 -1
View File
@@ -2,7 +2,7 @@ services:
renovate: renovate:
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update" # Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
# package rule in renovate/renovate.json. # package rule in renovate/renovate.json.
image: renovate/renovate:44.147.0 image: renovate/renovate:44.136.0
container_name: renovate container_name: renovate
restart: "no" restart: "no"
env_file: env_file:
+36
View File
@@ -18,6 +18,31 @@
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
}, },
"customManagers": [ "customManagers": [
{
"customType": "regex",
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
"managerFilePatterns": ["edu_master/k8s/playwright.yaml", "edu_master/compose.yaml"],
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
"datasourceTemplate": "npm",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: PLAYWRIGHT_VERSION file",
"managerFilePatterns": ["edu_master/PLAYWRIGHT_VERSION"],
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n)?$"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: playwright Python client version pinned in Dockerfile ARG",
"managerFilePatterns": ["edu_master/webinar-checker/Dockerfile"],
"matchStrings": ["(?:^|\\n)ARG PLAYWRIGHT_VERSION=(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n|$)"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright",
"versioningTemplate": "pep440"
},
{ {
"customType": "regex", "customType": "regex",
"description": "kube-prometheus-stack chart version pinned in the deploy workflow", "description": "kube-prometheus-stack chart version pinned in the deploy workflow",
@@ -176,6 +201,17 @@
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"], "matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
"enabled": false "enabled": false
}, },
{
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"groupName": "playwright singlesource",
"groupSlug": "playwright"
},
{
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"automerge": false
},
{ {
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file", "description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
"matchPackageNames": ["renovate/renovate"], "matchPackageNames": ["renovate/renovate"],
-22
View File
@@ -1,22 +0,0 @@
# SearXNG
Search frontend with a separate Valkey cache.
Kubernetes keeps the application settings in a ConfigMap and starts Valkey as a
StatefulSet. Set the secret from the example before exposing the search endpoint.
There is no active marker.
Compose expects local configuration under `core-config/`, which is ignored.
Prepare it before starting the stack; a container image alone does not supply
this lab's settings.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n searxng
kubectl get events -n searxng --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-18
View File
@@ -1,18 +0,0 @@
# Media stack
Docker services for playback, requests, library management, and downloads.
Compose runs Jellyfin, Jellyseerr, Sonarr, Radarr, Prowlarr, qBittorrent, and the
other services declared in the file. Kubernetes only routes to host endpoints;
update `k8s/routing/external-service.yaml` when the Docker host or ports change.
Prepare the paths, user/group IDs, and credentials from `.env.example`. Service
configuration and media/download directories are bind mounts. Preserve their
permissions when moving data, and keep the application databases with backups.
Review device mounts for hardware acceleration before starting on another host.
The Compose and Kubernetes routing files have no active markers, so automatic
deploys do not select this stack. Start the Compose project or apply its routing
resources manually when needed.
See the [repository README](../README.md) for deployment selection.
View File
Whitespace-only changes.
View File
Whitespace-only changes.
-20
View File
@@ -1,20 +0,0 @@
# Termix
Terminal and SSH connection manager with persistent application data.
Kubernetes stores state in `termix-pvc`; Compose mounts `termix-data/`.
The application config and routes are committed separately under `k8s/`.
There is no active marker. Review access control and retain the application data
needed to recover saved connections before enabling it.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n termix
kubectl get events -n termix --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
Loaded 100 of 111 files, more files were not shown because too many files have changed in this diff. Show more