Compare commits
35
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d85bdf5dbd | ||
|
|
3ea181e966 | ||
|
|
eb2f6f7d5d | ||
|
|
67d08fc33e | ||
|
|
f748a3c7aa | ||
|
|
8c8ff47241 | ||
|
|
ed2ba44bee | ||
|
|
bce653ebaf | ||
|
|
624ae84da1 | ||
|
|
f00c044f3d | ||
|
|
0e3035ed74 | ||
|
|
1d8eda6e5e | ||
|
|
fade5439c7 | ||
|
|
c4cbd87590 | ||
|
|
4c7c53e0f2 | ||
|
|
ba934265ac | ||
|
|
c7155808d9 | ||
|
|
fc4d64bdb2 | ||
|
|
a6f7fc6030 | ||
|
|
11de1d1468 | ||
|
|
64962d1a63 | ||
|
|
b08a0a927d | ||
|
|
8203ba1b0b | ||
|
|
48033b5495 | ||
|
|
73d2af73e5 | ||
|
|
64253e005e | ||
|
|
69accd1752 | ||
|
|
0cf4b08a95 | ||
|
|
86df5d9048 | ||
|
|
d7441bbbc2 | ||
|
|
2be089e048 | ||
|
|
83b2e68371 | ||
|
|
597f64cbb0 | ||
|
|
5c8bc15e60 | ||
|
|
3c4732ae20 |
No files matched your search
+38
-30
@@ -1,42 +1,50 @@
|
||||
# EDU ownership handoff
|
||||
|
||||
## Current status
|
||||
## Status
|
||||
|
||||
EDU PR #1 merged at 2026-10-07 08:04:30 UTC. Main release `5094952464ce315130839303985fd04d721bc1f2` passed CI run 1585 and deploy run 1586. The workstation checkout `/srv/edu-master` is at that SHA. The release changed the application image digests:
|
||||
The EDU ownership handoff is complete. The homelab repository no longer owns
|
||||
EDU workloads, images, routes, alerts, or deployment selection. The EDU
|
||||
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
|
||||
|
||||
- Session keeper: `sha256:1e59473bd40fe4c22622017d808a8927a68788275fe073dc23d718c44b2fd5dd`
|
||||
- Webinar checker: `sha256:987d9bf0770272766523ea5b94c7f3f849175d737551d46591ae55e058cf9f12`
|
||||
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
|
||||
build, deploy, rollback, verification, route-probe, and registry references.
|
||||
It also added the serial image build matrix for the homelab services. This
|
||||
handoff record is the only remaining EDU-specific file in homelab Git.
|
||||
|
||||
The live workloads remain healthy in context `Default`, namespace `edu-master`. Both health and live probes return 200. Redis AUTH passes, session TTL is 1178 seconds, the delivery backlog is zero, all nine EDU alert rules are healthy, and the scrape target is UP. The unauthorized-pod Redis check passed. The Redis PVC UID and Secret UID and values, including the Fernet key, match their pre-release state.
|
||||
The dedicated workstation checkout is `/srv/edu-master`, at release
|
||||
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
|
||||
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
|
||||
moved outside the homelab repository to
|
||||
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
|
||||
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
|
||||
checkout has no EDU marker or tracked EDU application/deployment files.
|
||||
`AUTODEPLOY=false` remains in place for homelab deployment.
|
||||
|
||||
The homelab EDU active marker was present after the EDU deployment. It was moved to the private snapshot as `homelab-k8s-active.marker` while holding `/tmp/homelab-apply.lock`. The homelab checkout at `/srv/homelab` is at `5f9354b` and has the tracked marker deletion. Its deploy preflight blocks a dirty checkout until this removal is reconciled. Preserve private ignored configuration when syncing that checkout.
|
||||
## Release evidence
|
||||
|
||||
The remaining homelab change is PR #105, branch `feat/edu-handoff-matrix`, based on `codex/ci-visible-checks`. Its eight protected checks passed. Renovate runs 1587 and 1588 passed. Image publishing was skipped for the PR. The EDU runtime changes are in PR #3 from `fix/handoff-runtime` to `main`; CI run 1589 is in progress. Those runtime changes have not been released.
|
||||
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
|
||||
all validation and both image builds. Deploy run 1653 passed for the exact main
|
||||
SHA above.
|
||||
|
||||
## Approval gate and next steps
|
||||
The workstation rollout completed for both Deployments. The deployment
|
||||
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
|
||||
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
|
||||
matching expressions and healthy evaluation.
|
||||
|
||||
PR #99 must merge before PR #105 can target `main`. A merge attempt for PR #99 returned HTTP 405 because it needs one approval; the protected branch has `required_approvals=1` and whitelist approval is enabled. This approval gate prevents the remaining transfer steps.
|
||||
The images now run by digest:
|
||||
|
||||
After the required approval:
|
||||
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
|
||||
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
|
||||
|
||||
1. Merge PR #99.
|
||||
2. Retarget PR #105 to `main`. Complete CI and review, then approve and merge it.
|
||||
3. Under the homelab apply lock, sync `/srv/homelab` to the merged removal. Preserve private ignored configuration and keep the active marker removed. Confirm the deploy preflight is clean.
|
||||
4. Merge the EDU runtime PR after its CI and review pass. The main-push CI run must complete successfully before its exact SHA can deploy.
|
||||
5. Verify the new release SHA, image digests, workload health, Redis AUTH and TTL, backlog, PVC and Secret identity, and monitoring. Record the results in the EDU PR.
|
||||
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
|
||||
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
|
||||
runtime Secret and Fernet key were preserved during the handoff. Notification
|
||||
delivery was verified before closeout, as confirmed by the operator. The
|
||||
deployment did not record downtime.
|
||||
|
||||
`AUTODEPLOY=false` is explicitly configured. The EDU repository path and port secrets are confirmed, and `EDU_KUBE_CONTEXT=Default` is configured as a repository variable. Keep deployment and registry credentials outside Git. Never run both homelab and EDU deployment paths at the same time.
|
||||
|
||||
## Change summary
|
||||
|
||||
The homelab PR removes the EDU subtree, deployment and image selection, rollback and verification cases, route probes, Renovate references, and external-image exceptions. It adds a serial dynamic matrix for the three homelab images. Each job builds an image or reuses a matching immutable digest. The final job checks all image results and publishes full-SHA tags and the existing release artifact only after they pass. PRs do not publish images. The protected check names from PR #99 are preserved. PR #100's service-metrics work is independent of this handoff.
|
||||
|
||||
The EDU runtime PR adds the Playwright service manifest, reconciles Redis storage and Secret reload annotations, and adds pre-apply Redis backup and identity checks. It verifies application endpoints, Redis AUTH, session TTL, metrics, and all nine vmalert rules. Rollback checks workload and application health and reports when manual recovery is needed. Its deployment guard rejects an unexpected or dirty checkout and refuses deployment while either legacy homelab EDU marker exists.
|
||||
|
||||
## Rollback and limits
|
||||
|
||||
The private snapshot is `/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z` on the workstation. It contains the pre-handoff Redis RDB and recovery data. RDB checksum verification confirmed twelve keys. Keep the snapshot outside Git. For an EDU release failure, restore the saved Kubernetes resources and inspect application health. The rollback does not automatically restore the Redis RDB; restore old Redis data only when recovery requires it.
|
||||
|
||||
For an ownership rollback, stop EDU deployment triggers first, restore the reviewed homelab source and marker, then reapply recorded immutable image digests. Verify both workload and application health. Never delete or recreate the Redis PVC.
|
||||
|
||||
The initial EDU release and the homelab marker move are complete. PR #99 approval and merge, PR #105 retarget and merge, homelab checkout reconciliation, EDU runtime PR merge, and release of those runtime changes remain pending. Synthetic Telegram delivery and Alertmanager-to-Telegram notification were not tested.
|
||||
The release rollback snapshot is
|
||||
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
|
||||
The handoff data snapshot remains at
|
||||
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
|
||||
Both snapshots are outside Git. Do not restore old Redis data unless recovery
|
||||
requires it. Never delete or recreate the Redis PVC.
|
||||
@@ -0,0 +1,64 @@
|
||||
# CI and deployment
|
||||
|
||||
Gitea Actions validates changes, builds the repository's custom images, and can
|
||||
deploy selected services to the workstation. CI and production deployment use
|
||||
separate workflows. See the [runner and recovery guide](runner/README.md) for
|
||||
installation, configuration, and operator commands.
|
||||
|
||||
## CI
|
||||
|
||||
`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit
|
||||
test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs
|
||||
use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool
|
||||
versions are pinned in `workflows/tool-versions.env`.
|
||||
|
||||
Compose CI checks every committed Compose file without requiring ignored `.env`
|
||||
files. Kubernetes checks validate known schemas; unknown CRDs are skipped.
|
||||
|
||||
On main, CI plans builds for the three owned images: `error-pages`,
|
||||
`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a
|
||||
digest from a successful earlier main run. The successful build job publishes a
|
||||
release artifact for the exact commit SHA. Pull requests do not publish images.
|
||||
|
||||
## Deployment gate
|
||||
|
||||
`workflows/deploy.yaml` starts a deployment after successful main CI when the
|
||||
`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the
|
||||
requested `main` ref or commit must have successful main CI and its matching
|
||||
release artifact. A manual dispatch does not bypass validation.
|
||||
|
||||
The workflow supports these modes:
|
||||
|
||||
- `changed`: select active services changed since the last successful deploy.
|
||||
- `full`: select all active services; use this for the first baseline.
|
||||
- `plan`: validate and show the selection without applying production resources.
|
||||
|
||||
`refresh_images=true` explicitly refreshes mutable third-party Compose tags.
|
||||
|
||||
## Selection and rollout
|
||||
|
||||
The active markers define automatic deployment. `<service>/active` selects a
|
||||
standard Compose file; `<service>/k8s/active` selects Kubernetes resources. Helm
|
||||
releases have their own markers in `workflows/deploy-lib.sh`. Service
|
||||
dependencies are declared in `deploy-dependencies.json`. Removed resources are
|
||||
reported for manual review; the workflow does not prune them automatically.
|
||||
|
||||
The workstation controller runs the checked source in a per-SHA worktree. It
|
||||
validates configuration, applies Kubernetes and Compose changes in sequence,
|
||||
verifies changed Kubernetes workloads, and checks public routes. A durable
|
||||
systemd service continues the rollout if the Actions SSH client disconnects.
|
||||
The workflow checks the exact CI release before it submits a deployment.
|
||||
|
||||
Kubernetes recovery uses captured workload revisions. It does not restore
|
||||
ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is
|
||||
manual and does not restore volume data or reverse migrations. Keep backups for
|
||||
stateful services. The runner guide documents status, retry, logs, and recovery
|
||||
commands.
|
||||
|
||||
## Settings
|
||||
|
||||
Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified
|
||||
`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`,
|
||||
`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation
|
||||
also needs its existing registry authentication. Set `AUTODEPLOY=false` until
|
||||
automatic production deploys are intended.
|
||||
@@ -7,4 +7,5 @@ self-hosted-runner:
|
||||
labels:
|
||||
- arch
|
||||
- homelab
|
||||
- homelab-pr
|
||||
- prod
|
||||
+57
-6
@@ -1,17 +1,20 @@
|
||||
# Homelab CI/CD
|
||||
|
||||
The native Gitea runner runs on **vps**; production runs on **workstation**.
|
||||
Compose, workflow, shell, Python, formatting, YAML, Dockerfile and Kubernetes
|
||||
checks appear as separate jobs. Jobs run on `homelab:host`, one at a time; the
|
||||
build waits for every check to pass. CI and deploy runs also show a summary with
|
||||
The native Gitea runners run on **vps**; production runs on **workstation**.
|
||||
Main-branch checks and image builds use `homelab:host`. Pull request and
|
||||
non-main checks use `homelab-pr:host` under a separate account without Docker
|
||||
access. The `homelab-pr` runner is registered at User scope for `forust`, so
|
||||
any repository under that account can schedule jobs that request this label.
|
||||
Each runner accepts one job at a time; the build waits for every check to pass.
|
||||
CI and deploy runs also show a summary with
|
||||
the release SHA, image build or reuse results, deploy mode, selected services,
|
||||
and image digests. Failed runs keep a summary of completed image builds, stage
|
||||
results, apply results, and recorded Kubernetes recovery. The final deploy
|
||||
summary is in the smoke job; earlier jobs show the state observed at that time.
|
||||
Apply success is separate from health and recovery. Update the installed
|
||||
workstation controller with `setup-workstation.sh` when no deploy is running.
|
||||
No job images or Kubernetes credentials
|
||||
are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
|
||||
No job images or Kubernetes credentials are needed on the VPS. Builds use one
|
||||
pinned BuildKit helper container. CI and deploy are separate workflows.
|
||||
|
||||
## Runner installation
|
||||
|
||||
@@ -37,6 +40,32 @@ pushes directly to the registry, and caps retained local cache at 1 GiB with a
|
||||
2 GiB free-space target. This is not a hard limit on peak build disk usage.
|
||||
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
|
||||
|
||||
### Pull request runner
|
||||
|
||||
Install the unprivileged host runner on the VPS:
|
||||
|
||||
```sh
|
||||
sudo bash .gitea/runner/setup-pr-runner.sh
|
||||
```
|
||||
|
||||
Get a registration token from the user Actions runner settings. Run the
|
||||
installer in a terminal. It asks for the token without echoing it, registers the
|
||||
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
|
||||
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
|
||||
runner as User scope before merging the workflow change. An unmatched label can
|
||||
fall back to the default job image.
|
||||
|
||||
Renovate PR validation uses `pull_request_target`, which reads the workflow from
|
||||
the base branch. It checks out the PR head only after runner selection and runs
|
||||
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
|
||||
|
||||
The PR runner has a separate home and tool cache. Do not add it to the `docker`
|
||||
group or give it access to `/var/run/docker.sock`. It runs repository code from
|
||||
pull requests, so keep its registration and permissions separate from the
|
||||
trusted `homelab` runner. This separates users and host permissions, but both
|
||||
runners still share the VPS kernel and network. Use a disposable VM if PRs from
|
||||
untrusted external authors must be fully isolated.
|
||||
|
||||
## Workstation setup
|
||||
|
||||
As the existing SSH deploy user on workstation:
|
||||
@@ -128,3 +157,25 @@ run first. Restore the runner config/unit from `.before-<timestamp>` backups,
|
||||
reload systemd and restart the runner. Restore the prior workflows from Git.
|
||||
Production data and persistent volumes stay where they were. Do not remove run
|
||||
state or Compose recovery files until recovery is confirmed.
|
||||
|
||||
### Compose configuration recovery
|
||||
|
||||
Successful deploys save the complete resolved Compose configuration in
|
||||
`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets.
|
||||
Keep them private and do not commit or upload them.
|
||||
The next deploy uses this configuration for its recovery file, including old
|
||||
commands, environment, mounts, ports, and removed services. The recovery command
|
||||
uses `--remove-orphans` to remove services added by the failed deploy. It does
|
||||
not restore volume data or reverse database migrations.
|
||||
|
||||
On the first run after this update, the controller can use the Compose file
|
||||
from the previous successful run. If that file is absent, it reads the persistent
|
||||
checkout and checks its service configuration hashes against existing containers.
|
||||
A mismatch stops preflight. Restore the previous configuration before retrying.
|
||||
Update the installed controller with `bash .gitea/runner/setup-workstation.sh`
|
||||
from the reviewed checkout before using this change.
|
||||
|
||||
New namespaces are checked during preflight. Server validation of their resources
|
||||
runs after namespace creation and before application resources are applied.
|
||||
Plan mode does not create namespaces. A failed deferred check can leave an empty
|
||||
namespace; inspect it before removing it.
|
||||
@@ -0,0 +1,8 @@
|
||||
runner:
|
||||
file: /var/lib/gitea-pr-runner/.runner
|
||||
capacity: 1
|
||||
timeout: 5h
|
||||
labels:
|
||||
- homelab-pr:host
|
||||
cache:
|
||||
enabled: false
|
||||
@@ -0,0 +1,27 @@
|
||||
[Unit]
|
||||
Description=Gitea Actions untrusted pull request runner
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
User=gitea-pr-runner
|
||||
Group=gitea-pr-runner
|
||||
WorkingDirectory=/var/lib/gitea-pr-runner
|
||||
Environment=HOME=/var/lib/gitea-pr-runner
|
||||
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
||||
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
NoNewPrivileges=yes
|
||||
PrivateTmp=yes
|
||||
ProtectSystem=full
|
||||
ProtectHome=yes
|
||||
ProtectKernelTunables=yes
|
||||
ProtectKernelModules=yes
|
||||
ProtectControlGroups=yes
|
||||
RestrictSUIDSGID=yes
|
||||
LockPersonality=yes
|
||||
UMask=0077
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
Executable
+56
@@ -0,0 +1,56 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install a native runner for untrusted PR jobs without Docker access.
|
||||
set -euo pipefail
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
||||
for tool in cp cut date getent id install runuser systemctl useradd; do
|
||||
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
||||
done
|
||||
command -v /usr/local/bin/gitea-runner >/dev/null || {
|
||||
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
id gitea-pr-runner >/dev/null 2>&1 || \
|
||||
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
|
||||
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
|
||||
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
|
||||
echo 'Unexpected PR runner home; inspect the existing service first' >&2
|
||||
exit 1
|
||||
}
|
||||
case " $(id -nG gitea-pr-runner) " in
|
||||
*' docker '*)
|
||||
echo 'The PR runner account must not belong to the docker group' >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
install -d -m 0755 /etc/gitea-pr-runner
|
||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
|
||||
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
||||
done
|
||||
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
|
||||
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
|
||||
|
||||
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
|
||||
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
|
||||
printf '\n'
|
||||
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
|
||||
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
|
||||
unset runner_token
|
||||
runuser --preserve-environment -u gitea-pr-runner -- \
|
||||
/usr/local/bin/gitea-runner register \
|
||||
--config /etc/gitea-pr-runner/config.yaml \
|
||||
--instance https://gitea.forust.xyz \
|
||||
--name homelab-pr \
|
||||
--labels homelab-pr:host \
|
||||
--no-interactive
|
||||
unset GITEA_RUNNER_REGISTRATION_TOKEN
|
||||
fi
|
||||
chmod 0600 /var/lib/gitea-pr-runner/.runner
|
||||
|
||||
systemctl daemon-reload
|
||||
systemctl enable --now gitea-pr-runner.service
|
||||
systemctl restart gitea-pr-runner.service
|
||||
echo "PR runner ready. Configuration backups: *.before-$stamp"
|
||||
@@ -91,4 +91,66 @@ if check_referenced_secrets >"$scratch/secrets.log"; then
|
||||
echo 'Secret check accepted a failed manifest render' >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# New declared namespaces defer only their own resources during preflight.
|
||||
render_selected_resources() {
|
||||
cat <<'JSON'
|
||||
{"apiVersion":"v1","kind":"List","items":[
|
||||
{"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}},
|
||||
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}},
|
||||
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}}
|
||||
]}
|
||||
JSON
|
||||
}
|
||||
kubectl() {
|
||||
case "$1" in
|
||||
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;;
|
||||
apply) cat >"$scratch/server-input.json" ;;
|
||||
*) return 1 ;;
|
||||
esac
|
||||
}
|
||||
validate_server_resources true
|
||||
jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null
|
||||
if validate_server_resources false 2>"$scratch/deferred.log"; then
|
||||
echo 'Post-namespace validation accepted a missing namespace' >&2
|
||||
exit 1
|
||||
fi
|
||||
kubectl() {
|
||||
case "$1" in
|
||||
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;;
|
||||
apply) cat >"$scratch/server-input.json" ;;
|
||||
*) return 1 ;;
|
||||
esac
|
||||
}
|
||||
validate_server_resources false
|
||||
jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null
|
||||
render_selected_resources() {
|
||||
printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}'
|
||||
}
|
||||
if validate_server_resources true 2>"$scratch/undeclared.log"; then
|
||||
echo 'Preflight accepted an undeclared missing namespace' >&2
|
||||
exit 1
|
||||
fi
|
||||
# Count services, not characters in the newline-separated service names.
|
||||
compose() {
|
||||
case "$*" in
|
||||
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;;
|
||||
*'ps --status running --services') printf '%s\n' headscale headplane web ;;
|
||||
*) return 1 ;;
|
||||
esac
|
||||
}
|
||||
verify_compose_stack example.yaml >"$scratch/compose-count.log"
|
||||
grep -qF 'all 3 service(s) running' "$scratch/compose-count.log"
|
||||
compose() {
|
||||
case "$*" in
|
||||
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;;
|
||||
*'ps --status running --services') printf '%s\n' headscale headplane ;;
|
||||
*) return 0 ;;
|
||||
esac
|
||||
}
|
||||
if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then
|
||||
echo 'Compose verification accepted a missing service' >&2
|
||||
exit 1
|
||||
fi
|
||||
grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log"
|
||||
printf '%s\n' 'Deploy validation regressions passed.'
|
||||
+13
-13
@@ -14,7 +14,7 @@ concurrency:
|
||||
jobs:
|
||||
compose:
|
||||
name: Compose
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -66,7 +66,7 @@ jobs:
|
||||
fi
|
||||
workflows:
|
||||
name: Workflows
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -102,7 +102,7 @@ jobs:
|
||||
fi
|
||||
shell:
|
||||
name: Shell
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -146,7 +146,7 @@ jobs:
|
||||
fi
|
||||
formatting:
|
||||
name: Formatting
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -194,7 +194,7 @@ jobs:
|
||||
fi
|
||||
python:
|
||||
name: Python and tests
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -232,7 +232,7 @@ jobs:
|
||||
fi
|
||||
yaml:
|
||||
name: YAML
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -280,7 +280,7 @@ jobs:
|
||||
fi
|
||||
dockerfiles:
|
||||
name: Dockerfiles
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -326,7 +326,7 @@ jobs:
|
||||
fi
|
||||
kubernetes:
|
||||
name: Kubernetes
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -395,7 +395,7 @@ jobs:
|
||||
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
|
||||
- name: Store the image plan
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: build-plan
|
||||
path: build-plan.json
|
||||
@@ -435,7 +435,7 @@ jobs:
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
- name: Download the checked image plan
|
||||
id: inputs
|
||||
uses: actions/download-artifact@9bc31d5ccc31df68ecc42ccf4149144866c47d8a
|
||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||
with:
|
||||
name: build-plan
|
||||
- name: Build or reuse this image
|
||||
@@ -447,7 +447,7 @@ jobs:
|
||||
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
|
||||
- name: Store the image result
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: image-${{ matrix.name }}
|
||||
path: image.json
|
||||
@@ -482,7 +482,7 @@ jobs:
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
- name: Download all image results
|
||||
id: inputs
|
||||
uses: actions/download-artifact@9bc31d5ccc31df68ecc42ccf4149144866c47d8a
|
||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||
with:
|
||||
path: artifacts
|
||||
- name: Pin SHA tags and write the complete release
|
||||
@@ -495,7 +495,7 @@ jobs:
|
||||
--plan artifacts/build-plan/build-plan.json
|
||||
- name: Store commit release
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: release-${{ github.sha }}
|
||||
path: release.json
|
||||
|
||||
@@ -42,15 +42,50 @@ def prepare(source_file):
|
||||
images_file = directory / 'compose-images.json'
|
||||
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
|
||||
release = json.loads((directory / 'release.json').read_text())
|
||||
before = json.loads(json.dumps(config))
|
||||
state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
|
||||
baseline = state / 'compose-configs' / f'{relative.parent.name}.json'
|
||||
if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')):
|
||||
baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name
|
||||
bootstrap = not baseline.exists()
|
||||
if not bootstrap:
|
||||
before = json.loads(baseline.read_text())
|
||||
else:
|
||||
# Bootstrap from the persistent configuration, never from the new source.
|
||||
persistent_file = config_repo / relative
|
||||
if persistent_file.exists():
|
||||
before = json.loads(
|
||||
output(
|
||||
'docker',
|
||||
'compose',
|
||||
'--project-directory',
|
||||
str(project_dir),
|
||||
'-f',
|
||||
str(persistent_file),
|
||||
'config',
|
||||
'--format',
|
||||
'json',
|
||||
cwd=config_repo,
|
||||
)
|
||||
)
|
||||
elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'):
|
||||
raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy')
|
||||
else:
|
||||
before = {'name': project, 'services': {}}
|
||||
if before['name'] != project:
|
||||
raise ValueError('Compose project name changed; manual migration is required')
|
||||
for service, settings in config['services'].items():
|
||||
reference = settings.get('image')
|
||||
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
||||
if not reference or settings.get('build'):
|
||||
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
|
||||
image_repo = reference.split('@')[0].rsplit('/', 1)
|
||||
image_repo[-1] = image_repo[-1].split(':')[0]
|
||||
image_repo = '/'.join(image_repo)
|
||||
if image_repo in release['images']:
|
||||
# Nextcloud AIO validates the mastercontainer image reference and rejects
|
||||
# a digest. Keep its configured tag so AIO can start and manage its stack.
|
||||
if nextcloud_aio_master:
|
||||
pinned = reference
|
||||
elif image_repo in release['images']:
|
||||
pinned = image_repo + '@' + release['images'][image_repo]
|
||||
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
|
||||
pinned = locks[reference]
|
||||
@@ -58,6 +93,12 @@ def prepare(source_file):
|
||||
pinned = resolve(reference)
|
||||
settings['image'] = pinned
|
||||
locks[reference] = pinned
|
||||
for service, settings in before['services'].items():
|
||||
reference = settings['image']
|
||||
image_repo = reference.split('@')[0].rsplit('/', 1)
|
||||
image_repo[-1] = image_repo[-1].split(':')[0]
|
||||
image_repo = '/'.join(image_repo)
|
||||
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
||||
# Capture what is running, not the current value of its mutable tag.
|
||||
ids = output(
|
||||
'docker',
|
||||
@@ -69,12 +110,42 @@ def prepare(source_file):
|
||||
f'label=com.docker.compose.service={service}',
|
||||
).splitlines()
|
||||
actual = set()
|
||||
if bootstrap and ids:
|
||||
expected_hash = output(
|
||||
'docker',
|
||||
'compose',
|
||||
'--project-directory',
|
||||
str(project_dir),
|
||||
'-f',
|
||||
str(persistent_file),
|
||||
'config',
|
||||
'--hash',
|
||||
service,
|
||||
cwd=config_repo,
|
||||
).split()[-1]
|
||||
for container in ids:
|
||||
running_hash = output(
|
||||
'docker',
|
||||
'inspect',
|
||||
container,
|
||||
'--format',
|
||||
'{{ index .Config.Labels "com.docker.compose.config-hash" }}',
|
||||
)
|
||||
if running_hash != expected_hash:
|
||||
raise ValueError(
|
||||
f'{project}/{service}: persistent config differs from running config; restore the previous config'
|
||||
)
|
||||
for container in ids:
|
||||
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
|
||||
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
|
||||
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
|
||||
if len(actual) > 1:
|
||||
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
|
||||
# AIO also rejects a digest in its recovery config. Preserve its tag in
|
||||
# both deploy and recovery files.
|
||||
if nextcloud_aio_master:
|
||||
before['services'][service]['image'] = reference
|
||||
else:
|
||||
before['services'][service]['image'] = next(iter(actual)) if actual else reference
|
||||
for name, data in (('compose', config), ('compose-before', before)):
|
||||
folder = directory / name
|
||||
@@ -85,7 +156,7 @@ def prepare(source_file):
|
||||
images_file.write_text(json.dumps(locks, indent=2) + '\n')
|
||||
print(f'Compose {project}: images pinned; local paths preserved')
|
||||
print(
|
||||
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
|
||||
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans'
|
||||
)
|
||||
|
||||
|
||||
|
||||
@@ -141,7 +141,8 @@ def make_plan(directory):
|
||||
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
|
||||
request = json.loads((directory / 'request.json').read_text())
|
||||
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
|
||||
helm = json.loads(command('helm', 'list', '--all', '-A', '-o', 'json'))
|
||||
# Helm 4 lists every release status by default and removed the --all flag.
|
||||
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
|
||||
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
|
||||
if request['refresh_images']:
|
||||
plan['selected']['compose'] = plan['active']['compose']
|
||||
@@ -164,6 +165,10 @@ def finish_success(directory, plan):
|
||||
if previous.exists()
|
||||
else {}
|
||||
)
|
||||
configs = STATE / 'compose-configs'
|
||||
configs.mkdir(mode=0o700, exist_ok=True)
|
||||
for config in (directory / 'compose').glob('*.json'):
|
||||
atomic_json(configs / config.name, json.loads(config.read_text()))
|
||||
atomic_json(STATE / 'last-success.json', plan)
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['state'] = 'success'
|
||||
|
||||
@@ -180,7 +180,8 @@ save_snapshot() {
|
||||
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
|
||||
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
|
||||
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
|
||||
releases="$(helm list --all -A -o json)" || return 1
|
||||
# Helm 4 lists every release status by default and removed the --all flag.
|
||||
releases="$(helm list -A -o json)" || return 1
|
||||
for entry in "${HELM_RELEASES[@]}"; do
|
||||
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
|
||||
if ! jq -e --arg r "$release" --arg n "$namespace" \
|
||||
@@ -329,11 +330,11 @@ rollback_workloads() {
|
||||
# written straight into a `helm upgrade` command would never be updated: these
|
||||
# have to be declared as custom.regex managers in renovate/renovate.json.
|
||||
HELM_RELEASES=(
|
||||
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
||||
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.3.2|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
||||
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
|
||||
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
|
||||
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
|
||||
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
||||
"reloader|stakater/reloader|reloader|2.2.18|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
||||
)
|
||||
|
||||
# "name url" for the Helm repository hosting a chart, empty if unknown.
|
||||
@@ -570,6 +571,41 @@ skip_uninstalled_vmagent_crd() {
|
||||
return 1
|
||||
}
|
||||
|
||||
# Render one complete resource list so new namespaces can be identified across
|
||||
# files and Kustomize apps. A missing undeclared namespace remains an error.
|
||||
render_selected_resources() {
|
||||
local m k
|
||||
{
|
||||
for m in "${K8S_MANIFESTS[@]}"; do
|
||||
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
|
||||
kubectl create --dry-run=client --validate=false -f "$m" -o json || return 1
|
||||
done
|
||||
for k in "${KUSTOMIZE_APPS[@]}"; do
|
||||
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json || return 1
|
||||
done
|
||||
} | jq -s '{apiVersion: "v1", kind: "List", items: [ .[] | if .kind == "List" then .items[] else . end ]}'
|
||||
}
|
||||
|
||||
validate_server_resources() {
|
||||
local defer_new="$1" resources existing filtered
|
||||
resources="$(render_selected_resources)" || return 1
|
||||
existing="$(kubectl get namespaces -o json)" || return 1
|
||||
filtered="$(jq --argjson existing "$existing" --argjson defer "$defer_new" '
|
||||
[.items[] | select(.kind == "Namespace") | .metadata.name] as $declared
|
||||
| [$existing.items[].metadata.name] as $present
|
||||
| .items |= map(
|
||||
(.metadata.namespace // "default") as $ns
|
||||
| if .kind == "Namespace" or ($present | index($ns)) != null then .
|
||||
elif ($declared | index($ns)) == null then error("Undeclared missing namespace: " + $ns)
|
||||
elif $defer then empty
|
||||
else error("Namespace still missing after namespace apply: " + $ns)
|
||||
end)
|
||||
' <<<"$resources")" || return 1
|
||||
if [ "$(jq '.items | length' <<<"$filtered")" -gt 0 ]; then
|
||||
kubectl apply --dry-run=server -f - <<<"$filtered" >/dev/null
|
||||
fi
|
||||
}
|
||||
|
||||
stage_validate() {
|
||||
check_prune_mode || return 1
|
||||
cd "$REPO"
|
||||
@@ -596,15 +632,7 @@ stage_validate() {
|
||||
kubectl apply -k "$k" --dry-run=client >/dev/null
|
||||
done
|
||||
log "Validate k8s manifests (kubectl dry-run=server)"
|
||||
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||
if skip_uninstalled_vmagent_crd "$m"; then
|
||||
continue
|
||||
fi
|
||||
kubectl apply --dry-run=server -f "$m" >/dev/null
|
||||
done
|
||||
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||
kubectl apply -k "$k" --dry-run=server >/dev/null
|
||||
done
|
||||
validate_server_resources true
|
||||
log "Checking referenced Secrets exist"
|
||||
echo " (deploy never applies *secret*.yaml; create missing ones manually)"
|
||||
check_referenced_secrets
|
||||
@@ -656,6 +684,14 @@ stage_apply_k8s() {
|
||||
record_apply kubectl "${m#"$REPO"/}" success
|
||||
done
|
||||
fi
|
||||
# Kustomize may declare namespaces inside its rendered resources too.
|
||||
local namespace_resources
|
||||
namespace_resources="$(render_selected_resources | jq '.items |= map(select(.kind == "Namespace"))')" || return 1
|
||||
if [ "$(jq '.items | length' <<<"$namespace_resources")" -gt 0 ]; then
|
||||
kubectl apply -f - <<<"$namespace_resources" || return 1
|
||||
fi
|
||||
# Complete the deferred server checks before Helm or application resources change.
|
||||
validate_server_resources false || return 1
|
||||
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
|
||||
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
|
||||
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
|
||||
@@ -791,12 +827,13 @@ stage_verify_k8s() {
|
||||
# actually be running.
|
||||
verify_compose_stack() {
|
||||
local cf="$1"
|
||||
local expected running missing=()
|
||||
local expected running svc missing=() service_count=0
|
||||
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
|
||||
running="$(compose "$cf" ps --status running --services | sort)" || return 1
|
||||
[ -n "$expected" ] || return 0
|
||||
while IFS= read -r svc; do
|
||||
[ -n "$svc" ] || continue
|
||||
service_count=$((service_count + 1))
|
||||
# restart:"no" services are allowed to have exited.
|
||||
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
|
||||
missing+=("$svc")
|
||||
@@ -807,7 +844,7 @@ verify_compose_stack() {
|
||||
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
|
||||
return 1
|
||||
fi
|
||||
echo " all ${#expected} service(s) running"
|
||||
echo " all $service_count service(s) running"
|
||||
return 0
|
||||
}
|
||||
|
||||
|
||||
@@ -83,20 +83,20 @@ def make_plan(repo, config_repo, release, previous, mode, live_helm):
|
||||
removed = []
|
||||
else:
|
||||
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
|
||||
changed = {path.split('/')[0] for path in paths}
|
||||
changed = {service for service in all_services for path in paths if path.startswith(service + '/')}
|
||||
if any(path.startswith('.gitea/') for path in paths):
|
||||
changed |= all_services
|
||||
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
|
||||
for file in tracked(repo):
|
||||
service = file.split('/')[0]
|
||||
if service not in all_services or not file.endswith(('.yaml', '.yml')):
|
||||
owners = {service for service in all_services if file.startswith(service + '/')}
|
||||
if not owners or not file.endswith(('.yaml', '.yml')):
|
||||
continue
|
||||
text = (repo / file).read_text()
|
||||
if any(
|
||||
image in text and previous.get('images', {}).get(image) != digest
|
||||
for image, digest in release['images'].items()
|
||||
):
|
||||
changed.add(service)
|
||||
changed |= owners
|
||||
removed = sorted(
|
||||
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
|
||||
- all_services
|
||||
|
||||
@@ -80,7 +80,7 @@ jobs:
|
||||
apply:
|
||||
needs: [gate]
|
||||
runs-on: homelab
|
||||
timeout-minutes: 100
|
||||
timeout-minutes: 120
|
||||
steps:
|
||||
- name: Checkout checked commit
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
@@ -305,7 +305,6 @@ def build_images(output, report, name, plan):
|
||||
if exists:
|
||||
print(f'Reuse {name}: inputs unchanged')
|
||||
digest = old_digest
|
||||
report['reused'].append(name)
|
||||
else:
|
||||
print(f'Build {name}', flush=True)
|
||||
metadata = Path(docker_config) / 'metadata.json'
|
||||
@@ -332,14 +331,14 @@ def build_images(output, report, name, plan):
|
||||
env=env,
|
||||
)
|
||||
digest = json.loads(metadata.read_text())['containerimage.digest']
|
||||
report['built'].append(name)
|
||||
if not isinstance(digest, str) or not DIGEST.fullmatch(digest):
|
||||
raise ValueError('Image job returned an invalid digest')
|
||||
release['images'][image] = digest
|
||||
release['inputs'][image] = inputs
|
||||
if not DIGEST.fullmatch(digest):
|
||||
raise ValueError('Image job returned an invalid digest')
|
||||
report['reused' if exists else 'built'].append(name)
|
||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||
report['current'] = None
|
||||
report['phase'] = 'Release file saved'
|
||||
report['phase'] = 'Image result file saved'
|
||||
finally:
|
||||
# Cleanup errors must neither leak credentials nor mask the original build error.
|
||||
try:
|
||||
@@ -395,13 +394,17 @@ def build(output, name, plan):
|
||||
result = 'success'
|
||||
finally:
|
||||
lines = [
|
||||
f'## Image release `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||
f'## Image build result `{name}`',
|
||||
'',
|
||||
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||
'',
|
||||
f'- Result: **{result}**',
|
||||
f'- Last stage: {report["phase"]}',
|
||||
]
|
||||
if result == 'failure':
|
||||
lines.append('- No release from this build can be deployed. Open the failed step log.')
|
||||
lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.')
|
||||
if result == 'success':
|
||||
lines.append('- This is one image result. The final build job must publish the complete release.')
|
||||
if report['current']:
|
||||
lines.append(f'- Image at the failure: `{report["current"]}`')
|
||||
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
|
||||
|
||||
@@ -1,7 +1,9 @@
|
||||
name: renovate-ci
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
# Read the workflow from the trusted base branch. PR code runs only on the
|
||||
# unprivileged runner selected below.
|
||||
pull_request_target:
|
||||
paths:
|
||||
- "renovate/**"
|
||||
- ".gitea/workflows/renovate-ci.yaml"
|
||||
@@ -26,37 +28,47 @@ permissions:
|
||||
|
||||
jobs:
|
||||
validate-renovate:
|
||||
runs-on: homelab
|
||||
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
|
||||
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
|
||||
# so the same version that runs in the cluster is the one validated here.
|
||||
- name: Resolve the deployed Renovate image
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the version.
|
||||
- name: Resolve the deployed Renovate version
|
||||
id: image
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||
renovate/k8s/cronjob.yaml | head -1)"
|
||||
if [ -z "$image" ]; then
|
||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
|
||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||
exit 1
|
||||
fi
|
||||
echo "using $image"
|
||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||
version="${BASH_REMATCH[1]}"
|
||||
echo "using Renovate $version"
|
||||
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Validate Renovate repository config
|
||||
- name: Prepare pinned validation tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
docker run --rm \
|
||||
-v "$PWD/renovate:/opt/renovate:ro" \
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
"${{ steps.image.outputs.image }}" \
|
||||
renovate-config-validator /opt/renovate/renovate.json
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
|
||||
- name: Validate Renovate repository config
|
||||
shell: bash
|
||||
env:
|
||||
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
|
||||
trap 'rm -rf "$npm_cache"' EXIT
|
||||
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
|
||||
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
|
||||
|
||||
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
||||
# carries an inlined copy of the config. Fail if it no longer matches.
|
||||
@@ -70,8 +82,6 @@ jobs:
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
kubeconform \
|
||||
-strict \
|
||||
-ignore-missing-schemas \
|
||||
|
||||
@@ -32,11 +32,14 @@ concurrency:
|
||||
|
||||
jobs:
|
||||
run-renovate:
|
||||
if: github.ref == 'refs/heads/main'
|
||||
runs-on: homelab
|
||||
timeout-minutes: 60
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: refs/heads/main
|
||||
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
||||
# Reading it here means this workflow validates and runs the exact version
|
||||
@@ -48,21 +51,23 @@ jobs:
|
||||
set -euo pipefail
|
||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||
renovate/k8s/cronjob.yaml | head -1)"
|
||||
if [ -z "$image" ]; then
|
||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
|
||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||
exit 1
|
||||
fi
|
||||
echo "using $image"
|
||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Validate Renovate config
|
||||
shell: bash
|
||||
env:
|
||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
docker run --rm \
|
||||
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
"${{ steps.image.outputs.image }}" \
|
||||
"$RENOVATE_IMAGE" \
|
||||
renovate-config-validator
|
||||
|
||||
- name: Run Renovate
|
||||
@@ -73,6 +78,7 @@ jobs:
|
||||
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
||||
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
||||
LOG_LEVEL: ${{ inputs.log_level }}
|
||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
@@ -89,4 +95,4 @@ jobs:
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
||||
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
||||
"${{ steps.image.outputs.image }}"
|
||||
"$RENOVATE_IMAGE"
|
||||
@@ -14,7 +14,7 @@ ACTIONLINT_VERSION="1.7.7"
|
||||
SHELLCHECK_VERSION="0.11.0"
|
||||
KUBECONFORM_VERSION="0.8.0"
|
||||
PRETTIER_VERSION="3.8.1"
|
||||
RUFF_VERSION="0.16.8"
|
||||
RUFF_VERSION="0.16.10"
|
||||
YAMLLINT_VERSION="1.38.0"
|
||||
HADOLINT_VERSION="2.14.0"
|
||||
# pip-audit reads the advisory database over the network, so a floating version
|
||||
|
||||
@@ -0,0 +1,163 @@
|
||||
# Homelab
|
||||
|
||||
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
|
||||
Gitea Actions that build and deploy them. Most applications have both deployment
|
||||
formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes
|
||||
ingress; the media stack has Compose and Kubernetes routing configuration.
|
||||
|
||||
These files contain this lab's domains, IP addresses, storage paths, and private
|
||||
registry names. Running them on another machine takes some editing.
|
||||
|
||||
## Start here
|
||||
|
||||
- [Service list](#services) — what each directory contains.
|
||||
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
|
||||
- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status.
|
||||
- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository.
|
||||
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
|
||||
[cert-manager](cert-manager/README.md) — common dependencies.
|
||||
|
||||
## What gets deployed
|
||||
|
||||
The `active` files are switches for the deploy workflow, not health indicators.
|
||||
|
||||
| File | Effect |
|
||||
| ---------------------- | ----------------------------------------------------------- |
|
||||
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
|
||||
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
|
||||
| Both | Run the Compose stack and apply the Kubernetes resources. |
|
||||
| Neither | Keep the configuration in Git without automatic deployment. |
|
||||
|
||||
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
|
||||
manual entry points. The deploy script does not discover them.
|
||||
|
||||
Kubernetes selection excludes secret files, examples, Helm values, and patches.
|
||||
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
|
||||
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
|
||||
does not install their charts.
|
||||
|
||||
The table below describes committed configuration. It does not claim that a
|
||||
service is currently healthy or running.
|
||||
|
||||
## Services
|
||||
|
||||
| Service | Configuration | Selected by markers |
|
||||
| ---------------------------------------------- | ---------------------------- | ------------------- |
|
||||
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
|
||||
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
|
||||
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
|
||||
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
||||
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
|
||||
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
|
||||
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
||||
| [Penpot](penpot/README.md) | Compose | Manual |
|
||||
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
|
||||
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual |
|
||||
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
|
||||
|
||||
## Running a Compose stack
|
||||
|
||||
Use the service README first. Where a service has an env example, copy it inside
|
||||
that service's directory and replace the placeholders. The root `.env.example`
|
||||
is an older collection of variables, not a complete configuration for every stack.
|
||||
|
||||
For example, from the repository root:
|
||||
|
||||
```sh
|
||||
cd netbox
|
||||
cp .env.example .env
|
||||
$EDITOR .env
|
||||
docker compose config --quiet
|
||||
docker compose up -d
|
||||
docker compose ps
|
||||
```
|
||||
|
||||
Stacks that attach to `proxy` require an existing Docker network of that name and
|
||||
an appropriate reverse proxy. Published host ports still work independently of
|
||||
Traefik. Check port conflicts before starting an alternative to a Kubernetes
|
||||
service: DNS, STUN, and HTTP listeners can share the same host.
|
||||
|
||||
`docker compose down` keeps named volumes. Adding `-v` removes them.
|
||||
|
||||
## Preparing Kubernetes
|
||||
|
||||
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
|
||||
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
|
||||
Replace the lab's hosts and addresses before using the configuration elsewhere.
|
||||
|
||||
Create a service's namespace, then prepare its ignored Secret from the example.
|
||||
For example:
|
||||
|
||||
```sh
|
||||
kubectl apply -f netbox/k8s/namespace.yaml
|
||||
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
|
||||
$EDITOR netbox/k8s/secrets.yaml
|
||||
kubectl apply -f netbox/k8s/secrets.yaml
|
||||
```
|
||||
|
||||
The deploy workflow applies the tracked resources for marked services. Avoid
|
||||
applying an entire `k8s/` directory blindly: some directories contain Helm values,
|
||||
examples, and alternative routes. For a manual change, apply the selected manifest
|
||||
explicitly and check the resulting rollout.
|
||||
|
||||
Shared database passwords must agree between the `database` namespace and each
|
||||
application's Secret. Updating the PostgreSQL Secret does not change an existing
|
||||
role's password; see the database README.
|
||||
|
||||
## Local checks
|
||||
|
||||
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
|
||||
|
||||
```fish
|
||||
set tools_dir (bash .gitea/workflows/install-ci-tools.sh)
|
||||
set -gx PATH $tools_dir $PATH
|
||||
ruff check .
|
||||
ruff format --check .
|
||||
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
|
||||
.gitea/workflows/sync-renovate-configmap.sh --check
|
||||
```
|
||||
|
||||
The [workflow README](.gitea/README.md#ci) lists the rest of the checks.
|
||||
Structure checks do not establish that local Secrets, mounted files, storage,
|
||||
or external services are ready.
|
||||
|
||||
## Data and recovery
|
||||
|
||||
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
|
||||
configuration. Keep backups of application data and the keys needed to read it.
|
||||
An image rollback does not roll back database migrations or ConfigMap contents.
|
||||
|
||||
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
|
||||
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
|
||||
The manifests do not provide a repository-wide backup schedule.
|
||||
|
||||
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
|
||||
is a planning document, not evidence that NFS has been installed.
|
||||
@@ -0,0 +1,22 @@
|
||||
# AdGuard Home
|
||||
|
||||
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
|
||||
|
||||
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
|
||||
configuration and working data, and mounts the `adguard-certs` TLS Secret.
|
||||
The LoadBalancer Service exposes DNS separately from the web ingress.
|
||||
|
||||
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
|
||||
and `certs/` before starting it. Starting both DNS deployments on the same address
|
||||
can cause a port conflict.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n adguard
|
||||
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Authentik
|
||||
|
||||
Identity provider with separate server and worker deployments.
|
||||
|
||||
Kubernetes connects to the shared PostgreSQL service in `database`. Set
|
||||
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
|
||||
Keep `AUTHENTIK_SECRET_KEY` with the backups.
|
||||
|
||||
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
|
||||
Its image defaults differ from Kubernetes; check both before an upgrade.
|
||||
The worker mounts the Docker socket for Docker outpost management.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n authentik
|
||||
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,18 @@
|
||||
# cert-manager
|
||||
|
||||
Public ACME issuers and an internal certificate authority.
|
||||
|
||||
This directory contains chart values and issuer resources, not the controller
|
||||
installation. Install the cert-manager chart with CRDs and the settings in
|
||||
`k8s/cert-manager-values.yaml` before applying the issuers.
|
||||
|
||||
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
|
||||
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
|
||||
reachability must work for the requested names before issuance.
|
||||
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
|
||||
up; the tracked `.crt` is only a public certificate.
|
||||
|
||||
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
|
||||
`kubectl apply` does not interpret the Helm values file.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Cloudflare DDNS
|
||||
|
||||
Updates the lab DNS records when the public address changes.
|
||||
|
||||
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
|
||||
The Compose stack also uses host networking. Configure the API token and domain
|
||||
list from the relevant example; keep DNS names consistent with the ingress rules.
|
||||
|
||||
`config.json.example` is a separate configuration example. The current Compose
|
||||
file does not mount a config.json file. Check configuration against the pinned
|
||||
DDNS image when changing between environment and file-based settings.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n default
|
||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Checkmk
|
||||
|
||||
Checkmk Raw monitoring site with web and agent-receiver ingress.
|
||||
|
||||
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
|
||||
volume on Compose. The agent receiver has a separate TCP route; enabling the
|
||||
web route alone does not expose it.
|
||||
|
||||
Prepare the password in the service env or Secret example. Inspect the Checkmk
|
||||
container logs during the first site creation.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n checkmk
|
||||
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Cloudflare Tunnel
|
||||
|
||||
A Kubernetes connector for an existing Cloudflare tunnel.
|
||||
|
||||
The Deployment runs in `default` and reads its token from the ignored Secret
|
||||
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
|
||||
in Cloudflare before starting the connector.
|
||||
|
||||
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
|
||||
`k8s/deployment.yaml` when this tunnel is needed.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n default
|
||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,22 @@
|
||||
# File converters
|
||||
|
||||
ConvertX for server-side conversion and BentoPDF for PDF tools.
|
||||
|
||||
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
|
||||
Kubernetes configuration includes a local `config.yaml.example`, excluded from
|
||||
normal deployment. Copy and apply the real ConfigMap separately where required.
|
||||
|
||||
Compose publishes ConvertX on host port 9992 as well as attaching it to the
|
||||
proxy network. Replace the authentication settings from `.env.example` before
|
||||
exposing it outside the lab.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n converters
|
||||
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,25 @@
|
||||
# CrowdSec
|
||||
|
||||
Helm values, dashboards, network policy, and a maintenance CronJob.
|
||||
|
||||
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
|
||||
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
|
||||
marker in this directory.
|
||||
|
||||
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
|
||||
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
|
||||
namespace Role. Review its script and schedule before enabling cleanup.
|
||||
|
||||
Traefik's values state that enforcement moved to a host firewall bouncer. This
|
||||
repository does not install that host component.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n crowdsec
|
||||
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Dockmon
|
||||
|
||||
Docker management UI that talks to the host Docker daemon.
|
||||
|
||||
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
|
||||
to the node hosting the pod, so this is not a cluster-wide container manager.
|
||||
|
||||
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
|
||||
with a volume claim template. Its ServersTransport is specific to the upstream
|
||||
connection; keep it with the ingress resources.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n dockmon
|
||||
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,134 @@
|
||||
# Repository review (6 October 2026 baseline)
|
||||
|
||||
This records the tracked tree at `cc9c3de` and the workstation state observed on
|
||||
6 October 2026. It is a historical review, not a current runtime inventory. The
|
||||
listed code fixes have since merged into `main`; EDU ownership has moved to the
|
||||
separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md).
|
||||
See the [CI and deployment guide](../.gitea/README.md) and
|
||||
[runner and recovery guide](../.gitea/runner/README.md) for the current workflow.
|
||||
No deployment was performed during the original review.
|
||||
|
||||
## Findings at the baseline and current status
|
||||
|
||||
| Priority | Finding at the baseline | Current status |
|
||||
| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
|
||||
| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. |
|
||||
| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. |
|
||||
| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. |
|
||||
| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. |
|
||||
| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. |
|
||||
| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. |
|
||||
| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. |
|
||||
| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. |
|
||||
| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. |
|
||||
|
||||
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
|
||||
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
|
||||
The fix retains the DoH path constraint for both hostnames.
|
||||
|
||||
The current deploy workflow deliberately rejects the unsafe prune option. It
|
||||
does not introduce automatic deletion under a different implementation. The
|
||||
baseline finding was a configuration risk, not evidence of a live deletion
|
||||
incident.
|
||||
|
||||
The former session fix bounded HTTP and Redis calls, validated credentials, set
|
||||
a cookie lifetime of two refresh intervals, and marked success only after
|
||||
publishing the verified cookie. The service is now owned by the EDU repository;
|
||||
see that repository for its current implementation.
|
||||
|
||||
The deployment fix extracts required pod Secret references from rendered JSON,
|
||||
checks their namespaces, includes init containers, image-pull credentials, and
|
||||
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
|
||||
issued by cert-manager are not treated as pre-existing pod prerequisites.
|
||||
It checks existence/access, not every key's contents or application validity.
|
||||
|
||||
## Live workstation observations
|
||||
|
||||
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
|
||||
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
|
||||
no pods were Pending or in another non-running, non-completed phase. This is a
|
||||
point-in-time observation, not a complete application health test.
|
||||
|
||||
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
|
||||
reviewed local commit. It has untracked host configuration and a separate
|
||||
`userbot/` directory. It was not reset or cleaned.
|
||||
|
||||
| Observed difference | Implication |
|
||||
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
|
||||
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
|
||||
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
|
||||
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
|
||||
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
|
||||
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
|
||||
|
||||
The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The
|
||||
first row above records the state before that change. Read
|
||||
[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current
|
||||
tracked monitoring configuration; the live observations in this section remain
|
||||
a snapshot from 6 October.
|
||||
|
||||
## Current recovery limits
|
||||
|
||||
The deployment controller and its recovery process changed after this review.
|
||||
The current operator workflow is documented in the
|
||||
[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries
|
||||
are:
|
||||
|
||||
- Kubernetes recovery can restore captured workload revisions. It does not
|
||||
restore ConfigMaps, Secrets, database schemas, or persistent data.
|
||||
- Compose recovery is manual. It uses saved resolved configuration, but it does
|
||||
not restore volume data or reverse database migrations.
|
||||
- Removed resources require manual review and removal; the deploy workflow does
|
||||
not prune them automatically.
|
||||
- Plan mode does not create namespaces. During apply, server validation for new
|
||||
namespaces runs after namespace creation and chart installation; a failed
|
||||
check can leave an empty namespace.
|
||||
- Storage policy and backup coverage remain service-specific. Check the live PV,
|
||||
PVC, and backup state before changing stateful workloads.
|
||||
|
||||
## Validation
|
||||
|
||||
At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose
|
||||
files, and Kubernetes resources with available schemas. Kubeconform found 347
|
||||
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
|
||||
That skip count matters: passing schema validation does not validate Traefik rule
|
||||
strings or other controller-specific behavior.
|
||||
|
||||
Fix validation covers:
|
||||
|
||||
- Compose discovery of manual entry points, rejection of required-variable gaps,
|
||||
namespace-scoped and optional Secret references, and API/render failures.
|
||||
- NetBird setup idempotence, preservation of existing keys, file permissions,
|
||||
runtime rendering, and rejection of invalid trusted proxy CIDRs.
|
||||
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
|
||||
missing credentials, and nonpositive refresh intervals.
|
||||
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
|
||||
- YAML and Compose structure for the corrected router rules, compared with the
|
||||
documented Traefik grammar. They were not exercised on the live proxy.
|
||||
- Prune rejection before any cluster invocation.
|
||||
|
||||
At the time of review, all seven fix branches and the documentation branch
|
||||
merged together in a disposable validation worktree. That combined tree passed the
|
||||
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
|
||||
Compose structure checks, and 11 Python regression tests plus the shell
|
||||
validation regressions. CRD server-side validation and live rollout tests were
|
||||
not run.
|
||||
|
||||
Runtime tests use fixtures and mocks, not production credentials. Live checks read
|
||||
workload metadata, storage policies, chart versions, and container state only.
|
||||
They did not read Secret contents or change services.
|
||||
|
||||
## Reloader follow-up (baseline)
|
||||
|
||||
`fix/reloader-integration` added the active marker and opt-in annotations to
|
||||
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
|
||||
It corrected AdGuard's misplaced pod-template annotation. The Helm settings use
|
||||
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
|
||||
CronJobs. PostgreSQL workloads are excluded because their credential variables
|
||||
and init scripts are only effective on an empty data directory.
|
||||
|
||||
The controller was running on the workstation when inspected. The original
|
||||
review checked configuration against the pinned chart with Helm rendering and
|
||||
manifest validation; it did not change production configuration to provoke a
|
||||
test restart or confirm every application's live reload behavior.
|
||||
@@ -0,0 +1,20 @@
|
||||
# Downtify
|
||||
|
||||
Download UI with a persistent downloads directory.
|
||||
|
||||
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
|
||||
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
|
||||
so check certificate and middleware availability before enabling them.
|
||||
|
||||
Back up downloads separately if they need to survive storage replacement.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n downtify
|
||||
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Error pages
|
||||
|
||||
Static HTTP error pages served by an Nginx image built in CI.
|
||||
|
||||
Edit the HTML in `html/`; the Dockerfile copies it into the image.
|
||||
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
|
||||
middleware. Keep the middleware's namespace and port aligned with that Service.
|
||||
|
||||
For a local build, run `docker build -t homelab-error-pages .` from this directory.
|
||||
Compose references the private registry image rather than a build context.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n error-pages
|
||||
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Gitea
|
||||
|
||||
Git hosting with HTTP and a separate SSH route.
|
||||
|
||||
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
|
||||
and application data. Match the Gitea database password with the shared database
|
||||
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
|
||||
|
||||
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
|
||||
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
|
||||
its own database, not a second frontend for the Kubernetes instance.
|
||||
|
||||
Back up repositories, application configuration, and a consistent database dump
|
||||
together. Gitea Actions definitions for this repository live in `../.gitea/`.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n gitea
|
||||
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -20,6 +20,8 @@ data:
|
||||
|
||||
GITEA__mailer__ENABLED: "false"
|
||||
|
||||
GITEA__metrics__ENABLED: "true"
|
||||
|
||||
# No code/issue search needed: bleve reindexes the whole issue index on
|
||||
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
||||
# the rotational disk for an hour. "db" serves issue search from postgres.
|
||||
|
||||
@@ -3,6 +3,8 @@ kind: Service
|
||||
metadata:
|
||||
name: gitea-service
|
||||
namespace: gitea
|
||||
labels:
|
||||
app: gitea
|
||||
spec:
|
||||
selector:
|
||||
app: gitea
|
||||
|
||||
@@ -7,7 +7,8 @@ spec:
|
||||
entryPoints:
|
||||
- websecure
|
||||
routes:
|
||||
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
|
||||
# Metrics are scraped directly through the cluster Service.
|
||||
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
|
||||
kind: Rule
|
||||
services:
|
||||
- name: gitea-service
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: gitea
|
||||
namespace: gitea
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: gitea
|
||||
endpoints:
|
||||
- port: http
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -0,0 +1,26 @@
|
||||
# Glance
|
||||
|
||||
Dashboard pages for links, service checks, and Docker containers.
|
||||
|
||||
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
|
||||
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
|
||||
`user.css`. Update both copies when changing shared content.
|
||||
|
||||
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
|
||||
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
|
||||
no tracked Secret example; create that Secret in `glance` before starting it.
|
||||
Compose expects a local `.env` with the same password.
|
||||
|
||||
The pod mounts `user.css` from `glance-assets`, which is the ConfigMap that
|
||||
contains that key.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n glance
|
||||
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,18 @@
|
||||
# Headscale
|
||||
|
||||
Headscale, Headplane, and a separate web administration UI on Docker.
|
||||
|
||||
Kubernetes only provides routes to the Docker host. Update the addresses in
|
||||
`k8s/routing/external-service.yaml` if the host moves.
|
||||
|
||||
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
|
||||
`config/policy.json.example` to their names without `.example`. Set the public
|
||||
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
|
||||
The example URLs are placeholders.
|
||||
|
||||
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
|
||||
13000, and the other UI on 10080. The data volumes store the Headscale database,
|
||||
keys, and Headplane state. The embedded DERP configuration needs reachable
|
||||
addresses; Compose does not publish its UDP 3478 listener.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -7,7 +7,7 @@
|
||||
# # Dev server_url
|
||||
# server_url: https://hs.dev_internal_domain.internal
|
||||
listen_addr: 0.0.0.0:8080
|
||||
metrics_listen_addr: 127.0.0.1:9090
|
||||
metrics_listen_addr: 0.0.0.0:9090
|
||||
grpc_listen_addr: 127.0.0.1:50443
|
||||
grpc_allow_insecure: false
|
||||
noise:
|
||||
|
||||
@@ -3,6 +3,8 @@ kind: Service
|
||||
metadata:
|
||||
name: headscale-server-external
|
||||
namespace: headscale
|
||||
labels:
|
||||
app: headscale
|
||||
spec:
|
||||
ports:
|
||||
- port: 8080
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
apiVersion: operator.victoriametrics.com/v1beta1
|
||||
kind: VMServiceScrape
|
||||
metadata:
|
||||
name: headscale
|
||||
namespace: headscale
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
# The external Service has a manually managed EndpointSlice, not Endpoints.
|
||||
discoveryRole: endpointslice
|
||||
selector:
|
||||
matchLabels:
|
||||
app: headscale
|
||||
endpoints:
|
||||
- port: metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -0,0 +1,24 @@
|
||||
# Homarr
|
||||
|
||||
Dashboard with Kubernetes integration and persistent application state.
|
||||
|
||||
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
|
||||
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
|
||||
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
|
||||
|
||||
The committed ingress is internal. There is no `k8s/active` marker even though
|
||||
manifests exist, so the workflow does not select Homarr automatically.
|
||||
|
||||
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
|
||||
expects a local kubeconfig. Check these host ports against Traefik before use.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n homarr
|
||||
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Homepages
|
||||
|
||||
Two static sites: Forust and xdfnx.
|
||||
|
||||
The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with
|
||||
its own Dockerfile and publishes it to the private registry. Kubernetes serves
|
||||
the image contents; Compose overlays the source directories as bind mounts.
|
||||
|
||||
Both Traefik IngressRoute and Gateway API route manifests are committed.
|
||||
Keep their hostnames and backend Services aligned when changing routes.
|
||||
Certificate resources cover public and internal hostnames.
|
||||
|
||||
Build either site locally with `docker build -f Dockerfile.forust .` or
|
||||
`docker build -f Dockerfile.xdfnx .` from this directory.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n homepages
|
||||
kubectl get events -n homepages --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,29 @@
|
||||
# Immich
|
||||
|
||||
Photo library with its own vector-enabled PostgreSQL and machine-learning service.
|
||||
|
||||
This database is separate from the shared PostgreSQL instance. Keep the server
|
||||
and machine-learning versions aligned when upgrading.
|
||||
|
||||
Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must
|
||||
already exist and contain the intended library; moving the pod to a different
|
||||
node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and
|
||||
the model cache has its own PVC.
|
||||
|
||||
Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example
|
||||
uses the same library path as Kubernetes. Run one writer against that library;
|
||||
do not start both deployments as independent instances over the same files.
|
||||
|
||||
Back up the library and a consistent database dump together. The model cache
|
||||
can be rebuilt; the photo database cannot.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n immich
|
||||
kubectl get events -n immich --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -6,6 +6,10 @@ metadata:
|
||||
data:
|
||||
TZ: "Europe/Bratislava"
|
||||
|
||||
IMMICH_TELEMETRY_INCLUDE: "all"
|
||||
IMMICH_API_METRICS_PORT: "8081"
|
||||
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
|
||||
|
||||
# The database in this namespace, not the shared one in the database
|
||||
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
||||
DB_HOSTNAME: "immich-postgres"
|
||||
|
||||
@@ -3,6 +3,8 @@ kind: Service
|
||||
metadata:
|
||||
name: immich-service
|
||||
namespace: immich
|
||||
labels:
|
||||
app: immich
|
||||
spec:
|
||||
selector:
|
||||
app: immich
|
||||
@@ -10,6 +12,12 @@ spec:
|
||||
- name: http
|
||||
port: 2283
|
||||
targetPort: 2283
|
||||
- name: api-metrics
|
||||
port: 8081
|
||||
targetPort: api-metrics
|
||||
- name: worker-metrics
|
||||
port: 8082
|
||||
targetPort: worker-metrics
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
@@ -41,6 +49,10 @@ spec:
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 2283
|
||||
- name: api-metrics
|
||||
containerPort: 8081
|
||||
- name: worker-metrics
|
||||
containerPort: 8082
|
||||
volumeMounts:
|
||||
- name: immich-data
|
||||
mountPath: /data
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: immich
|
||||
namespace: immich
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: immich
|
||||
endpoints:
|
||||
- port: api-metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
- port: worker-metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -0,0 +1,21 @@
|
||||
# Kener
|
||||
|
||||
Status page with Redis and persistent database and upload directories.
|
||||
|
||||
Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet.
|
||||
Compose keeps the corresponding directories in named volumes. Set the signing
|
||||
and other credentials from the env or Secret example.
|
||||
|
||||
The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`.
|
||||
There is no active marker for either runtime.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n kener
|
||||
kubectl get events -n kener --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
kener:
|
||||
image: rajnandan1/kener:4.1.5
|
||||
image: rajnandan1/kener:v4.1.7
|
||||
container_name: kener
|
||||
restart: unless-stopped
|
||||
# ports:
|
||||
|
||||
@@ -31,7 +31,7 @@ spec:
|
||||
spec:
|
||||
containers:
|
||||
- name: kener
|
||||
image: rajnandan1/kener:4.1.5
|
||||
image: rajnandan1/kener:v4.1.7
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: kener-config
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
# Loki and Alloy
|
||||
|
||||
Loki log storage and Alloy collection, both deployed through Helm.
|
||||
|
||||
The deploy library lists separate `loki` and `alloy` releases in `prometheus`,
|
||||
controlled by this directory's `k8s/active` marker. Chart versions are pinned in
|
||||
`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`.
|
||||
|
||||
Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the
|
||||
monitoring stack. Review Loki retention and storage settings before enabling
|
||||
collection on a new cluster.
|
||||
|
||||
Check releases with `helm list -n prometheus` and inspect collector logs before
|
||||
assuming that an empty Grafana query means there were no events.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# MeTube
|
||||
|
||||
Web downloader behind Traefik.
|
||||
|
||||
Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB
|
||||
`emptyDir` for `/downloads`: completed downloads disappear when the pod is
|
||||
replaced. Download files from the UI promptly if this temporary storage is intended.
|
||||
|
||||
Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes
|
||||
would require changing the volume to a PVC and choosing a storage policy.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n metube
|
||||
kubectl get events -n metube --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# n8n
|
||||
|
||||
Workflow automation with persistent application and file storage.
|
||||
|
||||
Kubernetes keeps application state in `n8n-node-pvc` and files in
|
||||
`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes.
|
||||
Webhook URLs and proxy settings are committed in the application config.
|
||||
|
||||
There is no active marker. Review the URLs before enabling the stack, and retain
|
||||
the credential encryption key with the database or application-data backup.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n n8n
|
||||
kubectl get events -n n8n --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
n8n:
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.2
|
||||
container_name: n8n
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
|
||||
+1
-1
@@ -31,7 +31,7 @@ spec:
|
||||
spec:
|
||||
containers:
|
||||
- name: n8n
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.2
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: n8n-config
|
||||
|
||||
+16
-4
@@ -1,12 +1,24 @@
|
||||
# NetBird
|
||||
|
||||
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
|
||||
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind Traefik. The Compose configuration uses the external Docker `proxy` network and publishes only STUN UDP `3478` directly.
|
||||
|
||||
The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
|
||||
The single-instance server uses SQLite. Back up its data and datastore encryption key together.
|
||||
|
||||
## Kubernetes
|
||||
|
||||
`k8s/active` selects the Kubernetes deployment. It runs the server and dashboard
|
||||
in namespace `netbird`; the server stores SQLite data in `netbird-pvc`. The
|
||||
configuration renderer and template are in `k8s/`. Prepare
|
||||
`k8s/secrets.yaml` from `k8s/secrets.yaml.example` before the first deploy.
|
||||
|
||||
## Compose alternative
|
||||
|
||||
The Compose files are available for manual use. There is no root `active` marker,
|
||||
so the automatic deploy workflow selects Kubernetes only.
|
||||
|
||||
## Files
|
||||
|
||||
- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
|
||||
- `compose.yaml`: dashboard and combined server; start it manually when using Compose.
|
||||
- `config.template.yaml`: non-secret server configuration rendered at startup.
|
||||
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
|
||||
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
|
||||
@@ -15,7 +27,7 @@ The deployment uses SQLite for a single-instance homelab server. The persistent
|
||||
|
||||
## First deployment
|
||||
|
||||
Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
|
||||
Run these commands on the Docker host before the first Compose start.
|
||||
|
||||
```bash
|
||||
cd /srv/homelab/netbird
|
||||
|
||||
@@ -3,6 +3,8 @@ kind: Service
|
||||
metadata:
|
||||
name: netbird-server-service
|
||||
namespace: netbird
|
||||
labels:
|
||||
app: netbird-server
|
||||
spec:
|
||||
selector:
|
||||
app: netbird-server
|
||||
@@ -11,6 +13,10 @@ spec:
|
||||
name: http
|
||||
targetPort: 80
|
||||
protocol: TCP
|
||||
- port: 9090
|
||||
name: metrics
|
||||
targetPort: metrics
|
||||
protocol: TCP
|
||||
- port: 3478
|
||||
name: stun
|
||||
targetPort: 3478
|
||||
@@ -59,6 +65,9 @@ spec:
|
||||
- containerPort: 80
|
||||
name: http
|
||||
protocol: TCP
|
||||
- containerPort: 9090
|
||||
name: metrics
|
||||
protocol: TCP
|
||||
- containerPort: 3478
|
||||
name: stun
|
||||
protocol: UDP
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: netbird-server
|
||||
namespace: netbird
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: netbird-server
|
||||
endpoints:
|
||||
- port: metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
+35
-88
@@ -1,96 +1,43 @@
|
||||
# NetBox
|
||||
|
||||
NetBox for homelab documentation and visualization. Two runtimes are available:
|
||||
Inventory and network documentation with a web process, worker, and Valkey.
|
||||
|
||||
| Runtime | Manifest | Purpose |
|
||||
| ------- | -------------- | -------------------------------------------------------------- |
|
||||
| Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) |
|
||||
| k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) |
|
||||
Kubernetes uses the shared PostgreSQL service at
|
||||
`postgres.database.svc.cluster.local:5432`, database and role `netbox`.
|
||||
The database and application Secrets must contain the same password.
|
||||
Media, reports, scripts, and Valkey have persistent storage.
|
||||
|
||||
Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks
|
||||
plus a second logical database for caching. The Docker stand keeps its own
|
||||
PostgreSQL container, while the k8s deployment uses the shared `database` cluster
|
||||
(`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey
|
||||
stays a per-service StatefulSet.
|
||||
Compose has its own PostgreSQL container and Valkey instances. It publishes the
|
||||
web UI on `127.0.0.1:8000`; its Traefik labels can also expose it while a Docker
|
||||
proxy is running. Copy `.env.example` to `.env`, replace the credentials, and run
|
||||
`docker compose config --quiet` before starting it.
|
||||
|
||||
## Docker Compose
|
||||
## First Kubernetes start
|
||||
|
||||
```bash
|
||||
cp .env.example .env
|
||||
# replace CHANGE_ME
|
||||
docker compose up -d
|
||||
Create the namespace and application Secret. Provision the database through the
|
||||
shared database initializer on a fresh instance, or create the role and database
|
||||
manually on an existing instance; see [PostgreSQL](../postgres/README.md).
|
||||
The database NetworkPolicy already includes `netbox`.
|
||||
|
||||
Apply the selected application manifests after the database is ready. Startup
|
||||
runs schema migrations, so the probes allow a longer first boot. Inspect web and
|
||||
worker logs before retrying a slow migration.
|
||||
|
||||
## Settings and backup
|
||||
|
||||
`configuration/configuration.py` is the Compose settings file. Its Kubernetes
|
||||
copy is embedded in `k8s/settings.yaml`; keep them aligned.
|
||||
Back up the database and media together. Keep `SECRET_KEY` and
|
||||
`API_TOKEN_PEPPER_1`: changing them invalidates sessions or API tokens.
|
||||
A container rollback cannot undo a database migration.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n netbox
|
||||
kubectl get events -n netbox --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
The UI is available at <http://localhost:8000>. The port is bound to `127.0.0.1`
|
||||
intentionally, so this stand is not exposed on the LAN or public interfaces.
|
||||
|
||||
The `netbox` service is also attached to the external `proxy` network and carries
|
||||
Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those
|
||||
labels only take effect while the Docker Traefik stack is running; it is currently
|
||||
stopped, and the live ingress path in this homelab is the k8s Traefik.
|
||||
|
||||
Inspect startup and health with:
|
||||
|
||||
```bash
|
||||
docker compose ps
|
||||
docker compose logs -f netbox
|
||||
```
|
||||
|
||||
Stop it with `docker compose down`; data is kept in the named volumes
|
||||
`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`,
|
||||
`netbox-scripts-files` and `netbox-redis-data`.
|
||||
|
||||
## Kubernetes
|
||||
|
||||
`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly
|
||||
plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To
|
||||
rebuild it from scratch:
|
||||
|
||||
```bash
|
||||
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
|
||||
kubectl -n database patch secret postgres-shared-secrets \
|
||||
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
|
||||
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
|
||||
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
|
||||
|
||||
# 2. secrets first: the deploy workflow never applies *secret*.yaml
|
||||
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
|
||||
kubectl apply -f k8s/secrets.yaml
|
||||
|
||||
# 3. manifests
|
||||
kubectl apply -f k8s/
|
||||
```
|
||||
|
||||
The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its
|
||||
NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace
|
||||
or connections are dropped, and `postgres/initdb/01-create-databases.sh` already
|
||||
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
|
||||
StatefulSet of its own — only `netbox-valkey`.
|
||||
|
||||
`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS`
|
||||
list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the
|
||||
`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`.
|
||||
|
||||
Resources are permanent again now that the first-boot migrations are complete:
|
||||
the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the
|
||||
worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves
|
||||
`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps
|
||||
leave enough headroom for future schema migrations without letting one process
|
||||
consume the whole node.
|
||||
|
||||
The first start applies ~810 migrations, each in its own transaction with DDL and
|
||||
a commit; every later start is a no-op. The startup probe allows 15 minutes and
|
||||
`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod
|
||||
and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would
|
||||
replace the probe destination with that public hostname and bypass the pod.
|
||||
|
||||
## Secrets
|
||||
|
||||
- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only
|
||||
`.env.example` and `k8s/secrets.yaml.example` are committed.
|
||||
- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and
|
||||
the Django keys all come from the environment, so the same settings file works in
|
||||
both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap
|
||||
(`k8s/settings.yaml`) and must be kept in sync with the file.
|
||||
- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1`
|
||||
invalidates every API token.
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Netronome
|
||||
|
||||
Network monitoring application using the shared PostgreSQL instance on Kubernetes.
|
||||
|
||||
Kubernetes reads application settings from its ConfigMap and Secret. Match the
|
||||
Netronome role password with `NETRONOME_DB_PASSWORD` in the shared database Secret.
|
||||
Its namespace is included in the PostgreSQL NetworkPolicy.
|
||||
|
||||
The Compose configuration is a separate deployment; review its local database
|
||||
settings and env example before starting it. Keep monitoring history in the
|
||||
database backup.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n netronome
|
||||
kubectl get events -n netronome --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,17 @@
|
||||
# Nextcloud AIO
|
||||
|
||||
Nextcloud All-in-One on Docker, with Kubernetes routes to the Docker host.
|
||||
|
||||
The master container manages its own child containers through the Docker
|
||||
socket. Kubernetes does not run the Nextcloud application; the EndpointSlices
|
||||
under `k8s/routing/` point to host services.
|
||||
|
||||
Compose publishes the AIO administration interface on 8888. The Apache frontend
|
||||
uses host port 11000. `NEXTCLOUD_DATADIR` is `/mnt/nextcloud/ncdata`; prepare that
|
||||
storage before first setup and do not change the path casually afterwards.
|
||||
|
||||
Use AIO's backup and restore tools for the managed application. Keep the master
|
||||
configuration volume and the data directory with the recovery plan. Do not
|
||||
remove child containers just because they do not appear as Compose services.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,12 @@
|
||||
# Penpot
|
||||
|
||||
A Compose-only design application with frontend, backend, exporter, database, and cache.
|
||||
|
||||
There is no active marker or Kubernetes deployment here. Configure the public
|
||||
URL and credentials from `.env.example` before starting `compose.yaml`.
|
||||
|
||||
Penpot has its own PostgreSQL container. The shared database initializer still
|
||||
contains a Penpot role, but this Compose stack does not use it.
|
||||
Back up the application assets and database together.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Portainer
|
||||
|
||||
Container management UI backed by the host Docker socket.
|
||||
|
||||
Kubernetes mounts the node's Docker socket and persists application data in
|
||||
`portainer-data-pvc`. This targets Docker on that node, not Kubernetes workloads.
|
||||
Compose uses the `portainer_data` volume for its state.
|
||||
|
||||
Review initial administrator setup and route access before exposing the UI.
|
||||
Neither deployment has an active marker.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n portainer
|
||||
kubectl get events -n portainer --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
portainer:
|
||||
image: portainer/portainer-ce:2.45.1
|
||||
image: portainer/portainer-ce:2.45.2
|
||||
container_name: portainer
|
||||
restart: always
|
||||
volumes:
|
||||
|
||||
@@ -29,7 +29,7 @@ spec:
|
||||
spec:
|
||||
containers:
|
||||
- name: portainer
|
||||
image: portainer/portainer-ce:2.45.1
|
||||
image: portainer/portainer-ce:2.45.2
|
||||
ports:
|
||||
- containerPort: 9000
|
||||
volumeMounts:
|
||||
|
||||
+56
-28
@@ -1,35 +1,63 @@
|
||||
# Shared PostgreSQL
|
||||
|
||||
This directory contains the shared PostgreSQL 17 deployment for Authentik,
|
||||
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
|
||||
per service. Per-service standalone databases were removed after the
|
||||
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
|
||||
not part of the shared instance).
|
||||
PostgreSQL 17 for the Kubernetes deployments of Authentik, Gitea, NetBox, and Netronome.
|
||||
|
||||
## Compatibility baseline
|
||||
The server runs in `database` as StatefulSet `postgres17`, with data in
|
||||
`postgres17-data`. Applications connect to
|
||||
`postgres.database.svc.cluster.local:5432`. The NetworkPolicy allows only the
|
||||
listed application namespaces; add a new consumer there as well as provisioning
|
||||
its database.
|
||||
|
||||
| Service | Current application | Shared PostgreSQL 17 |
|
||||
| ---------- | ------------------- | -------------------------------------- |
|
||||
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
|
||||
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
|
||||
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
|
||||
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
|
||||
| Statuspage | custom | Supported |
|
||||
## Initialization
|
||||
|
||||
A major-version change must use a logical dump/restore; changing only the
|
||||
image tag while keeping a data directory is not supported.
|
||||
`initdb/01-create-databases.sh` creates roles and databases on an empty data
|
||||
directory. The Kubernetes copy is embedded in `k8s/postgres.yaml`.
|
||||
It also provisions Penpot and Statuspage roles, even though those are not active
|
||||
consumers in the current Kubernetes manifests.
|
||||
|
||||
For Compose, copy `.env.example` to `.env`, set all passwords, and start it with
|
||||
`docker compose -f shared-compose.yaml up -d`. This file is intentionally not
|
||||
named `compose.yaml`, so the repository deploy workflow does not start a second
|
||||
database accidentally.
|
||||
Applications that use this database must also join that external network and use
|
||||
`homelab-postgres:5432`.
|
||||
The initializer requires every listed password. Prepare `k8s/secrets.yaml` from
|
||||
the example before applying the StatefulSet. Existing application Secrets keep
|
||||
copies of their own database passwords; they must match the corresponding role.
|
||||
|
||||
For Kubernetes, create `k8s/secrets.yaml` from the example before applying the
|
||||
manifests. The `k8s/active` marker makes the normal deploy workflow include the
|
||||
namespace, StatefulSet, ConfigMap, and NetworkPolicy. Applications use
|
||||
`postgres.database.svc.cluster.local:5432`.
|
||||
Migrate each existing database with a tested logical dump/restore before
|
||||
switching an application. Do not reuse a PostgreSQL 14 or 17 data directory
|
||||
with PostgreSQL 15.
|
||||
The init scripts do not run again when an existing data directory is mounted.
|
||||
Changing a Secret does not rotate the PostgreSQL role password. Rotate the role
|
||||
with SQL and update the application Secret together.
|
||||
|
||||
## Compose alternative
|
||||
|
||||
From this directory:
|
||||
|
||||
```sh
|
||||
cp .env.example .env
|
||||
$EDITOR .env
|
||||
docker compose -f shared-compose.yaml config --quiet
|
||||
docker compose -f shared-compose.yaml up -d
|
||||
```
|
||||
|
||||
The example includes `NETBOX_DB_PASSWORD`; fill it and every other required
|
||||
password before starting the stack.
|
||||
This stack creates the `homelab-database` Docker network and the
|
||||
`homelab-postgres` container. Compose applications need to join that network
|
||||
explicitly to use it; several committed Compose stacks use their own databases.
|
||||
|
||||
The filename is intentional: the automatic deploy discovery does not start this
|
||||
stack just because the Kubernetes database is active.
|
||||
|
||||
## Backup and upgrades
|
||||
|
||||
Keep database dumps and role definitions, including ownership and grants.
|
||||
Take a logical backup before changing a major PostgreSQL version. A new image
|
||||
tag over the existing data directory is not a major-version migration.
|
||||
Test restores separately before changing application connection settings.
|
||||
Immich uses its own vector-enabled database and is outside this shared instance.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n database
|
||||
kubectl get events -n database --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,32 @@
|
||||
# Monitoring stack
|
||||
|
||||
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
|
||||
and vmalert. The `k8s/active` marker selects the stack. The
|
||||
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
|
||||
Prometheus Operator, and related components. Its Prometheus server is configured
|
||||
with zero replicas while VMAgent collects metrics and writes them to the
|
||||
single-node VictoriaMetrics instance.
|
||||
|
||||
The `victoria-operator` Helm release converts selected Prometheus Operator
|
||||
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
|
||||
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
|
||||
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
|
||||
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
|
||||
validation commands.
|
||||
|
||||
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
|
||||
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
|
||||
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
|
||||
keep their credentials out of the values file. Persistent volumes store data for
|
||||
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
|
||||
backups before changing storage. VictoriaMetrics currently retains 30 days of
|
||||
data.
|
||||
|
||||
A separate Compose configuration is present for manual use. There is no root
|
||||
`active` marker, so the automatic deploy workflow does not select it.
|
||||
|
||||
The deploy workflow does not remove resources when manifests are deleted. For a
|
||||
rollback of application-metrics changes, follow the explicit cleanup steps in
|
||||
the [Kubernetes monitoring notes](k8s/README.md).
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -15,3 +15,29 @@ The VictoriaMetrics Operator chart and its CRDs are installed before the
|
||||
Kubernetes manifests by the normal deploy workflow. On a cluster where the
|
||||
operator CRDs are not installed yet, CI skips the server-side dry-run of the
|
||||
`VMAgent` resource; the deploy installs the chart before applying that resource.
|
||||
|
||||
## Application metrics
|
||||
|
||||
The application ServiceMonitors use a 30s interval and a 10s timeout:
|
||||
|
||||
- Headscale: the external Service points to the Compose host on port 19090.
|
||||
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
|
||||
The Compose configuration must bind metrics to `0.0.0.0:9090`.
|
||||
- NetBird: the combined server exports `/metrics` on port 9090. The existing
|
||||
`server.metricsPort` setting enables the listener.
|
||||
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
|
||||
ingress excludes this path. The monitor uses the internal Service directly.
|
||||
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
|
||||
8081 and 8082. The monitor scrapes both ports on each server replica.
|
||||
|
||||
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
|
||||
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
|
||||
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
|
||||
VictoriaMetrics. All targets should report 1.
|
||||
|
||||
For rollback, revert the application metrics changes, run CI, and deploy the
|
||||
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
|
||||
workflow applies manifests and does not prune removed resources.
|
||||
|
||||
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
|
||||
previous Compose metrics bind address, and restart only the Headscale service.
|
||||
@@ -0,0 +1,20 @@
|
||||
# RackPeek
|
||||
|
||||
Rack inventory UI behind Traefik.
|
||||
|
||||
Kubernetes stores configuration in `rackpeek-pvc`. The Compose alternative uses
|
||||
its own data mount. Keep rack descriptions and inventory data in the backup.
|
||||
|
||||
Public and internal certificates and routes are in `k8s/`. There are no tracked
|
||||
Secret examples for this service.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n rackpeek
|
||||
kubectl get events -n rackpeek --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Reloader
|
||||
|
||||
Restarts opted-in workloads when the ConfigMaps or Secrets they consume change.
|
||||
The deploy workflow upgrades the `reloader` Helm release in namespace `reloader`;
|
||||
`k8s/active` enables it. The chart version is pinned in `deploy-lib.sh`.
|
||||
|
||||
## Workload integration
|
||||
|
||||
Put this annotation on the Deployment or StatefulSet metadata:
|
||||
|
||||
```yaml
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
```
|
||||
|
||||
The annotation belongs to the workload, not `spec.template.metadata`.
|
||||
Reloader discovers references in environment variables and mounted volumes.
|
||||
This covers startup-only settings and ConfigMaps or Secrets mounted with `subPath`.
|
||||
See the [upstream usage guide](https://github.com/stakater/Reloader/blob/v1.4.22/README.md#usage).
|
||||
|
||||
The application manifests opt in workloads including AdGuard's TLS files,
|
||||
NetBird, both NetBox processes, and the password-protected Valkey servers.
|
||||
Inactive services have the same annotations ready for later activation.
|
||||
|
||||
## Controller policy
|
||||
|
||||
The controller watches all namespaces but only restarts annotated workloads.
|
||||
It uses the `annotations` reload strategy, so changes trigger a pod-template
|
||||
annotation rather than injecting extra environment variables.
|
||||
|
||||
Jobs and CronJobs are excluded: their next execution reads current configuration.
|
||||
PostgreSQL is intentionally not opted in. Its password variables and init scripts
|
||||
apply to first initialization; restarting an existing database does not rotate
|
||||
roles or rerun those scripts. Rotate database credentials with SQL and update the
|
||||
clients' Secrets together.
|
||||
|
||||
Helm-managed monitoring components already have their own configuration reload
|
||||
paths; Traefik watches its file-provider configuration. They are not globally
|
||||
opted in. The controller does not react to files in PVCs or changes to external
|
||||
services unless a watched ConfigMap or Secret changes.
|
||||
|
||||
## Verify
|
||||
|
||||
```sh
|
||||
kubectl -n reloader rollout status deployment/reloader-reloader
|
||||
kubectl -n reloader logs deployment/reloader-reloader --since=10m
|
||||
kubectl -n netbird get deployment netbird-server-deployment \
|
||||
-o jsonpath='{.metadata.annotations.reloader\.stakater\.com/auto}'
|
||||
```
|
||||
|
||||
A changed configuration can briefly interrupt a single-replica service, especially
|
||||
one using `Recreate`. Installing annotations does not validate the configuration
|
||||
or migrate database data. Keep changes to shared Secrets coordinated across consumers.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
+38
-88
@@ -1,101 +1,51 @@
|
||||
# Renovate for Gitea
|
||||
# Renovate
|
||||
|
||||
Renovate runs as a Kubernetes CronJob and creates container image update pull
|
||||
requests in Gitea. It does not deploy changes itself.
|
||||
Container and chart dependency updates for the Gitea repository.
|
||||
|
||||
## Kubernetes
|
||||
The Kubernetes CronJob runs in `renovate` every six hours with overlapping
|
||||
CronJob executions forbidden. Prepare the bot PAT from the Secret example.
|
||||
Give the dedicated Gitea user access to the repositories it should update.
|
||||
|
||||
Create a dedicated Gitea user named `renovate-bot`, create a repository access
|
||||
token, and grant it repository read/write plus issue read/write permissions.
|
||||
Add `read:packages` if Renovate must inspect private Gitea registry images.
|
||||
|
||||
Create the ignored Secret locally; never commit the PAT:
|
||||
`renovate.json` is the source configuration. The ConfigMap is a generated copy:
|
||||
|
||||
```sh
|
||||
cp renovate/k8s/secrets.yaml.example renovate/k8s/secrets.yaml
|
||||
$EDITOR renovate/k8s/secrets.yaml
|
||||
kubectl apply -f renovate/k8s/namespace.yaml
|
||||
kubectl apply -f renovate/k8s/secrets.yaml
|
||||
kubectl apply -f renovate/k8s/configmap.yaml
|
||||
kubectl apply -f renovate/k8s/cronjob.yaml
|
||||
.gitea/workflows/sync-renovate-configmap.sh
|
||||
.gitea/workflows/sync-renovate-configmap.sh --check
|
||||
```
|
||||
|
||||
The `renovate/k8s/active` marker makes the normal deployment workflow include
|
||||
the namespace, ConfigMap, and CronJob. The Secret is intentionally excluded
|
||||
from Git and must be applied separately after every new cluster.
|
||||
Run those commands from the repository root. The `renovate-ci` workflow checks
|
||||
that the generated configuration agrees with the source.
|
||||
|
||||
Run it immediately instead of waiting for the six-hour schedule.
|
||||
## Run manually
|
||||
|
||||
Two options, both use the same `renovate/renovate.json`:
|
||||
From the repository root:
|
||||
|
||||
```fish
|
||||
kubectl create job --from=cronjob/renovate renovate-manual-(date +%s) -n renovate
|
||||
kubectl get jobs,pods -n renovate
|
||||
```
|
||||
|
||||
Alternatively use the `renovate-run` Actions workflow. It reads the image tag
|
||||
from the CronJob and accepts repository, log-level, and dry-run inputs. Actions
|
||||
requires `RENOVATE_TOKEN`; `RENOVATE_GITHUB_COM_TOKEN` is optional.
|
||||
The Actions concurrency group and the CronJob policy are separate, so avoid
|
||||
starting both against the same repository at once.
|
||||
|
||||
For Compose, copy `.env.example` to `.env` in this directory and run
|
||||
`docker compose -f renovate-compose.yaml run --rm renovate`. That file is a
|
||||
manual entry point and is not selected by the deploy workflow.
|
||||
|
||||
The config also tracks chart versions in `deploy-lib.sh` and tool versions in
|
||||
`.gitea/workflows/tool-versions.env`. Renovate opens pull requests; the normal CI and deploy
|
||||
workflows handle changes after merge.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl create job --from=cronjob/renovate renovate-manual-$(date +%s) -n renovate
|
||||
kubectl get pods,svc,pvc -n renovate
|
||||
kubectl get events -n renovate --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
or the `renovate-run` Actions workflow (Actions tab → `renovate-run` →
|
||||
Run workflow). It runs the same image as the CronJob on the self-hosted runner
|
||||
via Docker — the tag is read out of `renovate/k8s/cronjob.yaml` at run time
|
||||
rather than hardcoded, so the two cannot drift apart. Required Actions secrets
|
||||
(repo or org settings):
|
||||
|
||||
- `RENOVATE_TOKEN` — renovate-bot PAT (repository + issue read/write).
|
||||
- `RENOVATE_GITHUB_COM_TOKEN` — optional, for changelogs and GitHub rate limits.
|
||||
|
||||
Inputs: `repositories` (default `forust/homelab`), `log_level`
|
||||
(`info`/`debug`). Only one run at a time (concurrency group
|
||||
`renovate-run`), same as the CronJob `Forbid` policy.
|
||||
|
||||
Inspect runs with:
|
||||
|
||||
```sh
|
||||
kubectl get cronjob,jobs,pods -n renovate
|
||||
kubectl logs -n renovate job/<job-name>
|
||||
```
|
||||
|
||||
`RENOVATE_GITHUB_COM_TOKEN` is optional but recommended for changelogs and
|
||||
GitHub API rate limits. Set it in the Kubernetes Secret if available.
|
||||
|
||||
## Compose
|
||||
|
||||
Copy `.env.example` to `.env`, set the PAT, and run:
|
||||
|
||||
```sh
|
||||
docker compose -f renovate-compose.yaml run --rm renovate
|
||||
```
|
||||
|
||||
The Compose file is intentionally named `renovate-compose.yaml`, so the
|
||||
repository's automatic deployment discovery does not start it accidentally.
|
||||
|
||||
## Configuration
|
||||
|
||||
`renovate/renovate.json` is the single source of truth. The Compose file and the
|
||||
`renovate-run` workflow mount that file directly.
|
||||
|
||||
A ConfigMap cannot read from the repository, so the CronJob needs the config
|
||||
inlined. `renovate/k8s/configmap.yaml` is therefore a **generated** copy:
|
||||
|
||||
```sh
|
||||
.gitea/workflows/sync-renovate-configmap.sh # regenerate after editing
|
||||
.gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
|
||||
```
|
||||
|
||||
The `renovate-ci` workflow runs the `--check` form on every PR and push, so a
|
||||
config edit that forgets to regenerate the ConfigMap cannot be merged.
|
||||
|
||||
Beyond images, `customManagers` in the config track:
|
||||
|
||||
- Helm chart versions pinned in `.gitea/workflows/deploy-lib.sh`. The built-in
|
||||
`helmv3` manager only reads `Chart.yaml` and `helm-values` only reads values
|
||||
files, so neither sees a version written into a `helm upgrade` command —
|
||||
these are declared as `custom.regex` managers against the `helm` datasource.
|
||||
- CI linter versions in `.gitea/workflows/tool-versions.env`.
|
||||
|
||||
The Renovate image tag is deliberately _not_ in `tool-versions.env`:
|
||||
`renovate/k8s/cronjob.yaml` owns it, and the workflows read it from there.
|
||||
|
||||
## How updates flow
|
||||
|
||||
Renovate scans both `compose.yaml` files and Kubernetes manifests, opens a
|
||||
branch and PR with image tag changes, and waits for CI. After merge, the
|
||||
existing deployment workflow applies Kubernetes changes or redeploys Compose
|
||||
stacks. Renovate never updates running workloads directly.
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -19,7 +19,7 @@ spec:
|
||||
restartPolicy: Never
|
||||
containers:
|
||||
- name: renovate
|
||||
image: renovate/renovate:44.140.0
|
||||
image: renovate/renovate:44.147.0
|
||||
env:
|
||||
- name: RENOVATE_PLATFORM
|
||||
value: gitea
|
||||
|
||||
@@ -2,7 +2,7 @@ services:
|
||||
renovate:
|
||||
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
|
||||
# package rule in renovate/renovate.json.
|
||||
image: renovate/renovate:44.136.0
|
||||
image: renovate/renovate:44.147.0
|
||||
container_name: renovate
|
||||
restart: "no"
|
||||
env_file:
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# SearXNG
|
||||
|
||||
Search frontend with a separate Valkey cache.
|
||||
|
||||
Kubernetes keeps the application settings in a ConfigMap and starts Valkey as a
|
||||
StatefulSet. Set the secret from the example before exposing the search endpoint.
|
||||
There is no active marker.
|
||||
|
||||
Compose expects local configuration under `core-config/`, which is ignored.
|
||||
Prepare it before starting the stack; a container image alone does not supply
|
||||
this lab's settings.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n searxng
|
||||
kubectl get events -n searxng --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,18 @@
|
||||
# Media stack
|
||||
|
||||
Docker services for playback, requests, library management, and downloads.
|
||||
|
||||
Compose runs Jellyfin, Jellyseerr, Sonarr, Radarr, Prowlarr, qBittorrent, and the
|
||||
other services declared in the file. Kubernetes only routes to host endpoints;
|
||||
update `k8s/routing/external-service.yaml` when the Docker host or ports change.
|
||||
|
||||
Prepare the paths, user/group IDs, and credentials from `.env.example`. Service
|
||||
configuration and media/download directories are bind mounts. Preserve their
|
||||
permissions when moving data, and keep the application databases with backups.
|
||||
|
||||
Review device mounts for hardware acceleration before starting on another host.
|
||||
The Compose and Kubernetes routing files have no active markers, so automatic
|
||||
deploys do not select this stack. Start the Compose project or apply its routing
|
||||
resources manually when needed.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
Whitespace-only changes.
Whitespace-only changes.
@@ -0,0 +1,20 @@
|
||||
# Termix
|
||||
|
||||
Terminal and SSH connection manager with persistent application data.
|
||||
|
||||
Kubernetes stores state in `termix-pvc`; Compose mounts `termix-data/`.
|
||||
The application config and routes are committed separately under `k8s/`.
|
||||
|
||||
There is no active marker. Review access control and retain the application data
|
||||
needed to recover saved connections before enabling it.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n termix
|
||||
kubectl get events -n termix --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,24 @@
|
||||
"""Keep unit-test workflow commands out of the real CI job files."""
|
||||
|
||||
import os
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
CI_COMMAND_FILES = ('GITHUB_STEP_SUMMARY', 'GITHUB_OUTPUT', 'GITHUB_ENV', 'GITHUB_PATH', 'GITHUB_STATE')
|
||||
|
||||
|
||||
class IsolatedCITestCase(unittest.TestCase):
|
||||
def setUp(self):
|
||||
super().setUp()
|
||||
directory = tempfile.TemporaryDirectory(prefix='homelab-test-ci-')
|
||||
self.addCleanup(directory.cleanup)
|
||||
paths = {}
|
||||
for variable in CI_COMMAND_FILES:
|
||||
path = Path(directory.name) / variable
|
||||
path.touch()
|
||||
paths[variable] = str(path)
|
||||
environment = patch.dict(os.environ, paths)
|
||||
environment.start()
|
||||
self.addCleanup(environment.stop)
|
||||
@@ -0,0 +1,60 @@
|
||||
"""Run the real unit tests with external CI files and detect leaked writes."""
|
||||
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from ci_test_case import CI_COMMAND_FILES, IsolatedCITestCase
|
||||
|
||||
|
||||
class CIOutputIsolationTests(IsolatedCITestCase):
|
||||
def test_all_command_files_are_private_and_environment_is_restored(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
external = {variable: str(Path(scratch) / variable) for variable in CI_COMMAND_FILES}
|
||||
for path in external.values():
|
||||
Path(path).write_text('external CI file\n')
|
||||
with patch.dict(os.environ, external):
|
||||
probe = IsolatedCITestCase()
|
||||
probe.setUp()
|
||||
private = []
|
||||
try:
|
||||
for variable in CI_COMMAND_FILES:
|
||||
self.assertNotEqual(os.environ[variable], external[variable])
|
||||
path = Path(os.environ[variable])
|
||||
private.append(path)
|
||||
path.write_text('test-only command\n')
|
||||
finally:
|
||||
probe.doCleanups()
|
||||
for variable in CI_COMMAND_FILES:
|
||||
self.assertEqual(os.environ[variable], external[variable])
|
||||
self.assertEqual(Path(external[variable]).read_text(), 'external CI file\n')
|
||||
self.assertTrue(all(not path.exists() for path in private))
|
||||
|
||||
def test_unit_suite_preserves_external_ci_files(self):
|
||||
tests = Path(__file__).resolve().parent
|
||||
modules = sorted(p.stem for p in tests.glob('test_*.py') if p.name != Path(__file__).name)
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
environment = os.environ.copy()
|
||||
environment['PYTHONPATH'] = str(tests) + os.pathsep + environment.get('PYTHONPATH', '')
|
||||
expected = {}
|
||||
for variable in CI_COMMAND_FILES:
|
||||
path = Path(scratch) / variable
|
||||
content = f'external {variable}\n'
|
||||
path.write_text(content)
|
||||
environment[variable] = str(path)
|
||||
expected[path] = content
|
||||
result = subprocess.run( # noqa: S603 -- Run local test modules with the current Python interpreter.
|
||||
[sys.executable, '-m', 'unittest', *modules, '-q'],
|
||||
cwd=tests.parent,
|
||||
env=environment,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=False,
|
||||
timeout=60,
|
||||
)
|
||||
self.assertEqual(result.returncode, 0, result.stdout + result.stderr)
|
||||
for path, content in expected.items():
|
||||
self.assertEqual(path.read_text(), content, f'Unit tests wrote to external {path.name}')
|
||||
+101
-7
@@ -9,6 +9,8 @@ import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from ci_test_case import IsolatedCITestCase
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
@@ -34,7 +36,7 @@ def release(sha='a' * 40):
|
||||
}
|
||||
|
||||
|
||||
class ReleaseGateTests(unittest.TestCase):
|
||||
class ReleaseGateTests(IsolatedCITestCase):
|
||||
def test_release_rejects_wrong_sha_missing_images_and_mutable_tags(self):
|
||||
for mutation in ('sha', 'missing', 'tag'):
|
||||
data = release()
|
||||
@@ -91,8 +93,9 @@ class ReleaseGateTests(unittest.TestCase):
|
||||
api.release({'id': 1, 'head_sha': 'a' * 40})
|
||||
|
||||
|
||||
class SelectionTests(unittest.TestCase):
|
||||
class SelectionTests(IsolatedCITestCase):
|
||||
def setUp(self):
|
||||
super().setUp()
|
||||
self.scratch = tempfile.TemporaryDirectory()
|
||||
self.addCleanup(self.scratch.cleanup)
|
||||
self.repo = Path(self.scratch.name)
|
||||
@@ -128,6 +131,23 @@ class SelectionTests(unittest.TestCase):
|
||||
self.assertEqual(result['selected']['k8s'], ['one'])
|
||||
self.assertEqual(result['helm'], [])
|
||||
|
||||
def test_nested_service_change_and_owned_image_are_selected(self):
|
||||
directory = self.repo / 'vpn/xui/k8s'
|
||||
directory.mkdir(parents=True)
|
||||
(directory / 'active').touch()
|
||||
image = next(iter(release()['images']))
|
||||
(directory / 'app.yaml').write_text('image: ' + image + ':main\n')
|
||||
baseline_sha = self.commit()
|
||||
baseline = planner.make_plan(self.repo, self.repo, release(baseline_sha), None, 'full', [])
|
||||
(directory / 'app.yaml').write_text('image: ' + image + ':prod\n')
|
||||
result = planner.make_plan(self.repo, self.repo, release(self.commit()), baseline, 'changed', [])
|
||||
self.assertEqual(result['selected']['k8s'], ['vpn/xui'])
|
||||
baseline = result
|
||||
updated = release(result['sha'])
|
||||
updated['images'][image] = 'sha256:' + 'e' * 64
|
||||
result = planner.make_plan(self.repo, self.repo, updated, baseline, 'changed', [])
|
||||
self.assertEqual(result['selected']['k8s'], ['vpn/xui'])
|
||||
|
||||
def test_failed_intermediate_deploy_does_not_lose_changes(self):
|
||||
(self.repo / 'one/k8s/app.yaml').write_text('kind: StatefulSet\n')
|
||||
self.commit() # This commit failed deploy: baseline must remain initial.
|
||||
@@ -154,7 +174,7 @@ class SelectionTests(unittest.TestCase):
|
||||
self.assertEqual(result['selected']['k8s'], ['one', 'postgres', 'two'])
|
||||
|
||||
|
||||
class ComposeConfigurationTests(unittest.TestCase):
|
||||
class ComposeConfigurationTests(IsolatedCITestCase):
|
||||
def test_pin_preserves_project_volumes_paths_and_previous_image(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
root = Path(scratch)
|
||||
@@ -162,7 +182,8 @@ class ComposeConfigurationTests(unittest.TestCase):
|
||||
source = run / 'source'
|
||||
config_repo = root / 'persistent'
|
||||
(source / 'headscale').mkdir(parents=True)
|
||||
config_repo.mkdir()
|
||||
(config_repo / 'headscale').mkdir(parents=True)
|
||||
(config_repo / 'headscale/compose.yaml').touch()
|
||||
(run / 'release.json').write_text(json.dumps(release()))
|
||||
old = 'busybox@sha256:' + 'd' * 64
|
||||
new = 'busybox@sha256:' + 'e' * 64
|
||||
@@ -180,19 +201,41 @@ class ComposeConfigurationTests(unittest.TestCase):
|
||||
'volumes': {'data': {'name': 'headscale_data'}},
|
||||
}
|
||||
|
||||
previous_config = json.loads(json.dumps(config))
|
||||
previous_config['services']['app']['command'] = ['old-command']
|
||||
previous_config['services']['app']['environment'] = {'VALUE': 'old'}
|
||||
previous_config['services']['removed'] = {'image': 'busybox:latest'}
|
||||
config['services']['app']['command'] = ['new-command']
|
||||
config['services']['app']['environment'] = {'VALUE': 'new'}
|
||||
config['services']['added'] = {'image': 'busybox:latest'}
|
||||
|
||||
def fake_output(*args, **kwargs):
|
||||
if args[:2] == ('docker', 'compose'):
|
||||
self.assertEqual(kwargs['cwd'], config_repo)
|
||||
self.assertIn(str(config_repo / 'headscale'), args)
|
||||
return json.dumps(config)
|
||||
if '--hash' in args:
|
||||
return 'app matching-hash'
|
||||
return json.dumps(
|
||||
previous_config if str(config_repo / 'headscale/compose.yaml') in args else config
|
||||
)
|
||||
if args[:2] == ('docker', 'ps'):
|
||||
return 'container'
|
||||
if args[:2] == ('docker', 'inspect'):
|
||||
if 'com.docker.compose.config-hash' in args[-1]:
|
||||
return 'matching-hash'
|
||||
return 'sha256:' + 'f' * 64
|
||||
return json.dumps([old])
|
||||
|
||||
with (
|
||||
patch.dict(os.environ, {'CONFIG_REPO': str(config_repo), 'REPO': str(source), 'RUN_DIR': str(run)}),
|
||||
patch.dict(
|
||||
os.environ,
|
||||
{
|
||||
'CONFIG_REPO': str(config_repo),
|
||||
'REPO': str(source),
|
||||
'RUN_DIR': str(run),
|
||||
'HOMELAB_STATE': str(root / 'state'),
|
||||
},
|
||||
),
|
||||
patch.object(compose_module, 'output', side_effect=fake_output),
|
||||
patch.object(compose_module, 'resolve', return_value=new),
|
||||
):
|
||||
@@ -204,6 +247,57 @@ class ComposeConfigurationTests(unittest.TestCase):
|
||||
self.assertEqual(pinned['services']['app']['volumes'], config['services']['app']['volumes'])
|
||||
self.assertEqual(pinned['services']['app']['image'], new)
|
||||
self.assertEqual(before['services']['app']['image'], old)
|
||||
self.assertEqual(before['services']['app']['command'], ['old-command'])
|
||||
self.assertEqual(before['services']['app']['environment'], {'VALUE': 'old'})
|
||||
self.assertIn('removed', before['services'])
|
||||
self.assertNotIn('added', before['services'])
|
||||
|
||||
def mismatched_output(*args, **kwargs):
|
||||
if args[:2] == ('docker', 'inspect') and 'com.docker.compose.config-hash' in args[-1]:
|
||||
return 'different-hash'
|
||||
return fake_output(*args, **kwargs)
|
||||
|
||||
with (
|
||||
patch.dict(
|
||||
os.environ,
|
||||
{
|
||||
'CONFIG_REPO': str(config_repo),
|
||||
'REPO': str(source),
|
||||
'RUN_DIR': str(run),
|
||||
'HOMELAB_STATE': str(root / 'state'),
|
||||
},
|
||||
),
|
||||
patch.object(compose_module, 'output', side_effect=mismatched_output),
|
||||
patch.object(compose_module, 'resolve', return_value=new),
|
||||
self.assertRaisesRegex(ValueError, 'differs from running config'),
|
||||
):
|
||||
compose_module.prepare(source / 'headscale/compose.yaml')
|
||||
state = root / 'state'
|
||||
with patch.object(controller, 'STATE', state):
|
||||
state.mkdir()
|
||||
(run / 'status.json').write_text('{"state": "running", "stages": {}}')
|
||||
with patch.object(controller, 'retain_completed'):
|
||||
controller.finish_success(run, {})
|
||||
self.assertEqual(json.loads((state / 'compose-configs/headscale.json').read_text()), pinned)
|
||||
# A stale persistent checkout must not replace the successful baseline.
|
||||
with (
|
||||
patch.dict(
|
||||
os.environ,
|
||||
{
|
||||
'CONFIG_REPO': str(config_repo),
|
||||
'REPO': str(source),
|
||||
'RUN_DIR': str(run),
|
||||
'HOMELAB_STATE': str(state),
|
||||
},
|
||||
),
|
||||
patch.object(compose_module, 'output', side_effect=fake_output),
|
||||
patch.object(compose_module, 'resolve', return_value=new),
|
||||
):
|
||||
compose_module.prepare(source / 'headscale/compose.yaml')
|
||||
before = json.loads((run / 'compose-before/headscale.json').read_text())
|
||||
self.assertEqual(before['services']['app']['command'], ['new-command'])
|
||||
self.assertIn('added', before['services'])
|
||||
self.assertNotIn('removed', before['services'])
|
||||
self.assertEqual((run / 'compose/headscale.json').stat().st_mode & 0o777, 0o600)
|
||||
|
||||
def test_registry_index_and_single_image_descriptors(self):
|
||||
@@ -214,7 +308,7 @@ class ComposeConfigurationTests(unittest.TestCase):
|
||||
)
|
||||
|
||||
|
||||
class ControllerTests(unittest.TestCase):
|
||||
class ControllerTests(IsolatedCITestCase):
|
||||
def test_completed_stage_cannot_apply_again(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
directory = Path(scratch)
|
||||
|
||||
@@ -10,10 +10,11 @@ import zipfile
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
from ci_test_case import IsolatedCITestCase
|
||||
from test_cicd import ROOT, controller, release, release_module
|
||||
|
||||
|
||||
class ArtifactTests(unittest.TestCase):
|
||||
class ArtifactTests(IsolatedCITestCase):
|
||||
def test_archive_rejects_nested_or_extra_files(self):
|
||||
api = object.__new__(release_module.Gitea)
|
||||
api.base = 'https://example.test/api/v1/repos/a/b'
|
||||
@@ -103,7 +104,7 @@ class ArtifactTests(unittest.TestCase):
|
||||
self.assertEqual(json.loads((root / 'error-pages.json').read_text())['sha'], 'e' * 40)
|
||||
|
||||
|
||||
class DurableRunTests(unittest.TestCase):
|
||||
class DurableRunTests(IsolatedCITestCase):
|
||||
def test_duplicate_start_only_reattaches(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
state = Path(scratch)
|
||||
@@ -203,7 +204,7 @@ class DurableRunTests(unittest.TestCase):
|
||||
self.assertEqual(json.loads((directory / 'status.json').read_text())['state'], 'failure')
|
||||
|
||||
|
||||
class FailureSummaryTests(unittest.TestCase):
|
||||
class FailureSummaryTests(IsolatedCITestCase):
|
||||
def test_build_failure_keeps_progress_and_does_not_expose_exception_text(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
summary = Path(scratch) / 'summary.md'
|
||||
@@ -225,6 +226,77 @@ class FailureSummaryTests(unittest.TestCase):
|
||||
self.assertIn('xdfnx-homepage', content)
|
||||
self.assertNotIn('private value', content)
|
||||
|
||||
def test_invalid_digest_is_not_reported_as_a_completed_image(self):
|
||||
for digest in ('invalid-private-metadata', None, ['invalid']):
|
||||
with self.subTest(digest=digest), tempfile.TemporaryDirectory() as scratch:
|
||||
root = Path(scratch)
|
||||
summary = root / 'summary.md'
|
||||
name = 'error-pages'
|
||||
context, dockerfile = release_module.IMAGES[name]
|
||||
plan = {
|
||||
'sha': 'a' * 40,
|
||||
'targets': [
|
||||
{
|
||||
'name': name,
|
||||
'context': context,
|
||||
'dockerfile': dockerfile,
|
||||
'inputs': 'c' * 64,
|
||||
'reuse_digest': None,
|
||||
}
|
||||
],
|
||||
}
|
||||
|
||||
def fake_command(*args, digest=digest, **_kwargs):
|
||||
if args[:3] == ('docker', 'buildx', 'build'):
|
||||
Path(args[args.index('--metadata-file') + 1]).write_text(
|
||||
json.dumps({'containerimage.digest': digest})
|
||||
)
|
||||
return ''
|
||||
|
||||
with (
|
||||
patch.dict(
|
||||
os.environ,
|
||||
{
|
||||
'GITHUB_STEP_SUMMARY': str(summary),
|
||||
'GITHUB_SHA': 'a' * 40,
|
||||
'REGISTRY_USERNAME': 'test',
|
||||
'REGISTRY_PASSWORD': 'placeholder',
|
||||
},
|
||||
),
|
||||
patch.object(release_module, 'checked_plan', return_value=plan),
|
||||
patch.object(release_module.Path, 'home', return_value=root),
|
||||
patch.object(release_module, 'command', side_effect=fake_command),
|
||||
patch.object(subprocess, 'run', return_value=subprocess.CompletedProcess([], 0)),
|
||||
self.assertRaisesRegex(ValueError, 'invalid digest'),
|
||||
):
|
||||
release_module.build(root / 'image.json', name, root / 'plan.json')
|
||||
self.assertFalse((root / 'image.json').exists())
|
||||
content = summary.read_text()
|
||||
self.assertIn('**failure**', content)
|
||||
self.assertIn('### Built\n- None', content)
|
||||
self.assertIn('### Completed image digests\n- None', content)
|
||||
self.assertNotIn('invalid-private-metadata', content)
|
||||
|
||||
def test_successful_image_result_does_not_claim_complete_release(self):
|
||||
def complete_image(_output, report, _name, _plan):
|
||||
report.update(phase='Image result file saved', built=['error-pages'])
|
||||
report['images']['gcr.forust.xyz/forust/error-pages'] = 'sha256:' + 'b' * 64
|
||||
|
||||
with (
|
||||
patch.dict(os.environ, {'GITHUB_SHA': 'a' * 40}),
|
||||
patch.object(
|
||||
release_module,
|
||||
'build_images',
|
||||
side_effect=complete_image,
|
||||
),
|
||||
):
|
||||
release_module.build(Path('unused.json'), 'error-pages', Path('unused-plan.json'))
|
||||
content = Path(os.environ['GITHUB_STEP_SUMMARY']).read_text()
|
||||
self.assertIn('## Image build result `error-pages`', content)
|
||||
self.assertIn('Commit: `' + 'a' * 40 + '`', content)
|
||||
self.assertIn('final build job must publish the complete release', content)
|
||||
self.assertNotIn('## Image release', content)
|
||||
|
||||
def test_deploy_failure_reports_completed_apply_and_rollback_result(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
state = Path(scratch)
|
||||
@@ -261,7 +333,7 @@ class FailureSummaryTests(unittest.TestCase):
|
||||
self.assertIn('Compose requires manual recovery', content)
|
||||
|
||||
|
||||
class InstallerTests(unittest.TestCase):
|
||||
class InstallerTests(IsolatedCITestCase):
|
||||
def test_version_comparison_is_exact_without_network_or_host_packages(self):
|
||||
with tempfile.TemporaryDirectory() as scratch:
|
||||
root = Path(scratch)
|
||||
|
||||
@@ -3,10 +3,10 @@
|
||||
import json
|
||||
import os
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, call, patch
|
||||
|
||||
from ci_test_case import IsolatedCITestCase
|
||||
from test_cicd import release, release_module
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ def plan_data(changed):
|
||||
return {'sha': 'a' * 40, 'targets': targets}
|
||||
|
||||
|
||||
class MatrixTests(unittest.TestCase):
|
||||
class MatrixTests(IsolatedCITestCase):
|
||||
def test_no_change_one_image_all_images_and_missing_baseline(self):
|
||||
for changed in (set(), {'error-pages'}, set(release_module.IMAGES)):
|
||||
with self.subTest(changed=changed), tempfile.TemporaryDirectory() as scratch:
|
||||
|
||||
@@ -7,11 +7,14 @@ import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
from ci_test_case import IsolatedCITestCase
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
class NetbirdRuntimeTests(unittest.TestCase):
|
||||
class NetbirdRuntimeTests(IsolatedCITestCase):
|
||||
def setUp(self):
|
||||
super().setUp()
|
||||
self.temp = tempfile.TemporaryDirectory()
|
||||
self.addCleanup(self.temp.cleanup)
|
||||
self.root = Path(self.temp.name)
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
# Traefik
|
||||
|
||||
Ingress for HTTP, gRPC, TCP, and UDP services, with public and internal TLS.
|
||||
|
||||
Kubernetes uses the Helm settings in `k8s/traefik-values.yaml`. The deploy
|
||||
library applies supporting resources in this directory but does not install or
|
||||
upgrade the Traefik chart. Bootstrap the chart and CRDs separately.
|
||||
|
||||
The LoadBalancer address is set to `192.168.80.2`. Change it for another network.
|
||||
Entrypoints include web traffic, Gitea SSH, NetBird STUN, and other lab protocols.
|
||||
Public certificates come from cert-manager; internal certificates use the lab CA.
|
||||
The file provider reads `traefik-dynamic` through an additional volume and flags.
|
||||
|
||||
## API access
|
||||
|
||||
The committed chart values enable `api.insecure` and expose TCP 8080 through the
|
||||
LoadBalancer for Homarr integration. That listener has no Traefik authentication.
|
||||
Its reachability depends on external network controls. Review those controls
|
||||
before deploying these values outside the trusted network.
|
||||
|
||||
The normal dashboard IngressRoute is a separate path; protecting that route does
|
||||
not protect the direct port 8080 listener.
|
||||
|
||||
## Compose alternative
|
||||
|
||||
Compose mounts static and dynamic config, certificates, ACME state, and the
|
||||
Docker socket. It needs the external `proxy` network. Local file-server routing
|
||||
and TLS files have `.example` templates; copy only the ones needed for the host.
|
||||
|
||||
Keep ACME state and private keys with backups. Changing ingress values can affect
|
||||
every service at once, so inspect routes and entrypoints after an upgrade.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
traefik:
|
||||
image: traefik:v3.7.13
|
||||
image: traefik:v3.7.14
|
||||
container_name: traefik
|
||||
restart: unless-stopped
|
||||
command:
|
||||
|
||||
@@ -3,7 +3,7 @@ hostNetwork: false
|
||||
image:
|
||||
registry: docker.io/library
|
||||
repository: traefik
|
||||
tag: v3.7.13
|
||||
tag: v3.7.14
|
||||
|
||||
securityContext:
|
||||
capabilities:
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# Uptime Kuma
|
||||
|
||||
Service checks, status pages, and Prometheus metrics.
|
||||
|
||||
Kubernetes keeps state in `uptime-kuma-pvc` and exposes metrics through a
|
||||
ServiceMonitor. The metrics credentials come from the local Secret example.
|
||||
`alerts.yaml` adds Prometheus rules; a running Kuma UI alone does not establish
|
||||
that Prometheus is scraping it.
|
||||
|
||||
Compose stores state in `data/`. Back up that application database and verify
|
||||
notification delivery after restoring it.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n uptime-kuma
|
||||
kubectl get events -n uptime-kuma --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,23 @@
|
||||
# Vaultwarden
|
||||
|
||||
Password vault server with persistent data and public/internal ingress.
|
||||
|
||||
Kubernetes uses `vaultwarden-pvc` and the public URL from a ConfigMap.
|
||||
Compose has a separate data volume. Preserve the database, attachments, and keys
|
||||
as part of the same backup.
|
||||
|
||||
The env example only sets `DOMAIN`; there is no tracked administrator Secret
|
||||
example. Configure any administrator token separately and keep it out of Git. Verify
|
||||
sign-in and client synchronization after any update. Do not use a successful
|
||||
container restart as the only restore check.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n vaultwarden
|
||||
kubectl get events -n vaultwarden --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,5 @@
|
||||
# VPN services
|
||||
|
||||
[x-ui](xui/README.md) contains the Kubernetes deployment for 3x-ui. This directory
|
||||
has no shared Compose stack. Active markers are checked at each service's `k8s/`
|
||||
level, including nested paths.
|
||||
@@ -0,0 +1,18 @@
|
||||
# 3x-ui
|
||||
|
||||
Kubernetes deployment for the 3x-ui management panel in namespace `xui`.
|
||||
The `k8s/active` marker includes it in normal deploy selection.
|
||||
|
||||
The workload, data mounts, and ports are in `k8s/xui.yaml`; panel routing and TLS
|
||||
are in the ingress and certificate files. Keep panel access and proxy protocol
|
||||
ports separate when changing the configuration.
|
||||
|
||||
Back up the application's database and keys before upgrades. Check the actual
|
||||
host and volume paths before moving the workload to another node.
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n xui
|
||||
kubectl get events -n xui --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../../README.md) for deployment selection.
|
||||
Reference in new issue
Block a user