Compare commits

..
1 Commits
Author SHA1 Message Date
renovate-bot 183daf14d4 chore(deps): update renovate/renovate docker tag to v44.101.0
ci / lint-prettier (pull_request) Successful in 7s
ci / lint-ruff (pull_request) Successful in 4s
ci / lint-yaml (pull_request) Successful in 8s
ci / lint-dockerfiles (pull_request) Successful in 5s
ci / validate (pull_request) Successful in 5s
ci / lint-prettier (push) Successful in 8s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 7s
ci / build (push) Has been skipped
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 4s
renovate-ci / validate-renovate (pull_request) Successful in 8s
ci / build (pull_request) Has been skipped
ci / deploy-userbot-panel (pull_request) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
2026-09-18 10:18:01 +00:00
437 changed files with 26641 additions and 10971 deletions

No files matched your search

-10
View File
@@ -1,10 +0,0 @@
# actionlint configuration. Passed explicitly from the ci workflow:
# actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
#
# The self-hosted act_runner registers custom labels that actionlint cannot know
# about, so declare them here instead of silencing the whole runner-label check.
self-hosted-runner:
labels:
- arch
- homelab
- prod
-3
View File
@@ -1,3 +0,0 @@
{
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
}
-122
View File
@@ -1,122 +0,0 @@
# Homelab CI/CD
The native Gitea runner runs on **vps**; production runs on **workstation**.
Jobs run on `homelab:host`, one at a time. No job images or Kubernetes credentials
are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
From this checkout on the VPS:
```sh
sudo bash .gitea/runner/setup-runner.sh
```
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
For a new host, install the same runner binary and register as `gitea-runner`
using the registration token interactively, label `homelab:host`, and working
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
## Workstation setup
As the existing SSH deploy user on workstation:
```sh
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
```
The controller uses `/srv/homelab` as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
`~/.config/homelab-deploy/environment`. Check these before installing.
Configure Gitea Actions Variables:
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
## Releases and deployment
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. EDU images remain pinned to the
digests released by their application repository. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA:
- `full`: required for the first baseline; reconcile all active components.
- `changed`: compare with the last fully successful production deploy.
- `plan`: validate configuration and show selection without changing production resources.
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
## Status and recovery
`--retry` repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
```sh
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
```
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using `compose-before/<stack>.json`, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures.
Update the workstation dispatcher only when no deploy is running.
## Validation and migration rollback
```sh
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
```
Test on a separate namespace before the initial production `full` run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
-11
View File
@@ -1,11 +0,0 @@
[worker.oci]
gc = true
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
[[worker.oci.gcpolicy]]
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
all = true
-10
View File
@@ -1,10 +0,0 @@
runner:
file: /var/lib/gitea-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab:host
cache:
enabled: false
container:
docker_host: unix:///var/run/docker.sock
-18
View File
@@ -1,18 +0,0 @@
[Unit]
Description=Gitea Actions runner
After=network-online.target docker.service
Wants=network-online.target
[Service]
User=gitea-runner
Group=gitea-runner
SupplementaryGroups=docker
WorkingDirectory=/var/lib/gitea-runner
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
Restart=on-failure
RestartSec=5
UMask=0077
[Install]
WantedBy=multi-user.target
-12
View File
@@ -1,12 +0,0 @@
[Unit]
Description=Homelab deploy %i
[Service]
Type=exec
EnvironmentFile=%h/.config/homelab-deploy/environment
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
RuntimeMaxSec=5h
TimeoutStopSec=135min
KillMode=control-group
UMask=0077
-48
View File
@@ -1,48 +0,0 @@
#!/usr/bin/env bash
# Native host runner, with pinned user-space tools and no extra CI images.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in docker curl python3 git tar xz flock runuser systemctl; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
docker info >/dev/null
docker compose version >/dev/null
docker buildx version >/dev/null
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
# Reuse the established service account and runner registration.
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
mkdir -p /etc/gitea-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
chmod 755 "$scratch"
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
python3 - <<'PYLABELS'
import json
from pathlib import Path
registration = Path('/var/lib/gitea-runner/.runner')
if registration.exists():
labels = json.loads(registration.read_text()).get('labels', [])
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
labels.append('homelab:host')
config = Path('/etc/gitea-runner/config.yaml')
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
PYLABELS
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
if [ ! -f /var/lib/gitea-runner/.runner ]; then
echo 'Register once as gitea-runner with homelab:host before starting the service.'
exit 0
fi
systemctl daemon-reload
systemctl enable --now gitea-runner.service
systemctl restart gitea-runner.service
echo "Runner ready. Configuration backups: *.before-$stamp"
-32
View File
@@ -1,32 +0,0 @@
#!/usr/bin/env bash
# Run as the existing deploy user on workstation. Never resets the working tree.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="${HOMELAB_REPO:-/srv/homelab}"
for tool in python3 git kubectl helm docker flock timeout; do
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
done
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
echo "Run once: sudo loginctl enable-linger $USER" >&2
exit 1
fi
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
chmod 700 "$config"
if [ ! -f "$config/environment" ]; then
context="$(kubectl config current-context)"
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
chmod 600 "$config/environment"
fi
# Do not replace a dispatcher while an existing deploy uses it.
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
echo 'An existing deploy is running; wait before updating the controller' >&2
exit 1
fi
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
systemctl --user daemon-reload
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
-94
View File
@@ -1,94 +0,0 @@
#!/usr/bin/env bash
# Local regressions only: kubectl is mocked and Docker is used for config parsing.
set -euo pipefail
repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate"
git -C "$scratch/repo" init -q
for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do
touch "$scratch/repo/$file"
done
git -C "$scratch/repo" add .
# shellcheck source=../workflows/compose-lint.sh
source "$repo/.gitea/workflows/compose-lint.sh"
actual="$(cd "$scratch/repo" && compose_files)"
expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml'
[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; }
cat >"$scratch/compose.yaml" <<'YAML'
services:
example:
image: busybox:1.37.0
environment:
REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression}
YAML
unset HOMELAB_TEST_REQUIRED
if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then
echo 'Full Compose validation accepted a missing variable' >&2
exit 1
fi
grep -q 'required for this regression' "$scratch/config.log"
HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml"
cat >"$scratch/resources.json" <<'JSON'
{"kind":"List","items":[
{"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{
"containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}],
"initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}],
"imagePullSecrets":[{"name":"registry"}],
"volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}]
}}}},
{"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}},
{"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}}
]}
JSON
actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)"
expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron'
[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; }
REPO="$repo"
# shellcheck source=../workflows/deploy-lib.sh
source "$repo/.gitea/workflows/deploy-lib.sh"
K8S_MANIFESTS=("$scratch/resources.json")
KUSTOMIZE_APPS=()
# No live cluster access. Reject credentials in app even if they exist elsewhere.
kubectl() {
case "$1" in
create) cat "$scratch/resources.json" ;;
get)
if [ "$3" = credentials ] && [ "$5" = app ]; then
return 1
fi
return 0
;;
*) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;;
esac
}
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Namespace-scoped Secret check accepted a missing Secret' >&2
exit 1
fi
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
# API/rendering errors must not produce an empty reference list and pass.
kubectl() { return 1; }
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
exit 1
fi
kubectl() { return 0; }
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight skipped an installed CRD' >&2
exit 1
fi
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
echo 'VMAgent preflight skipped an unrelated manifest' >&2
exit 1
fi
kubectl() { return 1; }
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2
exit 1
fi
printf '%s\n' 'Deploy validation regressions passed.'
+300 -102
View File
@@ -1,80 +1,25 @@
name: ci name: ci
"on":
on:
push: push:
branches: branches:
- main - "**"
pull_request: null pull_request:
workflow_dispatch: null workflow_dispatch:
permissions:
contents: read env:
actions: read REGISTRY: gcr.forust.xyz
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
jobs: jobs:
checks: lint-prettier:
runs-on: homelab runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 30
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@v4
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
echo "$tools_dir" >> "$GITHUB_PATH"
- name: Validate Compose files
shell: bash
run: |
set -euo pipefail
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
echo "docker compose config ${safe_flags[*]-}"
mapfile -t files < <(compose_files)
if [ "${#files[@]}" -eq 0 ]; then
echo "No Compose files found."
exit 0
fi
failed=0
for f in "${files[@]}"; do
if ! out="$(validate_compose_file "$f" ${safe_flags[@]+"${safe_flags[@]}"} 2>&1)"; then
failed=1
echo "::error file=${f}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Compose validation failed."
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash'
)
if [ "${#scripts[@]}" -eq 0 ]; then
echo "No shell scripts found."
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
bash .gitea/tests/deploy-validation.sh
- name: Check formatting with Prettier - name: Check formatting with Prettier
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t prettier_files < <( mapfile -t prettier_files < <(
git ls-files \ git ls-files \
| grep -E '\.(md|json|ya?ml|html|css)$' \ | grep -E '\.(md|json|ya?ml|html|css)$' \
@@ -86,19 +31,36 @@ jobs:
exit 0 exit 0
fi fi
prettier --check --ignore-unknown "${prettier_files[@]}" docker run --rm \
- name: Lint and format-check Python with Ruff -v "$PWD:/work" \
-w /work \
node:22-alpine \
sh -lc 'npx --yes prettier@3 --check --ignore-unknown "$@"' sh "${prettier_files[@]}"
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint Python with Ruff
shell: bash shell: bash
run: | run: |
set -euo pipefail docker run --rm \
ruff check . .gitea/workflows -v "$PWD:/work" \
ruff format --check . .gitea/workflows -w /work \
python3 -m unittest discover -s tests -v ghcr.io/astral-sh/ruff:latest \
check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint YAML syntax - name: Lint YAML syntax
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t yaml_files < <( mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \ git ls-files '*.yaml' '*.yml' \
':!node_modules/**' \ ':!node_modules/**' \
@@ -110,12 +72,21 @@ jobs:
exit 0 exit 0
fi fi
yamllint -c .yamllint "${yaml_files[@]}" docker run --rm \
-v "$PWD:/work" \
-w /work \
cytopia/yamllint:latest \
-c .yamllint "${yaml_files[@]}"
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint Dockerfiles - name: Lint Dockerfiles
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t dockerfiles < <( mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*' git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
) )
@@ -125,12 +96,22 @@ jobs:
exit 0 exit 0
fi fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}" docker run --rm \
- name: Validate Kubernetes manifests against JSON schemas -v "$PWD:/work" \
-w /work \
--entrypoint hadolint \
hadolint/hadolint:latest-debian \
-c .hadolint.yaml "${dockerfiles[@]}"
validate:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Validate Kubernetes manifests
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t manifests < <( mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \ git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' | grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
@@ -141,32 +122,249 @@ jobs:
exit 0 exit 0
fi fi
kubeconform \ docker run --rm \
-v "$PWD:/work" \
-w /work \
ghcr.io/yannh/kubeconform:latest \
-strict \ -strict \
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
"${manifests[@]}" "${manifests[@]}"
build: build:
needs: needs: [lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
- checks if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev')
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main' runs-on: [self-hosted, linux, arch, homelab]
runs-on: homelab outputs:
timeout-minutes: 60 services: ${{ steps.services.outputs.services }}
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@v4
with: with:
fetch-depth: 0 fetch-depth: 0
- name: Build changed images and write release
env: - name: Detect changed docker-built services
GITEA_TOKEN: ${{ github.token }} id: services
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} shell: bash
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} run: |
run: python3 .gitea/workflows/release.py build base="${{ github.event.before }}"
- name: Store commit release if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf base="$(git rev-list --max-parents=0 HEAD)"
with: fi
name: release-${{ github.sha }}
path: release.json mapfile -t changed_files < <(git diff --name-only "$base" "${GITHUB_SHA}")
if-no-files-found: error
retention-days: 30 services=()
add_service() {
local name="$1"
local seen=0
for existing in "${services[@]}"; do
if [ "$existing" = "$name" ]; then
seen=1
break
fi
done
if [ "$seen" -eq 0 ]; then
services+=("$name")
fi
}
for file in "${changed_files[@]}"; do
case "$file" in
dtek_notif/*)
add_service dtek_notif
;;
errorpages/*)
add_service errorpages
;;
userbot/*)
add_service userbot
;;
homepages/*)
add_service homepages
;;
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
add_service edu_master
;;
esac
done
if [ "${#services[@]}" -eq 0 ]; then
echo "No docker-built services changed."
echo "services=" >> "$GITHUB_OUTPUT"
exit 0
fi
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
if: steps.services.outputs.services != ''
shell: bash
run: |
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "${REGISTRY}" \
-u "${{ secrets.REGISTRY_USERNAME }}" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
for service in "${services[@]}"; do
case "$service" in
dtek_notif)
image="${REGISTRY}/forust/dtek-notif"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" dtek_notif
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
errorpages)
image="${REGISTRY}/forust/error-pages"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
userbot)
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
for target in runtime panel; do
case "$target" in
runtime)
context="userbot"
image="${REGISTRY}/forust/userbot"
;;
panel)
context="userbot/panel"
image="${REGISTRY}/forust/userbot-panel"
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
homepages)
for service in forust xdfnx; do
case "$service" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" -f "homepages/Dockerfile.${service}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for service in session-keeper webinar-checker; do
case "$service" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
deploy-userbot-panel:
needs: build
if: github.ref_name == 'main' && contains(needs.build.outputs.services, 'userbot')
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Apply and roll out userbot panel
shell: bash
run: |
kubectl apply -f userbot/k8s/base/panel.yaml
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
# Keep legacy deployments (forust/anna) in sync with manifests; they have no replicas field, so apply leaves scaling to the user manager only.
kubectl apply -f userbot/k8s/base/userbots.yaml
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
-45
View File
@@ -1,45 +0,0 @@
#!/usr/bin/env bash
# Shared helpers for validating Compose files. Sourced both by steps in
# .gitea/workflows/ci.yaml and by deploy-lib.sh on the workstation.
#
# Two levels of checking, matching how the repo is structured:
#
# general every committed Compose file, active or not. Pure structure check:
# no ${VAR} interpolation, no .env lookup, no bind-mount path
# resolution. Disabled stacks deliberately have no .env in the repo
# and no values on the CI runner, so a full `config` run would fail on
# their `${VAR:?}` guards for reasons that have nothing to do with the
# change under review.
#
# full active stacks only, with interpolation and env-file resolution, so
# required variables and referenced files are actually resolved. Needs
# the gitignored .env files, so this only runs in the deploy workflow
# on the workstation.
#
# This file is meant to be sourced, not executed.
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*compose.yaml' '*compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
# Probed rather than hardcoded so an older Compose without --no-env-resolution
# still gets the flags it does support.
compose_safe_flags() {
local help flag
help="$(docker compose config --help 2>/dev/null || true)"
for flag in --no-interpolate --no-env-resolution --no-path-resolution; do
if printf '%s' "$help" | grep -q -- "$flag"; then
printf '%s\n' "$flag"
fi
done
}
# validate_compose_file <file> [extra docker compose config flags...]
validate_compose_file() {
local file="$1"
shift
docker compose -f "$file" config --quiet "$@"
}
-93
View File
@@ -1,93 +0,0 @@
#!/usr/bin/env python3
"""Resolve Compose images without changing project names or local bind paths."""
import json
import os
import re
import subprocess
import sys
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def resolve(reference):
if '@sha256:' in reference:
return reference
descriptor = json.loads(
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
)
digest = descriptor['digest']
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError(f'Invalid registry digest for {reference}')
# Strip tag only from the final path segment (registry ports are preserved).
repository = reference.rsplit('/', 1)
repository[-1] = repository[-1].split(':')[0]
return '/'.join(repository) + '@' + digest
def prepare(source_file):
config_repo = Path(os.environ['CONFIG_REPO'])
source_repo = Path(os.environ['REPO'])
directory = Path(os.environ['RUN_DIR'])
relative = source_file.relative_to(source_repo)
project_dir = config_repo / relative.parent
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
project = config['name']
previous_file = directory / 'previous.json'
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text())
before = json.loads(json.dumps(config))
for service, settings in config['services'].items():
reference = settings.get('image')
if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
if image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference]
else:
pinned = resolve(reference)
settings['image'] = pinned
locks[reference] = pinned
# Capture what is running, not the current value of its mutable tag.
ids = output(
'docker',
'ps',
'-aq',
'--filter',
f'label=com.docker.compose.project={project}',
'--filter',
f'label=com.docker.compose.service={service}',
).splitlines()
actual = set()
for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True)
destination = folder / f'{relative.parent.name}.json'
destination.write_text(json.dumps(data, indent=2) + '\n')
destination.chmod(0o600)
images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved')
print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
)
if __name__ == '__main__':
prepare(Path(sys.argv[1]))
-323
View File
@@ -1,323 +0,0 @@
#!/usr/bin/env python3
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
import argparse
import contextlib
import fcntl
import importlib.util
import json
import math
import os
import re
import shutil
import subprocess
import sys
import time
from pathlib import Path
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
def command(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def atomic_json(path, data):
temporary = path.with_suffix('.tmp')
temporary.write_text(json.dumps(data, indent=2) + '\n')
temporary.chmod(0o600)
temporary.replace(path)
@contextlib.contextmanager
def lock(name):
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
with (STATE / name).open('a') as stream:
fcntl.flock(stream, fcntl.LOCK_EX)
yield
def load_module(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def run_directory(run_id):
if not RUN_ID.fullmatch(run_id):
raise ValueError('Run ID must be numeric workflow-id and attempt')
return STATE / 'runs' / run_id
def start(run_id):
payload = sys.stdin.buffer.read(256 * 1024 + 1)
if len(payload) > 256 * 1024:
raise ValueError('Deploy request exceeds 256 KiB')
request = json.loads(payload)
sha = request['release']['sha']
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
raise ValueError('Invalid deploy SHA or mode')
if not isinstance(request['refresh_images'], bool):
raise ValueError('refresh_images must be boolean')
directory = run_directory(run_id)
with lock('prepare.lock'):
if (directory / 'request.json').exists():
if json.loads((directory / 'request.json').read_text()) != request:
raise ValueError('Run ID already belongs to a different request')
else:
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
if not (directory / 'source').exists():
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
raise ValueError('Prepared source does not match deploy SHA')
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
release_module.validate_release(request['release'], sha)
atomic_json(directory / 'release.json', request['release'])
atomic_json(directory / 'request.json', request)
if not (directory / 'status.json').exists():
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
# Starting an existing active or finished ID is idempotent; never re-apply it.
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
print(f'Accepted deploy {run_id} ({sha})')
def environment(directory):
request = json.loads((directory / 'request.json').read_text())
return {
**os.environ,
'REPO': str(directory / 'source'),
'CONFIG_REPO': str(CONFIG_REPO),
'RUN_DIR': str(directory),
'DEPLOY_SHA': request['release']['sha'],
'RELEASE_FILE': str(directory / 'release.json'),
'DEPLOY_PLAN': str(directory / 'plan.json'),
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
'ROLLOUT_PARALLELISM': '4',
}
def stage(directory, name, budget):
status = json.loads((directory / 'status.json').read_text())
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
return status['stages'][name]['result'] == 'success'
started = time.time()
status['stages'][name] = {'result': 'running', 'started': started}
atomic_json(directory / 'status.json', status)
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
with (directory / f'{name}.log').open('a') as log:
# timeout kills the whole stage process group, including children, before recovery.
result = subprocess.run( # noqa: S603, S607
[
shutil.which('timeout') or '/usr/bin/timeout',
'--signal=TERM',
'--kill-after=30s',
str(budget),
'bash',
str(script),
name,
],
env=environment(directory),
stdout=log,
stderr=subprocess.STDOUT,
check=False,
).returncode
status = json.loads((directory / 'status.json').read_text())
status['stages'][name].update(
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
)
atomic_json(directory / 'status.json', status)
return result == 0
def make_plan(directory):
source = directory / 'source'
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
helm = json.loads(command('helm', 'list', '--all', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose']
atomic_json(directory / 'plan.json', plan)
if previous:
atomic_json(directory / 'previous.json', previous)
# Local config is deliberately separate from the immutable Git source.
return plan
def finish_success(directory, plan):
# Repeating finalization after a crash is safe while holding deploy.lock.
plan['run_id'] = directory.name
path = directory / 'compose-images.json'
previous = directory / 'previous.json'
plan['compose-images'] = (
json.loads(path.read_text())
if path.exists()
else json.loads(previous.read_text()).get('compose-images', {})
if previous.exists()
else {}
)
atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success'
atomic_json(directory / 'status.json', status)
try:
retain_completed(directory)
except (OSError, subprocess.CalledProcessError) as error:
print(f'Retention deferred: {error}', flush=True)
def recover(directory, retry=False):
status = json.loads((directory / 'status.json').read_text())
if status['state'] in ('success', 'planned'):
return
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
return
if retry:
for name in ('verify-k8s', 'smoke'):
if status['stages'].get(name, {}).get('result') == 'failure':
del status['stages'][name]
atomic_json(directory / 'status.json', status)
snapshot = directory / 'snapshot/current'
if snapshot.exists():
stage(directory, 'verify-k8s', 7200)
stage(directory, 'smoke', 600)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'failure'
atomic_json(directory / 'status.json', status)
def execute(run_id):
directory = run_directory(run_id)
with lock('deploy.lock'):
status = json.loads((directory / 'status.json').read_text())
if status['state'] != 'queued':
return
# A crashed predecessor must be recovered before another apply begins.
for other in (STATE / 'runs').iterdir():
if (
other != directory
and (other / 'status.json').exists()
and json.loads((other / 'status.json').read_text())['state'] == 'running'
):
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running'
atomic_json(directory / 'status.json', status)
try:
plan = make_plan(directory)
print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True,
)
if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed')
if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'planned'
atomic_json(directory / 'status.json', status)
return
# Budget includes both rollout checks and rollback waves, plus API overhead.
count = int(
command(
'bash',
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
'workload-count',
env=environment(directory),
)
)
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy')
k8s_ok = stage(directory, 'apply-k8s', 2700)
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
verify_ok = stage(directory, 'verify-k8s', verify_budget)
smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
finish_success(directory, plan)
except Exception as error:
with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n')
recover(directory)
raise
def retain_completed(current):
finished = []
for directory in (STATE / 'runs').iterdir():
status_file = directory / 'status.json'
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
finished.append(directory)
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
if directory == current:
continue
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
shutil.rmtree(directory)
def follow(run_id, phase):
directory = run_directory(run_id)
groups = {
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
'verify': ('verify-k8s',),
'smoke': ('smoke',),
}
names = groups[phase]
offsets = {}
while True:
status = json.loads((directory / 'status.json').read_text())
for name in (*names, 'controller'):
path = directory / f'{name}.log'
if path.exists():
with path.open() as stream:
stream.seek(offsets.get(name, 0))
content = stream.read()
if content:
print(content, end='', flush=True)
offsets[name] = stream.tell()
stages = status['stages']
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
return all(stages[name]['result'] == 'success' for name in names)
if status['state'] in ('success', 'failure', 'planned'):
return status['state'] in ('success', 'planned')
time.sleep(3)
def main():
os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow'))
parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
args = parser.parse_args()
directory = run_directory(args.run_id)
if args.action == 'start':
start(args.run_id)
elif args.action == 'execute':
execute(args.run_id)
elif args.action == 'recover':
with lock('deploy.lock'):
recover(directory, retry=args.retry)
elif args.action == 'status':
print((directory / 'status.json').read_text())
if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif not follow(args.run_id, args.phase):
sys.exit(1)
if __name__ == '__main__':
main()
-981
View File
@@ -1,981 +0,0 @@
#!/usr/bin/env bash
# Workstation deploy stages; invoked by the durable controller against pinned source.
# REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF'
# source "$REPO/.gitea/workflows/deploy-lib.sh"
# run_stage "$STAGE"
# EOF
set -euo pipefail
: "${REPO:?REPO must be set}"
APPLY_PRUNE="${APPLY_PRUNE:-false}"
CONFIG_REPO="${CONFIG_REPO:-$REPO}"
# Exact SHA accepted by the CI gate for both manual and automatic deploys.
DEPLOY_SHA="${DEPLOY_SHA:-}"
# Handoff point between the apply stage (writes) and the verify stage (reads).
# Under the deploy user's own XDG state directory rather than /var/backups: the
# deploy is unprivileged, /var/backups does not exist on a minimal Arch host, and
# creating it would need root — which is why the first real deploy died here with
# "is not writable" before touching a single workload. $HOME comes from sshd.
DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state}/homelab-deploy}"
# Per-workload rollout budget and how many workloads to watch at once. The whole
# apply job has its own timeout-minutes as a backstop.
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-4}"
WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps"
log() {
echo "== $* =="
}
warn() {
echo "WARNING: $*" >&2
}
# Prune needs the complete desired set in one invocation. Per-file pruning
# treats resources from the other files as absent and can delete them.
check_prune_mode() {
if [ "$APPLY_PRUNE" = "true" ]; then
echo "ERROR: APPLY_PRUNE=true is unsupported by the per-file deploy loop." >&2
echo "Disable it; remove obsolete resources explicitly after review." >&2
return 1
fi
}
collect_k8s() {
git -C "$REPO" ls-files -- "$1" \
| grep -E '\.ya?ml$' \
| grep -Ev '/overlays/' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' \
| grep -Ev '(^|/)[^/]*secret[^/]*\.ya?ml$' \
| sort
}
kustomize_overlay() {
if [ -f "$1/overlays/prod/kustomization.yaml" ]; then
echo "$1/overlays/prod"
elif [ -f "$1/base/kustomization.yaml" ]; then
echo "$1/base"
elif [ -f "$1/kustomization.yaml" ]; then
echo "$1"
fi
}
selected_service() {
local kind="$1" service="$2" section=selected
[ -n "${DEPLOY_PLAN:-}" ] || return 0
if [ "${DEPLOY_SMOKE_ALL:-false}" = true ]; then section=active; fi
jq -e --arg kind "$kind" --arg service "$service" --arg section "$section" \
'.[$section][$kind] | index($service) != null' "$DEPLOY_PLAN" >/dev/null
}
# Resolve .env and relative binds on the persistent workstation tree. Locked
# JSON configs keep the same Compose project name and volume names.
compose() {
local cf="$1" locked project_dir
project_dir="$CONFIG_REPO/$(basename "$(dirname "$cf")")"
shift
locked="${RUN_DIR:-/nonexistent}/compose/$(basename "$(dirname "$cf")").json"
if [ -f "$locked" ]; then cf="$locked"; fi
(cd "$CONFIG_REPO" && docker compose --project-directory "$project_dir" -f "$cf" "$@")
}
select_manifests() {
K8S_MANIFESTS=()
KUSTOMIZE_APPS=()
COMPOSE_STACKS=()
local kd_rel kd overlay cf_rel cf f
while IFS= read -r kd_rel; do
kd="$REPO/$kd_rel"
selected_service k8s "${kd_rel%/k8s}" || continue
if [ ! -f "$kd/active" ]; then
echo "skip (no k8s/active): $kd_rel"
continue
fi
overlay="$(kustomize_overlay "$kd" || true)"
if [ -n "${overlay:-}" ]; then
echo "kustomize app: ${overlay#"$REPO"/}"
KUSTOMIZE_APPS+=("$overlay")
else
while IFS= read -r f; do
[ -n "$f" ] && K8S_MANIFESTS+=("$REPO/$f")
done < <(collect_k8s "$kd_rel" || true)
fi
done < <(
git -C "$REPO" ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
while IFS= read -r cf_rel; do
cf="$REPO/$cf_rel"
selected_service compose "$(dirname "$cf_rel")" || continue
if [ -f "$(dirname "$cf")/active" ]; then
echo "compose: $cf_rel"
COMPOSE_STACKS+=("$cf")
else
echo "skip (no root active): $cf_rel"
fi
done < <(git -C "$REPO" ls-files '*/compose.yaml' '*/compose.yml' compose.yaml compose.yml | sort)
}
# --- post-apply verification and rollback -------------------------------------
#
# A green `kubectl apply` says nothing about the cluster being healthy. These
# helpers watch exactly the workloads whose spec changed during this apply, and
# on failure roll them back to the revision that was running before, so a bad
# push to main cannot leave a service crash-looping.
#
# The workstation controller runs apply and verification as separate durable
# stages. Runner jobs only follow their logs. ExecStopPost recovers interrupted
# runs using the per-run snapshot, even when the SSH connection has gone away.
#
# Creates this run's snapshot directory and publishes it as the handoff point for
# the verify stage. Fails hard by design: a deploy that cannot record what it is
# about to change must not start, because then nothing can be rolled back for it
# automatically. Publishing happens before the first apply, so an apply killed
# mid-flight still leaves a usable baseline behind.
snapshot_dir() {
local stamp dir
stamp="$(date -u +%Y%m%dT%H%M%SZ)-${DEPLOY_SHA:-$(git -C "$REPO" rev-parse --short HEAD 2>/dev/null || echo unknown)}"
dir="$DEPLOY_SNAPSHOT_DIR/$stamp"
if ! mkdir -p "$DEPLOY_SNAPSHOT_DIR" 2>/dev/null || [ ! -w "$DEPLOY_SNAPSHOT_DIR" ]; then
echo "ERROR: $DEPLOY_SNAPSHOT_DIR is not writable." >&2
echo "The verify job needs it to learn which workloads this deploy touches." >&2
echo "Refusing to deploy without a way to roll back." >&2
return 1
fi
if ! mkdir -p "$dir" 2>/dev/null || [ ! -w "$dir" ]; then
echo "ERROR: cannot create snapshot dir $dir" >&2
return 1
fi
if ! printf '%s\n' "$dir" >"$DEPLOY_SNAPSHOT_DIR/current" 2>/dev/null; then
echo "ERROR: cannot publish the snapshot pointer at $DEPLOY_SNAPSHOT_DIR/current" >&2
return 1
fi
printf '%s\n' "$dir"
}
save_snapshot() {
local dir="$1" releases revision status
log "Saving pre-apply snapshot to $dir"
workload_generations >"$dir/generations.before" || return 1
kubectl get "$WORKLOAD_KINDS" -A -o json >"$dir/workloads.json" || return 1
kubectl get controllerrevisions.apps -A -o json >"$dir/controller-revisions.json" || return 1
jq --slurpfile revisions "$dir/controller-revisions.json" '
[.items[] | . as $w | {
kind: (.kind | ascii_downcase), namespace: .metadata.namespace, name: .metadata.name, uid: .metadata.uid,
revision: (if .kind == "Deployment" then (.metadata.annotations["deployment.kubernetes.io/revision"] // "0" | tonumber)
else ([$revisions[0].items[] | select(.metadata.namespace == $w.metadata.namespace)
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
releases="$(helm list --all -A -o json)" || return 1
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
if ! jq -e --arg r "$release" --arg n "$namespace" \
'any(.[]; .name == $r and .namespace == $n)' <<<"$releases" >/dev/null; then
continue
fi
helm status "$release" -n "$namespace" -o json >"$dir/helm-$release.json" || return 1
status="$(jq -r '.info.status' "$dir/helm-$release.json")"
if [ "$status" != deployed ]; then
# Never capture a pending/failed revision as the recovery target.
helm history "$release" -n "$namespace" -o json >"$dir/helm-$release.history.json" || return 1
revision="$(jq '[.[] | select(.status == "deployed" or .status == "superseded") | .revision] | max // 0' \
"$dir/helm-$release.history.json")"
jq --argjson revision "$revision" '.version = $revision' "$dir/helm-$release.json" >"$dir/helm-$release.tmp"
mv "$dir/helm-$release.tmp" "$dir/helm-$release.json"
fi
done
# The verify stage compares this against the commit it is deploying, to refuse
# rolling back against a baseline left by an earlier run. A snapshot we cannot
# attribute to a commit is unusable for that, so fail before anything is applied.
if ! git -C "$REPO" rev-parse HEAD >"$dir/commit" 2>/dev/null; then
echo "ERROR: cannot record the deploy commit in $dir/commit" >&2
return 1
fi
}
# Prints "<ns> <name> <kind> <generation>" for every workload in the cluster.
workload_generations() {
kubectl get "$WORKLOAD_KINDS" -A \
-o 'custom-columns=NS:.metadata.namespace,NAME:.metadata.name,KIND:.kind,GEN:.metadata.generation' \
--no-headers 2>/dev/null \
| awk 'NF >= 4 { printf "%s %s %s %s\n", $1, $2, tolower($3), $4 }'
}
# Prints "<kind> <ns> <name>" for every workload that is new or whose generation
# moved since the snapshot, i.e. the ones this apply actually touched.
changed_workloads() {
local before="$1"
local ns name kind gen old current
current="$(workload_generations)" || return 1
while read -r ns name kind gen; do
[ -n "${gen:-}" ] || continue
old="$(awk -v want_ns="$ns" -v want_name="$name" -v want_kind="$kind" \
'$1 == want_ns && $2 == want_name && $3 == want_kind { print $4; exit }' "$before" 2>/dev/null || true)"
if [ -n "${RUN_DIR:-}" ] && ! grep -qxF "$kind $ns $name" "$RUN_DIR/workload-refs"; then
continue
fi
if [ "$old" != "$gen" ]; then
printf '%s %s %s\n' "$kind" "$ns" "$name"
fi
done <<<"$current"
}
# Resolve owned image references exclusively from the checked CI artifact.
render_pinned() {
python3 "$REPO/.gitea/workflows/release.py" render
}
# verify_workloads <failed-file> <kind> <ns> <name> ...
# Watches every workload in parallel and records the ones that never became
# healthy. Returns non-zero if any of them failed.
verify_workloads() {
local failed_file="$1"
shift
[ "$#" -gt 0 ] || return 0
: >"$failed_file"
local running=0 pid kind ns name
local -a pids=()
for entry in "$@"; do
read -r kind ns name <<<"$entry"
(
if kubectl rollout status "${kind}/${name}" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" >/dev/null 2>&1; then
echo " ok: ${kind}/${ns}/${name}"
else
echo " FAILED: ${kind}/${ns}/${name}"
printf '%s %s %s\n' "$kind" "$ns" "$name" >>"$failed_file"
fi
) &
pids+=($!)
running=$((running + 1))
if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then
wait -n 2>/dev/null || true
running=$((running - 1))
fi
done
for pid in ${pids[@]+"${pids[@]}"}; do
wait "$pid" || true
done
# Non-zero when the file holds at least one failure, i.e. a workload never
# became healthy. `[ -s ]` alone is the opposite test and silently disabled
# every rollback this stage is meant to perform.
[ ! -s "$failed_file" ]
}
# rollback_workloads <failed-file>
# Restores the previous revision of every failed workload and waits for it to
# settle. Prints a report and returns non-zero if any workload is still unhealthy,
# so the operator knows manual recovery is required.
rollback_workloads() {
local failed_file="$1" snapshot kind ns name index=0 running=0 pid revision uid
local -a pids=()
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
while read -r kind ns name; do
[[ "$kind" =~ ^(deployment|statefulset|daemonset)$ ]] || continue
index=$((index + 1))
(
if kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.annotations}' | grep -q 'meta.helm.sh/release-name'; then
echo " skip (Helm recovery owns this workload): $kind/$ns/$name"
exit 1
fi
revision="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .revision' "$snapshot/revisions.json")"
uid="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .uid' "$snapshot/revisions.json")"
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]] || [ "$uid" != "$(kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.uid}')" ]; then
echo " no safe previous revision: $kind/$ns/$name (new or replaced workload)"
exit 1
fi
kubectl rollout undo "$kind/$name" -n "$ns" --to-revision="$revision" \
&& kubectl rollout status "$kind/$name" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s"
) >"$snapshot/rollback-$index.log" 2>&1 &
pids+=($!)
running=$((running + 1))
if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then
wait -n 2>/dev/null || true
running=$((running - 1))
fi
done <"$failed_file"
local recovered=0 unrecovered=0 i=0
for pid in "${pids[@]}"; do
i=$((i + 1))
if wait "$pid"; then recovered=$((recovered + 1)); else unrecovered=$((unrecovered + 1)); fi
cat "$snapshot/rollback-$i.log"
done
echo "ROLLED_BACK=$recovered" >>"$failed_file"
echo "UNRECOVERED=$unrecovered" >>"$failed_file"
[ "$unrecovered" -eq 0 ]
}
# Helm releases owned by this stage, one line each:
#
# release|chart|namespace|chart version|values file (rel. to $REPO)|active marker
#
# The chart version is the field Renovate keeps current. The helmv3 manager only
# understands Chart.yaml and the helm-values manager only values files, so a pin
# written straight into a `helm upgrade` command would never be updated: these
# have to be declared as custom.regex managers in renovate/renovate.json.
HELM_RELEASES=(
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
)
# "name url" for the Helm repository hosting a chart, empty if unknown.
helm_repo_for() {
case "$1" in
prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;;
grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;;
stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;;
victoriametrics/*) echo "victoriametrics https://victoriametrics.github.io/helm-charts" ;;
esac
}
# helm_release_status <release> <namespace>
# Prints the release status in lowercase (deployed, failed, pending-rollback,
# ...) or "not-found" when the release does not exist yet.
helm_release_status() {
local out
if ! out="$(helm status "$1" -n "$2" 2>&1)"; then
if [[ "$out" == *"release: not found"* ]]; then
echo "not-found"
return 0
fi
printf 'ERROR: cannot read Helm status: %s\n' "$out" >&2
return 1
fi
awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]'
}
# recover_pending_release <release> <namespace>
# Rolls a release out of a pending-* state left by a failed upgrade with --rollback-on-failure
# whose own rollback never completed. Without this every future upgrade errors
# out until a human runs `helm rollback`. Passes through releases that are not
# pending (deployed, failed, not-found). Returns non-zero when the release is
# still not recoverable, so the pipeline fails loud instead of wedging.
recover_pending_release() {
local release="$1" namespace="$2" status revision snapshot
status="$(helm_release_status "$release" "$namespace")" || return 1
case "$status" in
pending-upgrade|pending-rollback|pending-install)
log "Release $release is $status, rolling back to the last deployed revision"
revision=""
if [ -s "$DEPLOY_SNAPSHOT_DIR/current" ]; then
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
if [ -s "$snapshot/helm-$release.json" ]; then
revision="$(jq -r '.version' "$snapshot/helm-$release.json")"
fi
fi
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]]; then
echo "ERROR: no captured Helm revision for $release; manual recovery required"
return 1
fi
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
echo "WARN: helm rollback of $release did not complete"
return 1
fi
status="$(helm_release_status "$release" "$namespace")" || return 1
if [ "$status" != "deployed" ]; then
echo "WARN: $release is $status after rollback"
return 1
fi
;;
esac
return 0
}
# wait_for_calm <stage>
# The deploy itself is heavy enough to melt this single node (helm churn plus
# apply churn drove load past 40, killed netbird/ssh, left helm pending-*).
# Never pile a heavy step onto an already-hot node: wait up to 10 minutes for
# the 1-minute load average to drop below the ceiling, then proceed anyway
# with a warning so a permanently busy node cannot wedge the pipeline forever.
wait_for_calm() {
local load waited=0
while [ "$waited" -lt 600 ]; do
load="$(cut -d' ' -f1 /proc/loadavg | cut -d. -f1)"
if [ "$load" -lt 28 ]; then
return 0
fi
if [ "$((waited % 60))" -eq 0 ]; then
log "$1: load $load, waiting for calm (<28)..."
fi
sleep 15
waited=$((waited + 15))
done
echo "WARN: $1: node still loaded ($load) after 10m, proceeding anyway"
}
upgrade_helm_releases() {
local entry release chart namespace version values marker repo
for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
if [ -n "${DEPLOY_PLAN:-}" ] && ! jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null; then
echo "skip (unchanged Helm release): $release"
continue
fi
if [ ! -f "$REPO/$values" ] && [ -f "$CONFIG_REPO/$values" ]; then values="$CONFIG_REPO/$values"; else values="$REPO/$values"; fi
if [ ! -f "$REPO/$marker" ]; then
echo "skip (no $marker): $release"
continue
fi
if [ ! -f "$values" ]; then
echo "ERROR: $values is gitignored but missing on the workstation, restore it first."
return 1
fi
repo="$(helm_repo_for "$chart")"
if [ -z "$repo" ]; then
echo "ERROR: no Helm repository configured for chart $chart"
return 1
fi
helm repo add "${repo%% *}" "${repo#* }" >/dev/null
helm repo update "${repo%% *}" >/dev/null
log "Upgrading $release ($chart $version)"
wait_for_calm "helm $release"
# A previous run with --rollback-on-failure whose own rollback never finished leaves the
# release in pending-*, which blocks every future upgrade. Recover first
# so one wedged revision cannot wedge the pipeline forever.
if ! recover_pending_release "$release" "$namespace"; then
echo "ERROR: $release is stuck and automatic rollback did not recover it, run 'helm rollback $release -n $namespace' by hand."
return 1
fi
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade
# times out or the workloads it touches never become ready, so a bad chart
# bump is not left half applied. (--atomic was this combo; deprecated.)
if ! helm upgrade --install "$release" "$chart" \
--namespace "$namespace" \
--version "$version" \
--values "$values" \
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
echo "WARN: upgrade of $release failed, checking release state"
# --rollback-on-failure already attempted its own rollback; finish the job when that
# rollback never completed, otherwise the release stays pending-* and
# blocks every future run.
if ! recover_pending_release "$release" "$namespace"; then
echo "ERROR: upgrade of $release failed and the release did not recover, run 'helm rollback $release -n $namespace' by hand."
else
echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
fi
return 1
fi
done
}
stage_doctor() {
local tool entry release chart namespace version values marker
for tool in git docker kubectl helm jq curl timeout flock python3; do
command -v "$tool" >/dev/null || { echo "Missing workstation tool: $tool"; return 1; }
done
docker compose version >/dev/null
docker buildx version >/dev/null
[ "$(kubectl config current-context)" = "${KUBE_CONTEXT:?configure KUBE_CONTEXT}" ] || { echo "Unexpected Kubernetes context"; return 1; }
[ "$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" = "${EXPECTED_CLUSTER_UID:?configure EXPECTED_CLUSTER_UID}" ] || { echo "Unexpected Kubernetes cluster"; return 1; }
kubectl get --raw=/readyz --request-timeout=10s >/dev/null
[ "$(git -C "$REPO" rev-parse HEAD)" = "$DEPLOY_SHA" ] || return 1
select_manifests
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
[ -f "$REPO/$marker" ] || continue
[ -f "$REPO/$values" ] || [ -f "$CONFIG_REPO/$values" ] || { echo "Missing values: $values"; return 1; }
done
jq '{sha, selected, helm, removed}' "$DEPLOY_PLAN"
local cf
for cf in "${COMPOSE_STACKS[@]}"; do
compose "$cf" config --quiet
while IFS= read -r network; do
docker network inspect "$network" >/dev/null || return 1
done < <(compose "$cf" config --format json | jq -r '.networks // {} | to_entries[] | select(.value.external == true) | .value.name')
python3 "$REPO/.gitea/workflows/compose-release.py" "$cf"
done
local image refs m k
refs="$(
for m in "${K8S_MANIFESTS[@]}"; do render_pinned <"$m" || return 1; done
for k in "${KUSTOMIZE_APPS[@]}"; do kubectl kustomize "$k" | render_pinned || return 1; done
)" || return 1
refs="$(grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+@sha256:[0-9a-f]{64}' <<<"$refs" | sort -u || true)"
while IFS= read -r image; do
[ -n "$image" ] || continue
timeout 60s docker buildx imagetools inspect "$image" >/dev/null
done <<<"$refs"
}
# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are
# created by cert-manager and are not prerequisites for applying a Certificate.
check_referenced_secrets() {
local m k objects refs extracted ns name
local missing=()
refs=""
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
refs+="$extracted"$'\n'
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
objects="$(kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json)" || return 1
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
refs+="$extracted"$'\n'
done
while read -r ns name; do
[ -n "${name:-}" ] || continue
if kubectl get secret "$name" -n "$ns" -o name >/dev/null 2>&1; then
echo " ok: $ns/$name"
else
echo " MISSING OR UNREADABLE: $ns/$name"
missing+=("$ns/$name")
fi
done < <(printf '%s' "$refs" | sort -u)
if [ "${#missing[@]}" -gt 0 ]; then
echo "ERROR: required pod Secrets are missing or unreadable:"
printf ' - %s\n' "${missing[@]}"
echo "Create them in the listed namespaces from the service's secret example."
return 1
fi
}
# The VMAgent CRD is installed by the VictoriaMetrics Operator Helm release in
# stage_apply_k8s, after this preflight stage. Skip only its dry-run until then.
skip_uninstalled_vmagent_crd() {
local manifest="$1"
if [[ "$manifest" == "$REPO/prometheus-stack/k8s/vmagent.yaml" ]] \
&& ! kubectl get crd vmagents.operator.victoriametrics.com >/dev/null 2>&1; then
echo " skip: VMAgent CRD is installed by Helm during apply: ${manifest#"$REPO"/}"
return 0
fi
return 1
}
stage_validate() {
check_prune_mode || return 1
cd "$REPO"
select_manifests
local m k cf
# The deploy host has the local .env and secret files. Resolve them here so
# missing configuration fails before either apply job changes workloads.
# CI keeps the structure-only check for inactive stacks.
# shellcheck source=compose-lint.sh
source "$REPO/.gitea/workflows/compose-lint.sh"
log "Validate compose stacks"
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " config: $cf"
compose "$cf" config --quiet
done
log "Validate k8s manifests (kubectl dry-run=client)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=client -f "$m" >/dev/null
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
kubectl apply -k "$k" --dry-run=client >/dev/null
done
log "Validate k8s manifests (kubectl dry-run=server)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=server -f "$m" >/dev/null
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
kubectl apply -k "$k" --dry-run=server >/dev/null
done
log "Checking referenced Secrets exist"
echo " (deploy never applies *secret*.yaml; create missing ones manually)"
check_referenced_secrets
}
selected_workload_refs() {
local m k
for m in "${K8S_MANIFESTS[@]}"; do
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
kubectl create --dry-run=client --validate=false -f "$m" -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
done
for k in "${KUSTOMIZE_APPS[@]}"; do
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
done
}
stage_apply_k8s() {
check_prune_mode || return 1
cd "$REPO"
select_manifests >/dev/null
local ns_files=() other_files=() m k
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
case "$m" in
*/namespace.yaml|*/namespace.yml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
# Record what is about to change, and publish it for the verify job, before
# the first apply. Both are fatal on failure: see snapshot_dir.
selected_workload_refs >"$RUN_DIR/workload-refs"
local snapshot
snapshot="$(snapshot_dir)" || return 1
save_snapshot "$snapshot" || return 1
touch "$snapshot/ready"
if [ "${#ns_files[@]}" -gt 0 ]; then
log "Applying namespaces (${#ns_files[@]} files)"
for m in "${ns_files[@]}"; do
kubectl apply -f "$m"
done
fi
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
exit 1
fi
fi
upgrade_helm_releases
wait_for_calm "apply resources"
if [ "${#other_files[@]}" -gt 0 ]; then
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
for m in "${other_files[@]}"; do
log "Applying ${m#"$REPO"/}"
if ! render_pinned <"$m" | kubectl apply -f -; then
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
exit 1
fi
done
fi
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
exit 1
fi
done
# No verification here on purpose. This stage may be killed at any point by
# timeout-minutes, by the runner cancelling the job, or by a dropped SSH
# connection, and any code below that line would simply not run. stage_verify_k8s
# picks the work up from the snapshot instead.
log "Applied. Verification and rollback are the verify job's job, not this one's."
}
# Runs as its own workflow job, after apply-k8s (and apply-compose) are done —
# including when they failed, timed out or were cancelled. Reads the baseline the
# apply stage published and works out what it changed, watches those workloads,
# and rolls back the ones that never became healthy.
stage_verify_k8s() {
local pointer="$DEPLOY_SNAPSHOT_DIR/current"
local snapshot want have generations changed
local -a touched=()
if [ ! -s "$pointer" ]; then
echo "ERROR: no snapshot pointer at $pointer."
echo "The apply stage died before publishing any state, so there is no baseline to"
echo "tell which workloads it touched. Nothing can be rolled back automatically —"
echo "inspect the cluster by hand."
return 1
fi
snapshot="$(head -1 "$pointer")"
if [ ! -d "$snapshot" ] || [ ! -f "$snapshot/ready" ]; then
echo "ERROR: snapshot pointer refers to a missing directory: $snapshot"
return 1
fi
# Never trust the pointer blindly. If the apply stage was killed before it
# published its own snapshot, `current` still points at the previous deploy's
# baseline. Verifying against that would watch the wrong workloads and the
# rollback would revert the wrong revisions, so refuse instead.
want="${DEPLOY_SHA:-}"
if [ -z "$want" ]; then
want="$(git -C "$REPO" rev-parse HEAD 2>/dev/null || true)"
fi
have="$(cat "$snapshot/commit" 2>/dev/null || true)"
if [ -z "$want" ] || [ "$have" != "$want" ]; then
echo "ERROR: refusing to verify or roll back against a stale snapshot."
echo " snapshot: $snapshot"
echo " snapshot commit: ${have:-<missing>}"
echo " deploy commit: ${want:-<unknown>}"
return 1
fi
echo " snapshot: $snapshot (commit ${have:0:12})"
local entry release chart namespace version values marker
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null || continue
recover_pending_release "$release" "$namespace" || return 1
done
generations="$snapshot/generations.before"
if [ ! -s "$generations" ]; then
# Without a baseline we cannot tell which workloads the apply touched, so
# fall back to watching everything rather than silently skipping the check.
warn "no pre-apply baseline, verifying every workload in the cluster"
: >"$generations"
fi
changed="$(changed_workloads "$generations")" || return 1
while read -r kind ns name; do
[ -n "${kind:-}" ] && touched+=("$kind $ns $name")
done <<<"$changed"
log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)"
if [ "${#touched[@]}" -eq 0 ]; then
echo " nothing to verify"
return 0
fi
printf ' watching: %s\n' "${touched[@]/#/ }"
local failed_file="$snapshot/failed-workloads"
if ! verify_workloads "$failed_file" ${touched[@]+"${touched[@]}"}; then
echo
echo "ERROR: ${#touched[@]} workload(s) changed by this deploy, and these never became healthy:"
grep -v -E '^(ROLLED_BACK|UNRECOVERED)=' "$failed_file" | sed 's/^/ - /'
echo
log "Rolling back to the previous revision"
if rollback_workloads "$failed_file"; then
echo
echo "Rolled back successfully. The cluster is back on the pre-deploy revision."
echo "Nothing else was reverted: Git holds desired state only, so config changes, PVCs and"
echo "externally created resources from this commit are still in place. Review the failed"
echo "workload, then re-run the deploy (Actions -> deploy -> Run workflow)."
else
echo
echo "Rollback did NOT fully recover the cluster. Manual intervention required:"
grep -E '^(ROLLED_BACK|UNRECOVERED)=' "$failed_file" | sed 's/^/ /'
echo "Pre-apply snapshot: $snapshot"
fi
return 1
fi
}
# verify_compose_stack <compose-file>
# `docker compose up -d` exits 0 as soon as containers are created, so a stack can
# come back broken with a green pipeline. Require every long-running service to
# actually be running.
verify_compose_stack() {
local cf="$1"
local expected running missing=()
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
running="$(compose "$cf" ps --status running --services | sort)" || return 1
[ -n "$expected" ] || return 0
while IFS= read -r svc; do
[ -n "$svc" ] || continue
# restart:"no" services are allowed to have exited.
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
missing+=("$svc")
fi
done <<<"$expected"
if [ "${#missing[@]}" -gt 0 ]; then
echo " NOT RUNNING: ${missing[*]}"
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
return 1
fi
echo " all ${#expected} service(s) running"
return 0
}
# The public hostname of every active service, one per line.
#
# Comments are stripped first, and deliberately so: a route that someone
# disabled by commenting it out is not a service to probe, and naio and xui are
# both still in the tree that way. A `#` only starts a comment when it is at the
# start of a line or after whitespace, so `s/#.*//` alone would also cut a
# legitimate value in half.
#
# Only the public names. The *.internal names are the same Traefik and the same
# Services, reached by a different label, so probing both would double the run
# to learn the same thing. The public name is also the one a user types.
smoke_hosts() {
local m k
# The backticks below are literal. They are Traefik's Host() delimiter, and the
# single quotes are precisely what keeps the shell from reading them as a
# command substitution, so the warning is the opposite of a real problem.
# shellcheck disable=SC2016
{
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
[ -f "$m" ] && cat "$m"
done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
kubectl kustomize "$k" 2>/dev/null || true
done
} | sed -E 's/(^|[[:space:]])#.*$//' \
| grep -oE 'Host\(`[^`]+`\)' \
| sed -E 's/^Host\(`//; s/`\)$//' \
| grep -E '(^|\.)forust\.xyz$' \
| grep -v '\${' \
| sort -u
}
# Traefik's own list of the routes it actually built. The Kubernetes CRs are the
# wrong source for this: when a middleware fails to load, Traefik drops the
# router that referenced it and leaves the CR behind looking perfectly healthy.
#
# api.insecure is already on for the internal `traefik` entrypoint, but the pod
# IP is not routable from the node, so read it through kubectl exec rather than
# standing up a port-forward. HTTP only: the TCP routers match on HostSNI(`*`)
# and the UDP ones carry no rule at all, both selected by entrypoint and port,
# so neither can answer whether a given host has a route.
traefik_http_routes() {
kubectl -n traefik exec deploy/traefik -- \
wget -qO- --timeout=10 http://127.0.0.1:8080/api/http/routers 2>/dev/null \
| jq -c '[.[] | {status, rule: (.rule // "")}]'
}
# The hosts Traefik currently routes to, one per line. Every backticked token of
# an enabled rule counts, which is a superset of the hosts -- PathPrefix values
# land here too, harmlessly -- but it keeps the host syntax in one place instead
# of a matcher per host. The scan keeps the delimiters, so strip them: what
# belongs in a comparison against a hostname is the bare name.
traefik_routed_hosts() {
jq -r '[.[] | select(.status == "enabled") | (.rule // "")
| scan("`[^`]+`") | ltrimstr("`") | rtrimstr("`")]
| unique | .[]' <<<"$1"
}
# stage_verify_k8s watches the rollout, which reports that the pods converged.
# It cannot tell a converged pod from a serving one: a route pointing at the
# wrong port, a Service selector that matches nothing the app listens on, a 500
# from the app itself, an OOMKill loop that still counts as Available for long
# enough to pass. All of those are green at the rollout level.
#
# So ask the thing users ask. Any HTTP response proves Traefik matched the
# host, the Service resolved to a pod and the pod answered -- a 302 to a login
# or a 404 from a path the service does not serve still means the chain is
# intact. Only a transport failure (no DNS, refused, timeout) or a 5xx means
# the service is not serving, and only those fail the run.
#
# Except that a 404 is not evidence on its own. A router Traefik refused to
# build answers with the same 404 and nothing behind it, so a middleware that
# fails to load takes down every route that referenced it while
# this stage reports `ok` for all of them. No status code separates those two
# cases, so ask Traefik which routes it built and fail on the difference.
stage_smoke() {
cd "$REPO"
if [ -n "${DEPLOY_PLAN:-}" ] && jq -e '.full_smoke' "$DEPLOY_PLAN" >/dev/null; then
DEPLOY_SMOKE_ALL=true
fi
select_manifests >/dev/null
local -a hosts=()
local h code rc bad=0
while IFS= read -r h; do
[ -n "$h" ] && hosts+=("$h")
done < <(smoke_hosts)
if [ "${#hosts[@]}" -eq 0 ]; then
# Nothing to probe means the extraction broke, not that the cluster is empty.
echo "No public routes in the selected components"
return 0
fi
log "Probing ${#hosts[@]} public route(s)"
for h in "${hosts[@]}"; do
code="$(curl -sS -o /dev/null --max-time 20 -w '%{http_code}' "https://$h/" 2>/dev/null)" && rc=0 || rc=$?
if [ "$rc" -ne 0 ]; then
echo " UNREACHABLE $h (curl exit $rc)"
bad=1
continue
fi
# A glob, not a string compare. `case` on the leading digit is the only one
# of these that survives a three-digit code, and the obvious expansion to
# try first -- ${code%%[0-9]*} -- is empty for every input, so it silently
# reports a 500 as healthy.
case "$code" in
5*)
echo " SERVER ERROR $h $code"
bad=1
;;
000)
# curl exited 0 and still no status, so nothing on the far end replied.
# Not a pass, whatever the transport thought.
echo " NO RESPONSE $h"
bad=1
;;
*)
echo " ok $h $code"
;;
esac
done
# Second gate. The probe above only means something if a router matched the
# host in the first place, so compare the hosts we expect against the hosts
# Traefik reports and fail on the difference.
local routes routed
if ! routes="$(traefik_http_routes)"; then
echo "ERROR: could not read Traefik's router list, refusing to report success"
return 1
fi
routed="$(traefik_routed_hosts "$routes")"
local -a unrouted=()
local tries=3
while :; do
unrouted=()
for h in "${hosts[@]}"; do
grep -qxF "$h" <<<"$routed" || unrouted+=("$h")
done
if [ "${#unrouted[@]}" -eq 0 ]; then
break
fi
# A router mid-rollout is legitimately absent for a moment. A middleware
# that failed to load stays absent, so waiting cannot paper over it.
if [ "$tries" -le 1 ]; then
break
fi
tries=$((tries - 1))
warn "${#unrouted[@]} host(s) have no enabled route yet, re-checking in 10s"
sleep 10
if ! routes="$(traefik_http_routes)"; then
break
fi
routed="$(traefik_routed_hosts "$routes")"
done
if [ "${#unrouted[@]}" -ne 0 ]; then
for h in "${unrouted[@]}"; do
echo " NO ROUTE $h (Traefik has no enabled router for this host)"
done
bad=1
fi
if [ "$bad" -ne 0 ]; then
echo "ERROR: at least one active service is not serving over its public route"
return 1
fi
echo "all ${#hosts[@]} route(s) answered and have a router"
}
stage_apply_compose() {
cd "$REPO"
select_manifests >/dev/null
local cf
for cf in "${COMPOSE_STACKS[@]}"; do
log "Applying Compose ${cf#"$REPO"/}"
compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans
verify_compose_stack "$cf"
done
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
}
run_stage() {
case "${1:?stage required}" in
doctor) stage_doctor ;;
validate) stage_validate ;;
apply-k8s) stage_apply_k8s ;;
verify-k8s) stage_verify_k8s ;;
smoke) stage_smoke ;;
apply-compose) stage_apply_compose ;;
*)
echo "ERROR: unknown stage: $1"
exit 1
;;
esac
}
-124
View File
@@ -1,124 +0,0 @@
#!/usr/bin/env python3
"""Calculate selected components against the last fully successful deploy."""
import hashlib
import json
import re
import subprocess
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def tracked(repo):
return output('git', '-C', str(repo), 'ls-files').splitlines()
def helm_releases(repo):
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
def inventory(repo):
files = tracked(repo)
k8s = sorted(
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
)
compose = sorted(
{
str(Path(f).parent)
for f in files
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
}
)
return {'k8s': k8s, 'compose': compose}
def file_hash(path):
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
def make_plan(repo, config_repo, release, previous, mode, live_helm):
active = inventory(repo)
all_services = set(active['k8s'] + active['compose'])
helm_inputs = {}
helm_selected = []
for name, chart, namespace, version, values, marker in helm_releases(repo):
if not (repo / marker).is_file():
continue
value_path = repo / values if (repo / values).is_file() else config_repo / values
if not value_path.is_file():
raise ValueError(f'Missing Helm values: {values}')
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
helm_inputs[name] = stamp
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
if (
mode == 'full'
or previous is None
or previous.get('helm_inputs', {}).get(name) != stamp
or live is None
or live.get('status') != 'deployed'
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
):
helm_selected.append(name)
local_inputs = {}
for service in all_services:
candidates = [config_repo / service / '.env']
if service in active['compose']:
candidates.append(config_repo / '.env')
cfg = config_repo / service / 'config'
if cfg.is_dir():
candidates.extend(
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
)
local_inputs[service] = hashlib.sha256(
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
).hexdigest()
if previous is None:
if mode == 'changed':
raise ValueError('No successful baseline; run deploy in full mode first')
changed = set(all_services)
removed = []
else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {path.split('/')[0] for path in paths}
if any(path.startswith('.gitea/') for path in paths):
changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo):
service = file.split('/')[0]
if service not in all_services or not file.endswith(('.yaml', '.yml')):
continue
text = (repo / file).read_text()
if any(
image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items()
):
changed.add(service)
removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services
)
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
if mode == 'full':
changed = set(all_services)
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
while True:
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
if expanded == changed:
break
changed = expanded
return {
'version': 1,
'sha': release['sha'],
'images': release['images'],
'active': active,
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
'helm': helm_selected,
'helm_inputs': helm_inputs,
'local_inputs': local_inputs,
'removed': sorted(set(removed)),
'full_smoke': mode == 'full' or 'traefik' in changed,
}
-10
View File
@@ -1,10 +0,0 @@
#!/usr/bin/env bash
set -euo pipefail
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
case "${1:?stage required}" in
workload-count)
select_manifests >/dev/null
selected_workload_refs | sort -u | wc -l
;;
*) run_stage "$1" ;;
esac
+137 -89
View File
@@ -1,104 +1,152 @@
name: deploy name: deploy
on: on:
workflow_run:
workflows: [ci]
branches: [main]
types: [completed]
workflow_dispatch: workflow_dispatch:
inputs:
deploy_ref:
description: "Commit already checked by successful main CI (main or SHA)"
default: main
required: true
deploy_mode:
description: "First deploy requires full; plan changes no production resources"
type: choice
options: [changed, full, plan]
default: changed
refresh_images:
description: "Explicitly refresh mutable third-party Compose tags"
type: boolean
default: false
permissions:
contents: read
actions: read
concurrency: concurrency:
group: deploy-main group: deploy-main
cancel-in-progress: false cancel-in-progress: false
env:
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
jobs: jobs:
gate: redeploy:
if: >- runs-on: [self-hosted, linux, arch, homelab, prod]
github.ref == 'refs/heads/main' &&
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
runs-on: homelab
timeout-minutes: 10
outputs:
sha: ${{ steps.release.outputs.sha }}
steps: steps:
- name: Checkout repository - name: Redeploy workstation
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 shell: bash
with:
fetch-depth: 0
- name: Check successful CI and download the exact commit release
id: release
env: env:
GITEA_TOKEN: ${{ github.token }} DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }} DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
EVENT_SHA: ${{ github.event.workflow_run.head_sha }} DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA" DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
- name: Submit durable deploy to workstation DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
run: bash .gitea/workflows/ssh-run.sh start # Set APPLY_PRUNE=true to enable kubectl apply --prune. Requires every
# manifest to carry label app.kubernetes.io/managed-by=homelab-deploy,
# otherwise previously applied resources get deleted on the next run.
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
run: |
set -euo pipefail
apply: : "${DEPLOY_HOST:?missing DEPLOY_HOST}"
needs: [gate] : "${DEPLOY_USER:?missing DEPLOY_USER}"
runs-on: homelab : "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
timeout-minutes: 100
steps:
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply
verify: deploy_port="${DEPLOY_PORT:-22}"
needs: [gate, apply] deploy_path="${DEPLOY_PATH:-/srv/homelab}"
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 130
steps:
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify
smoke: ssh_key="$RUNNER_TEMP/deploy_key"
needs: [gate, verify] mkdir -p "$RUNNER_TEMP"
if: always() && needs.gate.result == 'success' printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
runs-on: homelab chmod 600 "$ssh_key"
timeout-minutes: 15
steps: ssh_opts=(
- name: Checkout checked commit -i "$ssh_key"
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 -p "$deploy_port"
with: -o BatchMode=yes
ref: ${{ needs.gate.outputs.sha }} -o StrictHostKeyChecking=accept-new
- name: Follow public route checks )
run: bash .gitea/workflows/ssh-run.sh smoke
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
"DEPLOY_PATH=$(printf '%q' \"$deploy_path\") APPLY_PRUNE=$(printf '%q' \"${APPLY_PRUNE:-false}\") bash -se" <<'EOF'
set -euo pipefail
repo="${DEPLOY_PATH:-/srv/homelab}"
if [ ! -d "$repo/.git" ]; then
echo "Repository not found at $repo"
exit 1
fi
git -C "$repo" fetch origin main
git -C "$repo" reset --hard origin/main
# Runtime selection: a service is k8s-managed when $SERVICE/k8s/active
# exists. Otherwise it is compose-managed, and only k8s/routing/*
# manifests (external Services / EndpointSlices / ServersTransport /
# Ingresses that route to docker backends) are applied.
# migrate: touch SERVICE/k8s/active (+ move routing files up)
# rollback: rm SERVICE/k8s/active
collect_k8s() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
! -path '*/routing/*' ! -path '*/overlays/*' \
! -name 'kustomization.y*ml' ! -name '*.example.y*ml' \
! -name '*values.y*ml' ! -name 'patch-*.y*ml' \
| sort
}
collect_k8s_inactive() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
\( -name 'namespace.y*ml' -o -path '*/routing/*' \) \
! -path '*/overlays/*' ! -name '*.example.y*ml' \
| sort
}
mapfile -t compose_stacks < <(
find "$repo" -type f \( -name 'compose.yaml' -o -name 'compose.yml' \) | sort
)
mapfile -t k8s_manifests < <(
for kd in $(find "$repo" -type d -name k8s ! -path '*/.git/*' | sort); do
if [ -f "$kd/active" ]; then
collect_k8s "$kd"
else
collect_k8s_inactive "$kd"
fi
done
)
echo "== Validate compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " config: $cf"
docker compose -f "$cf" config --quiet
done
echo "== Validate k8s manifests (kubectl dry-run) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=client $m"
kubectl apply --dry-run=client -f "$m" >/dev/null
done
echo "== Applying Kubernetes manifests =="
ns_files=()
other_files=()
for m in "${k8s_manifests[@]}"; do
case "$m" in
*/namespace.y?ml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
prune_opts=()
if [ "${APPLY_PRUNE:-false}" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
if [ "${#ns_files[@]}" -gt 0 ]; then
echo " namespaces first: ${ns_files[*]}"
kubectl apply -f "${ns_files[@]}"
fi
if [ "${#other_files[@]}" -gt 0 ]; then
echo " resources: ${other_files[*]}"
kubectl apply "${prune_opts[@]}" -f "${other_files[@]}"
fi
echo "== Redeploying docker compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " compose: $dir"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
fi
docker compose -f "$cf" up -d --pull always --remove-orphans
done
EOF
-285
View File
@@ -1,285 +0,0 @@
#!/usr/bin/env bash
# Installs the pinned CI tools into "$TOOLS_DIR/bin" and echoes that directory
# on stdout, so callers can do:
#
# export PATH="$(bash .gitea/workflows/install-ci-tools.sh kubeconform shellcheck):$PATH"
#
# Versions come from tool-versions.env next to this script and are kept fresh by
# Renovate. Re-running is cheap: an already-installed tool at the pinned version
# is left alone.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
# A runner may accept overlapping workflows even though each workflow is sequential.
exec 9>"$TOOLS_DIR/install.lock"
flock -w 300 9
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
# The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner).
export PATH="$BIN_DIR:$PATH"
arch="$(uname -m)"
# Upstream projects disagree on arch spelling: kubeconform and actionlint use
# Go names (amd64/arm64), shellcheck uses uname names (x86_64/aarch64), node
# uses neither (x64/arm64), and hadolint mixes the two in a single release
# (x86_64 but arm64).
case "$arch" in
x86_64 | amd64)
goarch=amd64
sharch=x86_64
nodearch=x64
hadolintarch=x86_64
;;
aarch64 | arm64)
goarch=arm64
sharch=aarch64
nodearch=arm64
hadolintarch=arm64
;;
*)
echo "install-ci-tools: unsupported architecture: $arch" >&2
exit 1
;;
esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
echo "install-ci-tools: neither curl nor wget is available" >&2
exit 1
fi
}
# resolve <command>
# Absolute path to use for invoking a tool: the copy in BIN_DIR when present,
# otherwise the name for PATH lookup. Every version check and every in-script
# invocation goes through this, so a tool missing from both places reads as
# "not installed" instead of dying with 127 under `set -e`.
resolve() {
if [ -x "$BIN_DIR/$1" ]; then
printf '%s' "$BIN_DIR/$1"
else
printf '%s' "$1"
fi
}
# installed_version <command>
# Prints the version of an already-installed tool, or nothing. Each tool spells
# its version flag differently, hence the case.
installed_version() {
local bin out
bin="$(resolve "$1")"
if ! command -v "$bin" >/dev/null 2>&1; then
return 0
fi
case "$1" in
kubeconform) out="$("$bin" -v 2>/dev/null | head -1 || true)" ;;
*) out="$("$bin" --version 2>/dev/null | head -1 || true)" ;;
esac
printf '%s' "$out"
}
# at_version <command> <expected>
at_version() {
local version expected="${2#v}"
version="$(installed_version "$1")"
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
[ "${BASH_REMATCH[2]}" = "$expected" ]
else
return 1
fi
}
install_kubeconform() {
if at_version kubeconform "v${KUBECONFORM_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/yannh/kubeconform/releases/download/v${KUBECONFORM_VERSION}/kubeconform-linux-${goarch}.tar.gz" \
"$tmp/kubeconform.tar.gz"
tar -xzf "$tmp/kubeconform.tar.gz" -C "$tmp" kubeconform
install -m 0755 "$tmp/kubeconform" "$BIN_DIR/kubeconform"
rm -rf "$tmp"
}
install_shellcheck() {
if at_version shellcheck "${SHELLCHECK_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/koalaman/shellcheck/releases/download/v${SHELLCHECK_VERSION}/shellcheck-v${SHELLCHECK_VERSION}.linux.${sharch}.tar.xz" \
"$tmp/shellcheck.tar.xz"
tar -xJf "$tmp/shellcheck.tar.xz" -C "$tmp" --strip-components=1 "shellcheck-v${SHELLCHECK_VERSION}/shellcheck"
install -m 0755 "$tmp/shellcheck" "$BIN_DIR/shellcheck"
rm -rf "$tmp"
}
install_jq() {
if at_version jq "${JQ_VERSION}"; then
return 0
fi
fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \
"$BIN_DIR/jq"
chmod 0755 "$BIN_DIR/jq"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
# uv release tags carry no leading v, unlike every other tool installed here.
fetch "https://github.com/astral-sh/uv/releases/download/${UV_VERSION}/uv-${sharch}-unknown-linux-gnu.tar.gz" \
"$tmp/uv.tar.gz"
tar -xzf "$tmp/uv.tar.gz" -C "$tmp" --strip-components=1 "uv-${sharch}-unknown-linux-gnu/uv"
install -m 0755 "$tmp/uv" "$BIN_DIR/uv"
rm -rf "$tmp"
}
install_hadolint() {
if at_version hadolint "${HADOLINT_VERSION}"; then
return 0
fi
# A bare binary, no archive: hadolint ships one file per platform.
fetch "https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-linux-${hadolintarch}" \
"$BIN_DIR/hadolint"
chmod 0755 "$BIN_DIR/hadolint"
}
# ruff and yamllint both come from PyPI as wheels, which uv unpacks for us.
install_uv_tool() {
# <package> <pinned version>
if at_version "$1" "$2"; then
return 0
fi
install_uv
UV_TOOL_BIN_DIR="$BIN_DIR" "$BIN_DIR/uv" tool install --force "$1==$2" >/dev/null
}
install_ruff() {
install_uv_tool ruff "${RUFF_VERSION}"
}
install_yamllint() {
install_uv_tool yamllint "${YAMLLINT_VERSION}"
}
install_pip_audit() {
install_uv_tool pip-audit "${PIP_AUDIT_VERSION}"
}
install_prettier() {
install_node
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
# Not a standalone binary: prettier's entry point requires ../package.json
# relative to its own real path, so the package directory has to survive
# next to it. Hence a versioned directory plus a relative symlink, rather
# than copying the one file out as the other installers do.
local dir="$BIN_DIR/prettier-${PRETTIER_VERSION}"
if [ ! -f "$dir/package/package.json" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://registry.npmjs.org/prettier/-/prettier-${PRETTIER_VERSION}.tgz" "$dir/prettier.tgz"
tar -xzf "$dir/prettier.tgz" -C "$dir"
rm -f "$dir/prettier.tgz"
# npm strips the exec bit from bin/ on the way into the tarball.
chmod 0755 "$dir/package/bin/prettier.cjs"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
ln -sfn "prettier-${PRETTIER_VERSION}/package/bin/prettier.cjs" "$BIN_DIR/prettier"
}
install_node() {
# npm gets checked by running it, not by looking it up: what matters is that
# it answers, so a stub, a half-removed Arch package or a name that resolves
# to something broken all have to read as "not installed". The runner's npm
# is a symlink into /usr/lib/node_modules/npm, which is exactly the kind of
# thing that disappears between runs.
if at_version node "v${NODE_VERSION}" && [ -n "$(installed_version npm)" ]; then
return 0
fi
# Same shape as prettier above: the tarball's bin/npm and bin/npx are links
# into lib/node_modules, so the whole tree has to survive next to them.
local dir="$BIN_DIR/node-${NODE_VERSION}"
if [ ! -x "$dir/bin/node" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://nodejs.org/dist/v${NODE_VERSION}/node-v${NODE_VERSION}-linux-${nodearch}.tar.xz" \
"$dir/node.tar.xz"
tar -xJf "$dir/node.tar.xz" -C "$dir" --strip-components=1 "node-v${NODE_VERSION}-linux-${nodearch}"
rm -f "$dir/node.tar.xz"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
for bin in node npm npx; do
ln -sfn "node-${NODE_VERSION}/bin/${bin}" "$BIN_DIR/${bin}"
done
}
install_actionlint() {
if at_version actionlint "${ACTIONLINT_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_${goarch}.tar.gz" \
"$tmp/actionlint.tar.gz"
tar -xzf "$tmp/actionlint.tar.gz" -C "$tmp" actionlint
install -m 0755 "$tmp/actionlint" "$BIN_DIR/actionlint"
rm -rf "$tmp"
}
main() {
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
jq) install_jq ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
[ -d "$old" ] || continue
case "$(basename "$old")" in
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
*) rm -rf "$old" ;;
esac
done
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
printf '%s\n' "$BIN_DIR"
}
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
-350
View File
@@ -1,350 +0,0 @@
#!/usr/bin/env python3
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
import argparse
import hashlib
import io
import itertools
import json
import os
import re
import shutil
import subprocess
import sys
import tempfile
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from pathlib import Path
SHA = re.compile(r'[0-9a-f]{40}')
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
IMAGES = {
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
}
# These images are released by the EDU application repository.
EXTERNAL_IMAGES = {'gcr.forust.xyz/forust/session-keeper', 'gcr.forust.xyz/forust/webinar-checker'}
def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell."""
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def validate_release(data, sha=None):
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
raise ValueError('Invalid release version or SHA')
if sha is not None and data['sha'] != sha:
raise ValueError('Release SHA does not match the checked CI commit')
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
if set(data.get('images', {})) != expected:
raise ValueError('Release must contain all owned images')
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
raise ValueError('Release has an invalid image digest')
if set(data.get('inputs', {})) != expected or not all(
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
):
raise ValueError('Release has invalid build input fingerprints')
return data
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
return None
class Gitea:
def __init__(self):
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
if urllib.parse.urlsplit(self.origin).scheme != 'https':
raise ValueError('Gitea API must use HTTPS')
self.repository = os.environ['GITHUB_REPOSITORY']
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
raise ValueError('Invalid Gitea repository')
self.token = os.environ['GITEA_TOKEN']
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
def request(self, url, *, archive=False):
if not url.startswith(self.base + '/'):
raise ValueError('Refusing to send the Actions token to another origin')
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
opener = urllib.request.build_opener(NoRedirect())
try:
response = opener.open(req, timeout=30) # noqa: S310
except urllib.error.HTTPError as error:
if not archive or error.code not in (301, 302, 303, 307, 308):
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
target = urllib.parse.urljoin(url, error.headers['Location'])
if urllib.parse.urlsplit(target).scheme != 'https':
raise ValueError('Artifact redirect must use HTTPS') from None
# Signed storage redirects must never receive the Gitea token.
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
with response:
payload = response.read(8 * 1024 * 1024 + 1)
if len(payload) > 8 * 1024 * 1024:
raise ValueError('Gitea response exceeds 8 MiB')
return payload if archive else json.loads(payload)
def pages(self, path, key, **params):
for page in range(1, 101):
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
data = self.request(f'{self.base}/{path}?{query}')
entries = data[key]
yield from entries
if len(entries) < 50:
return
raise RuntimeError('Gitea pagination limit exceeded')
def successful_runs(self, sha=None):
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
if sha:
params['head_sha'] = sha
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
if (
run.get('status') == 'completed'
and run.get('conclusion') == 'success'
and run.get('head_branch') == 'main'
and run.get('event') in ('push', 'workflow_dispatch')
and (run.get('repository') or {}).get('full_name') == self.repository
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
and (sha is None or run.get('head_sha') == sha)
):
yield run
def release(self, run):
sha = run['head_sha']
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
# A green workflow with a skipped build must not authorize a deploy.
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
raise ValueError('CI build job did not succeed')
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
if len(matching) != 1:
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
files = [entry for entry in archive.infolist() if not entry.is_dir()]
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
raise ValueError('Unexpected release archive contents')
return validate_release(json.loads(archive.read(files[0])), sha)
def fingerprint(context, dockerfile):
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
return hashlib.sha256(tree.encode()).hexdigest()
def gate(output, requested_ref, event_sha):
command('git', 'fetch', '--quiet', 'origin', 'main')
if event_sha:
if not SHA.fullmatch(event_sha):
raise ValueError('Invalid workflow_run SHA')
sha = event_sha
else:
if requested_ref == 'main':
requested_ref = 'origin/main'
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
if not SHA.fullmatch(sha):
raise ValueError('Invalid deploy SHA')
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
api = Gitea()
runs = list(api.successful_runs(sha))
if not runs:
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
release = api.release(max(runs, key=lambda run: run['id']))
output.write_text(json.dumps(release, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write(f'sha={sha}\n')
print(f'CI gate accepted {sha}')
def build(output):
sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA')
api = Gitea()
previous = None
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
continue
try:
previous = api.release(run)
break
except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
builder = 'homelab-ci'
versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
)
image = versions['BUILDKIT_IMAGE']
signature = builder_config / 'homelab-ci-image'
exists = (
subprocess.run( # noqa: S603
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
capture_output=True,
env=env,
).returncode
== 0
)
if exists and (not signature.exists() or signature.read_text().strip() != image):
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
exists = False
if not exists:
command(
'docker',
'buildx',
'create',
'--name',
builder,
'--driver',
'docker-container',
'--driver-opt',
f'image={image}',
'--buildkitd-config',
'.gitea/runner/buildkitd.toml',
env=env,
)
signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
exists = False
if old_digest and previous['inputs'].get(image) == inputs:
exists = (
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'imagetools',
'inspect',
f'{image}@{old_digest}',
],
capture_output=True,
env=env,
timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--push',
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--tag',
f'{image}:sha-{sha}',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
release['images'][image] = digest
release['inputs'][image] = inputs
validate_release(release, sha)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
# Cleanup errors must neither leak credentials nor mask the original build error.
try:
subprocess.run( # noqa: S603
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'prune',
'--builder',
'homelab-ci',
'--force',
'--max-used-space',
'1gb',
],
env=env,
timeout=60,
)
except (OSError, subprocess.TimeoutExpired):
print('CI builder cache cleanup deferred', flush=True)
finally:
shutil.rmtree(docker_config)
def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile(
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
)
rendered = []
for line in stream:
match = image_line.fullmatch(line.rstrip('\n'))
if match:
prefix, quote, image, tail = match.groups()
if image in EXTERNAL_IMAGES and f'{image}@sha256:' in line:
rendered.append(line)
continue
if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
rendered.append(line)
destination.writelines(rendered)
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('build', 'gate', 'render'))
parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='')
args = parser.parse_args()
if args.action == 'render':
render(sys.stdin, sys.stdout)
elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha)
else:
build(args.output)
if __name__ == '__main__':
main()
+20 -58
View File
@@ -2,77 +2,38 @@ name: renovate-ci
on: on:
pull_request: pull_request:
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
push: push:
branches: branches:
- main - main
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
workflow_dispatch: workflow_dispatch:
permissions:
contents: read
jobs: jobs:
validate-renovate: validate-renovate:
runs-on: homelab runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@v4
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag, - name: Validate Renovate Compose draft
# so the same version that runs in the cluster is the one validated here.
- name: Resolve the deployed Renovate image
id: image
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ trap 'rm -f renovate/.env' EXIT
renovate/k8s/cronjob.yaml | head -1)" printf '%s\n' \
if [ -z "$image" ]; then 'RENOVATE_ENDPOINT=https://gitea.example/api/v1' \
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml" 'RENOVATE_TOKEN=test-token' \
exit 1 'RENOVATE_REPOSITORIES=forust/homelab' \
fi > renovate/.env
echo "using $image" docker compose -f renovate/renovate-compose.yaml config --quiet
echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate repository config - name: Validate Kubernetes manifests
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
docker run --rm \ docker run --rm \
-v "$PWD/renovate:/opt/renovate:ro" \ -v "$PWD:/work" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -w /work \
"${{ steps.image.outputs.image }}" \ ghcr.io/yannh/kubeconform:latest \
renovate-config-validator /opt/renovate/renovate.json
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches.
- name: Check the generated Renovate ConfigMap
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/sync-renovate-configmap.sh --check
- name: Validate Renovate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \
-strict \ -strict \
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
@@ -80,11 +41,12 @@ jobs:
renovate/k8s/configmap.yaml \ renovate/k8s/configmap.yaml \
renovate/k8s/cronjob.yaml renovate/k8s/cronjob.yaml
- name: Validate Renovate Compose file - name: Validate Renovate repository config
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
source .gitea/workflows/compose-lint.sh docker run --rm \
mapfile -t safe_flags < <(compose_safe_flags) -v "$PWD:/work" \
validate_compose_file renovate/renovate-compose.yaml \ -w /work \
${safe_flags[@]+"${safe_flags[@]}"} renovate/renovate:44.83.2 \
renovate-config-validator renovate.json
-92
View File
@@ -1,92 +0,0 @@
name: renovate-run
on:
workflow_dispatch:
inputs:
repositories:
description: "Repositories to scan (comma-separated)"
required: false
default: "forust/homelab"
log_level:
description: "Renovate log level"
required: false
default: "info"
type: choice
options:
- info
- debug
dry_run:
description: "Plan only, do not open or update PRs"
required: false
default: false
type: boolean
# Renovate writes through its own bot PAT, passed in as RENOVATE_TOKEN, so the
# Actions token is only ever used to read the checkout.
permissions:
contents: read
concurrency:
group: renovate-run
cancel-in-progress: false
jobs:
run-renovate:
runs-on: homelab
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version
# that is deployed, instead of a copy that silently goes stale.
- name: Resolve the deployed Renovate image
id: image
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
renovate-config-validator
- name: Run Renovate
shell: bash
env:
RENOVATE_TOKEN: ${{ secrets.RENOVATE_TOKEN }}
RENOVATE_GITHUB_COM_TOKEN: ${{ secrets.RENOVATE_GITHUB_COM_TOKEN }}
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
LOG_LEVEL: ${{ inputs.log_level }}
run: |
set -euo pipefail
: "${RENOVATE_TOKEN:?missing RENOVATE_TOKEN secret — add a renovate-bot PAT in repo/org Actions secrets}"
docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_PLATFORM=gitea \
-e RENOVATE_ENDPOINT=https://git.forust.xyz/api/v1 \
-e RENOVATE_TOKEN="$RENOVATE_TOKEN" \
-e RENOVATE_GITHUB_COM_TOKEN="${RENOVATE_GITHUB_COM_TOKEN:-}" \
-e RENOVATE_REPOSITORIES="${RENOVATE_REPOSITORIES:-forust/homelab}" \
-e RENOVATE_DRY_RUN="${RENOVATE_DRY_RUN:-}" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
"${{ steps.image.outputs.image }}"
-13
View File
@@ -1,13 +0,0 @@
# kubectl emits a List for files containing multiple resources.
(if .kind == "List" then .items[] else . end)
| (.metadata.namespace // "default") as $ns
| [
(.. | objects
| (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty)
| select(.optional != true)
| .name // .secretName // empty),
(.. | objects | .imagePullSecrets[]?.name)
]
| unique[]
| select(. != null and . != "")
| "\($ns) \(.)"
-56
View File
@@ -1,56 +0,0 @@
#!/usr/bin/env bash
# The SSH client submits once and follows durable stages on workstation.
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT
chmod 700 "$key_dir"
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
chmod 600 "$key_dir/key" "$key_dir/known_hosts"
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
controller=.local/lib/homelab-deploy/controller.py
case "${1:?start, apply, verify or smoke required}" in
start)
python3 - <<'PY' >"$key_dir/request.json"
import json
import os
from pathlib import Path
release = json.loads(Path('release.json').read_text())
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
PY
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
sleep 5
done
exit "$rc"
;;
apply|verify|smoke)
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
sleep 5
done
exit "$rc"
;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac
@@ -1,55 +0,0 @@
#!/usr/bin/env bash
# Regenerates renovate/k8s/configmap.yaml from renovate/renovate.json.
#
# renovate/renovate.json is the single source of truth: the CronJob, the Compose
# file and the renovate-run workflow all mount that exact file. A ConfigMap cannot
# read a file from the repository, so the same bytes are inlined here as a literal
# block. This script keeps the copy honest:
#
# .gitea/workflows/sync-renovate-configmap.sh # rewrite in place
# .gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
#
# renovate-ci runs the --check form on every PR and push, so a config change that
# forgets to regenerate the ConfigMap cannot be merged.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="$(git -C "$here" rev-parse --show-toplevel)"
src="$repo/renovate/renovate.json"
dst="$repo/renovate/k8s/configmap.yaml"
[ -f "$src" ] || {
echo "missing $src" >&2
exit 1
}
render() {
cat <<'HEADER'
# GENERATED FILE - do not edit by hand.
# Source: renovate/renovate.json
# Regenerate: .gitea/workflows/sync-renovate-configmap.sh
# Verify: .gitea/workflows/sync-renovate-configmap.sh --check
apiVersion: v1
kind: ConfigMap
metadata:
name: renovate-config
namespace: renovate
data:
renovate.json: |
HEADER
sed 's/^/ /' "$src"
}
if [ "${1:-}" = "--check" ]; then
if ! diff -u "$dst" <(render) >/dev/null 2>&1; then
echo "ERROR: $dst is out of sync with renovate/renovate.json"
echo "Run: .gitea/workflows/sync-renovate-configmap.sh"
diff -u "$dst" <(render) || true
exit 1
fi
echo "renovate/k8s/configmap.yaml is in sync with renovate/renovate.json"
exit 0
fi
render >"$dst"
echo "wrote $dst"
-39
View File
@@ -1,39 +0,0 @@
# Pinned versions of the CI tools installed by install-ci-tools.sh.
# Renovate keeps these up to date (see customManagers in renovate/renovate.json).
#
# Every version here except NODE_VERSION matches what was already installed on
# the runner, so pinning them changes what CI does not at all. It changes what
# CI does when the runner is rebuilt with something else: today
# install-ci-tools.sh finds the pinned version already on PATH and installs
# nothing, and a runner that drifts gets the pinned one installed over it.
#
# The renovate image version is NOT pinned here: renovate/k8s/cronjob.yaml is the
# single source of truth and the workflows read the tag from it, so there is
# nothing to drift.
ACTIONLINT_VERSION="1.7.7"
SHELLCHECK_VERSION="0.11.0"
KUBECONFORM_VERSION="0.8.0"
PRETTIER_VERSION="3.8.1"
RUFF_VERSION="0.16.8"
YAMLLINT_VERSION="1.38.0"
HADOLINT_VERSION="2.14.0"
# pip-audit reads the advisory database over the network, so a floating version
# would make the same commit report different things on different days. Pin it
# like the rest: the advisories themselves are the moving part, not the tool.
PIP_AUDIT_VERSION="2.10.1"
# uv builds the throwaway venv the pytest job runs in, and unpacks the PyPI
# wheels for ruff, yamllint and pip-audit.
UV_VERSION="0.12.17"
# node runs `npm ci` for the frontend tests and the npm audit, and it is the one
# pin here that does NOT come from the runner: the runner's system node is a
# rolling Arch package (it was node 26 with no npm at all when this was pinned),
# and the panel image is node:22-alpine. Pinned to the image's major on purpose,
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
# Secret-reference regression tests parse rendered Kubernetes objects.
JQ_VERSION="1.8.1"
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
+370
View File
@@ -0,0 +1,370 @@
name: ci
on:
push:
branches:
- "**"
pull_request:
workflow_dispatch:
env:
REGISTRY: gcr.forust.xyz
jobs:
lint-prettier:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Check formatting with Prettier
shell: bash
run: |
mapfile -t prettier_files < <(
git ls-files \
| grep -E '\.(md|json|ya?ml|html|css)$' \
| grep -Ev '^(\.docs/|\.zed/|errorpages/html/|homepages/(forust_files|xdfnx_files)/)'
)
if [ "${#prettier_files[@]}" -eq 0 ]; then
echo "No Prettier-managed files found."
exit 0
fi
docker run --rm \
-v "$PWD:/work" \
-w /work \
node:22-alpine \
sh -lc 'npx --yes prettier@3 --check --ignore-unknown "$@"' sh "${prettier_files[@]}"
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint Python with Ruff
shell: bash
run: |
docker run --rm \
-v "$PWD:/work" \
-w /work \
ghcr.io/astral-sh/ruff:latest \
check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint YAML syntax
shell: bash
run: |
mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \
':!node_modules/**' \
':!**/.venv/**'
)
if [ "${#yaml_files[@]}" -eq 0 ]; then
echo "No YAML files found."
exit 0
fi
docker run --rm \
-v "$PWD:/work" \
-w /work \
cytopia/yamllint:latest \
-c .yamllint "${yaml_files[@]}"
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Lint Dockerfiles
shell: bash
run: |
mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
)
if [ "${#dockerfiles[@]}" -eq 0 ]; then
echo "No Dockerfiles found."
exit 0
fi
docker run --rm \
-v "$PWD:/work" \
-w /work \
--entrypoint hadolint \
hadolint/hadolint:latest-debian \
-c .hadolint.yaml "${dockerfiles[@]}"
validate:
runs-on: [self-hosted, linux, arch, homelab]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Validate Kubernetes manifests
shell: bash
run: |
mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
if [ "${#manifests[@]}" -eq 0 ]; then
echo "No Kubernetes manifests found."
exit 0
fi
docker run --rm \
-v "$PWD:/work" \
-w /work \
ghcr.io/yannh/kubeconform:latest \
-strict \
-ignore-missing-schemas \
-summary \
"${manifests[@]}"
build:
needs: [lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev')
runs-on: [self-hosted, linux, arch, homelab]
outputs:
services: ${{ steps.services.outputs.services }}
steps:
- name: Checkout repository
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Detect changed docker-built services
id: services
shell: bash
run: |
base="${{ github.event.before }}"
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
base="$(git rev-list --max-parents=0 HEAD)"
fi
mapfile -t changed_files < <(git diff --name-only "$base" "${GITHUB_SHA}")
services=()
add_service() {
local name="$1"
local seen=0
for existing in "${services[@]}"; do
if [ "$existing" = "$name" ]; then
seen=1
break
fi
done
if [ "$seen" -eq 0 ]; then
services+=("$name")
fi
}
for file in "${changed_files[@]}"; do
case "$file" in
dtek_notif/*)
add_service dtek_notif
;;
errorpages/*)
add_service errorpages
;;
userbot/*)
add_service userbot
;;
homepages/*)
add_service homepages
;;
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
add_service edu_master
;;
esac
done
if [ "${#services[@]}" -eq 0 ]; then
echo "No docker-built services changed."
echo "services=" >> "$GITHUB_OUTPUT"
exit 0
fi
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
if: steps.services.outputs.services != ''
shell: bash
run: |
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "${REGISTRY}" \
-u "${{ secrets.REGISTRY_USERNAME }}" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
for service in "${services[@]}"; do
case "$service" in
dtek_notif)
image="${REGISTRY}/forust/dtek-notif"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" dtek_notif
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
errorpages)
image="${REGISTRY}/forust/error-pages"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
userbot)
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
for target in runtime panel; do
case "$target" in
runtime)
context="userbot"
image="${REGISTRY}/forust/userbot"
;;
panel)
context="userbot/panel"
image="${REGISTRY}/forust/userbot-panel"
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
homepages)
for service in forust xdfnx; do
case "$service" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" -f "homepages/Dockerfile.${service}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for service in session-keeper webinar-checker; do
case "$service" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build "${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
deploy-userbot-panel:
needs: build
if: github.ref_name == 'main' && contains(needs.build.outputs.services, 'userbot')
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Apply and roll out userbot panel
shell: bash
run: |
kubectl apply -f userbot/k8s/base/panel.yaml
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
# Keep legacy deployments (forust/anna) in sync with manifests; they have no replicas field, so apply leaves scaling to the user manager only.
kubectl apply -f userbot/k8s/base/userbots.yaml
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
+161
View File
@@ -0,0 +1,161 @@
name: deploy
on:
workflow_dispatch:
concurrency:
group: deploy-main
cancel-in-progress: false
jobs:
redeploy:
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Redeploy workstation
shell: bash
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
# Set APPLY_PRUNE=true to enable kubectl apply --prune. Requires every
# manifest to carry label app.kubernetes.io/managed-by=homelab-deploy,
# otherwise previously applied resources get deleted on the next run.
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
run: |
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
ssh_key="$RUNNER_TEMP/deploy_key"
mkdir -p "$RUNNER_TEMP"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
ssh_opts=(
-i "$ssh_key"
-p "$deploy_port"
-o BatchMode=yes
-o StrictHostKeyChecking=accept-new
)
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
"DEPLOY_PATH=$(printf '%q' \"$deploy_path\") APPLY_PRUNE=$(printf '%q' \"${APPLY_PRUNE:-false}\") bash -se" <<'EOF'
set -euo pipefail
repo="${DEPLOY_PATH:-/srv/homelab}"
if [ ! -d "$repo/.git" ]; then
echo "Repository not found at $repo"
exit 1
fi
git -C "$repo" fetch origin main
git -C "$repo" reset --hard origin/main
# Runtime selection: a service is k8s-managed when $SERVICE/k8s/active
# exists. Otherwise it is compose-managed, and only k8s/routing/*
# manifests (external Services / EndpointSlices / ServersTransport /
# Ingresses that route to docker backends) are applied.
# migrate: touch SERVICE/k8s/active (+ move routing files up)
# rollback: rm SERVICE/k8s/active
collect_k8s() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
! -path '*/routing/*' ! -path '*/overlays/*' \
! -name 'kustomization.y*ml' ! -name '*.example.y*ml' \
! -name '*values.y*ml' ! -name 'patch-*.y*ml' \
| sort
}
collect_k8s_inactive() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
\( -name 'namespace.y*ml' -o -path '*/routing/*' \) \
! -path '*/overlays/*' ! -name '*.example.y*ml' \
| sort
}
mapfile -t compose_stacks < <(
find "$repo" -type f \( -name 'compose.yaml' -o -name 'compose.yml' \) | sort
)
mapfile -t k8s_manifests < <(
for kd in $(find "$repo" -type d -name k8s ! -path '*/.git/*' | sort); do
if [ -f "$kd/active" ]; then
collect_k8s "$kd"
else
collect_k8s_inactive "$kd"
fi
done
)
echo "== Validate compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " config: $cf"
docker compose -f "$cf" config --quiet
done
echo "== Validate k8s manifests (kubectl dry-run) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=client $m"
kubectl apply --dry-run=client -f "$m" >/dev/null
done
echo "== Applying Kubernetes manifests =="
ns_files=()
other_files=()
for m in "${k8s_manifests[@]}"; do
case "$m" in
*/namespace.y?ml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
prune_opts=()
if [ "${APPLY_PRUNE:-false}" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
if [ "${#ns_files[@]}" -gt 0 ]; then
echo " namespaces first: ${ns_files[*]}"
kubectl apply -f "${ns_files[@]}"
fi
if [ -f "$repo/prometheus-stack/k8s/active" ]; then
echo "== Upgrading kube-prometheus-stack =="
helm upgrade --install prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace prometheus \
--version 86.2.3 \
--values "$repo/prometheus-stack/k8s/grafana-values.yaml" \
--wait
fi
if [ "${#other_files[@]}" -gt 0 ]; then
echo " resources: ${other_files[*]}"
kubectl apply "${prune_opts[@]}" -f "${other_files[@]}"
fi
echo "== Redeploying docker compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " compose: $dir"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
fi
docker compose -f "$cf" up -d --pull always --remove-orphans
done
EOF
-9
View File
@@ -21,9 +21,6 @@ checkmk/checkmk/*
downtify/Downtify_downloads downtify/Downtify_downloads
headscale/config/* headscale/config/*
headscale/data/* headscale/data/*
# NetBird local hostnames and generated secrets
netbird/.env
netbird/secrets/
searxng/core-config/* searxng/core-config/*
# Steaming services files # Steaming services files
@@ -96,8 +93,6 @@ replacements.txt
# Temp files # Temp files
edu_master/temp/ edu_master/temp/
temp/* temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/
# Environment # Environment
.env .env
@@ -109,10 +104,6 @@ tmp/
# kubernetes # kubernetes
*/k8s/*secret* */k8s/*secret*
!*/k8s/*secret*.example !*/k8s/*secret*.example
**/k8s/*secret*
!**/k8s/*secret*.example
# Local-only tweaks, not for upstream
prometheus-stack/k8s/grafana-values.yaml
traefik/k8s/local-tls.yaml traefik/k8s/local-tls.yaml
converters/k8s/config.yaml converters/k8s/config.yaml
convertx/k8s/config.yaml convertx/k8s/config.yaml
+1 -1
View File
@@ -31,7 +31,7 @@ services:
- "traefik.http.routers.adguard-dev.entrypoints=websecure" - "traefik.http.routers.adguard-dev.entrypoints=websecure"
- "traefik.http.routers.adguard-dev.tls=true" - "traefik.http.routers.adguard-dev.tls=true"
# DoH Router # DoH Router
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)" - "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
- "traefik.http.routers.dns-over-https.entrypoints=websecure" - "traefik.http.routers.dns-over-https.entrypoints=websecure"
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt" - "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
+1 -12
View File
@@ -51,8 +51,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: adguard-deployment name: adguard-deployment
namespace: adguard namespace: adguard
spec: spec:
@@ -60,8 +58,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: adguard app: adguard
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -75,7 +71,7 @@ spec:
memory: "1.5Gi" memory: "1.5Gi"
cpu: "300m" cpu: "300m"
requests: requests:
memory: "512Mi" memory: "500Mi"
cpu: "50m" cpu: "50m"
ports: ports:
- containerPort: 3000 - containerPort: 3000
@@ -84,13 +80,6 @@ spec:
name: dns name: dns
- containerPort: 853 - containerPort: 853
name: dot name: dot
readinessProbe:
tcpSocket:
port: dns
initialDelaySeconds: 5
periodSeconds: 5
successThreshold: 1
failureThreshold: 3
volumeMounts: volumeMounts:
- name: adguard-data - name: adguard-data
mountPath: /opt/adguardhome/work mountPath: /opt/adguardhome/work
-12
View File
@@ -1,12 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: adguard-certs
namespace: adguard
spec:
secretName: adguard-certs
dnsNames:
- dns.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
+6 -5
View File
@@ -7,18 +7,21 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
- match: Host(`dns.forust.xyz`) - match: Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: adguard-service - name: adguard-service
port: 3000 port: 3000
- match: (Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`) - match: (Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`)
kind: Rule kind: Rule
services: services:
- name: adguard-service - name: adguard-service
port: 3000 port: 3000
tls: tls:
secretName: adguard-certs certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -39,5 +42,3 @@ spec:
services: services:
- name: adguard-service - name: adguard-service
port: 3000 port: 3000
tls:
secretName: internal-wildcard-tls
-15
View File
@@ -1,15 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: adguard
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+7 -15
View File
@@ -27,8 +27,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-server-deployment name: authentik-server-deployment
namespace: authentik namespace: authentik
spec: spec:
@@ -36,8 +34,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: authentik-server app: authentik-server
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -45,7 +41,7 @@ spec:
spec: spec:
containers: containers:
- name: authentik-server - name: authentik-server
image: ghcr.io/goauthentik/server:2026.8.3 image: ghcr.io/goauthentik/server:2026.8.2
args: ["server"] args: ["server"]
envFrom: envFrom:
- configMapRef: - configMapRef:
@@ -56,8 +52,8 @@ spec:
- containerPort: 9000 - containerPort: 9000
resources: resources:
requests: requests:
memory: "768Mi" memory: "700Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "1.5Gi" memory: "1.5Gi"
cpu: "1000m" cpu: "1000m"
@@ -65,8 +61,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-worker-deployment name: authentik-worker-deployment
namespace: authentik namespace: authentik
spec: spec:
@@ -74,8 +68,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: authentik-worker app: authentik-worker
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -83,7 +75,7 @@ spec:
spec: spec:
containers: containers:
- name: authentik-worker - name: authentik-worker
image: ghcr.io/goauthentik/server:2026.8.3 image: ghcr.io/goauthentik/server:2026.8.2
args: ["worker"] args: ["worker"]
securityContext: securityContext:
runAsUser: 0 runAsUser: 0
@@ -94,8 +86,8 @@ spec:
name: authentik-secrets name: authentik-secrets
resources: resources:
requests: requests:
memory: "320Mi" memory: "512Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "768Mi" memory: "1Gi"
cpu: "700m" cpu: "700m"
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: authentik-prod-tls
namespace: authentik
spec:
secretName: authentik-prod-tls
dnsNames:
- auth.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: authentik
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+4 -3
View File
@@ -9,11 +9,14 @@ spec:
routes: routes:
- match: Host(`auth.forust.xyz`) - match: Host(`auth.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: authentik-server-service - name: authentik-server-service
port: 9000 port: 9000
tls: tls:
secretName: authentik-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -29,5 +32,3 @@ spec:
services: services:
- name: authentik-server-service - name: authentik-server-service
port: 9000 port: 9000
tls:
secretName: internal-wildcard-tls
@@ -1,9 +0,0 @@
crds:
enabled: true
prometheus:
servicemonitor:
enabled: true
interval: 60s
scrapeTimeout: 30s
labels:
release: prometheus-stack
-29
View File
@@ -1,29 +0,0 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-staging
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-staging-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-staging-account-key
solvers:
- http01:
ingress:
class: traefik
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
class: traefik
@@ -1,30 +0,0 @@
-----BEGIN CERTIFICATE-----
MIIFFjCCAv6gAwIBAgIUetKpTfEDOn2985FFMu6G26itT+wwDQYJKoZIhvcNAQEN
BQAwIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRlcm5hbCByb290IENBMB4XDTI2MDky
MzEyNDA0N1oXDTM2MDkyMDEyNDA0N1owIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRl
cm5hbCByb290IENBMIICIjANBgkqhkiG9w0BAQEFAAOCAg8AMIICCgKCAgEAvmNP
ZCOoD8NtNuYJKVXBlTPjX7D7sJCSK5neH7ZbYV5+lmUlEErY8Mik7j37V5k5NfpF
Ig85pOjP7RckTPz5V6ek3yaN40s4AL053sN5ZPauDVYjalaEHTgj5sEMqlLACQWI
yZmJOZspZykae8dIpQnqCoFpRT4FurJ78v4a0ylnFVLMQn/lyCHedwTjkEdtYWYr
ccJy8vQwqkzs/rWvEH1lDqZhennLOrmcCjfonG7D/pruMn4z+6E28p4+ejkRrI6x
luak3KnpT1XMeHtgU21hiRGaMDBchHMFgAhnY1qosymKenXvfTZItwgjZbwa1hJI
GAiDm+jQDKMjzRZ3rH6Xfc0auUcykNz73PpNu1NGm78nndXwCXcXn1LFKNQJ+r1U
sJiyAmUZmXVn4aM4OMf2F38k7wTYIKg7nRGaUkNeKDlNkjA4HvgWw+jwO1KmdHQ/
mOem1rosDWHRK01wg+Gga9mQCnhNhxglg3t/UeSic6uOaRsvaz4qkzHq8MbCujVz
DpKQjqdikYOAXZOs4KlBLWrS7NaK4NzfSD02pBUErh54ruJfY/bWz9KyXzBD/lQZ
VUTKyvUVB0bkVHEdf1jJmX3H4IZRQSF5JPqOBotW6bJI5fEGNBvj9Zxy4nm2WWGz
yyP3uWsQz8U/Wdx9nXZLHInTZBsvgLYtKUAWA30CAwEAAaNCMEAwDgYDVR0PAQH/
BAQDAgKkMA8GA1UdEwEB/wQFMAMBAf8wHQYDVR0OBBYEFEkKm2rxPaK6+O9WD80z
BLC6F9QsMA0GCSqGSIb3DQEBDQUAA4ICAQAnFyHz97Umf5VIu+dKTJid7C73VugJ
TIar/xJBs/4CxP+znBxhJjXygRoyIfzoVGWcB2ZSL//vL78Qlts79K/Imc9a4RFF
wMvCxsRXAEQ4TpeWi3ophPNcs4rhsP+gQKQFtnyKP9519bqpfxp0bTqwOV2o18fn
za7rlQViiEnNV58j7CVoM9+mJvVVfBEX1Km+GyJL9GadzbIQ7FxClVJZefCbft93
zHVk9gDOw8ys1XGSR2OUCyCLinXO6mqS16CmBb2MAKXq/YyH7E0N8iotAPGtfA8V
M/0ddy947rY0xCrtECfWwvGQpJS7NRv/Z9b2jCfXrI5LXmL2nfQRg0y9GE4Vjwr+
WxtGU5jOeFt0jQ+xRzcgG0Op+qK3x55l5LSo2hOcOVYbiHxcHEJFgwNi1ADeBFwb
q/HdysfURSOghqjIpMMAUabBp+DBUg2EUF7pIaUqbdqExFYcr9EYisEMiNsmKmN+
8ZbcOeerFKDQj+t/R0bFXa7UBn2UWsjI8zlR74aa2kLDXwtyz/XlO/FlYm66eBFo
2/eYUSeU+S4ej+wUAs/dvjF7f190/DUQGuwOTlLTahqWDztmhCk7qzbECu56CwKT
E5Ect2P72UleYwdblkVOVd352AmiwEzdOaziIRrPh8uenEknH6JBYPO3mjk7cCg+
GPGcNIBctjdXhg==
-----END CERTIFICATE-----
-33
View File
@@ -1,33 +0,0 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: selfsigned
spec:
selfSigned: {}
---
# Homelab internal root CA (10y). Install the .crt on clients (see below).
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-ca-root
namespace: cert-manager
spec:
isCA: true
commonName: homelab internal root CA
duration: 87600h
renewBefore: 7200h
secretName: internal-ca-root
privateKey:
algorithm: RSA
size: 4096
issuerRef:
name: selfsigned
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: internal-ca
spec:
ca:
secretName: internal-ca-root
-4
View File
@@ -1,4 +0,0 @@
apiVersion: v1
kind: Namespace
metadata:
name: cert-manager
+2 -6
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cfddns name: cfddns
labels: labels:
app: cfddns app: cfddns
@@ -11,8 +9,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: cfddns app: cfddns
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -26,10 +22,10 @@ spec:
imagePullPolicy: Always imagePullPolicy: Always
resources: resources:
requests: requests:
memory: "32Mi" memory: "20Mi"
cpu: "30m" cpu: "30m"
limits: limits:
memory: "128Mi" memory: "64Mi"
cpu: "50m" cpu: "50m"
envFrom: envFrom:
- secretRef: - secretRef:
File renamed without changes.
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: checkmk-prod-tls
namespace: checkmk
spec:
secretName: checkmk-prod-tls
dnsNames:
- cmk.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: checkmk
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
-4
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: checkmk-deployment name: checkmk-deployment
namespace: checkmk namespace: checkmk
spec: spec:
@@ -26,8 +24,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: checkmk app: checkmk
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
+4 -3
View File
@@ -9,11 +9,14 @@ spec:
routes: routes:
- match: Host(`cmk.forust.xyz`) - match: Host(`cmk.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: checkmk-service - name: checkmk-service
port: 5000 port: 5000
tls: tls:
secretName: checkmk-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP kind: IngressRouteTCP
@@ -45,5 +48,3 @@ spec:
services: services:
- name: checkmk-service - name: checkmk-service
port: 5000 port: 5000
tls:
secretName: internal-wildcard-tls
-1
View File
@@ -1 +0,0 @@
secret.yaml
-41
View File
@@ -1,41 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cloudflared
labels:
app: cloudflared
spec:
replicas: 1
selector:
matchLabels:
app: cloudflared
strategy:
type: Recreate
template:
metadata:
labels:
app: cloudflared
spec:
containers:
- name: cloudflared
image: cloudflare/cloudflared:2026.10.0
imagePullPolicy: IfNotPresent
args:
- tunnel
- --no-autoupdate
- run
env:
- name: TUNNEL_TOKEN
valueFrom:
secretKeyRef:
name: cloudflared-secrets
key: TUNNEL_TOKEN
resources:
requests:
memory: "128Mi"
cpu: "30m"
limits:
memory: "256Mi"
cpu: "200m"
-7
View File
@@ -1,7 +0,0 @@
apiVersion: v1
kind: Secret
metadata:
name: cloudflared-secrets
type: Opaque
stringData:
TUNNEL_TOKEN: your_tunnel_token_here
+2 -2
View File
@@ -1,7 +1,7 @@
services: services:
convertx: convertx:
container_name: convertx container_name: convertx
image: ghcr.io/c4illin/convertx:v0.19.0 image: ghcr.io/c4illin/convertx:v0.18.0
restart: unless-stopped restart: unless-stopped
ports: ports:
- "9992:3000" - "9992:3000"
@@ -54,7 +54,7 @@ services:
- "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt" - "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt"
- "traefik.http.routers.bentopdf.tls=true" - "traefik.http.routers.bentopdf.tls=true"
# Local router # Local router
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.workstation.internal`)" - "traefik.http.routers.bentopdf-local.rule=Host(`pdf.wokstation.internal`)"
- "traefik.http.routers.bentopdf-local.entrypoints=websecure" - "traefik.http.routers.bentopdf-local.entrypoints=websecure"
- "traefik.http.routers.bentopdf-local.tls=true" - "traefik.http.routers.bentopdf-local.tls=true"
# Dev router # Dev router
+2 -3
View File
@@ -31,13 +31,12 @@ spec:
name: bentopdf name: bentopdf
ports: ports:
- containerPort: 8080 - containerPort: 8080
# p95 4M, max 11M over 7 days. Was 50Mi/700Mi.
resources: resources:
requests: requests:
memory: "32Mi" memory: "50Mi"
cpu: "50m" cpu: "50m"
ephemeral-storage: "100Mi" ephemeral-storage: "100Mi"
limits: limits:
memory: "128Mi" memory: "700Mi"
cpu: "700m" cpu: "700m"
ephemeral-storage: "5Gi" ephemeral-storage: "5Gi"
-42
View File
@@ -1,42 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: convertx-prod-tls
namespace: converters
spec:
secretName: convertx-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: bentopdf-prod-tls
namespace: converters
spec:
secretName: bentopdf-prod-tls
dnsNames:
- pdf.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: converters
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -8
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: convertx-deployment name: convertx-deployment
namespace: converters namespace: converters
spec: spec:
@@ -22,15 +20,13 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: convertx app: convertx
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
app: convertx app: convertx
spec: spec:
containers: containers:
- image: ghcr.io/c4illin/convertx:v0.19.0 - image: ghcr.io/c4illin/convertx:v0.18.0
name: convertx name: convertx
envFrom: envFrom:
- configMapRef: - configMapRef:
@@ -42,14 +38,13 @@ spec:
volumeMounts: volumeMounts:
- mountPath: /data - mountPath: /data
name: data name: data
# p95 85M, max 136M over 7 days, spikes while converting. Was 250Mi/1.5Gi.
resources: resources:
requests: requests:
memory: "128Mi" memory: "250Mi"
cpu: "100m" cpu: "100m"
limits: limits:
cpu: "1500m" cpu: "1500m"
memory: "512Mi" memory: "1.5Gi"
volumes: volumes:
- name: data - name: data
persistentVolumeClaim: persistentVolumeClaim:
+2 -7
View File
@@ -14,7 +14,7 @@ spec:
- name: convertx-service - name: convertx-service
port: 3000 port: 3000
tls: tls:
secretName: convertx-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -31,9 +31,6 @@ spec:
services: services:
- name: convertx-service - name: convertx-service
port: 3000 port: 3000
tls:
secretName: internal-wildcard-tls
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -50,7 +47,7 @@ spec:
- name: bentopdf-service - name: bentopdf-service
port: 8080 port: 8080
tls: tls:
secretName: bentopdf-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -66,5 +63,3 @@ spec:
services: services:
- name: bentopdf-service - name: bentopdf-service
port: 8080 port: 8080
tls:
secretName: internal-wildcard-tls
+14
View File
@@ -0,0 +1,14 @@
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: crowdsec-bouncer
namespace: crowdsec
spec:
plugin:
crowdsec-bouncer:
enabled: true
LogLevel: INFO
CrowdsecMode: live
CrowdsecLapiScheme: http
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
+40 -186
View File
@@ -1,203 +1,57 @@
container_runtime: containerd container_runtime: containerd
agent: agent:
acquisition: []
additionalAcquisition:
- labels:
type: traefik
limit: 1000
query: |
{namespace="traefik"}
source: loki
url: http://loki.prometheus.svc.cluster.local:3100/
wait_for_ready: 30s
env: env:
- name: COLLECTIONS - name: COLLECTIONS
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios value: "crowdsecurity/traefik crowdsecurity/base-http-scenarios"
- name: DISABLE_COLLECTIONS - name: DISABLE_COLLECTIONS
value: crowdsecurity/sshd value: "crowdsecurity/linux crowdsecurity/sshd"
# Bans on 401/403 bursts hurt more than they protect: with L3 enforcement
# a false positive cuts the IP off everything (SSH included), and past acquisition:
# incidents show legit automation (deploy runner, mesh peers, registry - namespace: traefik
# pulls) tripping this probe. Probing/XSS/SQLi/CVE scenarios stay. podName: "*traefik*"
- name: DISABLE_SCENARIOS program: traefik
value: crowdsecurity/http-generic-bf poll_without_inotify: true
metrics:
enabled: true
serviceMonitor:
additionalLabels:
release: prometheus-stack
enabled: true
# Static machine identity: agent pods mount pre-created LAPI credentials
# (Secret crowdsec-agent-credentials, key local_api_credentials.yaml)
# at the exact path the agent entrypoint expects. Together with the
# patched register-init (enforced by janitor-cronjob.yaml) the agent
# never calls `cscli lapi register` in steady state, so pod names,
# restarts and reboots can no longer break it.
extraVolumes:
- name: static-creds
secret:
secretName: crowdsec-agent-credentials
items:
- key: local_api_credentials.yaml
path: local_api_credentials.yaml
extraVolumeMounts:
- name: static-creds
mountPath: /tmp_config/local_api_credentials.yaml
subPath: local_api_credentials.yaml
readOnly: true
resources: resources:
limits:
cpu: 200m
memory: 500Mi
requests: requests:
cpu: 50m cpu: 50m
memory: 100Mi memory: 100Mi
limits:
config: cpu: 200m
parsers: memory: 500Mi
s02-enrich:
mobile-whitelist.yaml: |
name: forust/mobile-whitelist
description: "Whitelist SWAN/4ka mobile network"
whitelist:
reason: "Mobile IP whitelist"
cidr:
- "84.245.64.0/18"
# CrowdSec's own guidance: CIDR allowlisting belongs at the parser stage.
# A parser whitelist discards the event before it reaches a bucket, so
# these addresses never produce an overflow and never become a decision.
# A postoverflow whitelist is checked only *after* the ban exists, and
# the bouncer answers 403 for as long as it does - which is a window we
# do not want the deploy sitting in.
local-network.yaml: |
name: forust/local-network
description: "Whitelist loopback, private and VPN networks"
whitelist:
reason: "Local network"
cidr:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
# CGNAT range (RFC 6598). The workstation and the k0s node live
# here on WireGuard, and 100.64.0.0/10 is not covered by the
# RFC 1918 blocks above.
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
vps-whitelist.yaml: |
name: forust/vps-whitelist
description: "Whitelist static VPS"
whitelist:
reason: "VPS"
ip:
- "193.181.211.79"
postoverflows:
s01-whitelist:
# The one whitelist that has to stay here: resolving a hostname is a
# network call, and the docs put expensive lookups in postoverflows on
# purpose - it runs only when a bucket actually overflows.
# ddns.forust.xyz is the public home address, not a private one, so
# forust/local-network does not cover it.
home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP"
whitelist:
reason: "Home dynamic IP"
expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# LAPI-only main config override, merged over config.yaml. NOTE: the
# chart's own default for this key is REPLACED, not merged, so its
# auto_registration block is repeated verbatim below - drop it and the
# agent can no longer register itself.
config.yaml.local: |
api:
server:
auto_registration: # Activate if not using TLS for authentication
enabled: true
token: "${REGISTRATION_TOKEN}" # /!\ Do not modify this variable (auto-generated and handled by the chart)
allowed_ranges: # /!\ Make sure to adapt to the pod IP ranges used by your cluster
- "127.0.0.1/32"
- "192.168.0.0/16"
- "10.0.0.0/8"
- "172.16.0.0/12"
# This homelab has no egress to console.crowdsec.cloud: DNS does
# not resolve. The LAPI kept trying anyway ("Signal push: N
# signals to push", "capi metrics: sending" every 10s) and each
# attempt sat on a resolver timeout WHILE HOLDING A WRITE
# TRANSACTION, which is what kept stalling per-request decision
# lookups even with WAL enabled. Nothing to share and nothing to
# pull - turn the Central API off instead of letting it block the
# only database writer we have.
online_client:
sharing: false
pull:
community: false
blocklists: false
disable_usage_metrics_export: true
db_config:
# SQLite without WAL serialises every reader behind the writer's
# rollback journal, and the LAPI writes constantly: the agent pushes
# Traefik alerts read from Loki, the metrics collector counts
# decisions, the bouncer touches "last pull" on every request.
# Symptom: decision lookups taking 10-30s (and a second connection
# that could not even open the database) while the LAPI sat at 28m
# CPU - the process was blocked in fsync, not computing. Every
# bouncer-protected request then blew through the plugin timeout and
# fail-closed with 403, on every site at once.
# The PVC is local-path-retain (hostPath), not a network share, so
# WAL is safe here; the crowdsec docs recommend it for exactly this
# ("allowing more concurrency in SQLite that will improve
# performances in most scenarios").
use_wal: true
# Keeps the alert table bounded. At the 5000/7d default the file
# reached 54MB in 15 days off the Traefik access log alone, and the
# metrics collector counts decisions on a timer; a smaller working
# set means fewer full scans. Crowdsec only prunes - SQLite never
# shrinks the file, so the size stays until a manual VACUUM.
flush:
max_items: 1000
max_age: 24h
lapi: lapi:
env: env:
- name: COLLECTIONS - name: COLLECTIONS
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios value: "crowdsecurity/traefik crowdsecurity/base-http-scenarios"
- name: DISABLE_COLLECTIONS - name: DISABLE_COLLECTIONS
value: crowdsecurity/linux crowdsecurity/sshd value: "crowdsecurity/linux crowdsecurity/sshd"
metrics:
enabled: true
serviceMonitor:
additionalLabels:
release: prometheus-stack
enabled: true
persistentVolume:
config:
enabled: true
size: 100Mi
storageClassName: local-path-retain
data:
enabled: true
size: 1Gi
storageClassName: local-path-retain
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
# request (whole Traefik front door), so it is the hot path of the proxy.
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
# which pushed requests into the bouncer's fail-closed 403.
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
# credentials on the `config` PVC) - two replicas sharing those RWO
# volumes would corrupt the decision store. Scale up CPU, not replicas.
resources:
limits:
cpu: 1500m
memory: 1Gi
requests:
cpu: 250m
memory: 500Mi
service: service:
type: ClusterIP type: ClusterIP
persistentVolume:
data:
enabled: true
storageClassName: local-path-retain
size: 1Gi
config:
enabled: true
storageClassName: local-path-retain
size: 100Mi
storeLAPICscliCredentialsInSecret: true storeLAPICscliCredentialsInSecret: true
resources:
requests:
cpu: 50m
memory: 150Mi
limits:
cpu: 400m
memory: 500Mi
metrics:
enabled: true
serviceMonitor:
additionalLabels:
release: prometheus-stack
enabled: true
interval: 30s
scrapeTimeout: 10s
namespace: prometheus
-32
View File
@@ -1,32 +0,0 @@
apiVersion: v1
data:
crowdsec-overview.json: "{\n \"__inputs\": [\n {\n \"name\": \"DS_PROMETHEUS\",\n \"label\": \"Prometheus\",\n \"description\": \"\",\n \"type\": \"datasource\",\n \"pluginId\": \"prometheus\",\n \"pluginName\": \"Prometheus\"\n }\n ],\n \"__requires\": [\n {\n \"type\": \"grafana\",\n \"id\": \"grafana\",\n \"name\": \"Grafana\",\n \"version\": \"8.1.2\"\n },\n {\n \"type\": \"panel\",\n \"id\": \"graph\",\n \"name\": \"Graph (old)\",\n \"version\": \"\"\n },\n {\n \"type\": \"datasource\",\n \"id\": \"prometheus\",\n \"name\": \"Prometheus\",\n \"version\": \"1.0.0\"\n },\n {\n \"type\": \"panel\",\n \"id\": \"stat\",\n \"name\": \"Stat\",\n \"version\": \"\"\n },\n {\n \"type\": \"panel\",\n \"id\": \"timeseries\",\n \"name\": \"Time series\",\n \"version\": \"\"\n }\n ],\n \"annotations\": {\n \"list\": [\n {\n \"builtIn\": 1,\n \"datasource\": \"-- Grafana --\",\n \"enable\": true,\n \"hide\": true,\n \"iconColor\": \"rgba(0, 211, 255, 1)\",\n \"name\": \"Annotations & Alerts\",\n \"target\": {\n \"limit\": 100,\n \"matchAny\": false,\n \"tags\": [],\n \"type\": \"dashboard\"\n },\n \"type\": \"dashboard\"\n }\n ]\n },\n \"editable\": true,\n \"gnetId\": null,\n \"graphTooltip\": 0,\n \"id\": null,\n \"links\": [],\n \"panels\": [\n {\n \"collapsed\": false,\n \"datasource\": null,\n \"gridPos\": {\n \"h\": 1,\n \"w\": 24,\n \"x\": 0,\n \"y\": 0\n },\n \"id\": 24,\n \"panels\": [],\n \"title\": \"Summary\",\n \"type\": \"row\"\n },\n {\n \"cacheTimeout\": null,\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"fieldConfig\": {\n \"defaults\": {\n \"color\": {\n \"mode\": \"thresholds\"\n },\n \"mappings\": [\n {\n \"options\": {\n \"match\": \"null\",\n \"result\": {\n \"text\": \"N/A\"\n }\n },\n \"type\": \"special\"\n }\n ],\n \"thresholds\": {\n \"mode\": \"absolute\",\n \"steps\": [\n {\n \"color\": \"#E02F44\",\n \"value\": null\n },\n {\n \"color\": \"#E02F44\",\n \"value\": 10\n },\n {\n \"color\": \"#299c46\",\n \"value\": 10\n }\n ]\n },\n \"unit\": \"none\"\n },\n \"overrides\": []\n },\n \"gridPos\": {\n \"h\": 8,\n \"w\": 6,\n \"x\": 0,\n \"y\": 1\n },\n \"id\": 2,\n \"interval\": null,\n \"links\": [],\n \"maxDataPoints\": 100,\n \"options\": {\n \"colorMode\": \"background\",\n \"graphMode\": \"none\",\n \"justifyMode\": \"auto\",\n \"orientation\": \"horizontal\",\n \"reduceOptions\": {\n \"calcs\": [\n \"lastNotNull\"\n ],\n \"fields\": \"\",\n \"values\": false\n },\n \"text\": {},\n \"textMode\": \"auto\"\n },\n \"pluginVersion\": \"8.1.2\",\n \"targets\": [\n {\n \"exemplar\": true,\n \"expr\": \"count(cs_info)\",\n \"interval\": \"\",\n \"legendFormat\": \"\",\n \"refId\": \"A\"\n }\n ],\n \"timeFrom\": null,\n \"timeShift\": null,\n \"title\": \"Running Crowdsec\",\n \"transparent\": true,\n \"type\": \"stat\"\n },\n {\n \"aliasColors\": {},\n \"bars\": false,\n \"dashLength\": 10,\n \"dashes\": false,\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"decimals\": 1,\n \"fieldConfig\": {\n \"defaults\": {\n \"links\": []\n },\n \"overrides\": []\n },\n \"fill\": 1,\n \"fillGradient\": 0,\n \"gridPos\": {\n \"h\": 8,\n \"w\": 18,\n \"x\": 6,\n \"y\": 1\n },\n \"hiddenSeries\": false,\n \"id\": 8,\n \"legend\": {\n \"alignAsTable\": true,\n \"avg\": false,\n \"current\": false,\n \"max\": false,\n \"min\": false,\n \"rightSide\": true,\n \"show\": true,\n \"sort\": \"total\",\n \"sortDesc\": true,\n \"total\": true,\n \"values\": true\n },\n \"lines\": true,\n \"linewidth\": 1,\n \"nullPointMode\": \"null\",\n \"options\": {\n \"alertThreshold\": true\n },\n \"percentage\": false,\n \"pluginVersion\": \"8.1.2\",\n \"pointradius\": 2,\n \"points\": false,\n \"renLine truncated
kind: ConfigMap
metadata:
labels:
app.kubernetes.io/managed-by: manual
grafana_dashboard: "1"
name: crowdsec-crowdsec-overview
namespace: prometheus
---
apiVersion: v1
data:
crowdsec-lapi-metrics.json: "{\n \"__inputs\": [\n {\n \"name\": \"DS_PROMETHEUS\",\n \"label\": \"Prometheus\",\n \"description\": \"\",\n \"type\": \"datasource\",\n \"pluginId\": \"prometheus\",\n \"pluginName\": \"Prometheus\"\n }\n ],\n \"__requires\": [\n {\n \"type\": \"panel\",\n \"id\": \"bargauge\",\n \"name\": \"Bar gauge\",\n \"version\": \"\"\n },\n {\n \"type\": \"grafana\",\n \"id\": \"grafana\",\n \"name\": \"Grafana\",\n \"version\": \"8.1.2\"\n },\n {\n \"type\": \"datasource\",\n \"id\": \"prometheus\",\n \"name\": \"Prometheus\",\n \"version\": \"1.0.0\"\n }\n ],\n \"annotations\": {\n \"list\": [\n {\n \"builtIn\": 1,\n \"datasource\": \"-- Grafana --\",\n \"enable\": true,\n \"hide\": true,\n \"iconColor\": \"rgba(0, 211, 255, 1)\",\n \"name\": \"Annotations & Alerts\",\n \"target\": {\n \"limit\": 100,\n \"matchAny\": false,\n \"tags\": [],\n \"type\": \"dashboard\"\n },\n \"type\": \"dashboard\"\n }\n ]\n },\n \"editable\": true,\n \"gnetId\": null,\n \"graphTooltip\": 0,\n \"id\": null,\n \"iteration\": 1655915193937,\n \"links\": [],\n \"panels\": [\n {\n \"collapsed\": false,\n \"datasource\": null,\n \"gridPos\": {\n \"h\": 1,\n \"w\": 24,\n \"x\": 0,\n \"y\": 0\n },\n \"id\": 10,\n \"panels\": [],\n \"title\": \"Agents\",\n \"type\": \"row\"\n },\n {\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"fieldConfig\": {\n \"defaults\": {\n \"color\": {\n \"mode\": \"thresholds\"\n },\n \"mappings\": [],\n \"thresholds\": {\n \"mode\": \"absolute\",\n \"steps\": [\n {\n \"color\": \"green\",\n \"value\": null\n },\n {\n \"color\": \"red\",\n \"value\": 80\n }\n ]\n }\n },\n \"overrides\": []\n },\n \"gridPos\": {\n \"h\": 8,\n \"w\": 12,\n \"x\": 0,\n \"y\": 1\n },\n \"id\": 2,\n \"options\": {\n \"displayMode\": \"gradient\",\n \"orientation\": \"vertical\",\n \"reduceOptions\": {\n \"calcs\": [\n \"lastNotNull\"\n ],\n \"fields\": \"\",\n \"values\": false\n },\n \"showUnfilled\": false,\n \"text\": {}\n },\n \"pluginVersion\": \"8.1.2\",\n \"repeat\": \"query0\",\n \"repeatDirection\": \"h\",\n \"targets\": [\n {\n \"exemplar\": false,\n \"expr\": \"sum(rate(cs_lapi_request_duration_seconds_bucket{endpoint=\\\"/v1/watchers/login\\\", instance=\\\"$lapi\\\"}[$__rate_interval])) by (le)\",\n \"format\": \"heatmap\",\n \"interval\": \"\",\n \"legendFormat\": \"{{le}}\",\n \"refId\": \"A\"\n }\n ],\n \"title\": \"Agents Login\",\n \"type\": \"heatmap\"\n },\n {\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"fieldConfig\": {\n \"defaults\": {\n \"color\": {\n \"mode\": \"thresholds\"\n },\n \"mappings\": [],\n \"thresholds\": {\n \"mode\": \"absolute\",\n \"steps\": [\n {\n \"color\": \"green\",\n \"value\": null\n }\n ]\n },\n \"unit\": \"none\"\n },\n \"overrides\": []\n },\n \"gridPos\": {\n \"h\": 8,\n \"w\": 12,\n \"x\": 12,\n \"y\": 1\n },\n \"id\": 6,\n \"options\": {\n \"displayMode\": \"gradient\",\n \"orientation\": \"auto\",\n \"reduceOptions\": {\n \"calcs\": [\n \"lastNotNull\"\n ],\n \"fields\": \"\",\n \"values\": false\n },\n \"showUnfilled\": false,\n \"text\": {}\n },\n \"pluginVersion\": \"8.1.2\",\n \"targets\": [\n {\n \"exemplar\": true,\n \"expr\": \"sum(rate(cs_lapi_request_duration_seconds_bucket{endpoint=\\\"/v1/watchers/login\\\"}[$__rate_interval])) by (le)\",\n \"format\": \"heatmap\",\n \"interval\": \"\",\n \"legendFormat\": \"{{le}}\",\n \"refId\": \"A\"\n }\n ],\n \"title\": \"Heartbeat\",\n \"type\": \"heatmap\"\n },\n {\n \"collapsed\": false,\n \"datasource\": null,\n \"gridPos\": {\n \"h\": 1,\n \"w\": 24,\n \"x\": 0,\n \"y\": 9\n },\n \"id\": 12,\n \"panels\": [],\n \"title\": \"Decisions\",\n \"type\": \"row\"\n },\n {\n \"datasource\": \"${DS_PROMETHEUS}\",\n Line truncated
kind: ConfigMap
metadata:
labels:
app.kubernetes.io/managed-by: manual
grafana_dashboard: "1"
name: crowdsec-crowdsec-lapi-metrics
namespace: prometheus
---
apiVersion: v1
data:
crowdsec-insight.json: "{\n \"__inputs\": [\n {\n \"name\": \"DS_PROMETHEUS\",\n \"label\": \"Prometheus\",\n \"description\": \"\",\n \"type\": \"datasource\",\n \"pluginId\": \"prometheus\",\n \"pluginName\": \"Prometheus\"\n }\n ],\n \"__requires\": [\n {\n \"type\": \"panel\",\n \"id\": \"bargauge\",\n \"name\": \"Bar gauge\",\n \"version\": \"\"\n },\n {\n \"type\": \"panel\",\n \"id\": \"gauge\",\n \"name\": \"Gauge\",\n \"version\": \"\"\n },\n {\n \"type\": \"grafana\",\n \"id\": \"grafana\",\n \"name\": \"Grafana\",\n \"version\": \"8.1.2\"\n },\n {\n \"type\": \"datasource\",\n \"id\": \"prometheus\",\n \"name\": \"Prometheus\",\n \"version\": \"1.0.0\"\n },\n {\n \"type\": \"panel\",\n \"id\": \"stat\",\n \"name\": \"Stat\",\n \"version\": \"\"\n }\n ],\n \"annotations\": {\n \"list\": [\n {\n \"builtIn\": 1,\n \"datasource\": \"-- Grafana --\",\n \"enable\": true,\n \"hide\": true,\n \"iconColor\": \"rgba(0, 211, 255, 1)\",\n \"name\": \"Annotations & Alerts\",\n \"target\": {\n \"limit\": 100,\n \"matchAny\": false,\n \"tags\": [],\n \"type\": \"dashboard\"\n },\n \"type\": \"dashboard\"\n }\n ]\n },\n \"editable\": true,\n \"gnetId\": null,\n \"graphTooltip\": 0,\n \"id\": null,\n \"iteration\": 1655915159751,\n \"links\": [],\n \"panels\": [\n {\n \"collapsed\": true,\n \"datasource\": null,\n \"gridPos\": {\n \"h\": 1,\n \"w\": 24,\n \"x\": 0,\n \"y\": 0\n },\n \"id\": 22,\n \"panels\": [\n {\n \"cacheTimeout\": null,\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"fieldConfig\": {\n \"defaults\": {\n \"color\": {\n \"mode\": \"thresholds\"\n },\n \"mappings\": [\n {\n \"options\": {\n \"match\": \"null\",\n \"result\": {\n \"text\": \"N/A\"\n }\n },\n \"type\": \"special\"\n }\n ],\n \"thresholds\": {\n \"mode\": \"absolute\",\n \"steps\": [\n {\n \"color\": \"green\",\n \"value\": null\n },\n {\n \"color\": \"red\",\n \"value\": 80\n }\n ]\n },\n \"unit\": \"dateTimeAsIso\"\n },\n \"overrides\": []\n },\n \"gridPos\": {\n \"h\": 9,\n \"w\": 5,\n \"x\": 2,\n \"y\": 1\n },\n \"id\": 2,\n \"interval\": null,\n \"links\": [],\n \"maxDataPoints\": 100,\n \"options\": {\n \"colorMode\": \"none\",\n \"graphMode\": \"none\",\n \"justifyMode\": \"auto\",\n \"orientation\": \"horizontal\",\n \"reduceOptions\": {\n \"calcs\": [\n \"lastNotNull\"\n ],\n \"fields\": \"\",\n \"values\": false\n },\n \"text\": {},\n \"textMode\": \"auto\"\n },\n \"pluginVersion\": \"8.1.2\",\n \"targets\": [\n {\n \"exemplar\": true,\n \"expr\": \"(process_start_time_seconds{instance=\\\"$instance\\\"})*1000\",\n \"interval\": \"\",\n \"legendFormat\": \"{{instance}}\",\n \"refId\": \"A\"\n }\n ],\n \"timeFrom\": null,\n \"timeShift\": null,\n \"title\": \"Up since\",\n \"type\": \"stat\"\n },\n {\n \"datasource\": \"${DS_PROMETHEUS}\",\n \"fieldConfig\": {\n \"defaults\": {\n \"displayName\": \"\",\n \"mappings\": [],\n \"thresholds\": {\n \"mode\": \"absolute\",\n \"steps\": [\n {\n \"color\": \"green\",\n \"value\": null\n }\n ]\n },\n \"unit\": \"decbytes\"\n },\n \"overrides\": []\n },\n \"gridPos\": {\n \"h\": 9,\n \"w\": 5,\n \"x\": 7,\n \"y\": 1\n },\n \"id\": 4,\n \"options\": {\n \"orientation\": \"auto\",\n \"reduceOptions\": {\n \"calcs\": [\n \"mean\"\n ],\n \"fields\": \"\",\n \"values\": false\n },\Line truncated
kind: ConfigMap
metadata:
labels:
app.kubernetes.io/managed-by: manual
grafana_dashboard: "1"
name: crowdsec-crowdsec-insight
namespace: prometheus
-202
View File
@@ -1,202 +0,0 @@
# CrowdSec self-healing: static machine identity + enforcement loops.
#
# Problem it fixes: the chart's agent init container runs
# `cscli lapi register --machine "$POD_NAME" ...`
# unconditionally. Credentials live in an emptyDir, the machine row lives
# in LAPI's persistent DB. Any init re-run for an already-known pod name
# (kubelet restart, node reboot) dies with
# 403 Forbidden: user '<pod>' already exist
# and the DaemonSet pod sticks in Init forever. Every DS restart also
# leaves an orphan machine row that is never cleaned.
#
# Design (name-independent):
# * Agent identity is a STATIC machine `crowdsec-agent-workstation`
# whose password lives in Secret `crowdsec-agent-credentials`
# (created once, manually - like all other secrets in this repo).
# The secret is mounted into agent pods at
# /tmp_config/local_api_credentials.yaml (see extraVolumeMounts in
# crowdsec-values.yaml), which is exactly the path the agent's main
# container copies into place at startup.
# * The DS init command is patched (strategic merge, by container name)
# to SKIP registration when that file exists, keeping the legacy
# register path only as fallback. Detection marker in the patched
# command: `[ -s /tmp_config`.
# * This CronJob enforces the desired state hourly, so recovery is
# automatic even after `helm upgrade` reverts the DS patch or the
# LAPI database is wiped:
# 1. patch DS init if it still has the unconditional register
# (no-op otherwise - no restart churn);
# 2. prune machines with no heartbeat for 2h (orphan hygiene);
# 3. ensure the static machine exists, recreating it with the
# Secret password if missing (agent retry loops reconnect
# on their own - same name + same password);
# 4. prune bouncer entries idle for 30d.
#
# It used to also delete LePresidente/http-generic-403-bf decisions hourly.
# That was a workaround for the bouncer failing closed on a slow LAPI and
# 403-ing the deploy runner into a 4h ban. The bouncer now polls decisions
# into a cache and never blocks on an unreachable LAPI, so it cannot
# manufacture those 403s any more, and the scenario only fires against real
# scanners - deleting their decisions hourly was undoing a working ban.
#
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
# Force a run:
# kubectl create job -n crowdsec --from=cronjob/crowdsec-janitor janitor-now
#
# Helm upgrades: the janitor's strategic patch puts the DS field under
# the `kubectl-patch` field manager, so a plain `helm upgrade` FAILS
# with an SSA conflict on initContainers[].command. Procedure:
# 1. revert init to chart state (kills the conflict):
# helm template crowdsec crowdsec/crowdsec --version <ver> \
# -n crowdsec -f crowdsec/k8s/crowdsec-values.yaml > /tmp/r.yaml
# python3 -c "import yaml,json; ..." # build revert patch from
# the rendered DaemonSet init command, then
# kubectl patch ds crowdsec-agent -n crowdsec \
# --type strategic -p "\$(cat /tmp/revert_patch.json)"
# 2. helm upgrade --install crowdsec ... (no --force needed)
# 3. janitor-now right away (upgrade reverts init; new pods would
# sit in Init until the next hourly run otherwise).
#
# One-time bootstrap (order matters):
# 1. Create Secret + static machine (see commands in chat).
# 2. Apply this file, trigger janitor-now, wait for agent 1/1.
# 3. One-time orphan cleanup:
# kubectl exec -n crowdsec deploy/crowdsec-lapi -- \
# cscli machines prune --duration 1h --force
# 4. Only then `helm upgrade` crowdsec with the extraVolumes values.
# Upgrade reverts the DS patch; trigger janitor-now right after it
# (otherwise new pods sit in Init until the next hourly run, then
# self-heal anyway).
#
# Password rotation: update the Secret, delete the machine
# (`cscli machines delete crowdsec-agent-workstation`), trigger
# janitor-now (recreates it), then `kubectl rollout restart
# ds/crowdsec-agent -n crowdsec` (agent reads the file at startup only).
apiVersion: v1
kind: ServiceAccount
metadata:
name: crowdsec-janitor
namespace: crowdsec
labels:
app.kubernetes.io/part-of: crowdsec
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: crowdsec-janitor
namespace: crowdsec
labels:
app.kubernetes.io/part-of: crowdsec
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["pods/exec"]
verbs: ["create"]
- apiGroups: ["apps"]
resources: ["daemonsets"]
verbs: ["get", "patch"]
# `kubectl exec deploy/<name>` resolves deploy -> replicaset -> pod,
# which needs read access to these (exec itself is pods/exec above).
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: crowdsec-janitor
namespace: crowdsec
labels:
app.kubernetes.io/part-of: crowdsec
subjects:
- kind: ServiceAccount
name: crowdsec-janitor
namespace: crowdsec
roleRef:
kind: Role
name: crowdsec-janitor
apiGroup: rbac.authorization.k8s.io
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: crowdsec-janitor
namespace: crowdsec
labels:
app.kubernetes.io/part-of: crowdsec
spec:
schedule: "17 * * * *"
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
activeDeadlineSeconds: 300
template:
metadata:
labels:
app.kubernetes.io/part-of: crowdsec
spec:
serviceAccountName: crowdsec-janitor
restartPolicy: OnFailure
containers:
- name: janitor
# Same image the chart itself uses for registration jobs;
# IfNotPresent so it works while the node is offline
# (layer cached from the chart install).
image: alpine/kubectl:latest
imagePullPolicy: IfNotPresent
env:
- name: AGENT_PASSWORD
valueFrom:
secretKeyRef:
name: crowdsec-agent-credentials
key: password
command:
- /bin/sh
- -c
- |
set -eu
LAPI_EXEC="kubectl exec -n crowdsec deploy/crowdsec-lapi --"
echo "== 1. enforce patched agent init =="
CUR=$(kubectl get ds crowdsec-agent -n crowdsec \
-o jsonpath='{.spec.template.spec.initContainers[0].command[2]}')
case "$CUR" in
*'-s /tmp_config'*)
echo "init already patched"
;;
*)
echo "patching init"
WAIT='until nc "$LAPI_HOST" "$LAPI_PORT" -z'
WAIT="$WAIT; do echo waiting for lapi to start; sleep 5; done"
LINK='ln -s /staging/etc/crowdsec /etc/crowdsec'
REG='cscli lapi register --machine "$USERNAME"'
REG="$REG -u \"\$LAPI_URL\" --token \"\$REGISTRATION_TOKEN\""
CREDS=/tmp_config/local_api_credentials.yaml
CMD="$WAIT; $LINK; [ -s $CREDS ] || {"
CMD="$CMD $REG && cp"
CMD="$CMD /etc/crowdsec/local_api_credentials.yaml $CREDS; }"
ESC=$(printf '%s' "$CMD" | sed 's/"/\\"/g')
PATCH='{"spec":{"template":{"spec":{"initContainers":'
PATCH=$PATCH'[{"name":"wait-for-lapi-and-register",'
PATCH=$PATCH'"command":["sh","-c","'$ESC'"]}]}}}}'
kubectl patch ds crowdsec-agent -n crowdsec \
--type strategic -p "$PATCH"
;;
esac
echo "== 2. prune orphan machines (no heartbeat for 2h) =="
$LAPI_EXEC cscli machines prune --duration 2h --force
echo "== 3. ensure static machine exists =="
if $LAPI_EXEC cscli machines inspect \
crowdsec-agent-workstation >/dev/null 2>&1; then
echo "static machine present"
else
echo "recreating static machine"
$LAPI_EXEC cscli machines add crowdsec-agent-workstation \
--password "$AGENT_PASSWORD" --force
fi
echo "== 4. prune stale bouncers (no pull for 30d) =="
$LAPI_EXEC cscli bouncers prune -d 720h --force
-7
View File
@@ -25,10 +25,3 @@ spec:
ports: ports:
- protocol: TCP - protocol: TCP
port: 8080 port: 8080
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: prometheus
ports:
- protocol: TCP
port: 6060
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
dockmon: dockmon:
image: darthnorse/dockmon:2.5.0 image: darthnorse/dockmon:2.4.5
container_name: dockmon container_name: dockmon
restart: unless-stopped restart: unless-stopped
# ports: # ports:
File renamed without changes.
-28
View File
@@ -1,28 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: dockmon-prod-tls
namespace: dockmon
spec:
secretName: dockmon-prod-tls
dnsNames:
- dockmon.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: dockmon
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+1 -1
View File
@@ -29,7 +29,7 @@ spec:
spec: spec:
containers: containers:
- name: dockmon - name: dockmon
image: darthnorse/dockmon:2.5.0 image: darthnorse/dockmon:2.4.5
ports: ports:
- containerPort: 443 - containerPort: 443
volumeMounts: volumeMounts:
+3 -3
View File
@@ -18,13 +18,15 @@ spec:
- match: Host(`dockmon.forust.xyz`) - match: Host(`dockmon.forust.xyz`)
kind: Rule kind: Rule
middlewares: middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-headers@file - name: security-headers@file
services: services:
- name: dockmon-service - name: dockmon-service
port: 443 port: 443
serversTransport: dockmon-transport serversTransport: dockmon-transport
tls: tls:
secretName: dockmon-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -41,5 +43,3 @@ spec:
- name: dockmon-service - name: dockmon-service
port: 443 port: 443
serversTransport: dockmon-transport serversTransport: dockmon-transport
tls:
secretName: internal-wildcard-tls
+1 -1
View File
@@ -1,7 +1,7 @@
services: services:
downtify: downtify:
container_name: downtify container_name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0 image: ghcr.io/henriquesebastiao/downtify:2.13.0
restart: unless-stopped restart: unless-stopped
# ports: # ports:
# - '7077:8000' # - '7077:8000'
+1 -3
View File
@@ -20,8 +20,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: downtify app: downtify
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -29,7 +27,7 @@ spec:
spec: spec:
containers: containers:
- name: downtify - name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0 image: ghcr.io/henriquesebastiao/downtify:2.13.0
ports: ports:
- containerPort: 8000 - containerPort: 8000
volumeMounts: volumeMounts:
+3 -3
View File
@@ -10,12 +10,14 @@ spec:
- match: Host(`downtify.forust.xyz`) - match: Host(`downtify.forust.xyz`)
kind: Rule kind: Rule
middlewares: middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-chain@file - name: security-chain@file
services: services:
- name: downtify-service - name: downtify-service
port: 8000 port: 8000
tls: tls:
secretName: downtify-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -31,5 +33,3 @@ spec:
services: services:
- name: downtify-service - name: downtify-service
port: 8000 port: 8000
tls:
secretName: internal-wildcard-tls
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.9-alpine
WORKDIR /app
# Установка зависимостей
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Копирование кода
COPY main.py .
COPY .env .
# Запуск бота
CMD ["python", "-u", "main.py"]
+373
View File
@@ -0,0 +1,373 @@
Mozilla Public License Version 2.0
==================================
1. Definitions
--------------
1.1. "Contributor"
means each individual or legal entity that creates, contributes to
the creation of, or owns Covered Software.
1.2. "Contributor Version"
means the combination of the Contributions of others (if any) used
by a Contributor and that particular Contributor's Contribution.
1.3. "Contribution"
means Covered Software of a particular Contributor.
1.4. "Covered Software"
means Source Code Form to which the initial Contributor has attached
the notice in Exhibit A, the Executable Form of such Source Code
Form, and Modifications of such Source Code Form, in each case
including portions thereof.
1.5. "Incompatible With Secondary Licenses"
means
(a) that the initial Contributor has attached the notice described
in Exhibit B to the Covered Software; or
(b) that the Covered Software was made available under the terms of
version 1.1 or earlier of the License, but not also under the
terms of a Secondary License.
1.6. "Executable Form"
means any form of the work other than Source Code Form.
1.7. "Larger Work"
means a work that combines Covered Software with other material, in
a separate file or files, that is not Covered Software.
1.8. "License"
means this document.
1.9. "Licensable"
means having the right to grant, to the maximum extent possible,
whether at the time of the initial grant or subsequently, any and
all of the rights conveyed by this License.
1.10. "Modifications"
means any of the following:
(a) any file in Source Code Form that results from an addition to,
deletion from, or modification of the contents of Covered
Software; or
(b) any new file in Source Code Form that contains any Covered
Software.
1.11. "Patent Claims" of a Contributor
means any patent claim(s), including without limitation, method,
process, and apparatus claims, in any patent Licensable by such
Contributor that would be infringed, but for the grant of the
License, by the making, using, selling, offering for sale, having
made, import, or transfer of either its Contributions or its
Contributor Version.
1.12. "Secondary License"
means either the GNU General Public License, Version 2.0, the GNU
Lesser General Public License, Version 2.1, the GNU Affero General
Public License, Version 3.0, or any later versions of those
licenses.
1.13. "Source Code Form"
means the form of the work preferred for making modifications.
1.14. "You" (or "Your")
means an individual or a legal entity exercising rights under this
License. For legal entities, "You" includes any entity that
controls, is controlled by, or is under common control with You. For
purposes of this definition, "control" means (a) the power, direct
or indirect, to cause the direction or management of such entity,
whether by contract or otherwise, or (b) ownership of more than
fifty percent (50%) of the outstanding shares or beneficial
ownership of such entity.
2. License Grants and Conditions
--------------------------------
2.1. Grants
Each Contributor hereby grants You a world-wide, royalty-free,
non-exclusive license:
(a) under intellectual property rights (other than patent or trademark)
Licensable by such Contributor to use, reproduce, make available,
modify, display, perform, distribute, and otherwise exploit its
Contributions, either on an unmodified basis, with Modifications, or
as part of a Larger Work; and
(b) under Patent Claims of such Contributor to make, use, sell, offer
for sale, have made, import, and otherwise transfer either its
Contributions or its Contributor Version.
2.2. Effective Date
The licenses granted in Section 2.1 with respect to any Contribution
become effective for each Contribution on the date the Contributor first
distributes such Contribution.
2.3. Limitations on Grant Scope
The licenses granted in this Section 2 are the only rights granted under
this License. No additional rights or licenses will be implied from the
distribution or licensing of Covered Software under this License.
Notwithstanding Section 2.1(b) above, no patent license is granted by a
Contributor:
(a) for any code that a Contributor has removed from Covered Software;
or
(b) for infringements caused by: (i) Your and any other third party's
modifications of Covered Software, or (ii) the combination of its
Contributions with other software (except as part of its Contributor
Version); or
(c) under Patent Claims infringed by Covered Software in the absence of
its Contributions.
This License does not grant any rights in the trademarks, service marks,
or logos of any Contributor (except as may be necessary to comply with
the notice requirements in Section 3.4).
2.4. Subsequent Licenses
No Contributor makes additional grants as a result of Your choice to
distribute the Covered Software under a subsequent version of this
License (see Section 10.2) or under the terms of a Secondary License (if
permitted under the terms of Section 3.3).
2.5. Representation
Each Contributor represents that the Contributor believes its
Contributions are its original creation(s) or it has sufficient rights
to grant the rights to its Contributions conveyed by this License.
2.6. Fair Use
This License is not intended to limit any rights You have under
applicable copyright doctrines of fair use, fair dealing, or other
equivalents.
2.7. Conditions
Sections 3.1, 3.2, 3.3, and 3.4 are conditions of the licenses granted
in Section 2.1.
3. Responsibilities
-------------------
3.1. Distribution of Source Form
All distribution of Covered Software in Source Code Form, including any
Modifications that You create or to which You contribute, must be under
the terms of this License. You must inform recipients that the Source
Code Form of the Covered Software is governed by the terms of this
License, and how they can obtain a copy of this License. You may not
attempt to alter or restrict the recipients' rights in the Source Code
Form.
3.2. Distribution of Executable Form
If You distribute Covered Software in Executable Form then:
(a) such Covered Software must also be made available in Source Code
Form, as described in Section 3.1, and You must inform recipients of
the Executable Form how they can obtain a copy of such Source Code
Form by reasonable means in a timely manner, at a charge no more
than the cost of distribution to the recipient; and
(b) You may distribute such Executable Form under the terms of this
License, or sublicense it under different terms, provided that the
license for the Executable Form does not attempt to limit or alter
the recipients' rights in the Source Code Form under this License.
3.3. Distribution of a Larger Work
You may create and distribute a Larger Work under terms of Your choice,
provided that You also comply with the requirements of this License for
the Covered Software. If the Larger Work is a combination of Covered
Software with a work governed by one or more Secondary Licenses, and the
Covered Software is not Incompatible With Secondary Licenses, this
License permits You to additionally distribute such Covered Software
under the terms of such Secondary License(s), so that the recipient of
the Larger Work may, at their option, further distribute the Covered
Software under the terms of either this License or such Secondary
License(s).
3.4. Notices
You may not remove or alter the substance of any license notices
(including copyright notices, patent notices, disclaimers of warranty,
or limitations of liability) contained within the Source Code Form of
the Covered Software, except that You may alter any license notices to
the extent required to remedy known factual inaccuracies.
3.5. Application of Additional Terms
You may choose to offer, and to charge a fee for, warranty, support,
indemnity or liability obligations to one or more recipients of Covered
Software. However, You may do so only on Your own behalf, and not on
behalf of any Contributor. You must make it absolutely clear that any
such warranty, support, indemnity, or liability obligation is offered by
You alone, and You hereby agree to indemnify every Contributor for any
liability incurred by such Contributor as a result of warranty, support,
indemnity or liability terms You offer. You may include additional
disclaimers of warranty and limitations of liability specific to any
jurisdiction.
4. Inability to Comply Due to Statute or Regulation
---------------------------------------------------
If it is impossible for You to comply with any of the terms of this
License with respect to some or all of the Covered Software due to
statute, judicial order, or regulation then You must: (a) comply with
the terms of this License to the maximum extent possible; and (b)
describe the limitations and the code they affect. Such description must
be placed in a text file included with all distributions of the Covered
Software under this License. Except to the extent prohibited by statute
or regulation, such description must be sufficiently detailed for a
recipient of ordinary skill to be able to understand it.
5. Termination
--------------
5.1. The rights granted under this License will terminate automatically
if You fail to comply with any of its terms. However, if You become
compliant, then the rights granted under this License from a particular
Contributor are reinstated (a) provisionally, unless and until such
Contributor explicitly and finally terminates Your grants, and (b) on an
ongoing basis, if such Contributor fails to notify You of the
non-compliance by some reasonable means prior to 60 days after You have
come back into compliance. Moreover, Your grants from a particular
Contributor are reinstated on an ongoing basis if such Contributor
notifies You of the non-compliance by some reasonable means, this is the
first time You have received notice of non-compliance with this License
from such Contributor, and You become compliant prior to 30 days after
Your receipt of the notice.
5.2. If You initiate litigation against any entity by asserting a patent
infringement claim (excluding declaratory judgment actions,
counter-claims, and cross-claims) alleging that a Contributor Version
directly or indirectly infringes any patent, then the rights granted to
You by any and all Contributors for the Covered Software under Section
2.1 of this License shall terminate.
5.3. In the event of termination under Sections 5.1 or 5.2 above, all
end user license agreements (excluding distributors and resellers) which
have been validly granted by You or Your distributors under this License
prior to termination shall survive termination.
************************************************************************
* *
* 6. Disclaimer of Warranty *
* ------------------------- *
* *
* Covered Software is provided under this License on an "as is" *
* basis, without warranty of any kind, either expressed, implied, or *
* statutory, including, without limitation, warranties that the *
* Covered Software is free of defects, merchantable, fit for a *
* particular purpose or non-infringing. The entire risk as to the *
* quality and performance of the Covered Software is with You. *
* Should any Covered Software prove defective in any respect, You *
* (not any Contributor) assume the cost of any necessary servicing, *
* repair, or correction. This disclaimer of warranty constitutes an *
* essential part of this License. No use of any Covered Software is *
* authorized under this License except under this disclaimer. *
* *
************************************************************************
************************************************************************
* *
* 7. Limitation of Liability *
* -------------------------- *
* *
* Under no circumstances and under no legal theory, whether tort *
* (including negligence), contract, or otherwise, shall any *
* Contributor, or anyone who distributes Covered Software as *
* permitted above, be liable to You for any direct, indirect, *
* special, incidental, or consequential damages of any character *
* including, without limitation, damages for lost profits, loss of *
* goodwill, work stoppage, computer failure or malfunction, or any *
* and all other commercial damages or losses, even if such party *
* shall have been informed of the possibility of such damages. This *
* limitation of liability shall not apply to liability for death or *
* personal injury resulting from such party's negligence to the *
* extent applicable law prohibits such limitation. Some *
* jurisdictions do not allow the exclusion or limitation of *
* incidental or consequential damages, so this exclusion and *
* limitation may not apply to You. *
* *
************************************************************************
8. Litigation
-------------
Any litigation relating to this License may be brought only in the
courts of a jurisdiction where the defendant maintains its principal
place of business and such litigation shall be governed by laws of that
jurisdiction, without reference to its conflict-of-law provisions.
Nothing in this Section shall prevent a party's ability to bring
cross-claims or counter-claims.
9. Miscellaneous
----------------
This License represents the complete agreement concerning the subject
matter hereof. If any provision of this License is held to be
unenforceable, such provision shall be reformed only to the extent
necessary to make it enforceable. Any law or regulation which provides
that the language of a contract shall be construed against the drafter
shall not be used to construe this License against a Contributor.
10. Versions of the License
---------------------------
10.1. New Versions
Mozilla Foundation is the license steward. Except as provided in Section
10.3, no one other than the license steward has the right to modify or
publish new versions of this License. Each version will be given a
distinguishing version number.
10.2. Effect of New Versions
You may distribute the Covered Software under the terms of the version
of the License under which You originally received the Covered Software,
or under the terms of any subsequent version published by the license
steward.
10.3. Modified Versions
If you create software not governed by this License, and you want to
create a new license for such software, you may create and use a
modified version of this License if you rename the license and remove
any references to the name of the license steward (except to note that
such modified license differs from this License).
10.4. Distributing Source Code Form that is Incompatible With Secondary
Licenses
If You choose to distribute Source Code Form that is Incompatible With
Secondary Licenses under the terms of this version of the License, the
notice described in Exhibit B of this License must be attached.
Exhibit A - Source Code Form License Notice
-------------------------------------------
This Source Code Form is subject to the terms of the Mozilla Public
License, v. 2.0. If a copy of the MPL was not distributed with this
file, You can obtain one at https://mozilla.org/MPL/2.0/.
If it is not possible or desirable to put the notice in a particular
file, then You may include the notice in a location (such as a LICENSE
file in a relevant directory) where a recipient would be likely to look
for such a notice.
You may add additional accurate notices of copyright ownership.
Exhibit B - "Incompatible With Secondary Licenses" Notice
---------------------------------------------------------
This Source Code Form is "Incompatible With Secondary Licenses", as
defined by the Mozilla Public License, v. 2.0.
+15
View File
@@ -0,0 +1,15 @@
services:
dtek_notif:
build:
context: .
dockerfile: Dockerfile
image: gcr.forust.xyz/forust/dtek-notif:latest
pull_policy: build
restart: unless-stopped
environment:
- TZ=Europe/Kyiv
dns:
- 1.1.1.1
- 8.8.8.8
networks:
- default
+748
View File
@@ -0,0 +1,748 @@
import asyncio
import contextlib
import logging
import os
from datetime import datetime, timedelta
import requests
from aiogram import Bot, Dispatcher
from aiogram.filters import Command
from aiogram.types import KeyboardButton, Message
from aiogram.utils.keyboard import ReplyKeyboardBuilder
from bs4 import BeautifulSoup
from dotenv import load_dotenv
# Загрузка переменных окружения
load_dotenv()
# Настройки
TELEGRAM_TOKEN = os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE')
ALLOWED_CHAT_IDS = list(map(int, os.getenv('ALLOWED_CHAT_IDS', '').split(','))) if os.getenv('ALLOWED_CHAT_IDS') else []
CHECK_INTERVAL = int(os.getenv('CHECK_INTERVAL', '120'))
# Параметры для запроса
VOE_CITY_ID = int(os.getenv('VOE_CITY_ID', 'VOE_CITY_ID'))
VOE_STREET_ID = int(os.getenv('VOE_STREET_ID', 'VOE_STREET_ID'))
VOE_HOUSE_ID = int(os.getenv('VOE_HOUSE_ID', 'VOE_HOUSE_ID'))
# Настройка логирования
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Глобальные переменные
bot = Bot(token=TELEGRAM_TOKEN)
dp = Dispatcher()
last_schedule: list[dict] | None = None
last_notification_time: dict[str, datetime] = {}
# ============================================================================
# УТИЛИТЫ
# ============================================================================
def format_time_duration(minutes: int) -> str:
"""Форматирует время из минут в часы и минуты"""
hours = minutes // 60
mins = minutes % 60
if hours == 0:
return f'{mins}м'
elif mins == 0:
return f'{hours}ч'
return f'{hours}ч {mins}м'
def get_day_statistics(day_blocks: list[dict]) -> dict[str, int]:
"""Получает статистику по дню"""
total_minutes = 0
confirmed_minutes = 0
possible_minutes = 0
for block in day_blocks:
for half in [block['first_half'], block['second_half']]:
if half['status'] == 'off':
total_minutes += 30
if half['confirmed']:
confirmed_minutes += 30
else:
possible_minutes += 30
return {'total': total_minutes, 'confirmed': confirmed_minutes, 'possible': possible_minutes}
# ============================================================================
# ПАРСИНГ ДАННЫХ
# ============================================================================
def parse_html(html: str) -> list[dict]:
"""Парсит HTML с графиком отключений (логика от 15.11.2024)"""
soup = BeautifulSoup(html, 'html.parser')
cells = soup.select('.disconnection-detailed-table-cell.cell')
schedule = []
current_hour = 0
current_day = 0
for cell in cells:
if 'legend' in cell.get('class', []) or 'head' in cell.get('class', []):
continue
cell_classes = cell.get('class', [])
# ПРоверка статуса отключения на весь час
full_hour_off = 'has_disconnection' in cell_classes and 'full_hour' in cell_classes
hour_block = cell.select_one('.hour_block')
if not hour_block:
continue
# Проверка подтверждённости отключения для всего часа
cell_confirmed = None
if 'confirm_1' in cell_classes:
cell_confirmed = True
elif 'confirm_0' in cell_classes:
cell_confirmed = False
# Проверка половин часа
left = hour_block.select_one('.half.left')
right = hour_block.select_one('.half.right')
def parse_half(half, is_full_hour_off: bool, cell_confirmed: bool | None = None) -> dict:
"""Парсит половину часа"""
if not half:
return {'status': 'on', 'queue': None, 'confirmed': None}
half_classes = half.get('class', [])
# Если вся ячейка full_hour - используем статус ячейки
if is_full_hour_off:
return {'status': 'off', 'queue': None, 'confirmed': cell_confirmed}
# Определяем статус половины
if 'has_disconnection' in half_classes:
status = 'off'
elif 'no_disconnection' in half_classes:
status = 'on'
else:
status = 'on' # По умолчанию считаем включенным
# Если выключено - ищем подробности
queue = None
confirmed = None
if status == 'off':
disconnection_div = half.select_one('.disconnection')
if disconnection_div:
# Ищем номер черги в title
if disconnection_div.has_attr('title'):
title = disconnection_div['title']
if 'Номер черги' in title or 'Номер черги:' in title:
with contextlib.suppress(BaseException):
queue = title.split(':')[-1].strip()
# Определяем подтверждение
disc_classes = disconnection_div.get('class', [])
if 'disconnection_confirm_1' in disc_classes:
confirmed = True
elif 'disconnection_confirm_0' in disc_classes:
confirmed = False
return {'status': status, 'queue': queue, 'confirmed': confirmed}
first_half_data = parse_half(left, full_hour_off, cell_confirmed)
second_half_data = parse_half(right, full_hour_off, cell_confirmed)
schedule.append(
{
'hour': current_hour,
'day': current_day,
'first_half': first_half_data,
'second_half': second_half_data,
}
)
current_hour += 1
if current_hour >= 24:
current_hour = 0
current_day += 1
return schedule
def get_voe_html(city_id: int, street_id: int, house_id: int) -> str:
"""Получает HTML с сайта VOE"""
url = 'https://www.voe.com.ua/disconnection/detailed?ajax_form=1&_wrapper_format=drupal_ajax'
headers = {
'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8',
'X-Requested-With': 'XMLHttpRequest',
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
}
data = {
'search_type': 0,
'city_id': city_id,
'street_id': street_id,
'house_id': house_id,
'form_build_id': 'form-Irv5aHw1R2FT_Ik2apyHOZ47hTH5xPNH_LQnBrmpSTc',
'form_id': 'disconnection_detailed_search_form',
'_triggering_element_name': 'search',
'_triggering_element_value': 'Показати',
'_drupal_ajax': 1,
}
try:
response = requests.post(url, headers=headers, data=data, timeout=10)
response.raise_for_status()
resp_json = response.json()
insert_html = next((item['data'] for item in resp_json if item.get('command') == 'insert'), None)
if not insert_html:
raise ValueError('HTML не найден в ответе')
return insert_html
except requests.exceptions.RequestException as e:
logger.error(f'Ошибка запроса VOE: {e}')
raise
# ============================================================================
# ФОРМАТИРОВАНИЕ СООБЩЕНИЙ
# ============================================================================
def get_main_keyboard():
"""Создает главную клавиатуру"""
builder = ReplyKeyboardBuilder()
builder.row(KeyboardButton(text='📊 Графік'), KeyboardButton(text='🔄 Оновити'))
builder.row(KeyboardButton(text='📅 Сьогодні'), KeyboardButton(text='📅 Завтра'))
builder.row(KeyboardButton(text='ℹ️ Про бота'))
return builder.as_markup(resize_keyboard=True)
def format_schedule_message(schedule: list[dict], days_to_show: int = 2) -> str:
"""Форматирует полный график на несколько дней"""
lines = [
'⚡️ <b>Графік відключень світла</b>',
f'🕐 Оновлено: {datetime.now().strftime("%d.%m.%Y %H:%M:%S")}',
'─' * 30,
'',
]
start_date = datetime.now()
for day in range(min(days_to_show, 2)):
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
continue
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🌅 <b>Сьогодні</b>' if day == 0 else '🌄 <b>Завтра</b>'
lines.append(f'{day_name} ({date_str})')
# Статистика
stats = get_day_statistics(day_blocks)
if stats['total'] > 0:
lines.append(f'⏱ Всього: <code>{format_time_duration(stats["total"])}</code>')
if stats['confirmed'] > 0:
lines.append(f'🔴 Підтверджено: <code>{format_time_duration(stats["confirmed"])}</code>')
if stats['possible'] > 0:
lines.append(f'🟠 Можливо: <code>{format_time_duration(stats["possible"])}</code>')
else:
lines.append('🟢 <b>Відключень немає!</b>')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 = підтверджено • 🟠 = можливо • 🟢 = світло</i>')
return '\n'.join(lines)
def format_single_day_schedule(schedule: list[dict], day: int) -> str:
"""Форматирует график на один день"""
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
return '❌ Немає даних для цього дня'
start_date = datetime.now()
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🟠 <b>Сьогодні</b>' if day == 0 else '🔶 <b>Завтра</b>'
lines = [f'{day_name} • {date_str}', '']
# Статистика
lines.append('<b>📊 Статистика</b>')
stats = get_day_statistics(day_blocks)
if stats['total'] == 0:
lines.append('└ 🟢 <b>Відключень немає!</b>')
else:
total_time = format_time_duration(stats['total'])
lines.append(f'├ ⏱ Всього: <code>{total_time}</code>')
if stats['confirmed'] > 0:
confirmed_time = format_time_duration(stats['confirmed'])
lines.append(f'├ 🔴 Підтверджено: <code>{confirmed_time}</code>')
if stats['possible'] > 0:
possible_time = format_time_duration(stats['possible'])
lines.append(f'└ 🟠 Можливо: <code>{possible_time}</code>')
else:
lines.append('└ 🟢 Решта часу світло')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
lines.append('<b>⚡️ Розклад відключень</b>')
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 підтверджено • 🟠 можливо • 🟢 світло</i>')
return '\n'.join(lines)
def schedules_differ(old_schedule: list[dict] | None, new_schedule: list[dict] | None) -> bool:
"""Проверяет отличия между графиками"""
if old_schedule is None or new_schedule is None:
return True
if len(old_schedule) != len(new_schedule):
return True
for old, new in zip(old_schedule, new_schedule, strict=False):
if old['day'] >= 2:
break
if old['first_half'] != new['first_half'] or old['second_half'] != new['second_half']:
return True
return False
# ============================================================================
# УВЕДОМЛЕНИЯ
# ============================================================================
async def send_to_all_users(message_text: str, parse_mode: str = 'HTML'):
"""Отправляет сообщение всем пользователям"""
if not ALLOWED_CHAT_IDS:
logger.warning('Нет допущенных ID чатов для отправки уведомлений')
return
for chat_id in ALLOWED_CHAT_IDS:
try:
await bot.send_message(chat_id, message_text, parse_mode=parse_mode)
logger.info(f'✅ Сообщение отправлено пользователю {chat_id}')
except Exception as e:
logger.error(f'❌ Ошибка отправки пользователю {chat_id}: {e}')
await asyncio.sleep(0.5)
async def check_schedule():
"""Проверяет график и отправляет уведомления"""
global last_schedule
try:
logger.info('🔍 Проверка графика...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
if schedules_differ(last_schedule, new_schedule):
logger.info('✨ Обнаружены изменения!')
message = format_schedule_message(new_schedule, days_to_show=2)
if last_schedule is not None:
await send_to_all_users(f'🔄 <b>Графік оновлено!</b>\n\n{message}')
last_schedule = new_schedule
else:
logger.info('✓ Графік без змін')
except Exception as e:
logger.error(f'❌ Ошибка при проверке графика: {e}')
async def check_upcoming_disconnections():
"""Проверяет предстоящие события и отправляет предупреждения за 5 минут"""
global last_notification_time
if last_schedule is None:
return
now = datetime.now()
today_blocks = [b for b in last_schedule if b['day'] == 0]
# Создаем список всех переходов (off -> on или on -> off)
transitions = []
prev_status = None
for block in today_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
minute = 0 if half_idx == 0 else 30
time_str = f'{hour:02d}:{minute:02d}'
current_status = half['status']
# Если статус изменился - это переход
if prev_status is not None and prev_status != current_status:
transitions.append(
{
'hour': hour,
'minute': minute,
'time_str': time_str,
'from_status': prev_status,
'to_status': current_status,
'confirmed': half.get('confirmed'),
'queue': half.get('queue'),
}
)
prev_status = current_status
# Проверяем переходы
for transition in transitions:
event_time = now.replace(hour=transition['hour'], minute=transition['minute'], second=0, microsecond=0)
time_until = (event_time - now).total_seconds() / 60
notification_key = f'{transition["hour"]}:{transition["minute"]}_{transition["to_status"]}'
# Если за 5 минут до события (±1 минута) и еще не отправляли
if 4 <= time_until <= 6:
# Проверяем, не отправляли ли уже уведомление сегодня
if notification_key in last_notification_time:
last_notif_time = last_notification_time[notification_key]
if last_notif_time.date() == now.date():
continue # Уже отправляли сегодня
# Переход на ОТКЛЮЧЕНИЕ (on -> off)
if transition['from_status'] == 'on' and transition['to_status'] == 'off':
icon = '🔴' if transition['confirmed'] else '🟠'
status = 'підтверджено' if transition['confirmed'] else 'можливе'
queue_info = f' (Черга {transition["queue"]})' if transition['queue'] else ''
warning = (
f'⚠️ <b>УВАГА! ВІДКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин\n'
f'Час: <code>{transition["time_str"]}</code>\n'
f'Статус: {icon} {status}{queue_info}'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВІДКЛЮЧЕННЯ в {transition["time_str"]}')
# Переход на ВКЛЮЧЕНИЕ (off -> on)
elif transition['from_status'] == 'off' and transition['to_status'] == 'on':
warning = (
f'✅ <b>УВАГА! ВКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин буде світло\n'
f'Час: <code>{transition["time_str"]}</code>'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВКЛЮЧЕННЯ в {transition["time_str"]}')
async def monitoring_loop():
"""Основной цикл мониторинга"""
await check_schedule()
while True:
try:
await asyncio.sleep(CHECK_INTERVAL)
await check_schedule()
await check_upcoming_disconnections()
except Exception as e:
logger.error(f'Ошибка в цикле мониторинга: {e}')
await asyncio.sleep(5)
# ============================================================================
# ОБРАБОТЧИКИ КОМАНД
# ============================================================================
@dp.message(Command('start'))
async def cmd_start(message: Message):
"""Обработчик /start"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу до цього бота.')
return
await message.answer(
'👋 <b>Ласкаво просимо!</b>\n\n'
'🤖 <b>Бот для моніторингу графіку відключень світла</b>\n\n'
'✨ <b>Можливості:</b>\n'
'• 📊 Перегляд графіку на сьогодні і завтра\n'
'• 🔔 Автоматичні сповіщення за 5 хвилин до подій\n'
'• 🔄 Моніторинг змін графіку\n\n'
'Використовуйте кнопки нижче 👇',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(lambda msg: msg.text == 'ℹ️ Про бота')
async def cmd_info(message: Message):
"""Показывает информацию о боте"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
await message.answer(
'<b>ℹ️ Про бота</b>\n\n'
'🚀 <b>Версія:</b> 2.2 (Стабільна)\n\n'
'📝 <b>Реліз-ноути:</b>\n'
'├ 15.11.2024: Адаптація під оновлену логіку сайту VOE\n'
'├ Виправлено парсинг half.left та half.right\n'
'├ Покращено визначення підтвердження відключень\n'
'└ Оптимізовано обробку статусу для всієї години\n\n'
'⚡ <b>Функціональність:</b>\n'
'├ Моніторинг графіку 24/7\n'
'├ Сповіщення за 5 хвилин\n'
'├ Детальна статистика дня\n'
'└ Красива візуалізація\n\n'
'🔐 <b>Безпека:</b> Використовуються .env файли\n'
'💾 <b>Джерело:</b> voe.com.ua',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(Command('schedule'))
async def cmd_schedule(message: Message):
"""Показывает полный график"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_schedule_message(schedule, days_to_show=2)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('today'))
async def cmd_today(message: Message):
"""Показывает график на сегодня"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку сьогодні...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 0)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('tomorrow'))
async def cmd_tomorrow(message: Message):
"""Показывает график на завтра"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку завтра...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 1)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
# await message.answer("❌ Функція тимчасово недоступна. Чекаємо на оновлення сайту", parse_mode="HTML", reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('check'))
async def cmd_check(message: Message):
"""Принудительная проверка графика"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('🔄 <b>Перевіряю графік...</b>', parse_mode='HTML')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
prefix = (
'✅ <b>Знайдено зміни!</b>\n\n'
if schedules_differ(last_schedule, new_schedule)
else '✓ <b>Графік без змін</b>\n\n'
)
result = prefix + format_schedule_message(new_schedule, days_to_show=2)
await message.answer(result, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message()
async def handle_text(message: Message):
"""Обработчик текстовых сообщений и кнопок"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
text = message.text
# Кнопка "Графік"
if text == '📊 Графік':
await cmd_schedule(message)
# Кнопка "Сьогодні"
elif text == '📅 Сьогодні':
await cmd_today(message)
# Кнопка "Завтра"
elif text == '📅 Завтра':
await cmd_tomorrow(message)
# Кнопка "Оновити"
elif text == '🔄 Оновити':
await cmd_check(message)
# Кнопка "Про бота"
elif text == 'ℹ️ Про бота':
await cmd_info(message)
# Неизвестная команда
else:
await message.answer(
'❓ <b>Команда не розпізнана</b>\n\n'
'Використовуйте кнопки на клавіатурі або команди:\n'
'/start • /today • /tomorrow • /schedule • /check',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
# ============================================================================
# ГЛАВНАЯ ФУНКЦИЯ
# ============================================================================
async def main():
"""Главная функция"""
logger.info('=' * 50)
logger.info('ЗАПУСК БОТА V2.2 (stable 2.2, 15.11.2025)')
logger.info('=' * 50)
if not TELEGRAM_TOKEN or os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE') == TELEGRAM_TOKEN:
logger.error('❌ TELEGRAM_TOKEN не конфігурований! Напишіть токен в .env файл')
return
if not ALLOWED_CHAT_IDS:
logger.error('❌ ALLOWED_CHAT_IDS не конфігуровані! Напишіть ID в .env файл')
return
logger.info(f'📌 Allowed chat ids: {ALLOWED_CHAT_IDS}')
logger.info(f'⏱ Інтервал перевірки: {CHECK_INTERVAL} сек')
logger.info('=' * 50)
# Запускаем мониторинг
monitoring_task = asyncio.create_task(monitoring_loop())
try:
await dp.start_polling(bot)
except KeyboardInterrupt:
logger.info('⏹ Бот зупинений користувачем')
finally:
monitoring_task.cancel()
await bot.session.close()
logger.info('✓ Підключення закрито')
if __name__ == '__main__':
try:
asyncio.run(main())
except KeyboardInterrupt:
logger.info('⏹ Завершено')
+7
View File
@@ -0,0 +1,7 @@
[project]
name = "dtek-notif"
version = "0.1.0"
description = "Add your description here"
readme = "README.md"
requires-python = ">=3.13"
dependencies = []
+5
View File
@@ -0,0 +1,5 @@
requests>=2.31.0
beautifulsoup4>=4.12.0
aiogram>=3.3.0
python-dotenv>=1.0.0
aiohttp>=3.9.0
-1
View File
@@ -1 +0,0 @@
1.56.0
-14
View File
@@ -1,14 +0,0 @@
# EDU deployment ownership
Application source and release builds: `forust/edu-master`.
The homelab pipeline deploys `edu_master/k8s` and preserves explicit image digests.
The application copies in this directory are legacy and are not build inputs.
Do not publish EDU `prod` images from homelab or resolve releases from moving tags.
For an EDU release, validate both images, select their digests in the keeper and
checker manifests, and run the existing homelab validation/apply/verification
helpers against this service. Keep the existing Secret and Redis PVC.
Coordinate Redis authentication changes with both clients and all init/probes;
keep a pre-rollout Redis backup and both previous compatible image references.
The current HTTP checker does not depend on Playwright; check other consumers
before removing the separate browser service.
+4 -4
View File
@@ -1,6 +1,6 @@
services: services:
redis: redis:
image: redis:8.10.2-alpine image: redis:8.10.1-alpine
restart: unless-stopped restart: unless-stopped
volumes: volumes:
- redis-data:/data - redis-data:/data
@@ -11,13 +11,13 @@ services:
retries: 5 retries: 5
playwright-service: playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy image: mcr.microsoft.com/playwright:v1.63.0-jammy
restart: unless-stopped restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper: session-keeper:
build: ./phpsessid-bot build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:prod image: gcr.forust.xyz/forust/session-keeper:latest
pull_policy: build pull_policy: build
env_file: .env env_file: .env
restart: unless-stopped restart: unless-stopped
@@ -33,7 +33,7 @@ services:
webinar-checker: webinar-checker:
build: ./webinar-checker build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod image: gcr.forust.xyz/forust/webinar-checker:latest
pull_policy: build pull_policy: build
env_file: .env env_file: .env
restart: unless-stopped restart: unless-stopped
-115
View File
@@ -1,115 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / http error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
- alert: WebinarCheckerNeverStarted
expr: |
(time() - edu_process_start > 120)
and (webinar_check_last_run_timestamp_seconds == 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker job has not started"
description: "The process exposes metrics but its webinar job has never started."
- alert: WebinarDeliveryPending
expr: edu_delivery_pending > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Webinar notifications await delivery"
description: "Telegram delivery has pending recipients. Check delivery failures and retry status."
- alert: EduRedisUnavailable
expr: edu_redis_connected == 0
for: 2m
labels:
severity: critical
annotations:
summary: "EDU checker cannot reach Redis"
description: "Redis health checks are failing; checker commands and delivery may be unavailable."
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker deployment unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
+1 -13
View File
@@ -10,8 +10,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: edu-master-playwright app: edu-master-playwright
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -19,18 +17,8 @@ spec:
spec: spec:
containers: containers:
- name: playwright - name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker image: mcr.microsoft.com/playwright:v1.63.0-jammy
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command: command:
- npx - npx
- -y - -y
-22
View File
@@ -1,22 +0,0 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: redis-clients-only
namespace: edu-master
spec:
podSelector:
matchLabels:
app: edu-master-redis
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: edu-master-session-keeper
- podSelector:
matchLabels:
app: edu-master-webinar-checker
ports:
- protocol: TCP
port: 6379
+3 -24
View File
@@ -18,29 +18,8 @@ spec:
spec: spec:
containers: containers:
- name: redis - name: redis
image: redis:8.10.2-alpine image: redis:8.10.1-alpine
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
env:
- name: REDIS_PASSWORD
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command:
- /bin/sh
- -ec
- |
case "$REDIS_PASSWORD" in *[!0-9a-fA-F]*|'') echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1;; esac
[ "${#REDIS_PASSWORD}" -eq 64 ] || { echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1; }
umask 077
printf 'requirepass "%s"\n' "$REDIS_PASSWORD" > /tmp/redis-auth.conf
chown redis:redis /tmp/redis-auth.conf
exec docker-entrypoint.sh redis-server /tmp/redis-auth.conf
ports: ports:
- containerPort: 6379 - containerPort: 6379
volumeMounts: volumeMounts:
@@ -49,10 +28,10 @@ spec:
resources: resources:
requests: requests:
cpu: 25m cpu: 25m
memory: 32Mi memory: 64Mi
limits: limits:
cpu: 250m cpu: 250m
memory: 128Mi memory: 256Mi
readinessProbe: readinessProbe:
exec: exec:
command: ["redis-cli", "ping"] command: ["redis-cli", "ping"]
-2
View File
@@ -21,8 +21,6 @@ stringData:
WEBINAR_TELEGRAM_TOKEN: "" WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: "" WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60" WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database # Database
REDIS_HOST: "redis" REDIS_HOST: "redis"
REDIS_PORT: "6379" REDIS_PORT: "6379"
-15
View File
@@ -1,15 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
-16
View File
@@ -1,16 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: webinar-checker
namespace: edu-master
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: edu-master-webinar-checker
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+5 -22
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: session-keeper name: session-keeper
namespace: edu-master namespace: edu-master
labels: labels:
@@ -12,24 +10,14 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: edu-master-session-keeper app: edu-master-session-keeper
strategy:
type: Recreate
template: template:
metadata: metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels: labels:
app: edu-master-session-keeper app: edu-master-session-keeper
spec: spec:
initContainers: initContainers:
- name: wait-redis - name: wait-redis
image: redis:8.10.2-alpine image: redis:8.10.1-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command: command:
- /bin/sh - /bin/sh
- -ec - -ec
@@ -43,23 +31,18 @@ spec:
echo "redis is ready" echo "redis is ready"
containers: containers:
- name: session-keeper - name: session-keeper
image: gcr.forust.xyz/forust/session-keeper@sha256:49285e87cc5bc4cf4ffe190813d87927916c2df8a206daac0aeb7d227c636450 image: gcr.forust.xyz/forust/session-keeper:latest
imagePullPolicy: Always
envFrom: envFrom:
- secretRef: - secretRef:
name: edu-master-secrets name: edu-master-secrets
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
resources: resources:
requests: requests:
cpu: 25m cpu: 25m
memory: 32Mi memory: 96Mi
limits: limits:
cpu: 250m cpu: 250m
memory: 128Mi memory: 256Mi
readinessProbe: readinessProbe:
exec: exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"] command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
+12 -37
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: webinar-checker name: webinar-checker
namespace: edu-master namespace: edu-master
labels: labels:
@@ -12,26 +10,16 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: edu-master-webinar-checker app: edu-master-webinar-checker
strategy:
type: Recreate
template: template:
metadata: metadata:
annotations:
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
labels: labels:
app: edu-master-webinar-checker app: edu-master-webinar-checker
spec: spec:
# Enforces dependency order like compose depends_on: # Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) # redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers: initContainers:
- name: wait-deps - name: wait-deps
image: redis:8.10.2-alpine image: redis:8.10.1-alpine
env:
- name: REDISCLI_AUTH
valueFrom:
secretKeyRef:
name: edu-master-secrets
key: REDIS_PASSWORD
command: command:
- /bin/sh - /bin/sh
- -ec - -ec
@@ -49,29 +37,16 @@ spec:
sleep 2 sleep 2
done done
echo "PHPSESSID ok" echo "PHPSESSID ok"
until nc -z playwright-service 3000; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
sleep 2
done
echo "playwright ok"
containers: containers:
- name: webinar-checker - name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker@sha256:66c146f7b43cb9f0dc31ba9aa36d217e01df42ddafba5971b79c12ec215b2c01 image: gcr.forust.xyz/forust/webinar-checker:latest
ports: imagePullPolicy: Always
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
livenessProbe:
httpGet:
path: /live
port: metrics
initialDelaySeconds: 60
periodSeconds: 15
timeoutSeconds: 3
failureThreshold: 4
envFrom: envFrom:
- secretRef: - secretRef:
name: edu-master-secrets name: edu-master-secrets
@@ -81,7 +56,7 @@ spec:
resources: resources:
requests: requests:
cpu: "50m" cpu: "50m"
memory: "192Mi" memory: "128Mi"
limits: limits:
cpu: "600m" cpu: "600m"
memory: "384Mi" memory: "512Mi"
+1 -1
View File
@@ -1,4 +1,4 @@
FROM python:3.14-slim FROM python:3.11-slim
WORKDIR /app WORKDIR /app
+3 -6
View File
@@ -1,12 +1,9 @@
FROM python:3.14-slim FROM python:3.11-slim
WORKDIR /app WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440 # Install dependencies
ARG PLAYWRIGHT_VERSION=1.56.0 RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==1.56.0 redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py . COPY checker.py .
+144 -277
View File
@@ -1,15 +1,12 @@
import asyncio
import contextlib import contextlib
import json import json
import logging import logging
import os import os
import re import re
import tempfile import tempfile
import threading
import time import time
from datetime import datetime, timedelta from datetime import datetime, timedelta
from html import escape from html import escape
from http.server import BaseHTTPRequestHandler, HTTPServer
import redis import redis
from playwright.async_api import async_playwright from playwright.async_api import async_playwright
@@ -51,106 +48,6 @@ USER_AGENT = _env(
) )
WEBINAR_TELEGRAM_TOKEN = _env('WEBINAR_TELEGRAM_TOKEN') WEBINAR_TELEGRAM_TOKEN = _env('WEBINAR_TELEGRAM_TOKEN')
ADMIN_ID = int(_env('WEBINAR_ADMIN_ID', '0')) ADMIN_ID = int(_env('WEBINAR_ADMIN_ID', '0'))
METRICS_PORT = int(_env('METRICS_PORT', '8000'))
# --- Prometheus metrics (stdlib only, no extra deps) ---
# Scraped by prometheus-stack via ServiceMonitor (edu_master/k8s/servicemonitor.yaml).
# Critical alerts in edu_master/k8s/alerts.yaml fire to Telegram via Alertmanager.
_METRICS_LOCK = threading.Lock()
_METRICS = {
'last_run': 0.0, # Unix ts of last check start
'last_success': 0.0, # Unix ts of last successful check
'last_duration': 0.0, # Duration of last check in seconds
'success_total': 0,
'failure_total': 0,
'consecutive_failures': 0,
'phpsessid_present': 1, # 1 if EDU_PHPSESSID found in redis, else 0
}
def _metric_check_start():
with _METRICS_LOCK:
_METRICS['last_run'] = time.time()
def _metric_check_ok(duration: float):
now = time.time()
with _METRICS_LOCK:
_METRICS['last_success'] = now
_METRICS['last_duration'] = duration
_METRICS['success_total'] += 1
_METRICS['consecutive_failures'] = 0
_METRICS['phpsessid_present'] = 1
def _metric_check_fail(duration: float, phpsessid_missing: bool = False):
with _METRICS_LOCK:
_METRICS['last_duration'] = duration
_METRICS['failure_total'] += 1
_METRICS['consecutive_failures'] += 1
_METRICS['phpsessid_present'] = 0 if phpsessid_missing else 1
def _metrics_render() -> bytes:
with _METRICS_LOCK:
m = dict(_METRICS)
lines = [
'# HELP webinar_check_last_run_timestamp_seconds Unix timestamp of last webinar check start.',
'# TYPE webinar_check_last_run_timestamp_seconds gauge',
f'webinar_check_last_run_timestamp_seconds {m["last_run"]}',
'# HELP webinar_check_last_success_timestamp_seconds Unix timestamp of last successful webinar check.',
'# TYPE webinar_check_last_success_timestamp_seconds gauge',
f'webinar_check_last_success_timestamp_seconds {m["last_success"]}',
'# HELP webinar_check_last_duration_seconds Duration of last webinar check in seconds.',
'# TYPE webinar_check_last_duration_seconds gauge',
f'webinar_check_last_duration_seconds {m["last_duration"]}',
'# HELP webinar_check_success_total Total successful webinar checks.',
'# TYPE webinar_check_success_total counter',
f'webinar_check_success_total {m["success_total"]}',
'# HELP webinar_check_failure_total Total failed webinar checks (timeout, playwright error, page error).',
'# TYPE webinar_check_failure_total counter',
f'webinar_check_failure_total {m["failure_total"]}',
'# HELP webinar_check_consecutive_failures Consecutive failed webinar checks (reset on success).',
'# TYPE webinar_check_consecutive_failures gauge',
f'webinar_check_consecutive_failures {m["consecutive_failures"]}',
'# HELP edu_phpsessid_present 1 if EDU_PHPSESSID exists in redis, 0 otherwise.',
'# TYPE edu_phpsessid_present gauge',
f'edu_phpsessid_present {m["phpsessid_present"]}',
]
return ('\n'.join(lines) + '\n').encode()
class _MetricsHandler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path == '/metrics':
body = _metrics_render()
self.send_response(200)
self.send_header('Content-Type', 'text/plain; version=0.0.4')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
elif self.path in ('/healthz', '/health'):
body = b'ok\n'
self.send_response(200)
self.send_header('Content-Type', 'text/plain')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
else:
self.send_response(404)
self.end_headers()
def log_message(self, *args):
pass # keep bot logs clean
def start_metrics_server(port: int = METRICS_PORT):
server = HTTPServer(('0.0.0.0', port), _MetricsHandler) # noqa: S104 - k8s ServiceMonitor scrapes pod IP
thread = threading.Thread(target=server.serve_forever, name='metrics-server', daemon=True)
thread.start()
logger.info(f'Metrics server listening on :{port}/metrics')
return server
# Redis Keys # Redis Keys
KEY_WHITELIST = 'bot:whitelist' KEY_WHITELIST = 'bot:whitelist'
@@ -700,66 +597,59 @@ async def _collect_event_times(page) -> dict:
async def fetch_diary_data(phpsessid: str) -> dict | None: async def fetch_diary_data(phpsessid: str) -> dict | None:
logger.info('Fetching diary data via Playwright...') logger.info('Fetching diary data via Playwright...')
try: try:
async with asyncio.timeout(60): async with async_playwright() as p:
async with async_playwright() as p: browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15) try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
try: try:
context_browser = await browser.new_context(user_agent=USER_AGENT) await page.goto(DIARY_URL, wait_until='domcontentloaded')
await context_browser.add_cookies( await page.wait_for_selector('table.calendar', timeout=10000)
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}] await page.wait_for_timeout(1500)
)
page = await context_browser.new_page()
try: table_html = await page.evaluate("""
await asyncio.wait_for(page.goto(DIARY_URL, wait_until='domcontentloaded'), timeout=30) () => {
await page.wait_for_selector('table.calendar', timeout=10000) const t = document.querySelector('table.calendar');
await page.wait_for_timeout(1500) return t ? t.outerHTML : null;
}
table_html = await page.evaluate(""" """)
() => { if not table_html:
const t = document.querySelector('table.calendar'); logger.error('table.calendar not found in DOM')
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.calendar not found in DOM')
return None
# Debug: save HTML for troubleshooting
with contextlib.suppress(Exception), open('/tmp/diary_debug.html', 'w', encoding='utf-8') as f: # noqa: S108
f.write(table_html)
month_text, days = _parse_calendar_html(table_html)
# Read event times by opening each event's AJAX popup.
times_by_id = await _collect_event_times(page)
if times_by_id:
for day_data in days.values():
for ev in day_data.get('events', []):
eid = ev.get('id')
if eid and eid in times_by_id:
ev['time'] = times_by_id[eid]
logger.info(
f'Diary parsed: month={month_text!r}, days_with_events={sum(1 for d in days.values() if d["events"])}/{len(days)}'
)
return {'monthFullText': month_text, 'days': days}
except Exception as e:
logger.error(f'Error parsing diary: {e}')
return None return None
finally:
with contextlib.suppress(Exception): # Debug: save HTML for troubleshooting
await asyncio.wait_for(page.close(), timeout=5) with contextlib.suppress(Exception), open('/tmp/diary_debug.html', 'w', encoding='utf-8') as f: # noqa: S108
with contextlib.suppress(Exception): f.write(table_html)
await asyncio.wait_for(context_browser.close(), timeout=5)
month_text, days = _parse_calendar_html(table_html)
# Read event times by opening each event's AJAX popup.
times_by_id = await _collect_event_times(page)
if times_by_id:
for day_data in days.values():
for ev in day_data.get('events', []):
eid = ev.get('id')
if eid and eid in times_by_id:
ev['time'] = times_by_id[eid]
logger.info(
f'Diary parsed: month={month_text!r}, days_with_events={sum(1 for d in days.values() if d["events"])}/{len(days)}'
)
return {'monthFullText': month_text, 'days': days}
except Exception as e:
logger.error(f'Error parsing diary: {e}')
return None
finally: finally:
with contextlib.suppress(Exception): await page.close()
await asyncio.wait_for(browser.close(), timeout=5) await context_browser.close()
except TimeoutError: finally:
logger.error('Diary fetch timed out (60s)') await browser.close()
return None
except Exception as e: except Exception as e:
logger.error(f'Playwright error in diary fetch: {e}') logger.error(f'Playwright error in diary fetch: {e}')
return None return None
@@ -1041,55 +931,48 @@ def _parse_schedule_html(table_html: str) -> dict:
async def fetch_schedule_data(phpsessid: str) -> dict | None: async def fetch_schedule_data(phpsessid: str) -> dict | None:
logger.info('Fetching schedule data via Playwright...') logger.info('Fetching schedule data via Playwright...')
try: try:
async with asyncio.timeout(60): async with async_playwright() as p:
async with async_playwright() as p: browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15) try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
page = await context_browser.new_page()
try: try:
context_browser = await browser.new_context(user_agent=USER_AGENT) await page.goto(SCHEDULE_URL, wait_until='domcontentloaded')
await context_browser.add_cookies( await page.wait_for_selector('table.schedule-table', timeout=10000)
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}] await page.wait_for_timeout(1500)
)
page = await context_browser.new_page()
try: table_html = await page.evaluate("""
await asyncio.wait_for(page.goto(SCHEDULE_URL, wait_until='domcontentloaded'), timeout=30) () => {
await page.wait_for_selector('table.schedule-table', timeout=10000) const t = document.querySelector('table.schedule-table');
await page.wait_for_timeout(1500) return t ? t.outerHTML : null;
}
table_html = await page.evaluate(""" """)
() => { if not table_html:
const t = document.querySelector('table.schedule-table'); logger.error('table.schedule-table not found in DOM')
return t ? t.outerHTML : null;
}
""")
if not table_html:
logger.error('table.schedule-table not found in DOM')
return None
debug_path = os.path.join(tempfile.gettempdir(), 'schedule_debug.html')
with contextlib.suppress(Exception), open(debug_path, 'w', encoding='utf-8') as f:
f.write(table_html)
data = _parse_schedule_html(table_html)
logger.info(list(data['weekdays'].keys()))
logger.info(f'Schedule parsed: {len(data["weekdays"])} days, classes={data["classes"]}')
return data
except Exception as e:
logger.error(f'Error parsing schedule: {e}')
return None return None
finally:
with contextlib.suppress(Exception): debug_path = os.path.join(tempfile.gettempdir(), 'schedule_debug.html')
await asyncio.wait_for(page.close(), timeout=5) with contextlib.suppress(Exception), open(debug_path, 'w', encoding='utf-8') as f:
with contextlib.suppress(Exception): f.write(table_html)
await asyncio.wait_for(context_browser.close(), timeout=5)
data = _parse_schedule_html(table_html)
logger.info(list(data['weekdays'].keys()))
logger.info(f'Schedule parsed: {len(data["weekdays"])} days, classes={data["classes"]}')
return data
except Exception as e:
logger.error(f'Error parsing schedule: {e}')
return None
finally: finally:
with contextlib.suppress(Exception): await page.close()
await asyncio.wait_for(browser.close(), timeout=5) await context_browser.close()
except TimeoutError: finally:
logger.error('Schedule fetch timed out (60s)') await browser.close()
return None
except Exception as e: except Exception as e:
logger.error(f'Playwright error in schedule fetch: {e}') logger.error(f'Playwright error in schedule fetch: {e}')
return None return None
@@ -1602,13 +1485,10 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
int: Number of webinars found, or None if check failed int: Number of webinars found, or None if check failed
""" """
logger.info('Running webinar check...') logger.info('Running webinar check...')
_t0 = time.time()
_metric_check_start()
phpsessid = redis_client.get(KEY_PHPSESSID) phpsessid = redis_client.get(KEY_PHPSESSID)
if not phpsessid: if not phpsessid:
logger.warning('PHPSESSID missing. Skipping check.') logger.warning('PHPSESSID missing. Skipping check.')
_metric_check_fail(time.time() - _t0, phpsessid_missing=True)
# --- DEBUG LOGGING --- # --- DEBUG LOGGING ---
try: try:
with open('phpsessid_missing.log', 'a') as f: with open('phpsessid_missing.log', 'a') as f:
@@ -1622,88 +1502,78 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
content = '' content = ''
try: try:
async with asyncio.timeout(90): async with async_playwright() as p:
async with async_playwright() as p: # Connect to remote Playwright service
# Connect to remote Playwright service browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
try:
# Create browser context with user agent
context_browser = await browser.new_context(user_agent=USER_AGENT)
# Add PHPSESSID cookie
await context_browser.add_cookies(
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}]
)
# Create new page
page = await context_browser.new_page()
try: try:
# Create browser context with user agent # Navigate to webinar page
context_browser = await browser.new_context(user_agent=USER_AGENT) await page.goto(WEBINAR_URL, wait_until='domcontentloaded')
# Add PHPSESSID cookie # Wait for the table to load
await context_browser.add_cookies( await page.wait_for_selector('#meetings table', timeout=10000)
[{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}] await page.wait_for_timeout(2000)
)
# Create new page # Get page content
page = await context_browser.new_page() content = await page.content()
try: # Check if "no webinar" message is present
# Navigate to webinar page if NO_WEBINAR_MARKER not in content:
await asyncio.wait_for(page.goto(WEBINAR_URL, wait_until='domcontentloaded'), timeout=30) logger.info('!!! WEBINAR FOUND !!!')
# Wait for the table to load # Extract webinar details from table rows
await page.wait_for_selector('#meetings table', timeout=10000) rows = page.locator('#meetings table tbody tr')
await page.wait_for_timeout(2000) count = await rows.count()
# Get page content for i in range(count):
content = await page.content() row = rows.nth(i)
text = await row.inner_text()
# Check if "no webinar" message is present if NO_WEBINAR_MARKER not in text:
if NO_WEBINAR_MARKER not in content: # Extract name (topic) from first column
logger.info('!!! WEBINAR FOUND !!!') name_elem = row.locator('td').nth(0)
name = await name_elem.inner_text()
name = name.strip()
# Extract webinar details from table rows # Extract join URL from fourth column
rows = page.locator('#meetings table tbody tr') url_elem = row.locator('td').nth(3).locator('a[href*="/webinar/join/"]').first
count = await rows.count() url = await url_elem.get_attribute('href')
for i in range(count): if name and url:
row = rows.nth(i) current_webinars.append({'name': name, 'url': url, 'text': text.strip()})
text = await row.inner_text() logger.info(f'Found webinar: {name} -> {url}')
else:
logger.info('No webinars found (expected message present)')
if NO_WEBINAR_MARKER not in text: except Exception as e:
# Extract name (topic) from first column logger.error(f'Error checking page: {e}. Saving content for debug.')
name_elem = row.locator('td').nth(0) # If page content is available, save it on error
name = await name_elem.inner_text()
name = name.strip()
# Extract join URL from fourth column
url_elem = row.locator('td').nth(3).locator('a[href*="/webinar/join/"]').first
url = await url_elem.get_attribute('href')
if name and url:
current_webinars.append({'name': name, 'url': url, 'text': text.strip()})
logger.info(f'Found webinar: {name} -> {url}')
else:
logger.info('No webinars found (expected message present)')
except Exception as e:
logger.error(f'Error checking page: {e}. Saving content for debug.')
# If page content is available, save it on error
with contextlib.suppress(Exception):
if page and not content:
content = await page.content()
_metric_check_fail(time.time() - _t0)
return None
finally:
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
finally:
with contextlib.suppress(Exception): with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5) if page and not content:
content = await page.content()
return None
finally:
await page.close()
await context_browser.close()
finally:
await browser.close()
except TimeoutError:
logger.error('Webinar check timed out after 90s (playwright hang)')
_metric_check_fail(time.time() - _t0)
return None
except Exception as e: except Exception as e:
logger.error(f'Playwright error: {e}') logger.error(f'Playwright error: {e}')
_metric_check_fail(time.time() - _t0)
return None return None
# --- DEBUG LOGGING (Saving last response content) --- # --- DEBUG LOGGING (Saving last response content) ---
@@ -1767,7 +1637,6 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
else: else:
logger.info(f'Found {len(current_webinars)} webinar(s), but all are already known') logger.info(f'Found {len(current_webinars)} webinar(s), but all are already known')
_metric_check_ok(time.time() - _t0)
return len(current_webinars) return len(current_webinars)
@@ -1811,8 +1680,6 @@ def main():
job_queue = app.job_queue job_queue = app.job_queue
job_queue.run_repeating(check_webinars_job, interval=WEBINAR_CHECK_INTERVAL, first=10) job_queue.run_repeating(check_webinars_job, interval=WEBINAR_CHECK_INTERVAL, first=10)
start_metrics_server()
logger.info('Bot started polling...') logger.info('Bot started polling...')
app.run_polling() app.run_polling()
+1 -1
View File
@@ -1,7 +1,7 @@
services: services:
errorpage: errorpage:
build: . build: .
image: gcr.forust.xyz/forust/error-pages:prod image: gcr.forust.xyz/forust/error-pages:latest
pull_policy: build pull_policy: build
container_name: error-pages container_name: error-pages
restart: unless-stopped restart: unless-stopped
+1 -17
View File
@@ -20,8 +20,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: error-pages app: error-pages
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -29,21 +27,7 @@ spec:
spec: spec:
containers: containers:
- name: error-pages - name: error-pages
image: gcr.forust.xyz/forust/error-pages:prod image: gcr.forust.xyz/forust/error-pages:latest
# p95 6M, max 10M, no limit before.
resources:
requests:
cpu: "10m"
memory: "32Mi"
limits:
memory: "128Mi"
ports: ports:
- containerPort: 80 - containerPort: 80
readinessProbe:
httpGet:
path: /404.html
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
--- ---
+3 -6
View File
@@ -1,6 +1,6 @@
services: services:
server: server:
image: docker.gitea.com/gitea:28.0.0 image: docker.gitea.com/gitea:1.27.3
container_name: gitea container_name: gitea
restart: always restart: always
environment: environment:
@@ -13,12 +13,9 @@ services:
- GITEA__database__PASSWD=gitea - GITEA__database__PASSWD=gitea
- GITEA__database__NAME=gitea - GITEA__database__NAME=gitea
# Server # Server
- GITEA__server__ROOT_URL=https://git.forust.xyz - GITEA__server__ROOT_URL=https://gitea.forust.xyz
- GITEA__server__SSH_DOMAIN=gitssh.forust.xyz - GITEA__server__SSH_DOMAIN=gitssh.forust.xyz
- GITEA__server__SSH_PORT=2221 - GITEA__server__SSH_PORT=2221
# Pin 28.0 defaults explicitly (see k8s/config.yaml for rationale)
- GITEA__service__DISABLE_REGISTRATION=true
- GITEA__actions__RUN_RETENTION_DAYS=90
# Mailer # Mailer
- GITEA__mailer__ENABLED=true - GITEA__mailer__ENABLED=true
- GITEA__mailer__FROM=${SERVICE_EMAIL} - GITEA__mailer__FROM=${SERVICE_EMAIL}
@@ -37,7 +34,7 @@ services:
- "traefik.http.services.gitea.loadbalancer.server.port=3000" - "traefik.http.services.gitea.loadbalancer.server.port=3000"
# Prod Router # Prod Router
- "traefik.http.routers.gitea.rule=Host(`git.forust.xyz`) || Host(`gitea.forust.xyz`)" - "traefik.http.routers.gitea.rule=Host(`gitea.forust.xyz`)"
- "traefik.http.routers.gitea.entrypoints=websecure" - "traefik.http.routers.gitea.entrypoints=websecure"
- "traefik.http.routers.gitea.tls.certresolver" - "traefik.http.routers.gitea.tls.certresolver"
# Local Router # Local Router
-30
View File
@@ -1,30 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: gitea-prod-tls
namespace: gitea
spec:
secretName: gitea-prod-tls
dnsNames:
- gcr.forust.xyz
- gitea.forust.xyz
- git.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: gitea
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+2 -11
View File
@@ -4,14 +4,11 @@ metadata:
name: gitea-config name: gitea-config
namespace: gitea namespace: gitea
data: data:
GITEA__server__ROOT_URL: "https://git.forust.xyz" GITEA__server__DOMAIN: "gitea.forust.xyz"
GITEA__server__ROOT_URL: "https://gitea.forust.xyz"
GITEA__server__SSH_DOMAIN: "gitssh.forust.xyz" GITEA__server__SSH_DOMAIN: "gitssh.forust.xyz"
GITEA__server__SSH_PORT: "2221" GITEA__server__SSH_PORT: "2221"
GITEA__service__DISABLE_REGISTRATION: "true"
GITEA__actions__RUN_RETENTION_DAYS: "90"
GITEA__database__DB_TYPE: "postgres" GITEA__database__DB_TYPE: "postgres"
GITEA__database__HOST: "postgres.database.svc.cluster.local:5432" GITEA__database__HOST: "postgres.database.svc.cluster.local:5432"
GITEA__database__NAME: "gitea" GITEA__database__NAME: "gitea"
@@ -20,12 +17,6 @@ data:
GITEA__mailer__ENABLED: "false" GITEA__mailer__ENABLED: "false"
# No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres.
GITEA__indexer__ISSUE_INDEXER_TYPE: "db"
GITEA__indexer__REPO_INDEXER_ENABLED: "false"
GITEA__log__logger.access.MODE: "console, file" GITEA__log__logger.access.MODE: "console, file"
USER_UID: "1000" USER_UID: "1000"
USER_GID: "1000" USER_GID: "1000"
+4 -8
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: gitea-deployment name: gitea-deployment
namespace: gitea namespace: gitea
spec: spec:
@@ -26,8 +24,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: gitea app: gitea
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -35,7 +31,7 @@ spec:
spec: spec:
containers: containers:
- name: gitea - name: gitea
image: gitea/gitea:28.0.0 image: docker.gitea.com/gitea:1.27.3
envFrom: envFrom:
- configMapRef: - configMapRef:
name: gitea-config name: gitea-config
@@ -51,10 +47,10 @@ spec:
mountPath: /data mountPath: /data
resources: resources:
requests: requests:
memory: "320Mi" memory: "512Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "1Gi" memory: "1.5Gi"
cpu: "1300m" cpu: "1300m"
volumes: volumes:
- name: gitea-data - name: gitea-data
+9 -6
View File
@@ -7,18 +7,24 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`) - match: Host(`gitea.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
- match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`) - match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
tls: tls:
secretName: gitea-prod-tls certResolver: letsencrypt
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRoute kind: IngressRoute
@@ -29,7 +35,7 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
- match: (Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)) || (Host(`git.workstation.internal`) || Host(`git.gigaforust.internal`)) - match: Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)
kind: Rule kind: Rule
services: services:
- name: gitea-service - name: gitea-service
@@ -39,9 +45,6 @@ spec:
services: services:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
tls:
secretName: internal-wildcard-tls
--- ---
apiVersion: traefik.io/v1alpha1 apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP kind: IngressRouteTCP
File renamed without changes.
-29
View File
@@ -1,29 +0,0 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: glance-prod-tls
namespace: glance
spec:
secretName: glance-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: glance
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+2 -2
View File
@@ -177,8 +177,8 @@ data:
# url: https://gitssh.forust.xyz # url: https://gitssh.forust.xyz
# - title: gcr.forust.xyz # - title: gcr.forust.xyz
# url: https://gcr.forust.xyz/v2/ # url: https://gcr.forust.xyz/v2/
- title: git.forust.xyz - title: gitea.forust.xyz
url: https://git.forust.xyz url: https://gitea.forust.xyz
- title: nextcloud.forust.xyz - title: nextcloud.forust.xyz
url: https://nextcloud.forust.xyz url: https://nextcloud.forust.xyz
- title: mc.forust.xyz - title: mc.forust.xyz
+3 -7
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: glance-deployment name: glance-deployment
namespace: glance namespace: glance
spec: spec:
@@ -22,8 +20,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: glance app: glance
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -61,17 +57,17 @@ spec:
resources: resources:
requests: requests:
cpu: "50m" cpu: "50m"
memory: "32Mi" memory: "64Mi"
limits: limits:
cpu: "200m" cpu: "200m"
memory: "128Mi" memory: "256Mi"
volumes: volumes:
- name: glance-config - name: glance-config
configMap: configMap:
name: glance-config name: glance-config
- name: glance-assets - name: glance-assets
configMap: configMap:
name: glance-assets name: glance-config
- name: docker-socket - name: docker-socket
hostPath: hostPath:
path: /var/run/docker.sock path: /var/run/docker.sock
Loaded 100 of 437 files, more files were not shown because too many files have changed in this diff. Show more