Compare commits

...
Author SHA1 Message Date
forust 5cf0c90258 style: prettier formatting for edu_master alerts and postgres readme
renovate-ci / validate-renovate (push) Skipped
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 0s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 7s
2026-09-23 16:10:33 +02:00
forust 90a452a253 feat(deploy): auto-rollout on push to main (gated by branch protection)
ci / lint-prettier (push) Failing after 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Failing after 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-dockerfiles (pull_request) Successful in 0s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 14s
2026-09-23 15:58:15 +02:00
forust a9eabfba02 fix(userbot): include internal-certificate in kustomize base 2026-09-23 15:52:32 +02:00
forust 46c7e99b1d chore(deploy): rework k8s pipeline, monitoring and postgres 17
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
2026-09-23 15:47:26 +02:00
forust 8b2cf29771 feat(tls): internal CA wildcard for *.internal routes
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 10s
ci / build (push) Successful in 4m51s
ci / deploy-userbot-panel (push) Failing after 2s
2026-09-23 14:47:14 +02:00
forust 7b7fc3bcb0 fix(tls): drop internal certs for dormant namespaces (kener, downtify have no ns live)
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 1s
ci / lint-prettier (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
2026-09-23 14:46:09 +02:00
forust bc1e69ebe0 feat(tls): internal CA wildcard for *.internal routes
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
Selfsigned root (10y) + internal-ca issuer; per-namespace
internal-wildcard-tls certs referenced by all -local routers.
Root public cert committed for client trust stores.
2026-09-23 14:45:05 +02:00
forust f29bb3d580 revert: drop Grafana cert-manager dashboard 2026-09-23 14:42:09 +02:00
forust 1f7166026b feat(observability): cert-manager metrics, dashboard and alerts
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 1s
ci / deploy-userbot-panel (push) Skipped
2026-09-23 14:32:36 +02:00
forust 1cbdfc1d6f feat(observability): cert-manager metrics, dashboard and alerts
- ServiceMonitor for cert-manager (release: prometheus-stack)
- Grafana dashboard Cert-manager-Kubernetes (ID 22908, datasource
  refs fixed) via dashboard sidecar ConfigMap
- PrometheusRule: NotReady (crit), expiry <14d (warn) / <7d
  (crit), ACME 429 rate-limit (warn)
2026-09-23 14:32:36 +02:00
forust 6be3769288 feat(tls): migrate public ingress TLS from Traefik ACME to cert-manager
ci / validate (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 9s
ci / build (push) Successful in 42s
ci / deploy-userbot-panel (push) Skipped
Big-bang: 24 Certificates (HTTP-01, letsencrypt-prod) replace
Traefik certResolver on all prod IngressRoutes. Traefik
certificatesResolvers removed (its acme-http router hijacked
HTTP-01 for every host). AdGuard DoT shares the cert-manager
adguard-certs secret; sync CronJob retired. Dormant files
converted, n8n untouched (live-only deletion rule).
2026-09-23 14:16:23 +02:00
forust ef325cd3b1 feat(tls): migrate public ingress TLS from Traefik ACME to cert-manager
ci / lint-ruff (push) Successful in 1s
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
All prod IngressRoutes switch tls.certResolver to tls.secretName
backed by per-router Certificates (HTTP-01, letsencrypt-prod).
adguard-prod reuses the shared adguard-certs secret (also feeds
DoT :853); sync CronJob removed as redundant.

Traefik certificatesResolvers removed: its internal
acme-http@internal router hijacks HTTP-01 for every host while
enabled, blocking external solvers. Dormant files (kener,
downtify) converted for consistency but not applied; n8n
untouched per live-only rule.
2026-09-23 14:12:36 +02:00
forust aa81f1bf8a feat(adguard): Traefik-synced TLS for AdGuard DoT + reloader
ci / validate (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 1s
ci / deploy-userbot-panel (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
- CronJob mirrors Traefik prod cert dns.forust.xyz into
  adguard-certs (cert-manager HTTP-01 is hijacked by Traefik
  acme-http router, see adguardhome/k8s/cert-sync.yaml)
- reloader for auto-restart on secret rotation
- drop retired adguard.forust.xyz from prod route
- traefik: enable kubernetesIngress, drop unused staging resolver
2026-09-23 14:04:14 +02:00
forust 93d768e988 fix(adguard): split deployment RBAC rule for list/watch
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
Collection verbs cannot combine with resourceNames (grant would
be void). Instance verbs stay name-scoped to adguard-deployment;
list/watch is namespace-scoped (single Deployment in ns).
2026-09-23 13:54:06 +02:00
forust a8f7c79934 fix(adguard): sync job RBAC and idempotency
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
- grant list+watch on adguard-deployment (rollout status hung
  without it, job hit activeDeadline and failed)
- compare content digests only (old hash embedded filenames, so
  every run patched + restarted even when in sync)

Keeps explicit rollout restart alongside reloader annotation:
one extra restart per rotation (~60d) is accepted for
determinism if reloader is down.
2026-09-23 13:52:57 +02:00
forust c42bf14c9a fix(adguard): drop retired adguard.forust.xyz from prod route so dns.forust.xyz gets LE cert
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 1s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
adguard.forust.xyz is NXDOMAIN (host retired); the combo SAN cert kept
failing and dns.forust.xyz served TRAEFIK DEFAULT CERT. Scope prod route
to dns.forust.xyz only.

Also commit live traefik-values state (remove letsencrypt-staging
resolver, live since helm rev 33).
2026-09-23 13:42:04 +02:00
forust 1e8479b853 feat(adguard): sync Traefik prod cert for dns.forust.xyz into adguard-certs
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
CronJob adguard-cert-sync (daily 03:17) copies the public cert/key for
dns.forust.xyz from Traefik acme.json into Secret adguard-certs, which
AdGuard mounts for DNS-over-TLS on :853.

- least-privilege RBAC: read pods/exec in ns traefik, get/update/patch
  Secret adguard-certs and get/patch adguard-deployment in ns adguard
- script selects the PROD resolver entry only, matches main domain or
  SANs, compares sha256 hashes, patches the secret and restarts the
  deployment ONLY on change; exits non-zero and touches nothing when
  Traefik holds no cert yet (HTTP-01 currently cannot complete)
2026-09-23 13:39:38 +02:00
forust c139d700f1 feat(adguard,traefik,cert-manager,reloader): foundation for adguard cert sync
- traefik: enable kubernetesIngress provider (needed for cert-manager HTTP-01)
- adguard: add reloader auto annotation for secret-driven restarts
- cert-manager: namespace, helm values (crds), staging+prod ClusterIssuers
- reloader: namespace manifest
2026-09-23 13:19:26 +02:00
forust 67b9996911 fix(renovate): default endpoint to gitea API URL
ci / lint-prettier (push) Failing after 34s
ci / lint-ruff (push) Failing after 36s
ci / lint-yaml (push) Failing after 27s
ci / lint-dockerfiles (push) Failing after 34s
ci / validate (push) Failing after 21s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
renovate-ci / validate-renovate (push) Failing after 28s
Fall back to https://gitea.forust.xyz/api/v1 when RENOVATE_ENDPOINT is unset, in both local config and in-cluster ConfigMap.
2026-09-23 13:06:19 +02:00
forust 4e9ee567ca Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.106.0' (#41) from renovate/renovate-renovate-44.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 12s
ci / build (push) Successful in 2s
ci / deploy-userbot-panel (push) Skipped
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/41
2026-09-21 20:03:24 +00:00
renovate-bot 4863e13596 chore(deps): update renovate/renovate docker tag to v44.106.0
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-yaml (pull_request) Successful in 3s
ci / lint-dockerfiles (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 2s
renovate-ci / validate-renovate (pull_request) Successful in 7s
ci / build (push) Has been skipped
ci / build (pull_request) Has been skipped
ci / deploy-userbot-panel (push) Has been skipped
ci / deploy-userbot-panel (pull_request) Has been skipped
2026-09-21 16:18:03 +00:00
forust 9c4580a522 fix(prometheus): exclude xui services from TraefikServiceHighLatency
ci / lint-prettier (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 9s
ci / build (push) Successful in 1s
ci / deploy-userbot-panel (push) Has been skipped
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
Long-lived VPN WebSocket sessions inflate P95 request duration; keep 5xx/down/cert alerts for xui unchanged.
2026-09-19 19:06:42 +02:00
forust 4e3ad00202 feat(xui): add 3x-ui VPN panel behind cloudflared tunnel
VLESS+WS inbound (port 10000) via Traefik IngressRoute, panel on internal domains only with public route commented out. gitignore now covers nested k8s secrets and local-only grafana values.
2026-09-19 19:06:37 +02:00
85 changed files with 1582 additions and 166 deletions

No files matched your search

-20
View File
@@ -347,23 +347,3 @@ jobs:
;;
esac
done
deploy-userbot-panel:
needs: build
if: github.ref_name == 'main' && contains(needs.build.outputs.services, 'userbot')
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Apply and roll out userbot panel
shell: bash
run: |
kubectl apply -f userbot/k8s/base/panel.yaml
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
# Keep legacy deployments (forust/anna) in sync with manifests; they have no replicas field, so apply leaves scaling to the user manager only.
kubectl apply -f userbot/k8s/base/userbots.yaml
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
+172 -41
View File
@@ -1,6 +1,9 @@
name: deploy
on:
push:
branches:
- main
workflow_dispatch:
concurrency:
@@ -19,9 +22,6 @@ jobs:
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
# Set APPLY_PRUNE=true to enable kubectl apply --prune. Requires every
# manifest to carry label app.kubernetes.io/managed-by=homelab-deploy,
# otherwise previously applied resources get deleted on the next run.
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
run: |
set -euo pipefail
@@ -57,59 +57,170 @@ jobs:
fi
git -C "$repo" fetch origin main
git -C "$repo" reset --hard origin/main
# Runtime selection: a service is k8s-managed when $SERVICE/k8s/active
# exists. Otherwise it is compose-managed, and only k8s/routing/*
# manifests (external Services / EndpointSlices / ServersTransport /
# Ingresses that route to docker backends) are applied.
# migrate: touch SERVICE/k8s/active (+ move routing files up)
# rollback: rm SERVICE/k8s/active
echo "== Workstation state =="
echo " local: $(git -C "$repo" rev-parse --short HEAD)"
echo " remote: $(git -C "$repo" rev-parse --short origin/main)"
if [ -n "$(git -C "$repo" status --porcelain --untracked-files=no)" ]; then
echo "ERROR: workstation has local tracked modifications, refusing reset:"
git -C "$repo" status --porcelain --untracked-files=no
git -C "$repo" diff --stat
exit 1
fi
git -C "$repo" reset --hard origin/main
cd "$repo"
is_disabled() {
local target="$1"
if [ -f "$target" ]; then
target="$(dirname "$target")"
fi
while true; do
if [ -f "$target/DISABLED" ]; then
return 0
fi
if [ "$target" = "$repo" ]; then
break
fi
target="$(dirname "$target")"
case "$target" in
"$repo"/*) ;;
*) break ;;
esac
done
return 1
}
collect_k8s() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
! -path '*/routing/*' ! -path '*/overlays/*' \
! -name 'kustomization.y*ml' ! -name '*.example.y*ml' \
! -name '*values.y*ml' ! -name 'patch-*.y*ml' \
git ls-files -- "$1" \
| grep -E '\.ya?ml$' \
| grep -Ev '/routing/|/overlays/' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' \
| grep -Ev '(^|/)[^/]*secret[^/]*\.ya?ml$' \
| sort
}
collect_k8s_inactive() {
find "$1" -type f \( -name '*.yaml' -o -name '*.yml' \) \
\( -name 'namespace.y*ml' -o -path '*/routing/*' \) \
! -path '*/overlays/*' ! -name '*.example.y*ml' \
| sort
collect_k8s "$1" \
| grep -E '(^|/)namespace\.ya?ml$|/routing/'
}
mapfile -t compose_stacks < <(
find "$repo" -type f \( -name 'compose.yaml' -o -name 'compose.yml' \) | sort
kustomize_overlay() {
if [ -f "$1/overlays/prod/kustomization.yaml" ]; then
echo "$1/overlays/prod"
elif [ -f "$1/base/kustomization.yaml" ]; then
echo "$1/base"
fi
}
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
mapfile -t k8s_manifests < <(
for kd in $(find "$repo" -type d -name k8s ! -path '*/.git/*' | sort); do
k8s_manifests=()
kustomize_apps=()
for kd_rel in "${k8s_dirs[@]}"; do
kd="$repo/$kd_rel"
if is_disabled "$kd"; then
echo "skip (DISABLED): $kd_rel"
continue
fi
if [ -f "$kd/active" ]; then
collect_k8s "$kd"
overlay="$(kustomize_overlay "$kd" || true)"
if [ -n "${overlay:-}" ]; then
echo "kustomize app: ${overlay#$repo/}"
kustomize_apps+=("$overlay")
else
collect_k8s_inactive "$kd"
while IFS= read -r f; do
[ -n "$f" ] && k8s_manifests+=("$repo/$f")
done < <(collect_k8s "$kd_rel" || true)
fi
else
while IFS= read -r f; do
[ -n "$f" ] && k8s_manifests+=("$repo/$f")
done < <(collect_k8s_inactive "$kd_rel" || true)
fi
done
mapfile -t compose_rel < <(
git ls-files '*/compose.yaml' '*/compose.yml' compose.yaml compose.yml | sort
)
compose_stacks=()
for cf_rel in "${compose_rel[@]}"; do
cf="$repo/$cf_rel"
if is_disabled "$cf"; then
echo "skip (DISABLED): $cf_rel"
continue
fi
if [ -f "$(dirname "$cf")/k8s/active" ]; then
echo "skip (k8s-managed): $cf_rel"
continue
fi
compose_stacks+=("$cf")
done
echo "== Validate compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " config: $cf"
docker compose -f "$cf" config --quiet
done
echo "== Validate k8s manifests (kubectl dry-run) =="
echo "== Validate k8s manifests (kubectl dry-run=client) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=client $m"
kubectl apply --dry-run=client -f "$m" >/dev/null
done
for k in "${kustomize_apps[@]}"; do
echo " apply -k --dry-run=client $k"
kubectl apply -k "$k" --dry-run=client >/dev/null
done
echo "== Validate k8s manifests (kubectl dry-run=server) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=server $m"
kubectl apply --dry-run=server -f "$m" >/dev/null
done
for k in "${kustomize_apps[@]}"; do
echo " apply -k --dry-run=server $k"
kubectl apply -k "$k" --dry-run=server >/dev/null
done
echo "== Checking referenced Secrets exist =="
echo " (deploy never applies *secret*.yaml; create missing ones from the laptop)"
ref_secrets=()
if [ "${#k8s_manifests[@]}" -gt 0 ]; then
while IFS= read -r s; do
[ -n "$s" ] && ref_secrets+=("$s")
done < <(
{
grep -h -A1 -E 'secretRef:|secretKeyRef:' "${k8s_manifests[@]}" 2>/dev/null || true
grep -h -E 'secretName:' "${k8s_manifests[@]}" 2>/dev/null || true
} | grep -E 'name:' | sed -E 's/.*name:[[:space:]]*//' | tr -d '"'"'"' "'"'" | sed -E 's/[[:space:]]*#.*//' | awk 'NF' | sort -u || true
)
fi
missing_secrets=()
all_secrets="$(kubectl get secrets -A --no-headers -o custom-columns=:metadata.name 2>/dev/null || true)"
for s in "${ref_secrets[@]}"; do
if printf '%s\n' "$all_secrets" | grep -qx "$s"; then
echo " ok: $s"
else
echo " MISSING: $s"
missing_secrets+=("$s")
fi
done
if [ "${#missing_secrets[@]}" -gt 0 ]; then
echo "ERROR: ${#missing_secrets[@]} referenced Secret(s) not found in the cluster:"
printf ' - %s\n' "${missing_secrets[@]}"
echo "Create them manually from the laptop, e.g.:"
echo " kubectl apply -f SERVICE/k8s/secrets.yaml # see SERVICE/k8s/secrets.yaml.example"
exit 1
fi
echo "== Applying Kubernetes manifests =="
ns_files=()
@@ -130,15 +241,19 @@ jobs:
echo " namespaces first: ${ns_files[*]}"
kubectl apply -f "${ns_files[@]}"
fi
if [ -f "$repo/prometheus-stack/k8s/active" ]; then
if [ -f "$repo/prometheus-stack/k8s/active" ] && ! is_disabled "$repo/prometheus-stack/k8s"; then
if [ ! -f "$repo/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
exit 1
fi
echo "== Upgrading kube-prometheus-stack =="
helm upgrade --install prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace prometheus \
--version 86.2.3 \
--values "$repo/prometheus-stack/k8s/grafana-values.yaml" \
--wait
--wait --timeout 10m
fi
if [ -f "$repo/loki/k8s/active" ]; then
if [ -f "$repo/loki/k8s/active" ] && ! is_disabled "$repo/loki/k8s"; then
echo "== Upgrading loki/alloy =="
helm repo add grafana https://grafana.github.io/helm-charts >/dev/null 2>&1 || true
helm repo update grafana >/dev/null 2>&1 || true
@@ -146,12 +261,12 @@ jobs:
--version 7.3.0 \
--namespace prometheus \
--values "$repo/loki/k8s/loki-values.yaml" \
--wait
--wait --timeout 10m
helm upgrade --install alloy grafana/alloy \
--version 1.12.1 \
--namespace prometheus \
--values "$repo/loki/k8s/alloy-values.yaml" \
--wait
--wait --timeout 10m
fi
if [ "${#other_files[@]}" -gt 0 ]; then
@@ -159,14 +274,30 @@ jobs:
kubectl apply "${prune_opts[@]}" -f "${other_files[@]}"
fi
for k in "${kustomize_apps[@]}"; do
echo "== Applying kustomize app: ${k#$repo/} =="
kubectl apply -k "$k"
done
if [ -f "$repo/userbot/k8s/active" ] && ! is_disabled "$repo/userbot"; then
echo "== userbot panel hook =="
if kubectl get secret userbot-common-secrets -n userbot >/dev/null 2>&1; then
echo " userbot-common-secrets already present in userbot ns, not touching"
elif kubectl get secret userbot-common-secrets -n default >/dev/null 2>&1; then
echo " bootstrapping userbot-common-secrets into userbot ns"
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
else
echo " WARNING: userbot-common-secrets missing in both default and userbot ns; create it manually from the laptop"
fi
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
fi
echo "== Redeploying docker compose stacks =="
for cf in "${compose_stacks[@]}"; do
dir=$(dirname "$cf")
if [ -f "$dir/k8s/active" ]; then
echo " skip (k8s-managed): $dir"
continue
fi
echo " compose: $dir"
echo " compose: $cf"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
+4
View File
@@ -104,6 +104,10 @@ temp/*
# kubernetes
*/k8s/*secret*
!*/k8s/*secret*.example
**/k8s/*secret*
!**/k8s/*secret*.example
# Local-only tweaks, not for upstream
prometheus-stack/k8s/grafana-values.yaml
traefik/k8s/local-tls.yaml
converters/k8s/config.yaml
convertx/k8s/config.yaml
+2
View File
@@ -62,6 +62,8 @@ spec:
metadata:
labels:
app: adguard
annotations:
reloader.stakater.com/auto: "true"
spec:
containers:
- name: adguard
+12
View File
@@ -0,0 +1,12 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: adguard-certs
namespace: adguard
spec:
secretName: adguard-certs
dnsNames:
- dns.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
+5 -3
View File
@@ -7,7 +7,7 @@ spec:
entryPoints:
- websecure
routes:
- match: Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)
- match: Host(`dns.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
@@ -15,13 +15,13 @@ spec:
services:
- name: adguard-service
port: 3000
- match: (Host(`adguard.forust.xyz`) || Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`)
- match: (Host(`dns.forust.xyz`)) && PathPrefix(`/dns-query`)
kind: Rule
services:
- name: adguard-service
port: 3000
tls:
certResolver: letsencrypt
secretName: adguard-certs
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -42,3 +42,5 @@ spec:
services:
- name: adguard-service
port: 3000
tls:
secretName: internal-wildcard-tls
+15
View File
@@ -0,0 +1,15 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: adguard
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: authentik-prod-tls
namespace: authentik
spec:
secretName: authentik-prod-tls
dnsNames:
- auth.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: authentik
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: authentik-server-service
port: 9000
tls:
certResolver: letsencrypt
secretName: authentik-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: authentik-server-service
port: 9000
tls:
secretName: internal-wildcard-tls
@@ -0,0 +1,9 @@
crds:
enabled: true
prometheus:
servicemonitor:
enabled: true
interval: 60s
scrapeTimeout: 30s
labels:
release: prometheus-stack
+29
View File
@@ -0,0 +1,29 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-staging
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-staging-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-staging-account-key
solvers:
- http01:
ingress:
class: traefik
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod
spec:
acme:
email: bobrovod@national.shitposting.agency
server: https://acme-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-prod-account-key
solvers:
- http01:
ingress:
class: traefik
@@ -0,0 +1,30 @@
-----BEGIN CERTIFICATE-----
MIIFFjCCAv6gAwIBAgIUetKpTfEDOn2985FFMu6G26itT+wwDQYJKoZIhvcNAQEN
BQAwIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRlcm5hbCByb290IENBMB4XDTI2MDky
MzEyNDA0N1oXDTM2MDkyMDEyNDA0N1owIzEhMB8GA1UEAxMYaG9tZWxhYiBpbnRl
cm5hbCByb290IENBMIICIjANBgkqhkiG9w0BAQEFAAOCAg8AMIICCgKCAgEAvmNP
ZCOoD8NtNuYJKVXBlTPjX7D7sJCSK5neH7ZbYV5+lmUlEErY8Mik7j37V5k5NfpF
Ig85pOjP7RckTPz5V6ek3yaN40s4AL053sN5ZPauDVYjalaEHTgj5sEMqlLACQWI
yZmJOZspZykae8dIpQnqCoFpRT4FurJ78v4a0ylnFVLMQn/lyCHedwTjkEdtYWYr
ccJy8vQwqkzs/rWvEH1lDqZhennLOrmcCjfonG7D/pruMn4z+6E28p4+ejkRrI6x
luak3KnpT1XMeHtgU21hiRGaMDBchHMFgAhnY1qosymKenXvfTZItwgjZbwa1hJI
GAiDm+jQDKMjzRZ3rH6Xfc0auUcykNz73PpNu1NGm78nndXwCXcXn1LFKNQJ+r1U
sJiyAmUZmXVn4aM4OMf2F38k7wTYIKg7nRGaUkNeKDlNkjA4HvgWw+jwO1KmdHQ/
mOem1rosDWHRK01wg+Gga9mQCnhNhxglg3t/UeSic6uOaRsvaz4qkzHq8MbCujVz
DpKQjqdikYOAXZOs4KlBLWrS7NaK4NzfSD02pBUErh54ruJfY/bWz9KyXzBD/lQZ
VUTKyvUVB0bkVHEdf1jJmX3H4IZRQSF5JPqOBotW6bJI5fEGNBvj9Zxy4nm2WWGz
yyP3uWsQz8U/Wdx9nXZLHInTZBsvgLYtKUAWA30CAwEAAaNCMEAwDgYDVR0PAQH/
BAQDAgKkMA8GA1UdEwEB/wQFMAMBAf8wHQYDVR0OBBYEFEkKm2rxPaK6+O9WD80z
BLC6F9QsMA0GCSqGSIb3DQEBDQUAA4ICAQAnFyHz97Umf5VIu+dKTJid7C73VugJ
TIar/xJBs/4CxP+znBxhJjXygRoyIfzoVGWcB2ZSL//vL78Qlts79K/Imc9a4RFF
wMvCxsRXAEQ4TpeWi3ophPNcs4rhsP+gQKQFtnyKP9519bqpfxp0bTqwOV2o18fn
za7rlQViiEnNV58j7CVoM9+mJvVVfBEX1Km+GyJL9GadzbIQ7FxClVJZefCbft93
zHVk9gDOw8ys1XGSR2OUCyCLinXO6mqS16CmBb2MAKXq/YyH7E0N8iotAPGtfA8V
M/0ddy947rY0xCrtECfWwvGQpJS7NRv/Z9b2jCfXrI5LXmL2nfQRg0y9GE4Vjwr+
WxtGU5jOeFt0jQ+xRzcgG0Op+qK3x55l5LSo2hOcOVYbiHxcHEJFgwNi1ADeBFwb
q/HdysfURSOghqjIpMMAUabBp+DBUg2EUF7pIaUqbdqExFYcr9EYisEMiNsmKmN+
8ZbcOeerFKDQj+t/R0bFXa7UBn2UWsjI8zlR74aa2kLDXwtyz/XlO/FlYm66eBFo
2/eYUSeU+S4ej+wUAs/dvjF7f190/DUQGuwOTlLTahqWDztmhCk7qzbECu56CwKT
E5Ect2P72UleYwdblkVOVd352AmiwEzdOaziIRrPh8uenEknH6JBYPO3mjk7cCg+
GPGcNIBctjdXhg==
-----END CERTIFICATE-----
+33
View File
@@ -0,0 +1,33 @@
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: selfsigned
spec:
selfSigned: {}
---
# Homelab internal root CA (10y). Install the .crt on clients (see below).
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-ca-root
namespace: cert-manager
spec:
isCA: true
commonName: homelab internal root CA
duration: 87600h
renewBefore: 7200h
secretName: internal-ca-root
privateKey:
algorithm: RSA
size: 4096
issuerRef:
name: selfsigned
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: internal-ca
spec:
ca:
secretName: internal-ca-root
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: cert-manager
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: checkmk-prod-tls
namespace: checkmk
spec:
secretName: checkmk-prod-tls
dnsNames:
- cmk.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: checkmk
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: checkmk-service
port: 5000
tls:
certResolver: letsencrypt
secretName: checkmk-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP
@@ -48,3 +48,5 @@ spec:
services:
- name: checkmk-service
port: 5000
tls:
secretName: internal-wildcard-tls
+42
View File
@@ -0,0 +1,42 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: convertx-prod-tls
namespace: converters
spec:
secretName: convertx-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: bentopdf-prod-tls
namespace: converters
spec:
secretName: bentopdf-prod-tls
dnsNames:
- pdf.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: converters
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+7 -2
View File
@@ -14,7 +14,7 @@ spec:
- name: convertx-service
port: 3000
tls:
certResolver: letsencrypt
secretName: convertx-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -31,6 +31,9 @@ spec:
services:
- name: convertx-service
port: 3000
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -47,7 +50,7 @@ spec:
- name: bentopdf-service
port: 8080
tls:
certResolver: letsencrypt
secretName: bentopdf-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -63,3 +66,5 @@ spec:
services:
- name: bentopdf-service
port: 8080
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: dockmon-prod-tls
namespace: dockmon
spec:
secretName: dockmon-prod-tls
dnsNames:
- dockmon.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: dockmon
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -26,7 +26,7 @@ spec:
port: 443
serversTransport: dockmon-transport
tls:
certResolver: letsencrypt
secretName: dockmon-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -43,3 +43,5 @@ spec:
- name: dockmon-service
port: 443
serversTransport: dockmon-transport
tls:
secretName: internal-wildcard-tls
+3 -1
View File
@@ -17,7 +17,7 @@ spec:
- name: downtify-service
port: 8000
tls:
certResolver: letsencrypt
secretName: downtify-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -33,3 +33,5 @@ spec:
services:
- name: downtify-service
port: 8000
tls:
secretName: internal-wildcard-tls
+1
View File
@@ -0,0 +1 @@
1.56.0
+1 -1
View File
@@ -11,7 +11,7 @@ services:
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.63.0-jammy
image: mcr.microsoft.com/playwright:v1.56.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
+77
View File
@@ -0,0 +1,77 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
(time() - webinar_check_last_success_timestamp_seconds > 300)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
+2 -1
View File
@@ -17,7 +17,8 @@ spec:
spec:
containers:
- name: playwright
image: mcr.microsoft.com/playwright:v1.63.0-jammy
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
command:
- npx
+2
View File
@@ -21,6 +21,8 @@ stringData:
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
+15
View File
@@ -0,0 +1,15 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
+16
View File
@@ -0,0 +1,16 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: webinar-checker
namespace: edu-master
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: edu-master-webinar-checker
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+4
View File
@@ -47,6 +47,10 @@ spec:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:latest
imagePullPolicy: Always
ports:
- name: metrics
containerPort: 8000
protocol: TCP
envFrom:
- secretRef:
name: edu-master-secrets
+5 -2
View File
@@ -2,8 +2,11 @@ FROM python:3.11-slim
WORKDIR /app
# Install dependencies
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==1.56.0 redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
+148 -15
View File
@@ -1,12 +1,15 @@
import asyncio
import contextlib
import json
import logging
import os
import re
import tempfile
import threading
import time
from datetime import datetime, timedelta
from html import escape
from http.server import BaseHTTPRequestHandler, HTTPServer
import redis
from playwright.async_api import async_playwright
@@ -48,6 +51,106 @@ USER_AGENT = _env(
)
WEBINAR_TELEGRAM_TOKEN = _env('WEBINAR_TELEGRAM_TOKEN')
ADMIN_ID = int(_env('WEBINAR_ADMIN_ID', '0'))
METRICS_PORT = int(_env('METRICS_PORT', '8000'))
# --- Prometheus metrics (stdlib only, no extra deps) ---
# Scraped by prometheus-stack via ServiceMonitor (edu_master/k8s/servicemonitor.yaml).
# Critical alerts in edu_master/k8s/alerts.yaml fire to Telegram via Alertmanager.
_METRICS_LOCK = threading.Lock()
_METRICS = {
'last_run': 0.0, # Unix ts of last check start
'last_success': 0.0, # Unix ts of last successful check
'last_duration': 0.0, # Duration of last check in seconds
'success_total': 0,
'failure_total': 0,
'consecutive_failures': 0,
'phpsessid_present': 1, # 1 if EDU_PHPSESSID found in redis, else 0
}
def _metric_check_start():
with _METRICS_LOCK:
_METRICS['last_run'] = time.time()
def _metric_check_ok(duration: float):
now = time.time()
with _METRICS_LOCK:
_METRICS['last_success'] = now
_METRICS['last_duration'] = duration
_METRICS['success_total'] += 1
_METRICS['consecutive_failures'] = 0
_METRICS['phpsessid_present'] = 1
def _metric_check_fail(duration: float, phpsessid_missing: bool = False):
with _METRICS_LOCK:
_METRICS['last_duration'] = duration
_METRICS['failure_total'] += 1
_METRICS['consecutive_failures'] += 1
_METRICS['phpsessid_present'] = 0 if phpsessid_missing else 1
def _metrics_render() -> bytes:
with _METRICS_LOCK:
m = dict(_METRICS)
lines = [
'# HELP webinar_check_last_run_timestamp_seconds Unix timestamp of last webinar check start.',
'# TYPE webinar_check_last_run_timestamp_seconds gauge',
f'webinar_check_last_run_timestamp_seconds {m["last_run"]}',
'# HELP webinar_check_last_success_timestamp_seconds Unix timestamp of last successful webinar check.',
'# TYPE webinar_check_last_success_timestamp_seconds gauge',
f'webinar_check_last_success_timestamp_seconds {m["last_success"]}',
'# HELP webinar_check_last_duration_seconds Duration of last webinar check in seconds.',
'# TYPE webinar_check_last_duration_seconds gauge',
f'webinar_check_last_duration_seconds {m["last_duration"]}',
'# HELP webinar_check_success_total Total successful webinar checks.',
'# TYPE webinar_check_success_total counter',
f'webinar_check_success_total {m["success_total"]}',
'# HELP webinar_check_failure_total Total failed webinar checks (timeout, playwright error, page error).',
'# TYPE webinar_check_failure_total counter',
f'webinar_check_failure_total {m["failure_total"]}',
'# HELP webinar_check_consecutive_failures Consecutive failed webinar checks (reset on success).',
'# TYPE webinar_check_consecutive_failures gauge',
f'webinar_check_consecutive_failures {m["consecutive_failures"]}',
'# HELP edu_phpsessid_present 1 if EDU_PHPSESSID exists in redis, 0 otherwise.',
'# TYPE edu_phpsessid_present gauge',
f'edu_phpsessid_present {m["phpsessid_present"]}',
]
return ('\n'.join(lines) + '\n').encode()
class _MetricsHandler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path == '/metrics':
body = _metrics_render()
self.send_response(200)
self.send_header('Content-Type', 'text/plain; version=0.0.4')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
elif self.path in ('/healthz', '/health'):
body = b'ok\n'
self.send_response(200)
self.send_header('Content-Type', 'text/plain')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
else:
self.send_response(404)
self.end_headers()
def log_message(self, *args):
pass # keep bot logs clean
def start_metrics_server(port: int = METRICS_PORT):
server = HTTPServer(('0.0.0.0', port), _MetricsHandler) # noqa: S104 - k8s ServiceMonitor scrapes pod IP
thread = threading.Thread(target=server.serve_forever, name='metrics-server', daemon=True)
thread.start()
logger.info(f'Metrics server listening on :{port}/metrics')
return server
# Redis Keys
KEY_WHITELIST = 'bot:whitelist'
@@ -597,8 +700,9 @@ async def _collect_event_times(page) -> dict:
async def fetch_diary_data(phpsessid: str) -> dict | None:
logger.info('Fetching diary data via Playwright...')
try:
async with asyncio.timeout(60):
async with async_playwright() as p:
browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
@@ -607,7 +711,7 @@ async def fetch_diary_data(phpsessid: str) -> dict | None:
page = await context_browser.new_page()
try:
await page.goto(DIARY_URL, wait_until='domcontentloaded')
await asyncio.wait_for(page.goto(DIARY_URL, wait_until='domcontentloaded'), timeout=30)
await page.wait_for_selector('table.calendar', timeout=10000)
await page.wait_for_timeout(1500)
@@ -646,10 +750,16 @@ async def fetch_diary_data(phpsessid: str) -> dict | None:
logger.error(f'Error parsing diary: {e}')
return None
finally:
await page.close()
await context_browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
finally:
await browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
except TimeoutError:
logger.error('Diary fetch timed out (60s)')
return None
except Exception as e:
logger.error(f'Playwright error in diary fetch: {e}')
return None
@@ -931,8 +1041,9 @@ def _parse_schedule_html(table_html: str) -> dict:
async def fetch_schedule_data(phpsessid: str) -> dict | None:
logger.info('Fetching schedule data via Playwright...')
try:
async with asyncio.timeout(60):
async with async_playwright() as p:
browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
try:
context_browser = await browser.new_context(user_agent=USER_AGENT)
await context_browser.add_cookies(
@@ -941,7 +1052,7 @@ async def fetch_schedule_data(phpsessid: str) -> dict | None:
page = await context_browser.new_page()
try:
await page.goto(SCHEDULE_URL, wait_until='domcontentloaded')
await asyncio.wait_for(page.goto(SCHEDULE_URL, wait_until='domcontentloaded'), timeout=30)
await page.wait_for_selector('table.schedule-table', timeout=10000)
await page.wait_for_timeout(1500)
@@ -969,10 +1080,16 @@ async def fetch_schedule_data(phpsessid: str) -> dict | None:
logger.error(f'Error parsing schedule: {e}')
return None
finally:
await page.close()
await context_browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
finally:
await browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
except TimeoutError:
logger.error('Schedule fetch timed out (60s)')
return None
except Exception as e:
logger.error(f'Playwright error in schedule fetch: {e}')
return None
@@ -1485,10 +1602,13 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
int: Number of webinars found, or None if check failed
"""
logger.info('Running webinar check...')
_t0 = time.time()
_metric_check_start()
phpsessid = redis_client.get(KEY_PHPSESSID)
if not phpsessid:
logger.warning('PHPSESSID missing. Skipping check.')
_metric_check_fail(time.time() - _t0, phpsessid_missing=True)
# --- DEBUG LOGGING ---
try:
with open('phpsessid_missing.log', 'a') as f:
@@ -1502,9 +1622,10 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
content = ''
try:
async with asyncio.timeout(90):
async with async_playwright() as p:
# Connect to remote Playwright service
browser = await p.chromium.connect(PLAYWRIGHT_WS)
browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15)
try:
# Create browser context with user agent
@@ -1520,7 +1641,7 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
try:
# Navigate to webinar page
await page.goto(WEBINAR_URL, wait_until='domcontentloaded')
await asyncio.wait_for(page.goto(WEBINAR_URL, wait_until='domcontentloaded'), timeout=30)
# Wait for the table to load
await page.wait_for_selector('#meetings table', timeout=10000)
@@ -1564,16 +1685,25 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
if page and not content:
content = await page.content()
_metric_check_fail(time.time() - _t0)
return None
finally:
await page.close()
await context_browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(page.close(), timeout=5)
with contextlib.suppress(Exception):
await asyncio.wait_for(context_browser.close(), timeout=5)
finally:
await browser.close()
with contextlib.suppress(Exception):
await asyncio.wait_for(browser.close(), timeout=5)
except TimeoutError:
logger.error('Webinar check timed out after 90s (playwright hang)')
_metric_check_fail(time.time() - _t0)
return None
except Exception as e:
logger.error(f'Playwright error: {e}')
_metric_check_fail(time.time() - _t0)
return None
# --- DEBUG LOGGING (Saving last response content) ---
@@ -1637,6 +1767,7 @@ async def check_webinars_job(context: ContextTypes.DEFAULT_TYPE):
else:
logger.info(f'Found {len(current_webinars)} webinar(s), but all are already known')
_metric_check_ok(time.time() - _t0)
return len(current_webinars)
@@ -1680,6 +1811,8 @@ def main():
job_queue = app.job_queue
job_queue.run_repeating(check_webinars_job, interval=WEBINAR_CHECK_INTERVAL, first=10)
start_metrics_server()
logger.info('Bot started polling...')
app.run_polling()
+29
View File
@@ -0,0 +1,29 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: gitea-prod-tls
namespace: gitea
spec:
secretName: gitea-prod-tls
dnsNames:
- gcr.forust.xyz
- gitea.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: gitea
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+1 -1
View File
@@ -31,7 +31,7 @@ spec:
spec:
containers:
- name: gitea
image: docker.gitea.com/gitea:1.27.3
image: gitea/gitea:1.27.3
envFrom:
- configMapRef:
name: gitea-config
+4 -1
View File
@@ -24,7 +24,7 @@ spec:
- name: gitea-service
port: 3000
tls:
certResolver: letsencrypt
secretName: gitea-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -45,6 +45,9 @@ spec:
services:
- name: gitea-service
port: 3000
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRouteTCP
+29
View File
@@ -0,0 +1,29 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: glance-prod-tls
namespace: glance
spec:
secretName: glance-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: glance
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+4 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: glance-service
port: 8080
tls:
certResolver: letsencrypt
secretName: glance-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -35,6 +35,9 @@ spec:
services:
- name: glance-service
port: 8080
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
+41
View File
@@ -0,0 +1,41 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: headscale-prod-tls
namespace: headscale
spec:
secretName: headscale-prod-tls
dnsNames:
- hs.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: headplane-prod-tls
namespace: headscale
spec:
secretName: headplane-prod-tls
dnsNames:
- hp.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: headscale
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+7 -2
View File
@@ -38,7 +38,7 @@ spec:
- name: headscale-server-external
port: 9090
tls:
certResolver: letsencrypt
secretName: headscale-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -64,7 +64,7 @@ spec:
- name: headplane-external
port: 3000
tls:
certResolver: letsencrypt
secretName: headplane-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -90,6 +90,9 @@ spec:
services:
- name: headscale-server-external
port: 9090
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -110,3 +113,5 @@ spec:
services:
- name: headplane-external
port: 3000
tls:
secretName: internal-wildcard-tls
+42
View File
@@ -0,0 +1,42 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: forust-homepage-prod-tls
namespace: homepages
spec:
secretName: forust-homepage-prod-tls
dnsNames:
- forust.xyz
- www.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: xdfnx-homepage-prod-tls
namespace: homepages
spec:
secretName: xdfnx-homepage-prod-tls
dnsNames:
- xdfnx.cfd
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: homepages
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+7 -2
View File
@@ -17,7 +17,7 @@ spec:
- name: forust-homepage-service
port: 80
tls:
certResolver: letsencrypt
secretName: forust-homepage-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -34,6 +34,9 @@ spec:
services:
- name: forust-homepage-service
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -53,7 +56,7 @@ spec:
- name: xdfnx-homepage-service
port: 80
tls:
certResolver: letsencrypt
secretName: xdfnx-homepage-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -69,3 +72,5 @@ spec:
services:
- name: xdfnx-homepage-service
port: 80
tls:
secretName: internal-wildcard-tls
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: kener-service
port: 3000
tls:
certResolver: letsencrypt
secretName: kener-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: kener-service
port: 3000
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: metube-prod-tls
namespace: metube
spec:
secretName: metube-prod-tls
dnsNames:
- metube.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: metube
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -15,7 +15,7 @@ spec:
- name: metube-service
port: 8081
tls:
certResolver: letsencrypt
secretName: metube-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -31,3 +31,5 @@ spec:
services:
- name: metube-service
port: 8081
tls:
secretName: internal-wildcard-tls
+2
View File
@@ -32,3 +32,5 @@ spec:
services:
- name: n8n-service
port: 5678
tls:
secretName: internal-wildcard-tls
+15
View File
@@ -0,0 +1,15 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: n8n
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: netronome-prod-tls
namespace: netronome
spec:
secretName: netronome-prod-tls
dnsNames:
- nm.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: netronome
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: netronome-service
port: 7575
tls:
certResolver: letsencrypt
secretName: netronome-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: netronome-service
port: 7575
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: nextcloud-prod-tls
namespace: nextcloud
spec:
secretName: nextcloud-prod-tls
dnsNames:
- nextcloud.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: nextcloud
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+26 -21
View File
@@ -18,7 +18,7 @@ spec:
- name: nextcloud-apache
port: 11000
tls:
certResolver: letsencrypt
secretName: nextcloud-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -38,26 +38,29 @@ spec:
services:
- name: nextcloud-apache
port: 11000
# ---
# Nextcloud AIO
# apiVersion: traefik.io/v1alpha1
# kind: IngressRoute
# metadata:
# name: naio-prod
# namespace: nextcloud
# spec:
# entryPoints:
# - websecure
# routes:
# - match: Host(`naio.forust.xyz`)
# kind: Rule
# services:
# - name: nextcloud-aio
# port: 8888
# scheme: https
# serversTransport: insecure-transport
# tls:
# certResolver: letsencrypt
# ---
# Nextcloud AIO
# apiVersion: traefik.io/v1alpha1
# kind: IngressRoute
# metadata:
# name: naio-prod
# namespace: nextcloud
# spec:
# entryPoints:
# - websecure
# routes:
# - match: Host(`naio.forust.xyz`)
# kind: Rule
# services:
# - name: nextcloud-aio
# port: 8888
# scheme: https
# serversTransport: insecure-transport
# tls:
# certResolver: letsencrypt
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -75,3 +78,5 @@ spec:
port: 8888
scheme: https
serversTransport: insecure-transport
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: portainer-prod-tls
namespace: portainer
spec:
secretName: portainer-prod-tls
dnsNames:
- portainer.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: portainer
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: portainer-service
port: 9000
tls:
certResolver: letsencrypt
secretName: portainer-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: portainer-service
port: 9000
tls:
secretName: internal-wildcard-tls
+1
View File
@@ -3,3 +3,4 @@ AUTHENTIK_DB_PASSWORD=
GITEA_DB_PASSWORD=
NETRONOME_DB_PASSWORD=
PENPOT_DB_PASSWORD=
STATUSPAGE_DB_PASSWORD=
+13 -14
View File
@@ -1,23 +1,22 @@
# Shared PostgreSQL
This directory contains a PostgreSQL 15 deployment draft for Authentik, Gitea,
Netronome, and Penpot. It creates one database and one login role per service;
it does not migrate existing data or change application connection settings.
This directory contains the shared PostgreSQL 17 deployment for Authentik,
Gitea, Netronome, and Statuspage. It creates one database and one login role
per service. Per-service standalone databases were removed after the
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
not part of the shared instance).
## Compatibility baseline
| Service | Current application | Current standalone PostgreSQL | Common PostgreSQL 15 |
| --------- | ------------------- | ----------------------------: | ------------------------------------------------------------------------------------ |
| Authentik | 2025.10.2 | 15 | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | 14 | Supported (Gitea requires 12+) |
| Netronome | 0.14.0 | 17 | Validate in staging; upstream's example uses 17 but no 17-only feature is documented |
| Penpot | 2.17.2 | 15 | Supported by the official deployment |
| Service | Current application | Shared PostgreSQL 17 |
| ---------- | ------------------- | -------------------------------------- |
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
| Statuspage | custom | Supported |
PostgreSQL 15 is the conservative common major. A major-version downgrade or
change must use a logical dump/restore; changing only the image tag while
keeping a data directory is not supported. Back up and migrate one application
at a time, starting with Netronome because its current standalone deployment
uses PostgreSQL 17.
A major-version change must use a logical dump/restore; changing only the
image tag while keeping a data directory is not supported.
For Compose, copy `.env.example` to `.env`, set all passwords, and start it with
`docker compose -f shared-compose.yaml up -d`. This file is intentionally not
+2 -1
View File
@@ -5,12 +5,12 @@ set -euo pipefail
: "${GITEA_DB_PASSWORD:?GITEA_DB_PASSWORD is required}"
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
create_role_and_database() {
local role="$1"
local database="$2"
local password="$3"
psql --username "$POSTGRES_USER" --dbname postgres \
-v role="$role" -v database="$database" -v password="$password" \
<<'SQL'
@@ -25,3 +25,4 @@ create_role_and_database authentik authentik "$AUTHENTIK_DB_PASSWORD"
create_role_and_database gitea gitea "$GITEA_DB_PASSWORD"
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
+12
View File
@@ -60,11 +60,23 @@ spec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 10
periodSeconds: 10
startupProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
failureThreshold: 30
periodSeconds: 10
livenessProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 30
periodSeconds: 20
resources:
requests:
memory: "512Mi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "2000m"
volumes:
- name: postgres-data
persistentVolumeClaim:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
postgres:
image: postgres:15.19-alpine
image: postgres:17.11-alpine
container_name: homelab-postgres
restart: unless-stopped
env_file:
@@ -0,0 +1,43 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cert-manager
namespace: prometheus
labels:
release: prometheus-stack
spec:
groups:
- name: cert-manager
rules:
- alert: CertManagerCertNotReady
expr: certmanager_certificate_ready_status{condition="False"} == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Certificate {{ $labels.exported_namespace }}/{{ $labels.name }} is not ready"
description: "Certificate has been failing for more than 10 minutes. Check Challenges/Orders in that namespace."
- alert: CertManagerCertExpirySoon
expr: certmanager_certificate_expiration_timestamp_seconds - time() < 86400 * 14
for: 1h
labels:
severity: warning
annotations:
summary: "Certificate {{ $labels.exported_namespace }}/{{ $labels.name }} expires in less than 14 days"
description: "Renewal should have happened automatically. Check cert-manager logs if it persists."
- alert: CertManagerCertExpiryCritical
expr: certmanager_certificate_expiration_timestamp_seconds - time() < 86400 * 7
for: 1h
labels:
severity: critical
annotations:
summary: "Certificate {{ $labels.exported_namespace }}/{{ $labels.name }} expires in less than 7 days"
description: "Manual intervention likely needed: cmctl status certificate {{ $labels.name }} -n {{ $labels.exported_namespace }}."
- alert: CertManagerHittingRateLimits
expr: sum by (exported_namespace) (rate(certmanager_http_acme_client_request_count{status=~"429.*"}[10m])) > 0
for: 15m
labels:
severity: warning
annotations:
summary: "cert-manager is hitting ACME rate limits"
description: "Back off manual issuance retries and check failed Orders/Challenges."
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: grafana-prod-tls
namespace: prometheus
spec:
secretName: grafana-prod-tls
dnsNames:
- grafana.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: prometheus
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -3
View File
@@ -16,9 +16,7 @@ spec:
- name: prometheus-stack-grafana
port: 80
tls:
certResolver: letsencrypt
domains:
- main: grafana.forust.xyz
secretName: grafana-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -34,3 +32,5 @@ spec:
services:
- name: prometheus-stack-grafana
port: 80
tls:
secretName: internal-wildcard-tls
+1 -1
View File
@@ -42,7 +42,7 @@ spec:
- alert: TraefikServiceHighLatency
expr: |
histogram_quantile(0.95,
sum(rate(traefik_service_request_duration_seconds_bucket[5m])) by (le, service)) > 2
sum(rate(traefik_service_request_duration_seconds_bucket{service!~"xui-xui-service-.*"}[5m])) by (le, service)) > 2
for: 5m
labels:
severity: warning
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: reloader
+30 -1
View File
@@ -1,14 +1,43 @@
{
"$schema": "https://docs.renovatebot.com/renovate-schema.json",
"extends": ["config:recommended"],
"enabledManagers": ["docker-compose", "kubernetes", "helm-values"],
"enabledManagers": ["dockerfile", "docker-compose", "kubernetes", "helm-values", "custom.regex"],
"helm-values": {
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
},
"kubernetes": {
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
},
"customManagers": [
{
"customType": "regex",
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
"fileMatch": ["^edu_master/k8s/playwright\\.yaml$", "^edu_master/compose\\.yaml$"],
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
"datasourceTemplate": "npm",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: PLAYWRIGHT_VERSION file",
"fileMatch": ["^edu_master/PLAYWRIGHT_VERSION$"],
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)$"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright"
}
],
"packageRules": [
{
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"groupName": "playwright singlesource",
"groupSlug": "playwright"
},
{
"description": "playwright must not automerge - version skew breaks WS handshake (checker.py:1523 vs playwright.yaml:20)",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"automerge": false
},
{
"description": "Keep private homelab images unchanged",
"matchDatasources": ["docker"],
+1 -1
View File
@@ -1,6 +1,6 @@
module.exports = {
platform: 'gitea',
endpoint: process.env.RENOVATE_ENDPOINT,
endpoint: process.env.RENOVATE_ENDPOINT || 'https://gitea.forust.xyz/api/v1',
enabledManagers: ['docker-compose', 'kubernetes', 'helm-values'],
'helm-values': {
managerFilePatterns: ['/k8s/.+values\\.ya?ml$/'],
+1 -1
View File
@@ -7,7 +7,7 @@ data:
config.js: |
module.exports = {
platform: 'gitea',
endpoint: process.env.RENOVATE_ENDPOINT,
endpoint: process.env.RENOVATE_ENDPOINT || 'https://gitea.forust.xyz/api/v1',
enabledManagers: ['docker-compose', 'kubernetes', 'helm-values'],
'helm-values': {
managerFilePatterns: ['/k8s/.+values\\.ya?ml$/'],
+1 -1
View File
@@ -16,7 +16,7 @@ spec:
restartPolicy: Never
containers:
- name: renovate
image: renovate/renovate:44.103.0
image: renovate/renovate:44.106.0
env:
- name: RENOVATE_PLATFORM
value: gitea
+29
View File
@@ -0,0 +1,29 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: searxng-prod-tls
namespace: searxng
spec:
secretName: searxng-prod-tls
dnsNames:
- s.forust.xyz
- search.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: searxng
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: searxng-service
port: 8080
tls:
certResolver: letsencrypt
secretName: searxng-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: searxng-service
port: 8080
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: termix-prod-tls
namespace: termix
spec:
secretName: termix-prod-tls
dnsNames:
- termix.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: termix
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: termix-service
port: 8080
tls:
certResolver: letsencrypt
secretName: termix-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: termix-service
port: 8080
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: dashboard-prod-tls
namespace: traefik
spec:
secretName: dashboard-prod-tls
dnsNames:
- traefik.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: traefik
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -17,7 +17,7 @@ spec:
- name: api@internal
kind: TraefikService
tls:
certResolver: letsencrypt
secretName: dashboard-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -33,3 +33,5 @@ spec:
services:
- name: api@internal
kind: TraefikService
tls:
secretName: internal-wildcard-tls
+2 -16
View File
@@ -31,7 +31,7 @@ deployment:
providers:
kubernetesIngress:
enabled: false
enabled: true
kubernetesCRD:
enabled: true
kubernetesGateway:
@@ -121,21 +121,7 @@ persistence:
size: 100Mi
path: /data
certificatesResolvers:
letsencrypt-staging:
acme:
email: bobrovod@national.shitposting.agency
storage: /data/letsencrypt/acme.json
caServer: https://acme-staging-v02.api.letsencrypt.org/directory
httpChallenge:
entryPoint: web
letsencrypt:
acme:
email: bobrovod@national.shitposting.agency
storage: /data/letsencrypt/acme.json
caServer: https://acme-v02.api.letsencrypt.org/directory
httpChallenge:
entryPoint: web
# No certificatesResolvers: public TLS comes from cert-manager.
volumes:
- name: traefik-dynamic
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: uptime-kuma-prod-tls
namespace: uptime-kuma
spec:
secretName: uptime-kuma-prod-tls
dnsNames:
- uptime.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: uptime-kuma
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -16,7 +16,7 @@ spec:
- name: uptime-kuma-service
port: 3001
tls:
certResolver: letsencrypt
secretName: uptime-kuma-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -32,3 +32,5 @@ spec:
services:
- name: uptime-kuma-service
port: 3001
tls:
secretName: internal-wildcard-tls
@@ -0,0 +1,15 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: userbot
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+1
View File
@@ -4,4 +4,5 @@ kind: Kustomization
resources:
- userbots.yaml
- panel.yaml
- internal-certificate.yaml
- common-config.yaml
+2
View File
@@ -262,3 +262,5 @@ spec:
services:
- name: userbot-panel
port: 8080
tls:
secretName: internal-wildcard-tls
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: vaultwarden-prod-tls
namespace: vaultwarden
spec:
secretName: vaultwarden-prod-tls
dnsNames:
- vw.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: vaultwarden
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+3 -1
View File
@@ -13,7 +13,7 @@ spec:
- name: vaultwarden-service
port: 80
tls:
certResolver: letsencrypt
secretName: vaultwarden-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
@@ -29,3 +29,5 @@ spec:
services:
- name: vaultwarden-service
port: 80
tls:
secretName: internal-wildcard-tls
View File
Whitespace-only changes.
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: xray-prod-tls
namespace: xui
spec:
secretName: xray-prod-tls
dnsNames:
- xray.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: xui
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+11
View File
@@ -0,0 +1,11 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: xui-config
namespace: xui
data:
XUI_DB_FOLDER: "/etc/x-ui"
XUI_ENABLE_FAIL2BAN: "false"
XUI_INIT_WEB_BASE_PATH: "/"
XUI_LOG_LEVEL: "warning"
XUI_PORT: "30379"
+86
View File
@@ -0,0 +1,86 @@
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: xui-local
namespace: xui
spec:
entryPoints:
- websecure
routes:
- match: Host(`xui.workstation.internal`) || Host(`xui.gigaforust.internal`)
kind: Rule
services:
- name: xui-service
port: 30379
tls:
secretName: internal-wildcard-tls
---
# Public panel access (optional).
# Realistic, but intentionally disabled: the panel has its own login,
# security-chain adds Authentik in front of it.
# To enable: uncomment and add Public Hostname `xui.forust.xyz`
# in the Cloudflare tunnel (same as other *.forust.xyz hosts).
# ---
# apiVersion: traefik.io/v1alpha1
# kind: IngressRoute
# metadata:
# name: xui-prod
# namespace: xui
# spec:
# entryPoints:
# - websecure
# routes:
# - match: Host(`xui.forust.xyz`)
# kind: Rule
# middlewares:
# - name: security-chain@file
# services:
# - name: xui-service
# port: 30379
# tls:
# certResolver: letsencrypt
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: xray-prod
namespace: xui
spec:
entryPoints:
- websecure
routes:
- match: Host(`xray.forust.xyz`) && PathPrefix(`/pzzfpz6oi281f0u8`)
kind: Rule
services:
- name: xui-service
port: 2096
- match: Host(`xray.forust.xyz`)
kind: Rule
services:
- name: xui-service
port: 10000
tls:
secretName: xray-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: xray-local
namespace: xui
spec:
entryPoints:
- websecure
routes:
- match: (Host(`xray.workstation.internal`) || Host(`xray.gigaforust.internal`)) && PathPrefix(`/pzzfpz6oi281f0u8`)
kind: Rule
services:
- name: xui-service
port: 2096
- match: Host(`xray.workstation.internal`) || Host(`xray.gigaforust.internal`)
kind: Rule
services:
- name: xui-service
port: 10000
tls:
secretName: internal-wildcard-tls
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: xui
+74
View File
@@ -0,0 +1,74 @@
apiVersion: v1
kind: Service
metadata:
name: xui-service
namespace: xui
spec:
selector:
app: xui
ports:
- port: 30379
name: panel
targetPort: 30379
- port: 10000
name: xray
targetPort: 10000
- port: 2096
name: sub
targetPort: 2096
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: xui-deployment
namespace: xui
spec:
replicas: 1
selector:
matchLabels:
app: xui
template:
metadata:
labels:
app: xui
spec:
containers:
- name: xui
image: ghcr.io/mhsanaei/3x-ui:v3.8.5
envFrom:
- configMapRef:
name: xui-config
tty: true
ports:
- containerPort: 30379
name: panel
- containerPort: 10000
name: xray
- containerPort: 2096
name: sub
volumeMounts:
- name: x-ui-db
mountPath: /etc/x-ui
resources:
requests:
memory: "128Mi"
cpu: "100m"
limits:
memory: "1Gi"
cpu: "1000m"
volumes:
- name: x-ui-db
persistentVolumeClaim:
claimName: xui-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: xui-pvc
namespace: xui
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi