Files
homelab/docs/repository-review.md
T
forust d85bdf5dbd
ci / Compose (pull_request) Successful in 27s
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 34s
ci / Python and tests (pull_request) Successful in 19s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 36s
ci / Kubernetes (pull_request) Successful in 14s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s
docs: sync service guides with current main
2026-10-08 21:23:46 +02:00

11 KiB

Repository review (6 October 2026 baseline)

This records the tracked tree at cc9c3de and the workstation state observed on 6 October 2026. It is a historical review, not a current runtime inventory. The listed code fixes have since merged into main; EDU ownership has moved to the separate repository described in the handoff record. See the CI and deployment guide and runner and recovery guide for the current workflow. No deployment was performed during the original review.

Findings at the baseline and current status

Priority Finding at the baseline Current status
High Per-file APPLY_PRUNE=true could delete resources selected by a shared label. The deploy workflow rejects unsafe pruning before applying resources.
High Compose validation did not resolve the local configuration required at deploy time. Preflight resolves the selected Compose configuration before apply.
Medium Secret validation could miss namespace-specific and mounted Secret references. Preflight checks rendered references in their namespaces, including mounted and projected Secrets.
Medium Compose CI missed manual entry points such as shared-compose.yaml and client.compose.yaml. CI checks all tracked Compose files.
Medium NetBird Compose referenced missing setup and renderer files. The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment.
Medium Glance mounted its CSS from the wrong ConfigMap. The mount now uses the ConfigMap that contains user.css.
Medium The PostgreSQL env example omitted the required NetBox password. The example now includes the required variable.
Medium The former EDU code had stale Compose variable names and session reliability problems. EDU workloads and their fixes moved out of this repository; see the handoff record.
Medium AdGuard DoH and SearXNG Compose router expressions used invalid Host(...) syntax. The router expressions now follow Traefik's rule syntax.

Traefik matchers should be combined as Host(a) || Host(b); the rule syntax is described in the Traefik rules documentation. The fix retains the DoH path constraint for both hostnames.

The current deploy workflow deliberately rejects the unsafe prune option. It does not introduce automatic deletion under a different implementation. The baseline finding was a configuration risk, not evidence of a live deletion incident.

The former session fix bounded HTTP and Redis calls, validated credentials, set a cookie lifetime of two refresh intervals, and marked success only after publishing the verified cookie. The service is now owned by the EDU repository; see that repository for its current implementation.

The deployment fix extracts required pod Secret references from rendered JSON, checks their namespaces, includes init containers, image-pull credentials, and mounted/projected Secrets, and honors optional references. Ingress TLS Secrets issued by cert-manager are not treated as pre-existing pod prerequisites. It checks existence/access, not every key's contents or application validity.

Live workstation observations

The SSH alias workstation is reachable. It has one Ready control-plane node, Kubernetes v1.35.4+k0s, and a Docker daemon alongside containerd. At inspection, no pods were Pending or in another non-running, non-completed phase. This is a point-in-time observation, not a complete application health test.

The deployment checkout at /srv/homelab is on main commit 2adf17c, behind the reviewed local commit. It has untracked host configuration and a separate userbot/ directory. It was not reset or cleaned.

Observed difference Implication
VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings.
Homarr, Cloudflared, and Reloader are installed without their current Git active markers. Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped.
Cloudflare DDNS is running in both Docker and Kubernetes. Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected.
Traefik's LoadBalancer exposes port 8080 at 192.168.80.2. The direct API listener is deployed; its external reachability was not tested.
Default local-path has reclaim policy Delete, while many existing PVs have been changed to Retain. Current retention is partly live state. Recreating a claim can get a different policy from the old PV.
NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup.

The VictoriaMetrics monitoring trial later merged into main in PR #95. The first row above records the state before that change. Read prometheus-stack/README.md for the current tracked monitoring configuration; the live observations in this section remain a snapshot from 6 October.

Current recovery limits

The deployment controller and its recovery process changed after this review. The current operator workflow is documented in the runner and recovery guide. The remaining boundaries are:

  • Kubernetes recovery can restore captured workload revisions. It does not restore ConfigMaps, Secrets, database schemas, or persistent data.
  • Compose recovery is manual. It uses saved resolved configuration, but it does not restore volume data or reverse database migrations.
  • Removed resources require manual review and removal; the deploy workflow does not prune them automatically.
  • Plan mode does not create namespaces. During apply, server validation for new namespaces runs after namespace creation and chart installation; a failed check can leave an empty namespace.
  • Storage policy and backup coverage remain service-specific. Check the live PV, PVC, and backup state before changing stateful workloads.

Validation

At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose files, and Kubernetes resources with available schemas. Kubeconform found 347 resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources. That skip count matters: passing schema validation does not validate Traefik rule strings or other controller-specific behavior.

Fix validation covers:

  • Compose discovery of manual entry points, rejection of required-variable gaps, namespace-scoped and optional Secret references, and API/render failures.
  • NetBird setup idempotence, preservation of existing keys, file permissions, runtime rendering, and rejection of invalid trusted proxy CIDRs.
  • Session refresh success and failure paths, timeouts, cookie expiry, log redaction, missing credentials, and nonpositive refresh intervals.
  • Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
  • YAML and Compose structure for the corrected router rules, compared with the documented Traefik grammar. They were not exercised on the live proxy.
  • Prune rejection before any cluster invocation.

At the time of review, all seven fix branches and the documentation branch merged together in a disposable validation worktree. That combined tree passed the CI-equivalent local checks, Markdown formatting/lint and link checks, all 35 Compose structure checks, and 11 Python regression tests plus the shell validation regressions. CRD server-side validation and live rollout tests were not run.

Runtime tests use fixtures and mocks, not production credentials. Live checks read workload metadata, storage policies, chart versions, and container state only. They did not read Secret contents or change services.

Reloader follow-up (baseline)

fix/reloader-integration added the active marker and opt-in annotations to application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets. It corrected AdGuard's misplaced pod-template annotation. The Helm settings use annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and CronJobs. PostgreSQL workloads are excluded because their credential variables and init scripts are only effective on an empty data directory.

The controller was running on the workstation when inspected. The original review checked configuration against the pinned chart with Helm rendering and manifest validation; it did not change production configuration to provoke a test restart or confirm every application's live reload behavior.