26 Commits
Author SHA1 Message Date
forust 64ba4ce9f1 Revert "chore(prometheus): drop dead grafana routes and cert"
This reverts commit 314ec0cda7.
2026-10-03 10:40:39 +02:00
forust 0859479c0f feat(ingress): replace traefik crowdsec plugin with firewall bouncer
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts.
2026-09-30 20:14:25 +02:00
forust a2af854867 fix(alerts): measure traefik latency per route via exported_service
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 6s
The service label the rule grouped by is the scrape target name, not the backend: with honorLabels=false the per-route value traefik emits is renamed to exported_service, so the old rule measured one global aggregate and both exclusions matched nothing. Group the latency and 5xx alerts by exported_service with exclusions in traefik's real namespace-routename-hash format. Already applied live with kubectl; this commit keeps git in sync.
2026-09-29 09:55:02 +02:00
forust 314ec0cda7 chore(prometheus): drop dead grafana routes and cert
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 6s
No grafana is deployed from this stack, so grafana-prod/grafana-local only ever served 503s. Live objects already deleted directly; this keeps git from resurrecting them on the next deploy.
2026-09-29 09:33:38 +02:00
forust a564dd8b67 fix(alerts): exclude netbird streams from traefik latency alert
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 12s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 8s
ci / build (push) Successful in 10s
SignalExchange ConnectStream holds 60s gRPC streams by design; at night they exceed 5% of samples and pin P95 to the 5.0s bucket ceiling, flapping the alert. Also fixes the stale xui exclusion pattern, which matched no real service label.
2026-09-29 08:12:07 +02:00
forust 74adf38d63 feat(alerts): cover OOM kills, restart loops and evictions
The OOMKilled container behind the immich crash loop was invisible: PodCrashLooping only fires once kubelet has already given up and started the backoff.
2026-09-28 17:01:36 +02:00
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00
forust 16aaeb60c1 fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.

alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.

reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.

Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.

The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.

Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
2026-09-28 10:14:52 +02:00
forust 91c344fe2c fix(monitoring): disable control-plane alerts and scrapes on k0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 24s
deploy / validate (push) Successful in 1m42s
deploy / apply-k8s (push) Successful in 1m45s
ci / build (push) Successful in 2s
deploy / apply-compose (push) Failing after 9s
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own
process rather than as pods, so the chart's Services never get endpoints
and the targets stay permanently absent. Drop the matching ServiceMonitors
and their Down/HighCommitDurations rules; kube-proxy and kubelet do get
endpoints on k0s and stay enabled.
2026-09-26 17:16:55 +02:00
forust bc1e69ebe0 feat(tls): internal CA wildcard for *.internal routes
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
Selfsigned root (10y) + internal-ca issuer; per-namespace
internal-wildcard-tls certs referenced by all -local routers.
Root public cert committed for client trust stores.
2026-09-23 14:45:05 +02:00
forust f29bb3d580 revert: drop Grafana cert-manager dashboard 2026-09-23 14:42:09 +02:00
forust 1cbdfc1d6f feat(observability): cert-manager metrics, dashboard and alerts
- ServiceMonitor for cert-manager (release: prometheus-stack)
- Grafana dashboard Cert-manager-Kubernetes (ID 22908, datasource
  refs fixed) via dashboard sidecar ConfigMap
- PrometheusRule: NotReady (crit), expiry <14d (warn) / <7d
  (crit), ACME 429 rate-limit (warn)
2026-09-23 14:32:36 +02:00
forust ef325cd3b1 feat(tls): migrate public ingress TLS from Traefik ACME to cert-manager
ci / lint-ruff (push) Successful in 1s
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / deploy-userbot-panel (push) Skipped
All prod IngressRoutes switch tls.certResolver to tls.secretName
backed by per-router Certificates (HTTP-01, letsencrypt-prod).
adguard-prod reuses the shared adguard-certs secret (also feeds
DoT :853); sync CronJob removed as redundant.

Traefik certificatesResolvers removed: its internal
acme-http@internal router hijacks HTTP-01 for every host while
enabled, blocking external solvers. Dormant files (kener,
downtify) converted for consistency but not applied; n8n
untouched per live-only rule.
2026-09-23 14:12:36 +02:00
forust 9c4580a522 fix(prometheus): exclude xui services from TraefikServiceHighLatency
ci / lint-prettier (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 9s
ci / build (push) Successful in 1s
ci / deploy-userbot-panel (push) Has been skipped
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
Long-lived VPN WebSocket sessions inflate P95 request duration; keep 5xx/down/cert alerts for xui unchanged.
2026-09-19 19:06:42 +02:00
forust 5936de3e56 fix(alertmanager): quote null receiver and wire telegram template
Unquoted null parses as YAML null and breaks routing. Add shared telegram message template.
2026-09-18 19:02:41 +02:00
forust c98ea8b957 feat(monitoring): align storage to live pvc, add loki datasource and alerts
Match grafana/alertmanager storageClass to live PVCs to avoid immutable-field failures. Add Loki datasource and LokiDown/AlloyDown alerts.
2026-09-18 19:02:38 +02:00
forustandCopilot 24d3686f60 ci: fix formatting and YAML lint scope
ci / lint-prettier (push) Successful in 13s
ci / lint-ruff (push) Successful in 4s
ci / lint-yaml (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
renovate-ci / validate-renovate (push) Failing after 2s
ci / build (push) Successful in 2s
ci / deploy-userbot-panel (push) Has been skipped
Lint tracked YAML files without scanning generated dependencies.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-14 09:30:55 +02:00
forust 360a6fc5dc feat(prometheus): add alerting rules and alertmanager config 2026-09-14 01:21:27 +02:00
forustandCopilot ed1ddaad5d feat(crowdsec): restore web traffic protection
Protect public Traefik routes with CrowdSec HTTP decisions and restore access logging for web traffic analysis.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-12 21:05:12 +02:00
forust 861d89d36a ci(deploy): split runtime by k8s/active marker
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.

validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.
2026-09-06 20:39:12 +02:00
forust 92aa731e44 deleted crowdsec stack from the repo. will figure something else
ci / lint-prettier (push) Successful in 18s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 6s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
ci / build (push) Successful in 43s
Signed-off-by: mr-forust <vzlomdsisma@gmail.com>
2026-07-19 23:16:14 +02:00
forust a128523c24 Fix CrowdSec middleware references in Traefik ingresses 2026-06-29 22:47:21 +02:00
forust 1762962f32 chore(k8s): fix indentation in ingress manifests (3-space -> 2-space) 2026-06-19 12:03:02 +02:00
forust 0803f3efff chore(k8s): returned to Host || Host standart instead of regexp.
Deploy to Server / deploy (push) Has been cancelled
Yaml lint (yamllint)
2026-06-18 21:02:40 +02:00
forust 76853637bc feat: traefik prometheus metrics 2026-06-17 01:41:45 +02:00
forust ea483da645 feat: kube-prometeus-stack 2026-06-16 12:36:31 +02:00