Files
homelab/prometheus-stack/README.md
T
forust d85bdf5dbd
ci / Compose (pull_request) Successful in 27s
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 34s
ci / Python and tests (pull_request) Successful in 19s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 36s
ci / Kubernetes (pull_request) Successful in 14s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s
docs: sync service guides with current main
2026-10-08 21:23:46 +02:00

1.7 KiB

Monitoring stack

The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent, and vmalert. The k8s/active marker selects the stack. The kube-prometheus-stack Helm release installs Grafana, Alertmanager, the Prometheus Operator, and related components. Its Prometheus server is configured with zero replicas while VMAgent collects metrics and writes them to the single-node VictoriaMetrics instance.

The victoria-operator Helm release converts selected Prometheus Operator ServiceMonitor resources into VMServiceScrape resources. VMAgent selects those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates the rule ConfigMap and sends alerts to the stack's Alertmanager. See the Kubernetes monitoring notes for application metrics and validation commands.

The chart versions are pinned in .gitea/workflows/deploy-lib.sh. The tracked k8s/grafana-values.yaml contains the Helm values for the stack. Create the grafana-admin and alertmanager-config Secrets from the examples in k8s/; keep their credentials out of the values file. Persistent volumes store data for Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and backups before changing storage. VictoriaMetrics currently retains 30 days of data.

A separate Compose configuration is present for manual use. There is no root active marker, so the automatic deploy workflow does not select it.

The deploy workflow does not remove resources when manifests are deleted. For a rollback of application-metrics changes, follow the explicit cleanup steps in the Kubernetes monitoring notes.

See the repository README for deployment selection.