Files
homelab/prometheus-stack/README.md
T
forust d85bdf5dbd
ci / Compose (pull_request) Successful in 27s
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 34s
ci / Python and tests (pull_request) Successful in 19s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 36s
ci / Kubernetes (pull_request) Successful in 14s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s
docs: sync service guides with current main
2026-10-08 21:23:46 +02:00

33 lines
1.7 KiB
Markdown

# Monitoring stack
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
and vmalert. The `k8s/active` marker selects the stack. The
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
Prometheus Operator, and related components. Its Prometheus server is configured
with zero replicas while VMAgent collects metrics and writes them to the
single-node VictoriaMetrics instance.
The `victoria-operator` Helm release converts selected Prometheus Operator
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
validation commands.
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
keep their credentials out of the values file. Persistent volumes store data for
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
backups before changing storage. VictoriaMetrics currently retains 30 days of
data.
A separate Compose configuration is present for manual use. There is no root
`active` marker, so the automatic deploy workflow does not select it.
The deploy workflow does not remove resources when manifests are deleted. For a
rollback of application-metrics changes, follow the explicit cleanup steps in
the [Kubernetes monitoring notes](k8s/README.md).
See the [repository README](../README.md) for deployment selection.