ci / Compose (pull_request) Successful in 27s
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 34s
ci / Python and tests (pull_request) Successful in 19s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 36s
ci / Kubernetes (pull_request) Successful in 14s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s
33 lines
1.7 KiB
Markdown
33 lines
1.7 KiB
Markdown
# Monitoring stack
|
|
|
|
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
|
|
and vmalert. The `k8s/active` marker selects the stack. The
|
|
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
|
|
Prometheus Operator, and related components. Its Prometheus server is configured
|
|
with zero replicas while VMAgent collects metrics and writes them to the
|
|
single-node VictoriaMetrics instance.
|
|
|
|
The `victoria-operator` Helm release converts selected Prometheus Operator
|
|
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
|
|
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
|
|
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
|
|
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
|
|
validation commands.
|
|
|
|
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
|
|
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
|
|
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
|
|
keep their credentials out of the values file. Persistent volumes store data for
|
|
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
|
|
backups before changing storage. VictoriaMetrics currently retains 30 days of
|
|
data.
|
|
|
|
A separate Compose configuration is present for manual use. There is no root
|
|
`active` marker, so the automatic deploy workflow does not select it.
|
|
|
|
The deploy workflow does not remove resources when manifests are deleted. For a
|
|
rollback of application-metrics changes, follow the explicit cleanup steps in
|
|
the [Kubernetes monitoring notes](k8s/README.md).
|
|
|
|
See the [repository README](../README.md) for deployment selection.
|