Files
forust a2af854867
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 6s
fix(alerts): measure traefik latency per route via exported_service
The service label the rule grouped by is the scrape target name, not the backend: with honorLabels=false the per-route value traefik emits is renamed to exported_service, so the old rule measured one global aggregate and both exclusions matched nothing. Group the latency and 5xx alerts by exported_service with exclusions in traefik's real namespace-routename-hash format. Already applied live with kubectl; this commit keeps git in sync.
2026-09-29 09:55:02 +02:00

69 lines
3.1 KiB
YAML

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: traefik
namespace: prometheus
labels:
release: prometheus-stack
spec:
groups:
- name: traefik
rules:
- alert: TraefikDown
expr: absent(up{job="traefik"})
for: 10m
labels:
severity: critical
annotations:
summary: "Traefik metrics are unavailable"
description: "Prometheus has no Traefik target."
- alert: TraefikConfigReloadFailed
expr: traefik_config_last_reload_success == 0
for: 5m
labels:
severity: warning
annotations:
summary: "Traefik configuration reload failed"
description: "Traefik failed to apply its last configuration reload. Check Traefik logs and the dynamic config sources."
- alert: TraefikServiceHigh5xxRate
expr: |
sum(rate(traefik_service_requests_total{code=~"5.."}[5m])) by (exported_service)
/ sum(rate(traefik_service_requests_total[5m])) by (exported_service) * 100 > 5
and sum(rate(traefik_service_requests_total[5m])) by (exported_service) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "High 5xx error rate for service {{ $labels.exported_service }}"
description: "Service {{ $labels.exported_service }} is returning 5xx errors for more than 5% of requests over the last 5 minutes (current: {{ $value | humanizePercentage }})."
# NOTE on `exported_service`: traefik emits per-route series labelled
# `service`, but the prometheus scrape already carries a target label
# called `service` (the monitored Service object), and with the
# default honorLabels=false the scraped value loses the collision and
# is renamed to `exported_service`. Grouping by `service` therefore
# measures one global aggregate mislabelled per backend, and any
# exclusion written against it matches nothing. The generated names
# below are namespace-routename-hash as traefik builds them.
- alert: TraefikServiceHighLatency
expr: |
histogram_quantile(0.95,
sum(rate(traefik_service_request_duration_seconds_bucket{exported_service!~"netbird-netbird-(prod|local)-.*|xui-xray-(prod|local)-.*"}[5m])) by (le, exported_service)) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "High latency for service {{ $labels.exported_service }}"
description: "P95 latency of {{ $labels.exported_service }} exceeded 2 seconds over the last 5 minutes (current: {{ $value | humanizeDuration }})."
- alert: TraefikCertExpiringSoon
expr: min(traefik_tls_certs_not_after) - time() < 14 * 24 * 60 * 60
for: 10m
labels:
severity: warning
annotations:
summary: "Traefik TLS certificate expires soon"
description: "A TLS certificate managed by Traefik expires in less than 14 days (in {{ $value | humanizeDuration }})."