fix(alerts): measure traefik latency per route via exported_service
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 6s

The service label the rule grouped by is the scrape target name, not the backend: with honorLabels=false the per-route value traefik emits is renamed to exported_service, so the old rule measured one global aggregate and both exclusions matched nothing. Group the latency and 5xx alerts by exported_service with exclusions in traefik's real namespace-routename-hash format. Already applied live with kubectl; this commit keeps git in sync.
This commit is contained in:
forust committed 2026-09-29 09:55:02 +02:00
1 parent 314ec0cda7
commit a2af854867
1 file changed
+16 -8
+16 -8
View File
@@ -29,26 +29,34 @@ spec:
- alert: TraefikServiceHigh5xxRate - alert: TraefikServiceHigh5xxRate
expr: | expr: |
sum(rate(traefik_service_requests_total{code=~"5.."}[5m])) by (service) sum(rate(traefik_service_requests_total{code=~"5.."}[5m])) by (exported_service)
/ sum(rate(traefik_service_requests_total[5m])) by (service) * 100 > 5 / sum(rate(traefik_service_requests_total[5m])) by (exported_service) * 100 > 5
and sum(rate(traefik_service_requests_total[5m])) by (service) > 0 and sum(rate(traefik_service_requests_total[5m])) by (exported_service) > 0
for: 5m for: 5m
labels: labels:
severity: warning severity: warning
annotations: annotations:
summary: "High 5xx error rate for service {{ $labels.service }}" summary: "High 5xx error rate for service {{ $labels.exported_service }}"
description: "Service {{ $labels.service }} is returning 5xx errors for more than 5% of requests over the last 5 minutes (current: {{ $value | humanizePercentage }})." description: "Service {{ $labels.exported_service }} is returning 5xx errors for more than 5% of requests over the last 5 minutes (current: {{ $value | humanizePercentage }})."
# NOTE on `exported_service`: traefik emits per-route series labelled
# `service`, but the prometheus scrape already carries a target label
# called `service` (the monitored Service object), and with the
# default honorLabels=false the scraped value loses the collision and
# is renamed to `exported_service`. Grouping by `service` therefore
# measures one global aggregate mislabelled per backend, and any
# exclusion written against it matches nothing. The generated names
# below are namespace-routename-hash as traefik builds them.
- alert: TraefikServiceHighLatency - alert: TraefikServiceHighLatency
expr: | expr: |
histogram_quantile(0.95, histogram_quantile(0.95,
sum(rate(traefik_service_request_duration_seconds_bucket{service!~"xui-service|netbird-server-service"}[5m])) by (le, service)) > 2 sum(rate(traefik_service_request_duration_seconds_bucket{exported_service!~"netbird-netbird-(prod|local)-.*|xui-xray-(prod|local)-.*"}[5m])) by (le, exported_service)) > 2
for: 5m for: 5m
labels: labels:
severity: warning severity: warning
annotations: annotations:
summary: "High latency for service {{ $labels.service }}" summary: "High latency for service {{ $labels.exported_service }}"
description: "P95 latency of {{ $labels.service }} exceeded 2 seconds over the last 5 minutes (current: {{ $value | humanizeDuration }})." description: "P95 latency of {{ $labels.exported_service }} exceeded 2 seconds over the last 5 minutes (current: {{ $value | humanizeDuration }})."
- alert: TraefikCertExpiringSoon - alert: TraefikCertExpiringSoon
expr: min(traefik_tls_certs_not_after) - time() < 14 * 24 * 60 * 60 expr: min(traefik_tls_certs_not_after) - time() < 14 * 24 * 60 * 60