Eighteen containers had no memory limit at all, so nothing on the node could bound them. Three of the values files even claimed to set resources: Helm does not complain about a key it does not recognise, so the block sat there looking like a limit while the pod ran unbounded. alloy is the one that mattered. The chart reads `alloy.resources`; the file had `controller.resources`, so the DaemonSet that tails every pod log on the node shipped with nothing at all. `kubeStateMetrics` is the same trap in a different shape -- that is the condition key, the values live under `kube-state-metrics` -- and `configReloader` in the alloy chart sits at the top level rather than under `alloy`. Each one is verified by rendering the chart and reading the resources back off the containers, because a values key that is ignored looks exactly like one that works. reloader turned out to be set and still wrong: 64Mi request against a measured p95 of 73M, so the pod ran permanently above its own request and stayed a standing eviction candidate. That is the pod that restarts every other pod, so it is the last one that should be evicted. Raised to 96Mi. Requests are set at p95 throughout, grafana, playwright and alloy included. Left at the values first proposed they would have sat below their own p95 and queued for eviction ahead of everything smaller. CPU limits are deliberately absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk wait into runnable-throttled, which is the failure mode that took the node down. The prometheus and alertmanager configReloader sidecars are left open: chart 86.2.3 does not template the key, so reaching those two containers needs a postRenderer. Verified: all four charts render with the resources landing on the intended containers, and 16/16 local gates pass.
72 lines
1.9 KiB
YAML
72 lines
1.9 KiB
YAML
# Pinned chart: grafana/alloy 1.12.1 (app v1.19.2).
|
|
# Install (deferred to deploy task, namespace prometheus):
|
|
# helm upgrade --install alloy grafana/alloy --version 1.12.1 \
|
|
# --namespace prometheus --values loki/k8s/alloy-values.yaml --wait
|
|
# DaemonSet ships k8s pod logs (API-tailed, no hostPath mounts) to Loki.
|
|
# Scope phase 1: k8s only, compose leftovers out.
|
|
|
|
controller:
|
|
type: daemonset
|
|
|
|
# config-reloader sidecar: p95 33M, max 43M. The chart keeps it at the top level,
|
|
# not under `alloy:`.
|
|
configReloader:
|
|
resources:
|
|
requests:
|
|
memory: "32Mi"
|
|
cpu: "10m"
|
|
limits:
|
|
memory: "128Mi"
|
|
|
|
image:
|
|
tag: "v1.19.2"
|
|
|
|
alloy:
|
|
# p95 275M, max 287M. Alloy tails every pod log and ships it to Loki, so it sits
|
|
# on the same IronWolf read path the node is I/O bound on. Request is set at p95.
|
|
# The chart key is `alloy.resources`. `controller.resources` is ignored silently,
|
|
# which is why this pod shipped with no limits at all.
|
|
resources:
|
|
requests:
|
|
memory: "288Mi"
|
|
cpu: "50m"
|
|
limits:
|
|
memory: "512Mi"
|
|
|
|
configMap:
|
|
create: true
|
|
content: |
|
|
discovery.kubernetes "pods" {
|
|
role = "pod"
|
|
}
|
|
|
|
discovery.relabel "pods" {
|
|
targets = discovery.kubernetes.pods.targets
|
|
|
|
rule {
|
|
source_labels = ["__meta_kubernetes_namespace"]
|
|
target_label = "namespace"
|
|
}
|
|
|
|
rule {
|
|
source_labels = ["__meta_kubernetes_pod_name"]
|
|
target_label = "pod"
|
|
}
|
|
|
|
rule {
|
|
source_labels = ["__meta_kubernetes_pod_container_name"]
|
|
target_label = "container"
|
|
}
|
|
}
|
|
|
|
loki.source.kubernetes "pods" {
|
|
targets = discovery.relabel.pods.output
|
|
forward_to = [loki.write.default.receiver]
|
|
}
|
|
|
|
loki.write "default" {
|
|
endpoint {
|
|
url = "http://loki-gateway.prometheus.svc.cluster.local/loki/api/v1/push"
|
|
}
|
|
}
|