Repository navigation
fix: cap DNS domain labels per container, add opt-in --min-container-age (B2 of #369) - #371
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces mechanisms to reduce metric cardinality and handle missed process exit events. It adds a --min-container-age flag to suppress metrics for short-lived containers and a --max-fqdns-per-container flag to bucket excess unique FQDNs under a ~other label. Additionally, it updates updateDelays to detect and clean up dead processes. The review feedback highlights a critical race condition where recycled PIDs could lead to the incorrect termination of newly started processes. To resolve this, the reviewer suggests returning *Process pointers from updateDelays and implementing a safe onProcessExitIf helper to verify the process instance before cleanup.
ReviewThe code is sound, but I'd hold this on one product question before merge. Before merge: rebase onto
|
Port of coroot/coroot-node-agent@33c46ec ("cap unique FQDN labels per container in container_dns_requests_total"). After the first --max-fqdns-per-container (default 50) distinct domains a container resolves, new domains are counted under domain="~other". Upstream keeps DNS metrics in a separate per-container vector. Here DNS is recorded through L7Stats.observe with destination and workload labels, so the cap lives in L7Stats and applies to the counter and the histogram alike. L7Stats never deletes series, so without a cap the domain label grows for the container's lifetime.
(cherry picked from commit 34ea61f99a4e5acce57f582665adc87cafcb1083) Includes the createdAt half of coroot/coroot-node-agent@75d6656 ("fix: systemd services disappearing after a restart"), which only applies to this flag: age counts from the earlier of the first process start and container discovery, so a restarted unit is not hidden for the minimum age after each restart. Fork adaptations: Collect here does not hold c.lock, so the dead-pid sweep goes through onProcessExit and the age check reads its fields under RLock (youngerThan). A pid counts as dead only when /proc/<pid> is gone; upstream treats any taskstats error as an exit, which here would close a live process's uprobes. The default is upstream's 30s: containers that live less than 30s, such as short jobs, produce no series.
The sweep in Collect finds exited processes without holding c.lock. If the pid was reused and the new process registered before the exit is handled, onProcessExit would untrack the new process. updateDelays now returns the *Process it found gone, and onProcessExitIf handles the exit only while that same process is registered.
faeabcb to
9b45e84
Compare
With 30s, a pod in CrashLoopBackOff at the 5-minute maximum backoff stays a zombie past gcInterval, is removed between restarts and comes back as a new container each time. If it dies within 30s it never reports: no OOM kills, memory or CPU for the pods most worth looking at. Jobs that finish in under 30s never appear either. Keep the flag, and enable it per install where short-lived series are a measured problem.
|
Thanks, the crash-loop case is real.
|
Summary
Second batch of #369: two limits on how many series the agent creates, ported from upstream coroot-node-agent.
container_dns_requests_totalkeeps the first--max-fqdns-per-container(default 50) distinct domains per container. Requests for any later domain are counted underdomain="~other".createdAtpart of coroot/coroot-node-agent@75d6656--min-container-age, off by default (0): when set, no series for a container until it has existed that long. Age counts from the earlier of its first process start and when the agent found it, so a restarted service doesn't disappear for 30s after each restart.Default: upstream ships 30s; this PR ships
0(off). With 30s, a pod in CrashLoopBackOff at the 5-minute maximum backoff is removed between restarts and comes back as a new container each time. If it dies within 30s it never reports: no OOM kills, memory or CPU. Jobs that finish in under 30s would never appear either. Set it per install where short-lived series are a measured problem.Engineering detail
Domain cap: why it's re-implemented. Upstream records DNS in its own per-container vector. Here DNS goes through
L7Stats.observewith destination and workload labels, so the cap lives inL7Statsand applies to the counter and the histogram alike.L7Statsnever deletes series, so without a cap the domain label grows for as long as the container lives. The DNS payload is now parsed once per request instead of twice.Minimum age: changes from upstream.
c.lockfor the wholeCollect. This fork doesn't, so the age check readsstartedAt,zombieAtandcreatedAtunder RLock (youngerThan), and the sweep of exited processes goes throughonProcessExit./proc/<pid>is gone, because a failed netlink call for a live process would otherwise close its uprobes.Pid reuse (faeabcb, from review): the exited-process sweep removes a pid only while that exact
*Processis still registered, so a reused pid can't untrack the new process. CI checks cover it; the e2e below ran before this commit and never reached that path.Default switched to 0 (d43f186, from review): with default flags, a 20s transient unit that the agent detected had 15 series 8s after starting; with
MIN_CONTAINER_AGE=30sit had none.CI: gofmt, goimports, vet, golangci-lint,
go test(excluding/containers) and the build all pass in a Linux container with Go 1.26.5.Found during the e2e, not changed here. On every build,
mainincluded, a burst of new systemd units started within about a minute of the agent starting can go undetected: 1–2 of 10 per burst. A unit that does nothing after it starts then never appears. On warm agents this happened once in 60 starts. In a later run, 2 of 3 transientsystemd-rununits went undetected on warm agents. I'll follow it up separately.Local e2e: I built agent binaries from this branch and from #370 and ran both side by side as systemd services on a local Debian 12 VM (kernel 6.1, systemd 252). This branch ran with
MAX_FQDNS_PER_CONTAINER=5.~otheron this branch, and all 10 on fix: port upstream push-path and container-tracking fixes (B1 of #369) #370.