Repository navigation
fix: detect systemd units whose start event sees /init.scope - #374
Conversation
systemd forks a unit's process inside its own /init.scope and moves it to the unit's cgroup before exec. The registry meant not to cache such a pid as ignored (cg.Id == "/init.scope"), but cgroup parsing skips /init.scope and the root cgroup, so their Id is "" and the check never matched. When the start event was handled before the move, the pid was cached as ignored for 15s, its exec was dropped, and a unit that did nothing else was never detected. Handling proc events on wakeup (#366) made that ordering common: 2 of 6 transient units on a warm agent. Match the empty Id instead, and pin the parsing it relies on in a test.
There was a problem hiding this comment.
Code Review
This pull request updates the container registry to properly handle processes in systemd's /init.scope or the root cgroup, which parse to an empty ID, preventing them from being incorrectly cached as ignored. A new unit test is added to verify this behavior. Feedback suggests using require.Nil instead of assert.Nil in the test to avoid potential nil pointer dereferences if the setup or parsing fails.
|
Good catch: Two small things for a later pass, neither blocking:
|
|
Thanks. Both follow-ups are in #376.
|
…s skipped #374 stopped caching every pid whose cgroup Id is empty, which covers the root cgroup as well as /init.scope. On hosts without systemd, daemons can run in the root cgroup, and each of their connect, listen and file-open events then re-read /proc/<pid>/cgroup. Cgroup now records whether the process is in /init.scope, and only those pids skip the ignore cache. Also drop the inline cleanup in getOrCreateContainer: it checked the entry it had just written, so it never deleted anything.
…s skipped (#376) #374 stopped caching every pid whose cgroup Id is empty, which covers the root cgroup as well as /init.scope. On hosts without systemd, daemons can run in the root cgroup, and each of their connect, listen and file-open events then re-read /proc/<pid>/cgroup. Cgroup now records whether the process is in /init.scope, and only those pids skip the ignore cache. Also drop the inline cleanup in getOrCreateContainer: it checked the entry it had just written, so it never deleted anything.
Summary
New systemd units were sometimes never detected. On a warm agent, 4 of 10 short-lived
systemd-rununits never appeared. A unit that does nothing after it starts then stays invisible for good.Cause. systemd forks a unit's process inside its own
/init.scopeand moves it to the unit's cgroup just beforeexec. The registry was meant not to cache such a pid as ignored, through the checkcg.Id == "/init.scope". But cgroup parsing skips/init.scopeand the root cgroup, so for those the Id is""and that check never matched. When the agent handled the start event before systemd moved the process, it cached the pid as ignored for 15s. It then dropped theexec, the one event a sleeping service produces. #366, which handles process events as soon as they arrive, made that ordering common.Fix. The check now matches the empty Id, which is what
/init.scopeactually parses to. pid 1 stays cached as before. A new test incgrouppins that/init.scopeand/parse to an empty Id, so the condition can't go dead again.The same dead check exists upstream (coroot-node-agent
containers/registry.go, since its cgroup parsing started skipping/init.scopein #203).Engineering detail
How it was found: while testing #371, a few new units never showed up in either build. A diagnostic build that logged every process start and exec event with the cgroup read at handling time showed the missed units:
"ignoring";With the fix: the same build logged
"ignoring without persisting"at start, and detected the unit at exec.Cost: pids in
/init.scopeor the root cgroup other than pid 1 are no longer cached as ignored, so each of their events re-reads/proc/<pid>/cgroup. In practice these are systemd's freshly forked children, which move to their unit almost immediately, and kernel threads, which produce no events the registry handles.CI: gofmt, goimports, vet, golangci-lint,
go test(excluding/containers, including the new cgroup test) and the build all pass in a Linux container with Go 1.26.5.Local e2e: I built agent binaries from this branch and from main and ran them side by side as systemd services on a local Debian 12 VM (kernel 6.1, systemd 252).
systemd-rununits started 3s apart, this branch detected 10/10 and main 6/10.