Repository navigation
fix(profiler): keep partial results when a PID fails, and explain pprof failures - #139
Conversation
Every profiler fanned out over the container's leaf PIDs and returned the first error, which cancelled the PIDs not yet started and failed the whole run. A container whose process tree has a shell, a helper process or anything else the tool cannot attach to next to the real workload could therefore never be profiled without --pid, even though the workload's own PID would have produced a result. Run every PID through one shared helper instead. Each successful PID still publishes its own result event; the run succeeds when at least one PID did and emits a notice naming the skipped PIDs and why. It fails only when every PID failed, with each reason in the error: "PID <n>: <reason>; PID <m>: <reason>". An empty PID list is now an error rather than a success that publishes nothing. The Go pprof profiler also set the PID on the shared job from concurrent tasks, so a later PID could leak into an earlier PID's scrape; each PID now gets its own copy of the job. The helper uses a WaitGroup, so the worker-pool dependency is dropped. The test fakes for the per-PID managers now lock in invoke, which the helper calls from one goroutine per PID.
bpf, perf, py-spy and austin logged a flamegraph rendering error and returned nil without publishing anything. With a single PID the run then ended as a success with no result, leaving the caller waiting for a file that would never come. Return the error instead, so the PID counts as failed: the run still succeeds when another PID published, and otherwise fails with the reason.
The pprof profiler reported every failed scrape as the raw nsenter+wget command line and stderr, and when it could not find the target's listening port it silently scraped :8080 instead, so a failure there read as if the target's own pprof port were down. Map the failures busybox wget reports to the cause: - HTTP 401/403: the pprof endpoint on :<port> requires authentication - HTTP 404: no /debug/pprof handler on :<port> (is net/http/pprof registered?) - connection refused: nothing listening on :<port> Anything else is passed through as before. :8080 is still tried when the port cannot be detected, but a failure there now says the port could not be detected and the default was tried. The scrape now runs through the profiler's commander, like the other profilers, so it can be exercised in tests.
python-austin-multi-pid.sh runs a Python program next to a non-Python process, in both start orders, and lets the agent find the container's leaf PIDs itself. The Python PID's profile must be published, the other PID named in a notice, and the run must not fail. go-pprof.sh builds a small Go server and scrapes it with the bpf image's own wget: with pprof registered a profile is published; behind basic auth, without a pprof handler, and with nothing listening, the agent's error must give the matching reason. Both run in Code Verify.
There was a problem hiding this comment.
Code Review
This pull request refactors concurrent profiling across all supported languages and tools by replacing the third-party pond library with a new custom common.ProfilePIDs helper. This helper runs profiling concurrently with a staggered delay and handles partial failures gracefully, ensuring that a single failing PID does not fail the entire run. Additionally, error reporting has been improved (especially for Go pprof scrape failures), fake managers have been updated with mutexes to serialize concurrent invocations in tests, and new E2E tests have been introduced. However, a critical compilation error was identified in the new common.ProfilePIDs implementation, where a non-existent Go method is called on a standard sync.WaitGroup.
Summary
A PID the profiler can't attach to no longer fails the whole run. That can be a shell, a helper process, or a process that isn't the target runtime.
PID <n>: <reason>; PID <m>: <reason>.A flamegraph that can't be rendered now fails that PID (bpf, perf, py-spy, austin). Before, the run ended as a success with no result.
Go pprof scrape failures name the cause:
/debug/pprofhandler (404);When the listening port can't be detected, the error says the default
:8080was tried.Type of change
Test plan
test/e2e/python-austin-multi-pid.sh: a Python process next to a non-Python one, both start orders, with the PIDs found by the agent itself.test/e2e/go-pprof.sh: Go servers with pprof open, behind basic auth, not registered, and with nothing listening.Engineering detail
common.ProfilePIDsreplaces eight copies of the worker-pool fan-out.alitto/ponddependency is dropped.resultevent per published file. A partial success still emits the successful PID's result first, followed by the notice andprogress: ended. Before, a failing PID ended the run with anerrorevent, after any result already published, and PIDs not yet started were cancelled.PIDon a shared one.invoke, which-raceneeds once PIDs run on separate goroutines.Local e2e: built the python and bpf profiler images from this branch (arm64) and ran
test/e2e/python-austin-multi-pid.sh,test/e2e/go-pprof.shandtest/e2e/python-austin.sh; all passed.Profiled 1 of 2 PIDs; skipped PID 8: could not launch profiler: austin (PID 8, /usr/bin/sleep): ….the pprof endpoint on :6060 requires authentication,no /debug/pprof handler on :9090 (is net/http/pprof registered?)andcould not detect the listening port (…), so tried the default :8080: … nothing listening on :8080.main, both new scripts fail. The multi-process run ends with an error event and no notice, and the pprof errors are raw wget output.Checks:
go build,go vetandgo test -race ./internal/agent/profiler/...are clean on Linux.golangci-lintshows no new findings, and fixes one existing QF1003.Checklist
go build ./...andgo test ./...pass locallynoticeevent, which is an existing event type.