Repository navigation
fix: absorb main's task timeout clamp into integration - #29
Merged
Merged
Conversation
…lled Two halves of NetGuard have been finished on the server for months and unreachable from a node the whole time. The reality panel showed 35 machines as "never reported" because nothing ever posted a snapshot: the agent had no collector. And the apply script the server generates calls `lattice-agent --guard-managed-sha` to record what it installed — a flag this binary did not implement, so every NetGuard apply came back without the canonical hash the server needs and the drift anchor could never be anything but unknown. The collector reads /proc rather than shelling out to ss or lsof, which are not installed on a minimal box: listening TCP and UDP sockets with their owning process where /proc/<pid>/fd is readable, interfaces with their addresses, the live managed table's canonical hash, the foreign nftables tables that are in force but not ours, and the nft version. All of it is read-only and best-effort — a host without nft or without root still reports listeners and interfaces, because a partial snapshot is what distinguishes "I can see this machine but not its ruleset" from a machine that is gone. The canonical hash strips handle numbers, trailing whitespace and blank lines before hashing, and keeps rule order, which is what makes drift mean anything: nft reassigns handles on every reload, so hashing them raw would report drift on a table nobody touched. --guard-managed-sha and the periodic report compute it with the same function, so the value recorded at apply time is comparable with the value observed later by construction. Tested: address decoding (little-endian per word, v4-mapped v6, garbage rejected), LISTEN-only filtering for TCP versus every bound UDP socket, a missing /proc file reading as no sockets, handle-insensitive but rule-sensitive canonicalization, and the exact wire shape the server's handler decodes. Full agent suite green. Not-tested: against a live nftables ruleset (this host has no nft; the release verification on a fleet node covers it).
The guard sent any timeout above ten minutes to the 30 second default, which inverts the caller's intent: ask for more time, receive the least possible. Agent updates are the case that matters. The server gives them 900s precisely because a slow uplink needs minutes to fetch a 12 MiB binary, so 900 became 30 and every node slower than roughly 400 KiB/s failed with context deadline exceeded, forever, because the fix for that can only arrive through an update it cannot complete. Measured on a real fleet node: 227 KiB/s from github, so the binary needs about 55 seconds. Fast nodes finished inside 30 and updated; slow ones never could, which is exactly the split seen across the fleet. Only an absent or nonsensical value gets the default now. Extracted so the boundaries are asserted directly rather than through a script that would have to actually sleep to prove it.
origin/main carried three commits integration lacked: c8de23e (task timeout clamp), 18eaceb (firewall reality collector and the --guard-managed-sha flag, tagged v0.3.4) and its merge 5c5c451. This merge brings the clamp onto the working branch and records that the rest is already on integration in a later form, so main can fast-forward to integration again after the next release. Kept from main: resolveTaskTimeout. Integration still sent any timeout above ten minutes to the 30 second default, so the 900 s the server gives agent updates ran with a 30 s deadline, and the stable v0.3.9 (tag at 6faf3d1) ships that guard. internal/taskexec/taskexec.go merged cleanly: absent or non-positive gets 30 s, above the maximum clamps to ten minutes, and Runner.run (behind Run and RunContext) now calls it. cmd/lattice-agent/main.go: integration's side, whole file. Main changed two things here. The version pair 0.3.4 / stable is older than integration's 0.3.9-alpha.9 / alpha lane. The firewall reality wiring (GuardManagedSHA field, a second "guard-managed-sha" flag, reportGuardReality(cfg) before runTasks) is superseded by integration's guard reality: printGuardManagedSHA with writeGuardManagedSHA, which fails closed and rejects anything but 64 hex characters, and reportGuardReality(ctx, cfg, collect) under guardRealityReportTimeout at the end of each poll, posting the typed model.GuardNodeReality with sshd facts to the same /api/agent/guard-reality as {node_id, reality}. Git had auto-merged main's hunks beside integration's, which defined the flag twice (a runtime panic) and reportGuardReality twice. cmd/lattice-agent/main_test.go: integration's side. Main's version pin expects 0.3.4, and TestGuardRealityPayloadShape is covered by TestReportGuardReality in guard_reality_test.go, which asserts the path and the {node_id, reality} body against a test server. internal/guardreality/{reality,listeners,guardreality_test}.go from main: not taken. They define Collect(ctx) in the same package as integration's Collect(ctx, source, nodeID), so the package cannot compile with both, and main's ManagedSHA hashes the text ruleset with handles stripped while integration's ParseNFTRuleset hashes the handle-stripped JSON objects of inet lattice_guard. The two digests never agree, so mixing them would make every drift comparison wrong. internal/taskexec/taskexec_test.go: both sides. Integration's TestRunContextCancelKillsRunningTask and TestRunContextKeepsOwnTimeout stay; main's TestOverlargeTimeoutClampsToMaxNotDefault is added after them with its boundary cases (900 and 86400 clamp to ten minutes, 600 is kept, 0 and -5 get 30 s). Left out and not superseded: main's listener collector read /proc/net and /proc/<pid>/fd, so it worked without ss. Integration's Collect runs ss -tulpnH and ip -j addr and fails the whole snapshot when either is missing. That is a behaviour gap to decide separately, not part of this merge. Tested: gofmt -l . empty; go vet ./... (darwin and GOOS=linux); go test -race -cover ./... all packages ok; the three taskexec tests above pass verbosely; sh scripts/check-release-workflow.sh and sh scripts/test-install-integrity.sh pass.
…-6348 govulncheck now fails CI on GO-2026-6348, heap exhaustion from fragmented HTTP/2 DATA frames in google.golang.org/grpc before v1.83.1. The agent reaches it through proxyusage.grpcQueryStats, the client that reads the sing-box stats API, so a hostile or broken endpoint on that port could grow the agent's heap without bound. integration at 6faf3d1 has the same dependency and would fail the same check; this PR is simply the first CI run since the advisory was published. go get raised golang.org/x/net to v0.55.0, golang.org/x/sys to v0.45.0, golang.org/x/text to v0.37.0 and genproto/googleapis/rpc to the 2026-05-26 pseudo-version, which are grpc v1.83.1's own minimums. No source change was needed. Tested: govulncheck ./... reports 0 vulnerabilities affecting the code; gofmt -l . empty; go vet ./... on darwin and GOOS=linux; go test -race -cover ./... all packages ok.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Merges
origin/maininto a branch offintegrationso the working branch gets the task timeout clamp (c8de23e) and main can fast-forward to integration again after the next release.Integration still sends any task timeout above ten minutes to the 30 second default. The server gives agent updates 900 s, so they run with a 30 s deadline and any node slower than about 400 KiB/s can never update itself. The stable
v0.3.9tag sits on6faf3d1and carries that guard. After this mergeRunner.run(behindRunandRunContext) callsresolveTaskTimeout: absent or non-positive gets 30 s, above the maximum clamps to ten minutes, anything in range is kept.Conflict resolutions
cmd/lattice-agent/main.go: integration's file. Main's 0.3.4 / stable version pair is older than integration's 0.3.9-alpha.9 / alpha lane, and main's firewall reality wiring is superseded by integration's guard reality (printGuardManagedSHA+writeGuardManagedSHA, which fails closed and accepts only 64 hex characters;reportGuardReality(ctx, cfg, collect)underguardRealityReportTimeout, posting typedmodel.GuardNodeRealitywith sshd facts to the same/api/agent/guard-reality). Git had auto-merged main's hunks next to integration's, registering theguard-managed-shaflag twice and definingreportGuardRealitytwice.cmd/lattice-agent/main_test.go: integration's file. Main's version pin expects 0.3.4; its payload-shape test is covered byTestReportGuardRealityinguard_reality_test.go.internal/guardreality/{reality,listeners,guardreality_test}.go(added on main): not taken. They redeclareCollectin the same package with a different signature, and main'sManagedSHAhashes the text ruleset while integration hashes the handle-stripped JSON objects, so the two digests never agree.internal/taskexec/taskexec_test.go: both sides. Integration's twoRunContexttests stay, andTestOverlargeTimeoutClampsToMaxNotDefaultis added with its boundary cases.internal/taskexec/taskexec.go: merged cleanly; the clamp replaces the inverted guard.Left out
Main's listener collector read
/proc/netand/proc/<pid>/fdand needed noss. Integration'sCollectrunsss -tulpnHandip -j addrand fails the whole snapshot if either is missing. That gap is not addressed here.Second commit: grpc v1.83.1
The first CI run failed only at govulncheck: GO-2026-6348 (heap exhaustion from fragmented HTTP/2 DATA frames) in
google.golang.org/grpcv1.82.1, reached throughproxyusage.grpcQueryStats, the sing-box stats client. integration has the same dependency, so the merge did not cause it. 60d9c93 moves grpc to v1.83.1;go getraised x/net, x/sys, x/text and genproto/googleapis/rpc to grpc's own minimums. No source change.Test plan
gofmt -l .is emptygo vet ./...on darwin and withGOOS=linuxgo test -race -cover ./...: every package okgo test -race -run 'TestOverlargeTimeoutClampsToMaxNotDefault|TestRunContext' -v ./internal/taskexec/: all five clamp cases and both RunContext tests passsh scripts/check-release-workflow.shsh scripts/test-install-integrity.shgovulncheck ./...after the bump: 0 vulnerabilities affecting the codeNo version bump, tag or release in this PR. Shipping the clamp to the fleet needs a node-agent release that includes it.