Skip to content

fix: absorb main's task timeout clamp into integration - #29

Merged
lr00rl merged 5 commits into
integrationfrom
fix/absorb-main-task-timeout
Oct 2, 2026
Merged

lr00rl merged 5 commits into
integrationfrom
fix/absorb-main-task-timeout

Conversation

@lr00rl

@lr00rl lr00rl commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Merges origin/main into a branch off integration so the working branch gets the task timeout clamp (c8de23e) and main can fast-forward to integration again after the next release.

Integration still sends any task timeout above ten minutes to the 30 second default. The server gives agent updates 900 s, so they run with a 30 s deadline and any node slower than about 400 KiB/s can never update itself. The stable v0.3.9 tag sits on 6faf3d1 and carries that guard. After this merge Runner.run (behind Run and RunContext) calls resolveTaskTimeout: absent or non-positive gets 30 s, above the maximum clamps to ten minutes, anything in range is kept.

Conflict resolutions

  • cmd/lattice-agent/main.go: integration's file. Main's 0.3.4 / stable version pair is older than integration's 0.3.9-alpha.9 / alpha lane, and main's firewall reality wiring is superseded by integration's guard reality (printGuardManagedSHA + writeGuardManagedSHA, which fails closed and accepts only 64 hex characters; reportGuardReality(ctx, cfg, collect) under guardRealityReportTimeout, posting typed model.GuardNodeReality with sshd facts to the same /api/agent/guard-reality). Git had auto-merged main's hunks next to integration's, registering the guard-managed-sha flag twice and defining reportGuardReality twice.
  • cmd/lattice-agent/main_test.go: integration's file. Main's version pin expects 0.3.4; its payload-shape test is covered by TestReportGuardReality in guard_reality_test.go.
  • internal/guardreality/{reality,listeners,guardreality_test}.go (added on main): not taken. They redeclare Collect in the same package with a different signature, and main's ManagedSHA hashes the text ruleset while integration hashes the handle-stripped JSON objects, so the two digests never agree.
  • internal/taskexec/taskexec_test.go: both sides. Integration's two RunContext tests stay, and TestOverlargeTimeoutClampsToMaxNotDefault is added with its boundary cases.
  • internal/taskexec/taskexec.go: merged cleanly; the clamp replaces the inverted guard.

Left out

Main's listener collector read /proc/net and /proc/<pid>/fd and needed no ss. Integration's Collect runs ss -tulpnH and ip -j addr and fails the whole snapshot if either is missing. That gap is not addressed here.

Second commit: grpc v1.83.1

The first CI run failed only at govulncheck: GO-2026-6348 (heap exhaustion from fragmented HTTP/2 DATA frames) in google.golang.org/grpc v1.82.1, reached through proxyusage.grpcQueryStats, the sing-box stats client. integration has the same dependency, so the merge did not cause it. 60d9c93 moves grpc to v1.83.1; go get raised x/net, x/sys, x/text and genproto/googleapis/rpc to grpc's own minimums. No source change.

Test plan

  • gofmt -l . is empty
  • go vet ./... on darwin and with GOOS=linux
  • go test -race -cover ./...: every package ok
  • go test -race -run 'TestOverlargeTimeoutClampsToMaxNotDefault|TestRunContext' -v ./internal/taskexec/: all five clamp cases and both RunContext tests pass
  • sh scripts/check-release-workflow.sh
  • sh scripts/test-install-integrity.sh
  • govulncheck ./... after the bump: 0 vulnerabilities affecting the code
  • CI green on this PR (run 36997081919 on 60d9c93)

No version bump, tag or release in this PR. Shipping the clamp to the fleet needs a node-agent release that includes it.

lr00rl added 5 commits August 18, 2026 19:39
…lled

Two halves of NetGuard have been finished on the server for months and
unreachable from a node the whole time.

The reality panel showed 35 machines as "never reported" because nothing
ever posted a snapshot: the agent had no collector. And the apply script
the server generates calls `lattice-agent --guard-managed-sha` to record
what it installed — a flag this binary did not implement, so every
NetGuard apply came back without the canonical hash the server needs and
the drift anchor could never be anything but unknown.

The collector reads /proc rather than shelling out to ss or lsof, which
are not installed on a minimal box: listening TCP and UDP sockets with
their owning process where /proc/<pid>/fd is readable, interfaces with
their addresses, the live managed table's canonical hash, the foreign
nftables tables that are in force but not ours, and the nft version. All
of it is read-only and best-effort — a host without nft or without root
still reports listeners and interfaces, because a partial snapshot is
what distinguishes "I can see this machine but not its ruleset" from a
machine that is gone.

The canonical hash strips handle numbers, trailing whitespace and blank
lines before hashing, and keeps rule order, which is what makes drift
mean anything: nft reassigns handles on every reload, so hashing them
raw would report drift on a table nobody touched. --guard-managed-sha
and the periodic report compute it with the same function, so the value
recorded at apply time is comparable with the value observed later by
construction.

Tested: address decoding (little-endian per word, v4-mapped v6, garbage
rejected), LISTEN-only filtering for TCP versus every bound UDP socket, a
missing /proc file reading as no sockets, handle-insensitive but
rule-sensitive canonicalization, and the exact wire shape the server's
handler decodes. Full agent suite green.
Not-tested: against a live nftables ruleset (this host has no nft; the
release verification on a fleet node covers it).
The guard sent any timeout above ten minutes to the 30 second default,
which inverts the caller's intent: ask for more time, receive the least
possible. Agent updates are the case that matters. The server gives them
900s precisely because a slow uplink needs minutes to fetch a 12 MiB
binary, so 900 became 30 and every node slower than roughly 400 KiB/s
failed with context deadline exceeded, forever, because the fix for that
can only arrive through an update it cannot complete.

Measured on a real fleet node: 227 KiB/s from github, so the binary needs
about 55 seconds. Fast nodes finished inside 30 and updated; slow ones
never could, which is exactly the split seen across the fleet.

Only an absent or nonsensical value gets the default now. Extracted so
the boundaries are asserted directly rather than through a script that
would have to actually sleep to prove it.
origin/main carried three commits integration lacked: c8de23e (task
timeout clamp), 18eaceb (firewall reality collector and the
--guard-managed-sha flag, tagged v0.3.4) and its merge 5c5c451. This
merge brings the clamp onto the working branch and records that the
rest is already on integration in a later form, so main can fast-forward
to integration again after the next release.

Kept from main: resolveTaskTimeout. Integration still sent any timeout
above ten minutes to the 30 second default, so the 900 s the server
gives agent updates ran with a 30 s deadline, and the stable v0.3.9
(tag at 6faf3d1) ships that guard. internal/taskexec/taskexec.go merged
cleanly: absent or non-positive gets 30 s, above the maximum clamps to
ten minutes, and Runner.run (behind Run and RunContext) now calls it.

cmd/lattice-agent/main.go: integration's side, whole file. Main changed
two things here. The version pair 0.3.4 / stable is older than
integration's 0.3.9-alpha.9 / alpha lane. The firewall reality wiring
(GuardManagedSHA field, a second "guard-managed-sha" flag,
reportGuardReality(cfg) before runTasks) is superseded by integration's
guard reality: printGuardManagedSHA with writeGuardManagedSHA, which
fails closed and rejects anything but 64 hex characters, and
reportGuardReality(ctx, cfg, collect) under guardRealityReportTimeout at
the end of each poll, posting the typed model.GuardNodeReality with sshd
facts to the same /api/agent/guard-reality as {node_id, reality}. Git
had auto-merged main's hunks beside integration's, which defined the
flag twice (a runtime panic) and reportGuardReality twice.

cmd/lattice-agent/main_test.go: integration's side. Main's version pin
expects 0.3.4, and TestGuardRealityPayloadShape is covered by
TestReportGuardReality in guard_reality_test.go, which asserts the path
and the {node_id, reality} body against a test server.

internal/guardreality/{reality,listeners,guardreality_test}.go from main:
not taken. They define Collect(ctx) in the same package as integration's
Collect(ctx, source, nodeID), so the package cannot compile with both,
and main's ManagedSHA hashes the text ruleset with handles stripped while
integration's ParseNFTRuleset hashes the handle-stripped JSON objects of
inet lattice_guard. The two digests never agree, so mixing them would
make every drift comparison wrong.

internal/taskexec/taskexec_test.go: both sides. Integration's
TestRunContextCancelKillsRunningTask and TestRunContextKeepsOwnTimeout
stay; main's TestOverlargeTimeoutClampsToMaxNotDefault is added after
them with its boundary cases (900 and 86400 clamp to ten minutes, 600 is
kept, 0 and -5 get 30 s).

Left out and not superseded: main's listener collector read /proc/net
and /proc/<pid>/fd, so it worked without ss. Integration's Collect runs
ss -tulpnH and ip -j addr and fails the whole snapshot when either is
missing. That is a behaviour gap to decide separately, not part of this
merge.

Tested: gofmt -l . empty; go vet ./... (darwin and GOOS=linux); go test
-race -cover ./... all packages ok; the three taskexec tests above pass
verbosely; sh scripts/check-release-workflow.sh and sh
scripts/test-install-integrity.sh pass.
…-6348

govulncheck now fails CI on GO-2026-6348, heap exhaustion from fragmented
HTTP/2 DATA frames in google.golang.org/grpc before v1.83.1. The agent
reaches it through proxyusage.grpcQueryStats, the client that reads the
sing-box stats API, so a hostile or broken endpoint on that port could
grow the agent's heap without bound. integration at 6faf3d1 has the same
dependency and would fail the same check; this PR is simply the first CI
run since the advisory was published.

go get raised golang.org/x/net to v0.55.0, golang.org/x/sys to v0.45.0,
golang.org/x/text to v0.37.0 and genproto/googleapis/rpc to the
2026-05-26 pseudo-version, which are grpc v1.83.1's own minimums. No
source change was needed.

Tested: govulncheck ./... reports 0 vulnerabilities affecting the code;
gofmt -l . empty; go vet ./... on darwin and GOOS=linux; go test -race
-cover ./... all packages ok.
@lr00rl
lr00rl merged commit 867ee3b into integration Oct 2, 2026
1 check passed
@lr00rl
lr00rl deleted the fix/absorb-main-task-timeout branch October 2, 2026 10:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant