Skip to content

feat: heartbeat goroutine, self-armed watchdog, update health marker, monitor result batches - #30

Merged
lr00rl merged 8 commits into
integrationfrom
feat/heartbeat-watchdog-update-guard
Oct 3, 2026
Merged

lr00rl merged 8 commits into
integrationfrom
feat/heartbeat-watchdog-update-guard

Conversation

@lr00rl

@lr00rl lr00rl commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Node keepalive for the agent side (r1-keepalive P2 and P4), agent-side monitor result batching against the monitors lane's batch route, and the 0.3.10-alpha.1 bump with a changelog. Not tagged.

Commits:

  1. feat: send the heartbeat on its own goroutine with the work loop's health (81ae17e). The metrics beat runs every interval on its own goroutine with a 10 s deadline, so a slow control plane or a blocked report no longer flips a healthy node offline (four 30 s timeouts ahead of the old in-loop beat already exceeded the 90 s offline flip). Each beat carries loop_health: per-step last success, last error and consecutive errors, the step in progress, cycle timing, task batch age, and a blocked linechain recovery with its reason. A blocked recovery keeps the node visible and withholds durable-task-result-v1. Older servers ignore the field.
  2. feat: arm a systemd watchdog on local progress and mark a healthy first hello (a1977d4). With a notify socket (NotifyAccess=main), the agent sends READY=1, arms a 120 s watchdog itself with WATCHDOG_USEC=, and pets it only while the work loop and the heartbeat both moved within five minutes (or three intervals). Failing requests are progress; a step that never returns is not. After its first hello it writes its version to $RUNTIME_DIRECTORY/healthy. -compat-json gains features: ["sd-notify-v1","health-marker-v1"].
  3. feat: grant the agent a notify socket at install and supervise it under openrc (98801f7). The installer writes 10-lattice-keepalive.conf (NotifyAccess=main, RuntimeDirectory) only for a binary advertising sd-notify-v1 and removes it otherwise; openrc uses supervise-daemon (10 s respawn delay, unlimited) where present.
  4. feat: queue monitor results and send them in batches, one at a time on older servers (bd075a1). Results queue (2000 cap through an outage) and go out in batches of 200 on POST /api/agent/monitor-results, the route from lattice-server feat/monitor-results-hot-store (4cf84b0). 404 falls back to single posts for 30 minutes; 5xx keeps and resends; 400 drops the batch.
  5. docs: ... (d62a67e) README keepalive and supervision section.
  6. chore: bump to 0.3.10-alpha.1 and start a changelog (f449257).
  7. fix: arm the watchdog after startup recovery and judge the heartbeat only once it runs (20b67b9). Review finding: the watchdog was armed before durable linechain recovery and the heartbeat's progress was stamped at process start, so a healthy recovery longer than five minutes read as a stall and systemd killed the agent mid-recovery on every start. Startup recovery now runs to success before the watchdog arms (recoverThenSupervise), and a heartbeat that has not started is not judged.
  8. fix: send queued monitor results on a clean stop and count every result the server never stored (5b636ea). Review finding: the in-memory queue lost the last interval on every restart. A clean stop now cancels the flush loop and sends the queue once more within 5 s (after the 55 s task budget, under the default 90 s TimeoutStopSec). Every result the server never stored (overflow, a refused batch or single, a server-listed drop) is counted in loop_health.monitor_results_dropped and logged once per flush. A disk-backed outbox stays a follow-up.

Design change from the research: no Type=notify and no WatchdogSec in the unit. An older binary installed later under such a unit (a rollback, the v0.3.9 installer through a reconfigure, /install.sh moving an alpha node to stable) never sends READY and is killed in a loop. The drop-in only grants the socket and a runtime directory, and the binary arms its own watchdog, so the drop-in is inert under an old binary. Verified below.

Companion PRs: the update guard and the update-time drop-in are server-side, LatticeNet/lattice-server#133; the installer pin is LatticeNet/lattice-server#132 and LatticeNet/latticenet.github.io#4.

Evidence

  • go test -race -cover ./...: 21 packages ok. go vet ./... (darwin and linux/amd64) clean, gofmt -l . empty, sh scripts/check-release-workflow.sh and sh scripts/test-install-integrity.sh ok (new keepalive contract forbids Type=notify and WatchdogSec= in the installer).
  • Debian 12 container, systemd 252, installed by the new installer from a local release: Type=simple, NotifyAccess=main, WatchdogUSec=2min, /run/lattice-agent/healthy = version, loop_health on every beat, monitor results in batches. SIGSTOP at 19:01:29, then Watchdog timeout (limit 2min) at 19:02:48, restart at 19:02:59, re-armed. The released v0.3.9 binary under the same drop-in: active, WatchdogUSec=0.
  • Alpine 3.20 (OpenRC 0.54): supervise-daemon script, env from start_pre reaches the agent, kill -9 respawned in 10 s and hello again.
  • Real servers on loopback: lattice-server a379dc4 (monitors branch) accepted every batch, 7 results stored in 20 s from a 2 s monitor; lattice-server e07d2eb (a104) answered 404, the agent posted singles, 7 results stored. Node online with the new version on both.

Review round (2026-10-03), further evidence:

  • TestLongStartupRecoveryDoesNotTripTheWatchdog fails with the old stamp ("heartbeat has not moved for 10m30s"); TestBlockedStartupRecoveryBeatsBeforeTheWatchdogArms; three monitor queue tests for the drop count, the drain, and the cancelled flush.
  • Debian 12 container, systemd 252: a garbage journal blocks recovery; for 25 s WatchdogUSec=0, three beats carry linechain_blocked, no hello. Journal removed: "watchdog armed" and "connected" in the same second, WatchdogUSec=2min, NRestarts=0, the next beat says "watchdog":true. The guarded update with the new server script kept the new build after its hello.
  • Pre-hello beats against an a104 server (lattice-server e07d2eb on loopback): with recovery blocked the server accepted three metrics posts before any hello, showed the node online with the new version, logged no errors; after the journal was removed the agent said hello. handleAgentMetrics authenticates the node token only.
  • go test -race ./... (21 packages), vet darwin and linux/amd64, gofmt, check-release-workflow.sh, test-install-integrity.sh: ok.

Test plan

  • CI green (gofmt, release workflow contract, install integrity, vet, race, gosec, govulncheck).
  • Review together with the server guard PR; merge the server PR first so guarded updates exist when this binary is released.
  • Tag v0.3.10-alpha.1 (operator), canary one non-production node through a guarded update, confirm WatchdogUSec=2min, the marker, and loop_health in the server log or a debug read; then batches.
  • Server-side follow-up: store loop_health beside agent_runtime and derive "beating but stalled" (K2), not in this PR.

lr00rl added 6 commits October 2, 2026 12:23
…alth

The heartbeat was the metrics POST inside the serial work loop. Each
request there has a 30 s timeout, and four of them ahead of the beat
already exceed the server's 90 s offline flip, so a slow or half-hung
control plane turned healthy nodes offline. A blocked linechain recovery
skipped the cycle, beat included, so the node read "offline" while the
reason sat only in its journal, at startup as well as mid-life.

The beat now runs on its own goroutine every interval with a 10 s
deadline, independent of config, usage, inventory, tasks, monitors, log
sources, trace, debug and guard reality. It carries loop_health: process
start, the last cycle's start, end and duration, the step in progress
and since when, per step the last success, last error (one bounded line)
and consecutive errors, how long a task batch has been in flight, and
the linechain recovery block with its reason and start. Older servers
decode agent bodies leniently and ignore the field; the server-side
derivation ("beating but stalled", K2 in the keepalive research) is a
separate slice.

While recovery is blocked the beat withholds durable-task-result-v1, so
the server does not plan durable linechain work for a node that cannot
run it. In the normal start the beat begins right after hello, as the
first metrics report always did; only a blocked startup recovery starts
it before hello, with IPs resolved first so it does not report empty
addresses.

Rejected: keeping the beat in the loop and shortening the request
timeouts. Any step can still stall past the offline flip, and the
linechain-blocked path would still hide the node.

Tested: go test -race ./cmd/lattice-agent/ (a beat lands every interval
while /api/agent/config hangs; a blocked recovery beat carries the
reason and drops the durable capability, a cleared one restores it; step
outcomes, error bounding and cycle timing); gofmt; go vet (darwin and
linux).
…st hello

systemd restarted the agent only when its process exited. An agent whose
work loop wedged with the process alive stayed wedged, and a new binary
that never reached the control plane after an update crash-looped with
nothing on the node to put the old one back.

When its unit grants a notify socket (NotifyAccess=main), the agent now
sends READY=1 once its local state is open, arms a 120 s watchdog for
itself with WATCHDOG_USEC= (a unit that sets WatchdogSec decides the
timeout instead), and pets it at half the timeout only while the work
loop and the heartbeat both moved within the stall bound: five minutes,
or three intervals when the interval is longer. Every request has a
timeout, so a failing control plane is progress and never a restart; a
step that does not return is a stall, and systemd restarts the agent.
After the first successful hello the agent writes its version to
$RUNTIME_DIRECTORY/healthy, which an agent update's dead-man timer reads
before it keeps a new binary. -compat-json lists both contracts as
features, sd-notify-v1 and health-marker-v1, so an installer or update
writes the unit drop-in only for a binary that keeps them.

Rejected: Type=notify with WatchdogSec in the unit, as the keepalive
research proposed. Any older binary later installed under that unit (a
rollback, the stable installer through a reconfigure, /install.sh moving
an alpha node to the stable release) never sends READY=1 and is killed
at start in a loop. The unit stays Type=simple without WatchdogSec, the
drop-in only grants the socket and a runtime directory, and the binary
that speaks the protocol arms the watchdog itself, so the same drop-in
is inert under an old binary.

Tested: go test -race ./cmd/lattice-agent/ (stall judgement with a fake
clock; the watchdog pets while loops move and withholds while stalled;
READY over a real unixgram socket; WATCHDOG_USEC and WATCHDOG_PID
handling; arming needs a socket and prefers the unit's timeout; the
marker is written atomically at 0644 and only with a runtime directory;
-compat-json carries both features beside the fields the release
workflow greps). On systemd 252 (Debian 12 container, Type=simple plus
NotifyAccess=main): WatchdogUSec=2min after start, the marker holds the
version after hello, a SIGSTOPped agent hit "Watchdog timeout (limit
2min)" and was restarted and re-armed, and the released v0.3.9 binary
under the same drop-in ran with WatchdogUSec=0.
…er openrc

The installer's systemd unit gave the agent no notify socket, so the
watchdog the agent can now arm stayed off on every fresh install. On
openrc the script backgrounded the agent with a pidfile and nothing
restarted it after a crash, unlike systemd's Restart=always and
launchd's KeepAlive.

On systemd, when the installed binary lists sd-notify-v1 in
-compat-json, the installer writes
/etc/systemd/system/<service>.service.d/10-lattice-keepalive.conf with
NotifyAccess=main and RuntimeDirectory=<service> (0700), the same file
an agent update writes; for a binary without it the file is removed.
The base unit stays Type=simple without WatchdogSec, and the contract
test forbids either in the installer, so an older binary installed
under the drop-in later is never killed for silence. --uninstall
removes the drop-in with the unit.

On openrc, where supervise-daemon exists (OpenRC 0.21 and later), the
service runs under supervisor=supervise-daemon with respawn_delay=10 and
respawn_max=0; older OpenRC keeps command_background. Whether any fleet
node runs openrc is unverified.

Tested: sh scripts/test-install-integrity.sh (new keepalive contract,
existing contracts unchanged); sh -n. In a Debian 12 systemd container
the installer, fed a local release through a fake curl, wrote the
drop-in and the agent came up with WatchdogUSec=2min and its health
marker. In an Alpine 3.20 container (OpenRC 0.54) the installer wrote
the supervise-daemon script, the env from start_pre reached the agent
(LATTICE_SERVER, LATTICE_NODE_ID in /proc/<pid>/environ), and after
kill -9 the agent was respawned 10 s later and said hello again.
…n older servers

Each probe posted its own result the moment it finished, so an
all-nodes monitor cost one request per node per interval, and a result
whose post failed was logged and lost. During a control plane outage
every result taken in that window disappeared.

Probe goroutines now queue results and a flusher sends them every agent
interval on POST /api/agent/monitor-results ({"node_id", "results"}),
the batch route the monitors lane added to lattice-server
(feat/monitor-results-hot-store): oldest first, up to 200 per request
and five requests per flush. The server's answer is final for every
result in the batch (accepted, duplicate or dropped with a reason), so
the whole batch settles; a 5xx or a network error stores nothing and
the batch is sent again; a 400 refuses the batch as a whole and it is
dropped rather than left to block the queue. A server without the
route answers 404, and the agent then posts one result per request on
the single route for 30 minutes before it tries the batch route again,
so no compatibility floor is needed. On the single route a 4xx drops
that result and a 401, 429 or 5xx stops the flush with the rest kept.

The queue is the outage buffer, bounded at 2000 results. Past that the
oldest are dropped and counted (the server refuses results stamped more
than 24 hours before they arrive in any case), and loop_health reports
the queue depth and the drop count. Sequence numbers keep a send that
overlapped an overflow from settling results it never sent.

Rejected: a "results" array on the single route. Older servers decode
agent bodies leniently, read it as one empty result and answer 200, so
the agent would believe it had delivered what was dropped.

Tested: go test -race ./cmd/lattice-agent/ (ordered batches of 200,
200, 50; 503 and network errors keep and resend; 404 falls back to
singles and re-probes after the window; 400 and 422 are final, 401
stops with the rest kept; overflow bounds and settle under overflow;
the request on the wire carries the bearer header, node_id and the
results array). Against real servers on loopback with a 2 s tcp
monitor for 20 s: lattice-server a379dc4 (the monitors branch, bolt hot
store) accepted every batch and stored 7 results; lattice-server
e07d2eb (a104) answered 404, the agent posted singles, and 7 results
were stored.
…ult batching

The README said the metrics heartbeat carried agent_runtime and nothing
about how the agent keeps itself alive. It now has a keepalive and
supervision section: the heartbeat goroutine and every loop_health
field; the keepalive drop-in, why it grants a notify socket without
turning the unit into Type=notify, and what the watchdog judges; the
health marker and the server-side update guard that reads it; openrc
supervision; and the monitor result batches with their fallback and
outage buffer.

Tested: rendered as plain markdown; no em or en dashes.
Prepares the next prerelease without tagging it. The version constant
and its pinning test move to 0.3.10-alpha.1, and CHANGELOG.md records
what that tag would carry since v0.3.9: the heartbeat goroutine and
loop health, the self-armed watchdog and health marker, the installer
drop-in and openrc supervision, monitor result batches, and the two
fixes absorbed from main (the task-timeout clamp and grpc v1.83.1). No
compatibility floor moves. The release notes say that the update guard
needs the matching lattice-server change and that the roll goes canary
first.

Tested: go test -race -cover ./... (21 packages ok); go vet (darwin and
linux/amd64); gofmt; sh scripts/check-release-workflow.sh; sh
scripts/test-install-integrity.sh.
lr00rl added 2 commits October 3, 2026 00:21
…only once it runs

The watchdog was armed before durable linechain recovery, and loop health
stamped the heartbeat's progress at process start although the heartbeat
only starts after hello (or when recovery is blocked). Recovery restarts
sing-box once per interrupted journal under context.Background, and
health.run stamps progress only at the start and end of a step. A healthy
recovery that outlasted the five minute stall bound therefore left both
the work loop and the heartbeat stale, the watchdog stopped petting, and
systemd killed the agent about two minutes later, mid-recovery, on every
start.

recoverThenSupervise now runs startup recovery to success and only then
arms the watchdog, and the heartbeat's progress stays zero until
heartbeat.start stamps it, so stalled() does not judge a heartbeat that
has not started and does judge a first beat that never returns. The
recovery at the top of each work loop cycle stays under the watchdog: it
runs only while no task is in flight and every task poll recovers first,
so it meets at most the journal of the task that just ended.

Rejected: starting the heartbeat before recovery on every start. It
fixes the stale beat but not the work loop, which still sits in one
unbounded step, and it would send metrics before hello on every start
rather than only when recovery is blocked. Rejected: progress callbacks
inside the linechain manager; one sing-box restart is bounded by
systemd's own start and stop timeouts, and the startup path is the only
one that can meet several journals.

Tested: go test -race ./cmd/lattice-agent/; new
TestLongStartupRecoveryDoesNotTripTheWatchdog fails when beatProgress is
stamped at process start (stalled: heartbeat has not moved for 10m30s);
TestBlockedStartupRecoveryBeatsBeforeTheWatchdogArms.
…lt the server never stored

The monitor result queue is in memory, so the results of the last interval
were lost on every restart, and with the watchdog and the update guard a
restart is now more likely than before. The drop count on the heartbeat
covered queue overflow only: a batch the server refused whole, a single
result refused on the old route, and results a batch answer listed as
dropped were logged at debug level or not counted at all.

On SIGTERM the agent now stops the flush loop, which cancels the flush in
flight, and sends the queue once more within 5 s after the task shutdown
budget (55 s), staying under systemd's default 90 s TimeoutStopSec. What
cannot go out is logged with its count. Every result the server never
stored goes into the dropped counter that rides on each heartbeat as
loop_health.monitor_results_dropped, and new drops are logged once per
flush rather than once per result. A server-listed drop is logged at the
normal level with the first reason.

Not done here: a disk-backed outbox, which is what survives a crash or a
watchdog kill, and a server that shows the dropped count (the server
ignores loop_health today). Both stay on the lane's deferred list. Results
still wait up to one interval before they are sent; sending each result
at once would give up the batching that cuts request volume.

Tested: go test -race ./cmd/lattice-agent/;
TestMonitorResultsCountEveryResultTheServerNeverStored,
TestMonitorDrainSendsTheQueueOnStop,
TestMonitorFlushLoopStopsItsFlushInFlight.
@lr00rl
lr00rl merged commit 5d988c8 into integration Oct 3, 2026
1 check passed
@lr00rl
lr00rl deleted the feat/heartbeat-watchdog-update-guard branch October 3, 2026 08:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant