Repository navigation
feat: heartbeat goroutine, self-armed watchdog, update health marker, monitor result batches - #30
Merged
Conversation
…alth
The heartbeat was the metrics POST inside the serial work loop. Each
request there has a 30 s timeout, and four of them ahead of the beat
already exceed the server's 90 s offline flip, so a slow or half-hung
control plane turned healthy nodes offline. A blocked linechain recovery
skipped the cycle, beat included, so the node read "offline" while the
reason sat only in its journal, at startup as well as mid-life.
The beat now runs on its own goroutine every interval with a 10 s
deadline, independent of config, usage, inventory, tasks, monitors, log
sources, trace, debug and guard reality. It carries loop_health: process
start, the last cycle's start, end and duration, the step in progress
and since when, per step the last success, last error (one bounded line)
and consecutive errors, how long a task batch has been in flight, and
the linechain recovery block with its reason and start. Older servers
decode agent bodies leniently and ignore the field; the server-side
derivation ("beating but stalled", K2 in the keepalive research) is a
separate slice.
While recovery is blocked the beat withholds durable-task-result-v1, so
the server does not plan durable linechain work for a node that cannot
run it. In the normal start the beat begins right after hello, as the
first metrics report always did; only a blocked startup recovery starts
it before hello, with IPs resolved first so it does not report empty
addresses.
Rejected: keeping the beat in the loop and shortening the request
timeouts. Any step can still stall past the offline flip, and the
linechain-blocked path would still hide the node.
Tested: go test -race ./cmd/lattice-agent/ (a beat lands every interval
while /api/agent/config hangs; a blocked recovery beat carries the
reason and drops the durable capability, a cleared one restores it; step
outcomes, error bounding and cycle timing); gofmt; go vet (darwin and
linux).
…st hello systemd restarted the agent only when its process exited. An agent whose work loop wedged with the process alive stayed wedged, and a new binary that never reached the control plane after an update crash-looped with nothing on the node to put the old one back. When its unit grants a notify socket (NotifyAccess=main), the agent now sends READY=1 once its local state is open, arms a 120 s watchdog for itself with WATCHDOG_USEC= (a unit that sets WatchdogSec decides the timeout instead), and pets it at half the timeout only while the work loop and the heartbeat both moved within the stall bound: five minutes, or three intervals when the interval is longer. Every request has a timeout, so a failing control plane is progress and never a restart; a step that does not return is a stall, and systemd restarts the agent. After the first successful hello the agent writes its version to $RUNTIME_DIRECTORY/healthy, which an agent update's dead-man timer reads before it keeps a new binary. -compat-json lists both contracts as features, sd-notify-v1 and health-marker-v1, so an installer or update writes the unit drop-in only for a binary that keeps them. Rejected: Type=notify with WatchdogSec in the unit, as the keepalive research proposed. Any older binary later installed under that unit (a rollback, the stable installer through a reconfigure, /install.sh moving an alpha node to the stable release) never sends READY=1 and is killed at start in a loop. The unit stays Type=simple without WatchdogSec, the drop-in only grants the socket and a runtime directory, and the binary that speaks the protocol arms the watchdog itself, so the same drop-in is inert under an old binary. Tested: go test -race ./cmd/lattice-agent/ (stall judgement with a fake clock; the watchdog pets while loops move and withholds while stalled; READY over a real unixgram socket; WATCHDOG_USEC and WATCHDOG_PID handling; arming needs a socket and prefers the unit's timeout; the marker is written atomically at 0644 and only with a runtime directory; -compat-json carries both features beside the fields the release workflow greps). On systemd 252 (Debian 12 container, Type=simple plus NotifyAccess=main): WatchdogUSec=2min after start, the marker holds the version after hello, a SIGSTOPped agent hit "Watchdog timeout (limit 2min)" and was restarted and re-armed, and the released v0.3.9 binary under the same drop-in ran with WatchdogUSec=0.
…er openrc The installer's systemd unit gave the agent no notify socket, so the watchdog the agent can now arm stayed off on every fresh install. On openrc the script backgrounded the agent with a pidfile and nothing restarted it after a crash, unlike systemd's Restart=always and launchd's KeepAlive. On systemd, when the installed binary lists sd-notify-v1 in -compat-json, the installer writes /etc/systemd/system/<service>.service.d/10-lattice-keepalive.conf with NotifyAccess=main and RuntimeDirectory=<service> (0700), the same file an agent update writes; for a binary without it the file is removed. The base unit stays Type=simple without WatchdogSec, and the contract test forbids either in the installer, so an older binary installed under the drop-in later is never killed for silence. --uninstall removes the drop-in with the unit. On openrc, where supervise-daemon exists (OpenRC 0.21 and later), the service runs under supervisor=supervise-daemon with respawn_delay=10 and respawn_max=0; older OpenRC keeps command_background. Whether any fleet node runs openrc is unverified. Tested: sh scripts/test-install-integrity.sh (new keepalive contract, existing contracts unchanged); sh -n. In a Debian 12 systemd container the installer, fed a local release through a fake curl, wrote the drop-in and the agent came up with WatchdogUSec=2min and its health marker. In an Alpine 3.20 container (OpenRC 0.54) the installer wrote the supervise-daemon script, the env from start_pre reached the agent (LATTICE_SERVER, LATTICE_NODE_ID in /proc/<pid>/environ), and after kill -9 the agent was respawned 10 s later and said hello again.
…n older servers
Each probe posted its own result the moment it finished, so an
all-nodes monitor cost one request per node per interval, and a result
whose post failed was logged and lost. During a control plane outage
every result taken in that window disappeared.
Probe goroutines now queue results and a flusher sends them every agent
interval on POST /api/agent/monitor-results ({"node_id", "results"}),
the batch route the monitors lane added to lattice-server
(feat/monitor-results-hot-store): oldest first, up to 200 per request
and five requests per flush. The server's answer is final for every
result in the batch (accepted, duplicate or dropped with a reason), so
the whole batch settles; a 5xx or a network error stores nothing and
the batch is sent again; a 400 refuses the batch as a whole and it is
dropped rather than left to block the queue. A server without the
route answers 404, and the agent then posts one result per request on
the single route for 30 minutes before it tries the batch route again,
so no compatibility floor is needed. On the single route a 4xx drops
that result and a 401, 429 or 5xx stops the flush with the rest kept.
The queue is the outage buffer, bounded at 2000 results. Past that the
oldest are dropped and counted (the server refuses results stamped more
than 24 hours before they arrive in any case), and loop_health reports
the queue depth and the drop count. Sequence numbers keep a send that
overlapped an overflow from settling results it never sent.
Rejected: a "results" array on the single route. Older servers decode
agent bodies leniently, read it as one empty result and answer 200, so
the agent would believe it had delivered what was dropped.
Tested: go test -race ./cmd/lattice-agent/ (ordered batches of 200,
200, 50; 503 and network errors keep and resend; 404 falls back to
singles and re-probes after the window; 400 and 422 are final, 401
stops with the rest kept; overflow bounds and settle under overflow;
the request on the wire carries the bearer header, node_id and the
results array). Against real servers on loopback with a 2 s tcp
monitor for 20 s: lattice-server a379dc4 (the monitors branch, bolt hot
store) accepted every batch and stored 7 results; lattice-server
e07d2eb (a104) answered 404, the agent posted singles, and 7 results
were stored.
…ult batching The README said the metrics heartbeat carried agent_runtime and nothing about how the agent keeps itself alive. It now has a keepalive and supervision section: the heartbeat goroutine and every loop_health field; the keepalive drop-in, why it grants a notify socket without turning the unit into Type=notify, and what the watchdog judges; the health marker and the server-side update guard that reads it; openrc supervision; and the monitor result batches with their fallback and outage buffer. Tested: rendered as plain markdown; no em or en dashes.
Prepares the next prerelease without tagging it. The version constant and its pinning test move to 0.3.10-alpha.1, and CHANGELOG.md records what that tag would carry since v0.3.9: the heartbeat goroutine and loop health, the self-armed watchdog and health marker, the installer drop-in and openrc supervision, monitor result batches, and the two fixes absorbed from main (the task-timeout clamp and grpc v1.83.1). No compatibility floor moves. The release notes say that the update guard needs the matching lattice-server change and that the roll goes canary first. Tested: go test -race -cover ./... (21 packages ok); go vet (darwin and linux/amd64); gofmt; sh scripts/check-release-workflow.sh; sh scripts/test-install-integrity.sh.
3 tasks
…only once it runs The watchdog was armed before durable linechain recovery, and loop health stamped the heartbeat's progress at process start although the heartbeat only starts after hello (or when recovery is blocked). Recovery restarts sing-box once per interrupted journal under context.Background, and health.run stamps progress only at the start and end of a step. A healthy recovery that outlasted the five minute stall bound therefore left both the work loop and the heartbeat stale, the watchdog stopped petting, and systemd killed the agent about two minutes later, mid-recovery, on every start. recoverThenSupervise now runs startup recovery to success and only then arms the watchdog, and the heartbeat's progress stays zero until heartbeat.start stamps it, so stalled() does not judge a heartbeat that has not started and does judge a first beat that never returns. The recovery at the top of each work loop cycle stays under the watchdog: it runs only while no task is in flight and every task poll recovers first, so it meets at most the journal of the task that just ended. Rejected: starting the heartbeat before recovery on every start. It fixes the stale beat but not the work loop, which still sits in one unbounded step, and it would send metrics before hello on every start rather than only when recovery is blocked. Rejected: progress callbacks inside the linechain manager; one sing-box restart is bounded by systemd's own start and stop timeouts, and the startup path is the only one that can meet several journals. Tested: go test -race ./cmd/lattice-agent/; new TestLongStartupRecoveryDoesNotTripTheWatchdog fails when beatProgress is stamped at process start (stalled: heartbeat has not moved for 10m30s); TestBlockedStartupRecoveryBeatsBeforeTheWatchdogArms.
…lt the server never stored The monitor result queue is in memory, so the results of the last interval were lost on every restart, and with the watchdog and the update guard a restart is now more likely than before. The drop count on the heartbeat covered queue overflow only: a batch the server refused whole, a single result refused on the old route, and results a batch answer listed as dropped were logged at debug level or not counted at all. On SIGTERM the agent now stops the flush loop, which cancels the flush in flight, and sends the queue once more within 5 s after the task shutdown budget (55 s), staying under systemd's default 90 s TimeoutStopSec. What cannot go out is logged with its count. Every result the server never stored goes into the dropped counter that rides on each heartbeat as loop_health.monitor_results_dropped, and new drops are logged once per flush rather than once per result. A server-listed drop is logged at the normal level with the first reason. Not done here: a disk-backed outbox, which is what survives a crash or a watchdog kill, and a server that shows the dropped count (the server ignores loop_health today). Both stay on the lane's deferred list. Results still wait up to one interval before they are sent; sending each result at once would give up the batching that cuts request volume. Tested: go test -race ./cmd/lattice-agent/; TestMonitorResultsCountEveryResultTheServerNeverStored, TestMonitorDrainSendsTheQueueOnStop, TestMonitorFlushLoopStopsItsFlushInFlight.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Node keepalive for the agent side (r1-keepalive P2 and P4), agent-side monitor result batching against the monitors lane's batch route, and the 0.3.10-alpha.1 bump with a changelog. Not tagged.
Commits:
feat: send the heartbeat on its own goroutine with the work loop's health(81ae17e). The metrics beat runs every interval on its own goroutine with a 10 s deadline, so a slow control plane or a blocked report no longer flips a healthy node offline (four 30 s timeouts ahead of the old in-loop beat already exceeded the 90 s offline flip). Each beat carriesloop_health: per-step last success, last error and consecutive errors, the step in progress, cycle timing, task batch age, and a blocked linechain recovery with its reason. A blocked recovery keeps the node visible and withholdsdurable-task-result-v1. Older servers ignore the field.feat: arm a systemd watchdog on local progress and mark a healthy first hello(a1977d4). With a notify socket (NotifyAccess=main), the agent sendsREADY=1, arms a 120 s watchdog itself withWATCHDOG_USEC=, and pets it only while the work loop and the heartbeat both moved within five minutes (or three intervals). Failing requests are progress; a step that never returns is not. After its first hello it writes its version to$RUNTIME_DIRECTORY/healthy.-compat-jsongainsfeatures: ["sd-notify-v1","health-marker-v1"].feat: grant the agent a notify socket at install and supervise it under openrc(98801f7). The installer writes10-lattice-keepalive.conf(NotifyAccess=main,RuntimeDirectory) only for a binary advertisingsd-notify-v1and removes it otherwise; openrc usessupervise-daemon(10 s respawn delay, unlimited) where present.feat: queue monitor results and send them in batches, one at a time on older servers(bd075a1). Results queue (2000 cap through an outage) and go out in batches of 200 onPOST /api/agent/monitor-results, the route from lattice-serverfeat/monitor-results-hot-store(4cf84b0). 404 falls back to single posts for 30 minutes; 5xx keeps and resends; 400 drops the batch.docs: ...(d62a67e) README keepalive and supervision section.chore: bump to 0.3.10-alpha.1 and start a changelog(f449257).fix: arm the watchdog after startup recovery and judge the heartbeat only once it runs(20b67b9). Review finding: the watchdog was armed before durable linechain recovery and the heartbeat's progress was stamped at process start, so a healthy recovery longer than five minutes read as a stall and systemd killed the agent mid-recovery on every start. Startup recovery now runs to success before the watchdog arms (recoverThenSupervise), and a heartbeat that has not started is not judged.fix: send queued monitor results on a clean stop and count every result the server never stored(5b636ea). Review finding: the in-memory queue lost the last interval on every restart. A clean stop now cancels the flush loop and sends the queue once more within 5 s (after the 55 s task budget, under the default 90 s TimeoutStopSec). Every result the server never stored (overflow, a refused batch or single, a server-listed drop) is counted inloop_health.monitor_results_droppedand logged once per flush. A disk-backed outbox stays a follow-up.Design change from the research: no
Type=notifyand noWatchdogSecin the unit. An older binary installed later under such a unit (a rollback, the v0.3.9 installer through a reconfigure,/install.shmoving an alpha node to stable) never sends READY and is killed in a loop. The drop-in only grants the socket and a runtime directory, and the binary arms its own watchdog, so the drop-in is inert under an old binary. Verified below.Companion PRs: the update guard and the update-time drop-in are server-side, LatticeNet/lattice-server#133; the installer pin is LatticeNet/lattice-server#132 and LatticeNet/latticenet.github.io#4.
Evidence
go test -race -cover ./...: 21 packages ok.go vet ./...(darwin and linux/amd64) clean,gofmt -l .empty,sh scripts/check-release-workflow.shandsh scripts/test-install-integrity.shok (new keepalive contract forbidsType=notifyandWatchdogSec=in the installer).Type=simple,NotifyAccess=main,WatchdogUSec=2min,/run/lattice-agent/healthy= version,loop_healthon every beat, monitor results in batches. SIGSTOP at 19:01:29, thenWatchdog timeout (limit 2min)at 19:02:48, restart at 19:02:59, re-armed. The released v0.3.9 binary under the same drop-in: active,WatchdogUSec=0.supervise-daemonscript, env fromstart_prereaches the agent,kill -9respawned in 10 s and hello again.Review round (2026-10-03), further evidence:
TestLongStartupRecoveryDoesNotTripTheWatchdogfails with the old stamp ("heartbeat has not moved for 10m30s");TestBlockedStartupRecoveryBeatsBeforeTheWatchdogArms; three monitor queue tests for the drop count, the drain, and the cancelled flush.WatchdogUSec=0, three beats carrylinechain_blocked, no hello. Journal removed: "watchdog armed" and "connected" in the same second,WatchdogUSec=2min,NRestarts=0, the next beat says"watchdog":true. The guarded update with the new server script kept the new build after its hello.handleAgentMetricsauthenticates the node token only.go test -race ./...(21 packages), vet darwin and linux/amd64, gofmt,check-release-workflow.sh,test-install-integrity.sh: ok.Test plan
v0.3.10-alpha.1(operator), canary one non-production node through a guarded update, confirmWatchdogUSec=2min, the marker, andloop_healthin the server log or a debug read; then batches.loop_healthbesideagent_runtimeand derive "beating but stalled" (K2), not in this PR.