Repository navigation
feat: control-plane witness mode (-witness) and its status on the heartbeat - #32
Merged
Merged
Conversation
…bark-server when it stops answering Lattice cannot report its own outage. internal/witness is the dead-man for that one failure: it polls the control plane's public readiness URL, and after the hold window (3 minutes and at least two checks by default) of failures while a reference URL proves the node's own network, it pushes one message through the bark-server on the node's loopback interface, then one recovery once the control plane has answered without a break for the recovery window. When the control plane and every reference fail together the node's own network is down and nothing is pushed; such a check neither counts toward the hold nor ends the run. The state is written atomically after every check and read back at start, so a restart in the middle of an outage neither pushes twice nor forgets the recovery, and a witness that was itself away for longer than the hold starts counting again. A push the bark-server did not accept stays owed and is retried at every check, logged once. A changed config restarts the runs but keeps an announced alert, so its recovery still goes out. The device key lives only in a root-only file the config names: refused when group or others can read it, when another user owns it, or when it is not shaped like a Bark key; read at start and at every push. The Bark URL must be loopback, a reference must not share the control plane's host, and failures are stored as classified reasons, never raw error text (which can carry URLs and remote text). Standard library only. Rejected: a goroutine in the agent's loop. The agent exits when its first hello fails, so after any restart during an outage an in-process witness would be gone exactly when it is needed. Tested: go test -race ./internal/witness/ (hold, own network down, flapping before the hold and after the alert, restart inside the hold, after the alert and after a long gap, failed push retried once delivered, config change, key file and config validation, HTTP prober and Bark pusher against httptest servers); gofmt; go vet (darwin and linux).
…beat `lattice-agent -witness <config>` runs internal/witness as its own process under its own unit (lattice-witness.service, which an approved server plan writes); it needs no node token and stops on SIGTERM after saving its state. `-witness-check` validates the config and the key file and prints the config's SHA-256 and what it watches, never the key; the apply script runs it before enabling the unit. The agent's heartbeat attaches the witness status file (/var/lib/lattice-witness/status.json, LATTICE_WITNESS_STATUS_FILE moves it) as `witness`, bounded to 16 KiB and re-encoded from the typed state, so the console can show the last check and the last push and no field the witness does not define leaves the node. A missing file sends nothing; a broken one is logged once. The relay reads that file only and never starts, stops or configures the witness. Older servers ignore the field, as they ignore loop_health. Hello capabilities and -compat-json features list control-plane-witness-v1. It is reported even while linechain recovery withholds the durable task capability, because witness mode is a property of the binary and the server reads it before it plans a witness on the node. Tested: go test -race ./cmd/lattice-agent/ (heartbeat relays a valid file, drops a missing, corrupt or oversized one and every undefined field; -witness-check refuses a readable key file and never prints the key; the capability in both lists and in -compat-json); gofmt; go vet (darwin and linux).
README gains a "Control-plane witness" section: why it is its own process and unit, what it talks to (public readyz, references, loopback bark-server) and nothing else, the device key file rules, the push rules (hold, own network down, recovery window, owed pushes), the state file and restart behaviour, the heartbeat relay, -witness-check, and the capability the server plans against. Tested: read through against internal/witness and witness_mode.go.
v0.3.10-alpha.1 is tagged at 5d988c8, so the changelog dates it and the next prerelease constant moves to 0.3.10-alpha.2 with the witness entry. Not tagged: the tag, the prerelease and the fleet roll stay with the operator. Tested: go test -race ./... ; sh scripts/check-release-workflow.sh ; sh scripts/test-install-integrity.sh
The sb caps probe (6812194) and the bounded-call fix (4d11bfe) merged into integration after v0.3.10-alpha.1 without a version bump, so the next tag built from integration with this branch merged is 0.3.10-alpha.2 and ships them too. The alpha.2 entry described only the witness; it now lists both, so the release notes match what the tag contains. Rejected: merging origin/integration into this branch for the note alone. The merge is clean and its result passes the full suite, so the branch does not need it. Tested: a detached merge of this branch with origin/integration f1111e6: gofmt -l . clean, go vet ./... (darwin and linux), go test ./... all ok, scripts/check-release-workflow.sh and scripts/test-install-integrity.sh ok.
Contributor
Author
|
33795b4: the 0.3.10-alpha.2 changelog entry now also lists the sb script caps (6812194) and the bounded-call fix (4d11bfe), which reached integration after alpha.1 without a version bump, so the next tag ships them with the witness. A detached merge of this branch with origin/integration f1111e6 is clean and passes gofmt, go vet (darwin and linux), go test ./..., check-release-workflow.sh and test-install-integrity.sh; the branch is not merged with integration because it does not need to be. |
…lane-witness Integration tagged v0.3.10-alpha.2 at 587d41b with the sb capability report and no witness, so the witness can no longer ship under that number. The changelog takes integration's alpha.2 entry as tagged; the next commit moves this branch to 0.3.10-alpha.3 and gives the witness its own entry. main.go, README.md and the sb caps files merged on their own. Tested: gofmt -l . (clean); go vet ./...; go test -count=1 ./... (all ok).
…d witness can be told apart The agent relays whatever status.json is on disk at every heartbeat. When lattice-witness.service is stopped, wedged or disabled, the file stops changing but keeps being relayed, so the server saw a fresh relay and the console went on saying "Watching". Nothing compared the last check with the witness's interval. The relay now carries relayed_at, this node's clock when the agent read the file, beside the file's own fields (the witness.State is embedded, so the wire shape is unchanged apart from the new field). The witness and the agent share that clock, so relayed_at minus last_check_at is the real age of the last check whatever skew the node has against the server; the server judges staleness from it against the reported interval. Rejected: comparing last_check_at with the server's receive time (node clock skew makes that wrong in both directions); deciding staleness in the agent (the threshold would then ship in a node release and the server could not change it); dropping a stale file from the beat (the console would then say "not reporting", which reads as the agent being down, not the witness). Tested: gofmt -l . (clean); go vet ./cmd/lattice-agent/; go test -run 'Witness|Heartbeat' ./cmd/lattice-agent/ including TestWitnessRelayStampsTheNodeClockBesideAnOldStatus (an hour-old status relayed with relayed_at an hour later, in UTC whatever the node's zone).
…d confirm once after the node's own network returns /readyz answers 503 when the store check or the audit WAL check fails, and a proxy or CDN in front of the server answers with its own status. The witness pushed both as "Lattice control plane unreachable ... has not answered", which sends the operator after the network. When the last failed check got an HTTP answer the push is now titled "Lattice control plane not ready" and says the host "has answered without reporting ready since"; a timeout, refusal or DNS failure keeps "unreachable". The recovery push is "Lattice control plane ready again", which is true after either kind of outage, since it follows a run of 200s. Checks made while the node's own network was down are not counted and do not end a failure run, but the hold is measured on the clock from the run's first failure, so the first failed check after a long own-network outage could push on its own. That check now only confirms; the next failed check with the network up pushes. Restarting the run on leaving network_down was rejected: on a node whose network drops now and then, a real outage would never reach the hold (TestFlakyOwnNetworkStillReachesTheHold pins that). The comment that said network-down time does not count toward the hold now says what the code does. Tested: gofmt -l . (clean); go vet ./internal/witness/; go test -race ./internal/witness/ (new TestDownPushNamesUnreachableOrNotReady, TestFlakyOwnNetworkStillReachesTheHold, updated TestOwnNetworkDownNeitherCountsNorEndsTheRun); the guard mutated to always push turns both network tests red.
…ped witness shows, and that it has no mute The README now says when a push is titled "not ready" rather than "unreachable", that the first failed check after the node's own network returns only confirms, that the heartbeat relay carries relayed_at so the server can show a witness whose service stopped as stopped, and that the witness has no mute: any stop of the control plane longer than the hold pages, and only a longer-hold plan or a remove plan avoids it. Tested: read back the rendered section; no em or en dashes.
v0.3.10-alpha.2 was tagged from integration at 587d41b with the sb capability report and without the witness, so the witness ships as 0.3.10-alpha.3. The version constant and its pinning test move, the changelog dates alpha.2 and gives alpha.3 the witness entry, including what the review changed: the "not ready" title, the confirming check after the node's own network returns, relayed_at on the heartbeat relay, and the absence of a mute. No compatibility floor moves. Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race ./... (all ok); sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok).
…step neither pages early nor holds a page back clock() returned now().UTC(), which drops Go's monotonic reading, and observe measured the gap and both runs as differences of wall times. A forward step mid-run (chrony correcting a slow clock by 130 s) let a 60 s blip pass the 180 s hold and page at critical level; a backward step gave a negative gap that never exceeded maxGap, so the run was not reset and the page came late by the size of the step. The runs now carry their length (failing_for_ns, ok_for_ns, outage_for_ns) summed from the time between checks, and the hold, the recovery window and both push bodies use those. Inside one process that time comes from the monotonic clock. The wall clock still restarts the count when it moved ahead by more than maxGap while the monotonic clock stood still, as it does across a suspend, because restarting can only delay a page. The first check after a start has only the saved wall time, so a negative gap restarts the count like a gap past maxGap. Wall times stay in the state, in UTC, for display. The down push now leads with how long the control plane has been failing and gives the first failure by the node's clock second, labelled as such; the recovery push leads with the outage length. Rejected: keeping the run start times and relying on Go's monotonic reading surviving in them. It cannot be faked in a test, and it is lost on every restart, so a run that straddles a reboot would be measured on the wall clock exactly when chrony steps it. Not-tested: a real suspend; Linux's CLOCK_MONOTONIC stopping across one is taken from its documentation.
On SIGTERM during a check the probes fail because their context is cancelled, so observe recorded this node's network as down and Run saved it; a push cancelled the same way was stored as one the bark-server refused (last_push_ok false with an error kind). Tick now returns as soon as the check finds its context ended, without observing or saving. When the push is the part cut short, the check that asked for it stands and Run saves it on the way out, but the push outcome is not recorded: whether it arrived is unknown, and it is owed again at the first check after the restart.
LoadConfig read the config with no owner or mode check, although the config decides where the witness looks and which file it sends as the device key. It now refuses a config that is not a regular file, belongs to another user than the one the witness runs as, or is writable by group or others. A config readable by others is still accepted: it holds no secret, only the key file's path, and refusing it would turn a harmless chmod into a node without a witness. ReadDeviceKey checked the path with Stat and then read it with ReadFile, so the file read need not be the file checked. Both files now go through one helper that opens once, checks the opened descriptor with Fstat, and reads from it with a bound (64 KiB for the config, 1 KiB for the key). The open uses O_NONBLOCK on unix so a FIFO left at either path is refused at once instead of blocking the witness, or the apply script's -witness-check, until a writer appears. Rejected: refusing a group- or other-readable config as the key file is refused. The plan writes it 0600 either way; the stricter rule buys no secrecy and costs the dead-man on a node where someone loosened it.
… leads with the outage length, and how the config file is checked
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Control-plane witness mode for the agent (keepalive K11, proposal P7; operator decision 2026-10-02: the dead-man lives on the service node that runs bark-server).
internal/witness(standard library only): polls the control plane's public/readyzevery interval (30 s default); only HTTP 200 counts. On a failure it asks one to three reference URLs; any HTTP answer proves the node's own network. After the hold window (3 min and at least two checks) of failures with the network up, it pushes one message through the bark-server on the node's loopback interface (levelcriticalby default), and one recovery (levelactive) once the control plane has answered without a break for the recovery window. Control plane and every reference failing together means the node's own network is down: the check is not counted and does not end the run, nothing is pushed, and the first failed check after the network returns only confirms. A push the bark-server did not accept stays owed and is retried at every check, logged once./var/lib/lattice-witness/status.json, 0644, no secret) and read back at start: a restart in the middle of an outage neither pushes twice nor forgets the recovery; a witness that was itself away longer than the hold starts counting again; a changed config keeps an announced alert.lattice-agent -witness <config>runs it as its own process (unitlattice-witness.service, written by the server's approved plan). Its own process because the agent exits when its first hello fails, so an in-process witness would be gone after any restart during an outage.-witness-checkvalidates config and key and prints the config SHA-256, never the key.witness, bounded and re-encoded from the typed state. Older servers ignore it, as they ignoreloop_health.-compat-jsonfeatures listcontrol-plane-witness-v1, also while linechain recovery withholds the durable capability.Server side: LatticeNet/lattice-server#138. Console: LatticeNet/lattice-dashboard#86.
Evidence
go test -race ./...all ok; gofmt clean;go vetdarwin and linux;sh scripts/check-release-workflow.shexit 0;sh scripts/test-install-integrity.shok.-witness-checknever printing the key, capability in both lists.Test plan
[cd]-gomami-jpn-pulse-nanoonly, confirmlattice-agent -compat-jsonlistscontrol-plane-witness-v1and the node's hello reports it.Review fixes (round 3)
An independent review returned WARNING with one HIGH, one MEDIUM and lows; all addressed here or in the server and console PRs.
relayed_at, this node's clock when the agent read the status file, beside the file's own fields. The witness and the agent share that clock, so the server can tell a witness whose service stopped (the file keeps being relayed butlast_check_atstops moving) from one that is watching, without trusting the node's clock against its own. The server judges it (lattice-server#138) and the console says "Stopped" (lattice-dashboard#86). Test:TestWitnessRelayStampsTheNodeClockBesideAnOldStatus.TestDownPushNamesUnreachableOrNotReady.TestFlakyOwnNetworkStillReachesTheHold). Mutating the guard turns both network tests red.go vet ./...darwin and linux;go test -race ./...all ok;sh scripts/check-release-workflow.shexit 0;sh scripts/test-install-integrity.shok.