Skip to content

feat: control-plane witness mode (-witness) and its status on the heartbeat - #32

Merged
lr00rl merged 14 commits into
integrationfrom
feat/control-plane-witness
Oct 4, 2026
Merged

lr00rl merged 14 commits into
integrationfrom
feat/control-plane-witness

Conversation

@lr00rl

@lr00rl lr00rl commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Control-plane witness mode for the agent (keepalive K11, proposal P7; operator decision 2026-10-02: the dead-man lives on the service node that runs bark-server).

  • internal/witness (standard library only): polls the control plane's public /readyz every interval (30 s default); only HTTP 200 counts. On a failure it asks one to three reference URLs; any HTTP answer proves the node's own network. After the hold window (3 min and at least two checks) of failures with the network up, it pushes one message through the bark-server on the node's loopback interface (level critical by default), and one recovery (level active) once the control plane has answered without a break for the recovery window. Control plane and every reference failing together means the node's own network is down: the check is not counted and does not end the run, nothing is pushed, and the first failed check after the network returns only confirms. A push the bark-server did not accept stays owed and is retried at every check, logged once.
  • State is written atomically after every check (/var/lib/lattice-witness/status.json, 0644, no secret) and read back at start: a restart in the middle of an outage neither pushes twice nor forgets the recovery; a witness that was itself away longer than the hold starts counting again; a changed config keeps an announced alert.
  • The device key lives only in a root-only file the config names: refused when group or others can read it, when another user owns it, or when it is not shaped like a Bark key. The Bark URL must be loopback; a reference must not share the control plane's host; failures are stored as classified reasons, never raw error text.
  • lattice-agent -witness <config> runs it as its own process (unit lattice-witness.service, written by the server's approved plan). Its own process because the agent exits when its first hello fails, so an in-process witness would be gone after any restart during an outage. -witness-check validates config and key and prints the config SHA-256, never the key.
  • The heartbeat relays the status file as witness, bounded and re-encoded from the typed state. Older servers ignore it, as they ignore loop_health.
  • Hello capabilities and -compat-json features list control-plane-witness-v1, also while linechain recovery withholds the durable capability.
  • Version 0.3.10-alpha.3 and changelog. v0.3.10-alpha.2 was tagged from integration at 587d41b with the sb capability report and without the witness, so this branch merged integration (94ad4e3) and moved to alpha.3 (7e90a81). Not tagged.

Server side: LatticeNet/lattice-server#138. Console: LatticeNet/lattice-dashboard#86.

Evidence

  • go test -race ./... all ok; gofmt clean; go vet darwin and linux; sh scripts/check-release-workflow.sh exit 0; sh scripts/test-install-integrity.sh ok.
  • Unit tests: hold, own network down (neither counts nor ends the run), network down while alerted, flapping before the hold and after the alert, restart inside the hold, after the alert and after a long gap, failed push retried and delivered once, recovery before an undelivered alert, config change, config validation, key file rules, state round trip, HTTP prober against httptest, Bark pusher posting the key from its file, Start refusing an unsafe key; heartbeat relay (valid, missing, corrupt, oversized, undefined fields dropped), -witness-check never printing the key, capability in both lists.
  • Container drill (Ubuntu 22.04, systemd 249) with this branch's binary and the server-rendered apply script, against a fake control plane, reference and bark-server (interval 15 s, hold 30 s, recover 15 s): a 503 pushed once after 30 s; a restart mid-outage pushed nothing more and the recovery still went out; 70 s with control plane and reference both silent pushed nothing; reference back with the control plane still silent pushed after the hold; a flap while down sent nothing; the stable recovery went out.

Test plan

  • CI green (gofmt, release workflow contract, install integrity, vet, race, gosec, govulncheck).
  • Tag v0.3.10-alpha.3 as a prerelease (operator), roll it to [cd]-gomami-jpn-pulse-nano only, confirm lattice-agent -compat-json lists control-plane-witness-v1 and the node's hello reports it.
  • With the server change deployed, file the witness plan from Notifications, approve it, and run the drill on a precheck copy (steps in the lane log).

Review fixes (round 3)

An independent review returned WARNING with one HIGH, one MEDIUM and lows; all addressed here or in the server and console PRs.

  • HIGH, a dead witness never flagged (904dba5): the heartbeat relay now carries relayed_at, this node's clock when the agent read the status file, beside the file's own fields. The witness and the agent share that clock, so the server can tell a witness whose service stopped (the file keeps being relayed but last_check_at stops moving) from one that is watching, without trusting the node's clock against its own. The server judges it (lattice-server#138) and the console says "Stopped" (lattice-dashboard#86). Test: TestWitnessRelayStampsTheNodeClockBesideAnOldStatus.
  • MEDIUM, a 503 labelled "unreachable" (b230d43): when the last failed check got an HTTP answer the push is titled "Lattice control plane not ready" and says the host "has answered without reporting ready since"; a timeout, refusal or DNS failure keeps "unreachable". The recovery is "Lattice control plane ready again". README (7cc29b5) now says the witness has no mute: any control plane stop longer than the hold pages, and only a longer-hold plan or a remove plan avoids it. Test: TestDownPushNamesUnreachableOrNotReady.
  • LOW, own-network-down time counted toward the hold (b230d43): the first failed check after the node's own network returns no longer pushes on its own; the next one does. Restarting the run was rejected because a node whose network drops now and then would never reach the hold (TestFlakyOwnNetworkStillReachesTheHold). Mutating the guard turns both network tests red.
  • Checks on 7e90a81: gofmt clean; go vet ./... darwin and linux; go test -race ./... all ok; sh scripts/check-release-workflow.sh exit 0; sh scripts/test-install-integrity.sh ok.

lr00rl added 4 commits October 3, 2026 04:59
…bark-server when it stops answering

Lattice cannot report its own outage. internal/witness is the dead-man for
that one failure: it polls the control plane's public readiness URL, and
after the hold window (3 minutes and at least two checks by default) of
failures while a reference URL proves the node's own network, it pushes one
message through the bark-server on the node's loopback interface, then one
recovery once the control plane has answered without a break for the
recovery window. When the control plane and every reference fail together
the node's own network is down and nothing is pushed; such a check neither
counts toward the hold nor ends the run.

The state is written atomically after every check and read back at start,
so a restart in the middle of an outage neither pushes twice nor forgets the
recovery, and a witness that was itself away for longer than the hold starts
counting again. A push the bark-server did not accept stays owed and is
retried at every check, logged once. A changed config restarts the runs but
keeps an announced alert, so its recovery still goes out.

The device key lives only in a root-only file the config names: refused
when group or others can read it, when another user owns it, or when it is
not shaped like a Bark key; read at start and at every push. The Bark URL
must be loopback, a reference must not share the control plane's host, and
failures are stored as classified reasons, never raw error text (which can
carry URLs and remote text). Standard library only.

Rejected: a goroutine in the agent's loop. The agent exits when its first
hello fails, so after any restart during an outage an in-process witness
would be gone exactly when it is needed.

Tested: go test -race ./internal/witness/ (hold, own network down, flapping
before the hold and after the alert, restart inside the hold, after the
alert and after a long gap, failed push retried once delivered, config
change, key file and config validation, HTTP prober and Bark pusher against
httptest servers); gofmt; go vet (darwin and linux).
…beat

`lattice-agent -witness <config>` runs internal/witness as its own process
under its own unit (lattice-witness.service, which an approved server plan
writes); it needs no node token and stops on SIGTERM after saving its state.
`-witness-check` validates the config and the key file and prints the
config's SHA-256 and what it watches, never the key; the apply script runs
it before enabling the unit.

The agent's heartbeat attaches the witness status file
(/var/lib/lattice-witness/status.json, LATTICE_WITNESS_STATUS_FILE moves it)
as `witness`, bounded to 16 KiB and re-encoded from the typed state, so the
console can show the last check and the last push and no field the witness
does not define leaves the node. A missing file sends nothing; a broken one
is logged once. The relay reads that file only and never starts, stops or
configures the witness. Older servers ignore the field, as they ignore
loop_health.

Hello capabilities and -compat-json features list control-plane-witness-v1.
It is reported even while linechain recovery withholds the durable task
capability, because witness mode is a property of the binary and the server
reads it before it plans a witness on the node.

Tested: go test -race ./cmd/lattice-agent/ (heartbeat relays a valid file,
drops a missing, corrupt or oversized one and every undefined field;
-witness-check refuses a readable key file and never prints the key; the
capability in both lists and in -compat-json); gofmt; go vet (darwin and
linux).
README gains a "Control-plane witness" section: why it is its own process
and unit, what it talks to (public readyz, references, loopback
bark-server) and nothing else, the device key file rules, the push rules
(hold, own network down, recovery window, owed pushes), the state file and
restart behaviour, the heartbeat relay, -witness-check, and the capability
the server plans against.

Tested: read through against internal/witness and witness_mode.go.
v0.3.10-alpha.1 is tagged at 5d988c8, so the changelog dates it and the
next prerelease constant moves to 0.3.10-alpha.2 with the witness entry.
Not tagged: the tag, the prerelease and the fleet roll stay with the
operator.

Tested: go test -race ./... ; sh scripts/check-release-workflow.sh ; sh
scripts/test-install-integrity.sh
The sb caps probe (6812194) and the bounded-call fix (4d11bfe) merged into
integration after v0.3.10-alpha.1 without a version bump, so the next tag
built from integration with this branch merged is 0.3.10-alpha.2 and ships
them too. The alpha.2 entry described only the witness; it now lists both,
so the release notes match what the tag contains.

Rejected: merging origin/integration into this branch for the note alone.
The merge is clean and its result passes the full suite, so the branch
does not need it.

Tested: a detached merge of this branch with origin/integration f1111e6:
gofmt -l . clean, go vet ./... (darwin and linux), go test ./... all ok,
scripts/check-release-workflow.sh and scripts/test-install-integrity.sh ok.
@lr00rl

lr00rl commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

33795b4: the 0.3.10-alpha.2 changelog entry now also lists the sb script caps (6812194) and the bounded-call fix (4d11bfe), which reached integration after alpha.1 without a version bump, so the next tag ships them with the witness. A detached merge of this branch with origin/integration f1111e6 is clean and passes gofmt, go vet (darwin and linux), go test ./..., check-release-workflow.sh and test-install-integrity.sh; the branch is not merged with integration because it does not need to be.

lr00rl added 9 commits October 3, 2026 07:37
…lane-witness

Integration tagged v0.3.10-alpha.2 at 587d41b with the sb capability report and no witness, so the witness can no longer ship under that number. The changelog takes integration's alpha.2 entry as tagged; the next commit moves this branch to 0.3.10-alpha.3 and gives the witness its own entry. main.go, README.md and the sb caps files merged on their own.

Tested: gofmt -l . (clean); go vet ./...; go test -count=1 ./... (all ok).
…d witness can be told apart

The agent relays whatever status.json is on disk at every heartbeat. When lattice-witness.service is stopped, wedged or disabled, the file stops changing but keeps being relayed, so the server saw a fresh relay and the console went on saying "Watching". Nothing compared the last check with the witness's interval.

The relay now carries relayed_at, this node's clock when the agent read the file, beside the file's own fields (the witness.State is embedded, so the wire shape is unchanged apart from the new field). The witness and the agent share that clock, so relayed_at minus last_check_at is the real age of the last check whatever skew the node has against the server; the server judges staleness from it against the reported interval.

Rejected: comparing last_check_at with the server's receive time (node clock skew makes that wrong in both directions); deciding staleness in the agent (the threshold would then ship in a node release and the server could not change it); dropping a stale file from the beat (the console would then say "not reporting", which reads as the agent being down, not the witness).

Tested: gofmt -l . (clean); go vet ./cmd/lattice-agent/; go test -run 'Witness|Heartbeat' ./cmd/lattice-agent/ including TestWitnessRelayStampsTheNodeClockBesideAnOldStatus (an hour-old status relayed with relayed_at an hour later, in UTC whatever the node's zone).
…d confirm once after the node's own network returns

/readyz answers 503 when the store check or the audit WAL check fails, and a proxy or CDN in front of the server answers with its own status. The witness pushed both as "Lattice control plane unreachable ... has not answered", which sends the operator after the network. When the last failed check got an HTTP answer the push is now titled "Lattice control plane not ready" and says the host "has answered without reporting ready since"; a timeout, refusal or DNS failure keeps "unreachable". The recovery push is "Lattice control plane ready again", which is true after either kind of outage, since it follows a run of 200s.

Checks made while the node's own network was down are not counted and do not end a failure run, but the hold is measured on the clock from the run's first failure, so the first failed check after a long own-network outage could push on its own. That check now only confirms; the next failed check with the network up pushes. Restarting the run on leaving network_down was rejected: on a node whose network drops now and then, a real outage would never reach the hold (TestFlakyOwnNetworkStillReachesTheHold pins that). The comment that said network-down time does not count toward the hold now says what the code does.

Tested: gofmt -l . (clean); go vet ./internal/witness/; go test -race ./internal/witness/ (new TestDownPushNamesUnreachableOrNotReady, TestFlakyOwnNetworkStillReachesTheHold, updated TestOwnNetworkDownNeitherCountsNorEndsTheRun); the guard mutated to always push turns both network tests red.
…ped witness shows, and that it has no mute

The README now says when a push is titled "not ready" rather than "unreachable", that the first failed check after the node's own network returns only confirms, that the heartbeat relay carries relayed_at so the server can show a witness whose service stopped as stopped, and that the witness has no mute: any stop of the control plane longer than the hold pages, and only a longer-hold plan or a remove plan avoids it.

Tested: read back the rendered section; no em or en dashes.
v0.3.10-alpha.2 was tagged from integration at 587d41b with the sb capability report and without the witness, so the witness ships as 0.3.10-alpha.3. The version constant and its pinning test move, the changelog dates alpha.2 and gives alpha.3 the witness entry, including what the review changed: the "not ready" title, the confirming check after the node's own network returns, relayed_at on the heartbeat relay, and the absence of a mute. No compatibility floor moves.

Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race ./... (all ok); sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok).
…step neither pages early nor holds a page back

clock() returned now().UTC(), which drops Go's monotonic reading, and
observe measured the gap and both runs as differences of wall times. A
forward step mid-run (chrony correcting a slow clock by 130 s) let a 60 s
blip pass the 180 s hold and page at critical level; a backward step gave a
negative gap that never exceeded maxGap, so the run was not reset and the
page came late by the size of the step.

The runs now carry their length (failing_for_ns, ok_for_ns, outage_for_ns)
summed from the time between checks, and the hold, the recovery window and
both push bodies use those. Inside one process that time comes from the
monotonic clock. The wall clock still restarts the count when it moved
ahead by more than maxGap while the monotonic clock stood still, as it does
across a suspend, because restarting can only delay a page. The first check
after a start has only the saved wall time, so a negative gap restarts the
count like a gap past maxGap. Wall times stay in the state, in UTC, for
display.

The down push now leads with how long the control plane has been failing
and gives the first failure by the node's clock second, labelled as such;
the recovery push leads with the outage length.

Rejected: keeping the run start times and relying on Go's monotonic reading
surviving in them. It cannot be faked in a test, and it is lost on every
restart, so a run that straddles a reboot would be measured on the wall
clock exactly when chrony steps it.
Not-tested: a real suspend; Linux's CLOCK_MONOTONIC stopping across one is
taken from its documentation.
On SIGTERM during a check the probes fail because their context is
cancelled, so observe recorded this node's network as down and Run saved
it; a push cancelled the same way was stored as one the bark-server
refused (last_push_ok false with an error kind).

Tick now returns as soon as the check finds its context ended, without
observing or saving. When the push is the part cut short, the check that
asked for it stands and Run saves it on the way out, but the push outcome
is not recorded: whether it arrived is unknown, and it is owed again at the
first check after the restart.
LoadConfig read the config with no owner or mode check, although the
config decides where the witness looks and which file it sends as the
device key. It now refuses a config that is not a regular file, belongs to
another user than the one the witness runs as, or is writable by group or
others. A config readable by others is still accepted: it holds no secret,
only the key file's path, and refusing it would turn a harmless chmod into
a node without a witness.

ReadDeviceKey checked the path with Stat and then read it with ReadFile, so
the file read need not be the file checked. Both files now go through one
helper that opens once, checks the opened descriptor with Fstat, and reads
from it with a bound (64 KiB for the config, 1 KiB for the key). The open
uses O_NONBLOCK on unix so a FIFO left at either path is refused at once
instead of blocking the witness, or the apply script's -witness-check,
until a writer appears.

Rejected: refusing a group- or other-readable config as the key file is
refused. The plan writes it 0600 either way; the stricter rule buys no
secrecy and costs the dead-man on a node where someone loosened it.
… leads with the outage length, and how the config file is checked
@lr00rl
lr00rl merged commit 89b7cbc into integration Oct 4, 2026
1 check passed
@lr00rl
lr00rl deleted the feat/control-plane-witness branch October 4, 2026 12:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant