Skip to content

feat: report trace collector readiness, switch raw lines separately and shed whole connections - #34

Merged
lr00rl merged 3 commits into
integrationfrom
feature/20261005_d26-r1-collector-status
Oct 5, 2026
Merged

lr00rl merged 3 commits into
integrationfrom
feature/20261005_d26-r1-collector-status

Conversation

@lr00rl

@lr00rl lr00rl commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Design 26 slice R1, agent side (plan section 1.2, A0 to A5). Two commits: the feature, then the bump to 0.3.10-alpha.4, the exact string lattice-server's traceCollectorStatusMinAgent floor names.

What changes

The trace collector finds its Clash API without being told: when the policy has no address, it reads experimental.clash_api.external_controller from /etc/sing-box/config.json on every poll and accepts it only through singboxapi.ValidateLoopbackAddr (exported; the same rule the client enforces). A policy address wins.

The collector keeps a model.CollectorStatus and the metrics beat carries it as trace_collector, omitted until the collector has applied a policy. A state change nudges an immediate beat, at most one a second. States sent: off, no_clash_api, secret_unreadable, stream_failing, ready; never agent_too_old. Every detail goes through model.BoundCollectorDetail and never carries the secret. ready needs a /logs subscription that returned 2xx (new StreamHooks.Open) and fewer than three consecutive /connections failures; a drop after it was up has a 15 s grace.

Raw lines follow policy.raw.enabled when the server sends it: off clears the raw source and discards the queued raw lines. A nil raw (a117 and older) keeps today's behaviour. The switch lives on the collector, not on tracepolicy.Set.

internal/traceshed sheds whole new connections over the per-second parsed-line budget and never cuts one the assembler holds or has finished (Assembler.Tracked). A pre-parse ceiling at four times the budget stays as the CPU brake. Shed connections ride the batch as shed_connections and are also counted in dropped, so a117's gap audit still fires. Default budget 5,000; 0 means the default.

SDK pin: v0.2.24-0.20261005113750-cae46247ddb6 (lattice-sdk integration merge cae4624), which also crosses a8e48e9..29dacb5 (additive).

Deviation from design 26 section 4.2, accepted in plan section 4 question 1: status rides the metrics beat, not every TraceBatch.

Test plan

  • gofmt -l . empty; go vet ./... on darwin and GOOS=linux
  • go test -race -count=1 ./...: 23 packages ok, 0 FAIL
  • sh scripts/check-release-workflow.sh exit 0; sh scripts/test-install-integrity.sh "install integrity contract ok"
  • Every A-test named in plan 1.2 is present and passes (tracesource, keepalive, traceshed, sessionasm, singboxapi, traceship)
  • Compatibility against a local build of lattice-server 1a4e65d (alpha-0.2.2a117) through a logging proxy, with a fake Clash API printing one whole connection every 300 ms:
    • beats carrying trace_collector answered 200 (state off, then no_clash_api with detail "cannot read /etc/sing-box/config.json: no such file or directory", then off, then ready)
    • status-only changes (off to no_clash_api at 12:32:57Z, beat at :58; back to off at 12:33:17Z, beat at :18) produced zero /api/agent/trace and zero /api/agent/logs requests
    • policy {enabled: true, raw: {enabled: false}} (raw added to a117's trace-config answer by the proxy, a117's raw_source_id left in place): 27 /api/agent/trace posts, all 200, 106 records, 0 dropped, zero /api/agent/logs over about 27 flush cycles
    • control, native a117 policy (nil raw): 16 /api/agent/logs posts with 318 lines, all 200, raw_lines: true
  • CI green on this PR

Review fixes (f4905b6)

  • A config read or parse failure while recording keeps the pipeline and its open connections; three in a row report no_clash_api ("the stream to ... keeps running"). Only a parsed config without a usable controller stops it.
  • Config and secret reads: O_NONBLOCK|O_NOFOLLOW, fstat on the descriptor, regular file owned by root or the agent uid, no group/other write (secret: no group/other read either), 1 MiB cap.
  • A rotated secret is taken while /logs is down and the stream resubscribes, with the same assembler.
  • JSON errors read " cannot be parsed as JSON".
  • /logs has a 10 s first-answer deadline.
  • lines_per_sec counts lines before the pre-parse ceiling.
  • Each new test fails with its fix disabled; full gates rerun green.

lr00rl added 3 commits October 5, 2026 05:04
…nd shed whole connections

Design 26 slice R1, agent side. A node switched on in Collection could sit
"on" while collecting nothing, with the reason only in its journal; raw
lines had no switch of their own; and the line budget cut lines out of
connections the assembler was building.

Discovery: when the policy names no Clash API address, the collector reads
experimental.clash_api.external_controller from /etc/sing-box/config.json on
every applyConfig and accepts it only through singboxapi.ValidateLoopbackAddr,
the rule the client already enforces (exported for this). A policy address
wins. One parse (readClashAPIConfig) serves discovery and the secret fallback.

Status: the collector keeps a model.CollectorStatus and the metrics beat
carries it as trace_collector, omitted until a state has been decided. A
state change nudges an immediate beat, at most one per second. ready needs
an accepted /logs subscription (the new StreamHooks.Open, fired only after a
2xx) and fewer than three consecutive /connections failures. A subscription
refused before it ever opened is stream_failing at once; a drop after it was
up has a 15 s grace measured on an injectable clock. Hooks carry the
subscription epoch so a replaced stream cannot move the state. Every detail
goes through model.BoundCollectorDetail and carries error text only, never
the secret.

Raw lines: the raw source is cleared when policy.raw.enabled is false, and
the queue is discarded with it so nothing ships on a later flush. The switch
lives on the collector, not on tracepolicy.Set, which local session expiry
rebuilds from Enabled and Level only. A nil raw (a117 and older) keeps the
old behaviour.

Shedding: internal/traceshed decides per parsed line. Over the budget a log
id first seen that second is shed for its lifetime; ids the assembler holds
or has finished (the new Assembler.Tracked) keep every line. A pre-parse
ceiling of four times the budget stays as the CPU brake and counts into
dropped. Shed connections ride the batch as shed_connections and are also
added to dropped, so a117's gap audit still fires. The default budget is
5,000 parsed lines a second, and 0 means that default.

The SDK pin moves to cae46247ddb6, which also crosses a8e48e9..29dacb5
(additive model and proto changes).

Rejected: status on every TraceBatch (design 4.2). A node in no_clash_api has
no shipper, a status sent only on change is lost on a server restart, and the
beat already runs every interval whatever the pipeline does.
Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race -cover ./... (all ok); gosec on the module (68 findings, the same 68 as origin/integration); govulncheck ./... (no reachable vulnerabilities)
v0.3.10-alpha.3 was tagged from integration at 89b7cbc with the
control-plane witness, so the collector readiness slice ships as
0.3.10-alpha.4. That exact string is the floor lattice-server's
traceCollectorStatusMinAgent names for agents that send trace_collector. The
version constant and its pinning test move, the changelog dates alpha.3 and
gives alpha.4 its entry: Clash API discovery, the collector status on the
beat, the raw switch, connection shedding and the 5,000 parsed-line default.
No compatibility floor moves.

Tested: gofmt -l . (clean); go test -race ./... (all ok); sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok)
… node file reads and pick up a rotated secret

Review of the R1 collector. Discovery ran on every poll and tore the pipeline
down on any failed read, so a config caught mid-rewrite dropped the
connections being assembled. Now only a config that was read and has no usable
controller (no experimental.clash_api, no external_controller, or a
non-loopback one) stops a running pipeline. A read or parse failure while the
pipeline runs from the config keeps it, and after three in a row
(traceDiscoveryFailLimit, equal to traceConnPollFailLimit) the state is
no_clash_api with a detail saying the stream keeps running. The next good read
is ready again.

Every config and secret read goes through readNodeFile: O_NONBLOCK and
O_NOFOLLOW at open, then fstat on the descriptor. It must be a regular file
owned by root or the agent's uid with no group or other write bit; a secret
file also has no group or other read bit. Reads go through io.LimitReader at
1 MiB. A symlink, FIFO, directory, oversized or loosely permitted file is
refused with a fixed detail naming the path. The fallback secret path is a
collector field so tests never read the machine's own. A fallback file that is
missing or not permitted to the agent still falls through to the config, as
before. A JSON error is reported as "<path> cannot be parsed as JSON", because
the decoder's message quotes bytes of a file that holds the secret.

A rotated secret is picked up without a restart. The collector keeps a
SHA-256 of the secret in use and re-reads it on each poll. When it changed
and the /logs stream is not open, the client takes it (new
singboxapi.Client.SetSecret) and the stream resubscribes at once, keeping the
assembler. An open stream keeps its secret, because the running core still
holds it until it restarts.

The /logs subscription gets a first-answer deadline (10 s, StreamOpenTimeout)
on a child context, so a request the core accepts and never answers reaches
the Error hook and cannot leave ready standing after a resubscribe. The
deadline is stopped once the headers arrive, so an open stream stays
unbounded.

LinesPerSec now counts every line that arrives, before the pre-parse ceiling,
so the rate R0 measures is not capped at four times the budget. Lines over the
ceiling are still counted in dropped.

Rejected: rebuilding the whole pipeline on a secret change. That would drain
every open connection for what is only a header change.
Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race -count=1 ./... (23 ok); each new test fails with its fix disabled; sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok)
@lr00rl
lr00rl merged commit 9916b9a into integration Oct 5, 2026
1 check passed
@lr00rl
lr00rl deleted the feature/20261005_d26-r1-collector-status branch October 5, 2026 12:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant