Repository navigation
feat: report trace collector readiness, switch raw lines separately and shed whole connections - #34
Merged
Conversation
…nd shed whole connections Design 26 slice R1, agent side. A node switched on in Collection could sit "on" while collecting nothing, with the reason only in its journal; raw lines had no switch of their own; and the line budget cut lines out of connections the assembler was building. Discovery: when the policy names no Clash API address, the collector reads experimental.clash_api.external_controller from /etc/sing-box/config.json on every applyConfig and accepts it only through singboxapi.ValidateLoopbackAddr, the rule the client already enforces (exported for this). A policy address wins. One parse (readClashAPIConfig) serves discovery and the secret fallback. Status: the collector keeps a model.CollectorStatus and the metrics beat carries it as trace_collector, omitted until a state has been decided. A state change nudges an immediate beat, at most one per second. ready needs an accepted /logs subscription (the new StreamHooks.Open, fired only after a 2xx) and fewer than three consecutive /connections failures. A subscription refused before it ever opened is stream_failing at once; a drop after it was up has a 15 s grace measured on an injectable clock. Hooks carry the subscription epoch so a replaced stream cannot move the state. Every detail goes through model.BoundCollectorDetail and carries error text only, never the secret. Raw lines: the raw source is cleared when policy.raw.enabled is false, and the queue is discarded with it so nothing ships on a later flush. The switch lives on the collector, not on tracepolicy.Set, which local session expiry rebuilds from Enabled and Level only. A nil raw (a117 and older) keeps the old behaviour. Shedding: internal/traceshed decides per parsed line. Over the budget a log id first seen that second is shed for its lifetime; ids the assembler holds or has finished (the new Assembler.Tracked) keep every line. A pre-parse ceiling of four times the budget stays as the CPU brake and counts into dropped. Shed connections ride the batch as shed_connections and are also added to dropped, so a117's gap audit still fires. The default budget is 5,000 parsed lines a second, and 0 means that default. The SDK pin moves to cae46247ddb6, which also crosses a8e48e9..29dacb5 (additive model and proto changes). Rejected: status on every TraceBatch (design 4.2). A node in no_clash_api has no shipper, a status sent only on change is lost on a server restart, and the beat already runs every interval whatever the pipeline does. Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race -cover ./... (all ok); gosec on the module (68 findings, the same 68 as origin/integration); govulncheck ./... (no reachable vulnerabilities)
v0.3.10-alpha.3 was tagged from integration at 89b7cbc with the control-plane witness, so the collector readiness slice ships as 0.3.10-alpha.4. That exact string is the floor lattice-server's traceCollectorStatusMinAgent names for agents that send trace_collector. The version constant and its pinning test move, the changelog dates alpha.3 and gives alpha.4 its entry: Clash API discovery, the collector status on the beat, the raw switch, connection shedding and the 5,000 parsed-line default. No compatibility floor moves. Tested: gofmt -l . (clean); go test -race ./... (all ok); sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok)
… node file reads and pick up a rotated secret Review of the R1 collector. Discovery ran on every poll and tore the pipeline down on any failed read, so a config caught mid-rewrite dropped the connections being assembled. Now only a config that was read and has no usable controller (no experimental.clash_api, no external_controller, or a non-loopback one) stops a running pipeline. A read or parse failure while the pipeline runs from the config keeps it, and after three in a row (traceDiscoveryFailLimit, equal to traceConnPollFailLimit) the state is no_clash_api with a detail saying the stream keeps running. The next good read is ready again. Every config and secret read goes through readNodeFile: O_NONBLOCK and O_NOFOLLOW at open, then fstat on the descriptor. It must be a regular file owned by root or the agent's uid with no group or other write bit; a secret file also has no group or other read bit. Reads go through io.LimitReader at 1 MiB. A symlink, FIFO, directory, oversized or loosely permitted file is refused with a fixed detail naming the path. The fallback secret path is a collector field so tests never read the machine's own. A fallback file that is missing or not permitted to the agent still falls through to the config, as before. A JSON error is reported as "<path> cannot be parsed as JSON", because the decoder's message quotes bytes of a file that holds the secret. A rotated secret is picked up without a restart. The collector keeps a SHA-256 of the secret in use and re-reads it on each poll. When it changed and the /logs stream is not open, the client takes it (new singboxapi.Client.SetSecret) and the stream resubscribes at once, keeping the assembler. An open stream keeps its secret, because the running core still holds it until it restarts. The /logs subscription gets a first-answer deadline (10 s, StreamOpenTimeout) on a child context, so a request the core accepts and never answers reaches the Error hook and cannot leave ready standing after a resubscribe. The deadline is stopped once the headers arrive, so an open stream stays unbounded. LinesPerSec now counts every line that arrives, before the pre-parse ceiling, so the rate R0 measures is not capped at four times the budget. Lines over the ceiling are still counted in dropped. Rejected: rebuilding the whole pipeline on a secret change. That would drain every open connection for what is only a header change. Tested: gofmt -l . (clean); go vet ./... (darwin and linux); go test -race -count=1 ./... (23 ok); each new test fails with its fix disabled; sh scripts/check-release-workflow.sh (exit 0); sh scripts/test-install-integrity.sh (ok)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Design 26 slice R1, agent side (plan section 1.2, A0 to A5). Two commits: the feature, then the bump to 0.3.10-alpha.4, the exact string lattice-server's
traceCollectorStatusMinAgentfloor names.What changes
The trace collector finds its Clash API without being told: when the policy has no address, it reads
experimental.clash_api.external_controllerfrom/etc/sing-box/config.jsonon every poll and accepts it only throughsingboxapi.ValidateLoopbackAddr(exported; the same rule the client enforces). A policy address wins.The collector keeps a
model.CollectorStatusand the metrics beat carries it astrace_collector, omitted until the collector has applied a policy. A state change nudges an immediate beat, at most one a second. States sent:off,no_clash_api,secret_unreadable,stream_failing,ready; neveragent_too_old. Every detail goes throughmodel.BoundCollectorDetailand never carries the secret.readyneeds a/logssubscription that returned 2xx (newStreamHooks.Open) and fewer than three consecutive/connectionsfailures; a drop after it was up has a 15 s grace.Raw lines follow
policy.raw.enabledwhen the server sends it: off clears the raw source and discards the queued raw lines. A nilraw(a117 and older) keeps today's behaviour. The switch lives on the collector, not ontracepolicy.Set.internal/traceshedsheds whole new connections over the per-second parsed-line budget and never cuts one the assembler holds or has finished (Assembler.Tracked). A pre-parse ceiling at four times the budget stays as the CPU brake. Shed connections ride the batch asshed_connectionsand are also counted indropped, so a117's gap audit still fires. Default budget 5,000; 0 means the default.SDK pin:
v0.2.24-0.20261005113750-cae46247ddb6(lattice-sdk integration mergecae4624), which also crossesa8e48e9..29dacb5(additive).Deviation from design 26 section 4.2, accepted in plan section 4 question 1: status rides the metrics beat, not every
TraceBatch.Test plan
gofmt -l .empty;go vet ./...on darwin andGOOS=linuxgo test -race -count=1 ./...: 23 packages ok, 0 FAILsh scripts/check-release-workflow.shexit 0;sh scripts/test-install-integrity.sh"install integrity contract ok"1a4e65d(alpha-0.2.2a117) through a logging proxy, with a fake Clash API printing one whole connection every 300 ms:trace_collectoranswered 200 (stateoff, thenno_clash_apiwith detail "cannot read /etc/sing-box/config.json: no such file or directory", thenoff, thenready)/api/agent/traceand zero/api/agent/logsrequests{enabled: true, raw: {enabled: false}}(raw added to a117's trace-config answer by the proxy, a117'sraw_source_idleft in place): 27/api/agent/traceposts, all 200, 106 records, 0 dropped, zero/api/agent/logsover about 27 flush cycles/api/agent/logsposts with 318 lines, all 200,raw_lines: trueReview fixes (f4905b6)
no_clash_api("the stream to ... keeps running"). Only a parsed config without a usable controller stops it.O_NONBLOCK|O_NOFOLLOW, fstat on the descriptor, regular file owned by root or the agent uid, no group/other write (secret: no group/other read either), 1 MiB cap./logsis down and the stream resubscribes, with the same assembler./logshas a 10 s first-answer deadline.lines_per_seccounts lines before the pre-parse ceiling.