Skip to content

feat(llm): accurate LLM usage capture; fix TLS and HTTP/2 parsing - #348

Merged
blue4209211 merged 13 commits into
mainfrom
feat/llm-capture-rework
Oct 2, 2026
Merged

blue4209211 merged 13 commits into
mainfrom
feat/llm-capture-rework

Conversation

@mayankpande88

@mayankpande88 mayankpande88 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Reworks LLM traffic capture so usage numbers are accurate, and fixes the HTTP/2 parsing defects that blocked it.

Why

Measured against an LLM-calling service's own usage metrics, the previous pipeline:

  • reported no token usage for a Go service calling Gemini over HTTP/2;
  • labelled every request's model unknown;
  • saw about a quarter of the requests.

Usage sits at the end of a response body or in the last event of a stream. The generic L7 path keeps one bounded request and the first read of the response, so usage was structurally out of reach.

What changes

TLS single source (eBPF). A TLS connection is seen as plaintext by the uprobes and as ciphertext by the socket syscalls, under the same pid+fd. Once a connection was detected as HTTP/2, its ciphertext went into the same parser. A TLS record header reads as an HTTP/2 frame header, which caused:

  • phantom PUSH_PROMISE, PRIORITY and extension frames;
  • partial frames that swallowed the real ones after them;
  • HPACK desync for the rest of the connection.

Connections handled by a TLS hook are now marked, and their socket-level events are skipped in the kernel (node_agent_l7_tls_ciphertext_skipped_total).

LLM capture channel.

  • Identifying connections: by TLS SNI; by API path for gateways and self-hosted servers; then by destination, so per-request connections to a known endpoint are captured from their first write.
  • Capture: the kernel copies every read and write on these connections into the L7 ring buffer, so they stay ordered with the generic events that preceded the mark.
  • Reassembly: a new llm package rebuilds HTTP/1.1 and HTTP/2 exchanges with net/http and x/net/http2 and decodes gzip, deflate, zstd, SSE and AWS event streams. It handles responses the client stops reading (at SSE [DONE]) and streams cut off by a reset or a closed connection.
  • Usage: typed decoders per wire format: OpenAI and compatible, Anthropic, Gemini/Vertex, Bedrock (including InvokeModel token headers), Cohere. Token types are disjoint (input, cached_input, cache_write, output, reasoning), so they sum to billed tokens.

HTTP/2 fixes found along the way:

  • Raw kernel time was clamped to zero for any uptime over an hour, so HTTP/2 durations were always 0 and the parser's stale-stream GC never ran.
  • After a read truncated by the 4095-byte capture limit, the parser now skips the exact remainder of the cut frame instead of parsing it as frame headers.
  • HTTP/1 and HTTP/2 now share one metric series set. Real HTTP/2 durations exposed a duplicate-series collision that failed the whole scrape. /metrics also continues on gather errors now.

Removed:

  • the DNS/IP LLM detector;
  • the connectionless HTTP/2 path;
  • regex usage extraction and the static price table;
  • payload accumulation in the shared HTTP/2 parser.

Opt-in: --enable-llm-capture

LLM capture is off by default (--enable-llm-capture / ENABLE_LLM_CAPTURE=true), like Node.js and .NET tracing:

  • Off: costs one array lookup per socket read or write in the kernel, and nothing in userspace. The LLM metric families are not registered.
  • On: adds a destination check on each new connection's first write, and full capture and parsing of LLM API connections only.

The TLS ciphertext skip and the HTTP/2 fixes are unconditional: they reduce L7 work for everyone. Deployments that want LLM metrics need the flag set; the Helm chart needs a value for it.

Metric changes (breaking for LLM dashboards)

Before After
container_llm_token_usage_total, container_llm_cached_input_tokens_total container_llm_tokens_total{gen_ai_token_type}
container_llm_errors_total derive from container_llm_requests_total{http_response_status_code}
container_llm_cost_usd_total removed: join tokens with a price list
container_llm_tokens_per_second, container_llm_tool_calls_total, node_agent_llm_sni_tags_total removed
– node_agent_llm_capture_total{outcome}, node_agent_llm_capture_drops_total

docs/llm-observability.md is rewritten to match.

Validation

  • Replay tests record real Go net/http and x/net/http2 client traffic at the TLS boundary (and at the socket for plain HTTP) against local servers with known usage, then assert exact token counts. Each test was checked to fail when its fix is reverted.
  • Deployed from this branch to a test cluster (Linux 6.8, amd64) and measured over 30 minutes with no scrape gaps:
    • against the calling service's own per-pod usage metrics: requests 98.4%, input tokens 99.4%, output tokens 98.4%, cached tokens exact on the busiest model;
    • against an LLM gateway's access log: 98.3% of requests.
    • Cluster-wide: completed external HTTP/2 streams went from 0 to about 2.5k per 30 min, and HPACK errors dropped 55%.
  • Not yet validated: kernels older than 6.8, and arm64.

Reviewing

Suggested order:

  1. ebpftracer/ebpf/l7/l7.c and llm_capture.c (kernel)
  2. llm/ (parsing and usage, tests in replay_test.go)
  3. containers/llm_capture.go (glue)
  4. the ebpftracer/l7/http2.go and containers/l7.go fixes

ebpftracer/ebpf.go is regenerated from the C sources.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a mechanism to skip processing socket-level ciphertext events when a TLS library hook (such as Go crypto/tls or OpenSSL) is already delivering the connection's plaintext. It adds a tls field to connection tracking, updates L7 events to report whether they are TLS plaintext, and introduces a new Prometheus metric node_agent_l7_tls_ciphertext_skipped_total to track skipped ciphertext events. Feedback on the changes points out a high-severity issue in ebpftracer/ebpf/l7/l7.c where mark_tls is called on a stack-allocated connection fallback (conn_on_stack) without persisting it to the active_connections map, which would prevent subsequent ciphertext events from being skipped.

Comment thread ebpftracer/ebpf/l7/l7.c
@mayankpande88
mayankpande88 force-pushed the feat/llm-capture-rework branch from 9806e62 to 1cf0ee0 Compare October 1, 2026 19:03
@mayankpande88

Copy link
Copy Markdown
Contributor Author

/gemini review

A TLS connection is seen twice under the same pid+fd: as plaintext by the
Go crypto/tls and OpenSSL uprobes, and as ciphertext by the read/write
syscalls. Nothing told the two apart, so once plaintext identified a
connection as HTTP/2, its ciphertext took the same fast path into the same
userspace parser. A TLS record header parses as an HTTP/2 frame header whose
type is the high byte of the record length and whose length is 0x170303:
phantom extension and PUSH_PROMISE frames, partial frames that swallow the
real frames after them, and HPACK desynchronisation for the rest of the
connection.

Mark a connection once a TLS hook handles it and drop its socket-level
events in the kernel. A ClientHello on a marked fd starts a new session, so
it resets the mark. Dropped events are counted per direction as
node_agent_l7_tls_ciphertext_skipped_total, events carry is_tls, and
node_agent_l7_events_total gains a tls label.
… stream

LLM usage sits at the end of a response body or in the last event of a
stream. The generic L7 path keeps one bounded request and the first read of
the response, so token counts were structurally out of reach, and the HTTP/2
path never completed a stream.

Connections identified as LLM API connections (by TLS SNI, or by API path for
gateways and self-hosted servers) are now marked in the kernel, which copies
every read and write on them, in order, into the L7 ring buffer. Userspace
splices in the generic events that preceded the mark and rebuilds HTTP/1.1
and HTTP/2 exchanges with net/http and x/net/http2, decodes gzip, deflate,
zstd, server-sent events and AWS event streams, and reads usage with typed
decoders per wire format. Token types are disjoint (input, cached_input,
cache_write, output, reasoning) so they sum to billed tokens.

The DNS/IP detector, the connectionless HTTP/2 path, the regex extractor and
the static price table are removed, as is payload accumulation in the shared
HTTP/2 parser. LLM metrics go from nine families to four plus
node_agent_llm_capture_total, which reports capture completeness.

HTTP/2 events also now carry their raw kernel timestamp separately: it was
passed through a duration clamp that zeroes anything over an hour, so HTTP/2
request durations were always 0 and the parser's stale-stream GC never ran.
Gateways and self-hosted model servers have no hostname to recognise in a
ClientHello, and clients often open a connection per request, so marking a
connection after its first LLM API request captured almost nothing. Once a
request identifies an endpoint, its destination (address and port as the
socket sees them) is recorded in the kernel, and every new connection to it
is marked at its first write. Userspace creates the capture when the first
chunk arrives, and replaces one left over on a reused fd.
Go reads 4096 bytes at a time and the kernel keeps 4095 of them. The parser
dropped the frame cut short, as it must, but the next read then began in the
middle of that frame, and its remainder was parsed as frame headers: phantom
PRIORITY, RST_STREAM and SETTINGS frames, and, whenever a HEADERS frame was
lost, an HPACK table out of step for the rest of the connection.

Parse now takes the number of bytes the kernel did not capture. When they
all belong to the frame that was cut, its remainder at the start of the next
read is known exactly and is skipped.
Clients that open a connection per request were almost never captured: the
API path was only checked once the event reached the HTTP handler, which
requires the connection to be tracked with a matching timestamp, and a
short-lived connection's first event usually outruns its connect event.
Endpoints are now recognised from the request line and Host header before
that lookup, and the destination is marked from the event's own socket
tuple, plain HTTP only.

Closed captures are released on a timer rather than at the next gc, so a
response delimited by the connection closing is finalised promptly.

Also splits destination-marked connections into their own capture outcome
and exports the kernel's capture drop count.
A client that cancels a request sends RST_STREAM and the server never
finishes the stream, so it was held until the connection closed and the
request was never reported. The test harness now snapshots recorded chunks
under its lock, since a cancelled transport keeps reading.
A captured connection closed with a request in flight now counts as
abandoned, and the first such connections are logged with what was seen in
each direction. Plain-HTTP gateway captures are being created but are not
completing, and the aggregates alone cannot say why.

Adds replay tests for plaintext HTTP/1.1 streaming, with and without
keep-alive.
Streaming clients commonly stop at an SSE [DONE] event and close the
connection without reading the rest of a chunked body. The capture sees only
what the application read, so the body ended mid-stream, the reader returned
an unexpected EOF, and the request was dropped as a parse failure, silently,
since the connection was already closed. The part that was read is complete
and carries the usage; it is now reported. Truncated gzip decodes up to the
cut.

Abandoned captures are judged once the parsers have consumed everything,
since such responses complete only when the stream closes.
A stream still responding when its connection went away was dropped. Once
the parsers have consumed the closed connection, streams with a request and
a started response are reported with what arrived.
L7Stats kept a vector per protocol, so HTTP/1 and HTTP/2 each had their own
counter and latency histogram under the same metric names. The histogram has
no label that tells the two apart, so a destination reached over both
produced identical series, Gather returned a duplicate-series error, and the
metrics handler answered the whole scrape with a 500. It never surfaced
before only because HTTP/2 durations were always zero and were therefore
never observed. HTTP/2 is now folded into HTTP before the vectors are looked
up.

The metrics handler also continues on gather errors, so one bad series can
no longer blank a node's entire scrape.
@mayankpande88
mayankpande88 force-pushed the feat/llm-capture-rework branch from 252bd44 to 6ad1869 Compare October 2, 2026 07:06
@mayankpande88 mayankpande88 changed the title LLM capture rework (draft, deployed from branch) feat(llm): accurate LLM usage capture; fix TLS and HTTP/2 parsing Oct 2, 2026
@mayankpande88
mayankpande88 marked this pull request as ready for review October 2, 2026 07:06

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the LLM observability pipeline by replacing the previous DNS-based detection with an eBPF-based connection capture mechanism that copies plaintext reads and writes directly to userspace. It introduces a new llm package to handle stream reassembly, decompression, and token usage extraction for multiple providers, alongside new self-observability metrics and eBPF tracer enhancements to skip socket-level ciphertext events on TLS connections. The review feedback highlights critical concurrency issues in containers/container.go where several capture-related methods are called without holding the required container lock, as well as resource leaks in llm/body.go due to unclosed decompression readers.

Comment thread containers/container.go
Comment thread containers/container.go
Comment thread llm/body.go
blue4209211
blue4209211 previously approved these changes Oct 2, 2026
A TLS hook on a connection with no tracking entry marked only a stack copy,
so the socket-level ciphertext after it was not skipped. The connection is
now tracked from that point. Decompression readers are closed after use.
@blue4209211
blue4209211 merged commit 9317b7c into main Oct 2, 2026
7 of 8 checks passed
@blue4209211
blue4209211 deleted the feat/llm-capture-rework branch October 2, 2026 07:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants