Skip to content

fix(dns): stop naming DNS resolvers after the next TLS host on a reused fd - #372

Merged
mayankpande88 merged 2 commits into
mainfrom
fix/dns-destination-label
Oct 8, 2026
Merged

mayankpande88 merged 2 commits into
mainfrom
fix/dns-destination-label

Conversation

@mayankpande88

@mayankpande88 mayankpande88 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

DNS series were labelled with the wrong destination. Queries to a resolver came out as destination="<some TLS host>:53", one value for every host the container had recently connected to. Each new value created a new series for every domain, in both the counter and the histogram. On a busy node this pushed /metrics past a scraper's response size limit, and the scraper dropped that node's metrics entirely.

This PR fixes the attribution, so each resolver has one destination again.

Engineering detail

Cause. eBPF passes UDP DNS through with the socket's tuple, and createConnectionFromSocketInfo stored a connection for it in connectionsByPidFd. A UDP socket has no close event, so that entry outlived the socket. The application's next socket usually got the same fd number: the TCP connection to the address it had just resolved. L7 events are handled as they arrive, and connection events later. So the new connection's TLS ClientHello found the stale DNS entry, and the SNI path (which skips the timestamp checks) published ip2fqdn[resolver IP] = <SNI host>. It also renamed the DNS entry's key to that host. This only matters when IsIpExternal(resolver) is true, for example a cluster whose service CIDR lies in a public range. NewDestinationKey then named every later query's destination after the resolver's latest "FQDN".

The --max-fqdns-per-container cap (#371) does not bound this, because the multiplier is the destination, not the domain.

Fix.

  • ActiveConnection.dst records the socket's pre-NAT destination, from the open event or the tuple. isSocket(si) compares it with an L7 event's tuple, which is read from the fd when the event happens. It returns true when there is nothing to compare.
  • DNS queries take an untracked connection built from their own tuple (connectionFromSocketInfo) when there is no tracked entry, or when the entry is a different socket. Those conns now live for one event; the cost is two BPF map lookups per query.
  • The ClientHello path ignores a tracked entry that is not its own socket, and takes the address from the tuple instead.

Tests (containers/socket_connection_test.go, which CI excludes, see ci.yml; run locally in a Linux container):

  • TestDNSResolverKeepsItsNameAcrossReusedFd: a DNS query on fd 7, then a ClientHello on fd 7 to the resolved address, then another query. The resolver keeps its address and has no FQDN.
  • TestEarlierSocketOnFdIsNotUsed: a stale entry on the fd, in each direction.
  • TestConnectionIsSocket: IPv4-mapped, port mismatch, no tuple.

Mutation check: disabling the DNS branch fails 2 tests, and disabling the ClientHello check fails 1.

Review follow-up (9d4e951): isSocket compares the port before it parses the address. Its result is unchanged for any parseable tuple. The e2e below ran on 86a77f8, before this commit; the unit tests cover the reordered check.

CI-equivalent checks. In a Linux container with Go 1.26.5: gofmt, goimports, go vet ./..., golangci-lint v2.13.2 (0 issues) and go test ./... (including /containers) all pass.

Local e2e: in Docker Desktop (kernel 6.10, arm64), I built this branch and main and ran each twice as the agent. The client was a Python container started with --dns 8.8.8.8. Three times over, it resolved 5 public HTTPS hosts and opened a TLS connection to each.

  • main: ip_to_fqdn{ip="8.8.8.8"} was a TLS host (www.cloudflare.com), and the client's DNS series had 5–6 destinations (api.github.com:53, example.com:53, …) on both runs. The published 0.1.9 image behaves the same.
  • This branch: no ip_to_fqdn entry for 8.8.8.8, and every DNS series has destination="8.8.8.8:53". That is 5 series, one per domain, on both runs.
  • Capture was the same on both builds: 14–15 HTTP requests, 15 TCP connects, 13–15 DNS requests, and 0 L7 events dropped. HTTP/TCP destinations were the 5 hostnames on both builds.

…ed fd

A UDP DNS socket has no close event, so the connection built from its
tuple stayed tracked on its pid and fd after the socket closed. The next
socket the application opened usually got the same fd: the connection to
the address it had just resolved. That connection's TLS ClientHello took
the stale entry and recorded the resolver's address under the TLS server
name. A resolver with a non-private address, such as one in a cloud
provider's service range, then went by whichever host had last been
connected to. Every new name minted a new destination for each queried
domain, in the counter and in the 12-bucket histogram, so DNS series
grew with the number of hosts a container connected to.

DNS queries now take their connection from their own socket tuple and
are never tracked. A tracked entry whose destination is not the event's
socket is treated as an earlier socket's, both by DNS queries and by the
ClientHello path.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces changes to prevent stale connection tracking when file descriptors are reused, particularly for DNS over UDP and TLS ClientHello. It adds a dst field to ActiveConnection and introduces an isSocket helper to verify if a tracked connection matches the socket information of an L7 event. Comprehensive unit tests are also added to validate these scenarios. The review feedback suggests two key improvements: optimizing the isSocket check by comparing destination ports first to avoid expensive IP parsing on the hot path, and generalizing the stale connection check across all protocols to prevent incorrect reuse of connections.

Comment thread containers/container.go
Comment thread containers/container.go
@mayankpande88
mayankpande88 merged commit 98e2624 into main Oct 8, 2026
7 checks passed
@mayankpande88
mayankpande88 deleted the fix/dns-destination-label branch October 8, 2026 05:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants