Skip to content

mdns: don't stop a unicast scan before results stabilize - #2911

Open
forsethc wants to merge 1 commit into
postlund:masterfrom
forsethc:fix/unicast-scan-dynamic-port-race
Open

forsethc wants to merge 1 commit into
postlund:masterfrom
forsethc:fix/unicast-scan-dynamic-port-race

Conversation

@forsethc

@forsethc forsethc commented Sep 13, 2026 •

Copy link
Copy Markdown

Problem

UnicastDnsSdClientProtocol.datagram_received finishes a scan as soon as it has received one response packet per query sent, regardless of whether all requested services actually appear in those packets. A device that's slow to re-register a dynamic-port service with its own mDNS responder (e.g. Companion right after boot/wake) can answer for everything else in round 1, causing the scan to close the socket immediately — before that service ever gets a chance to show up in a later round. The service isn't just "found late," it's never discovered for that scan at all.

Fix

Track whether the set of discovered services is still growing round-to-round, and only finish early when:

  • every explicitly-requested service has answered, or
  • resend rounds are exhausted, or
  • UNICAST_MIN_STABLE_ROUNDS (2) consecutive rounds have passed with no new services appearing.

An empty round never counts as "stable" on its own — that's just as likely to mean the device isn't ready yet as it is to mean nothing's coming. The common case (everything answers immediately) is unaffected — no added latency there.

Tests

Three new functional tests:

  • a service that only starts answering in round 2 is still caught
  • two consecutive empty rounds don't get mistaken for stability
  • one service stabilizing doesn't short-circuit discovery while another requested service is still pending

Context

Found this while chasing an intermittent bug in a small third-party bridge app talking to an Apple TV over Companion: on some connections Companion's power-state protocol was silently never active, and it traced back to this scan race dropping Companion during discovery. Fix has been running in production there for a few days with no repro of the original issue.

Possibly related (same mdns.py scan/discovery area, similar symptoms) but not something this PR claims to fix, since neither matches this exact mechanism: #2696 (multicast scan across subnets returning only a random subset of devices) and #2904 (identifier-based multicast discovery getting zero replies at all). Worth a look in case either turns out to share a root cause, but this fix is narrowly about the unicast resend loop finishing before a slow-to-register service appears.


Diagnosis, fix, and tests written together with Claude Code (Anthropic).

UnicastDnsSdClientProtocol previously finished as soon as it had received
some reply for every outbound query packet, regardless of whether the reply
content was complete. A device can take a moment after boot/wake to
re-register a dynamic-port service (e.g. Companion) with its own mDNS
responder, so this could close the socket before that service ever had a
chance to answer.

Completion now requires either every explicitly requested service to have
answered (unambiguous, no wait needed), or the discovered service set to
have gone unchanged for UNICAST_MIN_STABLE_ROUNDS consecutive rounds
(reduces, doesn't eliminate, the risk of a still-registering service being
missed), or the final resend round to be reached.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant