Skip to content

fix(opds): classify listener failures, keep serving through interface hiccups - #159

Closed
phildenhoff wants to merge 1 commit into
mainfrom
fix-opds-listener-resilience
Closed

phildenhoff wants to merge 1 commit into
mainfrom
fix-opds-listener-resilience

Conversation

@phildenhoff

Copy link
Copy Markdown
Member

Problem

The sharing service's error handling was all-or-nothing in ways that produced wrong behavior in common situations:

  • A taken port strobes forever. If another app held the port, every 2-second reconcile tick flipped the service between Error and Running, with a message promising a retry that could never succeed.
  • One dead listener killed them all. A single listener crashing tore down every other address, interrupting transfers that were working fine.
  • An interface-enumeration hiccup stopped serving. A transient getifaddrs-style failure shut down all listeners instead of riding it out.
  • Globally-routable addresses were bindable by default, with no hook to gate them later on auth being enabled.

What this does

  • Classifies bind failures. Address-in-use and permission errors (including Windows' Hyper-V/WSL reserved ranges, which surface as EACCES) are permanent: the service reports a clear "pick a different port" error and stops retrying that address. Everything else is transient: retried with exponential backoff (2× base up to 60s), and only surfaced as an error after two consecutive failures — one flaky bind is not an incident.
  • Isolates failures per address. A bind failure or crashed listener affects only that address; surviving listeners keep serving, and the status reports the URLs that are actually bound.
  • Rides out enumeration failures. When a snapshot fails mid-run, listeners keep serving on the last known plan; advertisement refreshes on the next successful snapshot.
  • Adds BindPolicy. plan_bindings now takes a policy; globally-routable addresses are excluded from binding by default. The flag exists so auth-enabled sharing (feat(opds): HTTP Basic auth and credential management #149) can open it deliberately.

Testing

43 tests pass, including four new ones: persistent-failure isolation, transient-failure error-then-recover, enumeration-hiccup keep-serving, and the bind policy. The dead-listener test was updated for the new semantics (single failure no longer surfaces an error).

Notes for reviewers

This is the "PR-a" correctness fix from the sharing rework discussion; the bind-model rework (Local network / All networks modes) follows separately.

… hiccups

- Bind failures are now classified: address-in-use and permission errors are
  permanent (clear message, no retry loop); other failures retry with
  exponential backoff and only surface as errors after repeated attempts.
- A failing listener no longer tears down the others; partial success serves
  the bound addresses.
- Transient interface-enumeration failures no longer stop running listeners;
  advertisement refreshes on the next successful snapshot.
- plan_bindings takes a BindPolicy; globally-routable addresses are excluded
  from binding by default (reserved for auth-enabled sharing).
@github-actions

Copy link
Copy Markdown

libcalibre Test Coverage Report

Overall coverage: 80.16%

📊 Download HTML Report

Coverage breakdown available in the artifacts.

@phildenhoff

Copy link
Copy Markdown
Member Author

Superseded by #160: rather than hardening the reconcile loop, we removed it - sharing is now an explicit state machine (Stopped/Starting/Running/Waiting/Failed) with a Local network / All interfaces model. The error classification and BindPolicy from this PR were carried into the rework.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant