Skip to content

fix(connectors): bound source forwarding channel with backpressure - #3795

Open
mlevkov wants to merge 1 commit into
apache:masterfrom
mlevkov:bounded-source-channel
Open

fix(connectors): bound source forwarding channel with backpressure#3795
mlevkov wants to merge 1 commit into
apache:masterfrom
mlevkov:bounded-source-channel

Conversation

@mlevkov

@mlevkov mlevkov commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Summary

The channel between a source plugin's send callback and the runtime's forwarding loop was flume::unbounded(), so a slow or hung Iggy meant batches accumulated in memory without bound instead of propagating backpressure into the plugin's polling loop. This is the prerequisite runtime fix requested in the HTTP source discussion (#3039), and it applies to every source connector, including the source PRs currently in flight.

What changed

  • The forwarding channel is now a bounded crossfire channel (crossfire::mpsc::bounded_blocking_async), the same shape shard and server-ng use. flume is no longer a runtime dependency.
  • Capacity comes from a new optional SourceConfig field, channel_capacity, counted in batches (a single batch can be megabytes), defaulting to 1024 and clamped to [1, 65536] since crossfire eagerly allocates the ring and asserts capacity < 2^31. The existing ConfigEnv derive provides IGGY_CONNECTORS_SOURCE_<KEY>_CHANNEL_CAPACITY; configs without the field behave as before apart from the bound.
  • The FFI send callback applies backpressure with a try_send fast path and a send_timeout(10ms) retry loop that re-reads a per-instance shutdown flag between waits. The manager sets that flag before iggy_source_close so a hung Iggy cannot deadlock the close. Process shutdown sets every instance's flag (signal_shutdown_all) before the sequential stops, because instances loaded from one plugin library share a single tokio runtime and a wedged sibling would otherwise hold a worker an earlier close needs.
  • A full channel logs one warn! per backpressure episode (latched, cleared on genuine recovery). A batch that still cannot be enqueued after the stop signal is dropped and counted in iggy_connector_errors_total.
  • Unit tests pin the behavior shutdown relies on: buffered batches drain after the senders drop (crossfire's docs do not promise this), the retry loop unblocks when the flag flips mid-backoff, and shutdown drops are counted without being enqueued.
  • Docs updated across the runtime README, sources README, and the connector skills, which still described flume and the pre-feat(connectors): add agent docs, per-batch observability, atomic state #3321 shutdown order.

One correction to the discussion notes

@hubcio the spec assumed crossfire's blocking sender has no send_timeout. It does: blocking_tx.rs:288 on Tx, reachable from MTx via Deref. The loop is built on it instead of try_send plus sleep, so the sender wakes as soon as capacity frees while shutdown latency stays bounded by the retry interval.

Known limitation

Stopping a single connector via the runtime API while enough same-library sibling instances are saturated can delay that close until the siblings drain, because the callback parks a worker of the shared plugin runtime. The code comment and the connector skill document this. The complete fix is an SDK-side worker handoff (tokio::task::block_in_place around the callback invocation); happy to file it as a follow-up issue.

Test plan

  • cargo clippy -p iggy-connectors --all-targets -- -D warnings clean
  • cargo test -p iggy-connectors: 128 passed, including the new channel and shutdown tests
  • cargo build -p iggy_connector_stdout_sink -p iggy_connector_random_source
  • cargo test -p integration -- connectors::runtime:: could not run on this machine (hwlocality-sys needs pkg-config); relying on CI for the integration suite

@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown

Thanks for the PR. It is labeled S-waiting-on-review and queued for review.

Slash commands (own line, regular comment) move it around the queue:

  • /ready - back to S-waiting-on-review after addressing feedback
  • /author - flip to S-waiting-on-author while you finish changes
  • /request-review @user-or-team - request a reviewer

See CONTRIBUTING.md for details.

@github-actions github-actions Bot added the S-waiting-on-review PR is waiting on a reviewer label Aug 1, 2026
@codecov

codecov Bot commented Aug 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.53846% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 17.28%. Comparing base (1453114) to head (bd19da2).
⚠️ Report is 20 commits behind head on master.

Files with missing lines Patch % Lines
core/connectors/runtime/src/source.rs 94.39% 18 Missing ⚠️
core/connectors/runtime/src/configs/connectors.rs 0.00% 1 Missing ⚠️
core/connectors/runtime/src/main.rs 0.00% 1 Missing ⚠️
core/connectors/runtime/src/manager/source.rs 50.00% 1 Missing ⚠️
Additional details and impacted files
@@              Coverage Diff              @@
##             master    #3795       +/-   ##
=============================================
- Coverage     75.72%   17.28%   -58.45%     
  Complexity      969      969               
=============================================
  Files          1322     1320        -2     
  Lines        159363   137264    -22099     
  Branches     132746   110724    -22022     
=============================================
- Hits         120684    23727    -96957     
- Misses        35041   113022    +77981     
+ Partials       3638      515     -3123     
Components Coverage Δ
Rust Core 2.25% <93.53%> (-73.44%) ⬇️
Java SDK 62.71% <ø> (ø)
C# SDK 71.13% <ø> (-1.17%) ⬇️
Python SDK 93.10% <ø> (ø)
PHP SDK 84.52% <ø> (ø)
Node SDK 95.30% <ø> (+0.07%) ⬆️
Go SDK 43.08% <ø> (ø)
Files with missing lines Coverage Δ
core/connectors/runtime/src/configs/connectors.rs 0.00% <0.00%> (-41.08%) ⬇️
core/connectors/runtime/src/main.rs 31.75% <0.00%> (-53.97%) ⬇️
core/connectors/runtime/src/manager/source.rs 70.00% <50.00%> (-23.47%) ⬇️
core/connectors/runtime/src/source.rs 37.60% <94.39%> (-43.07%) ⬇️

... and 814 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@mlevkov

mlevkov commented Aug 2, 2026

Copy link
Copy Markdown
Contributor Author

/request-review @hubcio

@github-actions
github-actions Bot requested a review from hubcio August 2, 2026 02:21
mlevkov added a commit to mlevkov/iggy that referenced this pull request Aug 2, 2026
Iggy has no way to receive a webhook. Every provider that pushes events
over HTTP needs something in front of it, and today that means running a
separate service whose only job is to accept a POST and republish it.
This connector removes that hop: it runs an embedded HTTP server, accepts
authenticated POST bodies, and produces them to the instance's stream and
topic as raw bytes.

One plugin .so is loaded once no matter how many source entries reference
it, so the listener cannot live on any single instance. It lives in a
process-global registry keyed by listen address: the first open binds the
public and admin ports, later opens validate their body limit, admin
address, management token and instance name against the running listener
before joining, and the last close releases both ports. Mismatches fail
that instance's open rather than silently handing it a listener its
configuration does not describe. A single port can therefore serve many
providers, each routed to its own topic.

Requests resolve against an ArcSwap route table that is rebuilt whole on
every control-plane change, so one atomic load yields both the endpoint's
auth rules and the destination bridge. Secret paths carry 128 bits in the
URL itself, on the model of a Slack webhook, with optional bearer or HMAC
on top; HMAC is verified over the raw body in constant time. Revoked
endpoints answer 404 alongside paths that never existed, so a leaked URL
cannot be used to confirm it was once live.

Endpoints can be registered, re-keyed and revoked at runtime through a
token-guarded API on the admin listener, because revoking a compromised
endpoint is time-critical and provisioning one per tenant is inherently
programmatic. Those endpoints ride the SDK's ConnectorState, and state is
attached only to an empty batch: the runtime saves state solely on the
success branch of the Iggy send, and an empty send always succeeds, so a
mutation cannot be lost to an unrelated send failure. Revocation writes a
tombstone that outranks TOML on restore, so a stale config file cannot
resurrect an endpoint an operator revoked.

Delivery is best-effort in both directions and the README says so first,
before anything else: HTTP 200 means accepted into an in-memory buffer,
and both the loss and duplicate windows are enumerated with what mitigates
each. A full bridge answers 429 with Retry-After rather than blocking,
since holding the connection open would turn a slow Iggy into a retry
storm. Gateway metrics on the admin listener cover accept-to-200 latency,
which the runtime's own stage histograms begin too late to see.

Part of the webhook gateway design accepted in apache#3039. The backpressure
chain is only complete once the bounded runtime forwarding channel from
apache#3795 lands; until then a full bridge signals an arrival burst rather
than a slow Iggy, which the README documents.

Co-authored-by: Claude <[email protected]>
The channel between a source plugin's send callback and the runtime's
forwarding loop was flume::unbounded(), so a slow or hung Iggy meant
batches accumulated without bound instead of propagating backpressure
into the plugin's polling loop.

Swap it for a bounded crossfire channel (the shard and server-ng
standard), sized by an optional SourceConfig channel_capacity counted
in batches, defaulting to 1024. The FFI callback retries with
send_timeout while re-reading a shutdown flag, set by the manager
before iggy_source_close and for every instance ahead of the
sequential process-shutdown stops, since same-library instances share
one plugin runtime and a wedged sibling would otherwise hold a worker
an earlier close needs. A unit test pins that buffered batches drain
after the senders drop, which shutdown relies on and crossfire's docs
do not promise. This drops flume from the runtime.

Requested in the HTTP source discussion (apache#3039).

Co-authored-by: Claude <[email protected]>
@mlevkov
mlevkov force-pushed the bounded-source-channel branch from d7580d5 to bd19da2 Compare August 2, 2026 21:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

S-waiting-on-review PR is waiting on a reviewer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant