Skip to content

[None][feat] support C++ streaming KV events with multimodal payloads - #19359

Open
GuanLuo wants to merge 16 commits into
NVIDIA:mainfrom
GuanLuo:feat/streaming-kv-events-cpp-dto
Open

GuanLuo wants to merge 16 commits into
NVIDIA:mainfrom
GuanLuo:feat/streaming-kv-events-cpp-dto

Conversation

@GuanLuo

@GuanLuo GuanLuo commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Dev Engineer Review

The change adds a native C++ streaming event sink for the C++ KV-cache manager V2 backend. The sink captures stored and removed block events, filters by lifecycle and block state, coalesces events, limits pending entries, and records statistics. Nanobind exposes the sink and event DTOs. Python converts the DTOs into BlockStored and BlockRemoved events for the existing publisher path. The Python backend retains its own event source.

The event payload can contain hexadecimal digest strings in token_ids and mm_keys metadata for multimodal blocks. Consumers that require integer-only token IDs may not be compatible. Verify wire compatibility, event ordering, buffer-drop behavior, shutdown handling, and API parity across backends.

The PR description reports successful native builds, unit tests, formatting, lint, and pre-commit checks. It also reports that the routing E2E test was not rerun after the source-ownership refactor. CI runs PR_Github #75118 and L0_MergeRequest_PR pipeline #61871 failed. The supplied evidence does not identify the failed tests. A later run was triggered, but no result is supplied.

QA Engineer Review

Changed unit tests cover native event handling, stored and removed events, multimodal payloads, backend validation, attention-layer window filtering, unpublished-parent drops, and negative lifecycle IDs. Test-function names are not available for all added coverage. The changed files are unit tests, not integration tests. The KV-cache manager V2 test directory is already included in the CI lists l0_cpu.yml, l0_b200.yml, l0_h100.yml, and l0_a10.yml; no test-list files changed. Coverage verdict: needs follow-up, because CI failed and the routing E2E test was not rerun after the refactor.

Per-File QA Perspective

  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/CMakeLists.txt: Adds the event implementation files to the native build. Verify compilation and linking.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/streamingEventSink.cpp: Adds native event capture, filtering, coalescing, bounded buffering, and statistics. Verify lifecycle selection, block drops, multimodal handling, and removal behavior.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/streamingEventSink.h: Defines native event payloads and the sink API. Verify constructor defaults and consistency with nanobind.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/eventData.cpp: Decodes block tokens and multimodal digests. Verify digest encoding and mm_keys alignment.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/eventData.h: Adds shared event-token and multimodal-key types. Verify compatibility with Python event structures.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/eventManager.cpp: Uses shared decoding for stored blocks. Verify decoded token IDs and cache metadata in emitted events.
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/eventManager.h: Replaces duplicate event-data declarations with shared types. Verify dependent C++ interfaces.
  • cpp/tensorrt_llm/nanobind/batch_manager/kvCacheManagerV2.cpp: Exposes native sink types and accepts an EventSink in manager construction. Verify binding signatures and accepted sink types.
  • tensorrt_llm/_torch/pyexecutor/kv_cache/kv_cache_manager_v2.py: Passes the backend choice to the event manager. Verify C++ sink selection and attention-layer filtering.
  • tensorrt_llm/_torch/pyexecutor/kv_cache_events.py: Adds backend-specific event sources and native DTO conversion. Verify draining, publication, shutdown, and statistics synchronization.
  • tensorrt_llm/runtime/kv_cache_manager_v2/__init__.py: Exports native event types for non-Python backends. Verify imports and __all__.
  • tensorrt_llm/runtime/kv_cache_manager_v2/__init__.pyi: Updates public type declarations for the sink and manager inputs. Verify stub and binding parity.
  • tensorrt_llm/runtime/kv_cache_manager_v2/_introspection.py: Adds helpers for forwarding events to the native sink. Verify event forwarding and unavailable-backend errors.
  • docs/source/features/kvcache.md: Documents streaming event scope and multimodal token values. Verify the documented wire contract matches producer behavior.
  • tests/unittest/_torch/executor/kv_cache/test_kv_cache_manager_v2.py: Adds coverage for attention-layer window filtering. This is a unit test; integration test-list registration does not apply.
  • tests/unittest/kv_cache_manager_v2_tests/test_kv_cache_event_manager.py: Adds native event coverage, including unpublished-parent drops and negative lifecycle IDs. This is a unit test; the parent test directory is listed in the CI test lists noted above.
  • tests/unittest/kv_cache_manager_v2_tests/test_streaming_kv_events.py: Covers stored and removed events, multimodal payloads, backend validation, and parallelism constraints. This is a unit test; the parent test directory is listed in the CI test lists noted above.

Description

The Python KV cache manager backend is being removed, but the streaming KV-event publisher introduced by #17023 currently depends on Python-side KV manager events.

This change keeps streaming KV events available with the C++ KV cache manager V2 backend:

  • Add a native StreamingEventSink that captures stored and removed block events as compact C++ DTOs.
  • Expose batched DTO draining through nanobind.
  • Translate the DTOs to Python BlockStored and BlockRemoved msgspec structures and reuse the existing serialization and publisher path.
  • Select the native event lifecycle automatically when the C++ backend is active while preserving the Python backend behavior.
  • Add unit coverage and update the KV-cache documentation.

The boundary intentionally remains semantic rather than serialized:

C++ KV cache manager -> native event DTOs -> nanobind -> Python BlockStored/BlockRemoved -> existing publisher

Multimodal wire compatibility

Breaking change for V2 multimodal streaming consumers: previously, streaming token_ids contained integers and digest-bearing blocks were suppressed. This PR can emit hexadecimal digest strings in token_ids and add mm_keys to BlockStored. Consumers that assume every token ID is an integer are incompatible. Text-only payloads are unchanged.

The final cross-project interface is still being finalized by Dynamo #15095 and TensorRT-LLM #19529. We will revisit multimodal streaming handling separately after both PRs merge, when the interface is finalized, rather than define that final contract in this DTO/backend PR.

Test Coverage

  • Native C++ library and nanobindings build: passed.
  • Native KV-event manager tests: 55 passed.
  • Legacy Python streaming-event tests: 34 passed.
  • C++ backend lifecycle regression: 1 passed, 133 deselected.
  • Earlier Dynamo two-worker KV-aware routing E2E with the C++ backend and ZMQ events: 1 passed in 84.31 seconds; both workers received stored-block state. This was not rerun for the source-ownership refactor.
  • Relevant formatting, lint, and pre-commit checks: passed.

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions).

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities.

  • CODEOWNERS updated if ownership changes.

  • Documentation updated as needed.

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, comment /bot help.

@GuanLuo
GuanLuo marked this pull request as ready for review September 21, 2026 08:24
@GuanLuo
GuanLuo requested review from a team as code owners September 21, 2026 08:24
@GuanLuo
GuanLuo force-pushed the feat/streaming-kv-events-cpp-dto branch from c095803 to 4fc3763 Compare September 21, 2026 08:31
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/TensorRT-LLM/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: fca6fac8-f25a-4524-9c6a-6c6e7c535626

📥 Commits

Reviewing files that changed from the base of the PR and between b3971b0 and 99f1810.

📒 Files selected for processing (1)
  • cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/eventData.h

Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

The PR adds native C++ streaming events for KV-cache blocks. It exposes the sink through bindings, integrates native event draining with Python publication, supports both V2 backends, and adds multimodal event handling and tests.

Changes

Native streaming KV-cache events

Layer / File(s) Summary
Event data and collection
cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/*
Defines event payloads, digest and multimodal decoding, lifecycle filtering, event coalescing, bounded pending storage, and statistics.
Native bindings and exports
cpp/tensorrt_llm/nanobind/batch_manager/kvCacheManagerV2.cpp, tensorrt_llm/runtime/kv_cache_manager_v2/*
Exposes the sink and event data, adds event-submission helpers, exports runtime symbols, and accepts the common EventSink interface.
Python event integration
tensorrt_llm/_torch/pyexecutor/kv_cache_events.py, tensorrt_llm/_torch/pyexecutor/kv_cache/kv_cache_manager_v2.py
Selects the backend-specific sink, propagates it through cache-manager construction, drains native events, converts DTOs to wire events, and synchronizes statistics.
Validation and documented backend support
docs/source/features/kvcache.md, tests/unittest/...
Documents C++ backend support and tests multimodal payloads, lifecycle filtering, event counters, bounded publication, event-window filtering, and backend validation.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant KVCacheManager
  participant StreamingEventSink
  participant StreamingKVCacheEventManager
  participant WirePublisher
  KVCacheManager->>StreamingEventSink: submit block lifecycle events
  StreamingEventSink->>StreamingKVCacheEventManager: provide drained native DTOs
  StreamingKVCacheEventManager->>WirePublisher: publish stored or removed wire events
  StreamingKVCacheEventManager->>StreamingEventSink: synchronize native statistics
Loading

Suggested reviewers: juney-nvidia

Merge Risk: ⚪ Minimal · up to 99f18

This change declares the native KV-event data types and decoder interface consistently with their implementation and consumer. No merge-blocking risk is identified.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 16.07% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 112 functions across 15 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the feature: C++ streaming KV events with multimodal payload support. It follows the required ticket and type format and matches the main changes.
Description check ✅ Passed The description follows the required template. It explains the motivation and implementation, identifies the multimodal compatibility change, lists relevant test coverage, documents the affected inter…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@GuanLuo
GuanLuo force-pushed the feat/streaming-kv-events-cpp-dto branch from 9850306 to 003d83f Compare September 21, 2026 08:55
@GuanLuo
GuanLuo force-pushed the feat/streaming-kv-events-cpp-dto branch from 003d83f to 3e0a0df Compare September 21, 2026 08:58
Comment thread cpp/tensorrt_llm/batch_manager/kv_cache_manager_v2/streamingEventSink.h Outdated
Comment thread tensorrt_llm/_torch/pyexecutor/kv_cache_events.py
Comment thread tensorrt_llm/_torch/pyexecutor/kv_cache_events.py Outdated
Comment thread tensorrt_llm/_torch/pyexecutor/kv_cache/kv_cache_manager_v2.py

@zhaoyangwang-nvidia zhaoyangwang-nvidia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve with nits.

@tanmayv25

Copy link
Copy Markdown
Collaborator

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75118 [ run ] triggered by Bot. Commit: 7f9e71b Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75118 [ run ] completed with state FAILURE. Commit: 7f9e71b
/LLM/main/L0_MergeRequest_PR pipeline #61871 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@GuanLuo

GuanLuo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

/bot run

@GuanLuo

GuanLuo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Fixed the x86_64/SBSA -Wmismatched-tags build failure in 99f1810415 by aligning the Block forward declaration with its canonical struct definition. Verified the affected header combination with Clang on Linux using -Wmismatched-tags -Werror; the compile passes. A fresh /bot run is in progress.

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75203 [ run ] triggered by Bot. Commit: 99f1810 Link to invocation

@GuanLuo GuanLuo changed the title [None][feat] support streaming KV events with C++ backend [None][feat] support C++ streaming KV events with multimodal payloads Oct 1, 2026
Address review feedback with a simpler digest encoder and concise streaming documentation that distinguishes multimodal payload support from deferred routing compatibility.

Signed-off-by: Guan Luo <[email protected]>
@GuanLuo

GuanLuo commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Addressed the follow-ups from review #19359 (review) in 1aabf98:

  • Updated the title to explicitly mention multimodal payloads. This does not claim finalized end-to-end MM-aware routing support; the documented follow-up remains separate.
  • Added native equivalent coverage for test_streaming_removals_are_never_dropped_by_the_entry_cap. Both a saturated cap and an overflowed cap are tested, exercising whole-block and lifecycle removal entry points and verifying the published MessagePack payload retains both removal hashes.
  • Simplified digestToHex and shortened the streaming documentation.

Validation: native library and bindings rebuild passed; 2 targeted cap tests and 228 regression tests passed; Python lint/format and C++ formatting checks passed. No new model-serving or Dynamo routing E2E was run for this follow-up.

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

@GuanLuo

GuanLuo commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76010 [ run ] triggered by Bot. Commit: 1aabf98 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76010 [ run ] completed with state SUCCESS. Commit: 1aabf98
/LLM/main/L0_MergeRequest_PR pipeline #62674 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

@trtllm-agent

This comment has been minimized.

@coderabbitai

This comment has been minimized.

@GuanLuo

GuanLuo commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76136 [ run ] triggered by Bot. Commit: 1aabf98 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76136 [ run ] completed with state FAILURE. Commit: 1aabf98
/LLM/main/L0_MergeRequest_PR pipeline #62767 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@GuanLuo

GuanLuo commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76211 [ run ] triggered by Bot. Commit: 2bf0ea7 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76211 [ run ] completed with state SUCCESS. Commit: 2bf0ea7
/LLM/main/L0_MergeRequest_PR pipeline #62835 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@GuanLuo

GuanLuo commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76217 [ run ] triggered by Bot. Commit: 2bf0ea7 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76217 [ run ] completed with state SUCCESS. Commit: 2bf0ea7
/LLM/main/L0_MergeRequest_PR pipeline #62840 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@GuanLuo

GuanLuo commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76233 [ run ] triggered by Bot. Commit: 2bf0ea7 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76233 [ run ] completed with state SUCCESS. Commit: 2bf0ea7
/LLM/main/L0_MergeRequest_PR pipeline #62854 completed with status: 'SUCCESS'

CI Report

Link to invocation

@GuanLuo

GuanLuo commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

/bot reuse-pipeline

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76308 [ reuse-pipeline ] triggered by Bot. Commit: 526e0dd Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76308 [ reuse-pipeline ] completed with state SUCCESS. Commit: 526e0dd
Reusing PR_Github #76233 for commit 526e0dd

Link to invocation

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants