Conversation
CUDA-graph padding rows and warmup dummies share the drafter's dummy slot. Their accepted tokens grow its context length like any request's, but nothing reset it, while their page-table rows are the padding request's pages with every other entry mapped to page 0, another request's page. Once the dummy length passed the padding request's own pages, the padding rows' context K / V (k3_ctx_kv, or the torch path) landed on that page: a live request's drafter context overwritten on every padded step, costing acceptance. prepare() now writes 0 to the dummy slot every step, with the evicted slots. test_dflash_dummy_slot.py: padded steps keep the dummy slot at 0 and the real slots untouched; an evicted request and the dummy reset in one step. It fails without the reset. Signed-off-by: Vasanth Sabavat <[email protected]>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 🧰 Additional context used📚 Code guidelines (1)No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 7 remain after this review. WalkthroughDFlash preparation now resets the dummy slot’s context length on each step. CUDA tests cover padding rows across acceptance steps and cleanup of an evicted request slot. ChangesDFlash dummy-slot preparation
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to The dummy-slot reset addresses the reported padding-row behavior, and no merge-blocking issue is established. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Description
CUDA-graph padding rows and warmup dummies decode on DFlash's dummy slot (
worker._dummy_slot,tensorrt_llm/_torch/speculative/dflash.py).DFlashSpecMetadata.prepareresets only evicted slots, and_assign_slotnever runs for thedummy slot, so the dummy's context length grows by every padded step's accepted tokens.
mapped to page 0, which belongs to another, live request.
A live request's drafter context is overwritten on every padded step. This costs acceptance only, because the
target verifies.
It is reachable in any configuration whose CUDA graphs pad a batch.
prepare()now writes 0 to the dummy slot's context length every step, with the evicted slots.Test Coverage
tests/unittest/_torch/speculative/hw_agnostic/test_dflash_dummy_slot.py(new; needs a GPU, the slot lengths live onit):
Both fail without the reset: on GB200, 2 failed on main and 2 passed with this PR.
speculative/hw_agnostic/has no new failures.
l0_h100.ymlcollects the file throughunittest/_torch/speculative/hw_agnostic; onl0_cpuit skips (no GPU).PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title. (No API change.)Any new dependencies have been scanned for license and vulnerabilities (No new dependencies.)
CODEOWNERS updated if ownership changes (No ownership change.)
Documentation updated as needed (No user-facing change.)
Update tava architecture diagram if there is a significant design change in PR. (No design change.)
The reviewers assigned automatically/manually are appropriate for the PR. (To check once the PR is open.)
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.Dev Engineer Review
DFlashSpecMetadata.preparenow resets the dummy slot’s context length to zero on each step, alongside evicted slots. This prevents CUDA-graph padding and warmup dummy rows from extending context K/V writes into page 0, which may belong to a live request. The change does not alter public APIs.QA Engineer Review
Added two CUDA-gated tests. They cover repeated padded steps with real-slot lengths preserved, and simultaneous reset of the dummy slot and an evicted request. The tests simulate slot updates; they do not run a model forward. Coverage verdict: sufficient for the reset behavior. The supplied results do not establish test execution status.
Per-File QA Perspective
tensorrt_llm/_torch/speculative/dflash.py: Verify thatprepareresets the dummy context length without changing real request slots. The change also applies when a step has no evicted requests.tests/unittest/_torch/speculative/hw_agnostic/test_dflash_dummy_slot.py: Covers padded slot mappings and resets alongside eviction. Thehw_agnosticsuite is included in thel0_h100.ymlCI list; the test file skips when CUDA is unavailable.