Repository navigation
[None][chore] Align the prefix-tokenization cache switch with main (enable_tokenization_cache) - #19683
Draft
zheyuf wants to merge 3 commits into
Draft
[None][chore] Align the prefix-tokenization cache switch with main (enable_tokenization_cache)#19683zheyuf wants to merge 3 commits into
zheyuf wants to merge 3 commits into
Conversation
…ation_cache Sync feat/m3_with_msa with NVIDIA#18389 as merged to main (15c5995). NVIDIA#19373 cherry-picked the PR at 8228785; this applies the three commits it gained before the merge: b9c62bc (replace the TLLM_PREFIX_TOKEN_CACHE environment variable with the enable_tokenization_cache TorchLlmArgs field, plumbed through create_input_processor to DefaultInputProcessor; the sizing knobs stay environment variables), be33876 (move the page to docs/source/features/) and 7f59b77 (golden manifest row). prefix_token_cache.py, its unit test and the docs page are byte-identical to main; the registry.py, llm.py, llm_args.py and API-stability hunks are main's, and the golden-manifest row is the one this branch's generator produces (its manifest predates main's capture_policy format). Signed-off-by: Iman Tabrizian <[email protected]> Signed-off-by: Zheyu Fu <[email protected]>
…tokenization_cache With the cache now enabled by the enable_tokenization_cache LLM argument instead of TLLM_PREFIX_TOKEN_CACHE=1, replace the MiniMax-M3 hook from NVIDIA#19373 with the one proposed for main in NVIDIA#19613, so the branch and main enable the cache the same way. create_input_processor forwards enable_tokenization_cache to model-specific input processors that set supports_tokenization_cache (the others would reject the unknown kwarg). MiniMaxM3VLInputProcessor opts in and tokenizes text-only prompts through the cache, built on the HF processor's own tokenizer. It is used only if that tokenizer adds no special tokens, checked at construction; this replaces NVIDIA#19373's probe prompt. Requests with images or videos are unchanged. The NVIDIA#19373 unit test exercised the removed environment variable and probe and is dropped, as NVIDIA#19613 adds none. Signed-off-by: Zheyu Fu <[email protected]>
…utProcessor's rules Mirror the change made to NVIDIA#19613: use the cache for a text-only prompt only when sampling_params is given, add_special_tokens is False and the prompt is not truncated, the rule DefaultInputProcessor applies, and declare supports_tokenization_cache on BaseMultimodalInputProcessor. modeling_minimaxm3_vl.py stays byte-identical to NVIDIA#19613's. Signed-off-by: Zheyu Fu <[email protected]>
1 task done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
On
feat/m3_with_msa, the prefix-tokenization cache is still turned on with theTLLM_PREFIX_TOKEN_CACHE=1env var. That is because #19373 cherry-picked #18389 before #18389 switched to theenable_tokenization_cacheLLM argument. As a result, configs don't carry over between the branches: main's YAML key is rejected on this branch, and this branch's env var is ignored on main.This PR brings the branch in line with main:
b9c62bc70a,be33876a10,7f59b77ef5, by @Tabrizian). These add theenable_tokenization_cacheargument, move the docs page tofeatures/, and add the golden-manifest row (in this branch's manifest format).modeling_minimaxm3_vl.pyis the same file as in [None][perf] Allow prefix-tokenization cache for MiniMax-M3 text-only prompts #19613. It also drops [None][perf] Prefix-tokenization cache for MiniMax-M3 (cherry-pick #18389 + M3 hook) #19373's M3 unit test, which was built around the removed env var, and its M3 note in the docs.The caching logic is unchanged; only the switch moves:
The
TLLM_PREFIX_TOKEN_CACHE_*sizing env vars still apply.TLLM_PREFIX_TOKEN_CACHE=1no longer does anything.End-to-end testing
Setup: MiniMax-M3 NVFP4 on B300, AgentX at TP4 C25 with Eagle3 K3, and
TLLM_PREFIX_TOKEN_CACHE_MAX_CHARS=268435456in every arm. I ran each arm for 900 s, twice on different nodes; each cell is the average of the two runs.This PR performs the same as #19373, within run-to-run noise. The "this PR" runs were done before the last commit put back #19373's
add_special_tokens/ truncation check. Every replayed request is a chat completion and passes that check.Test Coverage
No new unit test. I ran the following on a CPU node with the branch wheel and this PR's Python:
tests/unittest/inputs/test_prefix_token_cache.pyandtests/unittest/api_stability/test_llm_api.pypass.PR Checklist