Skip to content

[None][chore] Align the prefix-tokenization cache switch with main (enable_tokenization_cache) - #19683

Draft
zheyuf wants to merge 3 commits into
NVIDIA:feat/m3_with_msafrom
zheyuf:zheyu/chore/m3-side-tokenization-cache-switch
Draft

zheyuf wants to merge 3 commits into
NVIDIA:feat/m3_with_msafrom
zheyuf:zheyu/chore/m3-side-tokenization-cache-switch

Conversation

@zheyuf

@zheyuf zheyuf commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description

On feat/m3_with_msa, the prefix-tokenization cache is still turned on with the TLLM_PREFIX_TOKEN_CACHE=1 env var. That is because #19373 cherry-picked #18389 before #18389 switched to the enable_tokenization_cache LLM argument. As a result, configs don't carry over between the branches: main's YAML key is rejected on this branch, and this branch's env var is ignored on main.

This PR brings the branch in line with main:

  1. It brings over the last three commits of [None][perf] Prefix-tokenization cache for the default input processor #18389 (b9c62bc70a, be33876a10, 7f59b77ef5, by @Tabrizian). These add the enable_tokenization_cache argument, move the docs page to features/, and add the golden-manifest row (in this branch's manifest format).
  2. It replaces [None][perf] Prefix-tokenization cache for MiniMax-M3 (cherry-pick #18389 + M3 hook) #19373's M3 hook with the one from [None][perf] Allow prefix-tokenization cache for MiniMax-M3 text-only prompts #19613, so modeling_minimaxm3_vl.py is the same file as in [None][perf] Allow prefix-tokenization cache for MiniMax-M3 text-only prompts #19613. It also drops [None][perf] Prefix-tokenization cache for MiniMax-M3 (cherry-pick #18389 + M3 hook) #19373's M3 unit test, which was built around the removed env var, and its M3 note in the docs.

The caching logic is unchanged; only the switch moves:

enable_tokenization_cache: true   # replaces TLLM_PREFIX_TOKEN_CACHE=1

The TLLM_PREFIX_TOKEN_CACHE_* sizing env vars still apply. TLLM_PREFIX_TOKEN_CACHE=1 no longer does anything.

End-to-end testing

Setup: MiniMax-M3 NVFP4 on B300, AgentX at TP4 C25 with Eagle3 K3, and TLLM_PREFIX_TOKEN_CACHE_MAX_CHARS=268435456 in every arm. I ran each arm for 900 s, twice on different nodes; each cell is the average of the two runs.

hit rate TTFT p50 (ms) TTFT p90 (ms) out tok/s/GPU
cache off – 782 1443 307
#19373 (env var) 86% 609 (−22%) 1251 (−13%) 310
this PR (YAML) 86% 611 (−22%) 1212 (−16%) 310

This PR performs the same as #19373, within run-to-run noise. The "this PR" runs were done before the last commit put back #19373's add_special_tokens / truncation check. Every replayed request is a chat completion and passes that check.

Test Coverage

No new unit test. I ran the following on a CPU node with the branch wheel and this PR's Python:

  • tests/unittest/inputs/test_prefix_token_cache.py and tests/unittest/api_stability/test_llm_api.py pass.
  • The golden manifest matches the generator.
  • On the real checkpoint, the YAML flag turns the cache on, the old env var doesn't, and the ids match the HF processor's.

PR Checklist

  • Please check this after reviewing the above items as appropriate for this PR.

Tabrizian and others added 2 commits September 28, 2026 22:20
…ation_cache

Sync feat/m3_with_msa with NVIDIA#18389 as merged to main (15c5995). NVIDIA#19373 cherry-picked the PR
at 8228785; this applies the three commits it gained before the merge: b9c62bc (replace the
TLLM_PREFIX_TOKEN_CACHE environment variable with the enable_tokenization_cache TorchLlmArgs field, plumbed through
create_input_processor to DefaultInputProcessor; the sizing knobs stay environment variables), be33876 (move the
page to docs/source/features/) and 7f59b77 (golden manifest row). prefix_token_cache.py, its unit test and the
docs page are byte-identical to main; the registry.py, llm.py, llm_args.py and API-stability hunks are main's, and
the golden-manifest row is the one this branch's generator produces (its manifest predates main's capture_policy
format).

Signed-off-by: Iman Tabrizian <[email protected]>
Signed-off-by: Zheyu Fu <[email protected]>
…tokenization_cache

With the cache now enabled by the enable_tokenization_cache LLM argument instead of TLLM_PREFIX_TOKEN_CACHE=1,
replace the MiniMax-M3 hook from NVIDIA#19373 with the one proposed for main in NVIDIA#19613, so the branch and
main enable the cache the same way. create_input_processor forwards enable_tokenization_cache to model-specific input
processors that set supports_tokenization_cache (the others would reject the unknown kwarg). MiniMaxM3VLInputProcessor
opts in and tokenizes text-only prompts through the cache, built on the HF processor's own tokenizer. It is used only
if that tokenizer adds no special tokens, checked at construction; this replaces NVIDIA#19373's probe prompt. Requests
with images or videos are unchanged. The NVIDIA#19373 unit test exercised the removed environment variable and probe and
is dropped, as NVIDIA#19613 adds none.

Signed-off-by: Zheyu Fu <[email protected]>
…utProcessor's rules

Mirror the change made to NVIDIA#19613: use the cache for a text-only prompt only when
sampling_params is given, add_special_tokens is False and the prompt is not truncated, the rule DefaultInputProcessor
applies, and declare supports_tokenization_cache on BaseMultimodalInputProcessor. modeling_minimaxm3_vl.py stays
byte-identical to NVIDIA#19613's.

Signed-off-by: Zheyu Fu <[email protected]>
@zheyuf zheyuf added the api-compatible Accepted LLM API contract change that is backwards-compatible label Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api-compatible Accepted LLM API contract change that is backwards-compatible

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants