Conversation
07c71d9 to
a88a792
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #75627 [ run ] triggered by Bot. Commit: |
|
PR_Github #75627 [ run ] completed with state |
|
|
||
| The shard component has two forms. A page whose bytes depend on the shard is | ||
| named per rank, while one whose bytes are identical on every rank is named | ||
| `REPLICATED_SHARD_KEY` instead, giving a TP group a single shared key. |
There was a problem hiding this comment.
so if TP is 8, (K head is 4), the change does not generalizes to that duplication because it is all ranks or one?
| keys, and the existence filter in `_put` collapses the overlap. Gating | ||
| them on rank would drop the pages of every owner but one. | ||
| """ | ||
| return self._attention_rank == 0 |
There was a problem hiding this comment.
no easy way to spread the writes with TP?
a88a792 to
1bf4152
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #76208 [ run ] triggered by Bot. Commit: |
Signed-off-by: Balaram Buddharaju <[email protected]>
Signed-off-by: Balaram Buddharaju <[email protected]>
1bf4152 to
ace7552
Compare
|
/bot run --disable-fail-fast |
|
PR_Github #76209 [ run ] triggered by Bot. Commit: |
|
PR_Github #76208 [ run ] completed with state |
|
PR_Github #76209 [ run ] completed with state |
Description
MiniMax-M3's index-K is computed from a replicated projection, so every TP rank
holds identical bytes. The Mooncake store connector keyed every page by
attention shard (
w<count>r<rank>), so a TP=8 deployment wrote and stored eightcopies of the same index-K.
This introduces the notion of a replicated role and threads it through the
connector stack:
KVCacheManagerV2.get_replicated_roles()names the roles whose bytes do notdepend on the shard. It is derived from the existing
get_disagg_role_mapper_kinds()declaration, so replication is declared onceand the native disagg path and the connector path cannot disagree.
KvCacheLayoutdescribes those bytes as a separatereplicated_regionssetper layer group. V2 may interleave the two classes inside one pool, so each
class is aggregated separately rather than sliced out of a merged range.
replicatedin placeof the rank component. One rank writes it and every rank reads it, since each
still needs the bytes in its own GPU memory. A prefix hit requires it
alongside every shard's page, so a block whose shared copy was evicted reads
as a miss rather than loading half-initialized.
For MiniMax-M3 at TP=8, where K, V and index-K are each 256 B/token, this
removes about 29% of the store footprint and of the bytes written.
Models that declare no replicated role are unaffected:
replicated_regionsisempty and every buffer stays in
regions.Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.