refactor(ops): name Infinilm replacements - #915
Closed
voltjia wants to merge 1 commit into
Closed
Conversation
Collaborator
Author
|
Closing per scope prioritization: this PR only improves deprecation diagnostics and is not required for the functional InfiniLM migration. Continuing with the remaining executable migration work first. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Infinilmcompatibility operators with their concrete canonical replacements.Motivation
The canonical destinations for convolution, paged cache writes, decode/prefill attention, copies, batched top-k/top-p sampling, and zero filling are now present on
master. Their deprecated compatibility APIs should name those destinations directly so compiler diagnostics provide actionable migration guidance.No issue is linked. PR #893 was closed separately after the post-#911 audit because it targeted the retired vLLM v0.6.3
paged_attention_v1wrapper and duplicated the maintained FlashAttention replacement path.Type of Change
feat- new feature / new operator / new platformfix- bug fixperf- performance improvement (no behavioral change)refactor- code restructuring without behavior changetest- adding or fixing tests onlydocs- documentation onlybuild/ci- build system or CI configurationchore- tooling, formatting, or other non-code changes!in the Conventional Commits prefix or aBREAKING CHANGE:footer)Platforms Affected
WITH_CPU)WITH_NVIDIA)WITH_ILUVATAR)WITH_METAX)WITH_CAMBRICON)WITH_MOORE)WITH_ASCEND)WITH_TORCH)Only compile-time deprecation text changes. Runtime behavior is unchanged on every platform.
Smoke Test Result
Remote environment:
ssh nvidia, imageaccelerator-dev/nvidia:latest, with the existing InfiniRT prefix and/tmpCUTLASS source.The skips are the existing unsupported device/operator combinations in the legacy test matrix.
Test Results on Supported Platforms
236 passed; no runtime changeAdditional checks:
Benchmark / Performance Impact
N/A. No executable implementation changes.
Notes for Reviewers
Replacement alignment
ConvInfinilmConvolutiontorch.convolutionand fixed ATen schematransposed=false, zerooutput_padding, and reorders stride/padding.PagedCachingInfinilmReshapeAndCacheFlashreshape_and_cache_flashPagedAttentionInfinilmFlashAttnWithKvcacheflash_attn_with_kvcachePagedAttentionPrefillInfinilmFlashAttnVarlenFuncflash_attn_varlen_funcRearrangeInfinilmCopyTensor.copy_non_blocking=falsesubset.TopKTopPSampleInfinilmTopKTopPSamplingFromLogitstop_k_top_p_sampling_from_logitstop_k/top_ptensors and pass the canonical explicit attributes.ZerosInfinilmFillwith value0Tensor.fill_CausalSoftmaxInfinilm,ScaledSoftmaxInfinilm,KvCachingInfinilm, andRandomSampleInfinilmretain the generic diagnostic because no single stable public operator completely represents their existing contracts. This PR does not invent an approximate replacement or add an overload.