Disaggregated Prefill-Decode serving engine connecting NVIDIA Blackwell (SM120) and Huawei Ascend 910B2 (CANN) across physical nodes.
cuda cann inference-engine heterogeneous-computing kv-cache vllm llm-inference pagedattention deepseek ai-infrastructure distributed-inference nvidia-blackwell disaggregated-serving huawei-ascend ascend-910b pytorch-npu davinci-architecture
-
Updated
Sep 24, 2026 - Python