馃殌 The feature, motivation and pitch
hi @kedarpotdar-nv @mnicely
on pytorch images, there is automatic AWS EFA & auto gcp nccl plugin detection & setup such that efa & gcp networks work out of the box. [source: https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-07.html ]
Unfortunately on NGC TRTLLM release images & NGC TensorRT LLM Develop
images, using NIXL EFA requires an massive undocumented script to download the correct EFA userspace libraries.
Any chance that EFA NIXL can work out of the box on NGC TensorTLLM release & develop images too?
We recognize that on NGC TRTLLM dynamo images there is curerntly images with efa tags but for those that wanna use the latest trtllm develop and trtllm release, efa doesnt work out of the box on trtllm
https://catalog.ngc.nvidia.com/orgs/nvidia/tensorrt-llm/containers/release/-/tags
Similar requests on vLLM Pollara & BXNT NICs where implemented in vllm-project/vllm#38687
and over the past 2 weeks, there is an global cross ecosystem effort to make efa work out of the box on NVIDIA GPUs
gaint runtime installation command from https://github.com/SemiAnalysisAI/InferenceX/pull/2993/changes/BASE..1df41b9f4a4709e7007f441e9ab0ec6aa9383514#diff-d5cf9a63d050c8f89f86e3b80f728cf8f563d148ce6a6b772ece84f5d415f243R224-R263
NIXL_LIBFABRIC_HOST_DIR="/data/home/sa-gha-runner/nixl-libfabric/nixl-1.4.0-efa-1.47.0"
mkdir -p "$(dirname "$NIXL_LIBFABRIC_HOST_DIR")"
(
exec 9>"${NIXL_LIBFABRIC_HOST_DIR}.lock"
flock -w 1800 9 || exit 1
if [[ ! -r "$NIXL_LIBFABRIC_HOST_DIR/nixl/libplugin_LIBFABRIC.so" ||
! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/opt/amazon/efa/lib/libfabric.so.1" ||
! -r "$NIXL_LIBFABRIC_HOST_DIR/efa/usr/lib/x86_64-linux-gnu/libibverbs/libefa-rdmav59.so" ]]; then
if [[ -e "$NIXL_LIBFABRIC_HOST_DIR" ]]; then
echo "Error: incomplete NIXL LIBFABRIC cache: $NIXL_LIBFABRIC_HOST_DIR" >&2
exit 1
fi
nixl_stage=$(mktemp -d "${NIXL_LIBFABRIC_HOST_DIR}.tmp.XXXXXX")
trap 'rm -rf -- "$nixl_stage"' EXIT
mkdir -p "$nixl_stage/runtime/nixl" "$nixl_stage/runtime/efa" \
"$nixl_stage/installer"
curl -LfsS --retry 3 -o "$nixl_stage/nixl.whl" \
"https://files.pythonhosted.org/packages/8b/7c/b79fb09e832233c90f1e9d9b953e88c2b92096d968f2444839c6aa92b645/nixl_cu13-1.4.0-cp312-cp312-manylinux_2_28_x86_64.whl"
echo "3e606fbe80c39ce14899726fad0cb0fec53c6bac9f34168492692c4166b2fabb $nixl_stage/nixl.whl" | sha256sum -c -
unzip -p "$nixl_stage/nixl.whl" \
nixl_cu13.libs/nixl/libplugin_LIBFABRIC.so \
> "$nixl_stage/runtime/nixl/libplugin_LIBFABRIC.so"
unzip -p "$nixl_stage/nixl.whl" \
nixl_cu13.libs/libnuma-3387f5e3.so.1.0.0 \
> "$nixl_stage/runtime/nixl/libnuma-3387f5e3.so.1.0.0"
curl -LfsS --retry 3 -o "$nixl_stage/efa.tar.gz" \
"https://efa-installer.amazonaws.com/aws-efa-installer-1.47.0.tar.gz"
echo "2df4201e046833c7dc8160907bee7f52b76ff80ed147376a2d0ed8a0dd66b2db $nixl_stage/efa.tar.gz" | sha256sum -c -
tar -xzf "$nixl_stage/efa.tar.gz" -C "$nixl_stage/installer" \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/libfabric1-aws_2.4.0amzn1.0_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/ibverbs-providers_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/libibverbs1_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/librdmacm1_61.0-1_amd64.deb \
aws-efa-installer/DEBS/UBUNTU2404/x86_64/rdma-core/rdma-core_61.0-1_amd64.deb
while IFS= read -r -d '' efa_deb; do
dpkg-deb -x "$efa_deb" "$nixl_stage/runtime/efa"
Alternatives
No response
Additional context
No response
Before submitting a new issue...
馃殌 The feature, motivation and pitch
hi @kedarpotdar-nv @mnicely
on pytorch images, there is automatic AWS EFA & auto gcp nccl plugin detection & setup such that efa & gcp networks work out of the box. [source: https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-07.html ]
Unfortunately on NGC TRTLLM release images & NGC TensorRT LLM Develop
images, using NIXL EFA requires an massive undocumented script to download the correct EFA userspace libraries.
Any chance that EFA NIXL can work out of the box on NGC TensorTLLM release & develop images too?
We recognize that on NGC TRTLLM dynamo images there is curerntly images with efa tags but for those that wanna use the latest trtllm develop and trtllm release, efa doesnt work out of the box on trtllm
https://catalog.ngc.nvidia.com/orgs/nvidia/tensorrt-llm/containers/release/-/tags
Similar requests on vLLM Pollara & BXNT NICs where implemented in vllm-project/vllm#38687
and over the past 2 weeks, there is an global cross ecosystem effort to make efa work out of the box on NVIDIA GPUs
gaint runtime installation command from https://github.com/SemiAnalysisAI/InferenceX/pull/2993/changes/BASE..1df41b9f4a4709e7007f441e9ab0ec6aa9383514#diff-d5cf9a63d050c8f89f86e3b80f728cf8f563d148ce6a6b772ece84f5d415f243R224-R263
Alternatives
No response
Additional context
No response
Before submitting a new issue...