Rework RHEL/manylinux tutorial for the manylinux-native build (26.08+) - #170
Conversation
build.py's rhel target expects a pypa manylinux base on 26.08+ and main; the base is manylinux_2_34 plus cuda-rhel9 rpms and gcc-toolset-13.
a5bc5d0 to
ed3deab
Compare
|
| bash -c 'auditwheel show /b/install/python/tritonserver-*.whl' | ||
| # ... is consistent with the following platform tag: "manylinux_2_27_x86_64" | ||
| # ... external versioned symbols in system libraries: libc.so.6 (GLIBC_2.2.5 ... 2.27) | ||
| find build/install -name '*.whl' # ...-cp312-cp312-manylinux_2_34_x86_64.whl |
There was a problem hiding this comment.
Wheel Contents Aren't Verified
This check only validates the wheel filename. A wheel tagged manylinux_2_34 could still depend on symbols newer than glibc 2.34 and pass this step, even though the tutorial says compatibility is verified. Retain an auditwheel show check, or another inspection of the wheel's actual symbol requirements, so the validation covers the artifact contents.
| RUN SP=/opt/pyenv_build/versions/3.12/lib/python3.12/site-packages/nvidia; \ | ||
| printf '%s\n' "$SP/cusparselt/lib" "$SP/nccl/lib" "$SP/nvshmem/lib" \ | ||
| > /etc/ld.so.conf.d/torch-cuda.conf && ldconfig | ||
| RUN dnf install -y libnccl libcusparselt0-cuda-13 && dnf clean all |
There was a problem hiding this comment.
Runtime Libraries Aren't Pinned
The Torch wheel is pinned, but NCCL and cuSPARSELt are installed without versions from a live CUDA repository. Rebuilding the same documented release can therefore select different library versions than the known-good run, reducing reproducibility and potentially causing model-load compatibility problems. Pin these RPM versions or use a versioned repository snapshot for each release row.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Overview
Since 26.08,
build.py --target-platform=rhelis manylinux-native: it expects a pypamanylinuxbase image (its/opt/_internalCPython layout and pipx shared venv) and installs its own build tooling into it, and the released RHEL artifacts moved frommanylinux_2_28tomanylinux_2_34. The previous tutorial reconstructed a pyenv-on-Rocky-8 base for 26.05 and no longer builds. This reworks it to build 26.08 and later (including currentmain) from public sources only.Changes
Dockerfile.base.rhelstarts fromquay.io/pypa/manylinux_2_34_x86_64(AlmaLinux 9) and adds the CUDA 13.4 toolkit, cuDNN and TensorRT from the publiccuda-rhel9repo, EPEL and CRB (RHEL 9's CodeReady Builder repo, the successor of PowerTools) for the-develpackages build.py installs, gcc-toolset-13 (the compiler NVIDIA's 26.08 artifacts use; EL9's RapidJSON 1.1.0 does not compile with GCC 14), a Python 3.12 pipx shared venv, and an empty libpython archive that Step 3 passes asPython_LIBRARYso the core's Python bindings link without pulling in pypa's static libpython.Dockerfile.pytorch.rhelinstalls the public torch wheel into the same image and backfills the files the backend's--image=pytorchextraction expects from NVIDIA's internal wheel image;Dockerfile.pytorch-runtime.rhelcompletes the served image with torch, NCCL and cuSPARSELt.build.pyderives versions and repo tags from the checkout); the Step 4 wheel extraction workaround is gone (26.08 stagesmanylinuxwheels itself), and each check shows its actual output.Testing
TRITON_REF=main): Steps 1 to 4, wheels taggedmanylinux_2_34, CPU serving of the python and onnxruntime models and GPU serving of the pytorch model all return[11, 22, 33, 44].mainon every nightly: the build, the CPU and GPU serve checks, and a subset of the PyTorch backend's libtorch tests (those that don't need torchvision). All pass.