Add AMD HIP/ROCm backend - #45
Open
KerelosGayed wants to merge 1 commit into
Open
KerelosGayed wants to merge 1 commit into
KerelosGayed wants to merge 1 commit into
Conversation
Add HIP build selection, runtime operations, hipCUB algorithms, hipRAND generators, and hipSOLVER linear algebra. Share the GPU test suite across CUDA and HIP and document the toolchain requirements and opt-in ROCm CI.
Contributor
Author
|
@karl-kes review pls |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds an AMD HIP/ROCm backend so applications can use xpu's allocation, buffers, structure-of-arrays storage, parallel algorithms, random generators, and optional linear algebra on AMD GPUs. The build selects one backend: CUDA takes precedence when available, HIP is selected next, and systems without either toolchain use the CPU implementation.
This is one PR because the runtime implementation, compiler configuration, and shared GPU tests together provide a usable backend.
Build configuration and requirements
XPU_ENABLE_HIP, enabled by default when CUDA is not selected. HIP targets AMD GPUs on Linux; CMake rejects a non-AMD HIP platform.hip,rocprim,hipcub, andhiprandCMake packages. Optional linear algebra additionally requireshipsolver.ROCM_PATH,CMAKE_PREFIX_PATH, and/opt/rocm. ReuseCMAKE_HIP_ARCHITECTURESfor package configuration.std::extentsin<mdspan>, which hipCUB needs. Configuration errors explain the missing dependency and how to select a CPU build.xpu::xpuandxpu::linalg. Import hipCUB's include directories without propagating its device compiler flags to every C++ source in a consuming project.For example, build and test the HIP backend for an MI300-class
gfx942target with:cmake -S . -B build-hip -G Ninja \ -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \ -DCMAKE_HIP_ARCHITECTURES=gfx942 \ -DXPU_ENABLE_CUDA=OFF \ -DXPU_ENABLE_HIP=ON \ -DXPU_ENABLE_LINALG=ON cmake --build build-hip ctest --test-dir build-hip --output-on-failureUse the architecture appropriate for the target GPU. A CPU-only build now explicitly sets both
XPU_ENABLE_CUDA=OFFandXPU_ENABLE_HIP=OFF.Backend configuration, memory, and math
XPU_HIP, the commonXPU_GPUandXPU_DEVICE_COMPILEmacros, and thexpu::xpu_hip/xpu::xpu_gpuconstants. Reject simultaneously enabled CUDA/HIP macros and HIP headers compiled outside HIP mode.cu_checkhelper to HIP, including backend-specific error messages and device-side traps for checked arithmetic failures.Parallel algorithms and random generation
hipcub::DeviceFor::Bulkforparallel_forandfill_n, with a fill functor because hipCUB does not provide the CUDADeviceTransform::Filloperation. Empty fills and ranges return without launching work.parallel_reduce_sumAPI. The HIP implementation queries its temporary-storage requirement, allocates and releases scratch in order on the default stream withhipMallocAsync/hipFreeAsync, and retains the default pool for reuse. Empty reductions overwrite the output with zero.[0, 1)interval. Sequences are backend-specific.Optional linear algebra
Share the GPU implementation between cuSOLVER and hipSOLVER's compatible dense APIs. This covers single- and double-precision LU factorization, solves, inversion, Cholesky factorization/solves, and the supporting identity/transpose operations. Matrix storage remains row-major with strides measured in elements. Cholesky scratch is passed as a
buffer_view, and diagnostics identify the selected solver.Tests, CI, and documentation
.cutest entry points as HIP, so both GPU backends exercise the same component cases and standalone-header checks../scripts/test.sh --hip. The default invocation attempts CPU, CUDA, and HIP; missing GPU compilers are skipped for automatic selection and cause an explicit error when their backend was requested. Compute Sanitizer remains specific to the CUDA suite.XPU_HIP_CI=trueand a self-hosted runner labeledlinux,x64, androcm. Like the CUDA job, it is excluded from pull-request events and runs through eligible pushes or manual dispatches.Validation
bash -n scripts/test.shgit diff --check origin/main...HEADThe local CPU build used
XPU_ENABLE_CUDA=OFF,XPU_ENABLE_HIP=OFF,XPU_ENABLE_LINALG=OFF, andXPU_ARCH_NATIVE=OFF. It also used-D_LIBCPP_ENABLE_EXPERIMENTALbecause the installed libc++ hidesstd::jthreadbehind that setting. This was a local CMake flag only. GPU execution and the hipSOLVER path still need validation on the corresponding hardware.