From 82b6b1dcc5177fb74cd83b2e05db39fa7c0d1b1d Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Fri, 18 Sep 2026 09:21:06 +0900 Subject: [PATCH] [CUDA] Do not throw from ~CudaHandle while the driver is unloading `~CudaHandle` calls `reset()`, which checks the result through `CHECK_CUDA_ERROR` and throws on failure. A throw that escapes a destructor calls `std::terminate`, so this aborts the process rather than reporting anything. It is reachable at ordinary shutdown. The existing guard tests `cudaPeekAtLastError()`, which catches a sticky per-context error but not the CUDA runtime unloading: once teardown has begun the peek still returns `cudaSuccess`, `reset()` proceeds, `Destroy` returns `cudaErrorCudartUnloading`, and the throw fires. The symptom is a SIGABRT with `Destroy(handle_) failed: driver shutting down` after the program has otherwise finished. The destructor now releases the handle without checking. `reset()` is unchanged, so `operator=` and explicit callers keep their error checking, where throwing is legal. Nothing is lost by not reporting in the destructor: the handle cannot be reclaimed once the runtime is gone. Related: #4480 leaked the global `CommandEncoder` map for the same class of teardown problem. The `thread_local` map in `get_command_encoders()` (`device.cpp`) is not leaked, and `Worker` detaches its thread while `~CommandEncoder` only signals `stop()` without waiting, so `CudaStream`, `CudaGraph` and `CudaGraphExec` instances can still be destroyed after the primary context is released. This change makes that safe rather than fatal; whether those owners should also be leaked or joined is a separate question. --- mlx/backend/cuda/cuda_utils.h | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/mlx/backend/cuda/cuda_utils.h b/mlx/backend/cuda/cuda_utils.h index f8a234ee65..2d4d9a9e16 100644 --- a/mlx/backend/cuda/cuda_utils.h +++ b/mlx/backend/cuda/cuda_utils.h @@ -30,7 +30,13 @@ class CudaHandle { if (cudaPeekAtLastError() != cudaSuccess) { return; } - reset(); + // Not reset(): it throws via CHECK_CUDA_ERROR, and a throw escaping a + // destructor terminates. At exit the runtime may already be unloading, so + // Destroy fails and the handle can no longer be reclaimed or reported on. + if (handle_ != nullptr) { + Destroy(handle_); + handle_ = nullptr; + } } CudaHandle(const CudaHandle&) = delete;