Skip to content

parallel_reduce_sum: drop caller scratch, keep default pool cached - #10

Merged
shsahiti merged 1 commit into
mainfrom
reduce-per-call-scratch
Sep 30, 2026
Merged

shsahiti merged 1 commit into
mainfrom
reduce-per-call-scratch

Conversation

@shsahiti

Copy link
Copy Markdown
Contributor

cub's env DeviceReduce::Sum now owns the scratch (cudaMallocAsync/ cudaFreeAsync per call), so the parallel_reduce_sum_bytes API is gone.

The default mem pool trims to 0 bytes at every sync, which made a synchronized reduce cost ~0.8 ms per call. retain_default_pool() raises the release threshold once, bringing it back to caller-scratch speed.

bench/reduce_scratch.cu compares against caller scratch and naive cudaMalloc per call.

cub's env DeviceReduce::Sum now owns the scratch (cudaMallocAsync/
cudaFreeAsync per call), so the parallel_reduce_sum_bytes API is gone.

The default mem pool trims to 0 bytes at every sync, which made a
synchronized reduce cost ~0.8 ms per call. retain_default_pool() raises
the release threshold once, bringing it back to caller-scratch speed.
@shsahiti
shsahiti force-pushed the reduce-per-call-scratch branch from f6b2569 to 7ceaef4 Compare September 28, 2026 17:17

@karl-kes karl-kes left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excellent reduction in code

@shsahiti
shsahiti merged commit 9d50722 into main Sep 30, 2026
2 checks passed
@karl-kes
karl-kes deleted the reduce-per-call-scratch branch September 30, 2026 21:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants