Bypassing Contiguous Buffer Allocation Caps for Large Diffusion/Video Pipelines on 16GB Apple Silicon (UMA Benchmarks & Workaround) #4543
Replies: 1 comment
|
Quick clarification on the post above: The benchmarks and watermark bypass (PYTORCH_MPS_HIGH_WATERMARK_RATIO) reflect the current limitations we hit when running heavy diffusion loops under PyTorch's MPS allocator on 16GB M4 nodes. The reason I brought this to the MLX discussion is to understand how MLX's native Metal allocator (mlx/backend/metal/allocator.cpp) compares when pushed past physical memory limits:
Would love to hear how the MLX architecture handles these edge cases differently from MPS. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
On Apple Silicon's unified memory pool, overriding this allocation ratio allows contiguous memory requests to overflow safely into system virtual memory / SSD swap space without abruptly crashing the driver runtime, enabling single-pass 2K video generation on baseline 16GB hardware.
2. Heterogeneous CPU/GPU Zero-Copy UMA Pointer Routing
To prevent driver deadlocks when handling unaligned precision types (such as
float8_e4m3fntext encoders) without incurring traditional PCIe copy overhead, we implemented an application-level routing layer [2, 4]:device='cpu'[4].Empirical Benchmarks on 16GB Apple M4
Below are the empirical metrics observed when running heavy diffusion loops under this heterogeneous UMA routing strategy [6, 7]:
System Stability & Throughput Overhead
float8_e4m3fnfloat8_e4m3fnfloat8_e4m3fnfloat8_e4m3fnResult: Retains ~92% of peak GPU throughput while completely eliminating driver crashes across testing pipelines [2, 8].
Memory Bus Saturation & Copy Overhead
Live Execution Proof
Short visual demo showing sampler execution and memory allocation logs running under this allocation bypass on a 16GB M4 Mac:
🎥 Execution Demo: https://youtube.com/shorts/JXFWwQ1Vhu0?si=KurgQ8Ii6IRA49Hd
Discussion Questions for the MLX Team & Community
Would love to hear thoughts, feedback, or similar benchmark observations from others experimenting with local video generation on Apple Silicon!
All reactions