☑️ I understand it is strictly prohibited to use AI to write issues.
Describe the bug
If the element is more than 2^31, it will reach offsets beyond the boundary, result in incorrect values.
To Reproduce
Requires approximately 5 GB of free memory. The script checks both GPU and CPU streams and verifies that the input is unchanged. Save it as repro.py.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install "mlx==0.32.2" numpy
python repro.py
# max along the last axis of a sliced view x[:, 7:] of a uint8 array with more than 2^31 elements.
# Every element of row r is r % 251, so the max of row r must be r % 251 and a wrong value names the row read.
# Needs about 5 GB of free memory.
import mlx.core as mx
import numpy as np
C, K = 65536, 7
print("mlx", mx.__version__)
failures = []
for R in (32769, 32775): # the view has fewer / more than 2^31 elements; x has more in both cases
a = np.empty((R, C), dtype=np.uint8)
a[:] = (np.arange(R) % 251).astype(np.uint8)[:, None]
x = mx.array(a)
expected = (np.arange(R) % 251).astype(np.uint8)
for name, stream in (("gpu", mx.gpu), ("cpu", mx.cpu)):
got = np.array(mx.max(x[:, K:], axis=-1, stream=stream))
bad = np.flatnonzero(got != expected)
print(f"rows={R} (view has {R * (C - K):,} elements) {name}: " + ("ok" if bad.size == 0 else
f"{bad.size} wrong row(s) {bad.tolist()}, got {got[bad].tolist()}, expected {expected[bad].tolist()}"))
if bad.size:
failures.append((R, name, bad.size))
assert np.array_equal(np.array(x), a), "the input was modified"
del x, a
assert not failures, f"mx.max over a sliced view disagrees with numpy: {failures}"
Observed behavior
Recorded output (exit code 1; only the local traceback path has been shortened):
ProductName: macOS ProductVersion: 26.6.1 BuildVersion: 25G76
Apple M1 Ultra
mlx 0.32.2
rows=32769 (view has 2,147,319,801 elements) gpu: 1 wrong row(s) [32768], got [0], expected [138]
rows=32769 (view has 2,147,319,801 elements) cpu: 1 wrong row(s) [32768], got [125], expected [138]
rows=32775 (view has 2,147,712,975 elements) gpu: ok
rows=32775 (view has 2,147,712,975 elements) cpu: 7 wrong row(s) [32768, 32769, 32770, 32771, 32772, 32773, 32774], got [1, 2, 3, 4, 5, 6, 7], expected [138, 139, 140, 141, 142, 143, 144]
Traceback (most recent call last):
File "repro.py", line 24, in <module>
assert not failures, f"mx.max over a sliced view disagrees with numpy: {failures}"
^^^^^^^^^^^^
AssertionError: mx.max over a sliced view disagrees with numpy: [(32769, 'gpu', 1), (32769, 'cpu', 1), (32775, 'cpu', 7)]
exit code: 1
Expected behavior
Both devices should return np.arange(R) % 251, cast to uint8, for every row. This is also the result of a[:, 7:].max(axis=-1) in NumPy.
Desktop (please complete the following information):
- OS Version: macOS 26.6.1 (25G76)
- Hardware: Apple M1 Ultra
- MLX: 0.32.2 (pip wheel)
- Unified memory: 128 GB
Additional context
AI assistance is was used to prepare the reproduction and analysis additional context.
- CPU: reduce.cpp stores
elem_to_loc(...) results in int variables in several reduction paths. Narrowing large offsets is a suspected cause. For R=32775, the final seven values match rows 1–7 instead of rows 32768–32774. The value 125 in the smaller CPU case is unexplained.
- GPU: failure when the view's element count is below
2^31, followed by success when it exceeds that threshold, suggests index-width selection based on element count rather than reachable offsets. This resembles #3836, but the reduction dispatch path has not been traced to confirm it.
☑️ I understand it is strictly prohibited to use AI to write issues.
Describe the bug
If the element is more than 2^31, it will reach offsets beyond the boundary, result in incorrect values.
To Reproduce
Requires approximately 5 GB of free memory. The script checks both GPU and CPU streams and verifies that the input is unchanged. Save it as
repro.py.Observed behavior
Recorded output (exit code 1; only the local traceback path has been shortened):
Expected behavior
Both devices should return
np.arange(R) % 251, cast touint8, for every row. This is also the result ofa[:, 7:].max(axis=-1)in NumPy.Desktop (please complete the following information):
Additional context
AI assistance is was used to prepare the reproduction and analysis additional context.
elem_to_loc(...)results inintvariables in several reduction paths. Narrowing large offsets is a suspected cause. ForR=32775, the final seven values match rows 1–7 instead of rows 32768–32774. The value 125 in the smaller CPU case is unexplained.2^31, followed by success when it exceeds that threshold, suggests index-width selection based on element count rather than reachable offsets. This resembles #3836, but the reduction dispatch path has not been traced to confirm it.