Skip to content

Add winarm support to CI - #84

Open
jdumas wants to merge 87 commits into
mainfrom
jdumas/embree4
Open

jdumas wants to merge 87 commits into
mainfrom
jdumas/embree4

Conversation

@jdumas

@jdumas jdumas commented Feb 20, 2026

Copy link
Copy Markdown
Contributor

No description provided.

jdumas and others added 27 commits May 6, 2026 16:16
The windows-11-arm hostedtoolcache Python ships only the interpreter
binary — no include/ headers and no libs/python3XX.lib — so CMake's
FindPython falls back to x64 Python for Development.Module, which the
ARM64 MSVC toolchain cannot link against.

Use astral-sh/setup-uv + uv python install to pull a
python-build-standalone distribution instead. These distributions
include full dev files. uv intentionally defaults to x64-emulated
Python on Windows ARM64 runners (uv PR #13724), so we pin the
cpython-3.13-windows-aarch64 specifier to get the native build.

Pass -DPython_EXECUTABLE and -DPython_FIND_REGISTRY=NEVER to CMake so
FindPython uses the uv-managed ARM64 Python and does not fall back to
the registry-registered x64 Python.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
Intel MKL has no Windows ARM64 support — skip MKL/BLAS setup on that
platform. DirectSolver.h already falls back to Eigen::SimplicialLDLT
when LA_SOLVER_MKL is not defined, so no header changes are needed.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
- modules/CMakeLists.txt: revert _user_disabled_module mechanism; the
  Eigen solver fix means no modules need explicit ARM64 exclusions
- CMakeLists.txt: revert LAGRANGE_MODULE_PYTHON guard back to the simple
  LAGRANGE_ALL check (same motivation)
- embree.cmake: replace string(REPLACE) hack with git apply on a corrected
  embree-winarm.patch (fixed hunk offsets 91→92, 141→142); use
  --ignore-whitespace to handle CRLF from git-cloned sources on Windows
- simde.cmake: remove global SIMDE_NO_NATIVE for MSVC ARM64; the issue
  was specific to WindingNumber, not all simde consumers
- winding_number.cmake + winding-number-winarm.patch: fix VM_SSEFunc.h at
  source level — on MSVC ARM64, simde__m128 and simde__m128i are both
  __n128, making plain typedefs identical and breaking all overloaded
  functions (vm_shuffle, vm_extract, vm_v4sf/vm_v4si). Replace typedefs
  with distinct wrapper structs that carry implicit conversions, and guard
  the union-based MSVC cast functions with !defined(_M_ARM64)

Co-Authored-By: Claude Opus 4.7 <[email protected]>
…sse-adobe fork

The dousse-adobe fork's intrinsics.h has trailing spaces on the AVX2
preprocessor guards at lines 94 and 143. The patch's minus lines must
match exactly (git apply --ignore-whitespace only ignores trailing
whitespace in context lines, not in minus lines).

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Force-include a non-fatal eigen_assert override that logs each alignment failure to stderr (first 30 hits) instead of aborting/popping a CRT dialog. Disable Unix and non-ARM64 Windows jobs to iterate fast on the only failing config.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
Replace the temporary EIGEN_DISABLE_UNALIGNED_ARRAY_ASSERT workaround with a
non-fatal eigen_assert override that uses cpptrace::generate_trace() to print
a full stack trace whenever the failing assertion originates from
DenseStorage.h (the plain_array<> alignment check). All other eigen_assert
failures still abort to keep unrelated invariants surfaced.

Trim CI to the smallest possible reproducer:
- Disable Unix matrix.
- Windows matrix reduced to windows-11-arm Debug only.
- Configure only LAGRANGE_MODULE_SERIALIZATION2 (no LAGRANGE_ALL, no Python).
- Scope ctest to -R 'serialization2|serialize_'.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
…ERFACE

install(EXPORT Eigen_Targets ...) was failing because Eigen3_Eigen INTERFACE
linked a target outside the export set. Wrap the diagnostic compile options,
defines, and link with $<BUILD_INTERFACE:> so they only apply during the
local build and are stripped from the install/export.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
…uard

Last run produced [EIGEN_DIAG #N] header lines but no cpptrace frames followed.
Switch to to_string(false) + fprintf(stderr) so output uses the same C stream
as the header, log frame count, and catch exceptions in case symbol resolution
or unwinding fails on Win ARM64.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
MSVC ARM64 Debug does not guarantee 16-byte alignment for parameter
slots on the stack. Passing fixed-size vectorizable Eigen objects
(alignof >= 16: Vector2d, Vector4f, Matrix4f, Matrix<S,1,2>, etc.)
by value triggers the plain_array<> alignment assertion at the call
site, before the function body even runs.

Switch the following public APIs to const-ref parameters, which
sidesteps the issue and matches the rest of the codebase:

- internal::point_on_segment_2d / 3d
- mesh_cleanup unflip_uv_triangles 'update_uv' lambda
- ui::ShaderValue::operator= overloads (Vec/Mat/Affine/Projective)
- ui::Camera::set_ortho_viewport / rotate_turntable / get_frustum
- ui::AABB::intersects_ray
- ui::utils::render::compute_perpendicular_plane
- python transform_mesh binding lambda

Also drop the temporary Win ARM64 Debug eigen_assert override
diagnostic introduced to chase this bug (Eigen3.cmake block plus
eigen_alignment_diag.{h,cpp}).
Now that the underlying root cause (pass-by-value of fixed-size
vectorizable Eigen types on MSVC ARM64 Debug) is fixed in e771745,
restore the diagnostic-trimmed CI configuration to its full matrix:

- Re-enable the Unix job (Linux + macOS, all compilers + sanitizers).
- Restore Windows matrix to [windows-2025, windows-11-arm] x [Release, Debug].
- Restore LAGRANGE_ALL=ON + Python on windows-11-arm (uv-managed ARM64
  Python pinned via Python_EXECUTABLE / Python_FIND_REGISTRY=NEVER).
- Run full ctest suite (drop the serialization2-only -R filter).

Verified by run 26006218827 (windows-11-arm Debug, serialization2 only)
which passed cleanly after the const-ref refactor.
The macos-15-intel runner has only 14.3 GB RAM / 4 cores. With
EMBREE_MAX_ISA=DEFAULT, embree compiles every ISA codepath (SSE2..AVX512);
those AVX translation units, built with ASan + RelWithDebInfo under llvm@17,
exceed available memory when run 4-wide. The host then kills the runner
(reporting "lost communication with the server"), which surfaced only on the
llvm config -- apple-clang's lower per-TU footprint stayed under the limit.

Cap macOS x64 at SSE2, matching the Linux x64 ISA selection.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
The dousse/arm-forreal branch now upstreams the Windows ARM64 (_M_ARM64)
intrinsic guards for BMI/LZCNT/PEXT that were previously applied via
embree-winarm.patch. Bump the pinned commit from 03d8ec8 to c5a6207 and
remove the now-redundant local patch and its git-apply plumbing.
Boost.Context's Windows ARM64 CMake support (armclang assembler, ASM_LANGUAGE
refactor, MSVC ARM64 masm handling) that we backported via Boost.winarm.patch
has been upstream since Boost 1.89.0, so the patch is now redundant. Bump from
1.84.0 to 1.91.0-1 (using the -cmake release variant) and remove the patch.

The EMSCRIPTEN Boost.wasm.patch is kept, as BOOST_NO_FENV_H is still defined
upstream.
@jdumas jdumas changed the title Testing Embree4 + WinArm Add winarm support to CI Oct 9, 2026
jdumas added 2 commits October 9, 2026 23:40
Boost.Asio no longer pulls in Coroutine transitively in Boost 1.92. Select it explicitly so downstream consumers can still use and install the boost_coroutine target.
Boost 1.92 Asio stackful coroutine support uses Boost.Context directly. Remove the unnecessary explicit Coroutine selection and its stale indirect-dependency entry while retaining Context.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant