Skip to content

executorch: Add version 1.4.1 - #2434

Open
luhenry wants to merge 2 commits into
mainfrom
executorch
Open

luhenry wants to merge 2 commits into
mainfrom
executorch

Conversation

@luhenry

@luhenry luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

PyTorch's on-device inference runtime: a CMake pybind11 extension with the XNNPACK delegate and portable/optimized/quantized kernels. Upstream publishes no riscv64 wheel.

Mirrors upstream's build-wheels-linux.yml.

Differs from upstream

  • openssl-devel added - cmake has no riscv64 wheel and builds from source.
  • Build isolation off, torch==2.13.0 preinstalled - CMake imports torch; matches torch_pin.py.
  • auditwheel excludes torch's libraries - the torch wheel provides them at runtime.

Matrix: cp312/cp313/cp314 - torch and pytorch-tokenizers have no riscv64 cp310/cp311; upstream ships no cp314t.

Testing

  • Hand-built model replaces upstream's MobileNetV3 example - torchvision has no riscv64 wheel.

License: Wheel bundles XNNPACK/cpuinfo/pthreadpool/FP16/FXdiv/flatbuffers/flatcc/Eigen/pocketfft/nlohmann-json/pybind11 (BSD/MIT/Apache-2.0/MPL-2.0); upstream ships no licence text for them, so the build adds it.

Patches

  • 0001-packaging-... - Inappropriate, distributor-only licence bundling. Any arch.
  • 0002-cmake-relax-source-directory-name-... - Inappropriate, cibuildwheel's fixed /project path fails the name check. Any arch.
  • 0003-cmake-build-pybind-extensions-as-MODULE-... - To upstream. Configure needs Development.Embed, absent on manylinux. Any arch.
  • 0004-cmake-keep-portable_lib-SHARED-... - To upstream. Keeps the LLM runner linking after 0003. Any arch.
  • 0005-cmake-drop-FFHT-simd-arch-fatal-error-... - To upstream. Configure hard-errors outside x86_64/aarch64. Any other arch.
  • 0006-xnnpack-only-query-cpuinfo-uarchs-... - To upstream (XNNPACK). Import aborts when cpuinfo cannot initialize. riscv64/ppc64.
  • 0007-threadpool-fall-back-to-the-online-CPU-count-... - To upstream. No threadpool when cpuinfo cannot initialize; XNNPACK delegate aborts. Any arch.

luhenry added a commit that referenced this pull request Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://riseproject-dev.github.io/python-wheels/pr-preview/pr-2434/

Built to branch gh-pages at 2026-10-10 03:14 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

luhenry added a commit that referenced this pull request Sep 28, 2026
…us bracket must stay on one line

Found while triaging PR #2434 (executorch): cibuildwheel building cmake
from source (no riscv64 wheel for the pinned version) fails configure
with "Could not find OpenSSL" until openssl-devel is installed; and
check_patch.py's Upstream-Status regex has no re.DOTALL, so a bracketed
Inappropriate reason wrapped across multiple lines fails the format
check even though it reads as correctly bracketed.

luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member Author

check_patches is red and will stay red until this branch's history is rewritten — no fix-forward commit can clear it.

Root cause: ci_scripts/check_patch.py validates every commit in the PR range independently, checking each commit's own snapshot of any patch file it touched (for commit in commits: ... check_upstream_status(...)). Commit dddfb9e83d62 introduced patches/executorch/1.4.1/0001-packaging-Include-vendored-third-party-licences-in-.patch with its Upstream-Status: Inappropriate reason wrapped onto a second physical line. The script's regex has no re.DOTALL, so it only reads the tag's first line, sees no closing bracket, and fails that commit's snapshot. Commit 161c3ab77cc later folded the header onto one line — which fixes the current file content and would pass check_patches on its own — but dddfb9e83d62's original snapshot is still in history, so the job still fails on it.

Fix: squash this branch into a single commit (or rebase -i to amend dddfb9e83d62 so the patch never carries the malformed header at any point in history), then force-push. A plain follow-up commit cannot fix this.

Unrelated and already fixed on this same branch: the cmake source-directory-name build failure (patch 0002-cmake-relax-source-directory-name-check-for-cibuild.patch, commit 53206a23d4dc) — that one only needed a normal fix-forward commit and is unaffected by the above.

@luhenry
luhenry force-pushed the main branch 3 times, most recently from 39fb7ba to a75cf68 Compare October 1, 2026 15:22

luhenry commented Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Status of the riscv64 build: it compiles and passes auditwheel repair on cp312/cp313/cp314. What blocks it is a deterministic runtime abort in the smoke test (run 36518868871, all three legs, exit 134):

Error in cpuinfo: failed to parse file /sys/devices/system/cpu/cpu0/topology/core_id: "-1
" is not an unsigned number
... (same for cpu1..cpu3)
Fatal error in cpuinfo: cpuinfo_get_uarch called before cpuinfo is initialized
Aborted (core dumped) ... python riscv64_smoke_test.py
  • These runners' kernel reports core_id = -1 in sysfs (the same quirk behind gotcha 14, where torch's own cpuinfo only prints to stderr). cpuinfo's riscv64 Linux init treats that as fatal, so cpuinfo_initialize() returns false.
  • Something in the executorch wheel then calls cpuinfo_get_uarch() without checking that return value, and cpuinfo's cpuinfo_log_fatal aborts the process. The likely candidates are XNNPACK's hardware/uarch config or executorch's threadpool cpuinfo_utils, but that's not confirmed yet.
  • The gdb run meant to confirm it (36693672004, cp312 only, RelWithDebInfo) produced nothing: the job lost its runner after 2h40m inside the cibuildwheel step and no log was uploaded. I dropped that diagnostic in 55fe9ac and restored the full matrix. The smoke test now runs under python -X faulthandler, so the next abort dumps the Python frame that triggers it.

This isn't a hard platform gap. Real riscv64 boards whose device tree has a cpu-map report a valid core_id. But the wheel aborts on any kernel that reports -1 (these runners, and likely QEMU virt without a cpu-map), so it shouldn't be published like this. The fix is a patch to the vendored backends/xnnpack/third-party/cpuinfo with two parts: fall back to core_id = processor index when the sysfs topology is unreadable, instead of failing init; and guard the unchecked cpuinfo_get_uarch() caller once the faulthandler/backtrace identifies it. The cpuinfo patch is worth sending upstream to pytorch/cpuinfo as well.

check_patches: in addition to dddfb9e83d62 (0001) described above, 993ab938dd0 introduced 0003-cmake-build-pybind-extensions-as-MODULE-not-SHARED.patch with its To upstream [...] comment wrapped over 5 lines. 08ab1ec fixed the file but not that snapshot. Both commits need amending (or the branch squashed) and a force-push before check_patches can go green; a fix-forward commit can't do it.

Build executorch's pybind wheel for riscv64 (cp312/cp313/cp314) from
the v1.4.1 checkout, mirroring upstream's build-wheels-linux.yml.

Patches 0006/0007 keep the wheel usable when cpuinfo cannot initialize,
as on runners whose kernel reports core_id -1: torch's libc10.so exports
cpuinfo and interposes the copy statically linked here, so its failed
init reaches XNNPACK's unguarded cpuinfo_get_uarch() (abort at import)
and executorch's threadpool (nullptr for every caller).

luhenry commented Oct 9, 2026

Copy link
Copy Markdown
Member Author

Root cause of the cp312/cp313/cp314 smoke-test abort (run 36947661582), and the fix in 346337d.

The faulthandler trace puts the abort inside create_module for _portable_lib (portable_lib.py:80), i.e. while the extension is being loaded. XNNPACKBackend.cpp registers a static XnnpackBackend whose constructor calls xnn_initialize(), which runs XNNPACK's init_hardware_config(). There, cpuinfo_initialize() is checked before reading cache sizes, but the loop filling hardware_config.uarch[] right after it calls cpuinfo_get_uarch() unconditionally. On riscv64 that code is reached with cpuinfo uninitialized, because xnn_init_hardware_config() deliberately skips its cpuinfo gate for XNN_ARCH_RISCV (RVV is detected via getauxval). Same bug on XNNPACK master.

Patching the vendored backends/xnnpack/third-party/cpuinfo (as suggested above) would not have helped. torch 2.13.0's libc10.so exports cpuinfo_initialize/cpuinfo_get_uarch/cpuinfo_isa (40 cpuinfo_* symbols), and torch loads it RTLD_GLOBAL, so the copy statically linked into _portable_lib.so is interposed by torch's. That is also why the core_id errors print only once: torch's init already failed and is cached. torch's cpuinfo (bc3c01e) is a descendant of executorch's (f9a0324), and their riscv/linux init is identical. So the wheel has to tolerate a failed cpuinfo_initialize(), the way torch itself does on these runners.

Guarding XNNPACK alone only moves the crash: get_threadpool() returns nullptr when cpuinfo fails, then get_pthreadpool() aborts ("Failed to acquire an instance of ThreadPool!") when every XNNPACK runtime is created, and parallel_for() dereferences it. Hence two patches:

  • 0006 (XNNPACK): only query uarchs when cpuinfo_initialize() succeeds, otherwise xnn_uarch_unknown.
  • 0007 (executorch threadpool): size the pool from std::thread::hardware_concurrency() when cpuinfo fails. pthreadpool is built with PTHREADPOOL_USE_CPUINFO=0, so it does not need cpuinfo itself.

Checked locally against the v1.4.1 sources with a stub cpuinfo whose init fails and whose getters abort like the real one. Unpatched init_hardware_config() aborts with exactly cpuinfo_get_uarch called before cpuinfo is initialized (exit 134); patched, it completes with uarch unknown (XNN_MAX_UARCH_TYPES 1 and >1). Unpatched get_threadpool() returns nullptr; patched, it gives a real 4-thread pthreadpool and parallel_for computes correctly. All 7 patches apply with the workflow's git apply.

check_patches: the branch is squashed into one commit and rebased onto main, so the wrapped Upstream-Status snapshots from dddfb9e (0001) and 993ab93 (0003) are gone; check_patch.py origin/main HEAD passes locally.


Generated by Claude Code

@luhenry
luhenry marked this pull request as ready for review October 10, 2026 09:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant