Skip to content

Ship a C++ SDK in the wheel - #21639

Open
shoumikhin wants to merge 1 commit into
gh/shoumikhin/91/headfrom
gh/shoumikhin/92/head
Open

Ship a C++ SDK in the wheel#21639
shoumikhin wants to merge 1 commit into
gh/shoumikhin/91/headfrom
gh/shoumikhin/92/head

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

The problem

The previous change split the runtime, kernels, delegate, thread pool and profiler into separate
shared libraries, and the wheel ships them. But nothing outside Python can use them, because the
installed CMake package names none of them. A C++ application would have to hard-code paths into
the wheel's private layout.

The headers have the same gap. The wheel installs only the subset a custom-operator build needs,
which leaves out extension/module, the entry point the documentation tells C++ callers to use. So
the wheel ships the libraries to run a model and no way to call them.

The change

Name each shipped library as a CMake component, so find_package locates them, and ship the
headers a caller needs. A component is just a name a consumer can ask for, and CMake reports a
missing one while configuring rather than at link time.

find_package(executorch 1.5 REQUIRED COMPONENTS kernels_optimized)
target_link_libraries(my_app PRIVATE executorch::runtime
                                     executorch::kernels_optimized)
component library it resolves to
executorch::runtime libexecutorch.so
executorch::kernels_optimized libexecutorch_kernels_optimized.so
executorch::backend_xnnpack libexecutorch_backend_xnnpack.so
executorch::threadpool libexecutorch_threadpool.so
executorch::etdump libexecutorch_etdump.so

Each component records where the wheel keeps its libraries, so an application built against it
finds them without the caller setting a library search path.

Headers include the module and tensor entry points, the CPU kernel helpers, the allocator and data
loader concrete classes Module's constructors take, the profiler entry points, and the
FlatTensorDataMap and MergedDataMap types plus the .ptd file header a caller writing a .ptd needs.

CMake 3.28 or newer gets these targets. Older versions do not, because they write the $ORIGIN
marker (the "look next to me" token in a library search path) incorrectly:

3.24.3, 3.27.9   Makefiles double the dollar sign, Ninja drops the name
3.28.4, 3.31.8   both write the token correctly

That would produce a target that runs where it was built and fails once the application is copied
elsewhere, so no target is defined below 3.28. Those versions get plain variables instead:
EXECUTORCH_LIBRARIES with the runtime and every shipped library by path, plus
EXECUTORCH_INCLUDE_DIRS, EXECUTORCH_COMPILE_DEFINITIONS and EXECUTORCH_CXX_STANDARD. All four
are needed, because an imported target carries the definitions and the C++ standard along with the
library and a plain path carries neither. Linking the libraries alone stops at
#error "You need C++17 to compile ExecuTorch".

ET_USE_THREADPOOL is added to EXECUTORCH_COMPILE_DEFINITIONS on the pre-3.28 route when the
thread pool library ships. Without it the runtime header supplies a local inline serial fallback
for parallel_for, so a consumer following the documented recipe linked the thread pool library
and still ran serial code with no diagnostic.

Test plan

Built the wheel, installed it into a clean environment, and built a C++ application against the
installed wheel alone:

  • the application links the runtime, runs a model, and matches eager PyTorch, and still runs after
    being copied away from the wheel.
  • asking for a component the wheel does not ship fails while configuring, naming the component.
  • a version request is honoured, including ranges.
  • shipped headers can be included on their own, and one entry point per shipped component also
    links against the shipped libraries. A small number are exempt because they need something outside
    the package: a Windows shim, a test framework, or a header that says in its own text not to
    include it directly. The exempt list is compiled too, so an entry that starts working is reported
    rather than left in place.
  • the thread pool probe compiles with ET_USE_THREADPOOL, on both the modern-CMake route (from the
    runtime target) and the pre-3.28 route (from EXECUTORCH_COMPILE_DEFINITIONS). Without it the
    header supplies a local inline definition and the probe linked identically whether or not the
    library was on the link line, so it could not detect the component being dropped. Measured both
    ways.
  • an application's runtime search path is recorded as DT_RUNPATH, not the older DT_RPATH. That
    matters because DT_RPATH is searched ahead of LD_LIBRARY_PATH and is inherited by
    dependencies, so a consumer could not point a locally built or instrumented runtime at their
    application. Verified by shadowing the runtime through LD_LIBRARY_PATH and watching the loader
    pick it up, which DT_RPATH ignores.
  • on real CMake 3.24 and 3.27, an application configures, builds and runs through the variables.
    Measured what each one contributes, with the consumer pinned to C++14 so its own standard does not
    hide the package's requirement: linking EXECUTORCH_LIBRARIES alone fails on a missing header,
    adding the include directories and definitions then fails on the C++ standard, and applying
    EXECUTORCH_CXX_STANDARD builds and loads a model. The kernels also need scoped retention there,
    because a registration-only library exports nothing the application references and the linker
    drops it, which showed up as "Missing operator" at run time rather than as a link error.
  • find_package succeeds when the interpreter on PATH is not the one the wheel was built for. The
    extension's own file name carries its suffix, so asking a different interpreter for it reported a
    complete install as not found.

Ran on Linux x86_64 and aarch64. The macOS wheel keeps the fused extension and ships no separate
libraries, so these checks do not apply there and its smoke test does not run them.

@pytorch-bot

pytorch-bot Bot commented Aug 7, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21639

Note: Links to docs will display an error until the docs builds have been completed.

❌ 3 New Failures, 5 Pending, 2 Unrelated Failures, 8 Unclassified Failures

As of commit 280b2e8 with merge base 43f89fb (image):

NEW FAILURES - The following jobs have failed:

UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 7, 2026
@github-actions github-actions Bot added ciflow/trunk module: arm Issues related to arm backend labels Aug 7, 2026
shoumikhin added a commit that referenced this pull request Aug 7, 2026
## Why

The previous change split the runtime, kernels, delegate, thread pool and profiler out
of the Python extension into five prebuilt shared libraries. The wheel ships them, but
nothing outside Python can use them: the installed CMake package config names none of
the five, so a C++ application has no way to link them without hard-coding paths into
the wheel's private layout.

    BEFORE                              AFTER

    pip install executorch              pip install executorch
      |                                   |
      v                                   v
    executorch/lib/*.so                 executorch/lib/*.so
      (shipped, but unnamed)              |
                                          v
    a C++ app must                      find_package(executorch REQUIRED)
    clone the repo and                    |
    build from source                     v
                                        target_link_libraries(app PRIVATE
                                          executorch::runtime)

## What this change does

Gives the shipped libraries a public contract:

    find_package(executorch 1.5 REQUIRED COMPONENTS kernels_optimized)
    target_link_libraries(my_app PRIVATE executorch::runtime
                                         executorch::kernels_optimized)

| target | library it resolves to |
| --- | --- |
| `executorch::runtime` | `libexecutorch.so` |
| `executorch::kernels_optimized` | `libexecutorch_kernels_optimized.so` |
| `executorch::backend_xnnpack` | `libexecutorch_backend_xnnpack.so` |
| `executorch::threadpool` | `libexecutorch_threadpool.so` |
| `executorch::etdump` | `libexecutorch_etdump.so` |

Namespaced rather than bare, because a name containing `::` must be an alias or
imported target, so CMake reports a missing one while configuring and names it. A bare
name is handed to the linker as `-lexecutorch`, which fails later with a worse message
or silently resolves to an unrelated system library. That matters more for a wheel than
for a source build: the wheel's contents depend on the options it was built with, so a
consumer asking for a delegate the wheel does not carry should be told during
configuration.

Each component target carries the retention its library needs. A registration-only
library has no symbol the application references, so the default `--as-needed` drops it
and its static initializer never runs, leaving a delegate that is linked and
unregistered. The options are scoped per library, because CMake removes duplicate
option text and a shared `--push-state` pair silently loses its scoping for the second
component.

## What to expect

Nothing changes for a Python user. This only adds a way to use the libraries the wheel
already shipped.

| | before | after |
| --- | --- | --- |
| C++ app links the runtime | build from source | `find_package(executorch)` |
| `find_package(executorch 1.5)` | any version accepted | version checked |
| headers for `Module` | not shipped | shipped |

The package also gains a version file, so `find_package(executorch 1.5 REQUIRED)`
answers correctly instead of accepting any request. Generated at packaging time rather
than checked in, because the version is only known then: `version.txt` gives the base
and a nightly overrides it. Without the file CMake reports the version as `unknown` and
accepts every request, so a consumer pinning a minimum silently gets whatever is
installed.

The headers move with the libraries. The package previously installed the subset a
custom-operator build needs, which does not include `extension/module`, the entry point
the documentation tells a C++ application to use. So the package shipped the libraries
to load and run a program and no way to call them. This adds `extension/module`, the
two directories holding the concrete allocator and loader a caller has to construct,
and `devtools/etdump`, whose library was already advertised as a component.

## Fixes from review of an earlier revision

The version file declared a variable for pinning an exact build that it never wrote, so
the config's own advice for that case compared against an empty string.

The thread pool switch sat on the thread pool target, while the header it guards is
exposed by every component and selects between a declaration and an inline definition.
A consumer naming that component in one translation unit and not another compiled two
definitions of the same function into one program, and the serial one silently won
wherever it was inlined. It now sits on the runtime, which every component depends on.

An interface link directory was carried with eleven lines defending it, while every
library already reaches the link line by absolute path. Removing it changes no build.

The relocation check skipped when `patchelf` was absent, which is indistinguishable
from a pass in the log. It now installs the tool and fails if it cannot.

Test plan:

A standalone application built from outside the wheel, in
`.ci/scripts/wheel/test_cpp_sdk.py`. It exports a real `.pte`, runs it from C++ through
`Module`, and compares the output against eager PyTorch, because a model that returns
wrong numbers without erroring satisfies every other check. Seven properties:

- `find_package` accepts the installed version and an older request, and rejects a
  newer one
- linking only the runtime loads a program and reports every operator missing, which is
  the split working rather than a defect, and it fails if the runtime starts carrying
  kernels again
- adding the kernels component runs the model and matches eager PyTorch
- adding the delegate runs a delegated model and matches
- the same delegated program fails in an application that linked the kernels but not
  the delegate, which is what shows the component is what registers it
- the application still runs after being copied away from the wheel with the absolute
  search path removed, so the package is relocatable rather than only working where it
  was built
- an application linking five components sees exactly one more backend than one linking
  two, so there is one registry in the process rather than one per component

Ran against an installed wheel on x86_64 and aarch64. All seven pass, with the C++
output matching eager PyTorch to 2.4e-07 in every executing case.

ghstack-source-id: bca8af0
ghstack-comment-id: 5215967468
Pull-Request: #21639
@shoumikhin shoumikhin added ciflow/periodic ciflow/binaries ciflow/binaries/all Release PRs with this label will build wheels for all python versions ciflow/nightly ciflow/cuda labels Aug 7, 2026
@linux-foundation-easycla

linux-foundation-easycla Bot commented Aug 11, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

  • ✅ login: shoumikhin / name: shoumikhin (42e4517)

The wheel ships the runtime, kernels, delegate, thread pool and profiler as separate shared
libraries, but nothing outside Python can use them, because the installed CMake package names
none of them. A C++ application would have to hard-code paths into
the wheel's private layout.

The headers have the same gap. The wheel installs only the subset a custom-operator build needs,
which leaves out `extension/module`, the entry point the documentation tells C++ callers to use. So
the wheel ships the libraries to run a model and no way to call them.

Name each shipped library as a CMake component, so `find_package` locates them, and ship the
headers a caller needs. A component is just a name a consumer can ask for, and CMake reports a
missing one while configuring rather than at link time.

```cmake
find_package(executorch 1.5 REQUIRED COMPONENTS kernels_optimized)
target_link_libraries(my_app PRIVATE executorch::runtime
                                     executorch::kernels_optimized)
```

| component | library it resolves to |
| --- | --- |
| `executorch::runtime` | `libexecutorch.so` |
| `executorch::kernels_optimized` | `libexecutorch_kernels_optimized.so` |
| `executorch::backend_xnnpack` | `libexecutorch_backend_xnnpack.so` |
| `executorch::threadpool` | `libexecutorch_threadpool.so` |
| `executorch::etdump` | `libexecutorch_etdump.so` |

Each component records where the wheel keeps its libraries, so an application built against it
finds them without the caller setting a library search path.

Headers include the module and tensor entry points, the CPU kernel helpers, the allocator and data
loader concrete classes Module's constructors take, the profiler entry points, and the
FlatTensorDataMap and MergedDataMap types plus the .ptd file header a caller writing a .ptd needs.

CMake 3.28 or newer gets these targets. Older versions do not, because they write the `$ORIGIN`
marker (the "look next to me" token in a library search path) incorrectly:

```
3.24.3, 3.27.9   Makefiles double the dollar sign, Ninja drops the name
3.28.4, 3.31.8   both write the token correctly
```

That would produce a target that runs where it was built and fails once the application is copied
elsewhere, so no target is defined below 3.28. Those versions get plain variables instead:
`EXECUTORCH_LIBRARIES` with the runtime and every shipped library by path, plus
`EXECUTORCH_INCLUDE_DIRS`, `EXECUTORCH_COMPILE_DEFINITIONS` and `EXECUTORCH_CXX_STANDARD`. All four
are needed, because an imported target carries the definitions and the C++ standard along with the
library and a plain path carries neither. Linking the libraries alone stops at
`#error "You need C++17 to compile ExecuTorch"`.

`ET_USE_THREADPOOL` is added to `EXECUTORCH_COMPILE_DEFINITIONS` on the pre-3.28 route when the
thread pool library ships. Without it the runtime header supplies a local inline serial fallback
for `parallel_for`, so a consumer following the documented recipe linked the thread pool library
and still ran serial code with no diagnostic.

Built the wheel, installed it into a clean environment, and built a C++ application against the
installed wheel alone:

- the application links the runtime, runs a model, and matches eager PyTorch, and still runs after
  being copied away from the wheel.
- asking for a component the wheel does not ship fails while configuring, naming the component.
- a version request is honoured, including ranges.
- shipped headers can be included on their own, and one entry point per shipped component also
  links against the shipped libraries. A small number are exempt because they need something outside
  the package: a Windows shim, a test framework, or a header that says in its own text not to
  include it directly. The exempt list is compiled too, so an entry that starts working is reported
  rather than left in place.
- the thread pool probe compiles with `ET_USE_THREADPOOL`, on both the modern-CMake route (from the
  runtime target) and the pre-3.28 route (from `EXECUTORCH_COMPILE_DEFINITIONS`). Without it the
  header supplies a local inline definition and the probe linked identically whether or not the
  library was on the link line, so it could not detect the component being dropped. Measured both
  ways.
- an application's runtime search path is recorded as `DT_RUNPATH`, not the older `DT_RPATH`. That
  matters because `DT_RPATH` is searched ahead of `LD_LIBRARY_PATH` and is inherited by
  dependencies, so a consumer could not point a locally built or instrumented runtime at their
  application. Verified by shadowing the runtime through `LD_LIBRARY_PATH` and watching the loader
  pick it up, which `DT_RPATH` ignores.
- on real CMake 3.24 and 3.27, an application configures, builds and runs through the variables.
  Measured what each one contributes, with the consumer pinned to C++14 so its own standard does not
  hide the package's requirement: linking `EXECUTORCH_LIBRARIES` alone fails on a missing header,
  adding the include directories and definitions then fails on the C++ standard, and applying
  `EXECUTORCH_CXX_STANDARD` builds and loads a model. The kernels also need scoped retention there,
  because a registration-only library exports nothing the application references and the linker
  drops it, which showed up as "Missing operator" at run time rather than as a link error. The
  smoke test now runs the same shape automatically when `EXECUTORCH_PRE_328_CMAKE` points at an
  older cmake binary, so a future change on the fallback path fails a check rather than only
  showing up on the first user with older cmake.
- `find_package` succeeds when the interpreter on PATH is not the one the wheel was built for. The
  extension's own file name carries its suffix, so asking a different interpreter for it reported a
  complete install as not found.

Ran on Linux x86_64 and aarch64. The macOS wheel keeps the fused extension and ships no separate
libraries, so these checks do not apply there and its smoke test does not run them.

ghstack-source-id: 490d066
ghstack-comment-id: 5215967468
Pull-Request: #21639
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/binaries/all Release PRs with this label will build wheels for all python versions ciflow/binaries ciflow/cuda ciflow/nightly ciflow/periodic ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: arm Issues related to arm backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant