Ship the quantized kernels as their own library - #21642
Open
shoumikhin wants to merge 1 commit into
Open
Conversation
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21642
Note: Links to docs will display an error until the docs builds have been completed. ❌ 8 New Failures, 1 Pending, 2 Unrelated Failures, 51 Unclassified FailuresAs of commit ac32360 with merge base c6213ae ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Aug 7, 2026
This was referenced Aug 12, 2026
A quantized model uses smaller numbers than a normal one, so the tensors take less memory.
Running one needs the quantized operator kernels.
The only copy the wheel shipped is the one torch loads to export a model, which a C++ application
cannot use. Such an application links the runtime, loads a quantized model, and the model fails at
run time with a missing operator, which looks like a model problem rather than a packaging one.
Build the quantized kernels as their own shared library and name it as a CMake component, the same
way the other kernel sets are named.
```cmake
find_package(executorch REQUIRED COMPONENTS kernels_quantized)
target_link_libraries(my_app PRIVATE executorch::runtime
executorch::kernels_quantized)
```
The wheel now ships `lib/libexecutorch_kernels_quantized.so`.
Note that the wheel also ships a second copy of these kernels, inside the library torch loads when
you export a model. That copy is built into the plugin rather than resolved from the shared library,
so a process holding both registers the same operators twice, and the runtime treats that as fatal:
```
Re-registering quantized_decomposed::add.out
```
This affects only a process that does both, for example an application that embeds a Python
interpreter. A plain C++ application can link the component freely.
Because of that, this is the one component `EXECUTORCH_LIBRARIES` does not include, so an
application that links whatever the package offers cannot end up in that position without asking.
A consumer that wants the quantized kernels names the component, or on CMake older than 3.28, where
no component targets exist, links `EXECUTORCH_QUANTIZED_KERNELS_LIBRARY` as well. That variable is
now populated on both CMake routes, so a consumer that adopts the older-CMake recipe and later
upgrades keeps the library on their link line instead of silently losing it.
Built the wheel, installed it into a clean environment, and:
- exported a quantized model and ran it from Python, matching eager PyTorch to within the
quantization step (measured worst difference 0.0048 against a tolerance of 0.02).
- built a C++ application that links `executorch::kernels_quantized`, ran the same program, and got
the same output as Python, byte for byte.
- confirmed the Python extension does not depend on the run-time copy, and that a process holding
the shipped library and the export plugin aborts in either load order.
- checked every shipped library the same way, to establish that this is the only pair that
collides: the CPU kernels, the delegate, the thread pool, the profiler and the runtime all
coexist with both the extension and the export plugin.
- an application linking only `EXECUTORCH_LIBRARIES` does not depend on the quantized library while
still depending on the CPU kernels, on CMake 3.28 and on real CMake 3.24. A new check asserts
this, and it fails on the previous behaviour.
- `EXECUTORCH_QUANTIZED_KERNELS_LIBRARY` resolves to the shipped library on both the modern-CMake
route (as the imported target) and the pre-3.28 route (as a file path).
- a missing quantized library now fails the checks instead of skipping them. The preset that builds
the wheel enables these kernels unconditionally, so their absence is a regression rather than a
configuration to tolerate, and both the ownership table and the C++ check previously treated it as
an acceptable state and reported coverage they had not run.
Ran on Linux x86_64 and aarch64.
ghstack-source-id: d5aa850
ghstack-comment-id: 5217087046
Pull-Request: #21642
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A quantized model uses smaller numbers than a normal one, so the tensors take less memory.
Running one needs the quantized operator kernels.
The only copy the wheel shipped is the one torch loads to export a model, which a C++ application
cannot use. Such an application links the runtime, loads a quantized model, and the model fails at
run time with a missing operator, which looks like a model problem rather than a packaging one.
Build the quantized kernels as their own shared library and name it as a CMake component, the same
way the other kernel sets are named.
The wheel now ships
lib/libexecutorch_kernels_quantized.so.Note that the wheel also ships a second copy of these kernels, inside the library torch loads when
you export a model. That copy is built into the plugin rather than resolved from the shared library,
so a process holding both registers the same operators twice, and the runtime treats that as fatal:
This affects only a process that does both, for example an application that embeds a Python
interpreter. A plain C++ application can link the component freely.
Because of that, this is the one component
EXECUTORCH_LIBRARIESdoes not include, so anapplication that links whatever the package offers cannot end up in that position without asking.
A consumer that wants the quantized kernels names the component, or on CMake older than 3.28, where
no component targets exist, links
EXECUTORCH_QUANTIZED_KERNELS_LIBRARYas well. That variable isnow populated on both CMake routes, so a consumer that adopts the older-CMake recipe and later
upgrades keeps the library on their link line instead of silently losing it.
Built the wheel, installed it into a clean environment, and:
quantization step (measured worst difference 0.0048 against a tolerance of 0.02).
executorch::kernels_quantized, ran the same program, and gotthe same output as Python, byte for byte.
the shipped library and the export plugin aborts in either load order.
collides: the CPU kernels, the delegate, the thread pool, the profiler and the runtime all
coexist with both the extension and the export plugin.
EXECUTORCH_LIBRARIESdoes not depend on the quantized library whilestill depending on the CPU kernels, on CMake 3.28 and on real CMake 3.24. A new check asserts
this, and it fails on the previous behaviour.
EXECUTORCH_QUANTIZED_KERNELS_LIBRARYresolves to the shipped library on both the modern-CMakeroute (as the imported target) and the pre-3.28 route (as a file path).
the wheel enables these kernels unconditionally, so their absence is a regression rather than a
configuration to tolerate, and both the ownership table and the C++ check previously treated it as
an acceptable state and reported coverage they had not run.
Ran on Linux x86_64 and aarch64.