Ship only the TensorRT delegate in the ExecuTorch runtime wheel - #4567
Open
shoumikhin wants to merge 20 commits into
Open
Ship only the TensorRT delegate in the ExecuTorch runtime wheel#4567shoumikhin wants to merge 20 commits into
shoumikhin wants to merge 20 commits into
Conversation
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
3 times, most recently
from
August 23, 2026 19:00
44796ff to
3c104cb
Compare
shoumikhin
marked this pull request as ready for review
August 23, 2026 19:08
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
from
August 23, 2026 19:27
3c104cb to
4adc20b
Compare
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
from
August 23, 2026 19:30
4adc20b to
7cac1af
Compare
shoumikhin
marked this pull request as draft
August 23, 2026 19:31
shoumikhin
marked this pull request as ready for review
August 23, 2026 19:57
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
7 times, most recently
from
August 24, 2026 17:00
ff3379e to
a008221
Compare
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
4 times, most recently
from
August 25, 2026 06:55
49b052e to
370c368
Compare
shoumikhin
force-pushed
the
executorch-slim-runtime-wheel
branch
10 times, most recently
from
August 31, 2026 00:11
2325f28 to
613be85
Compare
ExecuTorch 1.4.1 ships no linkable C++ runtime: its wheel contains zero shared libraries, its CMake package exports only a static _portable_lib, and no CUDA wheel exists for it on any channel. That is why the runtime wheel rebuilds ExecuTorch from source today, and it is the blocker for shipping only the TensorRT delegate. The prebuilt runtime landed on ExecuTorch main on 2026-08-20, six days after 1.4.1 was tagged, so no release carries it yet. Move the pin to the nightly line that does, keeping the release-line range on installable metadata so the same range prefers 1.5.0 over any dev build the day it ships, with no edit needed. The two pins now have to name one ExecuTorch rather than two that look close, because the delegate compiles headers from the source tree and links the runtime out of the wheel. Every wheel records its source commit, so add a test asserting the pinned commit is the pinned wheel's own git_version. Nothing else was enforcing that, and a mismatch is silent: both pins look plausible and the build succeeds. Deriving the range with a three-field split raised on the nightly form, so derive it from the release line the first two fields name. ExecuTorch's CUDA wheels are published only on the PyTorch nightly index, so every install site gains that channel. No site gains --pre: the pin names an exact dev build, which pip installs from an explicit version without it, and the existing --pre uses here are for torch. CI derives the channel from the row's own CU_VERSION, which keeps the runtime the delegate links to the same CUDA build as the rest of the job.
The CI installer globs torch_tensorrt*.whl, which also matches the ExecuTorch runtime wheel, whose install_requires names a dev build published only on the nightly channel. With no index on that line the whole pip invocation failed, and because line 1's set -e is commented out the failure was swallowed and the job died later with a confusing ImportError. The two range sites installed a range against the nightly channel, which gains a member every day, so they resolved to whatever was newest while the delegate is compiled from the commit the pin names. Both now request the pin exactly, which is the pairing the drift test exists to check; it was written to skip in exactly the state the ranges produced, so nothing reported it. setup.py keeps its range, because a published requirement has to stay resolvable for users off the same line. The two shapes now differ deliberately and test_derived_requirements_match_the_pin checks each for its own. Six printed install instructions gave a bare pip install of the executorch extra, which cannot resolve a dev pin from PyPI. They name the channel now. The discovery regex saw only == and >=, so a site added with any other PEP 440 operator was invisible to the drift check. It now recognises all of them.
The executorch suite is nightly-only, and the nightly matrix runs cu132 rows as
well as cu130 ones, so the fixed cu130 channel in tests/ci/runner.py would
install a CUDA 13.0 ExecuTorch into a CUDA 13.2 job. PR builds are pinned to
cu130 by filter-matrix.py, which is why watching PR CI could never show this.
Derived from CU_VERSION now, with cu130 as the local default, matching what the
two workflow files already did.
Three example docstrings still printed a bare `pip install -e ".[executorch]"`.
That resolved off PyPI before this pin moved to a dev build; it cannot now, so
they name the nightly index too. The runtime's ImportError advice and the
reference runner README already did.
The installer's hardcoded nightly channel gets the reason written down: a .dev
wheel exists on no other channel, so deriving it from ${CHANNEL} like the lines
above would break the install on exactly the test and release runs the index was
added for.
…alls from Raising the executorch floor to a dev build made `uv lock` fail outright. The lock's win32 required-environment in pyproject.toml resolved the extra from PyPI, whose executorch stops at 1.4.1 -- uv.lock records five win_amd64 wheels for it -- so nothing satisfied the new range and uv errors rather than falling back. Reproduced against a probe project: without a marker uv reports the win32 split unsatisfiable, with one it resolves. The requirement now carries `platform_system == 'Linux'`, the shape EXECUTORCH_RUNTIME_REQUIREMENT already uses, which also stops pip reporting no matching distribution for Windows users of the extra. The delegate is a Linux object and ExecuTorch publishes CUDA wheels for no other platform, so the marker states what was already true. docgen installed the extra with --pre against the nightly channel, so it resolved through the range and took whichever dev build was newest that morning while the delegate compiled from the pinned commit. It names the pin now, read out of dev_dep_versions.yml. The pip line that installs both wheels gets `|| exit 1`. linux-test.yml concatenates this installer ahead of the user script and line 1's `set -e` is commented out, so a failure there was discarded and the job died later with an unrelated-looking ImportError; measured with `false` in place of the pip call, exit was 0 and the user script still ran. Two tests were checking source text rather than behaviour. The CUDA-row test now calls _setup_commands with CU_VERSION set and unset and reads the URL, which catches keeping the os.environ.get line while hardcoding the channel -- the mutation the string match passed. The drift check now asserts the set of files that pin ExecuTorch, because a site changing to bare `executorch` stops matching the search entirely and left the old `assert found` satisfied. Also corrects two claims: 1.4.1 does ship _portable_lib.so, so the comment says its executorch/lib carries no standalone linkable runtime, and no install site gains --pre, since an exact .dev pin needs none.
…tup fails The reference-runner README and the runtime's ImportError advice both printed `pip install "torch-tensorrt[executorch]"` with no index, and the commit that introduced the pin claimed otherwise. That claim was checked against the wrong branch: the fix existed only on the stacked runtime-wheel change, so this branch kept shipping the bare command. It matters more here than a docs nit, because this branch is what raises the floor above PyPI's newest executorch, so the bare command now cannot resolve at all. All seven printed install instructions carry the channel. The drift checks were counting the wrong thing. The requirement test asserted a set of paths, but two files carry two sites each, so either could drop one and stay in the set: turning `executorch-build-linux.yml:88` or `:128` into bare `executorch` both survived. The commit test only asserted nonzero, so any single MODULE.bazel could switch to `branch = "nightly"` unnoticed. Both now assert a per-file site count through one helper, as a minimum rather than an exact number so it holds on the stacked branch too, which removes one README site. All five mutations are caught and each names the file. Counting also surfaced a fifth commit site the nonzero check could not see: the reference-runner README's EXECUTORCH_REF shell default, correctly pinned but unaccounted for. docgen's pin was invisible to both: it is built by a shell substitution, so `$(` is not a digit and the literal search never saw it, and deleting the line survived. The derived-requirement test now runs the command docgen embeds and compares what it prints. A failed setup step printed `::warning::` and fell through to pytest. Most of the executorch suite gates on pytest.importorskip, so a failed ExecuTorch install skipped those files, left the rest passing, and reported success with a populated junit xml -- green exactly when the suite could not test what it exists to test. Driving the real run_suite with a failing setup step reproduced it, and returning the code makes it red without invoking pytest. Pre-existing, but this branch makes it likely to fire, since a nightly pin is eventually pruned from the channel. The pin tests themselves ran on nightly only, so none of this drift machinery ran on a PR or a push to main -- when a pin actually goes stale. They need no GPU, no ExecuTorch and not even torch, so they move to their own l0 suite in every lane, and the nightly suite excludes them by keyword so nothing runs twice. Also: the reference-runner README no longer says the extra installs the runtime wheel, since that requirement is commented out in setup.py.
…U lane The drift checks read derived strings and never the values CI consumes, so three ways of silently shipping no ExecuTorch all stayed green. Dropping the requirement from the runner's setup command left the step succeeding with nothing installed, after which the suite skips on importorskip; emptying EXTRAS_REQUIRE["executorch"] broke every documented `pip install "torch-tensorrt[executorch]"`; and the runtime README was recorded as carrying one pin site when it carries two, so either could go bare while the other satisfied the count -- the exact hole the per-file counts were added to close. The checks now assert the argument list the runner builds, the extras entries by AST, and the true per-file counts. All five mutations fail now. run_suite had no test at all, so replacing its `return rc` with `continue` restored the silent-green behaviour the fail-closed change exists to prevent. It is driven directly now, asserting both the propagated exit code and that pytest never runs once setup has failed. The pin suite was landing on a GPU runner: Suite.runner defaults to the matrix validation runner, so a five-second text check became one CUDA-container job per python and CUDA row, behind a wheel build. It runs in the Python lint job instead, which is already ubuntu-latest and needs none of that. The claim that it needs "not even torch" was also wrong -- tests/py/dynamo/conftest.py imports torch at module scope, which is why the lint invocation passes --noconftest. The shell tier that runs the whole executorch directory now excludes the pin file too, so the dedup claim is true of both paths rather than just the manifest one. uv.lock still records the pre-bump range with no platform marker. uv-update.yml regenerates it on pushes to main touching setup.py, and only that workflow runs `uv sync --locked`, so this breaks nothing -- but the drift was invisible, since the lock writes a bare specifier the pin search cannot match. A strict=False xfail records it and turns into a real failure via XPASS once the lock is refreshed. Editing the lock by hand was the wrong fix: its resolved entry and hashes come from a resolver run against the nightly index. Also removes internal shorthand from the PR description, and corrects a line citation for the one deliberate range in executorch-build-linux.yml.
The lint step added for these checks could not execute. It invokes pytest, and the job installs .github/scripts/requirements.txt (PyGithub) plus the lint dependency group (black, clang-format); neither carries pytest, so the step exited 1 on "No module named pytest" before running a single assertion. pyyaml is needed too, because reading the pin file shells out to a yaml import. Both are installed now, and a test asserts the step exists and installs them, since deleting it is otherwise invisible: every assertion here still passes locally while nothing runs it on a pull request. Reproduced the failure in a stdlib-only venv and confirmed the fixed command passes with only those two. The step also gets if: always(), so an unrelated formatting failure earlier in the job no longer hides the pin check. Three properties the checks are supposed to protect had no coverage: Deleting both published extras from EXTRAS_REQUIRE left everything green. The loop iterated whatever keys existed, so removing them iterated nothing and was indistinguishable from them being correct. It now requires the two published keys to be present, and only those, which also stops an unrelated future extra from turning this red for naming no ExecuTorch. The workflow opt-out marker was ordinary prose, "verify the end user's workflow". Pasting that sentence above a requirement and widening it to a range passed. It is an explicit token now, and the upward scan walks through comment lines to find it, so a cosmetic line between the opt-out and the requirement neither reclassifies the site nor fails the build. Nothing asserted that printed install instructions name the nightly channel, which is why that regressed and was re-fixed three times in this change without anything noticing. One test covers all of them by reading whole blocks rather than single lines, since every instruction wraps and the index lands on a continuation. It catches the CI install of the locally built wheel too, which carries no extra and is the site that broke most often. Generated docs under docs/ are excluded: corrections belong in docsrc/, and the committed Sphinx output is stale there independently. Also: the executorch requirement now strips its local version label like the other four, so the wheel does not bind itself to one CUDA train; the lockfile xfail is strict, since a non-strict xfail reports XPASS and ignores it and so could never fail; the fail-closed comment says it covers every setup step rather than implying only executorch; an empty frozenset and the dead branch reading it are gone; and the sys.path mutations use monkeypatch so they do not leak between tests.
The check that the two pins name one ExecuTorch could not run anywhere. It skips unless the installed wheel is exactly the pinned version, so it means something only on the nightly GPU lane, and that lane deselected it. The deselection is written as "not test_executorch_pin" to skip the source-consistency checks in the same file, but -k matches the module name in the test id, so it dropped every test in the module including this one. Both deselection sites now keep it by name. Proved it on a host with the pinned wheel installed, whose recorded git_version is the pinned commit: the check passes at the correct pins, fails when the commit pin names a different tree, and fails when the commit pin is deleted outright. Before this it was deselected in all three states. Bumping the version alone still skips, correctly, because the installed wheel is then not the one the pin names and its provenance says nothing about whether the two pins agree. A test asserts both sites keep it, since re-tightening either one to a bare module name is a small and plausible edit that would silently restore the gap.
Every guard added in this change asserted that a string appeared somewhere in a file, so each certified the state it was written to prevent. The keyword guard grepped for the kept test's name. Changing "or" to "and" in both -k expressions left it green, and that expression collects nothing at all, which is worse than the bug the guard exists to catch. Reverting the expressions and leaving the name behind in a comment also left it green, and a comment explaining the keyword sits directly above it, which is where an editor would naturally write that name. It now runs pytest's own collection under each expression and requires exactly the pairing test to come back. The CI guard searched the workflow as one blob, so it could not tell which job it was reading. The same commit that fixed the lint failure also added pytest and pyyaml to cpp-linting, which has no pin check, so deleting them from the job that does run it stayed green and would have restored the original failure invisibly. Neutralising the command while leaving its filename in a shell comment, and setting a falsy step condition, were also green. It now parses the workflow, finds the job that actually invokes pytest on this file, and requires the installs in an earlier step of that same job. The unused installs are gone from cpp-linting. The requirement pattern captured an equality prefix and stopped, so "executorch==PIN,!=PIN", a specifier that excludes the version it appears to pin, compared equal to the pin. The same truncation rejected the legal PEP 508 spelling with spaces around the operator. Requirements are parsed now and compared as specifier sets, with a check that the pinned version actually satisfies them. The site scanner counted raw search hits, so gutting a pin to a bare "executorch" while putting the exact pin in a comment in the same file kept the per-file minimum satisfied. Comments no longer count, except in the bazel repositories, where the annotation beside the pinned commit is the only record of which wheel that commit belongs to. Also corrected two claims this change made: the executorch tier is reachable from a pull request through executorch-test-linux.yml as well as the nightly manifest, so it is not the only route, and the shell helper now says why one test is kept out of the deselection.
The uv.lock check was a strict xfail. uv.lock records ">=1.4.1,<1.5" while the pin derives ">=1.5.0.dev20260822,<1.6", so the assertion fails and the xfail is satisfied. Refresh the lock and the assertion passes, and a strict xfail reports that pass as a failure. The lint step runs this file with if: always() on every pull request, so one lock refresh would have made the lint job red on every subsequent pull request, for a file none of them touched, until someone edited this test. Measured: baseline 1 xfailed, and 1 failed once the specifier is bumped. My own docstring claimed the lock is machine-generated and not edited by hand. Two hand refreshes landed on 2026-08-23, inside ordinary version-bump changes, so that was wrong as well. It now accepts both resting states and only fails where something is actually wrong: a recorded range whose lower bound is above the pin, which means the lock names an ExecuTorch this repository does not pin. Behind the pin passes, the derived range passes, and ">=1.6,<1.7", an open-ended ">=1.7" and "==1.9.0" all fail. Comparing lower bounds rather than probing the specifier with sample versions: an upper-bound test missed the open-ended case, and a low sentinel version called the ordinary behind-the-pin state a failure.
test_derived_requirements_match_the_pin extracted the python3 -c one-liner from docgen.yml and ran it. Whatever that line said got executed on every pull request: rewriting it to write a file left the test green and the file written. Same class as the bash -c problem fixed in test_api.py last round, still live here. It now compares the command as text against the exact form that reads __executorch_version__ out of dev_dep_versions.yml. Four mutations caught, including a payload that writes a file and still prints the right version, with nothing executed. The CI reachability guard tested the raw string for "--collect-only", so it accepted "--co", pytest's own documented short form, which collects and asserts nothing. It also could not see an exit status being discarded. Now tokenised: --collect-only, --co, -h, --help, a "||" short-circuit and continue-on-error are all rejected, and all five are caught where four previously survived. The comment exemption for .md/.rst/.txt defeated exactly the threat its docstring names. Install commands live in prose files, so exempting them made a comment count as a pin there: the runtime README's install line gutted to a bare "executorch" passed as long as a decoy "# executorch==<pin>" sat beside it, and failed only with no comment present. The exemption is gone, and trailing comments no longer count either, since a decoy after a live requirement on the same line kept the per-file count satisfied. Five mutations caught, baseline green.
…it resolves The nightly-index guard matched only the named-distribution spelling, so the four sites that write "pip install .[executorch]" were unguarded: docgen.yml and the three export examples. The nightly index could be deleted from all four with the test green. Each of the four is now caught individually. Its second half was a bare substring test for the host, which proves a string sits nearby rather than that the instruction resolves. Rewriting every channel in the tree, 18 files, to a nonexistent cu999 left it green. The CUDA suffix is now checked against the set the project publishes for. Deliberately not compared against __cuda_version__: five sites legitimately say cu130 while the pin says 13.2, and I confirmed against the live index that cu130 and cu132 both carry 38 ExecuTorch wheels while cu999 carries none.
The printed install commands resolved no ExecuTorch. "torch-tensorrt[executorch]" with no version pin resolves the stable PyPI wheel, which carries no executorch extra, so the command exited 0 and installed nothing the feature needs. Add --pre to the six commands that name the extra and assert its presence in the guard that already reads them. Close four ways to neutralise the pin check while its guard stayed green: a ";" or "&" terminator after pytest, continue-on-error or a falsy if: on the owning job, and reducing the workflow trigger so it never runs on pull requests. The trigger check also handles PyYAML reading the unquoted "on" key as the boolean True. Close both ways to strip the pairing check while its guard stayed green: assert the workflow actually calls trt_tier_executorch, and validate suite lane names against the known set so a typo raises at import instead of silently dropping the suite from every matrix. Also: anchor the docgen pin check to a live line so a commented-out install no longer satisfies it; fix the lockfile range check crashing on a legal "==1.4.*" clause; correct the range comment to describe what the range admits; and note in the install advice that the feature is published for Linux only.
The delegate is built against one ExecuTorch: __executorch_version__ selects the wheel it links against and __executorch_commit__ selects the tree it compiles from. Those two values repeat across the build workflows, the bazel modules, the docker and toolchain copies, and the docs, so they can drift apart or fall behind upstream with nothing to notice. Add a script and a daily workflow that move both pins to the newest ExecuTorch wheel on the nightly index. The source commit is read from the chosen wheel's own version.py, so the two pins always name one ExecuTorch rather than two that happen to be close. The update lands as a pull request, so the pin consistency checks and the delegate build and test lane decide whether the new wheel is usable before it reaches main. A day with no new nightly rewrites nothing and opens nothing. On a release branch the schedule is a no-op and the pin moves only by a manual run pointed at the stable line, so a cut release does not drift. Back the mechanism with consistency checks that run under the linter. Every requirement and comment that names ExecuTorch is asserted to match the pinned version, including the variable-index install once the variable's assignment is resolved and extensionless install files like justfile. The source commit is checked against the wheel's own provenance wherever that wheel is installed, and commits left in comments are not mistaken for pins. The wheel-content and CI-invocation checks measure effect, running the workflow's own step against a passing and a failing stub and requiring the exit status to follow, rather than enumerating bypass spellings. Install the built wheel in the runtime README rather than an unpublished package. The guard that checks the pin runs in CI compares the step's command as text rather than executing it. Running the step's own shell body meant whatever that body said ran on every pull request: appending a line that writes a file left the test green and the file written. That is the same defect this file already avoids for the docgen one-liner, and the reasoning there applies here too. The delegate claims both names ExecuTorch has used for its pybind extension. It renamed _portable_lib to _C, and portable_lib.py imports whichever its own version carries, so aliasing only the old name is silently ineffective at the new pin: nothing imports it, the stock extension loads, and the backend is never registered. CI reported that as "TensorRTBackend is not registered" from the native runtime check.
The pin now names executorch 1.5.0.dev20260901 and the source commit that wheel records for itself. Every pin site moves together, which is what the pin checks assert: a version bumped in one place and not another is the failure mode they exist to catch. Carries the pin-site changes CUDA 12.6 support brought with it. The release lane builds the runtime wheel, so it installs ExecuTorch and is a pin site: it arrived naming a stale release off the default index, which resolves no ExecuTorch at all, and is now the pinned nightly from the nightly channel, registered in both the guard and the bumper so a future bump moves it too. Two guard bugs of my own that this surfaced. The trailing-comment strip cut at the first "//", so any line carrying an index URL was truncated before its requirement and the site read as missing rather than as wrong; it now skips a "//" that follows a colon. And the install-command helper still passed --upgrade, which on a named requirement replaces a user's released torch_tensorrt with a nightly when all they asked for was the extra. Also drops an assertion that pinned the runtime wheel's TensorRT distribution to the literal tensorrt-cu13. That was right while only CUDA 13 shipped and wrong once 12.6 returned, since a cu126 row would then declare the CUDA 13 distribution; the value is resolved from the build's own CUDA instead.
Read the commit from the wheel's own version.py rather than assuming it, and applied with
the repository's own writer so every pin site moves together.
1.5.0.dev20260902 5afeaa8130f68f2afa800e0743d4a73aec79bf15
Test plan: 12 sites rewritten, 11 files naming the new version, zero references left to
either the old version or the old commit. Pin coherence suite passes, 17 passed 1 skipped.
The rewriter's trailing boundary excluded '+', so a requirement written as executorch==<pin>+cu130 did not match and the bump left it on the old version. The guard's own requirement pattern does accept that spelling, so such a site would be counted as a pin and skipped by the rewriter, which is the shape that fails the generated pull request as a mismatch rather than as an operator the rewriter cannot see. No live site is written that way today, so this is latent rather than a present break. Test plan: exercised the rewrite across the label-free pin, the labelled pin, the range form with and without a label, the YAML key, and a bystander bazel_dep version. The label survives the bump rather than being dropped, and the bystander is untouched.
A new release workflow for aarch64 arrived with an ExecuTorch requirement of its own, and
the guard caught two problems with it at once.
It was not on the list of places the pin lives, so a nightly bump would have moved the ten
other sites and left this one behind. Registering it fixes that, and the same registration
is what makes the guard check it from now on.
It also asked for a different version from every other site, and asked for it without naming
the nightly channel:
python -m pip install pyyaml "executorch==1.4.1"
That resolves against the default index, which carries no ExecuTorch nightly at all, so the
release build either picks up an unrelated release or fails outright. Its x86_64 counterpart
already had the right shape, so this copies that exactly rather than inventing a third
spelling.
Test plan: the pin suite goes from two failures to 17 passed, and the two failures were the
real ones, a site absent from the census and a requirement that excludes the pinned version.
Ran the bumper afterwards and confirmed the new file moves with a simulated bump and comes
back cleanly, leaving no stale version anywhere in the tree.
Read the commit from the wheel's own version.py rather than assuming it, and applied with
the repository's own writer so every pin site moves together.
1.5.0.dev20260904 9379a885af0544c2ae87bb19b341c1c2b58d80c9
The version is taken from a CUDA channel and stored without its local label. A published
wheel is labelled by the CUDA build it came from, for example 1.5.0.dev20260904+cu130, and
keeping that label would bind every row to one CUDA version. Dropping it lets the same pin
resolve on each CUDA row, which is what the install lines do when they read the channel from
the environment.
Test plan: 13 sites rewritten, 12 files naming the new version, and no ExecuTorch line left
on the old version or the old commit. Pin coherence suite passes, 17 passed 1 skipped.
The torch-tensorrt-executorch-runtime wheel shipped a full ExecuTorch Python
runtime alongside the TensorRT delegate. This ships only the delegate: a single
shared library that registers TensorRTBackend with the ExecuTorch runtime that
the executorch distribution already provides, rather than bundling a second copy
of that runtime. Shipping a second copy is also what made the old wheel prone to
a libstdc++ clash, because two C++ runtimes could end up in one process.
The native build produces just the delegate library, its RUNPATH points at the
executorch package the delegate links against, and setup.py packages the one
shared object. The runtime dependency stays commented out in the top-level
setup.py because the delegate wheel is not published to any index yet, so the
docs and the load-time and save-time errors direct users to build it from
py/torch-tensorrt-executorch-runtime/README.md.
The delegate links the C++ runtime dynamically, the way every other shared
object in the process already does. The build toolchain is newer than the
libstdc++ on a user's machine, so an optimized build emits out-of-line calls
into the newer runtime, for example std::string::_M_replace_cold. Naming stdc++
as a link library puts the reference after the objects, where the toolchain's
own libstdc++.so linker script resolves it: the old, stable symbols bind
dynamically to the system libstdc++.so.6 and only the newer helpers are pulled
statically from the toolchain's companion archive. The delegate ends up needing
no C++ runtime version above what the ExecuTorch it loads beside already needs.
A static C++ runtime is deliberately avoided: this library is loaded next to
libtorch and ExecuTorch, and a private libstdc++ would give it its own exception
type_info and locale state, which breaks exceptions and dynamic_cast across the
boundary. The build guard checks the shape: the delegate keeps a dynamic
libstdc++ dependency, has no unversioned C++ runtime symbol left undefined, and
requires no symbol version above the paired runtime.
The wheel is tagged py3-none rather than per-interpreter, because the delegate
is a plain shared object with no Python ABI and one build serves every CPython.
test_api.py checks the shipped layout: the delegate resolves through the loader
in the layout that ships, the wheel's RUNPATH is compared whole against the one
the build asks for, the symbol versions and the C++ runtime dependency are
compared against the runtime the delegate links, and the wheel's own metadata is
checked. The reachability scans that assert the import and static-C++ checks run
in CI parse each language's grammar rather than matching text, and none of them
execute the workflow they inspect.
The wheel exposes no runtime API at all. Loading and running a program belongs to
ExecuTorch, which already ships Runtime, Program and Method, so the Python wrapper
this wheel used to carry is gone along with the load(format="executorch") entry
point that reached it. That wrapper duplicated ExecuTorch's own classes down to the
line that keeps the file buffer alive, and its CPU copy of top-level inputs quietly
defeated programs exported for device-resident inputs. A consumer now imports this
package and uses executorch.runtime directly.
Registration happens on import, so there is nothing to call. ExecuTorch's own
delegates register because they are linked into its pybindings extension, and
loading that extension pulls them in; a delegate in a separate wheel cannot join
that link and ExecuTorch has no discovery hook for out-of-tree backends, so this
package performs the equivalent step itself. A load it cannot complete raises from
the import rather than being swallowed, because the diagnosis here names the real
cause, a CPU-only ExecuTorch wheel or an ABI mismatch, which a later "backend not
available" cannot. TORCH_TENSORRT_SKIP_DELEGATE_REGISTRATION=1 imports the module
without the side effect, for tooling that wants the metadata only.
The wheel now follows the layout ExecuTorch uses for its own backends, so the
TensorRT delegate is an out-of-tree sibling of them rather than a Python-only
artifact. The shared library moves to lib/, next to where executorch keeps
libexecutorch_backend_cuda.so and friends, and the wheel ships a CMake package
under share/cmake so a C++ app can link it:
find_package(executorch REQUIRED COMPONENTS backend_cuda)
find_package(torchtrt_executorch REQUIRED)
target_link_libraries(app PRIVATE executorch::runtime torchtrt::executorch_backend)
Before this the shared library was reachable only from Python, even though it is
a drop-in sibling of ExecuTorch's backends: same naming, same soname convention,
register_backend imported rather than defined. What was missing was the discovery
layer, so the only way for C++ to get the delegate was add_subdirectory against a
source checkout of this repository.
The imported target links with --no-as-needed, bracketed by push-state and
pop-state. Nothing in a consumer references a symbol the delegate defines, so the
default would drop the dependency and the backend would never register: the app
would build, load the program, and fail with an unregistered backend. That is the
shared-library counterpart of the --whole-archive the in-repo source build needs
for the same reason. No headers ship, because a consumer calls no Torch-TensorRT
code; registration happens in the library's static initializer and the rest is
ExecuTorch's runtime API.
Moving the library under lib/ also moves what $ORIGIN means, so the delegate's
own RUNPATH gains a level: $ORIGIN/../../executorch/lib rather than
$ORIGIN/../executorch/lib, and likewise for tensorrt_libs and nvidia/cu13/lib.
Without that the entries resolve inside the package directory instead of
site-packages, the delegate cannot find libexecutorch.so, libcudart or libnvinfer,
and a C++ consumer fails to link it with undefined references to cudaMemcpyAsync
and friends. The depth and the install location are one decision, so the test that
reads the declaration now rejects the single-level form it used to require.
The CMake package installs to lib/cmake/torchtrt_executorch, which is where ExecuTorch
puts its own: find_package resolves executorch from
site-packages/executorch/lib/cmake/executorch, so following that layout rather than
share/ means a consumer points CMAKE_PREFIX_PATH at the two package roots and both
resolve the same way. The walk that locates the package root now looks for the delegate
itself instead of for a directory named lib, because the config now lives inside lib/ and
stopping at the first lib/ it meets would set IMPORTED_LOCATION to that directory.
The CMake package test now configures the package with real CMake and asks for the imported
target back, because a string search over the config cannot tell a working package from a
broken one: inserting return() after cmake_minimum_required makes the config define nothing
and every string assertion still passes. The README command locates ExecuTorch through its
distribution metadata, since it is a namespace package whose __file__ is None, so the
documented one-liner raised TypeError before CMake ran.
The delegate builds for CUDA 12 as well as CUDA 13, because torch-tensorrt publishes both
channels. Three places assumed one major. The version check now accepts either, since a minor
bump inside a major does not change the ABI the delegate links. The RUNPATH carries both
layout directories, because the two majors package their runtime differently: the CUDA 13
wheels install nvidia/cu13/lib while the CUDA 12 wheels install nvidia/cuda_runtime/lib. The
artifact check maps the CUDA runtime the delegate asks for to the directory that carries it and
fails when the RUNPATH has no matching entry, which is the case that would link cleanly and then
find nothing at load time.
The symbol version ceiling is compared against the manylinux platform the wheel ships under
rather than against the ExecuTorch distribution beside it, and that platform is passed in per
architecture because the two rows use different builder images. A symbol version requirement is a
floor on the host, not a ceiling a library imposes on its neighbours: two libraries in one
process may need different versions, and the loader only needs the host to satisfy the highest.
Comparing against the sibling rejected the delegate wherever TensorRT itself was built with a
newer toolchain than ExecuTorch, which is the case on aarch64 today, and by that rule the check
would reject TensorRT too.
The delegate is built for aarch64 as well as x86_64, matching the architectures the
torch-tensorrt wheel it pairs with already ships. The native build already selected the right
TensorRT per architecture; what was missing is that the only caller generated an x86_64 matrix,
and the architecture input defaults to x86_64, so nothing ever asked for the other rows. The
aarch64 workflow now calls the same build against its own matrix, ordered after the job that
uploads the wheel it downloads, and deliberately outside that workflow's gate so a delegate
failure cannot block pull requests that have nothing to do with the delegate.
The check that the downloaded wheel carries the C++ runtime looks for the library instead of
importing the compiler package. That import reaches torch.cuda.get_device_capability() while
deciding whether it is running on Tegra, so it needs a GPU, and the aarch64 builder has none.
The symbol version cases in the guard's own test pass the manylinux tag, without which the
ceiling is skipped and every one of them passes for the wrong reason. Three of them asserted the
old rule, that a version above the ExecuTorch distribution's own is a rejection, and now expect
the artifact to be accepted: all three sit below what the platform guarantees, and the host
provides the C++ runtime rather than the sibling wheel.
The export and the reference runner run only where a GPU is present. Both compile and execute a
TensorRT engine, and the aarch64 builders are CPU-only instances, which is why the wheel's own
aarch64 lanes build without running their tests. The delegate is still built and checked on
aarch64; its runtime behaviour stays covered by the x86_64 rows, which have a GPU. Keyed on
whether the device is usable rather than on the architecture, so a GPU runner never skips it.
The delegate follows the main wheel's CUDA versions, CUDA 12.6 included. Both read the
same matrix filter, so the rows agree by construction: 25 rows, with cu126 on x86_64
only and the Arm rows on CUDA 13, which is what the main wheel publishes. The TensorRT
distribution is resolved from the CUDA the build actually uses rather than hardcoded, so
a cu126 row declares tensorrt-cu12 and a cu13 row declares tensorrt-cu13. The RUNPATH
and the ELF guard already carried both CUDA layouts, nvidia/cuda_runtime/lib for 12 and
nvidia/cu13/lib for 13, so no packaging change was needed for the new rows.
Test plan:
Ran the real filter from this branch against a full three-CUDA, five-Python,
two-architecture input and compared it against main's filter on the same input: both
return the same 25 rows, and the two filter files are byte-identical, so the rows agree
by construction rather than by a list kept in step by hand. The test asserts the x86_64
rows keep CUDA 12.6 and says nothing about which CUDA versions the Arm rows carry, so
Arm gaining 12.6 upstream needs no edit here. Added a test that asserts cu126 is present on x86_64, absent on
aarch64, and that both CUDA 13 rows survive on each architecture. It fails when cu126 is
removed from the x86 list. Replaced the test that asserted a hardcoded tensorrt-cu13,
which would have kept passing while a cu126 row declared the wrong dependency; the
replacement fails when the resolution is hardcoded again.
The release lane installs ExecuTorch too, so it is a pin site and now names the pinned
nightly from the nightly channel. It arrived with the CUDA 12.6 rows naming a stale
release off the default index, which resolves no ExecuTorch at all, and the pin checks
caught it.
A URL is no longer mistaken for a comment. The trailing-comment strip cut at the first
"//", so any line carrying an index URL was truncated before its requirement and the
site was reported as missing rather than as wrong. It now skips a "//" that follows a
colon, which is a scheme rather than a comment.
The channel variable that install reaches through is asserted to stay in scope. The URL
is built from CU_VERSION, which the reusable build workflow exports from the matrix row;
if that export is renamed the URL collapses to a channel that does not exist, pip falls
back to the default index, and the install resolves the wrong ExecuTorch without failing.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
The
torch-tensorrt-executorch-runtimewheel now ships one thing: the TensorRT backend, as a prebuilt shared library. Nothing else.The goal is that it looks like one of the backends ExecuTorch already ships in its own wheel, such as the CUDA backend, only built and distributed separately. Same file naming, same layout, same linking rules. A Python user registers it by importing the package. A C++ app links it through CMake.
The problem
Two copies of the same runtime in one process
The old wheel carried its own
libexecutorch.soand its own copy of ExecuTorch's Python extension, then took overexecutorch.extension.pybindingsat import time.That works only while both copies match. Two C++ runtimes in one process can also pull in two different
libstdc++versions, which breaks in ways that are hard to read. It also meant the wheel had to get in first, before anything imported ExecuTorch, or the wrong copy won.An inference API that already existed
The wheel shipped a
runtime.pywith aProgramclass and aload()function. ExecuTorch already exportsRuntime,ProgramandMethod. Most ofruntime.pywas the same code written again, down to the line that keeps the file buffer alive so the program is not freed.One part was worse than redundant.
Program.run()copied a CUDA input to CPU every time. If you exported a program that keeps its inputs on the GPU, that copy quietly undid it, and you could not get a device-resident input through the wrapper at all.No way to use it from C++
The shared library was reachable only from Python. ExecuTorch ships each of its own backends as a prebuilt library plus a CMake package, so a C++ app writes
find_packageand links a target. This wheel shipped no CMake package, so the only way to get the delegate into a C++ program was to build this repository from source.The fix
The wheel is a backend library and its loader
This mirrors what ExecuTorch does with its own backends:
lib/cmake/<name>/is wherefind_packagelooks under a prefix, and it is whereExecuTorch's own package resolves from, so a consumer points
CMAKE_PREFIX_PATHat thetwo package roots and both are found the same way.
The library also matches theirs where it counts: its
SONAMEis its own filename, it hasDT_RUNPATHand noDT_RPATH, it has noPyInit_because it is not a Python extension, and it importsregister_backendrather than defining it. It linkslibexecutorch.sofrom the installedexecutorchwheel instead of bringing its own, so there is one runtime in the process.One thing differs on purpose. ExecuTorch's backend sits in the same wheel as
libexecutorch.so, so its search path reaches it with one../. This one sits in a different wheel, so it has to climb out tosite-packagesand back down, which takes two:Python: importing the package registers the backend
There is no API to call:
ExecuTorch's own backends register because they are linked into its Python extension, so loading that extension pulls them in and their static initializers run. A backend in a separate wheel cannot join that link, and ExecuTorch has no discovery hook for out-of-tree backends, so this package performs the equivalent step itself at import time.
The library exports no
PyInit_, so a plainimportcannot load it; something has todlopenit. That is allregister()does, and it runs once on import.If the load fails, the import raises with the real cause, for example a CPU-only
executorchwheel or an ABI mismatch. Failing loudly is on purpose: this wheel exists only to register the backend, so a load it cannot finish leaves nothing useful behind, and ExecuTorch's later "backend not available" cannot name the cause. SetTORCH_TENSORRT_SKIP_DELEGATE_REGISTRATION=1to import without the side effect, for tooling that only wants the metadata.C++: link it the way you link an ExecuTorch backend
Point CMake at both wheels, since they are separate packages, and use CMake 3.28 or newer because the
backend_cudacomponent requires it:cmake -DCMAKE_PREFIX_PATH="<site-packages>/executorch;<site-packages>/torch_tensorrt_executorch_runtime" ...There is nothing to include. The backend has no public header: it registers itself from a static initializer inside the shared library, and everything after that is ExecuTorch's own runtime API. The CMake target links the library with
--no-as-needed, so the dependency survives even if the consumer never names a symbol from it.Loading a
.pteis ExecuTorch's jobtorch_tensorrt.load(path, format="executorch")is removed, along with theformatargument. This matches how the other save formats already work: Torch-TensorRT saves the file, and the framework that owns the runtime loads it..pt2(AOTInductor)torch_tensorrt.savetorch._inductor.aoti_load_package.pte(ExecuTorch)torch_tensorrt.saveexecutorch.runtime.Runtimeoutput_format="executorch"ontorch_tensorrt.saveis unchanged. Only the load side moved.Device-resident inputs now work
With the copy in
Program.run()gone, nothing in the Python layer touches your tensors, so a program exported to keep its inputs and outputs on the GPU keeps them there. Two settings are needed for that export, not one:Skipping the copy is not enough on its own. Memory planning allocates graph inputs and outputs by default, so the runtime would still reserve its own buffer and fill it from your memory with a host copy, which puts the copy back.
examples/torchtrt_executorch_example/export_device_resident.pyexports such a program and checks the result rather than trusting the flags: it reads the operator table of the saved file and fails if either boundary copy operator is still there, and it checks that every method input and output is recorded as a CUDA tensor.Breaking changes
These names are gone from the runtime package:
runtime.py, includingProgramandload(). Useexecutorch.runtime.Runtime.get_runtime(). Import the package, then useRuntime.get().activate(), which is nowregister()and is called for you on import.torch_tensorrt.load(..., format="executorch"). Theformatargument no longer exists; passing a value raisesTypeErrornaming the replacement. PassingNone, the old default, still works.torch_tensorrt_executorch_runtime._portable_liband.data_loader, because the wheel no longer ships them.The delegate also now needs a CUDA build of the
executorchwheel at runtime, not only at build time, because it links a library that only the CUDA wheels ship. With a CPU-onlyexecutorchinstalled, the import fails with a message that says so.Test plan
Run on Linux x86_64 with CUDA 13 and an NVIDIA H100, against the wheel this change builds in CI:
TensorRTBackend, with nothing else called.et_copy::_h2d_copyandet_copy::_d2h_copy, the two operators the device-resident program is asserted not to have. Without that check the assertion could pass because those operators never appear.Not covered: running a device-resident program from C++. That program requires the caller to own both the input and the output buffers in device memory, and the C++ example here supplies host buffers.
CI builds the wheel, checks its contents and its search paths, runs the C++ reference runner against three saved programs, and runs the Python runner. Unit tests: 24 in
test_python_runtime.py, plus the pin and updater suites.