[Perf] Persist Triton autotune cache to skip cold-start benchmarking - #85
Merged
Merged
Conversation
cennn
force-pushed
the
feat/autotune-at-compile-time
branch
from
September 28, 2026 15:34
ac86606 to
39a63ad
Compare
cennn
force-pushed
the
feat/autotune-at-compile-time
branch
from
September 28, 2026 17:07
39a63ad to
bde5747
Compare
cennn
force-pushed
the
feat/autotune-at-compile-time
branch
from
September 29, 2026 03:17
bde5747 to
d18bef9
Compare
cennn
force-pushed
the
feat/autotune-at-compile-time
branch
3 times, most recently
from
September 29, 2026 03:59
5a5d6b8 to
3ca29f7
Compare
Root cause: autotune_at_compile_time=False defers Triton kernel autotuning to first forward pass. Results go to Triton's in-memory cache only — Triton's disk cache for autotune results (check_disk_cache / .autotune.json) is disabled by default (knobs.autotuning.cache=False). Each new pod re-benchmarks all kernel configs: ~3-5 min overhead. Fix: Set TRITON_CACHE_AUTOTUNING=1 in _compilation_context(). This enables Triton's built-in autotune disk cache (Autotuner.check_disk_cache), which writes .autotune.json files to TRITON_CACHE_DIR. Since MagiCompiler already points TRITON_CACHE_DIR to persistent AFS, bake warmup autotune results are automatically reused by verify pods. Flow: Bake: compile → warmup forward → autotune → .autotune.json saved to AFS Verify: load .so → first forward → read .autotune.json → skip benchmark Verified with isolated 2-process experiment: 2.6x speedup on cache hit (1.97s → 0.76s for a simple kernel; real 400B model speedup expected to be much larger due to many more kernels). Also updated piecewise_compiler.py comment to document the full mechanism.
cennn
force-pushed
the
feat/autotune-at-compile-time
branch
from
September 29, 2026 04:12
3ca29f7 to
c9b3e76
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
One-line fix: set
TRITON_CACHE_AUTOTUNING=1in_compilation_context()soTriton's built-in autotune disk cache is enabled. Bake-time warmup writes
.autotune.jsonto the persistentTRITON_CACHE_DIR(AFS); subsequent coldstarts read cached results and skip kernel benchmarking entirely.
Background
standalone_compiledefaultsautotune_at_compile_time=True, but unbackedSymInt dimensions cause CUDA illegal-memory-access when Triton benchmarks run
at compile time (confirmed on PT 2.9 / B300). MagiCompiler overrides this to
False, deferring autotuning to the first forward pass.Without this fix, every new process re-benchmarks all kernel configs from
scratch (~3-5 min overhead in production). Triton already has a disk cache
mechanism (
Autotuner.check_disk_cache→.autotune.json) but it isdisabled by default (
TRITON_CACHE_AUTOTUNINGenv var unset). SinceMagiCompiler already points
TRITON_CACHE_DIRat persistent storage, the onlymissing piece was flipping this flag.
Verification
3-process controlled experiment on B300 (single kernel, 4 autotune configs):
.cubincached +.autotune.jsonread.cubincached + re-benchmarkP2 vs P3 isolates the autotune cache effect (same compilation cache, same CUDA
init). Real models with hundreds of kernels accumulate ~3-5 min savings.
Tests
6 tests in 3 classes:
TestAutotuneOverriddenToFalse— config_patches containautotune_at_compile_time=False(CPU, no GPU needed)TestTritonCacheAutotuningEnvVar—_compilation_contextsetsTRITON_CACHE_AUTOTUNING=1andTRITON_CACHE_DIR(CPU)TestAutotuneCachePersistence— bug reproduction (0.autotune.jsonwithout env) + fix verification (≥1.autotune.json+ timing comparison, GPU required)