Skip to content

[Perf] Persist Triton autotune cache to skip cold-start benchmarking - #85

Merged
jiahy0825 merged 1 commit into
mainfrom
feat/autotune-at-compile-time
Sep 29, 2026
Merged

jiahy0825 merged 1 commit into
mainfrom
feat/autotune-at-compile-time

Conversation

@cennn

@cennn cennn commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

One-line fix: set TRITON_CACHE_AUTOTUNING=1 in _compilation_context() so
Triton's built-in autotune disk cache is enabled. Bake-time warmup writes
.autotune.json to the persistent TRITON_CACHE_DIR (AFS); subsequent cold
starts read cached results and skip kernel benchmarking entirely.

Background

standalone_compile defaults autotune_at_compile_time=True, but unbacked
SymInt dimensions cause CUDA illegal-memory-access when Triton benchmarks run
at compile time (confirmed on PT 2.9 / B300). MagiCompiler overrides this to
False, deferring autotuning to the first forward pass.

Without this fix, every new process re-benchmarks all kernel configs from
scratch (~3-5 min overhead in production). Triton already has a disk cache
mechanism (Autotuner.check_disk_cache → .autotune.json) but it is
disabled by default (TRITON_CACHE_AUTOTUNING env var unset). Since
MagiCompiler already points TRITON_CACHE_DIR at persistent storage, the only
missing piece was flipping this flag.

Verification

3-process controlled experiment on B300 (single kernel, 4 autotune configs):

Process Condition Elapsed
P1 (warm-up) compile + autotune 1.62s
P2 (cache HIT) .cubin cached + .autotune.json read 0.68s
P3 (cache MISS) .cubin cached + re-benchmark 1.10s

P2 vs P3 isolates the autotune cache effect (same compilation cache, same CUDA
init). Real models with hundreds of kernels accumulate ~3-5 min savings.

Tests

6 tests in 3 classes:

  • TestAutotuneOverriddenToFalse — config_patches contain autotune_at_compile_time=False (CPU, no GPU needed)
  • TestTritonCacheAutotuningEnvVar — _compilation_context sets TRITON_CACHE_AUTOTUNING=1 and TRITON_CACHE_DIR (CPU)
  • TestAutotuneCachePersistence — bug reproduction (0 .autotune.json without env) + fix verification (≥1 .autotune.json + timing comparison, GPU required)

@cennn
cennn force-pushed the feat/autotune-at-compile-time branch from ac86606 to 39a63ad Compare September 28, 2026 15:34
@cennn cennn changed the title [Perf] Expose autotune_at_compile_time config to eliminate first-forward autotune overhead [Docs] Document autotune_at_compile_time=False rationale and fix paths Sep 28, 2026
@cennn
cennn force-pushed the feat/autotune-at-compile-time branch from 39a63ad to bde5747 Compare September 28, 2026 17:07
@cennn cennn changed the title [Docs] Document autotune_at_compile_time=False rationale and fix paths [Perf] Per-subgraph autotune_at_compile_time: bake autotune into .so for subgraphs without unbacked SymInt Sep 28, 2026
@cennn
cennn force-pushed the feat/autotune-at-compile-time branch from bde5747 to d18bef9 Compare September 29, 2026 03:17
@cennn cennn changed the title [Perf] Per-subgraph autotune_at_compile_time: bake autotune into .so for subgraphs without unbacked SymInt [Perf] Persist Triton autotune results to eliminate cold start overhead Sep 29, 2026
@cennn
cennn force-pushed the feat/autotune-at-compile-time branch 3 times, most recently from 5a5d6b8 to 3ca29f7 Compare September 29, 2026 03:59
@cennn cennn changed the title [Perf] Persist Triton autotune results to eliminate cold start overhead [WIP] [Perf] Persist Triton autotune cache to skip cold-start benchmarking Sep 29, 2026
@cennn cennn added the ci:run Trigger CI integration tests label Sep 29, 2026
@github-actions github-actions Bot removed the ci:run Trigger CI integration tests label Sep 29, 2026
Root cause: autotune_at_compile_time=False defers Triton kernel autotuning
to first forward pass. Results go to Triton's in-memory cache only —
Triton's disk cache for autotune results (check_disk_cache /
.autotune.json) is disabled by default (knobs.autotuning.cache=False).
Each new pod re-benchmarks all kernel configs: ~3-5 min overhead.

Fix: Set TRITON_CACHE_AUTOTUNING=1 in _compilation_context(). This enables
Triton's built-in autotune disk cache (Autotuner.check_disk_cache), which
writes .autotune.json files to TRITON_CACHE_DIR. Since MagiCompiler already
points TRITON_CACHE_DIR to persistent AFS, bake warmup autotune results are
automatically reused by verify pods.

Flow:
  Bake: compile → warmup forward → autotune → .autotune.json saved to AFS
  Verify: load .so → first forward → read .autotune.json → skip benchmark

Verified with isolated 2-process experiment: 2.6x speedup on cache hit
(1.97s → 0.76s for a simple kernel; real 400B model speedup expected to
be much larger due to many more kernels).

Also updated piecewise_compiler.py comment to document the full mechanism.
@cennn
cennn force-pushed the feat/autotune-at-compile-time branch from 3ca29f7 to c9b3e76 Compare September 29, 2026 04:12
@cennn cennn added the ci:run Trigger CI integration tests label Sep 29, 2026
@github-actions github-actions Bot removed the ci:run Trigger CI integration tests label Sep 29, 2026
@cennn cennn changed the title [WIP] [Perf] Persist Triton autotune cache to skip cold-start benchmarking [Perf] Persist Triton autotune cache to skip cold-start benchmarking Sep 29, 2026

@jiahy0825 jiahy0825 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@jiahy0825
jiahy0825 merged commit 7ee2183 into main Sep 29, 2026
21 checks passed
@jiahy0825
jiahy0825 deleted the feat/autotune-at-compile-time branch September 29, 2026 13:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants