ci: guard H100-calibrated test assertions for B300 (SM103) runners - #83
Merged
Merged
Conversation
Verify that the 8 new magi-compiler CI runners on B300 (SM103) can build and run all test shards successfully.
assert_speedup already skips on non-H100 hardware, but assert_magi_vs_torch was missing the same guard. The conv channels-last thresholds are H100-specific; on B300 (SM103) the pass benefit is narrower and trips the assertion.
The max/min < 1.2 check compares sub-millisecond CUDA event medians. On B300 (SM103) the fastest entry point hits ~50μs while others stay at ~200μs, yielding 3-4x ratios that are noise at this timescale. Reuse the existing is_perf_calibrated_gpu() guard.
…d GPUs profile_sync does lockstep JIT measurement of every graph node; on a cold B300 (no Triton cache) this can exceed 900s. Raise to 1800s for profile_sync only; other cost modes keep 900s. The api test entry-point timing consistency check (max/min < 1.2) hits CUDA event noise at sub-millisecond medians on B300. Guard it with is_perf_calibrated_gpu() like the perf shard assertions.
…alibrated GPUs" This reverts commit 98468b0.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
MagiCompiler CI now runs on B300 (SM103) self-hosted runners. Perf thresholds and timing assertions are calibrated for H100 — on B300 the operator mix shifts and sub-millisecond CUDA event medians hit noise floors, causing spurious failures in
perfandapishards.assert_speedupalready had anis_perf_calibrated_gpu()guard, butassert_magi_vs_torchand the API entry-point timing consistency checks did not.Changes
Guard
assert_magi_vs_torchwithis_perf_calibrated_gpu()(tests/perf_tests/utils.py) — consistent with the existing guard onassert_speedup. On B300, magi-vs-torch ratio drops below the H100 threshold (e.g.0.98xvs1.05x) due to operator mix differences; silently passes on non-calibrated GPUs.Guard timing consistency assertions (
tests/api_tests/test_magi_compile.py) — themax/min < 1.2checks across compile entry points fail on B300 where CUDA event medians are ~50μs (vs ~200μs on H100), causing noise-dominated ratios of 3-4x.Result
All 16 test jobs (8 shards × 2 PyTorch versions) pass on B300 SM103 runners.