Skip to content

feat(intrinsics): admit standard-metadata m16n8k16 bf16 sparse MMA - #1355

Open
YeonwooSung wants to merge 2 commits into
NVIDIA:mainfrom
YeonwooSung:feat/sparse-mma-bf16-m16n8k16
Open

YeonwooSung wants to merge 2 commits into
NVIDIA:mainfrom
YeonwooSung:feat/sparse-mma-bf16-m16n8k16

Conversation

@YeonwooSung

Copy link
Copy Markdown

Relates to #1296.

Admit one standard-metadata Ampere sparse MMA:

int_nvvm_mma_sp_m16n8k16_row_col_bf16 → mma.sp.sync.aligned.m16n8k16.row.col.f32.bf16.bf16.f32

The ordered sibling stays mma_sp_ordered_metadata_m16n8k16_f32_bf16 (sp::ordered_metadata, PTX 8.5). The ordered e4m3 golden identity is unchanged.

runtime_validation is unexecuted. An llc candidate probe at sm_80/+ptx71 retained the instruction for selectors 0–3. This host has no GPU, ptxas, or libNVVM, so there is no runtime oracle and no device-link stage. The libNVVM evidence record stays lowered and says the toolkit was not invoked here.

cargo test -p cuda-intrinsics-gen passed, including arch_introduction and the new admission test. cuda-intrinsics-gen generate updated the catalog. The catalog SHA line in other generated files and probes is the mechanical stamp.

Admit int_nvvm_mma_sp_m16n8k16_row_col_bf16 as plain
mma.sp.sync.aligned.m16n8k16.row.col.f32.bf16.bf16.f32. The ordered
sibling and the ordered e4m3 identity stay unchanged. Runtime
validation is unexecuted. Catalog SHA stamps are the generator output.

Signed-off-by: YeonwooSung <neos960518@gmail.com>
Keep the one plain m16n8k16 bf16 sparse MMA
(mma.sp.sync.aligned.m16n8k16.row.col.f32.bf16.bf16.f32) under
cuda-oxide/ after the SIMT tree move. ABI i1030 is appended.
Regenerated catalog outputs match the generator SHA. Runtime
validation stays unexecuted. The ordered e4m3 kind::f8f6f4 identity
and the SM89 FP8 admission are unchanged.

Signed-off-by: YeonwooSung <neos960518@gmail.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant