Skip to content

Use a more optimal formulation of swizzle_dyn for AVX2 - #548

Merged
calebzulawski merged 1 commit into
rust-lang:masterfrom
Shnatsel:faster-avx2-swizzle
Sep 1, 2026
Merged

Use a more optimal formulation of swizzle_dyn for AVX2#548
calebzulawski merged 1 commit into
rust-lang:masterfrom
Shnatsel:faster-avx2-swizzle

Conversation

@Shnatsel

Copy link
Copy Markdown
Member

@dzaima got nerd-sniped by my blog post about swizzle_dyn and workshopped even faster formulations. I couldn't let that stand and tried optimizing these this myself.

All the formulations we tried, their godbolt links and their llvm-mca timings can be found at https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56

The one proposed in this PR is the last line of the table, "Blend using control-derived mask".

On my Zen2 laptop this improves performance dramatically: 33% less time taken which means 50% higher throughput.

@calebzulawski
calebzulawski merged commit 4cbc75b into rust-lang:master Sep 1, 2026
61 checks passed
Shnatsel added a commit to Shnatsel/fearless_simd that referenced this pull request Sep 3, 2026
@dzaima got nerd-sniped by [my blog
post](https://shnatsel.github.io/improving-std-simd-swizzle-dyn/) about
`swizzle_dyn_precise` and workshopped even faster formulations. I
couldn't let that stand and tried optimizing these this myself.

All the formulations we tried, their godbolt links and their llvm-mca
timings can be found at
https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56

The one proposed in this PR is the last line of the table, "Blend using
control-derived mask".

On my Zen2 laptop this improves performance dramatically: 33% less time
taken which means 50% higher throughput.

Corresponding `std::simd` PR:
rust-lang/portable-simd#548

----

This PR also includes an optimization to swizzle_dyn, the non-precise
variant. Details on its performance are in the commit message.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants