Use a more optimal formulation of swizzle_dyn for AVX2 - #548
Merged
Conversation
Shnatsel
added a commit
to Shnatsel/fearless_simd
that referenced
this pull request
Sep 3, 2026
@dzaima got nerd-sniped by [my blog post](https://shnatsel.github.io/improving-std-simd-swizzle-dyn/) about `swizzle_dyn_precise` and workshopped even faster formulations. I couldn't let that stand and tried optimizing these this myself. All the formulations we tried, their godbolt links and their llvm-mca timings can be found at https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56 The one proposed in this PR is the last line of the table, "Blend using control-derived mask". On my Zen2 laptop this improves performance dramatically: 33% less time taken which means 50% higher throughput. Corresponding `std::simd` PR: rust-lang/portable-simd#548 ---- This PR also includes an optimization to swizzle_dyn, the non-precise variant. Details on its performance are in the commit message.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@dzaima got nerd-sniped by my blog post about
swizzle_dynand workshopped even faster formulations. I couldn't let that stand and tried optimizing these this myself.All the formulations we tried, their godbolt links and their llvm-mca timings can be found at https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56
The one proposed in this PR is the last line of the table, "Blend using control-derived mask".
On my Zen2 laptop this improves performance dramatically: 33% less time taken which means 50% higher throughput.