Conversation
…gments, QWEN35 mirror block, signs per-width
…evert whole-dim+mirror-block
…or attention block. PPL 1.0323 (baseline 1.0321)
…-regroup layout mismatch. PPL 1.0316 (baseline 1.0321); pp512 405 vs 242 single (+68%)
…L/DBG_CONS stderr prints
Author
|
works on my machine 💯 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tensor split (row split) for qwen35 hybrid models
This enables multi-GPU tensor split (
--split-mode tensor) for theqwen35hybrid architecture (GDN/SSM + attention), e.g. Ternary-Bonsai-2-27B. Previously, splitting these models produced garbage output; now the split result matches the single-GPU baseline.Builds on the tensor-split machinery from #214.
The bug
The
prism.hadamardrotation (block_size 1024) was the breaking point: any split that touches a hadamard-transformed activation was computed incorrectly, because the meta backend cannot gather across the split.Root cause: the hadamard
perm_repregroup forssm_outturned the axis-0 split 6144 activation into a contiguous device-local half{3072x1}, whilessm_out.weightwas materialised viaget_split_segments={key_dim=2048, head_ratio=3}as{1024x3}-> layout mismatch -> PPL 4.32.The fix
pattern_ssm_out_weightnow uses a contiguous segment{{tensor->ne[axis], 1}}instead of{key_dim, head_ratio}.attn_gatekeeps{key_dim, head_ratio}(the GDN output{1024x3}must match).signs.6144split.Results
llama-bench: pp512 405 vs 242 single (+68%); tg128 30.6 vs 32.6 (batch-1 decode slightly worse).batched-bench: prefill +68%, decode +30% at npl>=4.Testing
2x RX 6800 (gfx1030), ROCm/HIP, tensor split, flash-attention.
This was made WiTH AI.