Skip to content

Qwen35 tensor split (created with deepseek V4.1 flash) - #268

Closed
Yvi71 wants to merge 7 commits into
PrismML-Eng:prismfrom
Yvi71:qwen35-tensor-split
Closed

Yvi71 wants to merge 7 commits into
PrismML-Eng:prismfrom
Yvi71:qwen35-tensor-split

Conversation

@Yvi71

@Yvi71 Yvi71 commented Sep 25, 2026

Copy link
Copy Markdown

Tensor split (row split) for qwen35 hybrid models

This enables multi-GPU tensor split (--split-mode tensor) for the qwen35 hybrid architecture (GDN/SSM + attention), e.g. Ternary-Bonsai-2-27B. Previously, splitting these models produced garbage output; now the split result matches the single-GPU baseline.

Builds on the tensor-split machinery from #214.

The bug

The prism.hadamard rotation (block_size 1024) was the breaking point: any split that touches a hadamard-transformed activation was computed incorrectly, because the meta backend cannot gather across the split.

Root cause: the hadamard perm_rep regroup for ssm_out turned the axis-0 split 6144 activation into a contiguous device-local half {3072x1}, while ssm_out.weight was materialised via get_split_segments={key_dim=2048, head_ratio=3} as {1024x3} -> layout mismatch -> PPL 4.32.

The fix

  • pattern_ssm_out_weight now uses a contiguous segment {{tensor->ne[axis], 1}} instead of {key_dim, head_ratio}.
  • attn_gate keeps {key_dim, head_ratio} (the GDN output {1024x3} must match).
  • SSM/attention MIRROR overrides disabled; signs.6144 split.
  • FFN weights split with a matching hadamard-signs split.

Results

  • PPL 1.0316 vs single-GPU baseline 1.0321 (effectively clean).
  • llama-bench: pp512 405 vs 242 single (+68%); tg128 30.6 vs 32.6 (batch-1 decode slightly worse).
  • batched-bench: prefill +68%, decode +30% at npl>=4.

Testing

2x RX 6800 (gfx1030), ROCm/HIP, tensor split, flash-attention.

This was made WiTH AI.

@github-actions github-actions Bot added the ggml label Sep 25, 2026
@Yvi71

Yvi71 commented Sep 25, 2026

Copy link
Copy Markdown
Author

works on my machine 💯

@Yvi71 Yvi71 closed this Oct 1, 2026
@Yvi71
Yvi71 deleted the qwen35-tensor-split branch October 1, 2026 12:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant