Conversation
GET_ROWS f32 split work by rows only. The recurrent state load of delta-net layers (build_rs) is one large row per sequence (about 3 MB for Qwen3.5/3.8-27B), so a single thread copied it while the others waited. Split such rows into column blocks of at least 1024 floats. CONCAT f32 split work over ne2 only, which is 1 for the conv_input concat of a single sequence. Split over flattened (i1, i2, i3) rows. Both ops only copy data, so results are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 23, 2026
Collaborator
|
Tested on an Intel laptop (AVX2-only x86). Setup: Intel Core Ultra X7 358H (Panther Lake, 16 cores, hybrid; AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 UCRT64 GCC 16.2,
(2B pp64 is noisy in both arms.) My guess is that the matmul dominates so heavily at 16 threads here that the saved single-threaded copy time doesn't show; I didn't profile per op. Combined with #245 + #249 + #250, the stack is 5–8% faster than #250 alone at 4 and 8 threads (see the #249 comment), but I didn't separate out this PR's share. Tested with Claude Code. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
CPU decode of Bonsai 2 27B (qwen35) spent ~11% of each token in single-threaded copies:
GET_ROWSf32 splits work by rows only.build_rsloads the recurrent state of each delta-net layer as one row per sequence (~3 MB here), so with one sequence a single thread copies it while the others wait at the barrier. 48 of these per token.CONCATf32 splits work overne2only. For theconv_inputconcat of a single sequencene2 == 1, so again one thread does it, element by element.Change:
GET_ROWSf32: when there are fewer rows than threads, split each row into column blocks of at least 1024 floats. With enough rows the split factor is 1 and the loop is the same as before.CONCATf32: split over flattened(i1, i2, i3)rows instead ofi2.Both ops only copy data, so results are bit-identical.
Additional information
2x Xeon Gold 6262 (Cascade Lake, 24 cores/socket), Linux, GCC 14,
-DGGML_NATIVE=ON, CPU only,Ternary-Bonsai-2-27B-PQ2_0.gguf. Base isprism+ #245 (needed to load PQ2_0 on this CPU); runs alternated A/B/A/B in the same session.Per-op decode profile (thread 0 wall time per node incl. barrier, 24 threads, ms per token):
llama-bench -p 64 -n 64 -ngl 0:numactl --interleave=all+--numa distributeCorrectness:
llama-perplexity -c 512 --chunks 8on an Italian novel: PPL 23.5011 before and after, identical for every chunk.test-backend-ops(it needs a second backend and this machine only has the CPU).Requirements
🤖 Generated with Claude Code