Skip to content

ggml-cpu : split GET_ROWS and CONCAT f32 work when there are few rows - #246

Open
lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-rows-parallel-copies
Open

lenny76 wants to merge 1 commit into
PrismML-Eng:prismfrom
lenny76:cpu-rows-parallel-copies

Conversation

@lenny76

@lenny76 lenny76 commented Sep 23, 2026

Copy link
Copy Markdown

Overview

CPU decode of Bonsai 2 27B (qwen35) spent ~11% of each token in single-threaded copies:

  • GET_ROWS f32 splits work by rows only. build_rs loads the recurrent state of each delta-net layer as one row per sequence (~3 MB here), so with one sequence a single thread copies it while the others wait at the barrier. 48 of these per token.
  • CONCAT f32 splits work over ne2 only. For the conv_input concat of a single sequence ne2 == 1, so again one thread does it, element by element.

Change:

  • GET_ROWS f32: when there are fewer rows than threads, split each row into column blocks of at least 1024 floats. With enough rows the split factor is 1 and the loop is the same as before.
  • CONCAT f32: split over flattened (i1, i2, i3) rows instead of i2.

Both ops only copy data, so results are bit-identical.

Additional information

2x Xeon Gold 6262 (Cascade Lake, 24 cores/socket), Linux, GCC 14, -DGGML_NATIVE=ON, CPU only, Ternary-Bonsai-2-27B-PQ2_0.gguf. Base is prism + #245 (needed to load PQ2_0 on this CPU); runs alternated A/B/A/B in the same session.

Per-op decode profile (thread 0 wall time per node incl. barrier, 24 threads, ms per token):

op before after
GET_ROWS (all) 20.4 3.7
CONCAT 4.9 3.8
MUL_MAT pq2_0 112.6 114.6

llama-bench -p 64 -n 64 -ngl 0:

config threads tg64 before tg64 after pp64 before pp64 after
numactl --interleave=all + --numa distribute 24 5.28 / 5.31 6.46 / 6.46 23.4 / 23.5 24.3 / 24.3
default 24 2.79 / 2.79 3.01 / 3.04 - -
default 48 2.91 / 2.96 3.24 / 3.20 - -

Correctness:

  • llama-perplexity -c 512 --chunks 8 on an Italian novel: PPL 23.5011 before and after, identical for every chunk.
  • Greedy 96-token completion (PQ2_0) and 48-token completion (PTQ1_0, this PR alone): identical text.
  • I could not run test-backend-ops (it needs a second backend and this machine only has the CPU).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code (Anthropic) profiled the decode graph, wrote the patch and ran the measurements above on my machine, under my direction.

🤖 Generated with Claude Code

GET_ROWS f32 split work by rows only. The recurrent state load of
delta-net layers (build_rs) is one large row per sequence (about 3 MB
for Qwen3.5/3.8-27B), so a single thread copied it while the others
waited. Split such rows into column blocks of at least 1024 floats.

CONCAT f32 split work over ne2 only, which is 1 for the conv_input
concat of a single sequence. Split over flattened (i1, i2, i3) rows.

Both ops only copy data, so results are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bri-prism

Copy link
Copy Markdown
Collaborator

Tested on an Intel laptop (AVX2-only x86).

Setup: Intel Core Ultra X7 358H (Panther Lake, 16 cores, hybrid; AVX2 + AVX-VNNI, no AVX-512), Windows 11, MSYS2 UCRT64 GCC 16.2, -DGGML_NATIVE=ON, CPU only (-ngl 0). Base = prism @ 3b19c377d; PR merged on top. llama-bench -p 64 -n 32 -r 2 -t 16, two rounds with the variant order reversed in round 2 (values are round 1 / round 2, t/s).

  • Correctness: greedy output (latest-2B-PTQ1_0, 128 tokens, temp 0) is byte-identical to base, as expected for a copy-only change.
  • Performance is neutral on this box:
model base pp64 PR pp64 base tg32 PR tg32
Bonsai 2 27B PQ2_0 18.67 / 17.68 18.65 / 18.48 2.75 / 2.71 2.75 / 2.81
2B PTQ1_0 40.26 / 27.21 38.71 / 27.18 8.03 / 7.08 7.95 / 7.05

(2B pp64 is noisy in both arms.) My guess is that the matmul dominates so heavily at 16 threads here that the saved single-threaded copy time doesn't show; I didn't profile per op. Combined with #245 + #249 + #250, the stack is 5–8% faster than #250 alone at 4 and 8 threads (see the #249 comment), but I didn't separate out this PR's share.

Tested with Claude Code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants