Skip to content

cuda: branch-free PTQ1_0 MMQ tile loader and full Ampere tile table (2x prefill) - #214

Merged
bri-prism merged 2 commits into
PrismML-Eng:prismfrom
professorpalmer:cuda-ptq1-mmq-prefill
Sep 22, 2026
Merged

bri-prism merged 2 commits into
PrismML-Eng:prismfrom
professorpalmer:cuda-ptq1-mmq-prefill

Conversation

@professorpalmer

Copy link
Copy Markdown

Summary

PTQ1_0 prefill on consumer NVIDIA (measured on an RTX 4070, sm_89) ran at half the speed of PQ2_0 built from the same trits: llama-bench pp2048 630 vs ~1300 t/s. A CUPTI trace of a pp2048 run put the whole gap in mul_mat_q<GGML_TYPE_PTQ1_0>: 2.7x the per-call time of the PQ2_0 instantiation for identical shapes. Two independent causes, both in the PTQ1_0 MMQ path:

  1. mmq-config-ampere.cuh: PTQ1_0 was the only type capped at mmq_x = 64. PQ2_0 and the rest go to 80/96/112/128. A 2048-token prompt therefore made twice the number of K-passes over the weight tiles. This PR adds the four missing CASE lines (same layout/config as the existing PTQ1_0 entries). 630 -> 914 t/s.

  2. mmq-load-tiles.cuh: ggml_cuda_mmq_load_tiles_ptq1_0 was warp-divergent. The 8 lanes of a block were split into if (lane < 4) / else if (lane < 6) / else if (lane == 6), running 5, 5 and 2 trit-unpack iterations respectively with lane 7 idle. Divergent branches execute serially within a warp, so every warp paid 12 iterations for 5 of work; the PQ2_0 loader is uniform. The loader now runs one identical 5-iteration loop on every lane's own 32-bit word of the 28-byte block (lanes 0-5: qs words 0-5; lane 6: word 6 = qh[0] | qh[1] << 8 | d << 16, walking the two qh bytes in both 16-bit halves and recombining adjacent digits with one __byte_perm; lane 7 computes and discards). Only the shared-memory store offsets differ per lane. 914 -> 1304 t/s. The now-unused ggml_cuda_mmq_decode_ptq1_0_qs4 helper is removed.

PTQ1_0 now prefills at PQ2_0 speed while keeping its smaller footprint and faster decode on Ada, so the two packings no longer trade off against each other on this class of card.

Measurements

RTX 4070 12 GB, Ternary-Bonsai-2-27B-PTQ1_0.gguf, -fa on, CUDA 13.3 build of prism @ 9a9394a. Same memory clock throughout; prefill is compute-bound.

llama-bench pp512 pp2048 tg128
prism @ 9a9394a 630 630 66.7
+ wide MMQ tiles (1) 915 913 66.7
+ branch-free loader (2) 1304 1301 66.9

Live llama-server (262k q4_0 KV window, one slot): 2048-token prefill at depth 0 617 -> 1275 t/s, at depth 35k 498 -> 847 t/s; TTFT on a 1611-token prompt 2.75 s -> 1.39 s. Decode unchanged (the MMVQ path is untouched).

Correctness

  • test-backend-ops -b CUDA0 -o MUL_MAT: 1283/1283, all 45 supported PTQ1_0 shapes vs CPU (the remaining PTQ1_0 lines are f16-src1 "not supported" on both backends, as before).
  • test-backend-ops -b CUDA0 -o MUL_MAT_ID: all PTQ1_0 shapes pass.
  • 300-token greedy generation byte-identical to the unpatched build.
  • llama-perplexity -c 2048 -b 2048 on a 4-chunk text: 7.6740 vs 7.6742 before (float noise), 1.80 vs 3.59 s per pass.

Scope

CUDA only; the PTQ1_0 loader is already inside #if !defined(GGML_USE_HIP), and the config file is Ampere-and-up. The wide CASE entries mirror what PQ2_0 already ships, so tile shared-memory budgets are unchanged. Blackwell/Hopper untested but the change is arch-independent (no new intrinsics).

Related: #200 does the equivalent loader cleanup for PQ2_0 on HIP.

…2x prefill)

Prefill takes the MMQ path and PTQ1_0 ran it at half PQ2_0's speed from the same weights
(RTX 4070: pp2048 630 vs ~1300 t/s). CUPTI put the whole gap in mul_mat_q<PTQ1_0>, 2.7x the
per-call time of the PQ2_0 instantiation for identical shapes.

- mmq-config-ampere.cuh: PTQ1_0 was the only type capped at mmq_x = 64; add the 80/96/112/128
  entries every other type has. 630 -> 914 t/s.
- mmq-load-tiles.cuh: load_tiles_ptq1_0 split a block's 8 lanes into three divergent branches
  (lanes 0-3: 5 trit-unpack iterations, lanes 4-5: 5 more, lane 6: 2, lane 7 idle), so each warp
  serialized 12 iterations for 5 of work. All lanes now run one identical 5-iteration loop on
  their own 32-bit word of the block; lane 6 walks the two qh bytes in both 16-bit halves and
  recombines adjacent digits with a single __byte_perm. Only smem store offsets differ per lane.
  914 -> 1304 t/s.

Decode untouched (66.9 t/s). test-backend-ops MUL_MAT 78/78, MUL_MAT_ID 75/75 PTQ1_0 shapes vs
CPU; greedy 300-token generation identical; perplexity -c 2048 -b 2048 7.6740 vs 7.6742 stock.

@bri-prism bri-prism left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agent review: posted by the maintainer's coding agent at their request.

No findings in this source pass. I emulated the new tile unpack on 10,000 arbitrary packed blocks and compared all 128 decoded values per block against the scalar codec: zero mismatches, including the qh tail.

That validates the unpack arithmetic only. CUDA compilation, tile dispatch/shared-memory behavior and the performance claims were not rerun. This looks suitable to keep independent of the competing decode changes in #215/#218.

Reviewed commit: b16ac95b7eb6714899ebacda96493ae9a1fd960f.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@professorpalmer

Copy link
Copy Markdown
Author

Thanks for the emulation pass. Nothing changed here; b16ac95 stands and stays independent of the #215/#218 layout question (this PR is MMQ only, batch >= 9, no activation layout involved).

One more data point from today's paired runs on the RTX 4070 at stock clocks: Prism release binary pp512 600 t/s, this branch 1256 t/s, same file and flags, test-backend-ops MUL_MAT ptq1_0 45/45. @sudoingX's 3060 sits at 268 pp512 on the release binary, so an Ampere number for this branch would be a useful second card.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The architecture-sensitive CUDA kernel change lacks runtime validation on several supported GPU generations.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

Comment thread ggml/src/ggml-cuda/mmq-load-tiles.cuh Outdated
Comment on lines +303 to +308
// Branch-free unpack. All 8 lanes of a block run the same 5-iteration trit-extraction loop
// on their own 32-bit word: lanes 0-5 take qs words 0-5, lane 6 takes word 6
// (qh[0] | qh[1] << 8 | d << 16) with both 16-bit halves walking the two qh bytes, lane 7
// computes and discards. Only the shared-memory store offsets differ per lane, so the warp
// never diverges. The previous lane<4 / lane<6 / lane==6 branch chain serialized 5+5+2
// iterations per warp and made this loader ~2.7x slower than the PQ2_0 one for the same tile.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Trimmed to the lane-mapping invariant in 1b1fb88. Measurements stay in the PR body.

Keep the lane-mapping invariant only; the measurements stay in the PR body.
@bri-prism

Copy link
Copy Markdown
Collaborator

Agent benchmark follow-up, posted at the maintainer's request.

Pinned head b16ac95 against merge-base 9a9394a. Same public PTQ1 model on both arms, full GPU offload, flash attention, q4_0 K/V, batch/microbatch 512, 8 CPU threads. Three alternating baseline/candidate pairs, three repetitions per invocation, short/empty starting context.

GPU pp512 before → after, tok/s Paired change tg128 before → after, tok/s Paired change
RTX 3090 756.15 → 1364.68 +80.5% 60.01 → 60.13 +0.2%
RTX 4090 1563.84 → 3038.67 +94.3% 88.12 → 88.11 -0.0%
H100 SXM 1229.29 → 2760.28 +124.5% 88.41 → 88.30 -0.1%
RTX 5090 1881.37 → 3961.57 +110.6% 119.77 → 119.69 -0.1%

Selected CPU-reference backend checks passed on the compared arms.

The current head 1b1fb88 differs from tested b16ac95 only in comments in the tile loader; executable code is unchanged in the reviewed delta. These are measurements of the earlier head, not a fresh rerun. The observed gain is concentrated in prefill. End-to-end logit parity was not measured.

The percentages describe these paired runs; small changes should not be interpreted as established improvements. No long-context, multi-slot serving, or end-to-end logit-parity claim is made.

@professorpalmer

Copy link
Copy Markdown
Author

The campaign matches the 4070 and 3060 receipts. Same public PTQ1 file, -fa 1 -ctk q4_0 -ctv q4_0, batch/ubatch 512.

Your four cards, plus the two consumer parts already on the thread:

GPU pp512 before → after paired tg128 paired source
RTX 3090 756 → 1365 +80.5% +0.2% this campaign, b16ac95
RTX 4090 1564 → 3039 +94.3% -0.0% this campaign, b16ac95
H100 SXM 1229 → 2760 +124.5% -0.1% this campaign, b16ac95
RTX 5090 1881 → 3962 +110.6% -0.1% this campaign, b16ac95
RTX 4070 12 GB 600 → 1256 +109% +0.2% stock clocks, three alternating rounds
RTX 3060 12 GB 269 → 515 +92% -1.0% @sudoingX, b16ac95

Agreed that 1b1fb88 is comments-only against tested b16ac95; executable code is unchanged. Decode staying in the noise on every NVIDIA generation is the intended split: this PR is MMQ only (batch >= 9). test-backend-ops -o MUL_MAT PTQ1_0 45/45 on the 4070 build. End-to-end logit parity was not our claim either.

Ampere through Blackwell all showing ~2x prefill is enough for this PR on its own. It does not depend on the #215 / #218 layout question.

@bri-prism
bri-prism merged commit bdc23b5 into PrismML-Eng:prism Sep 22, 2026
1 check passed
sudoingX added a commit to sudoingX/llama.cpp that referenced this pull request Sep 22, 2026
…MMQ tile path

With the branch-free PTQ1_0 MMQ tile loader (PrismML-Eng#214) in the tree, the MMQ path
is the faster one from 5 columns on. On an RTX 3060 with the K = 5120
projections of Bonsai 2 27B, llama-bench -p 8 runs a batch of 8 in 64.4 ms
through the new MMQ tiles against 112.7 ms through the 8-column mat-vec
(five-arm a/b, r=3, github.com/sudoingX/bonsai2-small-gpu
kernel/ab/ab_215_214_rtx3060.md). The mat-vec keeps its lead at 2 to 4
columns (1.20x / 1.49x / 1.77x of a single pass against 1.46x / 2.21x /
2.75x for the tile path), which is the range speculative verification uses.

ggml_cuda_should_use_mmvq now routes PTQ1_0 to the mat-vec for ne11 <= 4
(PTQ1_0_PT_MAX_COLS, was 8) and the 5 to 8 column instantiations of the
dedicated kernel are gone. GGML_CUDA_BATCH_INVARIANT keeps its 1 to 4 column
guarantee; the comment no longer mentions 5 to 8, which the flag never
covered once those batches take the tile path.
sudoingX added a commit to sudoingX/llama.cpp that referenced this pull request Sep 27, 2026
…MMQ tile path

With the branch-free PTQ1_0 MMQ tile loader (PrismML-Eng#214) in the tree, the MMQ path
is the faster one from 5 columns on. On an RTX 3060 with the K = 5120
projections of Bonsai 2 27B, llama-bench -p 8 runs a batch of 8 in 64.4 ms
through the new MMQ tiles against 112.7 ms through the 8-column mat-vec
(five-arm a/b, r=3, github.com/sudoingX/bonsai2-small-gpu
kernel/ab/ab_215_214_rtx3060.md). The mat-vec keeps its lead at 2 to 4
columns (1.20x / 1.49x / 1.77x of a single pass against 1.46x / 2.21x /
2.75x for the tile path), which is the range speculative verification uses.

ggml_cuda_should_use_mmvq now routes PTQ1_0 to the mat-vec for ne11 <= 4
(PTQ1_0_PT_MAX_COLS, was 8) and the 5 to 8 column instantiations of the
dedicated kernel are gone. GGML_CUDA_BATCH_INVARIANT keeps its 1 to 4 column
guarantee; the comment no longer mentions 5 to 8, which the flag never
covered once those batches take the tile path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants