cuda: branch-free PTQ1_0 MMQ tile loader and full Ampere tile table (2x prefill) - #214
Conversation
…2x prefill) Prefill takes the MMQ path and PTQ1_0 ran it at half PQ2_0's speed from the same weights (RTX 4070: pp2048 630 vs ~1300 t/s). CUPTI put the whole gap in mul_mat_q<PTQ1_0>, 2.7x the per-call time of the PQ2_0 instantiation for identical shapes. - mmq-config-ampere.cuh: PTQ1_0 was the only type capped at mmq_x = 64; add the 80/96/112/128 entries every other type has. 630 -> 914 t/s. - mmq-load-tiles.cuh: load_tiles_ptq1_0 split a block's 8 lanes into three divergent branches (lanes 0-3: 5 trit-unpack iterations, lanes 4-5: 5 more, lane 6: 2, lane 7 idle), so each warp serialized 12 iterations for 5 of work. All lanes now run one identical 5-iteration loop on their own 32-bit word of the block; lane 6 walks the two qh bytes in both 16-bit halves and recombines adjacent digits with a single __byte_perm. Only smem store offsets differ per lane. 914 -> 1304 t/s. Decode untouched (66.9 t/s). test-backend-ops MUL_MAT 78/78, MUL_MAT_ID 75/75 PTQ1_0 shapes vs CPU; greedy 300-token generation identical; perplexity -c 2048 -b 2048 7.6740 vs 7.6742 stock.
bri-prism
left a comment
There was a problem hiding this comment.
Agent review: posted by the maintainer's coding agent at their request.
No findings in this source pass. I emulated the new tile unpack on 10,000 arbitrary packed blocks and compared all 128 decoded values per block against the scalar codec: zero mismatches, including the qh tail.
That validates the unpack arithmetic only. CUDA compilation, tile dispatch/shared-memory behavior and the performance claims were not rerun. This looks suitable to keep independent of the competing decode changes in #215/#218.
Reviewed commit: b16ac95b7eb6714899ebacda96493ae9a1fd960f.
|
Thanks for the emulation pass. Nothing changed here; One more data point from today's paired runs on the RTX 4070 at stock clocks: Prism release binary pp512 600 t/s, this branch 1256 t/s, same file and flags, |
| // Branch-free unpack. All 8 lanes of a block run the same 5-iteration trit-extraction loop | ||
| // on their own 32-bit word: lanes 0-5 take qs words 0-5, lane 6 takes word 6 | ||
| // (qh[0] | qh[1] << 8 | d << 16) with both 16-bit halves walking the two qh bytes, lane 7 | ||
| // computes and discards. Only the shared-memory store offsets differ per lane, so the warp | ||
| // never diverges. The previous lane<4 / lane<6 / lane==6 branch chain serialized 5+5+2 | ||
| // iterations per warp and made this loader ~2.7x slower than the PQ2_0 one for the same tile. |
There was a problem hiding this comment.
Agreed. Trimmed to the lane-mapping invariant in 1b1fb88. Measurements stay in the PR body.
Keep the lane-mapping invariant only; the measurements stay in the PR body.
|
Agent benchmark follow-up, posted at the maintainer's request. Pinned head
Selected CPU-reference backend checks passed on the compared arms. The current head The percentages describe these paired runs; small changes should not be interpreted as established improvements. No long-context, multi-slot serving, or end-to-end logit-parity claim is made. |
|
The campaign matches the 4070 and 3060 receipts. Same public PTQ1 file, Your four cards, plus the two consumer parts already on the thread:
Agreed that Ampere through Blackwell all showing ~2x prefill is enough for this PR on its own. It does not depend on the #215 / #218 layout question. |
…MMQ tile path With the branch-free PTQ1_0 MMQ tile loader (PrismML-Eng#214) in the tree, the MMQ path is the faster one from 5 columns on. On an RTX 3060 with the K = 5120 projections of Bonsai 2 27B, llama-bench -p 8 runs a batch of 8 in 64.4 ms through the new MMQ tiles against 112.7 ms through the 8-column mat-vec (five-arm a/b, r=3, github.com/sudoingX/bonsai2-small-gpu kernel/ab/ab_215_214_rtx3060.md). The mat-vec keeps its lead at 2 to 4 columns (1.20x / 1.49x / 1.77x of a single pass against 1.46x / 2.21x / 2.75x for the tile path), which is the range speculative verification uses. ggml_cuda_should_use_mmvq now routes PTQ1_0 to the mat-vec for ne11 <= 4 (PTQ1_0_PT_MAX_COLS, was 8) and the 5 to 8 column instantiations of the dedicated kernel are gone. GGML_CUDA_BATCH_INVARIANT keeps its 1 to 4 column guarantee; the comment no longer mentions 5 to 8, which the flag never covered once those batches take the tile path.
…MMQ tile path With the branch-free PTQ1_0 MMQ tile loader (PrismML-Eng#214) in the tree, the MMQ path is the faster one from 5 columns on. On an RTX 3060 with the K = 5120 projections of Bonsai 2 27B, llama-bench -p 8 runs a batch of 8 in 64.4 ms through the new MMQ tiles against 112.7 ms through the 8-column mat-vec (five-arm a/b, r=3, github.com/sudoingX/bonsai2-small-gpu kernel/ab/ab_215_214_rtx3060.md). The mat-vec keeps its lead at 2 to 4 columns (1.20x / 1.49x / 1.77x of a single pass against 1.46x / 2.21x / 2.75x for the tile path), which is the range speculative verification uses. ggml_cuda_should_use_mmvq now routes PTQ1_0 to the mat-vec for ne11 <= 4 (PTQ1_0_PT_MAX_COLS, was 8) and the 5 to 8 column instantiations of the dedicated kernel are gone. GGML_CUDA_BATCH_INVARIANT keeps its 1 to 4 column guarantee; the comment no longer mentions 5 to 8, which the flag never covered once those batches take the tile path.

Summary
PTQ1_0 prefill on consumer NVIDIA (measured on an RTX 4070, sm_89) ran at half the speed of PQ2_0 built from the same trits: llama-bench pp2048 630 vs ~1300 t/s. A CUPTI trace of a pp2048 run put the whole gap in
mul_mat_q<GGML_TYPE_PTQ1_0>: 2.7x the per-call time of the PQ2_0 instantiation for identical shapes. Two independent causes, both in the PTQ1_0 MMQ path:mmq-config-ampere.cuh: PTQ1_0 was the only type capped atmmq_x = 64. PQ2_0 and the rest go to 80/96/112/128. A 2048-token prompt therefore made twice the number of K-passes over the weight tiles. This PR adds the four missingCASElines (same layout/config as the existing PTQ1_0 entries). 630 -> 914 t/s.mmq-load-tiles.cuh:ggml_cuda_mmq_load_tiles_ptq1_0was warp-divergent. The 8 lanes of a block were split intoif (lane < 4) / else if (lane < 6) / else if (lane == 6), running 5, 5 and 2 trit-unpack iterations respectively with lane 7 idle. Divergent branches execute serially within a warp, so every warp paid 12 iterations for 5 of work; the PQ2_0 loader is uniform. The loader now runs one identical 5-iteration loop on every lane's own 32-bit word of the 28-byte block (lanes 0-5: qs words 0-5; lane 6: word 6 =qh[0] | qh[1] << 8 | d << 16, walking the two qh bytes in both 16-bit halves and recombining adjacent digits with one__byte_perm; lane 7 computes and discards). Only the shared-memory store offsets differ per lane. 914 -> 1304 t/s. The now-unusedggml_cuda_mmq_decode_ptq1_0_qs4helper is removed.PTQ1_0 now prefills at PQ2_0 speed while keeping its smaller footprint and faster decode on Ada, so the two packings no longer trade off against each other on this class of card.
Measurements
RTX 4070 12 GB, Ternary-Bonsai-2-27B-PTQ1_0.gguf,
-fa on, CUDA 13.3 build ofprism@ 9a9394a. Same memory clock throughout; prefill is compute-bound.prism@ 9a9394aLive
llama-server(262k q4_0 KV window, one slot): 2048-token prefill at depth 0 617 -> 1275 t/s, at depth 35k 498 -> 847 t/s; TTFT on a 1611-token prompt 2.75 s -> 1.39 s. Decode unchanged (the MMVQ path is untouched).Correctness
test-backend-ops -b CUDA0 -o MUL_MAT: 1283/1283, all 45 supported PTQ1_0 shapes vs CPU (the remaining PTQ1_0 lines are f16-src1 "not supported" on both backends, as before).test-backend-ops -b CUDA0 -o MUL_MAT_ID: all PTQ1_0 shapes pass.llama-perplexity -c 2048 -b 2048on a 4-chunk text: 7.6740 vs 7.6742 before (float noise), 1.80 vs 3.59 s per pass.Scope
CUDA only; the PTQ1_0 loader is already inside
#if !defined(GGML_USE_HIP), and the config file is Ampere-and-up. The wideCASEentries mirror what PQ2_0 already ships, so tile shared-memory budgets are unchanged. Blackwell/Hopper untested but the change is arch-independent (no new intrinsics).Related: #200 does the equivalent loader cleanup for PQ2_0 on HIP.