Skip to content

Fix: wait on the PDL dependency in the PTQ1_0 mat-vec kernel (garbage output on Hopper/Blackwell) - #1

Open
LamplighterPaul wants to merge 1 commit into
sudoingX:pr-ptq1-mmvfrom
LamplighterPaul:fix/ptq1_0-mmv-pdl-sync
Open

LamplighterPaul wants to merge 1 commit into
sudoingX:pr-ptq1-mmvfrom
LamplighterPaul:fix/ptq1_0-mmv-pdl-sync

Conversation

@LamplighterPaul

Copy link
Copy Markdown

This is the fix from sudoingX/bonsai2-small-gpu#3, opened here so it can go straight into pr-ptq1-mmv and from there into PrismML-Eng#218.

Bug: on Hopper and newer, mul_mat_vec_ptq1_0_pt is launched with PDL through ggml_cuda_kernel_launch, but the kernel never calls ggml_cuda_pdl_sync(). It can therefore read vy before the q8_1 quantize kernel has finished writing it. On an RTX 5080 (sm_120), llama-server output breaks into !!!! or //// a few tokens in, for text and vision alike.

Why it went unseen: llama-bench runs at full speed with the bug (97 tok/s), and test-backend-ops -o MUL_MAT passes every PTQ1_0 case because it runs each op on its own. Ampere and Ada never take the PDL path, so the 3060, 3060 Ti and 4070 rows were never affected.

Isolation on the 5080 (greedy, thinking off):

build output
PrismML b10685 / b10735 correct
285542d garbage
285542d, GGML_CUDA_DISABLE_GRAPHS=1 correct
285542d, GGML_CUDA_PDL=0 correct
285542d + this commit correct, identical to b10735

Cost: none. tg128 is 96.92 with the fix against 97.42 without, within noise. Full llama-bench ladder: sweeps/rtx5080-16gb.md in sudoingX/bonsai2-small-gpu#3.

The change is one call at the top of the kernel, the same one the mul_mat_vec_q kernels in mmvq.cu make.

…eading vy

mul_mat_vec_ptq1_0_pt is launched through ggml_cuda_kernel_launch, which
opts into programmatic dependent launch on Hopper and newer. The kernel
never called ggml_cuda_pdl_sync(), so it could start reading vy before
the q8_1 activation quantization had finished writing it. On an RTX 5080
(sm_120) llama-server output broke into repeated '!' or '/' a few tokens
in, text and vision alike. GGML_CUDA_PDL=0 or GGML_CUDA_DISABLE_GRAPHS=1
hid it. llama-bench speed is unaffected; test-backend-ops runs each op
on its own and cannot catch it. The mul_mat_vec_q kernels in mmvq.cu
make the same call. Ampere and Ada never take the PDL path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant