Fix: wait on the PDL dependency in the PTQ1_0 mat-vec kernel (garbage output on Hopper/Blackwell) - #1
Open
LamplighterPaul wants to merge 1 commit into
Conversation
…eading vy mul_mat_vec_ptq1_0_pt is launched through ggml_cuda_kernel_launch, which opts into programmatic dependent launch on Hopper and newer. The kernel never called ggml_cuda_pdl_sync(), so it could start reading vy before the q8_1 activation quantization had finished writing it. On an RTX 5080 (sm_120) llama-server output broke into repeated '!' or '/' a few tokens in, text and vision alike. GGML_CUDA_PDL=0 or GGML_CUDA_DISABLE_GRAPHS=1 hid it. llama-bench speed is unaffected; test-backend-ops runs each op on its own and cannot catch it. The mul_mat_vec_q kernels in mmvq.cu make the same call. Ampere and Ada never take the PDL path.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the fix from sudoingX/bonsai2-small-gpu#3, opened here so it can go straight into
pr-ptq1-mmvand from there into PrismML-Eng#218.Bug: on Hopper and newer,
mul_mat_vec_ptq1_0_ptis launched with PDL throughggml_cuda_kernel_launch, but the kernel never callsggml_cuda_pdl_sync(). It can therefore readvybefore the q8_1 quantize kernel has finished writing it. On an RTX 5080 (sm_120),llama-serveroutput breaks into!!!!or////a few tokens in, for text and vision alike.Why it went unseen:
llama-benchruns at full speed with the bug (97 tok/s), andtest-backend-ops -o MUL_MATpasses every PTQ1_0 case because it runs each op on its own. Ampere and Ada never take the PDL path, so the 3060, 3060 Ti and 4070 rows were never affected.Isolation on the 5080 (greedy, thinking off):
285542d285542d,GGML_CUDA_DISABLE_GRAPHS=1285542d,GGML_CUDA_PDL=0285542d+ this commitCost: none. tg128 is 96.92 with the fix against 97.42 without, within noise. Full llama-bench ladder:
sweeps/rtx5080-16gb.mdin sudoingX/bonsai2-small-gpu#3.The change is one call at the top of the kernel, the same one the
mul_mat_vec_qkernels inmmvq.cumake.