Skip to content

sweeps: RTX 5080 16GB row, plus a PDL sync bug in pr-ptq1-mmv on Blackwell - #3

Open
LamplighterPaul wants to merge 1 commit into
sudoingX:mainfrom
LamplighterPaul:sweeps-rtx5080-16gb
Open

LamplighterPaul wants to merge 1 commit into
sudoingX:mainfrom
LamplighterPaul:sweeps-rtx5080-16gb

Conversation

@LamplighterPaul

Copy link
Copy Markdown

The 16GB row, as promised on X, on an RTX 5080: the first Blackwell card in sweeps/. It found a bug on the way.

Bug: pr-ptq1-mmv at 285542d serves garbage on sm_120. llama-bench looks fine (97 tok/s), but in llama-server the text breaks a few tokens in, into !!!! or ////, for text and vision alike. GGML_CUDA_DISABLE_GRAPHS=1 or GGML_CUDA_PDL=0 makes it correct again. mul_mat_vec_ptq1_0_pt goes through ggml_cuda_kernel_launch, which opts into PDL on Hopper and newer, but the kernel never calls ggml_cuda_pdl_sync(), so it can read vy before the q8_1 quantize kernel has finished writing it. One line fixes it, at no speed cost:

     float      * GGML_CUDA_RESTRICT dst = dst_ptr;
+    ggml_cuda_pdl_sync();
+
     extern __shared__ float partials_dyn[];

Ampere and Ada never take the PDL path, which is why the 3060, 3060 Ti and 4070 rows are clean. test-backend-ops passes every PTQ1_0 MUL_MAT case with the unfixed build, because it runs each op on its own. PrismML-Eng/llama.cpp#218 carries the same kernel, so it will hit every Hopper and Blackwell user once merged. I'm happy to open that as a PR on your fork if it helps.

Numbers (llama-bench r=3, q4_0 K/V, fa on, stock b10685 / newest release b10735 / kernel+fix):

test b10685 b10735 kernel+fix
tg128 87.58 87.51 96.92 (+10.8%)
tg128 @ d131072 41.96 41.82 44.19 (+5.7%)
pp512 995 2256 2179

Decode gains less than on the 3060: at 960 GB/s the stock kernel is less bandwidth-starved.

VRAM: the full 262144 window fits with the vision tower: 12,684 MiB for the server, with 2.4 GB still free on a card that also drives the desktop. 16gb-vision.sh could move from 131072 to 262144; I left the script alone for you to decide.

Full ladder, verbatim llama-bench output, the isolation table and the build notes are in sweeps/rtx5080-16gb.md. Every number comes from runs on this machine.

Stock b10685, stock b10735 and pr-ptq1-mmv at 285542d on the same card,
the served VRAM for the 16GB vision tier, and the one-line PDL sync the
PTQ1_0 mat-vec kernel needs on Hopper and newer: without it the branch
serves garbage on sm_120 while llama-bench still runs at full speed.
@LamplighterPaul

Copy link
Copy Markdown
Author

Fix opened against pr-ptq1-mmv: sudoingX/llama.cpp#1. Also flagged on PrismML-Eng/llama.cpp#218.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant