sweeps: RTX 5080 16GB row, plus a PDL sync bug in pr-ptq1-mmv on Blackwell - #3
Open
LamplighterPaul wants to merge 1 commit into
Open
LamplighterPaul wants to merge 1 commit into
LamplighterPaul wants to merge 1 commit into
Conversation
Stock b10685, stock b10735 and pr-ptq1-mmv at 285542d on the same card, the served VRAM for the 16GB vision tier, and the one-line PDL sync the PTQ1_0 mat-vec kernel needs on Hopper and newer: without it the branch serves garbage on sm_120 while llama-bench still runs at full speed.
Author
|
Fix opened against |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 16GB row, as promised on X, on an RTX 5080: the first Blackwell card in
sweeps/. It found a bug on the way.Bug:
pr-ptq1-mmvat285542dserves garbage on sm_120.llama-benchlooks fine (97 tok/s), but inllama-serverthe text breaks a few tokens in, into!!!!or////, for text and vision alike.GGML_CUDA_DISABLE_GRAPHS=1orGGML_CUDA_PDL=0makes it correct again.mul_mat_vec_ptq1_0_ptgoes throughggml_cuda_kernel_launch, which opts into PDL on Hopper and newer, but the kernel never callsggml_cuda_pdl_sync(), so it can readvybefore the q8_1 quantize kernel has finished writing it. One line fixes it, at no speed cost:Ampere and Ada never take the PDL path, which is why the 3060, 3060 Ti and 4070 rows are clean.
test-backend-opspasses every PTQ1_0 MUL_MAT case with the unfixed build, because it runs each op on its own. PrismML-Eng/llama.cpp#218 carries the same kernel, so it will hit every Hopper and Blackwell user once merged. I'm happy to open that as a PR on your fork if it helps.Numbers (llama-bench r=3, q4_0 K/V, fa on, stock b10685 / newest release b10735 / kernel+fix):
Decode gains less than on the 3060: at 960 GB/s the stock kernel is less bandwidth-starved.
VRAM: the full 262144 window fits with the vision tower: 12,684 MiB for the server, with 2.4 GB still free on a card that also drives the desktop.
16gb-vision.shcould move from 131072 to 262144; I left the script alone for you to decide.Full ladder, verbatim llama-bench output, the isolation table and the build notes are in
sweeps/rtx5080-16gb.md. Every number comes from runs on this machine.