Skip to content

fix(cuda/msvc): add CUDA 13.4 CUB fixes, Win32 CreateFileW, and graceful tensor name fallback - #300

Open
vrwallace wants to merge 14 commits into
PrismML-Eng:prismfrom
vrwallace:prism
Open

vrwallace wants to merge 14 commits into
PrismML-Eng:prismfrom
vrwallace:prism

Conversation

@vrwallace

Copy link
Copy Markdown

Summary of Changes

  1. CUDA 13.4 CUB Header Fixes:

    • Resolved compilation errors with NVIDIA CUDA Toolkit 13.4 (CCCL 3.4) in ggml/src/ggml-cuda/argsort.cu and ggml/src/ggml-cuda/top-k.cu.
    • Restricted version check conditions to safely fall back to native CUDA kernels on CUDA 13.4.
  2. Windows MSVC httplib Compatibility:

    • Replaced CreateFile2 with standard Win32 CreateFileW in vendor/cpp-httplib/httplib.cpp to fix MSVC desktop build linking.
  3. Graceful Model Tensor Name Fallback:

    • Updated src/llama-arch.cpp to return "unknown" instead of GGML_ABORT() on unrecognized tensor names, preventing process crashes on novel model architectures (e.g. Gemma 4).

Verification

  • Built and verified with Visual Studio 2022 (MSVC 19.44) + CUDA Toolkit 13.4 on NVIDIA RTX 4070 Ti.

khosravipasha and others added 14 commits March 2, 2026 10:49
…CUDA)

Adds two 1-bit quantization types:
- Q1_0: block size 32, ~1.5 bpw
- Q1_0_g128: block size 128, ~1.125 bpw

Backend support: CPU (x86 SSE/AVX + ARM NEON), Metal, CUDA.
Kernel implementations follow Q4_0 as boilerplate, adapted for
1-bit sign-based dequantization.

CUDA MMQ kernels included but disabled (cuBLAS fallback used for
prompt processing) pending accuracy debugging.

Made-with: Cursor
some cpu fixes; getting ready for upstream PR; e.g. id 40 is taken by…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants