Skip to content

bug: Vulkan used VRAM underflows to ~2^64 when the driver reports heapBudget > heap size (Intel iGPU on Windows): free VRAM = 0, gpuLayers 0, every context fails #653

Description

@cmehta-xpl

Issue description

On Windows with an Intel integrated GPU (Arc B390), the Vulkan backend loads and detects the device,
but getVramState() reports 0 bytes free. inspect gpu shows Vulkan used VRAM: 47458790820.41% (16384PB/36.2GB). Because of that, gpuLayers: "auto" offloads 0 layers, gpuLayers: "max" throws
InsufficientMemoryError, and every createContext / createEmbeddingContext /
createRankingContext fails with A context size of 2048 is too large for the available VRAM, even at
contextSize: 512.

Root cause: an unsigned subtraction in llama/addon/globals/getGpuInfo.cpp, triggered by a driver
that reports a memory budget larger than the heap.

  1. The Intel Windows driver (101.8860) reports, for its single DEVICE_LOCAL heap (Windows vulkaninfo):
    memoryHeaps[0]:
        size   = 38868971888 (36.20 GiB)
        budget = 50206003200 (46.76 GiB)
        usage  = 4096
    
    This breaks the spec: "The heapBudget value must be less than or equal to
    VkMemoryHeap::size for each heap"

    (VkPhysicalDeviceMemoryBudgetPropertiesEXT).
    So the driver is at fault first.
  2. ggml_backend_vk_get_device_memory() (ggml-vulkan.cpp, b10361) passes it through unchanged:
    free += heapBudget[i] - heapUsage[i] (46.76 GiB) and total += heap.size (36.20 GiB), which gives
    free > total.
  3. getGpuVramInfo() then computes used += deviceTotal - deviceFree; with uint64_t used
    (getGpuInfo.cpp:37). 36.20 GiB − 46.76 GiB = −11,337,027,216, which wraps to 2⁶⁴ − 11.34 GB.
    The raw binding returns used = 18446744062372528000 (2⁶⁴ − 11,337,023,616 after double
    rounding), with total = 38868971888.
  4. getBalancedVramState() then clamps free = max(0, total - used) = 0.

The driver is out of spec, but a single defensive clamp here would make node-llama-cpp robust to it.
Without the clamp, one bad driver value disables the GPU completely and breaks context creation. It
does not even degrade to CPU.

Expected Behavior

used stays within [0, total] whatever the driver reports. For example:

// getGpuInfo.cpp
if (deviceFree > deviceTotal) deviceFree = deviceTotal;   // driver reported budget > heap size
used += deviceTotal - deviceFree;

With that clamp, free would read as the full 36.2 GB here. That is correct for a unified-memory iGPU
at 4 KiB usage, and auto offload works.

Actual Behavior

  • inspect gpu: Vulkan used VRAM: 47458790820.41% (16384PB/36.2GB), Vulkan free VRAM: 0% (0B/36.2GB).
  • loadModel({ gpuLayers: "auto" }) offloads 0 of 25 layers (embeddinggemma-300M) and 0 of 29
    (qwen3-reranker-0.6b), with no warning.
  • loadModel({ gpuLayers: "max" }) throws InsufficientMemoryError: Not enough VRAM to fit the model with the specified settings.
  • createEmbeddingContext({ contextSize: 2048 }) and { contextSize: 512 } both throw A context size of … is too large for the available VRAM.
  • Workaround that proves the device works: loadModel({ gpuLayers: "max", ignoreMemorySafetyChecks: true })
    plus ignoreMemorySafetyChecks: true on the contexts. This offloads 29/29 and 25/25 layers and runs
    correctly. Engine counters show 89-100% GPU compute. Retrieval results are identical in topic and
    top-1 to CPU and to a second Vulkan implementation (Mesa Dozen under WSL2 on the same GPU). A
    hybrid-search workload ran about 3× faster than Dozen.

Steps to reproduce

  1. Windows 11, Intel iGPU, Intel driver 101.8860 (Arc B390; see Additional Context for a UHD 770 report).
  2. npx --yes node-llama-cpp@3.20.0 inspect gpu shows used at 16384PB and free at 0B.
  3. Save the script as a .mjs file and run it. Do not use node -e: the binding self-test spawns a
    child that inherits the eval flags and fails.
    import { getLlama } from "node-llama-cpp";
    const llama = await getLlama({ gpu: "auto", build: "never" });
    console.log(llama._bindings.getGpuVramInfo());      // used ≈ 1.8446744062e19 > total
    console.log(await llama.getVramState());             // free: 0
    const model = await llama.loadModel({ modelPath: "embeddinggemma-300M-Q8_0.gguf" });
    console.log(model.gpuLayers);                         // 0
    await model.createEmbeddingContext({ contextSize: 512 }); // throws "too large for the available VRAM"
  4. Compare with Windows vulkaninfo, section VkPhysicalDeviceMemoryProperties: budget > size on heap 0.

My Environment

Dependency Version
Operating System Windows 11 (10.0.26200, x64)
CPU Intel Core Ultra X7 358H (16 math cores)
GPU Intel Arc B390, integrated, PHYSICAL_DEVICE_TYPE_INTEGRATED_GPU. Single heap, 36.20 GiB, DEVICE_LOCAL. VK_EXT_memory_budget rev 1
GPU driver Intel 101.8860 (32.0.101.8860). Vulkan apiVersion 1.4.348. Same symptom on 32.0.101.8724
Node.js 24.15.0
node-llama-cpp 3.20.0, prebuilt binaries b10361 (@node-llama-cpp/win-x64-vulkan)
Consumer @tobilu/qmd 2.8.3

Additional Context

  • Same GPU under WSL2 through Mesa Dozen (dzn, Mesa 26.2.2) reports sane values (23.5 GB free of
    47.5 GB). Auto offload works there (25/25 layers), which isolates the problem to the Windows driver's
    budget reporting plus the missing clamp.
  • A downstream report shows the same failure chain on a different Intel iGPU: Windows + Vulkan: integrated Intel iGPU is selected over the RTX 4090, getVramState() reports 0 MB free, and embedding dies with 'Failed to create any embedding context' (workaround: GGML_VK_VISIBLE_DEVICES=0) tobi/qmd#969. There, an
    Intel UHD 770 (driver 32.0.101.7088) alongside an RTX 4090 gave VRAM 0 B free / 55.4 GB total and
    A context size of 2048 is too large for the available VRAM. Their workaround,
    GGML_VK_VISIBLE_DEVICES=0 to hide the iGPU, is not available on an iGPU-only machine. That
    suggests the bad budget spans Intel driver generations, and that the summed-devices explanation in
    that issue was masking this underflow.
  • ggml-vulkan on current master still has no clamp (free += heapBudget[i] - heapUsage[i]). Clamping
    there as well would protect every ggml consumer. I'm happy to mirror this report to llama.cpp if
    that's the better venue.
  • The used == 0 && vulkanDeviceUsed != 0 fallback below the loop never fires here, because used
    is the wrapped non-zero value.

Relevant Features Used

  • Metal support
  • CUDA support
  • Vulkan support
  • Grammar
  • Function calling

Are you willing to resolve this issue by submitting a Pull Request?

No, I don’t have the time and I’m okay to wait for the community / maintainers to resolve this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions