You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
bug: Vulkan used VRAM underflows to ~2^64 when the driver reports heapBudget > heap size (Intel iGPU on Windows): free VRAM = 0, gpuLayers 0, every context fails #653
On Windows with an Intel integrated GPU (Arc B390), the Vulkan backend loads and detects the device,
but getVramState() reports 0 bytes free. inspect gpu shows Vulkan used VRAM: 47458790820.41% (16384PB/36.2GB). Because of that, gpuLayers: "auto" offloads 0 layers, gpuLayers: "max" throws InsufficientMemoryError, and everycreateContext / createEmbeddingContext / createRankingContext fails with A context size of 2048 is too large for the available VRAM, even at contextSize: 512.
Root cause: an unsigned subtraction in llama/addon/globals/getGpuInfo.cpp, triggered by a driver
that reports a memory budget larger than the heap.
The Intel Windows driver (101.8860) reports, for its single DEVICE_LOCAL heap (Windows vulkaninfo):
This breaks the spec: "The heapBudget value must be less than or equal to VkMemoryHeap::size for each heap"
(VkPhysicalDeviceMemoryBudgetPropertiesEXT).
So the driver is at fault first.
ggml_backend_vk_get_device_memory() (ggml-vulkan.cpp, b10361) passes it through unchanged: free += heapBudget[i] - heapUsage[i] (46.76 GiB) and total += heap.size (36.20 GiB), which gives free > total.
getGpuVramInfo() then computes used += deviceTotal - deviceFree; with uint64_t used
(getGpuInfo.cpp:37). 36.20 GiB − 46.76 GiB = −11,337,027,216, which wraps to 2⁶⁴ − 11.34 GB.
The raw binding returns used = 18446744062372528000 (2⁶⁴ − 11,337,023,616 after double
rounding), with total = 38868971888.
getBalancedVramState() then clamps free = max(0, total - used) = 0.
The driver is out of spec, but a single defensive clamp here would make node-llama-cpp robust to it.
Without the clamp, one bad driver value disables the GPU completely and breaks context creation. It
does not even degrade to CPU.
Expected Behavior
used stays within [0, total] whatever the driver reports. For example:
loadModel({ gpuLayers: "auto" }) offloads 0 of 25 layers (embeddinggemma-300M) and 0 of 29
(qwen3-reranker-0.6b), with no warning.
loadModel({ gpuLayers: "max" }) throws InsufficientMemoryError: Not enough VRAM to fit the model with the specified settings.
createEmbeddingContext({ contextSize: 2048 }) and { contextSize: 512 } both throw A context size of … is too large for the available VRAM.
Workaround that proves the device works:loadModel({ gpuLayers: "max", ignoreMemorySafetyChecks: true })
plus ignoreMemorySafetyChecks: true on the contexts. This offloads 29/29 and 25/25 layers and runs
correctly. Engine counters show 89-100% GPU compute. Retrieval results are identical in topic and
top-1 to CPU and to a second Vulkan implementation (Mesa Dozen under WSL2 on the same GPU). A
hybrid-search workload ran about 3× faster than Dozen.
Steps to reproduce
Windows 11, Intel iGPU, Intel driver 101.8860 (Arc B390; see Additional Context for a UHD 770 report).
npx --yes node-llama-cpp@3.20.0 inspect gpu shows used at 16384PB and free at 0B.
Save the script as a .mjs file and run it. Do not use node -e: the binding self-test spawns a
child that inherits the eval flags and fails.
import{getLlama}from"node-llama-cpp";constllama=awaitgetLlama({gpu: "auto",build: "never"});console.log(llama._bindings.getGpuVramInfo());// used ≈ 1.8446744062e19 > totalconsole.log(awaitllama.getVramState());// free: 0constmodel=awaitllama.loadModel({modelPath: "embeddinggemma-300M-Q8_0.gguf"});console.log(model.gpuLayers);// 0awaitmodel.createEmbeddingContext({contextSize: 512});// throws "too large for the available VRAM"
Compare with Windows vulkaninfo, section VkPhysicalDeviceMemoryProperties: budget > size on heap 0.
Same GPU under WSL2 through Mesa Dozen (dzn, Mesa 26.2.2) reports sane values (23.5 GB free of
47.5 GB). Auto offload works there (25/25 layers), which isolates the problem to the Windows driver's
budget reporting plus the missing clamp.
ggml-vulkan on current master still has no clamp (free += heapBudget[i] - heapUsage[i]). Clamping
there as well would protect every ggml consumer. I'm happy to mirror this report to llama.cpp if
that's the better venue.
The used == 0 && vulkanDeviceUsed != 0 fallback below the loop never fires here, because used
is the wrapped non-zero value.
Relevant Features Used
Metal support
CUDA support
Vulkan support
Grammar
Function calling
Are you willing to resolve this issue by submitting a Pull Request?
No, I don’t have the time and I’m okay to wait for the community / maintainers to resolve this issue.
Issue description
On Windows with an Intel integrated GPU (Arc B390), the Vulkan backend loads and detects the device,
but
getVramState()reports 0 bytes free.inspect gpushowsVulkan used VRAM: 47458790820.41% (16384PB/36.2GB). Because of that,gpuLayers: "auto"offloads 0 layers,gpuLayers: "max"throwsInsufficientMemoryError, and everycreateContext/createEmbeddingContext/createRankingContextfails withA context size of 2048 is too large for the available VRAM, even atcontextSize: 512.Root cause: an unsigned subtraction in
llama/addon/globals/getGpuInfo.cpp, triggered by a driverthat reports a memory budget larger than the heap.
DEVICE_LOCALheap (Windowsvulkaninfo):heapBudgetvalue must be less than or equal toVkMemoryHeap::sizefor each heap"(VkPhysicalDeviceMemoryBudgetPropertiesEXT).
So the driver is at fault first.
ggml_backend_vk_get_device_memory()(ggml-vulkan.cpp, b10361) passes it through unchanged:free += heapBudget[i] - heapUsage[i](46.76 GiB) andtotal += heap.size(36.20 GiB), which givesfree > total.
getGpuVramInfo()then computesused += deviceTotal - deviceFree;withuint64_t used(
getGpuInfo.cpp:37). 36.20 GiB − 46.76 GiB = −11,337,027,216, which wraps to 2⁶⁴ − 11.34 GB.The raw binding returns
used = 18446744062372528000(2⁶⁴ − 11,337,023,616 after doublerounding), with
total = 38868971888.getBalancedVramState()then clampsfree = max(0, total - used) = 0.The driver is out of spec, but a single defensive clamp here would make node-llama-cpp robust to it.
Without the clamp, one bad driver value disables the GPU completely and breaks context creation. It
does not even degrade to CPU.
Expected Behavior
usedstays within[0, total]whatever the driver reports. For example:With that clamp,
freewould read as the full 36.2 GB here. That is correct for a unified-memory iGPUat 4 KiB usage, and auto offload works.
Actual Behavior
inspect gpu:Vulkan used VRAM: 47458790820.41% (16384PB/36.2GB),Vulkan free VRAM: 0% (0B/36.2GB).loadModel({ gpuLayers: "auto" })offloads0of 25 layers (embeddinggemma-300M) and 0 of 29(qwen3-reranker-0.6b), with no warning.
loadModel({ gpuLayers: "max" })throwsInsufficientMemoryError: Not enough VRAM to fit the model with the specified settings.createEmbeddingContext({ contextSize: 2048 })and{ contextSize: 512 }both throwA context size of … is too large for the available VRAM.loadModel({ gpuLayers: "max", ignoreMemorySafetyChecks: true })plus
ignoreMemorySafetyChecks: trueon the contexts. This offloads 29/29 and 25/25 layers and runscorrectly. Engine counters show 89-100% GPU compute. Retrieval results are identical in topic and
top-1 to CPU and to a second Vulkan implementation (Mesa Dozen under WSL2 on the same GPU). A
hybrid-search workload ran about 3× faster than Dozen.
Steps to reproduce
npx --yes node-llama-cpp@3.20.0 inspect gpushowsusedat 16384PB and free at 0B..mjsfile and run it. Do not usenode -e: the binding self-test spawns achild that inherits the eval flags and fails.
vulkaninfo, sectionVkPhysicalDeviceMemoryProperties:budget > sizeon heap 0.My Environment
PHYSICAL_DEVICE_TYPE_INTEGRATED_GPU. Single heap, 36.20 GiB,DEVICE_LOCAL.VK_EXT_memory_budgetrev 1node-llama-cpp@node-llama-cpp/win-x64-vulkan)@tobilu/qmd2.8.3Additional Context
dzn, Mesa 26.2.2) reports sane values (23.5 GB free of47.5 GB). Auto offload works there (25/25 layers), which isolates the problem to the Windows driver's
budget reporting plus the missing clamp.
Intel UHD 770 (driver 32.0.101.7088) alongside an RTX 4090 gave
VRAM 0 B free / 55.4 GB totalandA context size of 2048 is too large for the available VRAM. Their workaround,GGML_VK_VISIBLE_DEVICES=0to hide the iGPU, is not available on an iGPU-only machine. Thatsuggests the bad budget spans Intel driver generations, and that the summed-devices explanation in
that issue was masking this underflow.
free += heapBudget[i] - heapUsage[i]). Clampingthere as well would protect every ggml consumer. I'm happy to mirror this report to llama.cpp if
that's the better venue.
used == 0 && vulkanDeviceUsed != 0fallback below the loop never fires here, becauseusedis the wrapped non-zero value.
Relevant Features Used
Are you willing to resolve this issue by submitting a Pull Request?
No, I don’t have the time and I’m okay to wait for the community / maintainers to resolve this issue.