Thanks for node-llama-cpp — it's the backbone of our cross-platform host.
On an NVIDIA DGX Spark, CUDA is fully installed and the driver is loaded, but the linux-arm64 prebuilt has no CUDA support compiled in, so inference falls back to CPU on the Grace cores and the Blackwell GPU sits idle.
Environment
OS: Ubuntu 24.04.5 LTS (arm64), DGX OS, kernel 7.0.0-1019-nvidia
Node: 22.23.2 (arm64)
node-llama-cpp: 3.21.1 (prebuilt binaries v0.4.0)
GPU: NVIDIA GB10, compute capability 12.1 (sm_121)
Driver: 580.178.04, CUDA 13.0
CUDA target: /usr/local/cuda/targets/sbsa-linux (standard arm64 server CUDA, not Tegra)
RAM: 121 GB unified LPDDR5X (CPU+GPU coherent)
npx node-llama-cpp inspect gpu
OS: Ubuntu 24.04.5 LTS (arm64)
Node: 22.23.2 (arm64)
node-llama-cpp: 3.21.1
Prebuilt binaries: v0.4.0
CUDA: CUDA is detected, but using it failed
Vulkan: Vulkan is detected, but using it failed
CPU model: unknown
Math cores: 20
Used RAM: 6.32% (7.7GB/121.69GB)
Free RAM: 93.67% (113.99GB/121.69GB)
What's published
linux-x64, linux-x64-cuda, linux-x64-vulkan, linux-arm64, linux-armv7l, linux-riscv64, mac-arm64-metal, mac-x64, win-x64, win-x64-cuda, win-x64-vulkan, win-arm64.
There's a CUDA variant for x64 but not for arm64, so @node-llama-cpp/linux-arm64 installs cleanly and then runs CPU-only on hardware bought specifically for GPU inference.
Why this cell may be worth adding now
arm64 + CUDA on Linux used to be datacenter-only, but DGX Spark puts a Grace-Blackwell part on desks at a consumer-ish price, and GH200 / GB200 / Jetson Thor share the same shape. GitHub now offers free hosted ubuntu-24.04-arm runners, and the CUDA sbsa toolkit installs on them normally — so the matrix cell looks buildable in CI without self-hosted hardware.
One detail if you do add it
On GB10, nvidia-smi --query-gpu=memory.total,memory.free,utilization.gpu --format=csv,noheader,nounits returns literally:
and nvidia-smi shows Memory-Usage: Not Supported, because CPU and GPU share one coherent pool rather than the GPU having its own. Anything parsing VRAM gets NaN (or 0 via systeminformation) on these parts. That reading means "unified memory", not "probe failed" — we hit exactly this in our own memory accounting and had to special-case it.
Happy to test any build on real GB10 hardware, and to share what we learn if we end up doing a source build with CMAKE_CUDA_ARCHITECTURES=121 in the meantime.
Thanks for node-llama-cpp — it's the backbone of our cross-platform host.
On an NVIDIA DGX Spark, CUDA is fully installed and the driver is loaded, but the
linux-arm64prebuilt has no CUDA support compiled in, so inference falls back to CPU on the Grace cores and the Blackwell GPU sits idle.Environment
npx node-llama-cpp inspect gpuWhat's published
linux-x64,linux-x64-cuda,linux-x64-vulkan,linux-arm64,linux-armv7l,linux-riscv64,mac-arm64-metal,mac-x64,win-x64,win-x64-cuda,win-x64-vulkan,win-arm64.There's a CUDA variant for x64 but not for arm64, so
@node-llama-cpp/linux-arm64installs cleanly and then runs CPU-only on hardware bought specifically for GPU inference.Why this cell may be worth adding now
arm64 + CUDA on Linux used to be datacenter-only, but DGX Spark puts a Grace-Blackwell part on desks at a consumer-ish price, and GH200 / GB200 / Jetson Thor share the same shape. GitHub now offers free hosted
ubuntu-24.04-armrunners, and the CUDAsbsatoolkit installs on them normally — so the matrix cell looks buildable in CI without self-hosted hardware.One detail if you do add it
On GB10,
nvidia-smi --query-gpu=memory.total,memory.free,utilization.gpu --format=csv,noheader,nounitsreturns literally:and
nvidia-smishowsMemory-Usage: Not Supported, because CPU and GPU share one coherent pool rather than the GPU having its own. Anything parsing VRAM getsNaN(or0via systeminformation) on these parts. That reading means "unified memory", not "probe failed" — we hit exactly this in our own memory accounting and had to special-case it.Happy to test any build on real GB10 hardware, and to share what we learn if we end up doing a source build with
CMAKE_CUDA_ARCHITECTURES=121in the meantime.