Skip to content

Add a linux-arm64-cuda prebuilt (NVIDIA Grace-class: DGX Spark GB10, GH200, Jetson Thor) #651

Description

@hassanhabib

Thanks for node-llama-cpp — it's the backbone of our cross-platform host.

On an NVIDIA DGX Spark, CUDA is fully installed and the driver is loaded, but the linux-arm64 prebuilt has no CUDA support compiled in, so inference falls back to CPU on the Grace cores and the Blackwell GPU sits idle.

Environment

OS:     Ubuntu 24.04.5 LTS (arm64), DGX OS, kernel 7.0.0-1019-nvidia
Node:   22.23.2 (arm64)
node-llama-cpp: 3.21.1 (prebuilt binaries v0.4.0)
GPU:    NVIDIA GB10, compute capability 12.1 (sm_121)
Driver: 580.178.04, CUDA 13.0
CUDA target: /usr/local/cuda/targets/sbsa-linux  (standard arm64 server CUDA, not Tegra)
RAM:    121 GB unified LPDDR5X (CPU+GPU coherent)

npx node-llama-cpp inspect gpu

OS: Ubuntu 24.04.5 LTS (arm64)
Node: 22.23.2 (arm64)

node-llama-cpp: 3.21.1
Prebuilt binaries: v0.4.0

CUDA: CUDA is detected, but using it failed
Vulkan: Vulkan is detected, but using it failed

CPU model: unknown
Math cores: 20
Used RAM: 6.32% (7.7GB/121.69GB)
Free RAM: 93.67% (113.99GB/121.69GB)

What's published

linux-x64, linux-x64-cuda, linux-x64-vulkan, linux-arm64, linux-armv7l, linux-riscv64, mac-arm64-metal, mac-x64, win-x64, win-x64-cuda, win-x64-vulkan, win-arm64.

There's a CUDA variant for x64 but not for arm64, so @node-llama-cpp/linux-arm64 installs cleanly and then runs CPU-only on hardware bought specifically for GPU inference.

Why this cell may be worth adding now

arm64 + CUDA on Linux used to be datacenter-only, but DGX Spark puts a Grace-Blackwell part on desks at a consumer-ish price, and GH200 / GB200 / Jetson Thor share the same shape. GitHub now offers free hosted ubuntu-24.04-arm runners, and the CUDA sbsa toolkit installs on them normally — so the matrix cell looks buildable in CI without self-hosted hardware.

One detail if you do add it

On GB10, nvidia-smi --query-gpu=memory.total,memory.free,utilization.gpu --format=csv,noheader,nounits returns literally:

[N/A], [N/A], 0

and nvidia-smi shows Memory-Usage: Not Supported, because CPU and GPU share one coherent pool rather than the GPU having its own. Anything parsing VRAM gets NaN (or 0 via systeminformation) on these parts. That reading means "unified memory", not "probe failed" — we hit exactly this in our own memory accounting and had to special-case it.

Happy to test any build on real GB10 hardware, and to share what we learn if we end up doing a source build with CMAKE_CUDA_ARCHITECTURES=121 in the meantime.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions