Skip to content

sycl: FWHT for widths above 512 and Kronecker sizes (port of upstream #28254, #29243) - #302

Merged
bri-prism merged 2 commits into
prismfrom
sycl-fwht-wide-to-prism
Oct 2, 2026
Merged

bri-prism merged 2 commits into
prismfrom
sycl-fwht-wide-to-prism

Conversation

@bri-prism

Copy link
Copy Markdown
Collaborator

What

Brings two merged upstream SYCL FWHT changes to prism, cherry-picked with -x:

Why

On prism the SYCL FWHT dispatch stops at 512. Any Hadamard-hinted matmul wider than that falls back to a dense GEMM against the materialized rotation, which is O(n^2) per row instead of O(n log n). CUDA, Metal and Vulkan on prism already cover these widths, so SYCL was the only backend left on the dense path for them.

How

ggml/src/ggml-sycl/fwht.cpp is now byte-identical to upstream at c829670. The only manual change is the conflict resolution in tests/test-backend-ops.cpp: the upstream SYCL-only Kronecker cases are added next to the existing prism signed-FWHT and SwiGLU cases, and both are kept.

Testing

  • git diff --check clean; test-backend-ops builds on CPU with the merged test file.
  • Not yet built with oneAPI or run on an Intel GPU. The upstream kernel was verified on a SYCL CPU device only (see the sycl: FWHT kernels for block widths above 512 ggml-org/llama.cpp#29243 commit message). An Arc run of test-backend-ops and a decode comparison will follow here before this leaves draft.

philip-jingxin and others added 2 commits October 1, 2026 14:57
…T support (ggml-org#28016) (ggml-org#28254)

* Reapply "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (ggml-org#28184)

This reverts commit c845263.

* tests : fix unused variable M in test-backend-ops

* tests: fix trailing space error and isolate kronecker tests for sycl backend only

(cherry picked from commit 4d91760)
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus
384/640/768/1280 via the Kronecker/Paley construction added separately in
Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through
to the default case and run as a dense GEMM against the materialized
rotation tensor, correct but O(n^2) instead of O(n log n).

fwht_kernel_wide runs one row per work-group instead of per sub-group, so
each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below
the sub-group width still shuffle; those up to the work-group width go
through work-group local memory; the rest stay in registers. Same butterfly
and sign convention as the existing narrow kernel.

ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a
GPU-labeled platform to exist and throws before any op-level test can run,
so test-backend-ops could not be exercised on this box (a GPU-less pod) even
via the CPU device. Verified instead with a standalone harness: the same
kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1,
OpenCL CPU backend), checked against an independent recursive-doubling
Hadamard reference, cross-validated by first running the existing unmodified
narrow kernel through the identical harness and confirming it passes (rules
out a reference-convention bug before trusting a pass on the new code).
Random-input results for all four widths, single- and multi-row:

  N=1024 NT=256 rows=1  max_abs_err=1.7e-07  max_rel_err=4.9e-04  PASS
  N=2048 NT=256 rows=1  max_abs_err=1.9e-07  max_rel_err=2.0e-04  PASS
  N=4096 NT=256 rows=1  max_abs_err=2.0e-07  max_rel_err=1.4e-04  PASS
  N=8192 NT=256 rows=1  max_abs_err=2.5e-07  max_rel_err=3.8e-03  PASS
  N=1024 NT=256 rows=7  max_abs_err=2.4e-07  max_rel_err=1.0e-03  PASS
  N=2048 NT=256 rows=5  max_abs_err=3.0e-07  max_rel_err=9.4e-04  PASS
  N=4096 NT=256 rows=3  max_abs_err=2.7e-07  max_rel_err=1.7e-03  PASS
  N=8192 NT=256 rows=2  max_abs_err=2.5e-07  max_rel_err=1.9e-03  PASS

This covers the kernel algorithm itself; it does not exercise the ggml
dispatch/supports_op integration end to end, which needs a real GPU (or a
SYCL GPU plugin) to get past backend registration. test-backend-ops build
is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.

(cherry picked from commit c829670)
@bri-prism
bri-prism marked this pull request as ready for review October 2, 2026 07:35
@bri-prism
bri-prism merged commit f450c76 into prism Oct 2, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants