Skip to content

tests : GET_ROWS -> RESHAPE -> GATED_DELTA_NET case, eager + CUDA graph capture/replay - #296

Open
professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:tests-gdn-gather-fusion
Open

professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:tests-gdn-gather-fusion

Conversation

@professorpalmer

Copy link
Copy Markdown

Overview

The test owed from #220 (Copilot review there, and my note on #220's close): the GET_ROWS -> RESHAPE -> GATED_DELTA_NET graph that ggml_cuda_try_gdn_gather_skip folds into the GDN kernel, with a nonzero cache row, run eagerly and through CUDA graph capture and replay.

  • test_case::n_eval() (default 1): evaluate the same graph that many times, with initialize_tensors between runs. On CUDA the first run is eager, the second captures a graph, the rest replay it.
  • test_gated_delta_net_gathered_state: one sequence, state gathered from a cache of 5 (or 3) rows, n_eval() = 4, and ids points at a different nonzero row on every run. A replay that kept the captured row would fail on run 3.

No new test file; this is in test-backend-ops.

Additional information

RTX 4070 (sm_89), CUDA 13, Windows, -o GATED_DELTA_NET -b CUDA0: 41/41 with the fusion on (default), with GGML_CUDA_GDN_GATHER_FUSION=0, and with GGML_CUDA_DISABLE_GRAPHS=1.

test-backend-ops cannot see whether the gather was actually skipped or a graph actually captured, so I checked that once with a local program (not in this PR) on the same graph: with the fusion on the gathered tensor stays untouched (NaN-filled before each run) and the backend logs CUDA graph warmup complete; with the fusion off it holds the selected row; with graphs disabled no capture happens. All four runs match the CPU in every mode (NMSE ~1e-14).

Negative control: with the fused kernel hard-wired to cache row 1 (gated_delta_net.cu, the s_ids[sequence] read), both new cases fail (ERR 1.9e-5 and 1.02 against 1e-7) and pass again with GGML_CUDA_GDN_GATHER_FUSION=0. The first evaluation reads row 1, so a single-evaluation test would have passed; the failure comes from the later runs.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - the test case, the n_eval hook and the checks above were written and run with Claude Code (Claude Opus 5.5).

🤖 Generated with Claude Code

…imes

The CUDA backend folds this gather into the GDN kernel (ggml_cuda_try_gdn_gather_skip) and no existing
case builds the graph. test_case::n_eval() evaluates the same graph several times with new inputs, so a
backend that records graphs is checked on the eager run, the capture and the replays; the case reads a
different nonzero cache row each time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant