Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 30 additions & 27 deletions .gitlab-ci.yml
Original file line number Diff line number Diff line change
@@ -1,12 +1,13 @@
variables:
# Use the specialized MPCDF HPC image
CUDA_IMAGE: "gitlab-registry.mpcdf.mpg.de/mpcdf/ci-module-image/nvhpcsdk_26:2026"
PIP_DISABLE_PIP_VERSION_CHECK: "1"
ARRAY_BACKEND: "cupy"
CUNUMPY_REQUIRE_CUDA: "1"

stages:
- test

gpu_tests:
.gpu_job:
stage: test
image: ${CUDA_IMAGE}
tags:
Expand All @@ -16,32 +17,34 @@ gpu_tests:
- module load python-waterboa/2025.06
- module load nvhpcsdk/26
- module load fftw-serial/3.3.10
script:
- echo "--- CUDA Sanity Check ---"
- nvidia-smi

- echo "--- Detect Compute Capability ---"
- nvidia-smi --query-gpu=compute_cap --format=csv,noheader

- echo "--- Tool Versions ---"
- cmake --version
- python3 --version
- git --version

- echo "--- Pytest Execution ---"
# The MPCDF image likely has a specific python environment.
# We install our dependencies into the user directory or a virtualenv.
# One CUDA version only: nvhpcsdk/26 provides CUDA 13.2 (its headers are used
# when CuPy compiles kernels with NVRTC), so install CuPy for CUDA 13 with the
# CUDA 13.2 libraries and NVRTC. Mixing CUDA 12 (cupy-cuda12x) with the CUDA 13.2
# headers fails to compile CuPy's own kernels (e.g. CUB reductions).
# Keep NVRTC, CUDA libraries and headers on the same toolkit generation.
- python3 -m pip install --user "cupy-cuda13x[ctk]" "cuda-toolkit==13.2.*"
- python3 -m pip install --user -e .

# Add the user bin to PATH for pytest
- python3 -m pip install --user -e ".[test]"
- export PATH="$HOME/.local/bin:$PATH"

- export ARRAY_BACKEND=cupy
- python3 -c "import cupy; cupy.show_config()"

- pytest -xvs .

gpu_tests:
extends: .gpu_job
script:
- python3 -m pytest -q tests

gpu_sanitizers:
extends: .gpu_job
timeout: 45 minutes
parallel:
matrix:
- SANITIZER: [memcheck, racecheck, synccheck]
script:
- compute-sanitizer --version
# These tests intentionally contain no illegal-access demonstration kernels.
- >-
compute-sanitizer --tool "$SANITIZER" --error-exitcode 1
--target-processes all python3 -m pytest -q
tests/unit/test_cuda_collectives.py
tests/unit/test_cuda_launch_limits.py
tests/unit/test_reusable_streams.py
tests/unit/test_staging.py
tests/unit/test_mirror.py
tests/unit/test_mpi_staging.py
tests/unit/test_segment_plans.py
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Added
- Per-device eager CUDA compilation, compiler log streams and explicit `recompile()`.
- Launch validation against actual device/kernel dimensions, thread limits and
static plus dynamic shared memory, with cached per-device opt-in state.
- Transfer payload byte counts, mirror-refresh and kernel-output accounting, and
separate device-only conversion observations.
- Producer `stream=`/`event=` dependencies for output staging and mirror downloads;
device-bound retained storage and CPU-to-pinned staging upgrades.
- GPU CI requires real hardware and runs focused Compute Sanitizer memory,
shared-memory race and synchronization checks.
- Reusable `cuda.create_stream`, `create_event`, `record_event`, and `wait_event`,
with synchronous host equivalents; `cuda.stream(existing)` selects an existing stream.
- `mpi.MPIStaging` reuses host staging storage and rejects overlapping device uses.
Expand Down Expand Up @@ -35,8 +44,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Development
- Formatting is checked with `ruff format` and import order with ruff's isort rules (`ruff check`), on `src/` and `tests/`, instead of black and isort; the `dev` extra installs ruff.
- GPU matrix/FFT correctness tests check numerical parity without timing assertions;
performance comparisons belong in scope-profiler.

### Fixed
- Block min/max reductions synchronize warp shared-memory reads before lane 0
overwrites the result slot, fixing Compute Sanitizer racecheck hazards.
- `CudaKernel`'s header hash (`-DCUNUMPY_INCLUDE_HASH`) now covers the headers shipped with cunumpy (`cunumpy/atomic.cuh`, `reduce.cuh`, ...), also when included in angle brackets. Before, an upgrade of cunumpy that changed one of them left CuPy's kernel cache serving the kernel compiled with the old header. `resolve_includes(..., angle_dirs=...)` tracks angle-bracket includes found in the given directories.
- `xp.testing.assert_kernels_agree` reads `CudaStructArguments` objects and struct values through their struct fields, so their arrays get the same names as the attributes of the host argument object (before, arrays behind properties were named after the private attribute holding the owner, and the comparison failed with "do not have the same array arguments").

Expand Down
16 changes: 10 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,17 +149,20 @@ with xp.use_backend("cupy"):

To verify that a block, such as a time step, makes no transfer at all, count
them: `count_transfers()` records every `to_numpy()`, `to_cupy()` and
`to_cunumpy()` call that actually copies, every `PyccelKernel` call that
converts device arrays, and every `Kernel` fallback to the host kernel, with
the call site of each. `assert_no_transfers()` raises with that report if
anything was counted. Only transfers made through CuNumpy are seen; raw
`to_cunumpy()` call that actually copies, mirror/staging refreshes, argument
conversion and kernel output copy-back, with call sites and payload byte counts.
Host kernel conversions and fallbacks have separate explanatory markers.
`assert_no_transfers()` rejects host/device movement and permits device-only
conversions. Only CuNumpy execution/conversion helpers are counted; forwarded
backend calls such as `xp.asarray()` and raw
`cupy.ndarray.get()` or `cupy.asarray()` calls need a profiler such as `nsys`.

```python
with xp.profiling.count_transfers() as counter:
propagator(dt)

assert counter.total == 0, counter.report()
assert counter.to_host == counter.to_device == 0, counter.report()
print(counter.bytes_to_host, counter.bytes_to_device)
```

## Random numbers and dtypes
Expand Down Expand Up @@ -368,7 +371,8 @@ When building such argument objects, `xp.as_device_array(value, dtype,
ndim=None)` applies the "reference or copy once" rule: a CuPy array that
already has the dtype and is C-contiguous is returned as it is, anything else
(a tuple such as `degree = (3, 3, 3)`, a host array, another dtype, a
non-contiguous view) becomes one device copy. Call it once when the object is
non-contiguous view) is converted; dtype and layout changes can require separate
device copies. Call it once when the object is
built, not per kernel call; on the NumPy backend it raises, so host data is
never copied to the device implicitly.
When the host kernel takes such a group as one object too (e.g. a Pyccel class
Expand Down
74 changes: 53 additions & 21 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -283,30 +283,34 @@ with xp.profiling.count_transfers() as counter:
assert counter.total == 0, counter.report()
```

Four kinds of events are recorded:
Five kinds of events are recorded:

* `to_host`: `to_numpy()` (or `to_cunumpy()`) called with a CuPy array;
* `to_device`: `to_cupy()` (or `to_cunumpy()`) called with anything that is not
a CuPy array already;
* `to_host`: an actual device-to-host copy through conversion, mirror, staging,
serial MPI, or host kernel helpers;
* `to_device`: an actual host-to-device copy through those helpers, including
`as_device_array()` and `kernel_output()` copy-back;
* `kernel_conversion`: a `PyccelKernel` call that copied device arrays to the
host (and back), one event per call, naming the kernel and the number of
arrays converted;
* `fallback`: a `Kernel` without CUDA kernel calling its host kernel on the
CuPy backend (`missing_cuda="fallback"`), one event per call, naming the
kernel. The host copies it makes are counted as one `kernel_conversion`
event in addition.
kernel. Physical copies and a `kernel_conversion` marker are recorded separately;
* `device_copy`: a device-only dtype/layout conversion through CuNumpy helpers.

Only real transfers count: `to_numpy()` of a NumPy array or `to_cupy()` of a
CuPy array records nothing. The counter has the attributes `to_host`,
`to_device`, `kernel_conversions`, `fallbacks` (counts per kind), `total`,
`events` (a list of `TransferEvent(kind, description, where)`, where `where`
is the `file:line` of the caller outside CuNumpy) and
`kernel_conversion_calls` (the `kernel_conversion` events). `report()` returns
`device_copies`, `bytes_to_host`, `bytes_to_device`, and `bytes(kind)`.
`total` includes physical copies and conversion/fallback markers. Byte totals
include only physical copies, so markers do not double-count bytes.
`events` is a list of `TransferEvent(kind, description, where, nbytes=None)`;
`where` is the caller's `file:line` and `nbytes` is None for markers/unknown sizes.
`kernel_conversion_calls` selects the conversion markers. `report()` returns
a multi-line string with the events grouped by kind and call site, with
counts:

```text
4 transfer(s) through cunumpy (3 to_host, 1 to_device, 0 kernel_conversion, 0 fallback)
4 transfer(s) through cunumpy (3 to_host, 1 to_device, 0 kernel_conversion, 0 fallback, 0 device_copy)
to_host (3):
/home/me/sim/diagnostics.py:42: to_numpy(shape=(100000,), dtype=float64) (x3)
to_device (1):
Expand All @@ -320,13 +324,15 @@ not thread-safe.

**Limitation:** only transfers made through CuNumpy are seen. Raw
`cupy.ndarray.get()`, `cupy.asarray(numpy_array)`, `numpy.asarray(cupy_array)`,
`float(device_array)`, and implicit conversions inside other libraries are
`float(device_array)`, forwarded backend calls such as `xp.asarray()`, and
implicit conversions inside other libraries are
not counted. Use `nsys` (or CuPy's profiling hooks) to find those.

### `profiling.assert_no_transfers()`

Context manager that raises `AssertionError` with the counter's `report()` if
the block makes a transfer through CuNumpy. It yields the `TransferCounter`
the block makes a host/device transfer or host fallback through CuNumpy.
Device-only dtype/layout conversions are allowed. It yields the `TransferCounter`
too. An exception raised inside the block propagates as it is:

```python
Expand All @@ -344,11 +350,14 @@ argument object is built, never per kernel call:
* a CuPy array that already has `dtype` (any dtype if `dtype` is `None`) and
is C-contiguous is returned unchanged, the same object without a copy, so
kernels write into the caller's array;
* anything else becomes one C-contiguous device copy,
* anything else is converted to a C-contiguous device array using
`cupy.ascontiguousarray(cupy.asarray(value, dtype))`: tuples and lists
(`degree = (3, 3, 3)`), host NumPy arrays (one explicit transfer at build
time), device arrays of another dtype, and non-contiguous views.

Dtype and layout conversion can require separate device copies; each copy
through this helper appears in transfer accounting.

The result passes the pointer checks of `CudaKernel` and `CudaStruct`. On the
NumPy backend it raises `RuntimeError`: device arguments are only built when
running on CuPy, and host data is never copied to the device implicitly. If
Expand All @@ -368,7 +377,7 @@ class DeviceParticles(xp.cuda.CudaArguments):
For the arguments of a `Kernel` with `dispatch="arrays"`, whose choice follows
the arrays. `as_kernel_array` returns `value` on the side of `like` (a CuPy
array if `like` is one, a NumPy array otherwise), C-contiguous and with `dtype`
(any if `None`): `value` itself if it already is such an array, else one copy,
(any if `None`): `value` itself if it already is such an array, else a conversion,
moved across if needed (counted by `count_transfers()`). `kernel_output` is a
context manager yielding the buffer for an array the kernel writes: `out`
itself if `as_kernel_array` takes it unchanged, else a converted copy whose
Expand Down Expand Up @@ -871,7 +880,12 @@ Wraps the `__global__` function `name` in the CUDA C `source` (declared
through CuPy on the first call (or by `compile()`), and cached, also on disk by
CuPy. CuPy is imported only then, so kernels can be created and their
signatures parsed without CuPy; `compile()` raises `RuntimeError` without a
GPU.
GPU. `compile(log_stream=None)` compiles eagerly and returns the CuPy raw kernel;
repeated calls reuse successful state on the current device. `is_compiled`
reports that device's state. `recompile(log_stream=None)` refreshes headers and
options on the current device; finish in-flight launches before rebuilding.
Compiler failures remain retryable and a writable `log_stream` receives compiler
output. Catalog compilation therefore reports errors during setup.

`from_file` reads the source from a file; the kernel name defaults to the file
name without `suffix` (`axpy_cuda.cu` -> `axpy`), and the directory of the file
Expand Down Expand Up @@ -972,10 +986,10 @@ shape is given either by `n_threads` or by `grid`:
* `grid`: number of blocks per dimension, instead of `n_threads`.
* `block`: block shape for this call, instead of `block_size`.
* `shared_mem`: dynamic shared memory per block in bytes, for
`extern __shared__` arrays. Above 48 KiB (the limit every device has) the
compiled kernel's `max_dynamic_shared_size_bytes` is raised to `shared_mem`
once, up to the device's opt-in limit (`xp.cuda.max_shared_memory_per_block(
opt_in=True)`); a larger request raises `ValueError` before the launch.
`extern __shared__` arrays. Static plus dynamic storage is checked against
the device's opt-in limit. Dynamic storage above the default allowance after
subtracting static storage opts in through `max_dynamic_shared_size_bytes`.
Block/grid dimensions and device/kernel thread limits are also checked.

Without `n_threads` and `grid`, a launch uses `n_threads_from(args)` if the
kernel has one (constructor argument and settable property): a function of the
Expand Down Expand Up @@ -1175,7 +1189,7 @@ too); `xp.cuda.get_cuda_debug()` returns the current setting. A kernel created w
enabling it also affects kernels created earlier; `debug=True` or
`debug=False` fix the mode for one kernel. Only the compile options are fixed
at compile time: a kernel compiled before debug mode was enabled keeps its
options, so call `compile()` after enabling, or create the kernels after
options, so call `recompile()` after enabling, or create the kernels after
enabling. `kernel.debug_active()` tells whether debug mode applies to a
kernel now, and `kernel.compile_options()` returns the options a compilation
now would use.
Expand Down Expand Up @@ -1972,6 +1986,16 @@ buffer is reused `buffers` copies later; a stale result raises
the NumPy backend copy at once. Device copies are counted by
`count_transfers()`. The arrays must have the staging shape and dtype.

`copy(array, *, stream=None, event=None)` accepts an explicit producer stream or
event (at most one). An event queues a wait before snapshotting on the current
stream. Later writes on another stream must wait for the snapshot/copy;
`copy.result()` is a conservative completion point. Keep the source's device
current when submitting copies. Storage binds to that device and rejects another
device. `ready()`/`result()` temporarily select the owning device and restore the
caller's device. The first GPU use of CPU-initialized storage allocates fresh
pinned slots; previous completed CPU handles keep their snapshots. `shape` and
`dtype` are read-only.

## `memory.DeviceMirror`

```python
Expand All @@ -1989,7 +2013,7 @@ raises `TypeError`.
transfer). On the NumPy backend it is the host array itself, so the same
code runs without any copy on the CPU.
* `to_device()`: copies the host array into the existing device array;
`to_host()`: copies the device array into the host array, in place, so the
`to_host(stream=None, event=None)`: copies the device array into the host array, in place, so the
host array keeps its identity and the owning library sees the new values.
Both are no-ops on the NumPy backend.
* `zero()`: zeroes the device array (allocating it empty if needed), or the
Expand All @@ -2001,6 +2025,14 @@ raises `TypeError`.
* `host`, `shape`, `dtype` properties. `to_device()`, `to_host()`, `zero()`
and `rebind()` return the mirror, for chaining.

Make the mirror's device current before using its device storage. A mismatched
device raises before copying or zeroing. Pass an explicit producer `stream=` or
`event=` to `to_host()` when production happened outside the current stream;
the host copy is complete on return. Initial uploads and refreshes in both
directions appear in transfer accounting, with payload byte counts. Switching
to NumPy uses the host array; explicitly refresh with `to_device()` after CPU
changes before resuming GPU work.

The transfers are explicit so that one per accumulation is visible and
bounded:

Expand Down
Loading
Loading