Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added
- Reusable `cuda.create_stream`, `create_event`, `record_event`, and `wait_event`,
with synchronous host equivalents; `cuda.stream(existing)` selects an existing stream.
- `mpi.MPIStaging` reuses host staging storage and rejects overlapping device uses.
- `algorithms.cell_offsets`, `segment_boundaries`, and `SegmentPlan` for prepared grouping.
- Strict backend selection via `set_backend`/`use_backend(..., strict=True)` and
JSON-compatible `backend_info()` diagnostics.
- `cunumpy/scan.cuh` inclusive/exclusive warp and block prefix sums, mask-aware
warp reductions, partial-warp block reductions, and integer atomic-add helpers.

### Execution helpers
- `segment_sum(..., out=...)` supports arbitrary trailing component dimensions;
CUDA reduces components in one accumulation launch. Prepared plans avoid
repeated key validation and GPU scalar reads. Floating-point order may vary.
- MPI producer synchronization accepts a stream/event, or waits for all work on
the buffer's device. Receive staging waits for copy-back before releasing storage.

### Changed (breaking, with deprecation)
- The helpers moved from the top level of `cunumpy` to submodules, so that the top level is the NumPy/CuPy namespace plus backend selection and array conversion, and no helper hides a NumPy or CuPy name (`xp.fuse` hid `cupy.fuse`): `cunumpy.cuda` (CUDA only: `CudaKernel`, `CudaKernelVariants`, `CudaStruct*`, `CudaArguments`, `CudaParameter`, header tools, debug mode, device selection and memory, `stream`, `pin_memory`), `cunumpy.kernels` (`Kernel`, `KernelCatalog`, `PyccelKernel`, `KernelArguments`, `PyccelStructArguments`, host implementations, `as_kernel_array`, `kernel_output`, `fuse`), `cunumpy.rng` (`random_streams`, `RandomStreams`, `get_rng`, `philox_*`), `cunumpy.algorithms` (`morton_*`, `sort_by_key`, `segment_sum`), `cunumpy.mpi` (`mpi_buffer`, CUDA-aware MPI, `local_rank`, `synchronize_for_mpi`), `cunumpy.profiling` (`timed_region`, `Timing`, `nvtx_range`, transfer counting), `cunumpy.memory` (`HostStaging`, `StagedCopy`, `DeviceMirror`) and `cunumpy.petsc` (`petsc_vec`). All are imported by `import cunumpy`. The old top-level names still work and raise a `DeprecationWarning` naming the new place; they will be removed in 0.6.
- `cunumpy.testing` is now `cunumpy.kernel_testing`. Once imported, `cunumpy.testing` replaced NumPy's `xp.testing`, so `xp.testing.assert_allclose` failed in every test that ran after an `import cunumpy.testing`. `cunumpy.testing` still works, with a `DeprecationWarning`, until 0.6.
Expand Down
28 changes: 19 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,8 @@ never hide a NumPy name:
| `xp.cuda` | CUDA only: `CudaKernel`, `CudaStruct`, CUDA headers, devices, streams |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, host implementations, `fuse` |
| `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
| `xp.algorithms` | `morton_*`, `sort_by_key`, `segment_sum` |
| `xp.mpi` | `mpi_buffer`, CUDA-aware MPI |
| `xp.algorithms` | `morton_*`, `sort_by_key`, `cell_offsets`, `segment_boundaries`, `segment_sum`, `SegmentPlan` |
| `xp.mpi` | `mpi_buffer`, reusable `MPIStaging`, CUDA-aware MPI |
| `xp.profiling` | `timed_region`, `nvtx_range`, `count_transfers` |
| `xp.memory` | `HostStaging`, `DeviceMirror` |
| `xp.petsc` | `petsc_vec` |
Expand Down Expand Up @@ -72,6 +72,12 @@ but unavailable or not functional, CuNumpy falls back to NumPy. Always check
`get_backend()` when the effective backend matters, such as when reporting
configuration or deciding whether GPU-specific work will happen.

Use `xp.set_backend("cupy", strict=True)` to raise when CUDA is unavailable,
preserving the previous backend. `xp.backend_info()` returns structured backend,
dependency, and CUDA diagnostics. Reusable streams/events, MPI staging, cell
ranges, and prepared reductions are described in the
[execution helpers guide](docs/source/guides/execution-helpers.md).

Use `use_backend()` for a temporary selection. It restores the previous
selection when the block exits, including when an exception is raised:

Expand Down Expand Up @@ -253,8 +259,7 @@ print(timing.elapsed, timing.synced)


@xp.profiling.nvtx_range("step")
def step(dt):
...
def step(dt): ...
```

## Use NumPy-only kernels with CuPy arrays
Expand Down Expand Up @@ -402,11 +407,13 @@ class is the one definition, and written to a header that a test keeps in sync:

```python
class MarkerArguments:
def __init__(self, markers: "float[:, :]", n_markers: int, valid: "bool[:]"):
...
def __init__(self, markers: "float[:, :]", n_markers: int, valid: "bool[:]"): ...


MarkerArgs = xp.cuda.CudaStruct.from_signature(MarkerArguments.__init__, "MarkerArgs")
MarkerArgs.to_header("marker_args.cuh") # Array2D<double> markers; long long n_markers; ...
MarkerArgs.to_header(
"marker_args.cuh"
) # Array2D<double> markers; long long n_markers; ...
push = xp.cuda.CudaKernel(
r"""
#include "marker_args.cuh"
Expand All @@ -419,8 +426,11 @@ push = xp.cuda.CudaKernel(
structs=[MarkerArgs],
include_dirs=["."],
)
push(MarkerArgs(markers=markers, n_markers=markers.shape[0], valid=valid), 0.1,
n_threads=markers.shape[0])
push(
MarkerArgs(markers=markers, n_markers=markers.shape[0], valid=valid),
0.1,
n_threads=markers.shape[0],
)
```

Launches can be 1D to 3D (`n_threads=(nx, ny)`, `block_size=(16, 16)`) or use
Expand Down
Loading
Loading