Skip to content

Cuda 3 device binding - #87

Draft
max-models wants to merge 3 commits into
devel-tinyfrom
cuda-3-device-binding
Draft

max-models wants to merge 3 commits into
devel-tinyfrom
cuda-3-device-binding

Conversation

@max-models

@max-models max-models commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

Each MPI rank now uses its own GPU instead of always GPU 0.

Changes

  • feectools.ddm.cart: the CUDA setup that always picked GPU 0 is replaced by cunumpy.bind_local_device(). Each process uses GPU local_rank % device_count, where local_rank is the rank within the node that the MPI launcher exports.
  • Order: the CUDA context is created before MPI is initialized, as CUDA-aware MPI requires. Importing anything under feectools.ddm runs ddm/__init__.py first, which imports cart first, so the GPU is chosen before feectools.ddm.mpi starts MPI, whatever the user imports first.
  • Nothing changes on the NumPy backend.
  • New ddm/tests/test_device_binding.py: checks that the current device is local_rank % device_count after cart is imported. It is skipped without CuPy or a GPU.

Stack

This is step 3 of 4 for CUDA support. The PRs are stacked, and each one targets devel-tiny, so the diff here also contains the earlier steps. Review only commit 3a11ad1 in this PR, and merge them in order:

  1. Cuda 1 xp arrays #85 — Run feectools with the CuPy backend
  2. Cuda 2 mpi sync #86 — MPI with device buffers
  3. Cuda 3 device binding #87 — Bind each MPI rank to its own GPU ← this PR
  4. Cuda 4 device kernels #88 — Stencil operations on the device

🤖 Generated with Claude Code

max-models and others added 3 commits September 30, 2026 23:53
Make feectools work when cunumpy's backend is CuPy, without any device
kernels yet: kernels that stay on the host still copy their arrays.

- Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF
  kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays.
- Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart,
  partition, petsc) and fem.partitioning, Kronecker solver sizes, and
  index arithmetic with Python ints (compute_diag_len, math.prod).
- Stage data for host-only libraries by array, not by global backend:
  LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products.
- Fix calls that ran host kernels on device arrays: the second
  stencil2coo call in StencilMatrix.tosparse and the conjugate transpose.
- Vectorize the construction of the 1D collocation matrices in the
  global projectors (element-wise indexing was a device round trip per
  entry: 334 s of a 348 s Derham setup on the GPU).
- GMRES: take real scalars from CuPy views before modifying them.
- Fix StencilMatrix._update_ghost_regions_serial: the ghost region is
  pads * shifts wide (wrong whenever shifts > 1, on both backends).
- Tests: work with CuPy arrays; skip PETSc tests without petsc4py.

Serial tests pass on both backends (core, ddm, fem, linalg).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow MPI on the CuPy backend and make it correct with device buffers
(requires a CUDA-aware MPI library).

- feectools.ddm.mpi no longer disables MPI when ARRAY_BACKEND=cupy; the
  segfaults it guarded against come from MPI libraries that are not
  CUDA-aware.
- Call cunumpy.synchronize_for_mpi before every MPI call on device
  buffers: CuPy kernels run asynchronously and MPI does not know about
  CUDA streams, so a buffer still being written would be sent silently
  wrong. Covers the blocking, non-blocking and interface data exchangers,
  the Allreduce in StencilVectorSpace.inner and the Alltoallv calls of the
  parallel Kronecker solver. Requires cunumpy >= 0.3.0.
- Fix CuPy incompatibilities reached only by the MPI tests: xp.dot/vdot
  on .flat iterators in StencilInterfaceMatrix._dot and the pure-Python
  inner product, and test_cart_1d assigning Python lists to CuPy arrays.
- Add test_mpi_device.py: distributed results against global references,
  and that the exchangers synchronize before MPI.

With 2 MPI ranks and a CUDA-aware Open MPI, all MPI tests in ddm and
linalg pass on both backends; serial tests are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the initialization that always used GPU 0 by
cunumpy.bind_local_device(): each process uses GPU
local_rank % device_count, chosen from the node-local rank that the MPI
launcher exports, and its CUDA context is created before MPI is
initialized (as CUDA-aware MPI requires). No-op on the NumPy backend.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant