Conversation
…base into bugfix/BLS-kernel
…riginal-timescale phi convention - data() now applies t0=4.5 before the model is evaluated, so every test exercises a non-trivial epoch (floor(min(t)) = 4 or 5) and the injected transit stays at original-timescale phase phi0 (attila's suggested reproduction, applied as a permanent strengthening). - test_ignore_positive_sols: drop the old-convention manual phi shift (single_bls converts internally now); update the deterministic hardcoded value for the rotated fold. - test_single_bls_bjd_invariance: shifted runs use covariantly shifted phases (phi0 + offset*freq) mod 1 per the new convention. - new test_single_bls_phase_is_original_timescale: unshifted phi0 on shifted times must MISS the transit (guards the convention). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…airness rules) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…precision, sparse f64 phase conversion Root cause of the test_standard failures attila reported with non-zero T0 (61/216 on RTX A5000): NOT the new epoch re-referencing (the float64 phi round-trip is bit-exact at float32 precision), but a pre-existing precision asymmetry in single_bls exposed by the shifted fixture. single_bls subtracted phi0 from the UNWRAPPED float32 product t*f (magnitude ~360 for a 1-yr baseline, ulp 3.05e-5) while the GPU kernels wrap into [0,1) first (ulp ~1e-7): a point whose true phase sat 8.3e-6 below the best box edge rounded to phase exactly 0.0 on the CPU side and flipped into the box, a ~power/n_in_transit (=0.036) disagreement at one frequency, which trips the zero-violation mostly_ok criterion. Hardware probe confirms nvcc does NOT FMA-contract mod1(t*f), so wrapping first makes the reference fold bit-identical to the kernels'. - single_bls: wrap phase into [0,1) before subtracting phi0 (documented; shrinks the CPU-vs-GPU edge-disagreement window ~150x, and up to ~2000x for 10-yr baselines). - bin_and_phase_fold_custom: fold with the float32-cast frequency (folding with the double freq shifted phases by up to |f64-f32|*t ~1e-5 vs the reference); keep double freqs ONLY for the epoch re-referencing. phi_values now uploaded as float64 so the in-kernel (phi - epoch*freq) % 1 matches single_bls's float64 conversion bit for bit (store_best_sols_custom signature updated accordingly). - sparse_bls_cpu / sparse_bls_gpu: convert solutions to the original timescale with the caller's float64 frequencies, not the float32-cast copies (error epoch*|f64-f32| reaches ~0.07 cycles at BJD-scale epochs). - eebls_transit_gpu: uniform 3-tuple return (sols=None on the fast/optimized paths), matching eebls_transit. - eebls_transit(use_optimized=True): respect an explicit block_size instead of silently overriding it. - tests: test_transit uses a single 3-way mode axis (standard/fast/optimized) instead of a use_fast x use_optimized cross-product; test_standard/test_custom drop the use_optimized axis (only reduction_max differs) in favor of focused equivalence tests; new TestHoneSolution exercises hone_solution end-to-end at maximal epoch-phase rotation ((epoch*freq) % 1 = 0.5). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… decomposition Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…00, Jul 2026) Headline: warm kernel parity (9.8ms vs 9.8ms kernel-only); 34x per-LC on the naive loop path (0.2.6 recompiles per call); BJD-timestamp demo (0.2.6 fails, v1.0 recovers); 0.2.6 LS segfaults on modern pycuda; 0.2.6 never shipped to PyPI (last PyPI release is 0.2.5, Oct 2023). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… + pod validation), docstring touch-up - CHANGELOG.rst: original-input-timescale phi convention (PR #65) and the single_bls fold-order precision fix + custom-kernel/sparse conversion parity fixes. - analysis/pr65-resolution-jul2026/: SUMMARY.md (mechanism, evidence, draft reply to attila) + raw logs (61/216 reproduction, per-frequency diagnostics, flip-point drill-down, FMA hardware probe, full-suite 752/752-passed log, release-gate ALL-PASSED log). - eebls_transit use_optimized docstring: block-size auto-selection is conditional on no explicit block_size (matches the code fix). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix BLS bugs and return correct phi values after time stamp normalization
…1.0.0 release notes)
- CHANGELOG: correct release lineage (0.2.5 is the last PyPI release; the 0.2.6 tag was never published), add the measured v0.2.6 head-to-head summary, strip internal tracker IDs and analysis/ paths from user-facing entries, add a 0.2.6 stub entry - BENCHMARK_RESULTS: retract the compile-overhead-confounded 'vs pre-v1.0 kernel' column, add the measured July 2026 head-to-head table, correct the stale batch-mode guidance (batch now wins at every measured scale) - New docs/RELEASE_NOTES_v1.0.0.md: outward-facing release notes vs 0.2.5 - INSTALL.rst rewritten for Python 3.9+/CUDA 11-12 (was CUDA 8 + Python 2.7) - MANIFEST.in so CHANGELOG/INSTALL ship in the sdist; README changelog link absolutized for PyPI rendering; Development Status classifier -> Stable Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
bench_bls_survey.py: 4 survey configs (ZTF/HAT-Net/TESS/Kepler) on Keplerian grids (0.5q..2q), variants fast_naive/fast_reuse/kernel/ kernel_1pass/pieces/batch, warm medians >=5 runs, cold reported separately, $/lightcurve from pod rate, parity dumps for before/after gates. profile_bls_survey.py: minimal kernel-only driver for ncu/nsys. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… tracing) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…gnosis Closes the item deferred from commit 14b0d90 (cunfft.cu A3 fix): the same float32 PI literal in lomb.cu, tls.cu and nufft_lrt.cu, each assessed for actual exposure and given the same treatment (PI double literal under DOUBLE_PRECISION, float32 literal otherwise). lomb.cu (live defect, rigorous gate): PI feeds the un-reduced phase 2*pi*f*(t+0.5) in cossum/sinsum, used by the direct-sums kernels (use_fft=False; the production NFFT path never touches PI). In double mode the float32 pi (rel. err. +2.7828e-8) makes the kernel evaluate the exact periodogram on a frequency axis stretched by 1+2.78e-8 - observable error ~ eps*f*T of the local slope. Measured on an A5000 at f=90-110 c/d, T=30 d vs a float64 CPU port of the kernel (exact pi, same op order): before 1.162e-4 max abs power error in every f64 case; after 3.67e-10 (nonzero epoch, both low-level and run() paths; 316,517x closer) and 1.22e-8 (raw BJD 2.45e6 epoch through the low-level API; 9,505x). The buggy kernel matches the same port evaluated with float64(float32(pi)) to 3.7e-10 - the diagnosis exactly. float32-mode periodograms are bit-identical (the literal's value is unchanged). New regression test test_ls_kernel_direct_sums_double_pi (threshold 1e-7; measured after 1.1e-10, buggy 1.2e-4). nufft_lrt.cu (PI dead, double mode real): literal moved under the DOUBLE_PRECISION guard; hardcoded float32 helpers on FLT operands (fmaxf/fmodf/fabsf) retyped to the overloaded forms - the same class 14b0d90 fixed for modflt/diffmod. Double-mode matched filter now matches a float64 numpy port exactly (was rel. err. 3.5e-9 from the fmaxf truncation); float32 mode bit-identical. tls.cu (PI dead, kernel float32-only by design): dead macro removed; TLS search output bit-identical before/after (period/depth/SDE and full chi2/SR/power arrays). Validation (RTX A5000, CUDA 12.4): full GPU suite 753 passed / 7 skipped (752 baseline at ff326c2 + the new regression test); scripts/check_release_gate.py 14/14 PASS. Evidence + raw before/after runs archived in analysis/kernel-hygiene-jul2026/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BLS is a weighted fit (w_i = dy_i^-2 / sum dy_j^-2): a point with a near-zero reported uncertainty concentrates essentially all statistical weight in one phase box, so the best box absorbs ~all the weighted variance and the power P = 1 - chi2/chi2_0 saturates near 1 (~0.99 after binning) deterministically in pure noise, at nearly every trial frequency. Documented where BLS users will see it: a new "Data hygiene: near-zero uncertainties" section in docs/source/bls.rst (mechanism, symptoms - one point dominating sum(1/dy^2), suspiciously flat ~0.99 power on noise - and a copy-pasteable percentile error-floor + weight-check guard) and matching .. warning:: blocks in the eebls_gpu_fast / eebls_gpu / eebls_transit_gpu docstrings. Docs only - no runtime warnings or behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause found via nsys + cgroup counters: OpenBLAS spawns nproc(=96) threads inside per-LC np.dot calls; on a 7.65-CPU-quota RunPod container the CFS quota freezes the process ~90 ms per 100 ms period (GPU idle). TESS reuse path: 52 -> 6.4 ms/lc with OPENBLAS_NUM_THREADS=1 (throttle events +5 -> 0). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- ncu blocked (ERR_NVGPUCTRPERM, host driver restriction; attempts documented) -> nsys + event decomposition + one-axis sweeps. - Finding 0: OpenBLAS threadpool vs cgroup CPU quota freezes the host ~90ms/100ms period (TESS reuse 52 -> 6.4 ms/lc with pinned pools). - Kernel attribution: ZTF ~85% per-freq fixed cost; HAT ~55-60% histogram atomics; TESS ~95% histogram, CONFLICT-bound (shuffle test: 3.09x); Kepler ~92% histogram. - Ranked plan: (1) fused-noverlap kernel, (2) host-side scatter permutation, (3) host overhead (BLAS-free per-LC path + plan reuse), (4) ZTF fixed-cost via batch/fused. Skipped (b) L2-resident loads, (c) bin sizing noise on A5000, (e) deferred <=8%. Raw JSON + parity dumps + logs under benchmarks/results/bls_survey_speed_jul2026/raw/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One launch histograms at noverlap-times finer phase resolution and derives every dphi-shifted pass's box sums from runs of fine bins (full_bls_no_sol_fused in bls_common.cuh, full_bls_batch_fused in bls_batch.cu). Routed for power-of-two noverlap with dphi == 0 -- where fine-bin assignment is bit-identical to the multi-pass loop -- and when the finer histogram fits in shared memory; every other case keeps the host-side multi-pass loop. Effects: noverlap-x fewer shared-memory atomics + folds, per-freq fixed costs paid once, and the finer histogram halves atomic conflicts as a side effect (fused beats even the single-pass time on HAT-Net/TESS/Kepler). Kernel-only, warm medians, RTX A5000, Keplerian grids (vs bench_base_envfix): ZTF (150 x60K): 5.63 -> 3.15 ms/lc (1.79x) HAT-Net (6K x301K): 76.26 -> 36.44 ms/lc (2.09x) TESS (20K x1.8K): 4.49 -> 1.66 ms/lc (2.70x) Kepler (65K x131K): 367.4 -> 172.4 ms/lc (2.13x) Gates: cuvarbase/tests/test_bls.py 438/438 on pod (incl. 7 new fused tests: manual-pass parity for nov=2/4 on both kernel variants, dphi!=0 fallback, BJD-scale, batch); parity vs base_envfix corr=1.0000000, identical peaks, max|d| <= 5.6e-6 (float32 accumulation-order level) on fast/fast_bjd/batch x 4 surveys. Full-suite + release-gate run records land with the next commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Time-sorted dense-cadence input folds warp-adjacent samples into the same phase bin at nearly every trial frequency, serializing the shared-memory atomics (profile: TESS bin-scale sweep INVERSE, random shuffle 3.09x). Store (t, yw, w) in the staging buffers in a deterministic golden-ratio-stride order instead (utils.conflict_scatter_perm, applied in BLSMemory.setdata and BLSBatchMemory.set_lightcurve; no-op for n < 64). Histogram binning is a sum, so data order is semantically free -- box sums change only at the float32 accumulation-order level that shared atomics already leave undefined. Kernel-only, warm medians, RTX A5000 (vs opt1_fused): ZTF : 3.15 -> 3.33 ms/lc (-5%, within session noise) HAT-Net : 36.44 -> 35.34 ms/lc (1.03x) TESS : 1.66 -> 0.56 ms/lc (3.0x; 8.0x vs pre-fusion baseline) Kepler : 172.4 -> 155.4 ms/lc (1.11x) Gates: test_bls.py + test_utils.py 444/444 on pod; parity vs base_envfix corr=1.0000000, identical peaks, max|d| <= 5.6e-6 across fast/fast_bjd/batch x 4 surveys. Full suite + release gate run in progress, recorded in the next commit (Opt 1's run: 759 passed / 7 skipped, gate 14/14). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(i) np.dot -> np.einsum in the per-LC host path (BLSMemory.setdata,
_chi2_null, BLSBatchMemory.set_lightcurve): BLAS ddot spawns an
nproc threadpool; on CPU-quota-limited containers the burst trips
CFS throttling and froze the process ~90 ms per 100 ms period
(8x end-to-end at TESS scale under default env). einsum stays in
numpy core, single-threaded, no env vars needed.
(ii) builtin max() -> int(np.max()) for the per-launch bin-count
scan in _eebls_gpu_fast_impl: Python-level iteration over the
nbinsf array cost ~2.4 ms (ZTF, 60K freqs) to ~9 ms (HAT-Net, 301K)
PER CALL, hidden inside even kernel-only timings.
(iii) eebls_gpu_batch: one BLSBatchMemory + stream + set_freqs
hoisted out of the chunk loop, prefix-only transfers
(n_lcs_active), frequency grid uploaded once -- and a new
memory= kwarg so survey drivers reuse the staging/device
buffers across calls (per-call pinned allocation cost 0.3-23 ms).
Warm medians, RTX A5000 (vs opt2_scatter):
kernel fast_reuse best batch path
ZTF : 3.33 -> 0.62 3.67 -> 0.86 1.22 -> 0.76 (reuse)
HAT-Net : 35.34 -> 26.32 35.94 -> 26.67 31.72 -> 27.00 (reuse)
TESS : 0.56 -> 0.48 1.25 -> 1.15 1.05 -> 0.85 (reuse)
Kepler : 155.4 -> 151.5 158.9 -> 153.3 157.8 -> 156.4 (reuse)
Gates: test_bls.py + test_utils.py 447/447 on pod (3 new
batch-memory-reuse tests incl. chunked-vs-single-chunk and
too-small-memory ValueError); parity vs base_envfix
corr=1.0000000, identical peaks, all 12 arrays. Full suite +
release gate in flight; Opt 2's full run: 761 passed / 7 skipped,
gate 14/14.
FLAG (documented, not silent): einsum vs BLAS ddot changes the
float64 summation order of ybar/yy/chi2_0, so normalizations move
at the last-ulp level (~1e-16 relative); periodogram parity is
corr=1.0000000 with identical peaks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Launches size their shared memory by the max bin count of the frequencies they cover, and ascending Keplerian grids have monotonically decreasing bin counts -- but a single whole-grid launch pays the GLOBAL max everywhere. On Kepler-scale grids (nbf up to 1665; fused request 27.7 KB) that caps residency at 3 blocks/SM vs the 6-thread-limit, while only 3.4% of frequencies actually need big histograms. When shared memory is the occupancy limiter (_shmem_limits_occupancy: blocks-by-shmem < blocks-by-threads) and the caller left freq_batch_size=None, both fast and batch paths now launch in 8192-frequency chunks with per-chunk shared sizing (measured 151.2 -> 114.8 ms sweep; chunk 16384 within 2%). The batch kernels gain an explicit bls_stride (output row pitch) argument so chunked launches -- and reused memories allocated for more frequencies than a call uses -- index the padded output correctly; get_results(nfreq_active=...) trims stale row tails. Warm medians, RTX A5000 (vs opt3_host): Kepler kernel 151.5 -> 115.0 ms/lc; batch 158.0 -> 119.0; batch_reuse 156.4 -> 117.7; fast_naive 158.3 -> 121.9 ZTF / HAT-Net / TESS: unchanged (heuristic correctly dormant; kernel 0.63 / 26.1 / 0.49 ms/lc) Cumulative kernel-only vs env-fixed baseline: ZTF 5.63->0.63 (8.9x), HAT-Net 76.3->26.1 (2.9x), TESS 4.49->0.49 (9.2x), Kepler 367.4->115.0 (3.2x) Gates: test_bls.py 443/443 on pod (2 new: freq-chunked batch parity with odd chunk size, oversized-memory reuse row-pitch); parity vs base_envfix corr=1.0000000, identical peaks, all 12 arrays. Full suite + release gate in flight (Opt 3 run: 764 passed / 7 skipped, gate 14/14). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- benchmarks/results/bls_survey_speed_jul2026/SUMMARY.md: full stage-by-stage numbers, $/lightcurve, correctness-gate table, flagged numerical notes. - CHANGELOG.rst: Unreleased (feature/bls-survey-speed) entry. - Raw JSON + parity dumps for opt2_scatter/opt3_host/opt4_chunk. - Default-env robustness verified post-einsum: TESS reuse loop 3.10 ms/lc with 0 CFS throttle events (was 52 ms/lc, +5 events). - Final gate on opt4: full suite 766 passed / 7 skipped, release gate 14/14, parity corr=1.0000000 identical peaks (12/12 arrays). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…hase 3/4 items in the 1.1 queue Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NxdN3KdoNGAwUBvnx5wnUM
Replace the phase-binned default with GTLS's observation-level templates, sample windows and full refinement. Reuse native scans through CUDA graphs and fuse residual reductions; retain explicit binned and legacy options. Publish the 160-case independent comparison, 24 separately sealed nulls, thin-transit stress diagnostics, exact inputs and executed sources. Preserve the corrected/literal GTLS distinction and shared floating-point limits. Replace older TLS speed headlines with measured single-source and batch results, component timings, cost projections and one README figure. Retain the failed four-worker GTLS warmup and original campaign gate alongside the explicit post hoc assessment of completed configurations. Validation: 265 GPU TLS tests, 1,215 CPU/harness tests and 87 installed-wheel checks passed. Public-only replay reproduces the normalized timing bytes.
Keep the preserved TLS implementation as the default and require an explicit experimental selector for the optimization bundle. Preserve the failed 5,111/5,120 exactness qualification and all unavailable timing panels. Archive the frozen study, follow-up reports, source bindings and validation receipts. Keep evidence bytes unchanged by Git line-ending normalization. The follow-up reports seven strict TLS/GTLS panels and four BLS execution-only panels; the separately corrected installed-wheel gate passed. Validation: 2,091 prior GPU tests passed with one expected failure and no skips; all 14 additional checks and six dependency preflights passed. Final host and operational checks are recorded in the release-preparation commit.
Separate required step and numerical qualifications from successful evidence collection. Persist incident delivery, acknowledge queued follow-ups, recover the existing bounded guard or collector, and check observer heartbeats with an independent login watchdog. Never create rentals or rerun experiments. Document operation and restore requirements; keep local credentials ignored. Validation: all 14 monitoring tests passed, including recovery, deduplication, failed delivery and false-success regressions. Both real chat delivery paths were received and acknowledged; the deployed observer and watchdog are healthy.
Use a distinct release version while retaining the existing June v1.0.0 tag. Update the release notes and changelog, record package/source verification, and run benchmark qualification and monitoring tests in CI. The built wheel and sdist retain 85 GPU-validated package files byte for byte; the only changed package file updates __version__. Strict metadata checks, both installed-package smoke checks, 1,141 host tests and all 211 benchmark and monitoring tests pass. Host-only GPU/dependency skips and the known notebook xfail remain explicit. Docs have only five expected GPU plot warnings. Prepare a local tag and delivery bundle. Publication and remote ref changes are deferred at the owner's request.
Retain the reviewed release versions of README.rst and the LS, PDM and normalization helper conflicts. The current implementation includes the upstream copy-before-normalization fixes, including batched LS delegation, and also supports None uncertainties. The complete merge tree is identical to 3c1b5c8; four normalization tests pass. This merge incorporates the nine master-only commits without changing the validated source.
Linux ps can truncate command output to the terminal width, hiding the service script path and causing recovery to misidentify a running collector. Request unlimited width on Linux and macOS, and exercise the recovery test with a narrow COLUMNS setting. All 211 benchmark and monitoring tests pass locally; scientific source and prepared artifacts remain unchanged.
Keep readable reports, selected figures, concise results and checksum inventories in Git. Preserve all original failures and source identities in immutable R2 archives and complete rollback bundles. Add safe restoration and a CI guard against reintroducing generated evidence. All 86 package files and prepared distribution inputs remain byte-identical.
johnh2o2
force-pushed
the
release/v1.0.1
branch
from
September 28, 2026 17:25
403c75d to
720039c
Compare
This was referenced Sep 28, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR brings the completed 1.x development into
masteras the 1.0.1 release candidate. The preserved TLS baseline remains the default; experimental optimizations are opt-in. Raw benchmark output now lives in verified private R2 archives, while reports, small fixtures, selected figures and archive inventories stay in Git.The source branches
release/v1.0.1andv1.0-fixes, and the unpublished annotatedv1.0.1source tag, identify6b6c48945160ead43ecf9932f303aaec28035d91. Merging this PR, creating a GitHub release, uploading to PyPI and deploying documentation remain deferred.Changes to review
Start with the release notes, benchmark report, and archive access and restoration guide.
Repository cleanup
The PR now adds 76,024 lines across 372 changed files, down from 2,782,171 added lines across 2,478 files. A fresh clone of all ordinary branches and tags has approximately 59 MiB of Git objects, down from 296 MiB (80% reduction).
The owner-authorized history rewrite covers the two development branches and the unpublished
v1.0.1tag.master, earlier version tags (including Junev1.0.0), other feature branches and deployedgh-pagesremain unchanged. Complete original local and remote Git bundles, original ref inventories and all ten evidence archives passed full SHA256 read-back from R2. Restoring all 506 TLS survey archive members also passed individual hash checks. Use a fresh clone or reapply local patches; merging old development history would restore the removed objects.See the cleanup record and original-to-filtered commit map. Original scientific commit IDs remain original and resolve in the preserved bundles. GitHub may retain old PR refs and cached objects independently; the size comparison measures fresh clone contents.
Scientific qualifications retained
The experimental TLS aggregate exactness gate remains failed: 5,111/5,120 comparisons, with nine chi2/SDE discrepancies. The throughput follow-up has seven strictly qualified TLS/GTLS panels, four BLS execution-only panels and five unavailable panels. BLS execution rates retain their failed repeatability qualification. Passing software tests does not change those results. No scientific experiment was repeated for this cleanup.
Validation
See the GPU receipts, package verification, and release preparation record.
Issue disposition for this release
The following fixes are implemented here and will close when this PR merges into the default branch,
master:The following issues are linked as follow-ups and remain open after this merge:
The performance work closing #32 does not establish a best-competitor ranking; that broader coverage remains under #19. The PDM batching and memory-cap APIs previously deferred in #33 are now implemented, but its CPU-comparison benchmark remains. Issue #15 explicitly waits for release publication and deployment of the rebuilt docs, both still deferred. The original scientific failures and unavailable panels remain unchanged.