Skip to content

Prepare cuvarbase 1.0.1 with verified external benchmark archives - #69

Open
johnh2o2 wants to merge 539 commits into
masterfrom
release/v1.0.1
Open

johnh2o2 wants to merge 539 commits into
masterfrom
release/v1.0.1

Conversation

@johnh2o2

@johnh2o2 johnh2o2 commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

This PR brings the completed 1.x development into master as the 1.0.1 release candidate. The preserved TLS baseline remains the default; experimental optimizations are opt-in. Raw benchmark output now lives in verified private R2 archives, while reports, small fixtures, selected figures and archive inventories stay in Git.

The source branches release/v1.0.1 and v1.0-fixes, and the unpublished annotated v1.0.1 source tag, identify 6b6c48945160ead43ecf9932f303aaec28035d91. Merging this PR, creating a GitHub release, uploading to PyPI and deploying documentation remain deferred.

Changes to review

  • Observation-level GPU Transit Least Squares; BJD-safe, sparse and batched BLS; Keplerian frequency grids; multiharmonic Lomb–Scargle; and PDM/CE improvements, with input validation and migration guidance.
  • Python 3.9–3.14 packaging and CPU CI, wheel/sdist installation checks, documentation checks, and persistent job monitoring with durable incident delivery and recovery.
  • Benchmark reporting that preserves failed scientific qualifications separately from software validation, collection and shutdown.
  • Ten complete evidence archives with per-file SHA256 inventories, a verified restore helper, and a CI check that rejects raw evidence or files larger than 1 MiB in Git.

Start with the release notes, benchmark report, and archive access and restoration guide.

Repository cleanup

The PR now adds 76,024 lines across 372 changed files, down from 2,782,171 added lines across 2,478 files. A fresh clone of all ordinary branches and tags has approximately 59 MiB of Git objects, down from 296 MiB (80% reduction).

The owner-authorized history rewrite covers the two development branches and the unpublished v1.0.1 tag. master, earlier version tags (including June v1.0.0), other feature branches and deployed gh-pages remain unchanged. Complete original local and remote Git bundles, original ref inventories and all ten evidence archives passed full SHA256 read-back from R2. Restoring all 506 TLS survey archive members also passed individual hash checks. Use a fresh clone or reapply local patches; merging old development history would restore the removed objects.

See the cleanup record and original-to-filtered commit map. Original scientific commit IDs remain original and resolve in the preserved bundles. GitHub may retain old PR refs and cached objects independently; the size comparison measures fresh clone contents.

Scientific qualifications retained

The experimental TLS aggregate exactness gate remains failed: 5,111/5,120 comparisons, with nine chi2/SDE discrepancies. The throughput follow-up has seven strictly qualified TLS/GTLS panels, four BLS execution-only panels and five unavailable panels. BLS execution rates retain their failed repeatability qualification. Passing software tests does not change those results. No scientific experiment was repeated for this cleanup.

Validation

  • Full A40 suite on the frozen 1.0.0 candidate: 2,091 passed, one expected notebook failure, zero skipped. The separate installed-wheel gate passed 14 numerical/runtime checks and six dependency preflights; its original launcher failure is retained.
  • The prepared 1.0.1 wheel has 85 package files identical to that GPU candidate and only the exact version-string change in the remaining file. All 86 package files, 93 tracked sdist inputs, and both prepared artifacts remain byte-identical through this cleanup.
  • Local CPU suite: 1,141 passed, 866 GPU/optional-dependency skips, one expected failure. Benchmark, monitor and archive tooling: 222 passed. All 590 relative Markdown document links resolve; fatal lint and the repository artifact check pass.
  • All 40 GitHub Actions checks passed on the rewritten source: ten jobs each for the PR, release branch, development branch and tag. Each run covers Python 3.9–3.14, release tools, package smoke, docs and lint.
  • Previous Linux monitoring repair retains its failed CI receipt and now reads complete process commands under narrow display widths. The historical GPU, installed-artifact, metadata and docs validation receipts remain preserved.
  • Source and evidence backups are verified in private R2. The GPU rentals remain terminated.

See the GPU receipts, package verification, and release preparation record.

Issue disposition for this release

The following fixes are implemented here and will close when this PR merges into the default branch, master:

The following issues are linked as follow-ups and remain open after this merge:

Issue Remaining requirement
#15 PDM source documentation is complete; release publication and the docs-site rebuild remain.
#19 Transit/BLS/TLS benchmark plots exist; the broader LS/CE/PDM comparison remains incomplete.
#28 The refactor umbrella still has unfinished documentation, naming and PDM follow-ups.
#29 Documentation improved substantially; the full API-docstring/notebook follow-up remains.
#30 Contributor standards and modernization landed; public API renaming remains deferred.
#33 GPU correctness, batching and memory-capped runs are delivered; CPU-comparison benchmarks remain.

The performance work closing #32 does not establish a best-competitor ranking; that broader coverage remains under #19. The PDM batching and memory-cap APIs previously deferred in #33 are now implemented, but its CPU-comparison benchmark remains. Issue #15 explicitly waits for release publication and deployment of the rebuilt docs, both still deferred. The original scientific failures and unavailable panels remain unchanged.

astrobatty and others added 30 commits July 3, 2026 16:14
…riginal-timescale phi convention

- data() now applies t0=4.5 before the model is evaluated, so every
  test exercises a non-trivial epoch (floor(min(t)) = 4 or 5) and the
  injected transit stays at original-timescale phase phi0 (attila's
  suggested reproduction, applied as a permanent strengthening).
- test_ignore_positive_sols: drop the old-convention manual phi shift
  (single_bls converts internally now); update the deterministic
  hardcoded value for the rotated fold.
- test_single_bls_bjd_invariance: shifted runs use covariantly shifted
  phases (phi0 + offset*freq) mod 1 per the new convention.
- new test_single_bls_phase_is_original_timescale: unshifted phi0 on
  shifted times must MISS the transit (guards the convention).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…airness rules)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…precision, sparse f64 phase conversion

Root cause of the test_standard failures attila reported with non-zero
T0 (61/216 on RTX A5000): NOT the new epoch re-referencing (the
float64 phi round-trip is bit-exact at float32 precision), but a
pre-existing precision asymmetry in single_bls exposed by the shifted
fixture. single_bls subtracted phi0 from the UNWRAPPED float32 product
t*f (magnitude ~360 for a 1-yr baseline, ulp 3.05e-5) while the GPU
kernels wrap into [0,1) first (ulp ~1e-7): a point whose true phase
sat 8.3e-6 below the best box edge rounded to phase exactly 0.0 on the
CPU side and flipped into the box, a ~power/n_in_transit (=0.036)
disagreement at one frequency, which trips the zero-violation
mostly_ok criterion. Hardware probe confirms nvcc does NOT
FMA-contract mod1(t*f), so wrapping first makes the reference fold
bit-identical to the kernels'.

- single_bls: wrap phase into [0,1) before subtracting phi0
  (documented; shrinks the CPU-vs-GPU edge-disagreement window ~150x,
  and up to ~2000x for 10-yr baselines).
- bin_and_phase_fold_custom: fold with the float32-cast frequency
  (folding with the double freq shifted phases by up to |f64-f32|*t
  ~1e-5 vs the reference); keep double freqs ONLY for the epoch
  re-referencing. phi_values now uploaded as float64 so the in-kernel
  (phi - epoch*freq) % 1 matches single_bls's float64 conversion bit
  for bit (store_best_sols_custom signature updated accordingly).
- sparse_bls_cpu / sparse_bls_gpu: convert solutions to the original
  timescale with the caller's float64 frequencies, not the
  float32-cast copies (error epoch*|f64-f32| reaches ~0.07 cycles at
  BJD-scale epochs).
- eebls_transit_gpu: uniform 3-tuple return (sols=None on the
  fast/optimized paths), matching eebls_transit.
- eebls_transit(use_optimized=True): respect an explicit block_size
  instead of silently overriding it.
- tests: test_transit uses a single 3-way mode axis
  (standard/fast/optimized) instead of a use_fast x use_optimized
  cross-product; test_standard/test_custom drop the use_optimized
  axis (only reduction_max differs) in favor of focused equivalence
  tests; new TestHoneSolution exercises hone_solution end-to-end at
  maximal epoch-phase rotation ((epoch*freq) % 1 = 0.5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… decomposition

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…00, Jul 2026)

Headline: warm kernel parity (9.8ms vs 9.8ms kernel-only); 34x per-LC on the
naive loop path (0.2.6 recompiles per call); BJD-timestamp demo (0.2.6 fails,
v1.0 recovers); 0.2.6 LS segfaults on modern pycuda; 0.2.6 never shipped to
PyPI (last PyPI release is 0.2.5, Oct 2023).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… + pod validation), docstring touch-up

- CHANGELOG.rst: original-input-timescale phi convention (PR #65) and
  the single_bls fold-order precision fix + custom-kernel/sparse
  conversion parity fixes.
- analysis/pr65-resolution-jul2026/: SUMMARY.md (mechanism, evidence,
  draft reply to attila) + raw logs (61/216 reproduction, per-frequency
  diagnostics, flip-point drill-down, FMA hardware probe, full-suite
  752/752-passed log, release-gate ALL-PASSED log).
- eebls_transit use_optimized docstring: block-size auto-selection is
  conditional on no explicit block_size (matches the code fix).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fix BLS bugs and return correct phi values after time stamp normalization
- CHANGELOG: correct release lineage (0.2.5 is the last PyPI release; the
  0.2.6 tag was never published), add the measured v0.2.6 head-to-head
  summary, strip internal tracker IDs and analysis/ paths from user-facing
  entries, add a 0.2.6 stub entry
- BENCHMARK_RESULTS: retract the compile-overhead-confounded 'vs pre-v1.0
  kernel' column, add the measured July 2026 head-to-head table, correct
  the stale batch-mode guidance (batch now wins at every measured scale)
- New docs/RELEASE_NOTES_v1.0.0.md: outward-facing release notes vs 0.2.5
- INSTALL.rst rewritten for Python 3.9+/CUDA 11-12 (was CUDA 8 + Python 2.7)
- MANIFEST.in so CHANGELOG/INSTALL ship in the sdist; README changelog link
  absolutized for PyPI rendering; Development Status classifier -> Stable

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
bench_bls_survey.py: 4 survey configs (ZTF/HAT-Net/TESS/Kepler) on
Keplerian grids (0.5q..2q), variants fast_naive/fast_reuse/kernel/
kernel_1pass/pieces/batch, warm medians >=5 runs, cold reported
separately, $/lightcurve from pod rate, parity dumps for
before/after gates.

profile_bls_survey.py: minimal kernel-only driver for ncu/nsys.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… tracing)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…gnosis

Closes the item deferred from commit 14b0d90 (cunfft.cu A3 fix): the
same float32 PI literal in lomb.cu, tls.cu and nufft_lrt.cu, each
assessed for actual exposure and given the same treatment (PI double
literal under DOUBLE_PRECISION, float32 literal otherwise).

lomb.cu (live defect, rigorous gate): PI feeds the un-reduced phase
2*pi*f*(t+0.5) in cossum/sinsum, used by the direct-sums kernels
(use_fft=False; the production NFFT path never touches PI). In double
mode the float32 pi (rel. err. +2.7828e-8) makes the kernel evaluate
the exact periodogram on a frequency axis stretched by 1+2.78e-8 -
observable error ~ eps*f*T of the local slope. Measured on an A5000 at
f=90-110 c/d, T=30 d vs a float64 CPU port of the kernel (exact pi,
same op order): before 1.162e-4 max abs power error in every f64 case;
after 3.67e-10 (nonzero epoch, both low-level and run() paths; 316,517x
closer) and 1.22e-8 (raw BJD 2.45e6 epoch through the low-level API;
9,505x). The buggy kernel matches the same port evaluated with
float64(float32(pi)) to 3.7e-10 - the diagnosis exactly. float32-mode
periodograms are bit-identical (the literal's value is unchanged). New
regression test test_ls_kernel_direct_sums_double_pi (threshold 1e-7;
measured after 1.1e-10, buggy 1.2e-4).

nufft_lrt.cu (PI dead, double mode real): literal moved under the
DOUBLE_PRECISION guard; hardcoded float32 helpers on FLT operands
(fmaxf/fmodf/fabsf) retyped to the overloaded forms - the same class
14b0d90 fixed for modflt/diffmod. Double-mode matched filter now
matches a float64 numpy port exactly (was rel. err. 3.5e-9 from the
fmaxf truncation); float32 mode bit-identical.

tls.cu (PI dead, kernel float32-only by design): dead macro removed;
TLS search output bit-identical before/after (period/depth/SDE and
full chi2/SR/power arrays).

Validation (RTX A5000, CUDA 12.4): full GPU suite 753 passed /
7 skipped (752 baseline at ff326c2 + the new regression test);
scripts/check_release_gate.py 14/14 PASS. Evidence + raw before/after
runs archived in analysis/kernel-hygiene-jul2026/.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BLS is a weighted fit (w_i = dy_i^-2 / sum dy_j^-2): a point with a
near-zero reported uncertainty concentrates essentially all statistical
weight in one phase box, so the best box absorbs ~all the weighted
variance and the power P = 1 - chi2/chi2_0 saturates near 1 (~0.99
after binning) deterministically in pure noise, at nearly every trial
frequency. Documented where BLS users will see it: a new "Data
hygiene: near-zero uncertainties" section in docs/source/bls.rst
(mechanism, symptoms - one point dominating sum(1/dy^2), suspiciously
flat ~0.99 power on noise - and a copy-pasteable percentile error-floor
+ weight-check guard) and matching .. warning:: blocks in the
eebls_gpu_fast / eebls_gpu / eebls_transit_gpu docstrings. Docs only -
no runtime warnings or behavior changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause found via nsys + cgroup counters: OpenBLAS spawns
nproc(=96) threads inside per-LC np.dot calls; on a 7.65-CPU-quota
RunPod container the CFS quota freezes the process ~90 ms per 100 ms
period (GPU idle). TESS reuse path: 52 -> 6.4 ms/lc with
OPENBLAS_NUM_THREADS=1 (throttle events +5 -> 0).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- ncu blocked (ERR_NVGPUCTRPERM, host driver restriction; attempts
  documented) -> nsys + event decomposition + one-axis sweeps.
- Finding 0: OpenBLAS threadpool vs cgroup CPU quota freezes the host
  ~90ms/100ms period (TESS reuse 52 -> 6.4 ms/lc with pinned pools).
- Kernel attribution: ZTF ~85% per-freq fixed cost; HAT ~55-60%
  histogram atomics; TESS ~95% histogram, CONFLICT-bound (shuffle
  test: 3.09x); Kepler ~92% histogram.
- Ranked plan: (1) fused-noverlap kernel, (2) host-side scatter
  permutation, (3) host overhead (BLAS-free per-LC path + plan
  reuse), (4) ZTF fixed-cost via batch/fused. Skipped (b) L2-resident
  loads, (c) bin sizing noise on A5000, (e) deferred <=8%.

Raw JSON + parity dumps + logs under
benchmarks/results/bls_survey_speed_jul2026/raw/.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One launch histograms at noverlap-times finer phase resolution and
derives every dphi-shifted pass's box sums from runs of fine bins
(full_bls_no_sol_fused in bls_common.cuh, full_bls_batch_fused in
bls_batch.cu). Routed for power-of-two noverlap with dphi == 0 --
where fine-bin assignment is bit-identical to the multi-pass loop --
and when the finer histogram fits in shared memory; every other case
keeps the host-side multi-pass loop.

Effects: noverlap-x fewer shared-memory atomics + folds, per-freq
fixed costs paid once, and the finer histogram halves atomic
conflicts as a side effect (fused beats even the single-pass time on
HAT-Net/TESS/Kepler).

Kernel-only, warm medians, RTX A5000, Keplerian grids (vs
bench_base_envfix):
  ZTF     (150 x60K):  5.63 ->  3.15 ms/lc (1.79x)
  HAT-Net (6K x301K): 76.26 -> 36.44 ms/lc (2.09x)
  TESS    (20K x1.8K): 4.49 ->  1.66 ms/lc (2.70x)
  Kepler  (65K x131K): 367.4 -> 172.4 ms/lc (2.13x)

Gates: cuvarbase/tests/test_bls.py 438/438 on pod (incl. 7 new fused
tests: manual-pass parity for nov=2/4 on both kernel variants,
dphi!=0 fallback, BJD-scale, batch); parity vs base_envfix
corr=1.0000000, identical peaks, max|d| <= 5.6e-6 (float32
accumulation-order level) on fast/fast_bjd/batch x 4 surveys.
Full-suite + release-gate run records land with the next commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Time-sorted dense-cadence input folds warp-adjacent samples into the
same phase bin at nearly every trial frequency, serializing the
shared-memory atomics (profile: TESS bin-scale sweep INVERSE, random
shuffle 3.09x). Store (t, yw, w) in the staging buffers in a
deterministic golden-ratio-stride order instead
(utils.conflict_scatter_perm, applied in BLSMemory.setdata and
BLSBatchMemory.set_lightcurve; no-op for n < 64). Histogram binning
is a sum, so data order is semantically free -- box sums change only
at the float32 accumulation-order level that shared atomics already
leave undefined.

Kernel-only, warm medians, RTX A5000 (vs opt1_fused):
  ZTF     :  3.15 ->  3.33 ms/lc (-5%, within session noise)
  HAT-Net : 36.44 -> 35.34 ms/lc (1.03x)
  TESS    :  1.66 ->  0.56 ms/lc (3.0x; 8.0x vs pre-fusion baseline)
  Kepler  : 172.4 -> 155.4 ms/lc (1.11x)

Gates: test_bls.py + test_utils.py 444/444 on pod; parity vs
base_envfix corr=1.0000000, identical peaks, max|d| <= 5.6e-6
across fast/fast_bjd/batch x 4 surveys. Full suite + release gate
run in progress, recorded in the next commit (Opt 1's run: 759
passed / 7 skipped, gate 14/14).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(i) np.dot -> np.einsum in the per-LC host path (BLSMemory.setdata,
_chi2_null, BLSBatchMemory.set_lightcurve): BLAS ddot spawns an
nproc threadpool; on CPU-quota-limited containers the burst trips
CFS throttling and froze the process ~90 ms per 100 ms period
(8x end-to-end at TESS scale under default env). einsum stays in
numpy core, single-threaded, no env vars needed.

(ii) builtin max() -> int(np.max()) for the per-launch bin-count
scan in _eebls_gpu_fast_impl: Python-level iteration over the
nbinsf array cost ~2.4 ms (ZTF, 60K freqs) to ~9 ms (HAT-Net, 301K)
PER CALL, hidden inside even kernel-only timings.

(iii) eebls_gpu_batch: one BLSBatchMemory + stream + set_freqs
hoisted out of the chunk loop, prefix-only transfers
(n_lcs_active), frequency grid uploaded once -- and a new
memory= kwarg so survey drivers reuse the staging/device
buffers across calls (per-call pinned allocation cost 0.3-23 ms).

Warm medians, RTX A5000 (vs opt2_scatter):
              kernel           fast_reuse       best batch path
  ZTF     :  3.33 -> 0.62     3.67 -> 0.86     1.22 -> 0.76 (reuse)
  HAT-Net : 35.34 -> 26.32   35.94 -> 26.67   31.72 -> 27.00 (reuse)
  TESS    :  0.56 -> 0.48     1.25 -> 1.15     1.05 -> 0.85 (reuse)
  Kepler  : 155.4 -> 151.5   158.9 -> 153.3   157.8 -> 156.4 (reuse)

Gates: test_bls.py + test_utils.py 447/447 on pod (3 new
batch-memory-reuse tests incl. chunked-vs-single-chunk and
too-small-memory ValueError); parity vs base_envfix
corr=1.0000000, identical peaks, all 12 arrays. Full suite +
release gate in flight; Opt 2's full run: 761 passed / 7 skipped,
gate 14/14.

FLAG (documented, not silent): einsum vs BLAS ddot changes the
float64 summation order of ybar/yy/chi2_0, so normalizations move
at the last-ulp level (~1e-16 relative); periodogram parity is
corr=1.0000000 with identical peaks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Launches size their shared memory by the max bin count of the
frequencies they cover, and ascending Keplerian grids have
monotonically decreasing bin counts -- but a single whole-grid
launch pays the GLOBAL max everywhere. On Kepler-scale grids
(nbf up to 1665; fused request 27.7 KB) that caps residency at 3
blocks/SM vs the 6-thread-limit, while only 3.4% of frequencies
actually need big histograms.

When shared memory is the occupancy limiter
(_shmem_limits_occupancy: blocks-by-shmem < blocks-by-threads) and
the caller left freq_batch_size=None, both fast and batch paths now
launch in 8192-frequency chunks with per-chunk shared sizing
(measured 151.2 -> 114.8 ms sweep; chunk 16384 within 2%). The
batch kernels gain an explicit bls_stride (output row pitch)
argument so chunked launches -- and reused memories allocated for
more frequencies than a call uses -- index the padded output
correctly; get_results(nfreq_active=...) trims stale row tails.

Warm medians, RTX A5000 (vs opt3_host):
  Kepler kernel 151.5 -> 115.0 ms/lc; batch 158.0 -> 119.0;
  batch_reuse 156.4 -> 117.7; fast_naive 158.3 -> 121.9
  ZTF / HAT-Net / TESS: unchanged (heuristic correctly dormant;
  kernel 0.63 / 26.1 / 0.49 ms/lc)

Cumulative kernel-only vs env-fixed baseline:
  ZTF 5.63->0.63 (8.9x), HAT-Net 76.3->26.1 (2.9x),
  TESS 4.49->0.49 (9.2x), Kepler 367.4->115.0 (3.2x)

Gates: test_bls.py 443/443 on pod (2 new: freq-chunked batch parity
with odd chunk size, oversized-memory reuse row-pitch); parity vs
base_envfix corr=1.0000000, identical peaks, all 12 arrays. Full
suite + release gate in flight (Opt 3 run: 764 passed / 7 skipped,
gate 14/14).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- benchmarks/results/bls_survey_speed_jul2026/SUMMARY.md: full
  stage-by-stage numbers, $/lightcurve, correctness-gate table,
  flagged numerical notes.
- CHANGELOG.rst: Unreleased (feature/bls-survey-speed) entry.
- Raw JSON + parity dumps for opt2_scatter/opt3_host/opt4_chunk.
- Default-env robustness verified post-einsum: TESS reuse loop
  3.10 ms/lc with 0 CFS throttle events (was 52 ms/lc, +5 events).
- Final gate on opt4: full suite 766 passed / 7 skipped, release
  gate 14/14, parity corr=1.0000000 identical peaks (12/12 arrays).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
johnh2o2 and others added 17 commits September 6, 2026 15:49
…hase 3/4 items in the 1.1 queue

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NxdN3KdoNGAwUBvnx5wnUM
Replace the phase-binned default with GTLS's observation-level templates,
sample windows and full refinement. Reuse native scans through CUDA graphs
and fuse residual reductions; retain explicit binned and legacy options.

Publish the 160-case independent comparison, 24 separately sealed nulls,
thin-transit stress diagnostics, exact inputs and executed sources. Preserve
the corrected/literal GTLS distinction and shared floating-point limits.

Replace older TLS speed headlines with measured single-source and batch
results, component timings, cost projections and one README figure. Retain
the failed four-worker GTLS warmup and original campaign gate alongside the
explicit post hoc assessment of completed configurations.

Validation: 265 GPU TLS tests, 1,215 CPU/harness tests and 87 installed-wheel
checks passed. Public-only replay reproduces the normalized timing bytes.
Keep the preserved TLS implementation as the default and require an explicit
experimental selector for the optimization bundle. Preserve the failed
5,111/5,120 exactness qualification and all unavailable timing panels.

Archive the frozen study, follow-up reports, source bindings and validation
receipts. Keep evidence bytes unchanged by Git line-ending normalization.
The follow-up reports seven strict TLS/GTLS panels and four BLS execution-only
panels; the separately corrected installed-wheel gate passed.

Validation: 2,091 prior GPU tests passed with one expected failure and no
skips; all 14 additional checks and six dependency preflights passed. Final
host and operational checks are recorded in the release-preparation commit.
Separate required step and numerical qualifications from successful evidence
collection. Persist incident delivery, acknowledge queued follow-ups, recover
the existing bounded guard or collector, and check observer heartbeats with
an independent login watchdog. Never create rentals or rerun experiments.

Document operation and restore requirements; keep local credentials ignored.
Validation: all 14 monitoring tests passed, including recovery, deduplication,
failed delivery and false-success regressions. Both real chat delivery paths
were received and acknowledged; the deployed observer and watchdog are healthy.
Use a distinct release version while retaining the existing June v1.0.0 tag.
Update the release notes and changelog, record package/source verification,
and run benchmark qualification and monitoring tests in CI.

The built wheel and sdist retain 85 GPU-validated package files byte for byte;
the only changed package file updates __version__. Strict metadata checks,
both installed-package smoke checks, 1,141 host tests and all 211 benchmark
and monitoring tests pass. Host-only GPU/dependency skips and the known
notebook xfail remain explicit. Docs have only five expected GPU plot warnings.

Prepare a local tag and delivery bundle. Publication and remote ref changes
are deferred at the owner's request.
Retain the reviewed release versions of README.rst and the LS, PDM and normalization helper conflicts. The current implementation includes the upstream copy-before-normalization fixes, including batched LS delegation, and also supports None uncertainties. The complete merge tree is identical to 3c1b5c8; four normalization tests pass. This merge incorporates the nine master-only commits without changing the validated source.
Linux ps can truncate command output to the terminal width, hiding the service script path and causing recovery to misidentify a running collector. Request unlimited width on Linux and macOS, and exercise the recovery test with a narrow COLUMNS setting. All 211 benchmark and monitoring tests pass locally; scientific source and prepared artifacts remain unchanged.
Keep readable reports, selected figures, concise results and checksum inventories in Git. Preserve all original failures and source identities in immutable R2 archives and complete rollback bundles. Add safe restoration and a CI guard against reintroducing generated evidence. All 86 package files and prepared distribution inputs remain byte-identical.
@johnh2o2 johnh2o2 changed the title Prepare cuvarbase 1.0.1 with validated 1.x features and retained benchmark qualifications Prepare cuvarbase 1.0.1 with verified external benchmark archives Sep 28, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants