This repository was archived by the owner on Sep 24, 2026. It is now read-only.
Repository navigation
bench: recapture the laptop-test pipeline on the CSC DE route - #560
Merged
Merged
Conversation
… the CSC route The community capture ran on pyscx 0.18.0, before the benchmark fixtures carried CSC sidecars. Since then the pipeline's DE stage takes the column-major route with the exact-nnz Wilcoxon kernel (`cpu_csc_nnz`), so every pyscx number in pipeline_ooc_constrained, and every floor calibrated from them, described a path that no longer runs. Re-run the eight pyscx cells and the five scanpy cells that complete, one job at a time on one partition, so each SCX-vs-scanpy ratio comes from a single capture. The three scanpy OOM cells keep their results; nothing in them depends on pyscx. census_1m under the 16 GB ceiling: 3,186 s / 8,819 MB -> 1,227 s / 4,332 MB, with DE 2,022 s -> 88 s. census_500k: 1,494 s -> 548 s, now 4.6x faster than scanpy at 32 GB on 5.1x less memory (was 1.72x / 2.8x). pbmc10k has no sidecar (it is below the auto rule's 50,000 cells), so its DE still loses to scanpy 6x, and the docs say why. Floors: total_wall_s and pipeline_peak_rss_mb recalibrated at 1.4-1.5x the new medians. The old census_1m total_wall_s bound (5,500 s) would have passed a full fallback to cpu_csr. Two new stage_wall_s__rank_genes_groups floors cover a fallback from nnz to the densify kernel, which the route floor reads as 1.0 and the pipeline total absorbs inside node noise: checked against PR-F's densify results (DE 114 s / 228 s), which fail them while passing total_wall_s. accel_qc_filter, accel_score_genes and multimodal_atlas_streaming are not re-run. Their paths do not read a sidecar, and census_1m check runs measured 2,197 MB / 805 MB against the baseline's 2,196 / 821 MB, still bit-exact against scanpy. The atlases carry no sidecar. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
There was a problem hiding this comment.
Code Review
This pull request updates the benchmark results and documentation to reflect the performance improvements from the new cpu_csc_nnz kernel and the use of CSC sidecars. It includes updates to several JSON benchmark files, the thresholds.yaml configuration, and the performance.md and sharding.md documentation. The reviewer correctly identified that the JSON field name should be de_route rather than route in the documentation.
|
|
||
| DE used to be 63 % of this pipeline (2,022 s of 3,186 s on pyscx 0.18.0). | ||
| It is 88 s now because the DE stage takes the column-major route, recorded | ||
| as `route="cpu_csc_nnz"`: the census fixtures carry a CSC sidecar (the |
nick-youngblut
added a commit
that referenced
this pull request
Sep 24, 2026
All 16 versioned workspace members, `pyscx/pyproject.toml` and `rscx/DESCRIPTION`; `tests/scx-integration-tests` stays at `0.0.0`. README's `scx-cli` download snippet follows. (ROADMAP.md no longer exists, so there is no date stamp to move.) 0.20.0 collects the merged PRs since 0.19.0 (#544-#561): - CSC sidecar series: one-pass bucketed CSC builder (#555); `build-csc` as an in-place, rollback-able append (#556); same-pass CSC build on ingest and carry-through on rewrite ops (#557); parallel CSC build and `csc="auto"` as the ingest default (#558); CSC dispatch for row-indexed transforms and row-filtered / gene-subset handles (#553, #554); CSC route coverage and the exact-nnz Wilcoxon kernel as the 1-vs-rest CSC default (#559) - Correctness: Leiden parallel local moving keeps decliners eligible (#544); five CPU-accelerator parity / convergence fixes (#548); GPU cuVS stream sync, reduction determinism, decoder bounds and VRAM accounting (#549); CLI destination-overwrite safety and MTX multimodal / streaming / integer integrity (#550); unscoped whole-matrix reads of a multimodal file refused (#551); dictionary (categorical) output from append / merge / merge --sort-by (#546) and `scx sort`'s obs spill path (#547) - Benchmarks and docs: four community analytical benchmarks (#552), laptop-test recapture (#560), README refresh (#545), docs split into per-topic directories (#561) Behaviour changes worth calling out in the release notes: - **`csc="auto"` is the ingest default** on every entry point and preset (`from_mudata` keeps `off`): files with `n_obs >= 50000` and `n_vars >= 5000` now get a CSC sidecar, costing ingest wall time and +42-71 % on disk. `--csc off` / `csc="off"` opts out (#558). - **Rewrite ops carry the CSC sidecar by default** (`compact`, `merge`, `optimize`, `sort`, `subset`), rebuilt from the output's own shards; `--csc carry|always|off`. `--rebuild-csc` / `rebuild_csc=True` are deprecated aliases for `always` (#557). - **`scx build-csc` appends in place** rather than rewriting the file, and `scx rollback` removes the sidecar (#556). - **Ten CLI subcommands no longer silently overwrite an existing destination**; `--force` is required (convert, merge, subset, query --output, upgrade, cloud-optimize, explode, pack, pull; compact / optimize / sort migrated onto the same guard) and is refused where nothing is written (#550). - **MTX**: a declared-`integer` MTX with a value past 2^24, a non-integral value or `nan`/`inf` is refused on ingest (`--allow-lossy` restores the old behaviour); multimodal MTX export requires a modality (`pyscx.to_mtx` gains `modality=`); export streams one shard at a time (#550). - **An unscoped whole-matrix read of a multimodal file raises** `MultimodalRequiresModality` instead of folding every modality into one `n_obs x n_modalities`-row answer — `to_anndata()`, `to_memory()`, `read_all_csr_shards*`, `scx pull --filter` (#551). - **Numerical changes**: UMAP init scale and `random_init` now match umap-learn, the kNN sigma search, and NB-GLM dispersion shrinkage (the `above_min_disp` residual filter always runs; new `disp_outlier_sd=2.0` carve-out, `None` disables only the carve-out) (#548); parallel Leiden labels change (#544). - **Wilcoxon on the CPU CSC route** defaults to the exact-nnz kernel (route id `cpu_csc_nnz`); `SCX_ACCEL_WILCOXON_NNZ=0` restores densify (`cpu_csc`), and `reference=` / `rankby_abs=True` still use densify (#559). - **Categorical obs columns** keep their declared categories, unused levels and `ordered` bit through append / merge / sort (#546, #547). No benchmark recapture for this release. Pre-release gate on this tree: `cargo fmt --check`; `cargo clippy --workspace --exclude rscx --all-targets -D warnings`; `cargo test --workspace --exclude rscx` (4,220 passed, 0 failed); `cargo test -p scx-convert --features hdf5` (474 passed); `cargo test -p scx-cli --features hdf5 -- --test-threads=1` (253 passed); `maturin develop --release` + `pytest tests/` (3,247 passed, 140 skipped, 2 xfailed; `pyscx.__version__ == "0.20.0"`). Not run: rscx tests, `--features cloud`, GPU pytest, fuzzing, benchmarks. Claude-Session: https://claude.ai/code/session_01RwRKstgr4TNkXhLhkQwmjA Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
pipeline_ooc_constrained's numbers and floors were captured on pyscx 0.18.0, before the benchmark fixtures carried CSC sidecars. Its DE stage now takescpu_csc_nnzinstead ofcpu_csr, so the pipeline's published numbers and floors described a path that no longer runs. This PR re-captures the benchmark, recalibrates its floors, and rewrites the doc section.Changes
Recapture (13 raw results).
cpu_csr, no sidecar)Floors (
thresholds.yaml).total_wall_sandpipeline_peak_rss_mbare recalibrated at 1.4–1.5× the new medians. The old census_1m wall bound (5,500 s) would have passed a full fallback tocpu_csr.stage_wall_s__rank_genes_groupsfloors at census_500k (85 s) and census_1m (140 s) cover a fallback from the exact-nnz kernel to the densify kernel.de_route_is_cscreads 1.0 on both kernels, and the pipeline total absorbs that fallback within node noise.total_wall_s.Docs.
docs/performance.md, § The laptop test: new table and stage breakdown. DE is now 7 % of the census_1m run instead of 63 %.autorule's 50,000 cells and still loses DE about 6×.docs/sharding.mdgains the nnz figure beside the densify one.Notes
library_versions.pyscx = 0.18.0. That is the distribution metadatascx-benchhas installed. Each job asserted the repository's release build by path and build profile before timing anything.v0.18.0-community-benchmarksstill holds the pre-sidecar pipeline rows, andLATESTis untouched.🤖 Generated with Claude Code
https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84