Skip to content
This repository was archived by the owner on Sep 24, 2026. It is now read-only.

bench: recapture the laptop-test pipeline on the CSC DE route - #560

Merged
nick-youngblut merged 1 commit into
mainfrom
bench-community-pipeline-csc-recapture
Sep 24, 2026
Merged

nick-youngblut merged 1 commit into
mainfrom
bench-community-pipeline-csc-recapture

Conversation

@nick-youngblut

Copy link
Copy Markdown
Contributor

Summary

pipeline_ooc_constrained's numbers and floors were captured on pyscx 0.18.0, before the benchmark fixtures carried CSC sidecars. Its DE stage now takes cpu_csc_nnz instead of cpu_csr, so the pipeline's published numbers and floors described a path that no longer runs. This PR re-captures the benchmark, recalibrates its floors, and rewrites the doc section.

Changes

  • Recapture (13 raw results).

    • Covers the 8 pyscx cells and the 5 scanpy cells that complete, run one SLURM job at a time so every SCX-vs-scanpy ratio comes from a single capture.
    • The 3 scanpy OOM cells keep their results, since nothing in them depends on pyscx.
    dataset pyscx 16 GB, before → after DE stage scanpy 32 GB
    pbmc10k 61 s → 66 s 46 → 49 s (cpu_csr, no sidecar) 58 s
    tabula_sapiens_100k 267 s / 4,259 MB → 113 s / 2,587 MB 179 → 19 s 247 s / 5,005 MB
    census_500k 1,494 s / 6,336 MB → 548 s / 3,496 MB 992 → 54 s 2,547 s / 17,828 MB
    census_1m 3,186 s / 8,819 MB → 1,227 s / 4,332 MB 2,022 → 88 s OOM
  • Floors (thresholds.yaml).

    • total_wall_s and pipeline_peak_rss_mb are recalibrated at 1.4–1.5× the new medians. The old census_1m wall bound (5,500 s) would have passed a full fallback to cpu_csr.
    • New stage_wall_s__rank_genes_groups floors at census_500k (85 s) and census_1m (140 s) cover a fallback from the exact-nnz kernel to the densify kernel. de_route_is_csc reads 1.0 on both kernels, and the pipeline total absorbs that fallback within node noise.
    • Checked against saved densify results (DE 114 s / 228 s): they fail the new floors while passing total_wall_s.
  • Docs.

    • docs/performance.md, § The laptop test: new table and stage breakdown. DE is now 7 % of the census_1m run instead of 63 %.
    • The "where SCX loses" paragraph is scoped to files without a sidecar: pbmc10k is below the auto rule's 50,000 cells and still loses DE about 6×.
    • Source and manifest notes are updated.
    • docs/sharding.md gains the nnz figure beside the densify one.

Notes

  • The other three community benchmarks are not re-run.
    • Their paths do not read a sidecar, and check runs at census_1m reproduce the baseline's peak RSS (QC 2,197 vs 2,196 MB; score_genes 805 vs 821 MB), still bit-exact against scanpy.
    • The multimodal atlases carry no sidecar.
  • The raw JSONs record library_versions.pyscx = 0.18.0. That is the distribution metadata scx-bench has installed. Each job asserted the repository's release build by path and build profile before timing anything.
  • No baseline is promoted. v0.18.0-community-benchmarks still holds the pre-sidecar pipeline rows, and LATEST is untouched.

🤖 Generated with Claude Code

https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84

… the CSC route

The community capture ran on pyscx 0.18.0, before the benchmark fixtures
carried CSC sidecars. Since then the pipeline's DE stage takes the
column-major route with the exact-nnz Wilcoxon kernel (`cpu_csc_nnz`), so
every pyscx number in pipeline_ooc_constrained, and every floor calibrated
from them, described a path that no longer runs. Re-run the eight pyscx
cells and the five scanpy cells that complete, one job at a time on one
partition, so each SCX-vs-scanpy ratio comes from a single capture. The
three scanpy OOM cells keep their results; nothing in them depends on pyscx.

census_1m under the 16 GB ceiling: 3,186 s / 8,819 MB -> 1,227 s /
4,332 MB, with DE 2,022 s -> 88 s. census_500k: 1,494 s -> 548 s, now
4.6x faster than scanpy at 32 GB on 5.1x less memory (was 1.72x / 2.8x).
pbmc10k has no sidecar (it is below the auto rule's 50,000 cells), so its
DE still loses to scanpy 6x, and the docs say why.

Floors: total_wall_s and pipeline_peak_rss_mb recalibrated at 1.4-1.5x the
new medians. The old census_1m total_wall_s bound (5,500 s) would have
passed a full fallback to cpu_csr. Two new stage_wall_s__rank_genes_groups
floors cover a fallback from nnz to the densify kernel, which the route
floor reads as 1.0 and the pipeline total absorbs inside node noise:
checked against PR-F's densify results (DE 114 s / 228 s), which fail them
while passing total_wall_s.

accel_qc_filter, accel_score_genes and multimodal_atlas_streaming are not
re-run. Their paths do not read a sidecar, and census_1m check runs
measured 2,197 MB / 805 MB against the baseline's 2,196 / 821 MB, still
bit-exact against scanpy. The atlases carry no sidecar.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the benchmark results and documentation to reflect the performance improvements from the new cpu_csc_nnz kernel and the use of CSC sidecars. It includes updates to several JSON benchmark files, the thresholds.yaml configuration, and the performance.md and sharding.md documentation. The reviewer correctly identified that the JSON field name should be de_route rather than route in the documentation.

Comment thread docs/performance.md

DE used to be 63 % of this pipeline (2,022 s of 3,186 s on pyscx 0.18.0).
It is 88 s now because the DE stage takes the column-major route, recorded
as `route="cpu_csc_nnz"`: the census fixtures carry a CSC sidecar (the

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For accuracy, the recorded field in the JSON results is de_route, not route. It would be clearer to refer to it as de_route="cpu_csc_nnz".

@nick-youngblut
nick-youngblut merged commit 27d6347 into main Sep 24, 2026
11 checks passed
@nick-youngblut
nick-youngblut deleted the bench-community-pipeline-csc-recapture branch September 24, 2026 02:46
nick-youngblut added a commit that referenced this pull request Sep 24, 2026
All 16 versioned workspace members, `pyscx/pyproject.toml` and
`rscx/DESCRIPTION`; `tests/scx-integration-tests` stays at `0.0.0`. README's
`scx-cli` download snippet follows. (ROADMAP.md no longer exists, so there is
no date stamp to move.)

0.20.0 collects the merged PRs since 0.19.0 (#544-#561):

- CSC sidecar series: one-pass bucketed CSC builder (#555); `build-csc` as an
  in-place, rollback-able append (#556); same-pass CSC build on ingest and
  carry-through on rewrite ops (#557); parallel CSC build and `csc="auto"` as
  the ingest default (#558); CSC dispatch for row-indexed transforms and
  row-filtered / gene-subset handles (#553, #554); CSC route coverage and the
  exact-nnz Wilcoxon kernel as the 1-vs-rest CSC default (#559)
- Correctness: Leiden parallel local moving keeps decliners eligible (#544);
  five CPU-accelerator parity / convergence fixes (#548); GPU cuVS stream sync,
  reduction determinism, decoder bounds and VRAM accounting (#549); CLI
  destination-overwrite safety and MTX multimodal / streaming / integer
  integrity (#550); unscoped whole-matrix reads of a multimodal file refused
  (#551); dictionary (categorical) output from append / merge / merge
  --sort-by (#546) and `scx sort`'s obs spill path (#547)
- Benchmarks and docs: four community analytical benchmarks (#552), laptop-test
  recapture (#560), README refresh (#545), docs split into per-topic
  directories (#561)

Behaviour changes worth calling out in the release notes:

- **`csc="auto"` is the ingest default** on every entry point and preset
  (`from_mudata` keeps `off`): files with `n_obs >= 50000` and
  `n_vars >= 5000` now get a CSC sidecar, costing ingest wall time and
  +42-71 % on disk. `--csc off` / `csc="off"` opts out (#558).
- **Rewrite ops carry the CSC sidecar by default** (`compact`, `merge`,
  `optimize`, `sort`, `subset`), rebuilt from the output's own shards;
  `--csc carry|always|off`. `--rebuild-csc` / `rebuild_csc=True` are
  deprecated aliases for `always` (#557).
- **`scx build-csc` appends in place** rather than rewriting the file, and
  `scx rollback` removes the sidecar (#556).
- **Ten CLI subcommands no longer silently overwrite an existing destination**;
  `--force` is required (convert, merge, subset, query --output, upgrade,
  cloud-optimize, explode, pack, pull; compact / optimize / sort migrated onto
  the same guard) and is refused where nothing is written (#550).
- **MTX**: a declared-`integer` MTX with a value past 2^24, a non-integral
  value or `nan`/`inf` is refused on ingest (`--allow-lossy` restores the old
  behaviour); multimodal MTX export requires a modality (`pyscx.to_mtx` gains
  `modality=`); export streams one shard at a time (#550).
- **An unscoped whole-matrix read of a multimodal file raises**
  `MultimodalRequiresModality` instead of folding every modality into one
  `n_obs x n_modalities`-row answer — `to_anndata()`, `to_memory()`,
  `read_all_csr_shards*`, `scx pull --filter` (#551).
- **Numerical changes**: UMAP init scale and `random_init` now match
  umap-learn, the kNN sigma search, and NB-GLM dispersion shrinkage (the
  `above_min_disp` residual filter always runs; new `disp_outlier_sd=2.0`
  carve-out, `None` disables only the carve-out) (#548); parallel Leiden
  labels change (#544).
- **Wilcoxon on the CPU CSC route** defaults to the exact-nnz kernel (route id
  `cpu_csc_nnz`); `SCX_ACCEL_WILCOXON_NNZ=0` restores densify (`cpu_csc`), and
  `reference=` / `rankby_abs=True` still use densify (#559).
- **Categorical obs columns** keep their declared categories, unused levels and
  `ordered` bit through append / merge / sort (#546, #547).

No benchmark recapture for this release.

Pre-release gate on this tree: `cargo fmt --check`; `cargo clippy --workspace
--exclude rscx --all-targets -D warnings`; `cargo test --workspace --exclude
rscx` (4,220 passed, 0 failed); `cargo test -p scx-convert --features hdf5`
(474 passed); `cargo test -p scx-cli --features hdf5 -- --test-threads=1` (253
passed); `maturin develop --release` + `pytest tests/` (3,247 passed, 140
skipped, 2 xfailed; `pyscx.__version__ == "0.20.0"`). Not run: rscx tests,
`--features cloud`, GPU pytest, fuzzing, benchmarks.


Claude-Session: https://claude.ai/code/session_01RwRKstgr4TNkXhLhkQwmjA

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant