Repository navigation
sparse+format-io+engine: parallel CSC build, and csc="auto" as the ingest default - #558
Conversation
The same-pass CSC build ran almost entirely on one thread: the writer decoded each pre-encoded X shard serially and pushed it into the builder one nonzero at a time, and the emit drained one bucket and encoded one shard at a time. At census_500k's layout (87 shards of 715 columns, three row groups each) the emit kept about three threads busy. Measured on a 12-core node, convert --csc always went 12.0 -> 6.8 s at tabula_100k and 59 -> 40-44 s at census_500k; the CSC shards are byte-identical. - CscBuilder::push_shard routes a shard with one task per bucket (the `parallel` feature, turned on by scx-format-io's). Each task walks the rows in order and binary-searches its column range, so a bucket gets the same records in the same order and seals at the same records. A task that crosses the spill ceiling spills its own sealed blocks, which keeps the declared spill_after + 2 * n_buckets * block_capacity bound. - Fine buckets: with fewer shards than target_buckets, a shard splits into equal buckets that never straddle a shard boundary, so the push and the per-shard drain both fan out. Shard widths are unchanged. - CscEmitter::next_batch builds several emit groups concurrently. emit_csc_shards value-encodes and codec-encodes a batch in parallel through the new ScxWriter::write_csc_shards, and builds the next batch while the current one encodes. Each batch is half of CSC_EMIT_SHARE at 12 B/nnz, so the two in flight stay within the share. - write_shard_inner is split into a pure encode_shard_section and a sequential commit_shard_section. write_csc_shards runs the same two, which is what keeps a batch-encoded shard byte-identical. - The sink decodes a framed X shard's row groups in parallel. - SpillStore::append_all hands a whole spill to the store, so TempDirSpillStore opens the bucket file once per spill instead of once per 1 MiB block. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
An unset `csc` used to resolve to "off" unless the index preset was `training` or `perturbseq`, so a file got a CSC sidecar only if its author already knew to ask. With the builder in the same pass as X and now parallel, the default becomes "auto" (a sidecar when n_obs >= 50,000 and n_vars >= 5,000) on `scx convert` and every pyscx ingest entry point; `--csc off` / csc="off" is the opt-out. - CscPolicy's Default is Auto (and IngestOptions::default follows it, so library callers change behaviour too). CscPolicy::as_str added. - resolve_csc_policy(csc) takes no preset: with "auto" the default for every preset there is nothing left for one to decide, so preset_implies_csc_auto is gone rather than widened to `cellxgene`. - from_mudata keeps "off": it cannot build per-modality sidecars. - ResidentCscSource (the in-memory from_anndata path, now reached by default) finds each shard's columns by binary search per row on canonical input instead of testing every nonzero once per CSC shard, and fills a batch of shards in parallel. At census_500k's 87 shards the old scan read the matrix 87 times. - pyscx tests: the preset tests now assert that an unset `csc` is auto for every preset, cellxgene included, and that an explicit "off" wins. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
… timed arms
The PR-B gate half, plus what the ingest default flip needs from the suite.
- pipeline_ooc_constrained: the worker records the DE stage's route,
csc_available and fallback_reason plus the fixture's has_csc, and the
parent emits de_route / de_route_is_csc / de_csc_available /
fixture_has_csc. Floored at 1.0 on tabula_sapiens_100k, census_500k and
census_1m (the fixtures that clear the auto rule and so carry a
sidecar); the metric reads 0.0, not a vacuous pass, on one without.
- bench_csc_dispatch: bench_csc__de_csr_bounded, the CSR DE route at the
default shard cache — the regime a 16 GB pipeline runs in, which the
whole-matrix-cache CSR arm never measured.
- accel_de: the CPU Wilcoxon and pdex_ref variants on the bench-built CSC
fixture (accel_de__pyscx_{wilcoxon,pdex_ref}_cpu_csc), floored on
de_route_csc_cpu at pbmc3k and tabula_sapiens_100k.
- conversion_streaming: the base arms pin csc="off" (CLI: --csc off) so
their floors keep their meaning; a csc_auto arm passes no csc and
measures the new default, with a has_csc premise check and the disk
ratio in metadata.
- The fixture converter passes csc="auto" for scx_auto fixtures (so a
reconvert keeps the sidecar the floors depend on), the runner and the
timed rewrite/convert arms pin "off", and compression subtracts the
sidecar so its byte metric stays CSR-only.
- _add_csc_to_fixtures.sh is deleted: the reconvert carries the policy.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
The parallel emit encodes several CSC shards at once, and each holds its
arrays and encoder buffers. At census_1m build-csc (4 GiB default) that is
55.8 s / 6.75 GB against main's 94.6 s / 5.10 GB; batching whole shards
against the emit share instead is 77.1 s / 5.63 GB. Both are kept:
- Default: a batch hands the encoder every shard of the emit groups it
built (CscEmitOptions::whole_groups), the fast mode.
- A named budget — build-csc --memory-limit, --csc-memory-limit on the
rewrite ops and append --rebuild-csc, pyscx memory_limit= /
csc_memory_limit=, or --memory-budget on convert/sort/optimize (which
already hands the builder a spill share) — makes the emit shard-granular
against csc_emit_batch_nnz. CscBuildOptions::bounded_emit carries it.
- run_build_csc / rebuild_csc_inplace take memory_limit: Option<&str>;
None is DEFAULT_CSC_MEMORY_LIMIT ("4G"). The CLI flags and the pyscx
kwargs lose their "4G" default for the same reason (an unset flag is
now distinguishable from an explicit 4G). The shard layout, and so the
bytes, are the same for the same limit, named or not; a new build_csc
test pins that on a layout where the two modes batch differently.
- CscEmitter::next_batch hands out whole shards up to max_nnz unless
whole_groups is set; the old behaviour handed out every built shard,
which made max_nnz a floor of one group rather than a cap.
- The emit-share comments no longer claim the pair of batches fits the
share: only their arrays are charged, and the measured cost of each
mode is recorded on CscEmitOptions::whole_groups.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
…builder - sharding.md § Build policy: `auto` is the default on every entry point and preset, `off` the opt-out; the measured convert cost (wall and disk at tabula_100k / census_500k / census_1m) and the in-memory from_anndata cost, and which workloads the sidecar pays for. - operations.md: a second table for the parallel builder (serial -> parallel, and `--csc off`), and what naming a budget changes. - api.md / migrating-from-h5ad.md / skills conversion reference: every "csc=None resolves to off unless index_preset implies auto" is now "csc=None is auto"; `scx convert --csc off|auto|always`. - AGENTS.md CSC line and ROADMAP: the default, its cost, the parallel builder and the budget switch. - README: drop "CSC sidecars preserved through scx append" — append drops the sidecar, and refuses a multimodal target outright. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012LmrECbnnqEEHQtURZRn84
There was a problem hiding this comment.
Code Review
This pull request changes the default CSC (column-major sidecar) generation policy from off to auto across all ingest and conversion entry points, automatically building a sidecar for datasets with at least 50,000 cells and 5,000 genes. To support this, the CSC builder has been parallelized using Rayon to route shard pushes across column buckets in parallel and encode multiple CSC shards concurrently in batches. Optional memory budgets can be specified to bound peak memory during parallel encoding by switching the emit to shard-granular batches. Additionally, comprehensive benchmark suites, tests, and documentation have been updated to pin csc="off" where CSR-only baselines are measured, and to introduce new coverage for the parallel CSC build paths. I have no feedback to provide as no review comments were submitted.
|
codex - gpt-5.6-sol Defects, highest severity first
Over-engineering
Blast radius and verdictI reviewed the full 71-file diff at head
|
|
Cursor Agent - Grok 4.7 High Defects1. High — the canonical gate cannot go green on this commit
Fix: recapture 2. Medium —
|
|
Antigravity - Gemini 3.8 Flash Defects, ranked by severity1. High — CI failure:
|
…iptor cap - pyscx.from_mtx takes csc=None / csc_cols_per_shard=5000 and resolves them like every other ingest entry point; it was CSR-only at any size. The CLI's MTX post-pass moves into scx_ops::build_csc_for_policy, which both now call. (codex, Antigravity, Cursor) - The ingest guards key on the `from_*` names, with from_mudata the one named exclusion, instead of selecting functions that take both `csc` and `index_preset` — which by construction could not see an entry point missing `csc`. The signature guard fails on the pre-fix build. - scx convert --from mtx passes --memory-budget (through csc_sidecar_bytes, as streaming ingest does) and --temp-dir to its sidecar build; the MTX dispatch returned before the budget was parsed, so an invalid budget was ignored and a valid one never bounded the build. New CLI test fails without it. (Cursor) - Two hdf5-gated needless_borrow lints in pyscx/src/lib.rs, left by resolve_csc_policy returning &str; the Feature matrix (clippy, hdf5 legs) job. (all three) - The parallel emit drains spilled buckets concurrently, each holding its store reader, so the descriptor count grew with the rayon pool. Spill reads now take one of MAX_CONCURRENT_SPILL_READS (8) permits; a new test drains 200 spilled groups on a 32-thread pool under a descriptor limit 12 above baseline, and hits EMFILE without the cap. (Cursor) - ScxRunner has one CSC knob: `with_csc` is gone, SCX_BENCH_WITH_CSC sets csc="always", and compression forces csc="off". (codex, Antigravity) - docs: the build_csc / csc_memory_limit signatures show None rather than "4G" (an explicit "4G" now also bounds the emit); from_h5mu's API line says the streaming default cannot build a sidecar; from_mtx's csc.
|
Antigravity - Gemini 3.8 Flash Defects, ranked by severity1. High — CI failure:
|
|
Cursor Agent - Grok 4.7 High Reviewed head Defects1. High — the canonical gate is still red on this commit (untouched)
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: tabula_sapiens_100k
metric: de_route_is_csc
min: 1.0
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: census_500k
metric: de_route_is_csc
min: 1.0
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: census_1m
metric: de_route_is_csc
min: 1.0
2. Medium —
|
| Job | Result |
|---|---|
| Clippy | fail — defect 2 |
| Feature matrix (clippy, hdf5 legs) | fail — same lint |
| Feature matrix (clippy, no-hdf5 legs) | fail — same lint |
| Test | pass (7m50s) |
| Test (hdf5 features) | pass |
| Python bindings | pass (10m1s) |
| Base install | pass |
| Format, fuzz build, changed-path, dedup guards | pass |
| Docs anchors, fuzz matrix | skipping |
Nothing is still pending. Test and pytest passing does not cover the Clippy deny, and it does not run the uncommitted fd-limit test. I am not claiming this commit is regression-free.
Verdict: request changes. The MTX surface, the hdf5 borrows, the "4G" docs, and the dual ScxRunner knob are actually fixed in the code. What is still open is the gate (high, pre-existing) and a Clippy failure the fix commit introduced (medium).
|
codex - gpt-5.6-sol Defects, highest severity first
if output_path.exists() and not overwrite:
return output_path
...
if format_variant.key == "scx_auto":
runner.csc = FIXTURE_CSC_POLICY # "auto"while
assert_eq!(
mtx_csc(&["--memory-budget", "10MB"]).0,
false,
"decimal MB is rejected"
);Rust 1.98 rejects this under
scx_ops::build_csc_for_policy(
Path::new(scx_path),
csc_policy,
csc_cols_per_shard,
None,
None,
)This path is triggered by default precisely on large MTX inputs, yet callers cannot select the bounded emit or a spill directory; the PR's own census_1m figures put the unnamed emit at 6.75 GB peak. The CLI now forwards
let _permit = (store.spilled_bytes(b) > 0).then(SpillReadPermit::acquire);
pub const MAX_CONCURRENT_SPILL_READS: usize = 8;but
Over-engineering
Prior-round finding auditDuplicates reported by more than one reviewer are consolidated here; each distinct finding is checked against b036c5b.
Full blast radius and verdictThe behavior change reaches Below that policy layer, the PR changes bucket layout, Rayon push/drain scheduling, framed decode, batch encode/ordered commit, spill-file behavior, and bounded-vs-whole-group emit semantics. It also changes rewrite carry paths, benchmark fixture construction, route floors, and compression accounting, so the external fixture/baseline state is part of the functional blast radius, not merely benchmark housekeeping. I reviewed head Immediately before posting, Verdict: request changes. The round-1 MTX, borrow, CLI-budget, signature, and runner issues are materially improved, but the canonical gate remains unreproducible and b036c5b is not CI-clean. |
- scx-format-io/tests/csc_spill_parallel_fd_limit.rs was left untracked by b036c5b, so the descriptor cap it guards had no CI coverage. Committed. (Antigravity, Cursor, codex) - The new CLI MTX test's `assert_eq!(.., false, ..)` tripped clippy::bool_assert_comparison in all three clippy jobs. (all three) - pyscx.from_mtx gains memory_budget= / temp_dir=, passed to its sidecar build through csc_sidecar_bytes as from_h5ad does; an invalid budget is refused before anything is written. (Cursor, codex) - The parallel framed decode is folded into decode_framed_shard_scipy behind a parallel_groups flag: one group decode and one rebase, with the slot-filling half a helper beside it. The sink calls decode_shard_regions_scipy_with(.., true); every other caller keeps the serial walk. (codex, Antigravity; Cursor in round 1) - docs: the build-policy section lists from_mtx and says the auto default is every *single-modality* entry point; the skills reference no longer counts entry points it does not list. - Removed scx-sparse/proptest-regressions: the seeds were written by a deliberate mutation check, not by a real failure.
|
codex - gpt-5.6-sol Defects, highest severity first
if output_path.exists() and not overwrite:
return output_path
...
if format_variant.key == "scx_auto" and hasattr(runner, "csc"):
runner.csc = FIXTURE_CSC_POLICY
- benchmark: pipeline_ooc_constrained
dataset: tabula_sapiens_100k
metric: de_route_is_csc
min: 1.0(with the same floor for census_500k and census_1m). The description confirms that the shared
let _permit = (store.spilled_bytes(b) > 0).then(SpillReadPermit::acquire);
pub const MAX_CONCURRENT_SPILL_READS: usize = 8;and the new test deliberately grants
/// The default is `Auto` on every ingest entry point, since the builder runs
/// in the same pass as X and in parallel; `Off` is the opt-out.
but I found no additional defect introduced specifically by Over-engineering
pub use csc_builder::{ ..., MAX_CONCURRENT_SPILL_READS, ... };Make it private/
|
|
Antigravity - Gemini 3.8 Flash Defects, ranked by severity1. Medium — Canonical benchmark gate remains unversioned against
|
|
Cursor Agent - Grok 4.7 High Reviewed head Defects1. High — the canonical gate is still not reproducible on this commit (untouched)
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: tabula_sapiens_100k
metric: de_route_is_csc
min: 1.0
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: census_500k
metric: de_route_is_csc
min: 1.0
- benchmark: pipeline_ooc_constrained
format: pipeline_ooc_constrained__pyscx_16g
dataset: census_1m
metric: de_route_is_csc
min: 1.0A reconvert only applies if output_path.exists() and not overwrite:
logger.info(
"Skipping conversion (exists): %s -> %s at %s",
dataset.name, format_variant.key, output_path,
)
return output_path
2. Low —
|
| Job | Result |
|---|---|
| Clippy | pass (1m19s) |
| Feature matrix (clippy, hdf5 legs) | pass |
| Feature matrix (clippy, no-hdf5 legs) | pass |
| Test | pass (12m21s) |
| Test (hdf5 features) | pass |
| Python bindings | pass (9m15s) |
| Base install | pass |
| Format, fuzz build, changed-path, dedup guards | pass |
| Docs anchors, fuzz matrix | skipping |
Nothing is pending. That is the workspace suite, not gate_candidate.py. I am not claiming the benchmark gate is clean.
Verdict: request changes. The Clippy failure, the uncommitted descriptor test, the from_mtx budget kwargs, and the duplicate framed decoder are fixed in the code, and CI on 3a667411 is green. The gate floors still depend on fixtures and a baseline this commit does not version.
- CscPolicy's rustdoc limits the auto default to single-modality entry points, as the docs now do. (codex) - pyscx.from_mtx's comment no longer claims parity with from_h5ad beyond what holds: shard widths and the bounded emit match; the spill threshold is build-csc's half, not ingest's same-pass quarter. (Cursor) - MAX_CONCURRENT_SPILL_READS is pub(crate); the fd-limit test hard-codes its 8. (codex, Cursor, Antigravity) - decode_groups_into_slots takes the group decoder as `impl Fn` rather than a trait object. (Antigravity) - sharding.md says what naming --memory-limit changes besides the staging bound: the emit, measured, never the bytes. (Cursor)
All 16 versioned workspace members, `pyscx/pyproject.toml` and `rscx/DESCRIPTION`; `tests/scx-integration-tests` stays at `0.0.0`. README's `scx-cli` download snippet follows. (ROADMAP.md no longer exists, so there is no date stamp to move.) 0.20.0 collects the merged PRs since 0.19.0 (#544-#561): - CSC sidecar series: one-pass bucketed CSC builder (#555); `build-csc` as an in-place, rollback-able append (#556); same-pass CSC build on ingest and carry-through on rewrite ops (#557); parallel CSC build and `csc="auto"` as the ingest default (#558); CSC dispatch for row-indexed transforms and row-filtered / gene-subset handles (#553, #554); CSC route coverage and the exact-nnz Wilcoxon kernel as the 1-vs-rest CSC default (#559) - Correctness: Leiden parallel local moving keeps decliners eligible (#544); five CPU-accelerator parity / convergence fixes (#548); GPU cuVS stream sync, reduction determinism, decoder bounds and VRAM accounting (#549); CLI destination-overwrite safety and MTX multimodal / streaming / integer integrity (#550); unscoped whole-matrix reads of a multimodal file refused (#551); dictionary (categorical) output from append / merge / merge --sort-by (#546) and `scx sort`'s obs spill path (#547) - Benchmarks and docs: four community analytical benchmarks (#552), laptop-test recapture (#560), README refresh (#545), docs split into per-topic directories (#561) Behaviour changes worth calling out in the release notes: - **`csc="auto"` is the ingest default** on every entry point and preset (`from_mudata` keeps `off`): files with `n_obs >= 50000` and `n_vars >= 5000` now get a CSC sidecar, costing ingest wall time and +42-71 % on disk. `--csc off` / `csc="off"` opts out (#558). - **Rewrite ops carry the CSC sidecar by default** (`compact`, `merge`, `optimize`, `sort`, `subset`), rebuilt from the output's own shards; `--csc carry|always|off`. `--rebuild-csc` / `rebuild_csc=True` are deprecated aliases for `always` (#557). - **`scx build-csc` appends in place** rather than rewriting the file, and `scx rollback` removes the sidecar (#556). - **Ten CLI subcommands no longer silently overwrite an existing destination**; `--force` is required (convert, merge, subset, query --output, upgrade, cloud-optimize, explode, pack, pull; compact / optimize / sort migrated onto the same guard) and is refused where nothing is written (#550). - **MTX**: a declared-`integer` MTX with a value past 2^24, a non-integral value or `nan`/`inf` is refused on ingest (`--allow-lossy` restores the old behaviour); multimodal MTX export requires a modality (`pyscx.to_mtx` gains `modality=`); export streams one shard at a time (#550). - **An unscoped whole-matrix read of a multimodal file raises** `MultimodalRequiresModality` instead of folding every modality into one `n_obs x n_modalities`-row answer — `to_anndata()`, `to_memory()`, `read_all_csr_shards*`, `scx pull --filter` (#551). - **Numerical changes**: UMAP init scale and `random_init` now match umap-learn, the kNN sigma search, and NB-GLM dispersion shrinkage (the `above_min_disp` residual filter always runs; new `disp_outlier_sd=2.0` carve-out, `None` disables only the carve-out) (#548); parallel Leiden labels change (#544). - **Wilcoxon on the CPU CSC route** defaults to the exact-nnz kernel (route id `cpu_csc_nnz`); `SCX_ACCEL_WILCOXON_NNZ=0` restores densify (`cpu_csc`), and `reference=` / `rankby_abs=True` still use densify (#559). - **Categorical obs columns** keep their declared categories, unused levels and `ordered` bit through append / merge / sort (#546, #547). No benchmark recapture for this release. Pre-release gate on this tree: `cargo fmt --check`; `cargo clippy --workspace --exclude rscx --all-targets -D warnings`; `cargo test --workspace --exclude rscx` (4,220 passed, 0 failed); `cargo test -p scx-convert --features hdf5` (474 passed); `cargo test -p scx-cli --features hdf5 -- --test-threads=1` (253 passed); `maturin develop --release` + `pytest tests/` (3,247 passed, 140 skipped, 2 xfailed; `pyscx.__version__ == "0.20.0"`). Not run: rscx tests, `--features cloud`, GPU pytest, fuzzing, benchmarks. Claude-Session: https://claude.ai/code/session_01RwRKstgr4TNkXhLhkQwmjA Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An unset
cscnow means"auto"onscx convertand every pyscx ingest entry point: a CSC sidecar is built whenn_obs ≥ 50,000andn_vars ≥ 5,000, and--csc off/csc="off"opts out. Before this, it meant"off"unless the index preset wastrainingorperturbseq, so a file got a sidecar only if its author already knew to ask. Doing that cheaply needed a faster builder, so this PR also parallelises the CSC build. It then lands the gate half of PR-B (#554) that was deferred there: route floors under the pipeline's shape, and fixtures that carry a sidecar.Parallel CSC build (
b4dff202,9e8b1737)Measured before any of this, the default would have cost 3.9× (tabula_100k) to 5.5× (census_500k) of the convert wall time, because the same-pass build ran almost entirely on one thread. The writer decoded each X shard serially and pushed it one nonzero at a time, and the emit drained one bucket and encoded one shard at a time. At census_500k's layout (87 shards of 715 columns, three row groups each) the emit kept about three threads busy.
CscBuilder::push_shardroutes a shard with one task per column bucket. This is the newparallelfeature ofscx-sparse, turned on byscx-format-io's. Each task walks the rows in order and binary-searches its column range, so a bucket receives the same records in the same order and seals at the same records. A task that crosses the spill ceiling spills its own sealed blocks, which keeps the declaredspill_after + 2·n_buckets·block_capacitybound.target_buckets, a shard splits into equal buckets that never straddle a shard boundary. Shard widths are unchanged.CscEmitter::next_batchbuilds several emit groups concurrently.emit_csc_shardsvalue-encodes and codec-encodes a batch in parallel through the newScxWriter::write_csc_shards, and builds the next batch while the current one encodes.write_shard_inneris split into a pureencode_shard_sectionand a sequentialcommit_shard_section, and the batch path runs the same two functions.TempDirSpillStoreopens a bucket file once per spill instead of once per 1 MiB block.ResidentCscSource, which the in-memoryfrom_anndatapath now reaches by default, finds each shard's columns by binary search per row on canonical input instead of testing every nonzero once per shard. It also fills a batch of shards in parallel.Output is byte-identical to the serial build:
op_output_identitygolden and thecsc_same_passequivalence tests are unchanged and pass.main's only in the header, the provenance timestamp and the catalog checksums; I compared them byte for byte at tabula_100k and census_500k.The emit trades memory for speed, because each shard encoding at once holds its arrays and encoder buffers. I measured census_1m
build-csc(4 GiB default) on a 12-core node:main(serial)Both modes are kept. The default is whole groups. Naming a budget switches to shard-granular batches; that means any of
build-csc --memory-limit,--csc-memory-limit(rewrite ops,append --rebuild-csc),--memory-budget(convert/sort/optimize), or pyscxmemory_limit=/csc_memory_limit=. Those flags and kwargs therefore lose their literal"4G"default: unset is still 4G, but is now distinguishable from an explicit one.run_build_csc/rebuild_csc_inplacetakememory_limit: Option<&str>. The shard layout and the bytes are the same for the same limit, named or not.Measured on one node against
main(median of 2–3, no budget, job 2997546; the node reported 32 CPUs):main--csc offconverttabula_100kconvertcensus_500kconvertcensus_1mcompact --csc alwayscensus_1msort --csc alwayscensus_500kbuild-csccensus_1mThe default (
56865891)CscPolicy'sDefaultisAuto, andIngestOptions::default()follows it, so library callers change behaviour too.resolve_csc_policy(csc)no longer takes the preset. With"auto"the default for every preset, there is nothing left for a preset to decide, sopreset_implies_csc_autois deleted rather than widened tocellxgene.from_mudatakeeps"off": it cannot build per-modality sidecars; it refusesalwaysand degradesauto.autowith its existingCscSkippedStreamingMultimodalwarning. With the default flipped, that warning now fires on any large h5mu stream.What the default costs (same job): convert time is in the table above; disk grows +71% at tabula_100k, +54% at census_500k and +42% at census_1m. In-memory
from_anndatagoes from 15.5 s to 39 s at census_500k, where pyscx 0.18.0 took 207 s for the same sidecar.Gate and fixtures (
c6fe2acc)pipeline_ooc_constrained: the worker records the DE stage's route and carries it back through the stage JSON. It is floored asde_route_is_csc ≥ 1on tabula_sapiens_100k, census_500k and census_1m. Measured 1.0, routecpu_csc, on all three (job 2997709).bench_csc_dispatch: abench_csc__de_csr_boundedarm times the CSR DE route at the default shard cache, the regime a 16 GB pipeline runs in. It took 270 s at tabula.accel_de: the CPU Wilcoxon and pdex_ref variants now also run on the CSC fixture, with ade_route_csc_cpufloor. Measured 1.0, with p-values matching the in-memory CPU run, at pbmc3k and tabula.conversion_streaming:csc="off"(the CLI arms--csc off), so their floors keep their meaning.csc_autoarm measures the default: a 1.71× disk ratio at tabula.csc="auto"forscx_autofixtures, so a reconvert keeps the sidecar the floors depend on. The runner and the timed rewrite/convert arms pin"off", andcompressionsubtracts the sidecar so its byte metric stays CSR-only._add_csc_to_fixtures.shis deleted._auto.scxfixture that clears theautorule got its sidecar in place. That covers smartseq2 (+lognorm), tabula_sapiens_100k (+lognorm), census_500k/1m/5m, replogle_k562, tahoe_c38 and chemogenetic_rgfp. It usedscx build-csc, which is rollback-able. pbmc3k, pbmc10k and visium are below the threshold.LATESTbaseline has not been recaptured. Benchmarks that read these fixtures now take CSC routes, so a gate against the old baseline will see their numbers move.Found, not fixed
autocodec selection lands on ShufDeltaZstd, but it is worth reconsidering whetherpick_csc_encodingshould follow an Scx1 source.scx-benchconda env imports an installed pyscx 0.18.0 from site-packages, not the repo build. The first gate run of this PR silently tested 0.18.0, and acsc_autopremise check caught it. The rerun setPYTHONPATH=pyscx/python.Local verification
cargo test --workspace --exclude rscx,cargo test -p scx-convert --features hdf5andcargo test -p scx-cli --features hdf5 -- --test-threads=1pass.cargo clippy --workspace --exclude rscx --all-targets -D warnings,cargo fmt --checkandcargo check -p rscxare clean.pytest tests/: 3224 passed.main.