Skip to content

Compute website graph statistics on the canonical fold-0 split - #136

Merged
cornelislouisa merged 1 commit into
mainfrom
guille/fold0-dataset-stats
Sep 24, 2026
Merged

cornelislouisa merged 1 commit into
mainfrom
guille/fold0-dataset-stats

Conversation

@cornelislouisa

Copy link
Copy Markdown
Collaborator

Why

The numbers behind the dataset explorer were computed from a dataset loaded with default settings. That does not correspond to any graph a model is actually trained on, so the reported structure described an object that exists nowhere in our experiments.

This rebuilds them from fold 0 of the 5-fold rotation, so what the website reports is a real training graph.

What changed

load_dataset now requests the canonical fold. It passes split_type='k-fold', k=5, fold=0, plus the train-only corrections and grouping key each dataset requires. Without these the graph is built from a different sample set than training uses:

Dataset Required for the graph to match training
motrpac covariate adjustment
addneuromed ComBat
smoking promoter collapsing, median centering
parkinsons hybridization batches kept unmixed (grouping='batch')

Feature homophily no longer reads held-out data. It was averaging node features over every sample, including validation and test. The graph itself is built from training samples only, so this mixed held-out data into a statistic describing a train-derived object. It now averages over the training block, read from the split_info.json cut point.

Adjacency is read from the dataset's own raw_dir. The previous code used a hard-coded absolute path under /home/lcornelis, which resolved on exactly one machine and silently reported an empty graph everywhere else.

Tests

tests/tutorials/test_dataset_stats_analysis.py pins the loader arguments per dataset, so the corrections and grouping tables in this script cannot drift away from configs/dataset/*.yaml without a test failure. It stubs the dataset class, so there is no download and no graph construction.

6 passed

Notes for the reviewer

  • These statistics describe fold 0 specifically, not an average across folds. Per-fold variation exists, but reporting five sets of numbers per cell on the website is not useful.
  • The diff includes ~34 lines of ruff-format churn. That file already fails ruff-format on main, and CI runs pre-commit against changed files, so any PR touching it has to absorb that.
  • No regenerated artifacts here. webapp/public/data/stats.json and the CSVs are intentionally left out so this PR stays reviewable; they are a separate data-only change.
  • WGCNA still uses the explicit adjacency_threshold sweep, which is what the explorer's threshold slider needs. Train-fold density targeting is a separate PR.

The statistics behind the dataset explorer were computed on a dataset
loaded with default settings, which does not correspond to any graph a
model is actually trained on. This builds them from fold 0 of the 5-fold
rotation instead, so the reported structure matches a real training
graph.

Concretely, load_dataset now requests split_type='k-fold' with k=5 and
fold=0, and passes the train-only corrections and the grouping key each
dataset requires. Without these, the graph is built from a different
sample set than training uses: motrpac needs covariate adjustment,
addneuromed needs ComBat, smoking needs promoter collapsing and median
centering, and parkinsons must keep hybridization batches unmixed.

Feature homophily was averaging node features over every sample,
including validation and test. Since the graph itself is built from
training samples only, this mixed held-out data into a statistic
describing a train-derived object. It now averages over the training
block, read from the split_info.json cut point.

Graph statistics also read the adjacency matrix through a hard-coded
absolute path under /home/lcornelis, which only resolved on one machine
and silently reported an empty graph elsewhere. It now reads from the
dataset's own raw_dir.

The new test pins the loader arguments per dataset, so the correction
and grouping tables here cannot drift away from configs/dataset/*.yaml
without a failure.

Note this describes fold 0 specifically, not an average across folds.
Per-fold variation exists but reporting five sets of numbers per cell on
the website is not useful.

Co-authored-by: Cursor <cursoragent@cursor.com>
@cornelislouisa
cornelislouisa merged commit 8afeeaf into main Sep 24, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants