Repository navigation
Compute website graph statistics on the canonical fold-0 split - #136
Merged
Merged
Conversation
The statistics behind the dataset explorer were computed on a dataset loaded with default settings, which does not correspond to any graph a model is actually trained on. This builds them from fold 0 of the 5-fold rotation instead, so the reported structure matches a real training graph. Concretely, load_dataset now requests split_type='k-fold' with k=5 and fold=0, and passes the train-only corrections and the grouping key each dataset requires. Without these, the graph is built from a different sample set than training uses: motrpac needs covariate adjustment, addneuromed needs ComBat, smoking needs promoter collapsing and median centering, and parkinsons must keep hybridization batches unmixed. Feature homophily was averaging node features over every sample, including validation and test. Since the graph itself is built from training samples only, this mixed held-out data into a statistic describing a train-derived object. It now averages over the training block, read from the split_info.json cut point. Graph statistics also read the adjacency matrix through a hard-coded absolute path under /home/lcornelis, which only resolved on one machine and silently reported an empty graph elsewhere. It now reads from the dataset's own raw_dir. The new test pins the loader arguments per dataset, so the correction and grouping tables here cannot drift away from configs/dataset/*.yaml without a failure. Note this describes fold 0 specifically, not an average across folds. Per-fold variation exists but reporting five sets of numbers per cell on the website is not useful. Co-authored-by: Cursor <cursoragent@cursor.com>
johmathe
approved these changes
Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The numbers behind the dataset explorer were computed from a dataset loaded with default settings. That does not correspond to any graph a model is actually trained on, so the reported structure described an object that exists nowhere in our experiments.
This rebuilds them from fold 0 of the 5-fold rotation, so what the website reports is a real training graph.
What changed
load_datasetnow requests the canonical fold. It passessplit_type='k-fold',k=5,fold=0, plus the train-only corrections and grouping key each dataset requires. Without these the graph is built from a different sample set than training uses:motrpacaddneuromedsmokingparkinsonsgrouping='batch')Feature homophily no longer reads held-out data. It was averaging node features over every sample, including validation and test. The graph itself is built from training samples only, so this mixed held-out data into a statistic describing a train-derived object. It now averages over the training block, read from the
split_info.jsoncut point.Adjacency is read from the dataset's own
raw_dir. The previous code used a hard-coded absolute path under/home/lcornelis, which resolved on exactly one machine and silently reported an empty graph everywhere else.Tests
tests/tutorials/test_dataset_stats_analysis.pypins the loader arguments per dataset, so the corrections and grouping tables in this script cannot drift away fromconfigs/dataset/*.yamlwithout a test failure. It stubs the dataset class, so there is no download and no graph construction.Notes for the reviewer
ruff-formatchurn. That file already failsruff-formatonmain, and CI runs pre-commit against changed files, so any PR touching it has to absorb that.webapp/public/data/stats.jsonand the CSVs are intentionally left out so this PR stays reviewable; they are a separate data-only change.adjacency_thresholdsweep, which is what the explorer's threshold slider needs. Train-fold density targeting is a separate PR.