Skip to content

Select sklearn baseline hyperparameters once across all folds - #137

Open
cornelislouisa wants to merge 1 commit into
mainfrom
guille/kfold-sklearn-baselines
Open

cornelislouisa wants to merge 1 commit into
mainfrom
guille/kfold-sklearn-baselines

Conversation

@cornelislouisa

Copy link
Copy Markdown
Collaborator

Why

The sklearn baselines were not comparable to the GNNs, in two independent ways.

1. Hyperparameters were tuned separately inside each fold

Each fold ran its own grid search, so two folds of the same cell could be reported under different hyperparameters. That is not what the GNN side does: an Optuna trial evaluates one configuration across the whole rotation and is scored on statistics.fmean of the fold results.

It also biases the comparison. Picking a per-fold winner reports the best of several draws from a noisy validation estimate, which flatters the baseline.

2. The baselines saw different features than the GNNs

Baselines ran SelectKBest on top of whatever columns they were handed, while the GNNs consume nodes chosen by variance, correlation, distance_correlation, or random against the training split. Any measured accuracy gap was partly a feature-selection artifact rather than a modeling result.

What changed

Selection happens once per cell. Every grid candidate is scored on every fold's validation split, the candidate with the highest mean wins, and all five folds refit and evaluate with that single configuration. Ties resolve by grid order, matching GridSearchCV.

The choice does not depend on which fold's job computes it, so it is cached under a fold-independent key and the other four jobs reuse it. An fcntl lock makes that safe when Hydra starts the fold jobs concurrently.

Refitting still uses training data only, and test data never takes part in selection.

Baselines now reproduce the GNN node selection directly — same node budget derived from the training split size, same train-only selector, and the RNG seeding that makes method=random reproducible. SelectKBest is removed from the dataset configs so the feature set is not filtered a second time.

Folds are built once per job and shared between SVM and elastic net. Rebuilding per baseline meant running the distance-correlation scan over every gene twice, which dominates runtime on the wide datasets.

Verification

scripts/verify_baseline_features.py compares the baseline feature matrix against the selected_data.parquet artifacts the GNN pipeline writes. Across all six datasets the columns match in identical order with a maximum absolute difference of 0.0, including random.

tests/test_baseline_global_hparams.py covers the parts that are easy to get subtly wrong:

  • a candidate that wins one fold outright still loses on the mean
  • ties resolve by grid order
  • all folds are scored once, then the cache is reused
  • the shared fold cache prevents a second feature build per baseline
4 passed

Notes for the reviewer

  • Opt-in and scoped. baseline_hparam_selection in configs/baseline.yaml gates this, and it applies only to k-fold runs. Fixed splits, and anyone setting per_fold, keep the previous behavior.
  • best_cv_score now means the global mean validation score, not this fold's score from tuning. Worth knowing when comparing against older runs.
  • The selection cache key includes imputation_method. It changes the feature matrix, so a cached selection must not survive a change to it.
  • The dataset YAML diffs are SelectKBest removal only. Flipping split_type to k-fold by default affects GNN training too, so it is deliberately left for its own PR.
  • WGCNA density targeting is a separate PR. Nothing here builds a graph; baseline_filter=standard never touches adjacency.

The baselines were not comparable to the GNNs in two ways.

Hyperparameters were tuned independently inside each fold, so two folds
of the same cell could be reported under different hyperparameters. That
is not what the GNN side does: an Optuna trial evaluates one
configuration across the whole rotation and is scored on the mean. A
per-fold winner also reports the best of several draws from a noisy
validation estimate, which flatters the baseline.

Selection now happens once per cell. Every grid candidate is scored on
every fold's validation split, the candidate with the highest mean wins,
and all five folds are refit and evaluated with that single
configuration. Ties resolve by grid order, matching GridSearchCV. The
choice does not depend on which fold's job computes it, so it is cached
under a fold-independent key and the remaining jobs reuse it. Refitting
still uses training data only, and test data never takes part in
selection.

The feature sets also differed. Baselines ran SelectKBest on top of
whatever columns they were given, while the GNNs consume nodes chosen by
variance, correlation, distance correlation, or random sampling against
the training split. Any accuracy gap was partly a feature-selection
artifact. Baselines now reproduce the GNN node selection directly,
including the node budget and the RNG seeding that makes method=random
reproducible, and SelectKBest is gone from the dataset configs.
verify_baseline_features.py checks this against the parquet artifacts
the GNN pipeline writes.

Within one job the folds are built once and shared between SVM and
elastic net. Rebuilding per baseline meant running the
distance-correlation scan over every gene twice.

Selection is enabled by baseline_hparam_selection in configs/baseline.yaml
and applies only to k-fold runs. Fixed splits, and anyone who sets
per_fold, keep the previous behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant