Repository navigation
Select sklearn baseline hyperparameters once across all folds - #137
Open
cornelislouisa wants to merge 1 commit into
Open
cornelislouisa wants to merge 1 commit into
cornelislouisa wants to merge 1 commit into
Conversation
The baselines were not comparable to the GNNs in two ways. Hyperparameters were tuned independently inside each fold, so two folds of the same cell could be reported under different hyperparameters. That is not what the GNN side does: an Optuna trial evaluates one configuration across the whole rotation and is scored on the mean. A per-fold winner also reports the best of several draws from a noisy validation estimate, which flatters the baseline. Selection now happens once per cell. Every grid candidate is scored on every fold's validation split, the candidate with the highest mean wins, and all five folds are refit and evaluated with that single configuration. Ties resolve by grid order, matching GridSearchCV. The choice does not depend on which fold's job computes it, so it is cached under a fold-independent key and the remaining jobs reuse it. Refitting still uses training data only, and test data never takes part in selection. The feature sets also differed. Baselines ran SelectKBest on top of whatever columns they were given, while the GNNs consume nodes chosen by variance, correlation, distance correlation, or random sampling against the training split. Any accuracy gap was partly a feature-selection artifact. Baselines now reproduce the GNN node selection directly, including the node budget and the RNG seeding that makes method=random reproducible, and SelectKBest is gone from the dataset configs. verify_baseline_features.py checks this against the parquet artifacts the GNN pipeline writes. Within one job the folds are built once and shared between SVM and elastic net. Rebuilding per baseline meant running the distance-correlation scan over every gene twice. Selection is enabled by baseline_hparam_selection in configs/baseline.yaml and applies only to k-fold runs. Fixed splits, and anyone who sets per_fold, keep the previous behavior. Co-authored-by: Cursor <cursoragent@cursor.com>
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The sklearn baselines were not comparable to the GNNs, in two independent ways.
1. Hyperparameters were tuned separately inside each fold
Each fold ran its own grid search, so two folds of the same cell could be reported under different hyperparameters. That is not what the GNN side does: an Optuna trial evaluates one configuration across the whole rotation and is scored on
statistics.fmeanof the fold results.It also biases the comparison. Picking a per-fold winner reports the best of several draws from a noisy validation estimate, which flatters the baseline.
2. The baselines saw different features than the GNNs
Baselines ran
SelectKBeston top of whatever columns they were handed, while the GNNs consume nodes chosen byvariance,correlation,distance_correlation, orrandomagainst the training split. Any measured accuracy gap was partly a feature-selection artifact rather than a modeling result.What changed
Selection happens once per cell. Every grid candidate is scored on every fold's validation split, the candidate with the highest mean wins, and all five folds refit and evaluate with that single configuration. Ties resolve by grid order, matching
GridSearchCV.The choice does not depend on which fold's job computes it, so it is cached under a fold-independent key and the other four jobs reuse it. An
fcntllock makes that safe when Hydra starts the fold jobs concurrently.Refitting still uses training data only, and test data never takes part in selection.
Baselines now reproduce the GNN node selection directly — same node budget derived from the training split size, same train-only selector, and the RNG seeding that makes
method=randomreproducible.SelectKBestis removed from the dataset configs so the feature set is not filtered a second time.Folds are built once per job and shared between SVM and elastic net. Rebuilding per baseline meant running the distance-correlation scan over every gene twice, which dominates runtime on the wide datasets.
Verification
scripts/verify_baseline_features.pycompares the baseline feature matrix against theselected_data.parquetartifacts the GNN pipeline writes. Across all six datasets the columns match in identical order with a maximum absolute difference of0.0, includingrandom.tests/test_baseline_global_hparams.pycovers the parts that are easy to get subtly wrong:Notes for the reviewer
baseline_hparam_selectioninconfigs/baseline.yamlgates this, and it applies only to k-fold runs. Fixed splits, and anyone settingper_fold, keep the previous behavior.best_cv_scorenow means the global mean validation score, not this fold's score from tuning. Worth knowing when comparing against older runs.imputation_method. It changes the feature matrix, so a cached selection must not survive a change to it.SelectKBestremoval only. Flippingsplit_typetok-foldby default affects GNN training too, so it is deliberately left for its own PR.baseline_filter=standardnever touches adjacency.