Repository navigation
Conversation
Replace the Dataset Explorer data with the fold-0 graph statistics computed by tutorials/dataset_stats_analysis.py for motrpac, addneuromed, parkinsons, brca, tuberculosis and smoking (4 ratios x 4 selection methods x 10 thresholds x 2 graph constructions = 1920 graphs). The per-dataset CSVs under stats/ are the committed source; the build script regenerates webapp/public/data/stats.json from them. build_stats_from_csv.py now globs stats/*/graph_stats_comprehensive.csv, carries the new LCC metrics (clustering_coefficient, diameter, modularity) and feature homophily, drops the stale tutorials/stats inputs, and writes null rather than NaN so the JSON stays parseable in the browser. It also stops collapsing node_sample_ratio 'full' into '1'. Those are different graphs (n_nodes = n_train / ratio; 'full' keeps every feature), and the old key collision meant the site only ever showed the 1.0 graph. 'full' is now its own slider stop in the Explorer. Explorer shows all six datasets in a 3x3 metric grid with explicit subplot domains so row titles and angled tick labels no longer collide, and scales each y-axis to the bars on screen since edge counts span two orders of magnitude across thresholds. The Leaderboard keeps its own four-dataset list (LEADERBOARD_DATASETS) because results.json has no runs for the two new datasets yet. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Updates the website's Dataset Explorer with the fold-0 graph statistics for all six datasets (
motrpac,addneuromed,parkinsons,brca,tuberculosis,smoking): 4 ratios × 4 selection methods × 10 thresholds × 2 graph constructions = 1920 graphs, including the new LCC metrics and feature homophily from #136.Data
stats/<dataset>/graph_stats_comprehensive.csv— rewrote the four existing CSVs and addedsmoking/andtuberculosis/, in the column layouttutorials/dataset_stats_analysis.pyemits (bothstringandwgcnarows per file). These are the committed source;make buildregeneratesstats.jsonfrom them.webapp/public/data/stats.json— regenerated (960 → 1920 entries).Build script (
webapp/scripts/build_stats_from_csv.py)stats/*/graph_stats_comprehensive.csv; drops the staletutorials/stats/*_addneuro.csvlegacy inputs.clustering_coefficient,diameter,modularity,homophily.nullinstead of NaN sofetch().json()can't choke on it.node_sample_ratio=fullinto key1.full(all features) and1.0(n_nodes = n_train / 1.0) are different graphs; the old key collision meant the site only ever showed the1.0graph andfullwas unreachable.fullis now its own slider stop.Webapp
Explorer.tsx— six datasets in a 3×3 metric grid (nodes, edges, avg degree, density, degree std, LCC %, clustering coeff, homophily, modularity). Explicit subplot domains so row titles and angled tick labels don't collide; y-axes scaled to the bars on screen (edge counts span ~100× across τ, node counts ~50× betweenfulland subsampled).types.ts/constants.ts/data.ts—DatasetNamewidened to six; newLeaderboardDatasetName+LEADERBOARD_DATASETS(four) so the Leaderboard is unchanged untilresults.jsonhas runs for the new datasets.Leaderboard.tsx— dataset dropdown iteratesLEADERBOARD_DATASETSrather than all ofDATASETS.webapp/README.md— pipeline, metric list, key format.Verification
astro check: 0 errors.make buildpasses.full+ WGCNA.pre-commitpasses on all changed files.Follow-ups (not in this PR)
tutorials/stats/still holds stale pre-fold-0 CSVs (one is rows ofPermission denied: '/scratch'). Nothing reads them anymore; worth removing.smoking/tuberculosistoLEADERBOARD_DATASETSonce their benchmark results land inresults.json.