Skip to content

Update website graph statistics with fold-0 stats for all six datasets - #140

Open
johmathe wants to merge 1 commit into
mainfrom
johan/website-iclr-graph-stats
Open

johmathe wants to merge 1 commit into
mainfrom
johan/website-iclr-graph-stats

Conversation

@johmathe

Copy link
Copy Markdown
Collaborator

What does this PR do?

Updates the website's Dataset Explorer with the fold-0 graph statistics for all six datasets (motrpac, addneuromed, parkinsons, brca, tuberculosis, smoking): 4 ratios × 4 selection methods × 10 thresholds × 2 graph constructions = 1920 graphs, including the new LCC metrics and feature homophily from #136.

Data

  • stats/<dataset>/graph_stats_comprehensive.csv — rewrote the four existing CSVs and added smoking/ and tuberculosis/, in the column layout tutorials/dataset_stats_analysis.py emits (both string and wgcna rows per file). These are the committed source; make build regenerates stats.json from them.
  • webapp/public/data/stats.json — regenerated (960 → 1920 entries).

Build script (webapp/scripts/build_stats_from_csv.py)

  • Globs stats/*/graph_stats_comprehensive.csv; drops the stale tutorials/stats/*_addneuro.csv legacy inputs.
  • Carries clustering_coefficient, diameter, modularity, homophily.
  • Writes null instead of NaN so fetch().json() can't choke on it.
  • Bug fix: no longer collapses node_sample_ratio=full into key 1. full (all features) and 1.0 (n_nodes = n_train / 1.0) are different graphs; the old key collision meant the site only ever showed the 1.0 graph and full was unreachable. full is now its own slider stop.

Webapp

  • Explorer.tsx — six datasets in a 3×3 metric grid (nodes, edges, avg degree, density, degree std, LCC %, clustering coeff, homophily, modularity). Explicit subplot domains so row titles and angled tick labels don't collide; y-axes scaled to the bars on screen (edge counts span ~100× across τ, node counts ~50× between full and subsampled).
  • types.ts / constants.ts / data.ts — DatasetName widened to six; new LeaderboardDatasetName + LEADERBOARD_DATASETS (four) so the Leaderboard is unchanged until results.json has runs for the new datasets.
  • Leaderboard.tsx — dataset dropdown iterates LEADERBOARD_DATASETS rather than all of DATASETS.
  • webapp/README.md — pipeline, metric list, key format.

Verification

  • astro check: 0 errors. make build passes.
  • Local preview renders 54 bars (9 metrics × 6 datasets) for the default selection and for full + WGCNA.
  • Leaderboard unchanged: 4 datasets, 11 models, charts render.
  • pre-commit passes on all changed files.

Follow-ups (not in this PR)

  • tutorials/stats/ still holds stale pre-fold-0 CSVs (one is rows of Permission denied: '/scratch'). Nothing reads them anymore; worth removing.
  • Add smoking/tuberculosis to LEADERBOARD_DATASETS once their benchmark results land in results.json.

Replace the Dataset Explorer data with the fold-0 graph statistics
computed by tutorials/dataset_stats_analysis.py for motrpac, addneuromed,
parkinsons, brca, tuberculosis and smoking (4 ratios x 4 selection
methods x 10 thresholds x 2 graph constructions = 1920 graphs). The
per-dataset CSVs under stats/ are the committed source; the build script
regenerates webapp/public/data/stats.json from them.

build_stats_from_csv.py now globs stats/*/graph_stats_comprehensive.csv,
carries the new LCC metrics (clustering_coefficient, diameter,
modularity) and feature homophily, drops the stale tutorials/stats
inputs, and writes null rather than NaN so the JSON stays parseable in
the browser.

It also stops collapsing node_sample_ratio 'full' into '1'. Those are
different graphs (n_nodes = n_train / ratio; 'full' keeps every
feature), and the old key collision meant the site only ever showed the
1.0 graph. 'full' is now its own slider stop in the Explorer.

Explorer shows all six datasets in a 3x3 metric grid with explicit
subplot domains so row titles and angled tick labels no longer collide,
and scales each y-axis to the bars on screen since edge counts span two
orders of magnitude across thresholds. The Leaderboard keeps its own
four-dataset list (LEADERBOARD_DATASETS) because results.json has no
runs for the two new datasets yet.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant