Skip to content

Save IGV hosted genomes by ID in sessions - #1875

Merged
jrobinso merged 16 commits into
mainfrom
session-genome-id
Sep 17, 2026
Merged

jrobinso merged 16 commits into
mainfrom
session-genome-id

Conversation

@jrobinso

Copy link
Copy Markdown
Contributor

Sessions on an IGV hosted genome now record just the genome ID instead of an expanded reference object.

A frozen copy of a hosted genome definition goes stale as the server gains annotations or corrects URLs; the ID is smaller and always resolves to the current definition.

"genome": "hg38"

replaces the reference block (which was already written minus its tracks). Anything not on the genome server — GenArk, hubs, local files, user genomes — still writes the expanded reference exactly as before, and both forms are read.

Not loading the genome's tracks twice

A session lists every track, the genome's default annotations included, so when a session names a genome by ID those annotations are not loaded again with the genome: GenomeManager.loadGenomeById / loadGenome / setCurrentGenome / restoreGenomeTracks take a loadAnnotationTracks flag, and the session's track list is the sole authority. The sequence track and the legacy .genome/.gbk gene track are still restored.

A hand-written session that names a genome by ID and expects its default annotations without listing them can set "loadGenomeTracks": true, which merges the genome's tracks with the session's, matched on url. Default false. The reference branch is untouched and still merges inline reference.tracks.

Fixes found along the way

  • The by-ID branch skipped its session reset when the requested genome was already current. Harmless while IGV never wrote "genome" itself; once it does, loading a second session on the same genome leaks ROIs, sample attributes and frames from the first.
  • loadGenomeById's return value was discarded, so a genome that failed to resolve left the session's tracks loaded against whatever genome was still current — a session that looks loaded but is plotted on the wrong assembly. It now fails.
  • Genbank genomes could not be loaded from the genome chooser — a regression against 2.19.x. The dialog assumed every entry had a json or hub definition to rename, choose annotations from and download; .gbk has neither, and is now loaded directly.
  • Genbank genomes are identified by the versioned accession (NC_012920.1) from the VERSION line, matching the hosted genome list. Previously the unversioned ACCESSION was used, so a session on the mitochondrial genome wrote a reference carrying no sequence at all — GenomeConfig.toJSON serializes only URL fields, and a genbank genome's sequence lives in the file.
  • GenArk records were keyed by the assembly column — names like Loxafr3.0 that nothing else uses — and that list has no url column, so all 52,777 had a null path and could not be loaded by ID at all. They are now keyed by accession, with the path derived via HubGenomeLoader.convertToHubURL.
  • The genome chooser's legacy .genome cleanup was passed the accession, so it looked for a filename that never exists.

Keeping the genome list cheap

Answering "is this ID hosted?" used to download the 5.8 MB UCSC GenArk list along with the 4 KB IGV list, from the EDT, the autosave timer thread and the JVM shutdown hook. The two lists are now loaded independently: the session writer and ID resolution touch only the small one, and GenArk loads solely for the genome chooser or an ID that is not ours. The IGV list is also cached in the genome directory — fetched live every session, with the local copy used only when the server cannot be reached, so an outage does not silently cost sessions their genome IDs.

Test sessions

test/sessions/json/*.json converted to the ID form. The webapp-* sessions were written by igv.js and are left alone, as is blat-session.json, whose genome is GenArk rather than IGV hosted.


The matching igv.js change is not part of this PR.

🤖 Generated with Claude Code

jrobinso and others added 16 commits September 16, 2026 18:37
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Session saving runs on the timer and shutdown threads, where the
multi-MB list fetch and its error dialog could stall or hang exit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The list has no url column, so these records had a null path and could
not be loaded by ID; the assembly column is a name, not the ID a loaded
GenArk genome takes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A .gbk genome has no json or hub definition, so skip the config,
annotation-selection and download steps and load the file directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ACCESSION line is unversioned, so the genome ID did not match the
NC_012920.1 of the hosted genome list, and sessions fell back to an
expanded reference that carries no sequence for a genbank genome.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
removeDotGenomeFile was passed the accession, so it looked for a file
name that never exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A live list is read every session; the cached copy is only a fallback,
so an unreachable server does not cost sessions their genome IDs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@jrobinso
jrobinso merged commit e4e152c into main Sep 17, 2026
2 checks passed
@jrobinso
jrobinso deleted the session-genome-id branch September 17, 2026 06:20
@maximilianh

maximilianh commented Sep 17, 2026 via email

Copy link
Copy Markdown

@jrobinso

Copy link
Copy Markdown
Contributor Author

@maximilianh In the future we should move in that direction, but "hosted" in this context means it is one of the 30 or so genomes hosted on our server. By hosted I mean the genome json, most of the files are hosted at UCSC. A long term goal would be to eliminate this list and use the UCSC apis directly for all, as we do now for Genark. That will be easier if all annotations are eventually reachable as files, the genomes that have a mix of database and file tracks remain a problem for IGV.

@maximilianh

maximilianh commented Sep 17, 2026 via email

Copy link
Copy Markdown

@jrobinso

Copy link
Copy Markdown
Contributor Author

Actually we don't need all files, just at a minimum a 2 bit sequence and some annotations.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants