Save IGV hosted genomes by ID in sessions - #1875
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This reverts commit cfc3cf4.
Session saving runs on the timer and shutdown threads, where the multi-MB list fetch and its error dialog could stall or hang exit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The list has no url column, so these records had a null path and could not be loaded by ID; the assembly column is a name, not the ID a loaded GenArk genome takes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A .gbk genome has no json or hub definition, so skip the config, annotation-selection and download steps and load the file directly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ACCESSION line is unversioned, so the genome ID did not match the NC_012920.1 of the hosted genome list, and sessions fell back to an expanded reference that carries no sequence for a genbank genome. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
removeDotGenomeFile was passed the accession, so it looked for a file name that never exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A live list is read every session; the cached copy is only a fallback, so an unreachable server does not cost sessions their genome IDs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
To answer the question "is this ID hosted" you could also use the UCSC API,
that would be a much smaller request.
|
Contributor
Author
|
@maximilianh In the future we should move in that direction, but "hosted" in this context means it is one of the 30 or so genomes hosted on our server. By hosted I mean the genome json, most of the files are hosted at UCSC. A long term goal would be to eliminate this list and use the UCSC apis directly for all, as we do now for Genark. That will be easier if all annotations are eventually reachable as files, the genomes that have a mix of database and file tracks remain a problem for IGV. |
|
Ok let’s discuss that one day, I guess we should be able to provide files
for all tracks …
…On Thu, Sep 17, 2026 at 18:02 Jim Robinson ***@***.***> wrote:
*jrobinso* left a comment (igvteam/igv#1875)
<#1875 (comment)>
@maximilianh <https://github.com/maximilianh> In the future we should
move in that direction, but "hosted" in this context means it is one of the
30 or so genomes hosted on our server. By hosted I mean the genome json,
most of the files are hosted at UCSC. A long term goal would be to
eliminate this list and use the UCSC apis directly for all, as we do now
for Genark. That will be easier if all annotations are eventually reachable
as files, the genomes that have a mix of database and file tracks remain a
problem for IGV.
—
Reply to this email directly, view it on GitHub
<#1875?email_source=notifications&email_token=AACL4TN6KK2IQXXWRK6BGWD5PQDJZA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZRG42DGMRSGQ32M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5717432247>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AACL4TKB7MQX66JUTQ3UIY35PQDJZAVCNFSNUABDKJSXA33TNF2G64TZHMZTIMJWGU3TQO2JONZXKZJ3GU2DQMZXGA4TQMBSUF3AE>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/AACL4TL7QTWTHXINB7B4BN35PQDJZA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZRG42DGMRSGQ32M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJKTGN5XXIZLSL5UW64Y>
and Android
<https://github.com/notifications/mobile/android/AACL4TO47C6PE7XQO2XJJ2D5PQDJZA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZRG42DGMRSGQ32M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>.
Download it today!
You are receiving this because you were mentioned.Message ID:
***@***.***>
|
Contributor
Author
|
Actually we don't need all files, just at a minimum a 2 bit sequence and some annotations. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sessions on an IGV hosted genome now record just the genome ID instead of an expanded
referenceobject.A frozen copy of a hosted genome definition goes stale as the server gains annotations or corrects URLs; the ID is smaller and always resolves to the current definition.
replaces the
referenceblock (which was already written minus itstracks). Anything not on the genome server — GenArk, hubs, local files, user genomes — still writes the expandedreferenceexactly as before, and both forms are read.Not loading the genome's tracks twice
A session lists every track, the genome's default annotations included, so when a session names a genome by ID those annotations are not loaded again with the genome:
GenomeManager.loadGenomeById/loadGenome/setCurrentGenome/restoreGenomeTrackstake aloadAnnotationTracksflag, and the session's track list is the sole authority. The sequence track and the legacy.genome/.gbkgene track are still restored.A hand-written session that names a genome by ID and expects its default annotations without listing them can set
"loadGenomeTracks": true, which merges the genome'strackswith the session's, matched onurl. Default false. Thereferencebranch is untouched and still merges inlinereference.tracks.Fixes found along the way
"genome"itself; once it does, loading a second session on the same genome leaks ROIs, sample attributes and frames from the first.loadGenomeById's return value was discarded, so a genome that failed to resolve left the session's tracks loaded against whatever genome was still current — a session that looks loaded but is plotted on the wrong assembly. It now fails..gbkhas neither, and is now loaded directly.NC_012920.1) from theVERSIONline, matching the hosted genome list. Previously the unversionedACCESSIONwas used, so a session on the mitochondrial genome wrote areferencecarrying no sequence at all —GenomeConfig.toJSONserializes only URL fields, and a genbank genome's sequence lives in the file.assemblycolumn — names likeLoxafr3.0that nothing else uses — and that list has nourlcolumn, so all 52,777 had a null path and could not be loaded by ID at all. They are now keyed by accession, with the path derived viaHubGenomeLoader.convertToHubURL..genomecleanup was passed the accession, so it looked for a filename that never exists.Keeping the genome list cheap
Answering "is this ID hosted?" used to download the 5.8 MB UCSC GenArk list along with the 4 KB IGV list, from the EDT, the autosave timer thread and the JVM shutdown hook. The two lists are now loaded independently: the session writer and ID resolution touch only the small one, and GenArk loads solely for the genome chooser or an ID that is not ours. The IGV list is also cached in the genome directory — fetched live every session, with the local copy used only when the server cannot be reached, so an outage does not silently cost sessions their genome IDs.
Test sessions
test/sessions/json/*.jsonconverted to the ID form. Thewebapp-*sessions were written by igv.js and are left alone, as isblat-session.json, whose genome is GenArk rather than IGV hosted.The matching igv.js change is not part of this PR.
🤖 Generated with Claude Code