Repository navigation
Conversation
|
Thanks for your suggestion.
This is very low. What happens with datasets with higher particle density? How does the scaling behaviour compare with the original implementation? Finally, please declare the use of LLMs if you are using one for code generation and/or analysis. |
|
Hi, @biochem-fan, thanks for your suggestions. I use GPT-5.6-sol via Codex for most code generation, all benchmark scripts and numerical validation. The benchmark param matrix is made up myself. Your claim on particles size is fair, since the smoke micrograph is randomly picked. I rerun the benchmark with selected fiber rich micrographs. I'd like to test how it scales, however scaling legacy polish with MPI procs is impossible in my system since I have only 64 GiB RAM installed. I will test how thread scales. For each micrograph, particles are extracted with https://ftp.ebi.ac.uk/empiar/world_availability/12870/data/RAW_Tif/Peter_08032022_206-5_00010_X%201Y%200-1.tif (202 particles extracted)
https://ftp.ebi.ac.uk/empiar/world_availability/12870/data/RAW_Tif/Peter_08032022_191-14_00013_X%200Y%200-1.tif (242 particles extracted)
https://ftp.ebi.ac.uk/empiar/world_availability/12870/data/RAW_Tif/Peter_08032022_170-17_00027_X%201Y%201-3.tif (208 particles extracted)
Benchmark
|
Where does this deferring happen? Can you provide more details? Where are your movies stored? SSD? How does the access pattern change with your patch? In the original code, a full movie is read by the main thread before the multi-thread section starts, right? Or was it already parallel? I don't remember well. My concern is that if each thread reads (part of) movies, it generates many non-sequential disk reads and makes HDD-based systems very slow.
It is a pity. The code does not scale to many threads / MPI because the number of particles per movie is limited. MPI scaling behaviour is also important. |
They are in HDD, however this doesn't matter since the tests are run multiple times on same micrographs, they are in memory cache already. |
|
In Here, the
|
I made a mistake that memory saving/defer loading/streaming doesn't yield from the wrapper. The wrapper actually accounts for runtime optimization. In that case, particles or other are stored in a continuous memory block, reduce overhead of allocating/reclaiming on heap. The memory usage optimization and the wrapper are technically orthogonal to each other. The wrapper touchs too many things and should be separated.
It does scale in my system. For legacy, MPI mprocs scaling is limited since ~40GiB per process exceeds 64 GiB RAM easily, regardless of particle count of the micrograph. In this PR, the most significant memory footprint cost shifts to particles count and box size. I can run with |
|
I am getting confused. If deferred reading is not implemented here, where does the memory saving come from?
Is the core problem the copy assignment and life time of
I meant movie frame loading in |



This PR reduces the peak memory and runtime of Bayesian polishing by removing the full corrected detector-movie stack from the motion-refinement path and consolidating particle/frame images into reusable contiguous allocations.
Instead of load everything into RAM at once and index them with offset like
It now uses an wrapper and index them with wrapper param. The wrapper will load frames only when needed.
Benchmark
The benchmark selects one raw micrograph containing 12 particles, which are from K3 superresolutioned dataset EMPIAR-12870. The raw size is 11,520 x 8,184 x 48 (head and tail are dropped).
Numerical validation
The multithreaded polishing path is not bit identical for two reasons:
rand()from an OpenMP loop, so scheduling can assign different replacement samples to hot pixels.With one thread over all 48 production frames:
cc,w0, andw1MRC files are bit identical;The exact single-thread result establishes that the memory-layout and streaming changes preserve the polishing calculation when execution order is fixed.