Idea, not approved work. Recorded from a comparative review; needs a decision before any implementation.
Gap
#471 shipped the H3 latent upscale (dw/tasks/h3_latent_upscale.py), but the refine pass was deferred (docs/proposals/complete/h3-latent-upscale-complete.md §Deferred), and the #500 gate judged upscale-only output "close but soft".
Idea
Two linked pieces:
- Audio-hold in the H3 denoise loop ("audio drive"). VAE-encode a fixed audio track (resample to the audio VAE rate), fit it to the template audio latent (crop/pad), and keep it un-denoised: per-stream mask of 1 for video, 0 for audio, so only video is sampled to the fixed audio. Mux the original waveform at the end, not the VAE round-trip.
- Refine pass on the upscaled latent. Upscale (pixel or latent), concat the stage-1 audio latent held as in (1), re-noise to a low sigma and run a few steps. Their shipped settings: denoise ~0.2, 4-5 steps; stage 1 at 0.4-0.5 MP with the turbo LoRA. A 3-pass variant repeats stage 2 at 2 MP.
Possible value
Costs / open questions
What would unpark it
A lem A/B on the #500 workflows: upscale-only vs upscale + 5-step 0.2 refine.
Their files: VRGDG_MiniMaxH3AudioDrive.py, Workflows/UsedForUIDoNotTouch/minimax_*2pass*/3pass*_api.json, scripts/build_minimax_h3_ref2va_2pass_audio_api.py.
Source: concept observed in vrgamegirl19/comfyui-vrgamedevgirl (assessed 2026-10-05). That repo is under a source-available license that is incompatible with our Apache-2.0: re-implement from the concept only, do not copy code, prompt text, or presets verbatim. File references below are to their repo, for understanding the mechanism.
Idea, not approved work. Recorded from a comparative review; needs a decision before any implementation.
Gap
#471 shipped the H3 latent upscale (
dw/tasks/h3_latent_upscale.py), but the refine pass was deferred (docs/proposals/complete/h3-latent-upscale-complete.md§Deferred), and the #500 gate judged upscale-only output "close but soft".Idea
Two linked pieces:
Possible value
chain-matched-to-audioonly passes the slice as a reference and H3 still generates its own audio latent. That should tighten lip sync.Costs / open questions
modular_pipelines/minimax_h3/denoise.pystepsaudio_latents[num_condition_audio_rows:], so it must either skip the audio scheduler step or re-impose the encoded audio after each step.What would unpark it
A lem A/B on the #500 workflows: upscale-only vs upscale + 5-step 0.2 refine.
Their files:
VRGDG_MiniMaxH3AudioDrive.py,Workflows/UsedForUIDoNotTouch/minimax_*2pass*/3pass*_api.json,scripts/build_minimax_h3_ref2va_2pass_audio_api.py.Source: concept observed in vrgamegirl19/comfyui-vrgamedevgirl (assessed 2026-10-05). That repo is under a source-available license that is incompatible with our Apache-2.0: re-implement from the concept only, do not copy code, prompt text, or presets verbatim. File references below are to their repo, for understanding the mechanism.