Stage B of #602. Section Stages → Stage B: rewire upscale-clip and refine-clip of plan v4; the plan comment stays authoritative, including The design, Templates (stage B), Risks and Decisions. Needs stage A (#631, verified).
Filed by the lead agent, model claude-opus-5-5 via anthropic. Reshaped for plan v4 (reconciled 2026-10-06 by the same model): v2's ref_width/ref_height breaking change is gone (Q2′). upscale-clip keeps width/height as the output size, and the stage now also changes fit_to_model itself: a downscale argument, and video as an fps-carrying float32 ndarray. No breaking-change.
What it builds
- In
dw/tasks/fit.py:
fit_to_model gains downscale: an optional positive integer, default 1. It fits into width/downscale × height/downscale and records that as the model size. It refuses, naming both values, when width or height isn't divisible by it. Its get_task description says so.
fit_to_model's video becomes a float32 numpy.ndarray subclass carrying .fps (in dw/media_types.py, alongside AudioVideo), in place of an AudioVideo. Then previous_result:fit.video feeds a pipeline argument or a condition's frames directly.
- Unit tests for both. One realizes
previous_result:fit.video into a pipeline's arguments and a condition's frames through the real argument path, not a mock. Write that one first. If a condition's frames still won't take it, stop and re-plan (plan Risks); don't bend the engine.
- In both
upscale-clip and refine-clip, the same step sequence: fit (fit_to_model, mode from a new fit variable, default letterbox) → the pipeline fed previous_result:fit.video → restore (restore_to_source, fit: previous_result:fit.fit) → the existing pair_audio.
- upscale-clip:
width/height stay the output size (default 960×544, 32n) and go to the pipeline as today. fit gets width: variable:width, height: variable:height, downscale: 2, so the reference is 480×272, and diffusers' crop becomes a no-op. The IC reference is previous_result:fit.video in place of variable:source_video.
- refine-clip: fit to
width×height. loop_frames is replaced, and the upsampler's stretch becomes a no-op.
- If the GPU runs show an edge halo at the letterbox edge, flip the default to
stretch (Q1).
What it owes
- Rewritten descriptions and summaries in both templates, within the
COMPACT_BUDGET (tests/test_catalog_structure.py:411). They state the changed behaviour: a different aspect is letterboxed and restored, a short source is held and trimmed, and the output is exactly 2× the source at the source's length. They also say which size each template's variables name: the output for upscale-clip, the working size for refine-clip.
workflows/templates/ltx2/README.md, plugins/dw/skills/ltx-2.5/SKILL.md with a plugin version bump, and docs/RECIPES_24GB.md:186.
- Updated tests that pin today's shape:
tests/test_ltx2_ic_loras.py (:212, :174-185 for upscale-clip only, :216), tests/test_catalog_structure.py:566-645 (TestLtxRefineClip), and the tests/test_step_cache.py:706-726 hashes if the configuration changes.
Deploy
Server, plus a plugin release.
Passes when
Its own repro works, per the plan's stage B acceptance intent:
- both templates validate clean on a 640×480, 50-frame, 24 fps asset with a soundtrack;
final/ is 1280×960, 50 frames, 24 fps, with audio at the source's length, no edge content lost, no bars, and no lapped tail;
intermediate/ shows the model frame with bars;
- a source already at the working size gives the same output size as before;
downscale: get_task shows it, width=512, height=288, downscale=2 gives a 256×144 fit with that model size in the record, and downscale=3 with width=512 is refused naming both;
- the existing refusals are unchanged;
and every pending: #632 case the tester writes passes.
After it lands (lead files to Don)
- The Q3 follow-up (adopt the pair in restore-deblur / restore-decompression), once letterbox holds up in these runs.
- Removal of
loop_frames as a proposal, if no template consumer is left.
Stage B of #602. Section Stages → Stage B: rewire upscale-clip and refine-clip of plan v4; the plan comment stays authoritative, including The design, Templates (stage B), Risks and Decisions. Needs stage A (#631, verified).
Filed by the lead agent, model
claude-opus-5-5viaanthropic. Reshaped for plan v4 (reconciled 2026-10-06 by the same model): v2'sref_width/ref_heightbreaking change is gone (Q2′). upscale-clip keepswidth/heightas the output size, and the stage now also changesfit_to_modelitself: adownscaleargument, andvideoas an fps-carrying float32 ndarray. Nobreaking-change.What it builds
dw/tasks/fit.py:fit_to_modelgainsdownscale: an optional positive integer, default 1. It fits intowidth/downscale × height/downscaleand records that as the model size. It refuses, naming both values, whenwidthorheightisn't divisible by it. Itsget_taskdescription says so.fit_to_model'svideobecomes a float32numpy.ndarraysubclass carrying.fps(indw/media_types.py, alongsideAudioVideo), in place of anAudioVideo. Thenprevious_result:fit.videofeeds a pipeline argument or a condition'sframesdirectly.previous_result:fit.videointo a pipeline's arguments and a condition'sframesthrough the real argument path, not a mock. Write that one first. If a condition'sframesstill won't take it, stop and re-plan (plan Risks); don't bend the engine.upscale-clipandrefine-clip, the same step sequence:fit(fit_to_model, mode from a newfitvariable, defaultletterbox) → the pipeline fedprevious_result:fit.video→restore(restore_to_source,fit: previous_result:fit.fit) → the existingpair_audio.width/heightstay the output size (default 960×544, 32n) and go to the pipeline as today.fitgetswidth: variable:width,height: variable:height,downscale: 2, so the reference is 480×272, and diffusers' crop becomes a no-op. The IC reference isprevious_result:fit.videoin place ofvariable:source_video.width×height.loop_framesis replaced, and the upsampler's stretch becomes a no-op.stretch(Q1).What it owes
COMPACT_BUDGET(tests/test_catalog_structure.py:411). They state the changed behaviour: a different aspect is letterboxed and restored, a short source is held and trimmed, and the output is exactly 2× the source at the source's length. They also say which size each template's variables name: the output for upscale-clip, the working size for refine-clip.workflows/templates/ltx2/README.md,plugins/dw/skills/ltx-2.5/SKILL.mdwith a plugin version bump, anddocs/RECIPES_24GB.md:186.tests/test_ltx2_ic_loras.py(:212,:174-185for upscale-clip only,:216),tests/test_catalog_structure.py:566-645(TestLtxRefineClip), and thetests/test_step_cache.py:706-726hashes if the configuration changes.Deploy
Server, plus a plugin release.
Passes when
Its own repro works, per the plan's stage B acceptance intent:
final/is 1280×960, 50 frames, 24 fps, with audio at the source's length, no edge content lost, no bars, and no lapped tail;intermediate/shows the model frame with bars;downscale:get_taskshows it,width=512, height=288, downscale=2gives a 256×144 fit with that model size in the record, anddownscale=3withwidth=512is refused naming both;and every
pending: #632case the tester writes passes.After it lands (lead files to Don)
loop_framesas a proposal, if no template consumer is left.