Repository navigation
idea: overlap windowing for long video sources (split, process, cosine-blend join) #601
Description
Activity
- addedenhancementNew feature or requestNew feature or requestowner:donParked for human decisionParked for human decisionideaAn idea for the researcher agent to assessAn idea for the researcher agent to assessowner:leadFeature lead's turn (harnest R11)Feature lead's turn (harnest R11)priority:2Backlog rank: solid ROI, moderate scopeBacklog rank: solid ROI, moderate scopefeatureWork bigger than a fix: designed with Don, built in stages (harnest R11)Work bigger than a fix: designed with Don, built in stages (harnest R11)and removedowner:donParked for human decisionParked for human decisionideaAn idea for the researcher agent to assessAn idea for the researcher agent to assess
on Oct 5, 2026 Approved by Don (2026-10-05) at priority:2. Build it as a split/join task pair in pixel space, composed with per_entry steps. The window grid comes from the workflow's variable_constraints. Handing to the feature lead for a plan.
Plan v3: overlap windowing for long video sources (
window_video+join_windows)Feature plan by the lead agent, model
claude-opus-5-5viaanthropic, 2026-10-05. Design read againstdevelop@ 21cd845. v2 records Don's approval answers (Q1–Q3) and his change to the stage 3 preflight, re-read againstdevelop@ 837fd24. v3 (2026-10-06, from stage 3's build, #630): the template ships unpriced, a new stage 4 records the measuredper_entry, and the plugin version bump is dropped.Verdict: build smaller
Axis Finding Value Don approved the direction on 2026-10-05. There is no field-report demand yet: no driving session has reported a source that is too long for restore-*,refine-cliporupscale-clip. The gain is real but prospective. Today a source over one LTX bucket (121 frames ≈ 5 s) can't be restored or refined at all;refine-cliptrims it.Build cost 4 stages, about $4 + $5 + $4 + $1 ≈ $14. Inertia Two utility tasks and one template. MCP, REST and UI enumerate tasks dynamically ( dw/server/routes/system.py:86,dw_mcp/tools_catalog.py:135,ui/src/lib/api.ts:185). There is no new MCP tool, no schema change and no surface-budget cost (tests/test_mcp_server.py:1555is unaffected). Upkeep is twoAudioVideoconstruction sites (tests/test_shots.py:73), argument domains, aTASKS.mdsection, anARCHITECTURE.mdrow and a seam/sync contract forassess_output. No security boundary changes and no new concept: it composesfor_each/gather:.Reversibility High. No stored data depends on it. Removing it later would break only the template and any saved workflow that names the tasks. Cheaper alternatives (a) Do nothing: an agent splits by hand, with slice_audio-style trims per run anddissolve_videosacross the results. But no task cuts a video at a start frame, anddissolve_videos' linear ramp shortens the cut by the overlap, so frames are lost. (b) Join only: gives nothing without the split.Overlap #597 (latent spatial tiling) is parked and independent: neither blocks the other. #600 (music-video timing) doesn't touch this. What "smaller" cuts relative to the idea text:
- The H3 17n+5 grid. Every long-source consumer (
restore-*,refine-clip,upscale-clip) is LTX. The tasks are grid-agnostic, so H3 needs no task work, only a template. Brings it back: an H3 video-to-video template that needs it. - Per-window audio blending in the join. The join puts the source audio back, sliced exactly to the source (Q2, decided). Brings it back: a windowed template whose model generates audio worth keeping.
- One template, not three.
ltx2/restore-longcomes first (Q1, decided). Brings the others back: a field report asking for a long refine or upscale.
Problem and non-goals
Problem. LTX video-to-video templates take at most one model bucket. Nothing cuts a long source into overlapping model-sized windows and stitches the processed windows back to the source's exact length.
Non-goals:
- Latent-space or spatial tiling (idea (parked): tiled latent fusion for tile-trained LTX-2.5 IC-LoRAs (Refine-Details, Restore, Deblur) #597).
- Windows chosen at run time from the source's length.
for_eachexpands before validation (dw/for_each.py:61-124), so the window list is a workflow variable (Q3, decided). - More than 32 windows (
dw/for_each.py:43). At 121-frame windows with a 16-frame overlap that is 3,360 frames, about 140 s at 24 fps. - Copying anything from the cited repo: this is a clean re-implementation from the concept.
- Template-specific logic in validation (Don, 2026-10-05). Every validate-time rule here is owned by a task.
Design, with corrections to the idea
Correction 1: "per_entry" is cost pricing, not iteration. In dw,
per_entryonly prices a list (dw/plan.py:248). Iteration isfor_eachover avariable:list, withitem:andgather:. The design composes those.Correction 2: the split can't return a list of windows. A list result fans every later
previous_result:out per item (dw/previous_results.py:56-123), and nothing can iterate a list made at run time. Sowindow_videoreturns one window per call, picked byindex. The workflow drives it withfor_eachovervariable:windows, a list whose entries are{name, index}.Correction 3: the grid comes from the existing
num_framesconstraint, not a new frame-snap argument. The window length isnum_frames, andstride = num_frames - overlap. The template bindsnum_framestovariable:num_frames, which the LTX constraint already holds to 8n+1 (dw/variable_constraints.py). This needs no new constraint plumbing and matches Don's "grid from variable_constraints".Correction 4: the join can't read the window plan from its inputs. Pipeline outputs don't carry their input's shot records (
dw/result.py:650-711). Sojoin_windowstakessource(the samevariable:source_video) and re-derives the plan from the source's frame count. That also gives it the source audio to re-attach.window_video(video, index, num_frames, overlap, fps=None)returns oneAudioVideo- Range: window
icovers source frames[i·stride − overlap, i·stride + stride). - Synthetic frames:
- Frames before 0 are frame 0 repeated (the first window's prefix).
- Frames past the end are the last frame repeated (the final window's pad).
- Audio: when the source has audio, the window carries source samples
f2s(a)…f2s(b)for its real span[a, b), usingframes_to_samples(dw/task_domains.py:402).- Real frames are always measured on source-frame boundaries, so adjacent windows tile exactly (the shot records: dissolve_videos floors frame→sample, pair_audio/gain_audio round — re-pairing an episode shifts its shots by one sample #401 cumulative rule).
- Synthetic frames get silence of matching
f2slength.
- Frame format: float [0,1] like
loop_frames(dw/tasks/video_utils.py:163), because the LTX reference condition takes that form. - Refusals:
overlap >= num_frames;overlap < 0;index < 0;- an
indexwhose window starts at or past the source's last frame (index · stride >= source_frames).
- Not a probe (
assessment=False).
join_windows(videos, source, num_frames, overlap, curve="cosine", fps=None)returns oneAudioVideo- Inputs:
videosisgather:<step>, the processed windows in list order. - Count rule: the required window count is
ceil(source_frames / (num_frames − overlap)). A different count is refused (Q3, decided); the refusal names both numbers and the list entries to add or drop. The rule has one home, a function indw/task_domains.pybesidedissolve_shortfalls, called by both the task at run time (stage 2) and its static check at validate time (stage 3). - Blend: it emits
source_framesframes, each source frame exactly once.- Window 0 contributes its real frames, the ones after the synthetic prefix.
- Each later window's first
overlapframes are blended over the previous window's lastoverlapreal frames with weightw(t), wheret = (k+1)/(overlap+1). That is the same open ramp as_ramp(dw/tasks/dissolve_videos.py:259), so neither end is a hard copy. w(t)is cosine(1−cos πt)/2, smoothstep3t²−2t³, or lineart.- The final window's pad frames are dropped.
- More refusals:
- windows whose frame counts differ from
num_frames, naming which window and its count; - windows of different sizes. They may differ from the source's size, since
refine-clipdoubles it.
- windows whose frame counts differ from
- Audio (Q2, decided): the source's audio over all
source_frames. When the source has none, so does the output. The windows' own audio is discarded. - Shot records: one shot per window. Each is named from
videos, its span covers the frames it owns,overlap_framesis set on every seam, andstart_sample/num_samplesare cumulativef2s. That wayassess_output's seam probe treats the blends as dissolves (dw/tasks/assess.py:677), and the sync probe passes. - Catalog shape: added to
_CUT_TASKS(dw/server/catalog_shape.py:61).
join_windows' static window-count check (Don, 2026-10-05)A validate-time check owned by
join_windows, in the shape ofdw/dissolve_frame_errors.py(a module of its own, wired intovalidation_errors):- It applies to any
join_windowsstep, never to a named template, so a later windowed template gets it with no extra code. - It fires when
sourcenames a knowable file (asset:/output:, or a literal path the run may read;resolve_probe_path+probe_metadata, header only),num_framesandoverlapare literal after substitution, andvideosisgather:<step>whose member count is known afterfor_eachexpansion. - It errors when
len(gather list) ≠ ceil(source_frames / (num_frames − overlap)), naming the count needed, at thejoin_windowsstep's path, by calling the stage 2 rule. - Anything not knowable (a
previous_result:source, a non-literalnum_frames) is left to the run-time refusal. Silence there is correct, asdissolve_frame_errorsdocuments.
Surfaces touched
These are what
needs-approvalexists for, decided here:- Engine: two new utility tasks, and one validate-time check owned by
join_windows. - MCP: none new; they appear in
list_tasks/get_task. - REST: none new.
- Syntax: none new. They use
for_each,item:andgather:as they are. - Catalog: one new template (stage 3).
Stages
Stages are built in order, and each one lands on
developon its own. Stages 1–2 are unreferenced surface until stage 3 wires them in.Stage 1:
window_video(server)- Builds: the task as specified above, plus:
- its
TASK_ARGUMENT_DOMAINSentries (indexandoverlapnon-negative,num_framespositive), checked at run time likeslice_audio; - the
AudioVideosite intests/test_shots.pyEXPECTED_SITESand indw/shots.py's list.
- its
- Tests:
- window spans for first, middle and last windows;
- prefix and pad frames;
- audio sample spans tiling exactly across windows, at 24 and 25 fps with 44.1k and 48k audio;
- each refusal;
- a source with no audio.
- Docs:
docs/TASKS.md(Video Processing) and andocs/ARCHITECTURE.mdrow. - Deploy: server restart.
- Estimate: ~$4.
- Acceptance intent (over MCP, with a short
asset:clip, e.g. 50 frames,num_frames=17,overlap=4, stride 13):get_task("window_video")describes it.- A one-step workflow with
index=0succeeds with exactly 17 frames.- Its first 4 frames match source frame 0 (
get_output_frames). - Its audio is 17 frames long, and the first 4 frames' worth is silent.
- Its first 4 frames match source frame 0 (
index=3(frames 35–51) has 17 frames, and its last 2 repeat source frame 49.index=4is refused, naming the source's frame count and the last valid index (3).overlap=17is refused, and so is a negativeoverlap.- The refusals arrive at
validate_workflowwhere it can see the asset's length, or otherwise at run time with a clear error. The tester should accept either and note which.
Stage 2:
join_windows(server)- Builds: the task as specified above, including the curves, the count/size/length refusals, source-audio re-attachment, shot records with
overlap_frames, the_CUT_TASKSentry, and the window-count rule as a shared function indw/task_domains.py. - Tests:
- round-trip: split a synthetic ramp video, join the unprocessed windows, and get back exactly the source frames for all three curves (identical overlap content blends to itself);
- curve shapes at the endpoints;
- each refusal;
- shot spans and samples cumulative per shot records: dissolve_videos floors frame→sample, pair_audio/gain_audio round — re-pairing an episode shifts its shots by one sample #401;
- a no-audio source.
- Docs:
TASKS.mdandARCHITECTURE.md. - Deploy: server restart.
- Estimate: ~$5.
- Acceptance intent:
- A workflow that windows an asset clip with
for_each(no model step) and joins withgather:returns a video with the source's exact frame count.- It is visually identical to the source at sampled frames.
- Its audio length matches the source.
assess_outputraises no sync-drift finding, and its seams read as dissolves.
- A window list one entry short is refused with the required count (at run time in this stage).
- Windows of mixed length are refused, naming the odd window.
curve="bogus"is refused.
- A workflow that windows an asset clip with
Stage 3:
join_windowsstatic count check,ltx2/restore-longtemplate and skill note (server + plugin)- Builds:
- The static window-count check owned by
join_windows, as specified above: a module in the shape ofdw/dissolve_frame_errors.py, wired intovalidation_errors, calling the stage 2 rule. No template-specific logic in validate. workflows/templates/ltx2/restore-long.json. It holds:- a
windowfor_each overvariable:windows; - a
restorefor_each, therestore-deblurpipeline step withframes: previous_result:window; - a
joinstep withjoin_windows(gather:restore, source=variable:source_video); - no
costblock yet. A cost is measured, never derived, andrestore-deblurhas none to start from. Stage 4 adds it; windows(soindex) kept out ofcost_driversnow, so stage 4'sper_entrywon't be made unknown (validate_workflow: plan.estimate ignores a per-entry num_frames in a list-driven template (follow-up to #589) #593's_list_entry_field_shifted).
- a
- A short note in the
dw:ltx-2.5skill on choosingwindowsfor a long source.
- The static window-count check owned by
- Tests:
- the catalog-structure tests pass on the template;
- check unit tests on a plain
join_windowsworkflow (not the template): short, exact and long lists; an unknowable source or a non-literalnum_framesstays silent.
- Docs: the template's own description,
WORKFLOW_GUIDE.md's for_each example list, and the check inTASKS.md'sjoin_windowssection. - Deploy: server restart. No plugin version bump: the version is pinned to the engine's and moves only at release (
tests/test_plugin_skills.py). The skill note reaches the tester throughorigin/develop. - Estimate: ~$4.
- Acceptance intent:
list_workflows(shape=…)lists the template, andvalidate_workflowon it is clean. Itsplan.estimateis unknown (nocostyet), and validate doesn't refuse on that.- The tester's verify comment records the minutes of one
restore@wNstep and the whole run's minutes, with the window count and device, for stage 4. - Against a ~10 s asset clip with the right
windowslist, the run succeeds with the source's frame count and audio length.- No visible seam at the window boundaries (
get_output_framesaround each seam). assess_outputraises no sync-drift finding.
- No visible seam at the window boundaries (
- The same run with one entry too few fails at
validate_workflow, naming the needed count, before any GPU time. - A hand-written
window_video/join_windowsworkflow (no model step, not the template) with a short list is also refused atvalidate_workflow: the check belongs to the task. - Verified on the LTX-capable server (CUDA, 24 GB).
Stage 4:
per_entrycost forltx2/restore-long(catalog)- Builds: the template's
costentry from stage 3's acceptance run.minutesis the run's total, andper_entryis{variable: "windows", minutes: <one restore window plus its slice>, entries: 3}, measured as shots-batch's was (Add templates/minimax/shots-batch: list-driven H3 shot generation with no in-job assembly #352). If that run used a list other than the default 3, the total is the measured run's, andentriesand the default list match it. - Tests:
per_entry_problemsand the catalog-structure tests pass. - Docs: none beyond the template.
- Deploy: server restart.
- Estimate: ~$1.
- Acceptance intent:
validate_workflowonrestore-longwith 2, 3 and 4 windows returns aplan.estimatewithbasis: per_entrythat scales by window count. - Fallback: if stage 3's verify comment lacks the timings, Don runs the default template once on lem, and the figure comes from that job.
Risks
- Content drift between windows. A restore model may shift colour or detail per window, so a cosine blend over a few frames can show a slow "breathing". Mitigations are the default
overlap(16) and the stage 3 visual check. If it shows, a follow-up could condition each window on the previous window's tail, which is a bigger feature. - Pipeline frame count ≠
num_frames. H3 snaps up, and LTX should match.join_windowsrefuses rather than guesses. - Hand-written window lists. An agent writes the
windowslist from the source length. The static check and the join's count refusal both name the right number. - Already-flagged sync probe rounding.
_sample_span(dw/tasks/assess.py:255) and thefade_samplesrounding at:727don't follow shot records: dissolve_videos floors frame→sample, pair_audio/gain_audio round — re-pairing an episode shifts its shots by one sample #401's cumulative rule. If they trip stage 2's acceptance, the fix belongs to the implementer as its own bug, not to this plan.
Decisions (Don, 2026-10-05)
- Q1. Which template comes first?
ltx2/restore-long, built onrestore-deblur. - Q2. What audio does the join output? The source's audio, re-attached exactly.
- Q3. Refused or derived window count? Refused, naming the needed count.
- Stage 3 preflight: a static check owned by
join_windowsthat states its own rule. There is no template-specific logic in validate.
- The H3 17n+5 grid. Every long-source consumer (
- addedowner:donParked for human decisionParked for human decisionstatus:plan-reviewA feature plan is posted and waiting for DonA feature plan is posted and waiting for Donand removedowner:leadFeature lead's turn (harnest R11)Feature lead's turn (harnest R11)
on Oct 5, 2026 Plan v1 approved by Don (2026-10-05). Q1-Q3 defaults accepted:
ltx2/restore-longcomes first;- the join re-attaches the source audio;
- a wrong window count is refused, and the refusal names the needed count.
One change to stage 3: the validate-time window-count preflight should be a static check owned by
join_windows, which states its own rule:len(gather list) == ceil(source_frames / (num_frames - overlap))whensourceis an asset with a known length. Don't add template-specific logic to validate. Any later windowed template then gets the check with no extra code.40 remaining items
Verify of #601 (feature parent): FAILED on C-F233 step 2. Bouncing to
owner:lead. Tester: claude-opus-5-5 via anthropic, lem on develop @ b8aef18.Failing case: C-F233, step 2 (
output:source)The case says a wrong window count must be refused at
validate_workflowwhen the source is anoutput:reference, naming the needed count of 5 ("Anoutput:source is knowable"). It is not.- I made a 124-frame source with
run_workflow{"id":"cc","steps":[{"name":"cat","task":{"command":"concat_videos","arguments":{"videos":["asset:qa-cast/ep6-cold-open.mp4"]}},"result":{"content_type":"video/mp4","fps":24}}]}.- Job 1b73865a4a47 succeeded and wrote
cc/20261007-030343-8f1db730/cc-cat.0-0.0.mp4, one shot of 124 frames.
- Job 1b73865a4a47 succeeded and wrote
- I ran
validate_workflowon C-F229's RT (num_frames 33, overlap 8, cosine) withsource_video: "output:cc/20261007-030343-8f1db730/cc-cat.0-0.0.mp4"and 4 windows (w0–w3, indexes 0–3).- Result:
valid: true, no errors, no warnings,plan.list_entries.windows: 4.
- Result:
- The reference does resolve at validate. The same workflow with
output:cc/20261007-030343-8f1db730/nope.mp4is refused atvariables.source_videowith "Output '…/nope.mp4' not found under …/outputs". run_workflowon the step 2 workflow queued job 17b72f533155. It ran all 4 windows, then failed atjoinwith:- "join_windows needs 5 windows for a 124-frame source with num_frames 33 and overlap 8 (stride 25), got 4 - add 1 entry (index 4)".
- So the run-time check works, but the static check misses an
output:source.
For contrast, the same RT with the
asset:qa-cast/ep6-cold-open.mp4source is refused at validate atsteps[1]naming 5, for 4, 6 and 1 windows (observed this session). The static check seems to coverasset:sources only.Passed this session (all over MCP)
- C-F314: pass.
- Defaults validate.
- A wrong count is refused naming the needed count: ep13 with 2, 4 or 1, and ep11 with 3.
- ep11 with 5 is valid.
- run_workflow with 2 windows was refused with no job queued.
- C-F315: pass. Job 63da56c3f40e.
- Output: 282 frames, 24 fps, 960×544, 44.1 kHz stereo, 11.75 s.
- Shots: w0 [0, 89), w1 [89, 194) with overlap 16, w2 [194, 282) with overlap 16.
- sync_drift: no findings, max offset 0.01 ms.
- Both seams are dissolves with no findings.
- Timings, recorded in
regression-perf/C-F315.jsonl:- w0 3.04 min (cold)
- w1 1.48 min
- w2 1.51 min
- whole job 11.18 min
- Note: the
join_windowstask alone took about 5 min.
- C-F316: pass (the observed clause).
- per_entry is {windows, 1.56, 3}.
- Estimates for 1–5 windows were 10.1, 11.2, 12.2, 13.2 and 14.3: strictly increasing, none unknown.
- C-F317: pass.
- per_entry 1.56 against the warm median of 1.495.
- The 3-window estimate of 12.2 against the job's 11.18.
- SE-F042: pass.
- Every refusal is at the argument's own path, with the same wording for an existing and a nonexistent path. The
..andfile://probes are refused as well. - run_workflow on the /etc/passwd source was refused with no job.
- Wording ambiguity, not a failure: the arms that probe only
video, withsourcestill the asset, also return the count error. That count comes from the legitimate asset source, so nothing about the probed path leaks.
- Every refusal is at the argument's own path, with the same wording for an existing and a nonexistent path. The
- C-F229 (functional part): RT with the cosine, smoothstep and linear curves all succeeded (jobs 4a1b0541190c, 02c2622da816, a0d24251bf7f).
- Output: 124 frames, 24 fps, 960×544, 32 kHz stereo, 5.1667 s.
- 5 shots, overlap_frames 8, last shot ending at sample 122667 + 42666 = 165333.
- assess_output on the cosine join: no findings, with sync_drift and the seam rules applied.
- Not done: the per-frame source comparison.
- C-F231, arm 4: curve "bogus" is refused at
steps[1].task.arguments.curve, naming cosine, smoothstep and linear. - C-F233:
- Step 1 passes: 5 windows valid; 4, 6 and 1 refused at
steps[1]naming 5. - Steps 3 and 4 pass: no count error for
previous_result:framesor a non-literal num_frames. - Step 2 fails, as above.
- Step 1 passes: 5 windows valid; 4, 6 and 1 refused at
- C-F227 (refusal arms):
- (17, 17), (17, 20), (17, −1 / index −1) and (0, 0) are each refused at validate, naming the argument.
- (17, 4, idx10) is refused at run: "window 10 would start at source frame 130, but the source has 124 frames … the last window is index 9".
- (35, 4, idx4) is refused at run, naming the last index 3.
- Not done: the accepted arms 8–10.
Not run this session
C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 were stopped short on budget once a failing case had been confirmed. The next verify of the parent should run the full set again.
All outputs from this session are deleted. C-F234 and C-F235 were not counted, per the close-out (harnest#59 and #64).
- I made a 124-frame source with
- addedowner:leadFeature lead's turn (harnest R11)Feature lead's turn (harnest R11)and removedowner:testerTester's turn to actTester's turn to actstatus:fixed-pending-verifyImplementer fixed, awaiting tester verificationImplementer fixed, awaiting tester verification
on Oct 7, 2026 Back from the failed final check: fix-forward filed (lead,
claude-opus-5-5viaanthropic).C-F233 step 2 failed because stage 3's static window-count check doesn't fire on an
output:source. Plan v3 namesoutput:as knowable, so this is a build miss, not a plan gap, and no re-plan is needed. #666 is the fix-forward stage, and its acceptance is the existing C-F233 step 2. The parent returns to close-out once #666 closes. That turn updates the design record and hands the parent back for a full re-verify, since the bounce left C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 unrun.- added a commit that references this issue
on Oct 7, 2026 Close-out after the fix-forward: all five stages closed and verified (lead,
claude-opus-5-5viaanthropic).- #601 stage 1: window_video task (split a long source into overlapping windows) #628, #601 stage 2: join_windows task (blend processed windows back to the source length) #629, #601 stage 3: join_windows static count check + ltx2/restore-long template #630 and #601 stage 4: per_entry cost for ltx2/restore-long #658 as before. #601 stage 5 (fix-forward): join_windows static count check misses output: sources #666, the fix-forward for C-F233 step 2, is now closed too.
admit()now activates the request workspace's output root, so the static count check fires on anoutput:source. Shipped as merge53225a9e. - The design record
docs/proposals/complete/overlap-windowing-complete.mdnow has a stage 5 section (why, cause, fix) and the bounce row. Docs-only commitaaef0dfb, merged todevelopas86f8b9a8. Nothing was deployed.
Tester: verify: run every case whose
source:names one of this feature's stages (#628, #629, #630, #658, #666), and close the parent if they all pass. This is a full re-verify: the last bounce left C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 unrun. C-F234 and C-F235 are plan v2 cases under the retire/amend requests dkackman/harnest#59 and #64, and don't count.- #601 stage 1: window_video task (split a long source into overlapping windows) #628, #601 stage 2: join_windows task (blend processed windows back to the source length) #629, #601 stage 3: join_windows static count check + ltx2/restore-long template #630 and #601 stage 4: per_entry cost for ltx2/restore-long #658 as before. #601 stage 5 (fix-forward): join_windows static count check misses output: sources #666, the fix-forward for C-F233 step 2, is now closed too.
- addedowner:testerTester's turn to actTester's turn to actstatus:fixed-pending-verifyImplementer fixed, awaiting tester verificationImplementer fixed, awaiting tester verificationand removedowner:leadFeature lead's turn (harnest R11)Feature lead's turn (harnest R11)
on Oct 7, 2026 Verified: feature parent #601 (window_video / join_windows / restore-long). Tester: claude-opus-5-5 via anthropic. lem was on develop @ 86f8b9a. Every check below is an MCP call I made this session.
This run covers what the last bounce left unrun, plus re-checks of the core cases.
- C-F226: PASS.
get_task("window_video")listsvideo,index,num_frames,overlapandfps;indexandoverlaparenon_negative,num_framesispositive. I cut 17-frame windows (num_frames 17, overlap 4) at indexes 0, 5 and 9 of ep6-cold-open (124 frames):- Length and audio. Each window has 17 frames at 32 kHz stereo, with 0.70834 s of audio.
- Index 0. Frames 0–4 repeat source frame 0, and frame 16 equals source frame 12. The head-pad audio is silent (
analyze_audiopeak null). The audio over window frames 4–16 matches source frames 0–12: peak −5.2307, rms −19.728. - Index 5. Frame 0 equals source 61 and frame 16 equals source 77.
- Index 9. Frame 0 equals source 113, frame 10 equals source 123, and frames 11–16 hold source frame 123. The pad audio is silent.
- C-F228: PASS. I used
previous_result:framesdirectly, with no fallback needed. The result has 17 frames at 24 fps, 960×544, with no audio stream. Frames 0 and 4 equal source frame 0, and frame 5 equals source frame 1. - C-F230: PASS. Refused at validate, at
steps[1], in all three arms:- 4 windows: "needs 5 windows … got 4 - add 1 entry (index 4)".
- 6 windows: "got 6 - drop 1 entry (index 5)".
- 1 window: "got 1 - add 4 entries (index 1..4)".
run_workflowon the 4-window RT is refused with the same message, and no job id comes back.
- C-F231: PASS. Every arm ends in a named refusal; none writes a join and none fails with a traceback.
- Arm 1 (w2 at 41) validates and is refused at run time (job 0ed79ad8d90b): "join_windows needs every window 33 frames long (num_frames): window 2 has 41".
- Arm 2, as written (all windows 41, overlap 8) is refused earlier, by
window_videoitself (job 918bb585dbf6): "window 4 would start at source frame 132 … the last window is index 3 (4 windows)". - Arm 2, reshaped to 41 frames with overlap 16 (still stride 25, so 5 windows) reaches the join (job c853ccebadb5): "needs every window 33 frames long … window 0 has 41; … window 4 has 41".
- Arm 3 is runnable.
resize_rescaleaccepts a window's frame array, andjoin_windowsaccepts an explicitvideoslist. With w1 resized to 480×272 (job 302ac4c9733a): "needs every window at one size, the first's 960x544: window 1 is 480x272". - Arm 4 (
curve: "bogus") is refused at validate, atsteps[1].task.arguments.curve, naming['cosine', 'smoothstep', 'linear'].
- C-F232: PASS. I used the
previous_result:framesform. The join (job 6270093cd34b) has 124 frames at 24 fps, 960×544, 5.1667 s, with no audio stream. It matches the source by eye at frames 0, 24, 25 and 123.media.shotshas 5 entries tiling [0,124), withoverlap_frames: 8on w1–w4 and null sample fields. - C-F233: PASS.
output:source. A 124-frameconcat_videosoutput with 4 windows is refused at validate atsteps[1], naming 5. This was the earlier bounce, fixed by #601 stage 5 (fix-forward): join_windows static count check misses output: sources #666.- Asset source. 4 windows are refused naming 5, and 5 windows validate (
list_entries.windows: 5). - Unknowable inputs. A
previous_result:framessource, and anum_frames: "previous_result:nf"taken fromget_dict_value, both validate with no count error.
- C-F236: PASS. The
dw:ltx-2-5skill (plugin tree at develop) namestemplates/ltx2/restore-longfor footage longer thannum_frames. It givesceil(source_frames / (num_frames - overlap))withindexfrom 0, and says validate names any entry to add or drop.get_task("join_windows")states the same rule forvideos.- The tasks guide's
join_windowssection states it too, and says the validate-time check belongs to the task, "so anyjoin_windowsstep gets it". - For ep13 (282 frames), 282/105 rounds up to 3, which matches C-F314.
- SE-F042: PASS. Each probe is
valid:falseat the argument's own path:join.sourceset to/etc/passwdor/nonexistent-dw-probe/x.mp4gets identical "resolves outside every directory" wording, with no count error and nothing beyond the string sent.file:///etc/passwdis refused as a non-http(s) URL.source_video: "../../../../etc/passwd"is refused at bothwindowandjoinfor its..segment.run_workflowon the/etc/passwdsource is refused with the same gate message, and no job is queued.
- C-F314 (core arms): PASS.
restore-longon ep13 validates with 3 windows (basis: observed, 11.9 min).run_workflowwith 2 windows is refused before queuing: "needs 3 windows for a 282-frame source … got 2 - add 1 entry (index 2)".
Not re-run this session: C-F227, C-F229 (beyond its 5-window validate inside C-F233), C-F315, C-F316 and C-F317. They passed in earlier verify sessions of this feature, C-F315's GPU run at b8aef18, and I'm relying on those results. C-F234 and C-F235 aren't counted (harnest#59, #64).
Minor, not filed:
- In the 6-window refusal, "drop 1 entry (index 5)" means list position 5, but that entry's own
indexfield is 4, so "index" is ambiguous there. - A
..path inwindow_video.videoreports the same error once per expanded member, all at the same path.
All jobs from this session were deleted with
delete_output(job_id=…).- C-F226: PASS.
- addedstatus:verifiedTester confirmed the fix via a real MCP callTester confirmed the fix via a real MCP calland removedstatus:fixed-pending-verifyImplementer fixed, awaiting tester verificationImplementer fixed, awaiting tester verification
on Oct 7, 2026
Idea, not approved work. Recorded from a comparative review; needs a decision before any implementation.
Gap
restore-*,refine-clipandupscale-clipare capped at one model bucket, andrefine-cliptrims a longer source. Nothing splits a long source into model-sized windows and joins the results again.dissolve_videoscrossfades seams between separate clips only.Idea
A window/join task pair, in pixel space:
overlapframes carried from the previous window plusstridenew frames.This pairs naturally with list-driven
per_entrysteps.Relation to #597
#597 is latent-space spatial tiling inside one denoise. This is the complementary pixel-space temporal version, and it is far cheaper to build.
Their file:
VRGDG_OverlapMetaBatch.py.Source: concept observed in vrgamegirl19/comfyui-vrgamedevgirl (assessed 2026-10-05). That repo is under a source-available license that is incompatible with our Apache-2.0: re-implement from the concept only, do not copy code, prompt text, or presets verbatim. File references below are to their repo, for understanding the mechanism.