Skip to content

idea: overlap windowing for long video sources (split, process, cosine-blend join) #601

Description

@dkackman

Idea, not approved work. Recorded from a comparative review; needs a decision before any implementation.

Gap

restore-*, refine-clip and upscale-clip are capped at one model bucket, and refine-clip trims a longer source. Nothing splits a long source into model-sized windows and joins the results again. dissolve_videos crossfades seams between separate clips only.

Idea

A window/join task pair, in pixel space:

  • Split: each window is the last overlap frames carried from the previous window plus stride new frames.
    • The first window's prefix is frame 0 repeated.
    • A short final window is padded with its last frame.
    • The window length sits on the model grid (LTX 8n+1, H3 17n+5).
  • Audio: slice the audio per window, with silence for the synthetic frames.
  • Join: blend each window's head into the previous window's held tail with cosine, smoothstep or linear weights. Every source frame is emitted exactly once.

This pairs naturally with list-driven per_entry steps.

Relation to #597

#597 is latent-space spatial tiling inside one denoise. This is the complementary pixel-space temporal version, and it is far cheaper to build.

Their file: VRGDG_OverlapMetaBatch.py.

Source: concept observed in vrgamegirl19/comfyui-vrgamedevgirl (assessed 2026-10-05). That repo is under a source-available license that is incompatible with our Apache-2.0: re-implement from the concept only, do not copy code, prompt text, or presets verbatim. File references below are to their repo, for understanding the mechanism.

Activity

  1. added
    enhancementNew feature or request
    owner:donParked for human decision
    ideaAn idea for the researcher agent to assess
    owner:leadFeature lead's turn (harnest R11)
    priority:2Backlog rank: solid ROI, moderate scope
    featureWork bigger than a fix: designed with Don, built in stages (harnest R11)
    and removed
    owner:donParked for human decision
    ideaAn idea for the researcher agent to assess
    on Oct 5, 2026
  2. dkackman commented on Oct 5, 2026

    @dkackman
    OwnerAuthor

    Approved by Don (2026-10-05) at priority:2. Build it as a split/join task pair in pixel space, composed with per_entry steps. The window grid comes from the workflow's variable_constraints. Handing to the feature lead for a plan.

  3. dkackman commented on Oct 5, 2026

    @dkackman
    OwnerAuthor

    Plan v3: overlap windowing for long video sources (window_video + join_windows)

    Feature plan by the lead agent, model claude-opus-5-5 via anthropic, 2026-10-05. Design read against develop @ 21cd845. v2 records Don's approval answers (Q1–Q3) and his change to the stage 3 preflight, re-read against develop @ 837fd24. v3 (2026-10-06, from stage 3's build, #630): the template ships unpriced, a new stage 4 records the measured per_entry, and the plugin version bump is dropped.

    Verdict: build smaller

    Axis Finding
    Value Don approved the direction on 2026-10-05. There is no field-report demand yet: no driving session has reported a source that is too long for restore-*, refine-clip or upscale-clip. The gain is real but prospective. Today a source over one LTX bucket (121 frames ≈ 5 s) can't be restored or refined at all; refine-clip trims it.
    Build cost 4 stages, about $4 + $5 + $4 + $1 ≈ $14.
    Inertia Two utility tasks and one template. MCP, REST and UI enumerate tasks dynamically (dw/server/routes/system.py:86, dw_mcp/tools_catalog.py:135, ui/src/lib/api.ts:185). There is no new MCP tool, no schema change and no surface-budget cost (tests/test_mcp_server.py:1555 is unaffected). Upkeep is two AudioVideo construction sites (tests/test_shots.py:73), argument domains, a TASKS.md section, an ARCHITECTURE.md row and a seam/sync contract for assess_output. No security boundary changes and no new concept: it composes for_each/gather:.
    Reversibility High. No stored data depends on it. Removing it later would break only the template and any saved workflow that names the tasks.
    Cheaper alternatives (a) Do nothing: an agent splits by hand, with slice_audio-style trims per run and dissolve_videos across the results. But no task cuts a video at a start frame, and dissolve_videos' linear ramp shortens the cut by the overlap, so frames are lost. (b) Join only: gives nothing without the split.
    Overlap #597 (latent spatial tiling) is parked and independent: neither blocks the other. #600 (music-video timing) doesn't touch this.

    What "smaller" cuts relative to the idea text:

    • The H3 17n+5 grid. Every long-source consumer (restore-*, refine-clip, upscale-clip) is LTX. The tasks are grid-agnostic, so H3 needs no task work, only a template. Brings it back: an H3 video-to-video template that needs it.
    • Per-window audio blending in the join. The join puts the source audio back, sliced exactly to the source (Q2, decided). Brings it back: a windowed template whose model generates audio worth keeping.
    • One template, not three. ltx2/restore-long comes first (Q1, decided). Brings the others back: a field report asking for a long refine or upscale.

    Problem and non-goals

    Problem. LTX video-to-video templates take at most one model bucket. Nothing cuts a long source into overlapping model-sized windows and stitches the processed windows back to the source's exact length.

    Non-goals:

    • Latent-space or spatial tiling (idea (parked): tiled latent fusion for tile-trained LTX-2.5 IC-LoRAs (Refine-Details, Restore, Deblur) #597).
    • Windows chosen at run time from the source's length. for_each expands before validation (dw/for_each.py:61-124), so the window list is a workflow variable (Q3, decided).
    • More than 32 windows (dw/for_each.py:43). At 121-frame windows with a 16-frame overlap that is 3,360 frames, about 140 s at 24 fps.
    • Copying anything from the cited repo: this is a clean re-implementation from the concept.
    • Template-specific logic in validation (Don, 2026-10-05). Every validate-time rule here is owned by a task.

    Design, with corrections to the idea

    Correction 1: "per_entry" is cost pricing, not iteration. In dw, per_entry only prices a list (dw/plan.py:248). Iteration is for_each over a variable: list, with item: and gather:. The design composes those.

    Correction 2: the split can't return a list of windows. A list result fans every later previous_result: out per item (dw/previous_results.py:56-123), and nothing can iterate a list made at run time. So window_video returns one window per call, picked by index. The workflow drives it with for_each over variable:windows, a list whose entries are {name, index}.

    Correction 3: the grid comes from the existing num_frames constraint, not a new frame-snap argument. The window length is num_frames, and stride = num_frames - overlap. The template binds num_frames to variable:num_frames, which the LTX constraint already holds to 8n+1 (dw/variable_constraints.py). This needs no new constraint plumbing and matches Don's "grid from variable_constraints".

    Correction 4: the join can't read the window plan from its inputs. Pipeline outputs don't carry their input's shot records (dw/result.py:650-711). So join_windows takes source (the same variable:source_video) and re-derives the plan from the source's frame count. That also gives it the source audio to re-attach.

    window_video(video, index, num_frames, overlap, fps=None) returns one AudioVideo

    • Range: window i covers source frames [i·stride − overlap, i·stride + stride).
    • Synthetic frames:
      • Frames before 0 are frame 0 repeated (the first window's prefix).
      • Frames past the end are the last frame repeated (the final window's pad).
    • Audio: when the source has audio, the window carries source samples f2s(a)…f2s(b) for its real span [a, b), using frames_to_samples (dw/task_domains.py:402).
    • Frame format: float [0,1] like loop_frames (dw/tasks/video_utils.py:163), because the LTX reference condition takes that form.
    • Refusals:
      • overlap >= num_frames;
      • overlap < 0;
      • index < 0;
      • an index whose window starts at or past the source's last frame (index · stride >= source_frames).
    • Not a probe (assessment=False).

    join_windows(videos, source, num_frames, overlap, curve="cosine", fps=None) returns one AudioVideo

    • Inputs: videos is gather:<step>, the processed windows in list order.
    • Count rule: the required window count is ceil(source_frames / (num_frames − overlap)). A different count is refused (Q3, decided); the refusal names both numbers and the list entries to add or drop. The rule has one home, a function in dw/task_domains.py beside dissolve_shortfalls, called by both the task at run time (stage 2) and its static check at validate time (stage 3).
    • Blend: it emits source_frames frames, each source frame exactly once.
      • Window 0 contributes its real frames, the ones after the synthetic prefix.
      • Each later window's first overlap frames are blended over the previous window's last overlap real frames with weight w(t), where t = (k+1)/(overlap+1). That is the same open ramp as _ramp (dw/tasks/dissolve_videos.py:259), so neither end is a hard copy.
      • w(t) is cosine (1−cos πt)/2, smoothstep 3t²−2t³, or linear t.
      • The final window's pad frames are dropped.
    • More refusals:
      • windows whose frame counts differ from num_frames, naming which window and its count;
      • windows of different sizes. They may differ from the source's size, since refine-clip doubles it.
    • Audio (Q2, decided): the source's audio over all source_frames. When the source has none, so does the output. The windows' own audio is discarded.
    • Shot records: one shot per window. Each is named from videos, its span covers the frames it owns, overlap_frames is set on every seam, and start_sample/num_samples are cumulative f2s. That way assess_output's seam probe treats the blends as dissolves (dw/tasks/assess.py:677), and the sync probe passes.
    • Catalog shape: added to _CUT_TASKS (dw/server/catalog_shape.py:61).

    join_windows' static window-count check (Don, 2026-10-05)

    A validate-time check owned by join_windows, in the shape of dw/dissolve_frame_errors.py (a module of its own, wired into validation_errors):

    • It applies to any join_windows step, never to a named template, so a later windowed template gets it with no extra code.
    • It fires when source names a knowable file (asset:/output:, or a literal path the run may read; resolve_probe_path + probe_metadata, header only), num_frames and overlap are literal after substitution, and videos is gather:<step> whose member count is known after for_each expansion.
    • It errors when len(gather list) ≠ ceil(source_frames / (num_frames − overlap)), naming the count needed, at the join_windows step's path, by calling the stage 2 rule.
    • Anything not knowable (a previous_result: source, a non-literal num_frames) is left to the run-time refusal. Silence there is correct, as dissolve_frame_errors documents.

    Surfaces touched

    These are what needs-approval exists for, decided here:

    • Engine: two new utility tasks, and one validate-time check owned by join_windows.
    • MCP: none new; they appear in list_tasks/get_task.
    • REST: none new.
    • Syntax: none new. They use for_each, item: and gather: as they are.
    • Catalog: one new template (stage 3).

    Stages

    Stages are built in order, and each one lands on develop on its own. Stages 1–2 are unreferenced surface until stage 3 wires them in.

    Stage 1: window_video (server)

    • Builds: the task as specified above, plus:
      • its TASK_ARGUMENT_DOMAINS entries (index and overlap non-negative, num_frames positive), checked at run time like slice_audio;
      • the AudioVideo site in tests/test_shots.py EXPECTED_SITES and in dw/shots.py's list.
    • Tests:
      • window spans for first, middle and last windows;
      • prefix and pad frames;
      • audio sample spans tiling exactly across windows, at 24 and 25 fps with 44.1k and 48k audio;
      • each refusal;
      • a source with no audio.
    • Docs: docs/TASKS.md (Video Processing) and an docs/ARCHITECTURE.md row.
    • Deploy: server restart.
    • Estimate: ~$4.
    • Acceptance intent (over MCP, with a short asset: clip, e.g. 50 frames, num_frames=17, overlap=4, stride 13):
      • get_task("window_video") describes it.
      • A one-step workflow with index=0 succeeds with exactly 17 frames.
        • Its first 4 frames match source frame 0 (get_output_frames).
        • Its audio is 17 frames long, and the first 4 frames' worth is silent.
      • index=3 (frames 35–51) has 17 frames, and its last 2 repeat source frame 49.
      • index=4 is refused, naming the source's frame count and the last valid index (3).
      • overlap=17 is refused, and so is a negative overlap.
      • The refusals arrive at validate_workflow where it can see the asset's length, or otherwise at run time with a clear error. The tester should accept either and note which.

    Stage 2: join_windows (server)

    • Builds: the task as specified above, including the curves, the count/size/length refusals, source-audio re-attachment, shot records with overlap_frames, the _CUT_TASKS entry, and the window-count rule as a shared function in dw/task_domains.py.
    • Tests:
    • Docs: TASKS.md and ARCHITECTURE.md.
    • Deploy: server restart.
    • Estimate: ~$5.
    • Acceptance intent:
      • A workflow that windows an asset clip with for_each (no model step) and joins with gather: returns a video with the source's exact frame count.
        • It is visually identical to the source at sampled frames.
        • Its audio length matches the source.
        • assess_output raises no sync-drift finding, and its seams read as dissolves.
      • A window list one entry short is refused with the required count (at run time in this stage).
      • Windows of mixed length are refused, naming the odd window.
      • curve="bogus" is refused.

    Stage 3: join_windows static count check, ltx2/restore-long template and skill note (server + plugin)

    • Builds:
      • The static window-count check owned by join_windows, as specified above: a module in the shape of dw/dissolve_frame_errors.py, wired into validation_errors, calling the stage 2 rule. No template-specific logic in validate.
      • workflows/templates/ltx2/restore-long.json. It holds:
      • A short note in the dw:ltx-2.5 skill on choosing windows for a long source.
    • Tests:
      • the catalog-structure tests pass on the template;
      • check unit tests on a plain join_windows workflow (not the template): short, exact and long lists; an unknowable source or a non-literal num_frames stays silent.
    • Docs: the template's own description, WORKFLOW_GUIDE.md's for_each example list, and the check in TASKS.md's join_windows section.
    • Deploy: server restart. No plugin version bump: the version is pinned to the engine's and moves only at release (tests/test_plugin_skills.py). The skill note reaches the tester through origin/develop.
    • Estimate: ~$4.
    • Acceptance intent:
      • list_workflows(shape=…) lists the template, and validate_workflow on it is clean. Its plan.estimate is unknown (no cost yet), and validate doesn't refuse on that.
      • The tester's verify comment records the minutes of one restore@wN step and the whole run's minutes, with the window count and device, for stage 4.
      • Against a ~10 s asset clip with the right windows list, the run succeeds with the source's frame count and audio length.
        • No visible seam at the window boundaries (get_output_frames around each seam).
        • assess_output raises no sync-drift finding.
      • The same run with one entry too few fails at validate_workflow, naming the needed count, before any GPU time.
      • A hand-written window_video/join_windows workflow (no model step, not the template) with a short list is also refused at validate_workflow: the check belongs to the task.
      • Verified on the LTX-capable server (CUDA, 24 GB).

    Stage 4: per_entry cost for ltx2/restore-long (catalog)

    • Builds: the template's cost entry from stage 3's acceptance run. minutes is the run's total, and per_entry is {variable: "windows", minutes: <one restore window plus its slice>, entries: 3}, measured as shots-batch's was (Add templates/minimax/shots-batch: list-driven H3 shot generation with no in-job assembly #352). If that run used a list other than the default 3, the total is the measured run's, and entries and the default list match it.
    • Tests: per_entry_problems and the catalog-structure tests pass.
    • Docs: none beyond the template.
    • Deploy: server restart.
    • Estimate: ~$1.
    • Acceptance intent: validate_workflow on restore-long with 2, 3 and 4 windows returns a plan.estimate with basis: per_entry that scales by window count.
    • Fallback: if stage 3's verify comment lacks the timings, Don runs the default template once on lem, and the figure comes from that job.

    Risks

    • Content drift between windows. A restore model may shift colour or detail per window, so a cosine blend over a few frames can show a slow "breathing". Mitigations are the default overlap (16) and the stage 3 visual check. If it shows, a follow-up could condition each window on the previous window's tail, which is a bigger feature.
    • Pipeline frame count ≠ num_frames. H3 snaps up, and LTX should match. join_windows refuses rather than guesses.
    • Hand-written window lists. An agent writes the windows list from the source length. The static check and the join's count refusal both name the right number.
    • Already-flagged sync probe rounding. _sample_span (dw/tasks/assess.py:255) and the fade_samples rounding at :727 don't follow shot records: dissolve_videos floors frame→sample, pair_audio/gain_audio round — re-pairing an episode shifts its shots by one sample #401's cumulative rule. If they trip stage 2's acceptance, the fix belongs to the implementer as its own bug, not to this plan.

    Decisions (Don, 2026-10-05)

    • Q1. Which template comes first? ltx2/restore-long, built on restore-deblur.
    • Q2. What audio does the join output? The source's audio, re-attached exactly.
    • Q3. Refused or derived window count? Refused, naming the needed count.
    • Stage 3 preflight: a static check owned by join_windows that states its own rule. There is no template-specific logic in validate.
  4. added
    owner:donParked for human decision
    status:plan-reviewA feature plan is posted and waiting for Don
    and removed
    owner:leadFeature lead's turn (harnest R11)
    on Oct 5, 2026
  5. dkackman commented on Oct 5, 2026

    @dkackman
    OwnerAuthor

    Plan v1 approved by Don (2026-10-05). Q1-Q3 defaults accepted:

    • ltx2/restore-long comes first;
    • the join re-attaches the source audio;
    • a wrong window count is refused, and the refusal names the needed count.

    One change to stage 3: the validate-time window-count preflight should be a static check owned by join_windows, which states its own rule: len(gather list) == ceil(source_frames / (num_frames - overlap)) when source is an asset with a known length. Don't add template-specific logic to validate. Any later windowed template then gets the check with no extra code.

  6. 40 remaining items

  7. dkackman commented on Oct 7, 2026

    @dkackman
    OwnerAuthor

    Verify of #601 (feature parent): FAILED on C-F233 step 2. Bouncing to owner:lead. Tester: claude-opus-5-5 via anthropic, lem on develop @ b8aef18.

    Failing case: C-F233, step 2 (output: source)

    The case says a wrong window count must be refused at validate_workflow when the source is an output: reference, naming the needed count of 5 ("An output: source is knowable"). It is not.

    1. I made a 124-frame source with run_workflow {"id":"cc","steps":[{"name":"cat","task":{"command":"concat_videos","arguments":{"videos":["asset:qa-cast/ep6-cold-open.mp4"]}},"result":{"content_type":"video/mp4","fps":24}}]}.
      • Job 1b73865a4a47 succeeded and wrote cc/20261007-030343-8f1db730/cc-cat.0-0.0.mp4, one shot of 124 frames.
    2. I ran validate_workflow on C-F229's RT (num_frames 33, overlap 8, cosine) with source_video: "output:cc/20261007-030343-8f1db730/cc-cat.0-0.0.mp4" and 4 windows (w0–w3, indexes 0–3).
      • Result: valid: true, no errors, no warnings, plan.list_entries.windows: 4.
    3. The reference does resolve at validate. The same workflow with output:cc/20261007-030343-8f1db730/nope.mp4 is refused at variables.source_video with "Output '…/nope.mp4' not found under …/outputs".
    4. run_workflow on the step 2 workflow queued job 17b72f533155. It ran all 4 windows, then failed at join with:
      • "join_windows needs 5 windows for a 124-frame source with num_frames 33 and overlap 8 (stride 25), got 4 - add 1 entry (index 4)".
      • So the run-time check works, but the static check misses an output: source.

    For contrast, the same RT with the asset:qa-cast/ep6-cold-open.mp4 source is refused at validate at steps[1] naming 5, for 4, 6 and 1 windows (observed this session). The static check seems to cover asset: sources only.

    Passed this session (all over MCP)

    • C-F314: pass.
      • Defaults validate.
      • A wrong count is refused naming the needed count: ep13 with 2, 4 or 1, and ep11 with 3.
      • ep11 with 5 is valid.
      • run_workflow with 2 windows was refused with no job queued.
    • C-F315: pass. Job 63da56c3f40e.
      • Output: 282 frames, 24 fps, 960×544, 44.1 kHz stereo, 11.75 s.
      • Shots: w0 [0, 89), w1 [89, 194) with overlap 16, w2 [194, 282) with overlap 16.
      • sync_drift: no findings, max offset 0.01 ms.
      • Both seams are dissolves with no findings.
      • Timings, recorded in regression-perf/C-F315.jsonl:
        • w0 3.04 min (cold)
        • w1 1.48 min
        • w2 1.51 min
        • whole job 11.18 min
      • Note: the join_windows task alone took about 5 min.
    • C-F316: pass (the observed clause).
      • per_entry is {windows, 1.56, 3}.
      • Estimates for 1–5 windows were 10.1, 11.2, 12.2, 13.2 and 14.3: strictly increasing, none unknown.
    • C-F317: pass.
      • per_entry 1.56 against the warm median of 1.495.
      • The 3-window estimate of 12.2 against the job's 11.18.
    • SE-F042: pass.
      • Every refusal is at the argument's own path, with the same wording for an existing and a nonexistent path. The .. and file:// probes are refused as well.
      • run_workflow on the /etc/passwd source was refused with no job.
      • Wording ambiguity, not a failure: the arms that probe only video, with source still the asset, also return the count error. That count comes from the legitimate asset source, so nothing about the probed path leaks.
    • C-F229 (functional part): RT with the cosine, smoothstep and linear curves all succeeded (jobs 4a1b0541190c, 02c2622da816, a0d24251bf7f).
      • Output: 124 frames, 24 fps, 960×544, 32 kHz stereo, 5.1667 s.
      • 5 shots, overlap_frames 8, last shot ending at sample 122667 + 42666 = 165333.
      • assess_output on the cosine join: no findings, with sync_drift and the seam rules applied.
      • Not done: the per-frame source comparison.
    • C-F231, arm 4: curve "bogus" is refused at steps[1].task.arguments.curve, naming cosine, smoothstep and linear.
    • C-F233:
      • Step 1 passes: 5 windows valid; 4, 6 and 1 refused at steps[1] naming 5.
      • Steps 3 and 4 pass: no count error for previous_result:frames or a non-literal num_frames.
      • Step 2 fails, as above.
    • C-F227 (refusal arms):
      • (17, 17), (17, 20), (17, −1 / index −1) and (0, 0) are each refused at validate, naming the argument.
      • (17, 4, idx10) is refused at run: "window 10 would start at source frame 130, but the source has 124 frames … the last window is index 9".
      • (35, 4, idx4) is refused at run, naming the last index 3.
      • Not done: the accepted arms 8–10.

    Not run this session

    C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 were stopped short on budget once a failing case had been confirmed. The next verify of the parent should run the full set again.

    All outputs from this session are deleted. C-F234 and C-F235 were not counted, per the close-out (harnest#59 and #64).

  8. added
    owner:leadFeature lead's turn (harnest R11)
    and removed on Oct 7, 2026
  9. dkackman commented on Oct 7, 2026

    @dkackman
    OwnerAuthor

    Back from the failed final check: fix-forward filed (lead, claude-opus-5-5 via anthropic).

    C-F233 step 2 failed because stage 3's static window-count check doesn't fire on an output: source. Plan v3 names output: as knowable, so this is a build miss, not a plan gap, and no re-plan is needed. #666 is the fix-forward stage, and its acceptance is the existing C-F233 step 2. The parent returns to close-out once #666 closes. That turn updates the design record and hands the parent back for a full re-verify, since the bounce left C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 unrun.

  10. dkackman commented on Oct 7, 2026

    @dkackman
    OwnerAuthor

    Close-out after the fix-forward: all five stages closed and verified (lead, claude-opus-5-5 via anthropic).

    Tester: verify: run every case whose source: names one of this feature's stages (#628, #629, #630, #658, #666), and close the parent if they all pass. This is a full re-verify: the last bounce left C-F226, C-F228, C-F230's run part, C-F231 arms 1–3, C-F232 and C-F236 unrun. C-F234 and C-F235 are plan v2 cases under the retire/amend requests dkackman/harnest#59 and #64, and don't count.

  11. added and removed
    owner:leadFeature lead's turn (harnest R11)
    on Oct 7, 2026
  12. dkackman commented on Oct 7, 2026

    @dkackman
    OwnerAuthor

    Verified: feature parent #601 (window_video / join_windows / restore-long). Tester: claude-opus-5-5 via anthropic. lem was on develop @ 86f8b9a. Every check below is an MCP call I made this session.

    This run covers what the last bounce left unrun, plus re-checks of the core cases.

    • C-F226: PASS. get_task("window_video") lists video, index, num_frames, overlap and fps; index and overlap are non_negative, num_frames is positive. I cut 17-frame windows (num_frames 17, overlap 4) at indexes 0, 5 and 9 of ep6-cold-open (124 frames):
      • Length and audio. Each window has 17 frames at 32 kHz stereo, with 0.70834 s of audio.
      • Index 0. Frames 0–4 repeat source frame 0, and frame 16 equals source frame 12. The head-pad audio is silent (analyze_audio peak null). The audio over window frames 4–16 matches source frames 0–12: peak −5.2307, rms −19.728.
      • Index 5. Frame 0 equals source 61 and frame 16 equals source 77.
      • Index 9. Frame 0 equals source 113, frame 10 equals source 123, and frames 11–16 hold source frame 123. The pad audio is silent.
    • C-F228: PASS. I used previous_result:frames directly, with no fallback needed. The result has 17 frames at 24 fps, 960×544, with no audio stream. Frames 0 and 4 equal source frame 0, and frame 5 equals source frame 1.
    • C-F230: PASS. Refused at validate, at steps[1], in all three arms:
      • 4 windows: "needs 5 windows … got 4 - add 1 entry (index 4)".
      • 6 windows: "got 6 - drop 1 entry (index 5)".
      • 1 window: "got 1 - add 4 entries (index 1..4)".
      • run_workflow on the 4-window RT is refused with the same message, and no job id comes back.
    • C-F231: PASS. Every arm ends in a named refusal; none writes a join and none fails with a traceback.
      • Arm 1 (w2 at 41) validates and is refused at run time (job 0ed79ad8d90b): "join_windows needs every window 33 frames long (num_frames): window 2 has 41".
      • Arm 2, as written (all windows 41, overlap 8) is refused earlier, by window_video itself (job 918bb585dbf6): "window 4 would start at source frame 132 … the last window is index 3 (4 windows)".
      • Arm 2, reshaped to 41 frames with overlap 16 (still stride 25, so 5 windows) reaches the join (job c853ccebadb5): "needs every window 33 frames long … window 0 has 41; … window 4 has 41".
      • Arm 3 is runnable. resize_rescale accepts a window's frame array, and join_windows accepts an explicit videos list. With w1 resized to 480×272 (job 302ac4c9733a): "needs every window at one size, the first's 960x544: window 1 is 480x272".
      • Arm 4 (curve: "bogus") is refused at validate, at steps[1].task.arguments.curve, naming ['cosine', 'smoothstep', 'linear'].
    • C-F232: PASS. I used the previous_result:frames form. The join (job 6270093cd34b) has 124 frames at 24 fps, 960×544, 5.1667 s, with no audio stream. It matches the source by eye at frames 0, 24, 25 and 123. media.shots has 5 entries tiling [0,124), with overlap_frames: 8 on w1–w4 and null sample fields.
    • C-F233: PASS.
      • output: source. A 124-frame concat_videos output with 4 windows is refused at validate at steps[1], naming 5. This was the earlier bounce, fixed by #601 stage 5 (fix-forward): join_windows static count check misses output: sources #666.
      • Asset source. 4 windows are refused naming 5, and 5 windows validate (list_entries.windows: 5).
      • Unknowable inputs. A previous_result:frames source, and a num_frames: "previous_result:nf" taken from get_dict_value, both validate with no count error.
    • C-F236: PASS. The dw:ltx-2-5 skill (plugin tree at develop) names templates/ltx2/restore-long for footage longer than num_frames. It gives ceil(source_frames / (num_frames - overlap)) with index from 0, and says validate names any entry to add or drop.
      • get_task("join_windows") states the same rule for videos.
      • The tasks guide's join_windows section states it too, and says the validate-time check belongs to the task, "so any join_windows step gets it".
      • For ep13 (282 frames), 282/105 rounds up to 3, which matches C-F314.
    • SE-F042: PASS. Each probe is valid:false at the argument's own path:
      • join.source set to /etc/passwd or /nonexistent-dw-probe/x.mp4 gets identical "resolves outside every directory" wording, with no count error and nothing beyond the string sent.
      • file:///etc/passwd is refused as a non-http(s) URL.
      • source_video: "../../../../etc/passwd" is refused at both window and join for its .. segment.
      • run_workflow on the /etc/passwd source is refused with the same gate message, and no job is queued.
    • C-F314 (core arms): PASS. restore-long on ep13 validates with 3 windows (basis: observed, 11.9 min). run_workflow with 2 windows is refused before queuing: "needs 3 windows for a 282-frame source … got 2 - add 1 entry (index 2)".

    Not re-run this session: C-F227, C-F229 (beyond its 5-window validate inside C-F233), C-F315, C-F316 and C-F317. They passed in earlier verify sessions of this feature, C-F315's GPU run at b8aef18, and I'm relying on those results. C-F234 and C-F235 aren't counted (harnest#59, #64).

    Minor, not filed:

    • In the 6-window refusal, "drop 1 entry (index 5)" means list position 5, but that entry's own index field is 4, so "index" is ambiguous there.
    • A .. path in window_video.video reports the same error once per expanded member, all at the same path.

    All jobs from this session were deleted with delete_output(job_id=…).

  13. added
    status:verifiedTester confirmed the fix via a real MCP call
    and removed on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    backend:sharedenhancementNew feature or requestfeatureWork bigger than a fix: designed with Don, built in stages (harnest R11)owner:testerTester's turn to actpriority:2Backlog rank: solid ROI, moderate scopestatus:plan-approvedDon approved the feature plan (Don only)status:verifiedTester confirmed the fix via a real MCP call

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions