From d6443625f5a2f91ce5495371194d4499dfbc5e06 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 12:56:28 -0500 Subject: [PATCH 001/181] Give every run a version, so one of several can be named A file's name is per step, not per run, so four runs of one workflow write four files called AcornWarsCutAndScore-film.7-0.0.mp4 and the gallery drew four identical captions. The run id told them apart but is not something anyone says out loud, so there was no way for an agent to point a person at one of them and no way for the person to find it. Every run now takes an ordinal. assign_run_version runs when Workflow.run opens the run directory and records it as `version` in manifest.json; run_versions reads it back. Assigned once and never recomputed, which is the point - deleting a middle run leaves a gap rather than sliding every later number down, so a number quoted today still means the same run tomorrow. Assignment is max(recorded) + 1 over every sibling manifest rather than one past the newest: run ids are chronological only to the second, and within one second the spec digest decides the sort, which is exactly what three quick reruns hit. A run with no recorded number - made before this field, or killed before its manifest landed - is backfilled by rank, the unrecorded runs older than every recorded one taking the numbers beneath the lowest. GET /api/gallery and the metadata route carry `version` and `run_id` (run_versions read once per identity per listing, not per file). list_gallery teaches the vocabulary over MCP, which costs 44 tokens the surface budget now takes deliberately. The web UI reads the field only: a v4 chip ahead of the caption, the run id in the detail pane. Nothing on disk is renamed, so `output:` references, the step cache and keep_output are untouched. Two limits taken deliberately: deleting the newest run frees its number for reuse, and the flat layout has no runs to number. Co-Authored-By: Claude Opus 5 (1M context) --- CLAUDE.md | 25 +++++++ docs/MCP.md | 2 +- dw/runs.py | 104 +++++++++++++++++++++++++++ dw/server/app.py | 40 +++++++++-- dw/workflow.py | 22 +++++- dw_mcp/catalog.py | 8 ++- dw_mcp/server.py | 5 ++ tests/test_events.py | 25 +++++++ tests/test_mcp_server.py | 8 ++- tests/test_runs.py | 104 +++++++++++++++++++++++++++ tests/test_server.py | 73 +++++++++++++++++++ ui/CLAUDE.md | 7 ++ ui/src/lib/pages/GalleryPage.svelte | 33 ++++++++- ui/src/lib/pages/GalleryPage.test.ts | 45 ++++++++++++ ui/src/lib/proofs.test.ts | 2 + ui/src/lib/types.ts | 7 ++ 16 files changed, 501 insertions(+), 9 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index a8b7efef..337c4d7f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -368,6 +368,31 @@ same reason - default setup cannot load a pack. `JobManager.realized` finds the file. `exports` is a reserved workspace name: `POST /api/jobs/{id}/export` gathers one finished job into `/exports//` and `GET /exports/.zip` streams it. +- **A run has a number, and it is not derived from the listing** - a file's name + is per *step*, so four runs of one workflow write four files called + `AcornWarsCutAndScore-film.7-0.0.mp4` and the gallery drew four identical + captions: the run id told them apart but is not something anyone says out + loud, so an agent had no way to name one of them to a person. Every run now + takes an ordinal, `assign_run_version` (`dw/runs.py`) at the moment + `Workflow.run` opens the run directory, recorded as `version` in + `manifest.json` and read back by `run_versions`. Assigned once and never + recomputed, which is the point: deleting a middle run leaves a gap rather + than sliding every later number down, so "version 5" still means the same + run tomorrow. Assignment is `max(recorded) + 1` over *every* sibling + manifest, not one past the newest - run ids are chronological only to the + second, and within one second the spec digest decides the sort, which is + exactly what three quick reruns hit. A run with no recorded number (made + before the field, or killed before its manifest landed) is backfilled by + rank: the unrecorded runs older than every recorded one take the numbers + beneath the lowest, later ones continue from the run before. `GET + /api/gallery` and the metadata route carry `version` and `run_id` + (`run_versions` read once per identity per listing, not per file), MCP + `list_gallery` teaches the vocabulary, and the web UI reads the field only - + a `v4` chip on the card, the run id in the detail pane. Nothing on disk is + renamed, so `output:` references, the step cache and `keep_output` are + untouched. Two limits taken deliberately: deleting the *newest* run frees + its number for reuse (the high-water mark lived in the manifest that went + with it), and the flat layout has no runs, so `version` is null there. - **Result subfolders**: a step's `result.subfolder` (`dw/subfolders.py`) puts its files in a subfolder of the run directory - `/final/x.mp4` - by convention `final` or `intermediate`; the engine treats no name specially and there is no default. diff --git a/docs/MCP.md b/docs/MCP.md index a9ee2f1f..58f7762d 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -226,7 +226,7 @@ when no single workflow covers it. | `get_health()` | — | Check that the server is alive, and which machine answered: `version`, `device`, whether a model process is currently resident (`worker_alive`), the job running now and the queue depth. `worker_alive: false` is the normal idle state on a server that has not run a job since startup or the last memory clear - not a fault - the on-demand worker starts with the next job (#206) | | `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in | | `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds | -| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down, and it is `null` under the flat output layout, which has no runs. Tools take `name`, never `version`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | ### Media diff --git a/dw/runs.py b/dw/runs.py index ac42818b..7fa11c46 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -345,6 +345,110 @@ def strip_run_id(relative_path): return split_run_path(relative_path)[0] +# The key a run's ordinal is recorded under in its manifest. It is assigned +# once, when the run directory is opened, and never recomputed - which is the +# whole point: a number quoted in conversation has to still mean the same run +# after a sibling is deleted. Deleting a middle run leaves a gap +RUN_VERSION_KEY = "version" + + +def _recorded_version(run_dir): + """The ordinal a run recorded for itself, or None. + + None covers every way the number can be missing: a run made before this + field existed, one killed before its manifest landed, and one whose + manifest cannot be parsed. All three are ranked rather than trusted. + """ + try: + with open(os.path.join(run_dir, MANIFEST_FILE_NAME)) as file: + manifest = json.load(file) + except (OSError, ValueError): + return None + version = manifest.get(RUN_VERSION_KEY) if isinstance(manifest, dict) else None + return version if isinstance(version, int) and version > 0 else None + + +def _run_ids(identity_dir): + """Every run directory under one workflow identity, oldest first. + + Run ids sort by their UTC timestamp, so lexical order is chronological + to the second - the same property `latest` relies on. Within one second + the spec digest decides, which is arbitrary but stable; nothing here + needs finer ordering than that. + """ + try: + entries = os.listdir(identity_dir) + except OSError: + return [] + return sorted( + name + for name in entries + if is_run_id(name) and os.path.isdir(os.path.join(identity_dir, name)) + ) + + +def run_versions(identity_dir): + """Every run of one workflow mapped to its ordinal: {run id: version}. + + A run that recorded a version keeps it verbatim - that is what makes the + number survive a sibling being deleted. A run that recorded none (made + before the field existed, or killed before its manifest landed) is + ranked into the sequence around it: the unrecorded runs *older* than + every recorded one take the numbers just beneath the lowest recorded + one, so history that predates the field lands where it belongs, and an + unrecorded run anywhere later simply continues from the run before it. + Ordering is by run id, which is chronological. + """ + run_ids = _run_ids(identity_dir) + recorded = { + run_id: _recorded_version(os.path.join(identity_dir, run_id)) + for run_id in run_ids + } + # Room beneath the lowest recorded number for the unrecorded runs that + # come before it. Where there is not enough room the sequence starts at + # 1 and the recorded numbers stand: a duplicate is better than + # renumbering a run someone has already been told the number of + leading = 0 + for run_id in run_ids: + if recorded[run_id] is not None: + break + leading += 1 + next_number = 1 + if leading < len(run_ids): + next_number = max(1, recorded[run_ids[leading]] - leading) + versions = {} + for run_id in run_ids: + if recorded[run_id] is not None: + versions[run_id] = recorded[run_id] + next_number = recorded[run_id] + 1 + else: + versions[run_id] = next_number + next_number += 1 + return versions + + +def assign_run_version(output_dir, identity): + """The ordinal the run about to open under `identity` takes. + + One past the highest ordinal any sibling holds - not one past the newest + run's, because run ids are chronological only across seconds: two runs + started in the same second are ordered by their spec digest, so the last + id is not reliably the highest number. Three quick reruns are exactly + that case. + + It reads every sibling manifest, which is a small JSON file per run of + one workflow, once, against a run measured in minutes. Sharing + `run_versions` rather than deriving the maximum separately is what keeps + the number assigned here and the number the gallery reports from + drifting apart. + + Best effort, like everything else that writes a run's bookkeeping: a + directory that cannot be read yields 1 rather than failing the run. + """ + versions = run_versions(os.path.join(output_dir, identity)) + return max(versions.values(), default=0) + 1 + + def run_directory(output_dir, file_spec, workflow_id, run_id): """Where one execution writes: //. diff --git a/dw/server/app.py b/dw/server/app.py index a384c6a5..19511e1f 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -99,6 +99,7 @@ is_output_reference, is_run_id, resolve_output_reference, + run_versions, split_run_path, ) from ..workspace import ( @@ -2701,13 +2702,14 @@ def _iter_gallery_files(root, group_runs=True): continue relative_name = name if not directory else f"{directory}/{name}" if group_runs: - folder, _run_id, subfolder = split_run_path(relative_name) + folder, run_id, subfolder = split_run_path(relative_name) else: - folder, subfolder = directory, "" + folder, subfolder, run_id = directory, "", "" yield ( relative_name, folder, subfolder, + run_id, kind, os.path.join(current, name), ) @@ -2718,7 +2720,19 @@ def _gallery_entries(root, ws): files = list(_iter_gallery_files(root)) except OSError: files = [] - for relative_name, folder, subfolder, kind, path in files: + # One read of each workflow's run ordinals per listing, not per file: + # a run of fifty files would otherwise re-read the same manifests + # fifty times + versions_by_folder = {} + + def _version(folder, run_id): + if not run_id: + return None + if folder not in versions_by_folder: + versions_by_folder[folder] = run_versions(os.path.join(root, folder)) + return versions_by_folder[folder].get(run_id) + + for relative_name, folder, subfolder, run_id, kind, path in files: try: stat = os.stat(path) except OSError: @@ -2736,6 +2750,14 @@ def _gallery_entries(root, ws): "name": relative_name, "folder": folder, "subfolder": subfolder, + # Which run wrote it, and that run's ordinal among this + # workflow's runs - the 'v4' a person sees in the grid + # and an agent says out loud. Two runs write the same + # basename, so `label` cannot tell them apart and + # `name` is too long to quote. None under the flat + # layout, which has no runs to number + "run_id": run_id, + "version": _version(folder, run_id), # Quoted (slashes kept literal): a name carrying '#', '?' # or '%' would otherwise break the src the gallery # renders it into. The mtime still rides along for cache @@ -2884,6 +2906,7 @@ def gallery_metadata( the only way to read a wav's length was to run a job that copied it into the output directory. `job` is null for an asset (nothing here produced it) and `source` says which of the two roots answered.""" + run_id, version = "", None if is_asset_reference(name): path = _asset_file(name, ws) source, job = "asset", None @@ -2897,6 +2920,13 @@ def gallery_metadata( job = manager.history.job_for_file(name, workspace=ws.name) except Exception: job = None + # Which run wrote it, and that run's ordinal - the same 'v4' the + # listing reports. After "look at version 3" this is the next + # call, so it confirms the right file was reached rather than + # sending the caller back to the listing + folder, run_id, _subfolder = split_run_path(name) + if run_id: + version = run_versions(os.path.join(ws.outputs, folder)).get(run_id) metadata = read_embedded_metadata(path) extension = os.path.splitext(path)[1].lower() media = ( @@ -2909,6 +2939,8 @@ def gallery_metadata( "source": source, "metadata": metadata, "job": job, + "run_id": run_id, + "version": version, "media": media, } @@ -3597,7 +3629,7 @@ def list_assets(ws: Workspace = Depends(selected_workspace)): except OSError: files = [] origin = _asset_origin(ws, root) - for relative, folder, _subfolder, kind, path in files: + for relative, folder, _subfolder, _run_id, kind, path in files: try: stat = os.stat(path) except OSError: diff --git a/dw/workflow.py b/dw/workflow.py index 7077e7f1..4a99c34e 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -58,6 +58,7 @@ FLAT_LAYOUT, REALIZED_FILE_NAME, activate_output_root, + assign_run_version, deactivate_output_root, workflow_identity, manifest_relative_files, @@ -315,6 +316,11 @@ class Workflow: # leaves no manifest of its own, since its steps are already rolled up # into the parent's _run_dir_inherited = False + # That directory's ordinal among this workflow's runs - what the gallery + # shows as 'v4'. None in the flat layout, for a sub-workflow (which is + # part of the parent's run, not a run of its own), and before a run + # starts + _run_version = None # Where the parent step that delegated to this workflow sits in the # run the caller queued: {"step", "index", "total_steps"}. A child # counts its own steps from zero, so without this a composed run @@ -1113,7 +1119,18 @@ def run( self._run_dir = run_directory( self.output_dir, self.file_spec, workflow_id, run_id ) - logger.debug(f"Run directory: {self._run_dir}") + # The run's ordinal among this workflow's runs, taken + # once here and carried into the manifest. Assigning it + # at run time rather than deriving it when the gallery + # asks is what lets a sibling be deleted without + # renumbering the runs that outlive it + self._run_version = assign_run_version( + self.output_dir, + workflow_identity(self.file_spec, workflow_id), + ) + logger.debug( + f"Run directory: {self._run_dir} (v{self._run_version})" + ) # The record of what actually ran, written before the first step # so a crash or a cancel still leaves it. A sub-workflow inherits @@ -1483,6 +1500,9 @@ def _write_run_manifest( self._run_dir, { "run_id": run_id, + # This run's ordinal among the workflow's runs - 'v4' in the + # gallery. Recorded, never recomputed + "version": self._run_version, "status": status, "started_at": started_at, "finished_at": datetime.now(timezone.utc).isoformat(), diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 8da6ede1..98555712 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -230,7 +230,13 @@ def list_gallery(client, limit=50, subfolder=None, only_orphans=False, workspace Each file entry also carries `label`, a bare display basename for a UI grid - it is not a valid reference on its own (two runs can write the same basename) and is not accepted by `get_gallery_metadata` or - `delete_output`. Pass `name` to those, not `label`.""" + `delete_output`. Pass `name` to those, not `label`. + + `version` is that run's ordinal among the workflow's runs, and `run_id` + the run it came from. The version is what to quote to a person - the web + UI labels the same file `v5` - and is stable: it is assigned when the + run opens and a deleted sibling leaves a gap rather than renumbering + what is left. Null under the flat output layout, which has no runs.""" params = {"limit": limit} if subfolder is not None: params["subfolder"] = subfolder diff --git a/dw_mcp/server.py b/dw_mcp/server.py index e00806c7..44e02089 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -410,6 +410,11 @@ def list_gallery( file over HTTP, already scoped to the right workspace; use it as given rather than composing one from the name. + Entries also carry `run_id` and `version`, that run's ordinal among + the workflow's runs - stable, never renumbered. Quote the version + to a person: the web UI labels the same file `v5`. Tools still take + `name`. + `only_orphans=True` inverts the call: instead of files, it returns run directories holding nothing but their own bookkeeping (manifest.json, workflow.json, job.json) as `runs`, each diff --git a/tests/test_events.py b/tests/test_events.py index 4931f41f..8985d425 100644 --- a/tests/test_events.py +++ b/tests/test_events.py @@ -461,3 +461,28 @@ def test_watchdog_event_carries_the_required_fields(): assert stall["phase"] == "saving" assert isinstance(stall["seconds_since_phase_start"], (int, float)) assert "message" in stall + + +def test_each_run_records_its_own_version(tmp_path): + """Consecutive runs of one workflow number themselves 1, 2, 3 - the + ordinal the gallery shows as 'v2' and an agent quotes.""" + + def mock_load(self, shared_components): + self.pipeline = FakePipeline() + + versions = [] + for _ in range(3): + workflow_def = _workflow_def() + workflow_def["steps"][0]["result"] = {"content_type": "image/png"} + workflow = Workflow(workflow_def, str(tmp_path), "test.json") + with patch.object(Pipeline, "load", mock_load): + with patch("dw.workflow.empty_device_cache"): + workflow.run({}, previous_pipelines={}) + # The run's own directory, not the one its files came from: a + # cached step reports the earlier run's files while still being a + # run of its own with its own number + run_dir = pathlib.Path(workflow._run_dir) + manifest = json.loads((run_dir / "manifest.json").read_text()) + versions.append(manifest["version"]) + + assert versions == [1, 2, 3] diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 86f034bd..b4a45d5e 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -1435,7 +1435,13 @@ def test_the_stated_tool_count_is_the_registered_one(): # downscale" without the old space-filling clause, which pushed descriptions # over budget first (13_820). Measured 2026-09-21 at 13_790.8 (9_025.0 / # 3_751.8 / 1_014.0). 9.2 tokens of headroom left. -SURFACE_BUDGET = 13_800 +# Run versions added four sentences to list_gallery teaching `version` and +# `run_id` - the handle for naming one of several runs that wrote the same +# basename, which is the one thing the surface could not say before. Written +# as tightly as it can be said and still 44 tokens over, so the budget takes +# them deliberately rather than the sentence being cut to nothing. Measured +# 2026-09-22 at 13_844.0 (9_078.0 / 3_752.0 / 1_014.0). 6 tokens of headroom. +SURFACE_BUDGET = 13_850 @pytest.mark.asyncio diff --git a/tests/test_runs.py b/tests/test_runs.py index 18713d9a..069b2538 100644 --- a/tests/test_runs.py +++ b/tests/test_runs.py @@ -10,6 +10,7 @@ from dw.runs import ( FLAT_LAYOUT, + assign_run_version, OUTPUT_LAYOUT_ENV_VAR, RUN_LAYOUT, is_run_id, @@ -17,6 +18,7 @@ new_run_id, output_layout, resolve_output_reference, + run_versions, split_run_path, strip_run_id, workflow_identity, @@ -733,3 +735,105 @@ def test_a_chain_spill_lands_in_the_steps_subfolder(self, tmp_path, fake_pipelin ) segments = sorted((tmp_path / "Gyre" / "run" / "final").glob("*.segment-*.mp4")) assert len(segments) == 2 + + +class TestRunVersions: + """A run's ordinal among the runs of its workflow - the number a person + sees as 'v4' in the gallery and an agent says out loud. Assigned once, + recorded in the manifest, and never renumbered when a sibling is + deleted.""" + + @staticmethod + def _run(identity_dir, run_id, version=None): + """A run directory holding a manifest, with or without a version.""" + run_dir = os.path.join(identity_dir, run_id) + os.makedirs(run_dir, exist_ok=True) + manifest = {"run_id": run_id, "status": "completed"} + if version is not None: + manifest["version"] = version + with open(os.path.join(run_dir, "manifest.json"), "w") as file: + json.dump(manifest, file) + return run_dir + + def test_the_first_run_of_a_workflow_is_version_one(self, tmp_path): + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 1 + + def test_the_next_run_takes_the_number_after_the_newest(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + self._run(str(identity), "20260902-120000-bbbbbbbb", version=2) + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 3 + + def test_a_deleted_middle_run_leaves_a_gap_rather_than_renumbering(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + self._run(str(identity), "20260903-120000-cccccccc", version=3) + versions = run_versions(str(identity)) + # v2 is gone; v3 is still v3, and the next run is v4 + assert versions == { + "20260901-120000-aaaaaaaa": 1, + "20260903-120000-cccccccc": 3, + } + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 4 + + def test_runs_started_in_the_same_second_still_number_upward(self, tmp_path): + # Run ids are chronological only across seconds - within one second + # the spec digest decides the sort, so the highest number is not + # necessarily the last id. Three quick reruns are exactly this case + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-dddddddd", version=1) + self._run(str(identity), "20260901-120000-aaaaaaaa", version=2) + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 3 + + def test_runs_predating_the_field_are_ranked_by_run_id(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260903-120000-cccccccc") + self._run(str(identity), "20260901-120000-aaaaaaaa") + self._run(str(identity), "20260902-120000-bbbbbbbb") + assert run_versions(str(identity)) == { + "20260901-120000-aaaaaaaa": 1, + "20260902-120000-bbbbbbbb": 2, + "20260903-120000-cccccccc": 3, + } + + def test_backfilled_runs_sit_below_the_lowest_recorded_number(self, tmp_path): + # Two runs made before the field existed, then one that records it. + # The recorded number is authoritative; the older two are ranked + # beneath it so nothing collides + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa") + self._run(str(identity), "20260902-120000-bbbbbbbb") + self._run(str(identity), "20260903-120000-cccccccc", version=3) + assert run_versions(str(identity)) == { + "20260901-120000-aaaaaaaa": 1, + "20260902-120000-bbbbbbbb": 2, + "20260903-120000-cccccccc": 3, + } + + def test_a_run_whose_manifest_never_landed_is_still_numbered(self, tmp_path): + # Killed hard enough to write nothing: the directory is there, the + # manifest is not, and it is ranked like any pre-field run + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + os.makedirs(str(identity / "20260902-120000-bbbbbbbb")) + assert run_versions(str(identity)) == { + "20260901-120000-aaaaaaaa": 1, + "20260902-120000-bbbbbbbb": 2, + } + + def test_a_directory_that_is_not_a_run_is_ignored(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + os.makedirs(str(identity / "not-a-run-id")) + assert run_versions(str(identity)) == {"20260901-120000-aaaaaaaa": 1} + + def test_an_unreadable_manifest_does_not_lose_the_run(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + run_dir = identity / "20260901-120000-aaaaaaaa" + os.makedirs(str(run_dir)) + with open(os.path.join(str(run_dir), "manifest.json"), "w") as file: + file.write("{ not json") + assert run_versions(str(identity)) == {"20260901-120000-aaaaaaaa": 1} + + def test_a_workflow_with_no_runs_yet_has_none(self, tmp_path): + assert run_versions(str(tmp_path / "never" / "ran")) == {} diff --git a/tests/test_server.py b/tests/test_server.py index 56f1b165..5dc08882 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -1993,6 +1993,79 @@ def test_gallery_reports_and_filters_by_subfolder(server, tmp_path): assert finals["subfolders"] == full["subfolders"] +def test_gallery_reports_each_run_version(server, tmp_path): + """Four runs of one workflow write the same basename, so the gallery + label alone cannot tell them apart. Every entry carries the run it came + from and that run's ordinal - 'v4' - which is the handle an agent quotes + and a person finds in the grid.""" + import json as _json + + from PIL import Image + + def _run(identity, run_id, version=None): + run = tmp_path / "outputs" / identity / run_id + (run / "final").mkdir(parents=True) + Image.new("RGB", (2, 2)).save(run / "final" / "film.7-0.0.png") + manifest = {"run_id": run_id} + if version is not None: + manifest["version"] = version + (run / "manifest.json").write_text(_json.dumps(manifest)) + return run + + with server(success_script) as client: + _run("acorn/cut", "20260901-120000-aaaaaaaa", version=1) + # v2 was deleted; v3 keeps its number rather than sliding down + _run("acorn/cut", "20260903-120000-cccccccc", version=3) + # a run from before the field existed is ranked, not dropped + _run("acorn/score", "20260902-120000-bbbbbbbb") + # the flat layout has no runs at all + (tmp_path / "outputs" / "ltx").mkdir() + Image.new("RGB", (2, 2)).save(tmp_path / "outputs" / "ltx" / "flat.png") + + by_name = {f["name"]: f for f in client.get("/api/gallery").json()["files"]} + + first = by_name["acorn/cut/20260901-120000-aaaaaaaa/final/film.7-0.0.png"] + assert first["version"] == 1 + assert first["run_id"] == "20260901-120000-aaaaaaaa" + third = by_name["acorn/cut/20260903-120000-cccccccc/final/film.7-0.0.png"] + assert third["version"] == 3 + # the two runs are indistinguishable by label alone - which is the + # whole reason the version is here + assert first["label"] == third["label"] == "film.7-0.0.png" + # numbering is per workflow identity, so another workflow's first + # run is its own v1 + assert ( + by_name["acorn/score/20260902-120000-bbbbbbbb/final/film.7-0.0.png"][ + "version" + ] + == 1 + ) + assert by_name["ltx/flat.png"]["version"] is None + assert by_name["ltx/flat.png"]["run_id"] == "" + + +def test_gallery_metadata_names_the_run_and_its_version(server, tmp_path): + """After "look at version 3", the next call is usually this one - so it + answers with the run and the ordinal rather than making the caller go + back to the listing to confirm it read the right file.""" + import json as _json + + from PIL import Image + + with server(success_script) as client: + run = tmp_path / "outputs" / "acorn/cut" / "20260903-120000-cccccccc" + (run / "final").mkdir(parents=True) + Image.new("RGB", (2, 2)).save(run / "final" / "film.7-0.0.png") + (run / "manifest.json").write_text(_json.dumps({"version": 3})) + + body = client.get( + "/api/gallery/acorn/cut/20260903-120000-cccccccc" + "/final/film.7-0.0.png/metadata" + ).json() + assert body["run_id"] == "20260903-120000-cccccccc" + assert body["version"] == 3 + + def test_gallery_only_orphans_lists_media_less_run_directories(server, tmp_path): """#170: a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything, has no diff --git a/ui/CLAUDE.md b/ui/CLAUDE.md index 58aceef6..979067f3 100644 --- a/ui/CLAUDE.md +++ b/ui/CLAUDE.md @@ -61,6 +61,13 @@ workflow that has never run gets no frame at all rather than a grey placeholder (a fresh workspace would otherwise be a wall of empty plates). Within a folder, workflows that have produced something sort first. +Stripping the run id is also what makes four runs of one workflow four +identical captions, since a file's name is per step rather than per run. +The gallery entry carries `version` - the run's ordinal, assigned by the +engine and never renumbered - and the grid draws it as a `v4` chip ahead +of the label, with the run id in the detail pane beside it. The UI reads +the field only; nothing here computes or orders a version. + Every picture in the app sits in the global `.frame` (app.css): the media fills it edge to edge, with no inner padding and no rounding of its own, which is what makes it read as a proof on a sheet rather than as another diff --git a/ui/src/lib/pages/GalleryPage.svelte b/ui/src/lib/pages/GalleryPage.svelte index e9f0b1ac..3c6be3e3 100644 --- a/ui/src/lib/pages/GalleryPage.svelte +++ b/ui/src/lib/pages/GalleryPage.svelte @@ -319,7 +319,15 @@ {:else} ♪ {file.label} {/if} - {file.label} + + {#if file.version} + v{file.version} + {/if}{file.label} {/snippet} @@ -352,6 +360,15 @@ class="muted" title="open the file itself in a new tab">open file + {#if selected.version} + + version {selected.version} · {selected.run_id} + {/if} {formatBytes(selected.size)} · {formatMtime(selected.mtime)} @@ -524,6 +541,20 @@ white-space: normal; word-break: break-all; } + /* The one part of the caption that must not be broken or clamped away: + with four runs writing the same name it is the only thing on the card + that differs. Inline-block so word-break: break-all cannot split 'v10' + across lines */ + .version { + display: inline-block; + margin-right: 0.35rem; + padding: 0 0.3rem; + border-radius: 0.2rem; + background: var(--chip, rgb(255 255 255 / 0.08)); + color: var(--fg); + font-weight: 600; + word-break: keep-all; + } .detail { position: sticky; bottom: 1rem; diff --git a/ui/src/lib/pages/GalleryPage.test.ts b/ui/src/lib/pages/GalleryPage.test.ts index 16c2f495..68a859d6 100644 --- a/ui/src/lib/pages/GalleryPage.test.ts +++ b/ui/src/lib/pages/GalleryPage.test.ts @@ -6,6 +6,7 @@ import { within, } from '@testing-library/svelte' import { afterEach, beforeEach, expect, it, vi } from 'vitest' +import { fireEvent } from '@testing-library/dom' // Hoisted above the imports so the static import of the component below - // itself hoisted - sees an initialized mock. Importing the component inside // the test instead would charge its (multi-second) compile to the test timeout @@ -18,6 +19,8 @@ const file = (name: string, subfolder = ''): GalleryFile => ({ name, folder: name.includes('/') ? name.split('/')[0] : '', subfolder, + run_id: '', + version: null, url: `/outputs/${name}`, kind: 'image', size: 1024, @@ -383,3 +386,45 @@ it('leaves the detail open when Escape answers a confirm dialog', async () => { screen.getByLabelText('delete this file from the output directory'), ).toBeTruthy() }) + +it('marks each file with the version of the run that wrote it', async () => { + // Two runs of one workflow write the same basename - the case where the + // label alone tells a person nothing about which is which + listing.files = [ + { ...file('acorn/r1/film.mp4'), label: 'film.mp4', version: 1 }, + { ...file('acorn/r2/film.mp4'), label: 'film.mp4', version: 4 }, + ] + render(GalleryPage) + await waitFor(() => expect(screen.getAllByText('film.mp4')).toHaveLength(2)) + expect(screen.getByText('v1')).toBeTruthy() + expect(screen.getByText('v4')).toBeTruthy() +}) + +it('shows no version for a flat-layout file, which belongs to no run', async () => { + // The flat layout has no runs to number, and a card must not read + // 'vnull' or 'vundefined' because of it + listing.files = [{ ...file('ltx/flat.png'), label: 'flat.png' }] + render(GalleryPage) + await waitFor(() => expect(screen.getByText('flat.png')).toBeTruthy()) + expect(screen.queryByText(/^v\S+$/)).toBeNull() +}) + +it('names the run and its version in the details of the selected file', async () => { + // "Look at version 4" ends here: the pane says which run it reached, so + // the number in the grid can be checked against the one quoted + listing.files = [ + { + ...file('acorn/r2/film.mp4'), + label: 'film.mp4', + run_id: 'r2', + version: 4, + }, + ] + render(GalleryPage) + await waitFor(() => expect(screen.getByText('film.mp4')).toBeTruthy()) + await fireEvent.click(screen.getByText('film.mp4')) + const detail = document.querySelector('.detail') as HTMLElement + // One span holding both forms, so the query is over its whole text + expect(detail.textContent).toContain('version 4') + expect(detail.querySelector('code')?.textContent).toBe('r2') +}) diff --git a/ui/src/lib/proofs.test.ts b/ui/src/lib/proofs.test.ts index 895d00c2..59029b4b 100644 --- a/ui/src/lib/proofs.test.ts +++ b/ui/src/lib/proofs.test.ts @@ -10,6 +10,8 @@ const file = ( name, folder, subfolder: '', + run_id: '', + version: null, url: '/' + name, kind, size: 1, diff --git a/ui/src/lib/types.ts b/ui/src/lib/types.ts index d1dabb7c..357fcbaf 100644 --- a/ui/src/lib/types.ts +++ b/ui/src/lib/types.ts @@ -262,6 +262,13 @@ export interface GalleryFile { /** What followed the run id in the file's path - the `final` / * `intermediate` a step's `result.subfolder` chose, `''` for none. */ subfolder: string + /** The run that wrote the file, `''` under the flat layout. */ + run_id: string + /** That run's ordinal among the workflow's runs - what the grid shows as + * `v4`. Two runs write the same `label`, so this is what tells them + * apart at a glance. Assigned when the run opens and never renumbered, + * so a deleted sibling leaves a gap. Null when there is no run. */ + version: number | null url: string kind: 'image' | 'video' | 'audio' size: number From 7240ede218af4fa30321dd6be2193d347b34e28d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 16:06:04 -0500 Subject: [PATCH 002/181] feat: implement run versioning and manifest updates for workflow executions --- CLAUDE.md | 17 +++- docs/MCP.md | 2 +- dw/runs.py | 146 ++++++++++++++++++++++----- dw/server/app.py | 7 ++ dw/workflow.py | 21 +++- dw_mcp/catalog.py | 4 +- tests/test_events.py | 30 ++++++ tests/test_runs.py | 79 +++++++++++++++ tests/test_server.py | 38 +++++++ ui/src/lib/pages/GalleryPage.svelte | 9 +- ui/src/lib/pages/GalleryPage.test.ts | 6 +- 11 files changed, 320 insertions(+), 39 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 337c4d7f..f6849a83 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -381,10 +381,19 @@ same reason - default setup cannot load a pack. run tomorrow. Assignment is `max(recorded) + 1` over *every* sibling manifest, not one past the newest - run ids are chronological only to the second, and within one second the spec digest decides the sort, which is - exactly what three quick reruns hit. A run with no recorded number (made - before the field, or killed before its manifest landed) is backfilled by - rank: the unrecorded runs older than every recorded one take the numbers - beneath the lowest, later ones continue from the run before. `GET + exactly what three quick reruns hit. The number is on disk from the moment + the run opens - a `status: "running"` manifest is written before the first + step and rewritten in full at the end - so a hard kill does not lose it and + a second process opening a run of the same workflow sees it. A run with no + recorded number (made before the field, or killed before even that first + manifest) is ranked: the unrecorded runs older than every recorded one take + the numbers beneath the lowest, later ones continue from the highest before + them. A ranked number would move when an older sibling is deleted, so + `record_run_versions` writes it into the manifest on the two write paths - + a run opening and a run directory being deleted; the listing never writes, + and a run with no manifest at all is left ranked. A gap in the numbers is + not only a deletion: a failed run or a fully cached rerun takes a number + and may have no media for the gallery to show under it. `GET /api/gallery` and the metadata route carry `version` and `run_id` (`run_versions` read once per identity per listing, not per file), MCP `list_gallery` teaches the vocabulary, and the web UI reads the field only - diff --git a/docs/MCP.md b/docs/MCP.md index 58f7762d..f06956cb 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -226,7 +226,7 @@ when no single workflow covers it. | `get_health()` | — | Check that the server is alive, and which machine answered: `version`, `device`, whether a model process is currently resident (`worker_alive`), the job running now and the queue depth. `worker_alive: false` is the normal idle state on a server that has not run a job since startup or the last memory clear - not a fault - the on-demand worker starts with the next job (#206) | | `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in | | `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds | -| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down, and it is `null` under the flat output layout, which has no runs. Tools take `name`, never `version`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. Tools take `name`, never `version`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | ### Media diff --git a/dw/runs.py b/dw/runs.py index 7fa11c46..797369ca 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -118,6 +118,7 @@ def _runs_newest_first(directory): for name in os.listdir(directory) if is_run_id(name) and os.path.isdir(os.path.join(directory, name)) ), + key=run_id_sort_key, reverse=True, ) except OSError: @@ -348,9 +349,44 @@ def strip_run_id(relative_path): # The key a run's ordinal is recorded under in its manifest. It is assigned # once, when the run directory is opened, and never recomputed - which is the # whole point: a number quoted in conversation has to still mean the same run -# after a sibling is deleted. Deleting a middle run leaves a gap +# after a sibling is deleted. Deleting a middle run leaves a gap, and so does +# a run that wrote no media (it failed, or every step was reused from the +# cache): it took a number and has nothing in the gallery to show under it RUN_VERSION_KEY = "version" +# The length of a run id before any '-N' counter a same-second rerun takes +_RUN_ID_BASE_LENGTH = len("20260101-000000-00000000") + +# Recorded ordinals by manifest path, keyed on the manifest's stat so an +# edited or replaced manifest is read again. A recorded number never changes, +# so this is what keeps a gallery listing from parsing every manifest under +# the output root on every call +_recorded_versions = {} + + +def run_id_sort_key(run_id): + """Order run ids oldest first, with a rerun's '-N' counter compared as a + number - lexically '-10' would sort before '-2'.""" + base, counter = run_id[:_RUN_ID_BASE_LENGTH], run_id[_RUN_ID_BASE_LENGTH + 1 :] + return (base, int(counter) if counter.isdigit() else 1) + + +def _read_manifest(run_dir): + """A run's manifest as a dict, or None when it is missing or unreadable.""" + try: + with open(os.path.join(run_dir, MANIFEST_FILE_NAME)) as file: + manifest = json.load(file) + except (OSError, ValueError): + return None + return manifest if isinstance(manifest, dict) else None + + +def _valid_version(version): + # bool is an int subclass, and True is not version 1 + if isinstance(version, bool) or not isinstance(version, int): + return None + return version if version > 0 else None + def _recorded_version(run_dir): """The ordinal a run recorded for itself, or None. @@ -359,19 +395,26 @@ def _recorded_version(run_dir): field existed, one killed before its manifest landed, and one whose manifest cannot be parsed. All three are ranked rather than trusted. """ + path = os.path.join(run_dir, MANIFEST_FILE_NAME) try: - with open(os.path.join(run_dir, MANIFEST_FILE_NAME)) as file: - manifest = json.load(file) - except (OSError, ValueError): + stat = os.stat(path) + except OSError: + _recorded_versions.pop(path, None) return None - version = manifest.get(RUN_VERSION_KEY) if isinstance(manifest, dict) else None - return version if isinstance(version, int) and version > 0 else None + signature = (stat.st_mtime_ns, stat.st_size, stat.st_ino) + cached = _recorded_versions.get(path) + if cached is not None and cached[0] == signature: + return cached[1] + manifest = _read_manifest(run_dir) + version = _valid_version(manifest.get(RUN_VERSION_KEY)) if manifest else None + _recorded_versions[path] = (signature, version) + return version def _run_ids(identity_dir): """Every run directory under one workflow identity, oldest first. - Run ids sort by their UTC timestamp, so lexical order is chronological + Run ids sort by their UTC timestamp, so this order is chronological to the second - the same property `latest` relies on. Within one second the spec digest decides, which is arbitrary but stable; nothing here needs finer ordering than that. @@ -381,24 +424,17 @@ def _run_ids(identity_dir): except OSError: return [] return sorted( - name - for name in entries - if is_run_id(name) and os.path.isdir(os.path.join(identity_dir, name)) + ( + name + for name in entries + if is_run_id(name) and os.path.isdir(os.path.join(identity_dir, name)) + ), + key=run_id_sort_key, ) -def run_versions(identity_dir): - """Every run of one workflow mapped to its ordinal: {run id: version}. - - A run that recorded a version keeps it verbatim - that is what makes the - number survive a sibling being deleted. A run that recorded none (made - before the field existed, or killed before its manifest landed) is - ranked into the sequence around it: the unrecorded runs *older* than - every recorded one take the numbers just beneath the lowest recorded - one, so history that predates the field lands where it belongs, and an - unrecorded run anywhere later simply continues from the run before it. - Ordering is by run id, which is chronological. - """ +def _ranked_versions(identity_dir): + """({run id: version}, {run id: recorded version or None}).""" run_ids = _run_ids(identity_dir) recorded = { run_id: _recorded_version(os.path.join(identity_dir, run_id)) @@ -420,10 +456,64 @@ def run_versions(identity_dir): for run_id in run_ids: if recorded[run_id] is not None: versions[run_id] = recorded[run_id] - next_number = recorded[run_id] + 1 + # Never backwards: two runs of one second can sort in the + # opposite order to their numbers, and an unrecorded run after + # them must not take a number the higher one already holds + next_number = max(next_number, recorded[run_id] + 1) else: versions[run_id] = next_number next_number += 1 + return versions, recorded + + +def run_versions(identity_dir): + """Every run of one workflow mapped to its ordinal: {run id: version}. + + A run that recorded a version keeps it verbatim - that is what makes the + number survive a sibling being deleted. A run that recorded none (made + before the field existed, or killed before its manifest landed) is + ranked into the sequence around it: the unrecorded runs *older* than + every recorded one take the numbers just beneath the lowest recorded + one, so history that predates the field lands where it belongs, and an + unrecorded run anywhere later continues from the highest number before + it. Ordering is by run id, which is chronological. + + Read only. A ranked number is only as stable as its neighbours until + `record_run_versions` writes it down. + """ + return _ranked_versions(identity_dir)[0] + + +def record_run_versions(identity_dir): + """Write each ranked number into the manifest of a run that has one but + records no version, and return every run's ordinal. + + A ranked number moves when an older unrecorded sibling is deleted, so + runs made before the field existed are pinned the first time anything + writes under their workflow: a new run opening, or a run directory + being deleted. The listing never writes. A run with no manifest at all + is left alone - writing one would invent a record of a run nobody + recorded - and stays ranked. + + Best effort: a manifest that cannot be rewritten keeps its ranked number. + """ + versions, recorded = _ranked_versions(identity_dir) + for run_id, version in versions.items(): + if recorded[run_id] is not None: + continue + run_dir = os.path.join(identity_dir, run_id) + manifest = _read_manifest(run_dir) + if manifest is None: + continue + manifest[RUN_VERSION_KEY] = version + path = os.path.join(run_dir, MANIFEST_FILE_NAME) + partial = f"{path}.partial" + try: + with open(partial, "w") as file: + json.dump(manifest, file, indent=2, default=str) + os.replace(partial, path) + except OSError as e: + logger.warning(f"Could not record version {version} in {path}: {e}") return versions @@ -436,16 +526,16 @@ def assign_run_version(output_dir, identity): id is not reliably the highest number. Three quick reruns are exactly that case. - It reads every sibling manifest, which is a small JSON file per run of - one workflow, once, against a run measured in minutes. Sharing - `run_versions` rather than deriving the maximum separately is what keeps - the number assigned here and the number the gallery reports from + Pins the ranked numbers of older runs on the way (`record_run_versions`), + so history that predates the field stops moving once a new run joins it. + Sharing that ranking rather than deriving the maximum separately is what + keeps the number assigned here and the number the gallery reports from drifting apart. Best effort, like everything else that writes a run's bookkeeping: a directory that cannot be read yields 1 rather than failing the run. """ - versions = run_versions(os.path.join(output_dir, identity)) + versions = record_run_versions(os.path.join(output_dir, identity)) return max(versions.values(), default=0) + 1 diff --git a/dw/server/app.py b/dw/server/app.py index 19511e1f..28ee5929 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -98,6 +98,7 @@ REALIZED_FILE_NAME, is_output_reference, is_run_id, + record_run_versions, resolve_output_reference, run_versions, split_run_path, @@ -3387,6 +3388,9 @@ def _prune_empty_run_directory(name, root): continue return None + # Pin the siblings' numbers first: a run that predates versions is + # ranked, and removing one ahead of it would renumber it + record_run_versions(os.path.dirname(run_dir)) shutil.rmtree(run_dir, ignore_errors=True) # And the identity folders above it, while they are empty - a swept # workspace should not keep one directory per workflow it once ran @@ -3429,6 +3433,9 @@ def delete_output(name: str, ws: Workspace = Depends(selected_workspace)): """ run_dir = _run_directory(name, ws.outputs) if run_dir is not None: + # As in _prune_empty_run_directory: pin the siblings' numbers + # before one of them goes + record_run_versions(os.path.dirname(run_dir)) shutil.rmtree(run_dir, ignore_errors=True) parent = os.path.dirname(run_dir) while os.path.normpath(parent) != os.path.normpath(ws.outputs): diff --git a/dw/workflow.py b/dw/workflow.py index 4a99c34e..0845237e 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -1152,6 +1152,20 @@ def run( # Never fatal: the record is worth less than the run logger.warning(f"Could not realize workflow {workflow_id}: {e}") + # A manifest now, rewritten in full when the run ends: the + # version held only in memory until then was lost to a hard + # kill, and a second process opening a run of this workflow + # meanwhile could not see it and took the same number + self._write_run_manifest( + run_id, + "running", + started_at, + arguments, + resolved_seed, + realized_name, + annotations, + ) + # Which run this is, so a server job can find the directory # it wrote. Emitted even when the realized file did not land: # the manifest is still there, and so are the files @@ -1505,7 +1519,12 @@ def _write_run_manifest( "version": self._run_version, "status": status, "started_at": started_at, - "finished_at": datetime.now(timezone.utc).isoformat(), + # None on the manifest written as the run opens + "finished_at": ( + None + if status == "running" + else datetime.now(timezone.utc).isoformat() + ), "dw_version": __version__, "device": str(get_device()), "workflow": { diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 98555712..4e1c70bb 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -236,7 +236,9 @@ def list_gallery(client, limit=50, subfolder=None, only_orphans=False, workspace the run it came from. The version is what to quote to a person - the web UI labels the same file `v5` - and is stable: it is assigned when the run opens and a deleted sibling leaves a gap rather than renumbering - what is left. Null under the flat output layout, which has no runs.""" + what is left - as does a run that failed, or reused every step from + the cache, and so wrote nothing to list. Null under the flat output + layout, which has no runs.""" params = {"limit": limit} if subfolder is not None: params["subfolder"] = subfolder diff --git a/tests/test_events.py b/tests/test_events.py index 8985d425..1a60baa8 100644 --- a/tests/test_events.py +++ b/tests/test_events.py @@ -486,3 +486,33 @@ def mock_load(self, shared_components): versions.append(manifest["version"]) assert versions == [1, 2, 3] + + +def test_the_version_is_on_disk_before_the_first_step_runs(tmp_path): + """A run killed mid-step - which is how a stuck server gets restarted - + never reaches the closing manifest, so the number has to land when the + run opens. Also what lets a second process opening a run of the same + workflow see this one's number rather than taking it too.""" + seen = {} + + def mock_load(self, shared_components): + manifest_path = pathlib.Path(workflow._run_dir) / "manifest.json" + seen.update(json.loads(manifest_path.read_text())) + self.pipeline = FakePipeline() + + workflow_def = _workflow_def() + workflow_def["steps"][0]["result"] = {"content_type": "image/png"} + workflow = Workflow(workflow_def, str(tmp_path), "test.json") + with patch.object(Pipeline, "load", mock_load): + with patch("dw.workflow.empty_device_cache"): + workflow.run({}, previous_pipelines={}) + + assert seen["version"] == 1 + assert seen["status"] == "running" + assert seen["finished_at"] is None + closing = json.loads( + (pathlib.Path(workflow._run_dir) / "manifest.json").read_text() + ) + assert closing["status"] == "completed" + assert closing["version"] == 1 + assert closing["finished_at"] is not None diff --git a/tests/test_runs.py b/tests/test_runs.py index 069b2538..8fe7f208 100644 --- a/tests/test_runs.py +++ b/tests/test_runs.py @@ -11,6 +11,7 @@ from dw.runs import ( FLAT_LAYOUT, assign_run_version, + record_run_versions, OUTPUT_LAYOUT_ENV_VAR, RUN_LAYOUT, is_run_id, @@ -837,3 +838,81 @@ def test_an_unreadable_manifest_does_not_lose_the_run(self, tmp_path): def test_a_workflow_with_no_runs_yet_has_none(self, tmp_path): assert run_versions(str(tmp_path / "never" / "ran")) == {} + + def test_a_run_after_an_out_of_order_second_takes_no_held_number(self, tmp_path): + # Two runs of one second sort opposite to their numbers (v6 then + # v5); a killed run after them must not be ranked back down onto 6 + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-00000000", version=6) + self._run(str(identity), "20260901-120000-ffffffff", version=5) + os.makedirs(str(identity / "20260901-120005-aaaaaaaa")) + assert run_versions(str(identity)) == { + "20260901-120000-00000000": 6, + "20260901-120000-ffffffff": 5, + "20260901-120005-aaaaaaaa": 7, + } + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 8 + + def test_a_rerun_counter_sorts_as_a_number(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + for run_id in ( + "20260901-120000-aaaaaaaa-10", + "20260901-120000-aaaaaaaa", + "20260901-120000-aaaaaaaa-2", + ): + self._run(str(identity), run_id) + assert list(run_versions(str(identity)).items()) == [ + ("20260901-120000-aaaaaaaa", 1), + ("20260901-120000-aaaaaaaa-2", 2), + ("20260901-120000-aaaaaaaa-10", 3), + ] + + def test_a_boolean_is_not_a_recorded_version(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=True) + self._run(str(identity), "20260902-120000-bbbbbbbb", version=5) + assert run_versions(str(identity))["20260901-120000-aaaaaaaa"] == 4 + + def test_an_edited_manifest_is_read_again(self, tmp_path): + # Recorded numbers are cached against the manifest's stat, so a + # rewrite - the run's closing manifest, or a backfill - is seen + identity = tmp_path / "ltx2" / "Gyre" + run_dir = self._run(str(identity), "20260901-120000-aaaaaaaa") + assert run_versions(str(identity)) == {"20260901-120000-aaaaaaaa": 1} + with open(os.path.join(run_dir, "manifest.json"), "w") as file: + json.dump({"version": 9, "padding": "changes the size"}, file) + assert run_versions(str(identity)) == {"20260901-120000-aaaaaaaa": 9} + + def test_opening_a_run_pins_the_numbers_of_older_runs(self, tmp_path): + # History from before the field is ranked, and a ranked number moves + # when an older sibling goes - until a new run writes it down + identity = tmp_path / "ltx2" / "Gyre" + for day in (1, 2, 3): + self._run(str(identity), f"2026090{day}-120000-aaaaaaaa") + assert assign_run_version(str(tmp_path), "ltx2/Gyre") == 4 + for day, version in ((1, 1), (2, 2), (3, 3)): + manifest_path = identity / f"2026090{day}-120000-aaaaaaaa" / "manifest.json" + manifest = json.loads(manifest_path.read_text()) + assert manifest["version"] == version + # the rest of the record is untouched + assert manifest["status"] == "completed" + # now deleting the oldest renumbers nothing + import shutil + + shutil.rmtree(str(identity / "20260901-120000-aaaaaaaa")) + assert run_versions(str(identity)) == { + "20260902-120000-aaaaaaaa": 2, + "20260903-120000-aaaaaaaa": 3, + } + + def test_pinning_leaves_a_run_with_no_manifest_alone(self, tmp_path): + # Writing one would invent a record of a run nobody recorded + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + killed = identity / "20260902-120000-bbbbbbbb" + os.makedirs(str(killed)) + assert record_run_versions(str(identity)) == { + "20260901-120000-aaaaaaaa": 1, + "20260902-120000-bbbbbbbb": 2, + } + assert not (killed / "manifest.json").exists() diff --git a/tests/test_server.py b/tests/test_server.py index 5dc08882..cdccac7c 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -2066,6 +2066,44 @@ def test_gallery_metadata_names_the_run_and_its_version(server, tmp_path): assert body["version"] == 3 +def test_deleting_an_older_run_renumbers_none_of_its_siblings(server, tmp_path): + """Runs from before versions existed are ranked, so removing the oldest + would slide every later one down a number. The delete pins the + siblings' numbers into their manifests first - both for a whole run + directory and for the last file of a run, which sweeps the directory.""" + import json as _json + + from PIL import Image + + identity = tmp_path / "outputs" / "acorn" / "cut" + run_ids = [f"2026090{day}-120000-aaaaaaaa" for day in (1, 2, 3, 4)] + with server(success_script) as client: + for run_id in run_ids: + (identity / run_id).mkdir(parents=True) + Image.new("RGB", (2, 2)).save(identity / run_id / "film.png") + (identity / run_id / "manifest.json").write_text( + _json.dumps({"run_id": run_id}) + ) + + def versions(): + return { + f["run_id"]: f["version"] + for f in client.get("/api/gallery").json()["files"] + } + + assert versions() == dict(zip(run_ids, (1, 2, 3, 4))) + # the whole run directory + assert client.delete(f"/api/gallery/acorn/cut/{run_ids[0]}").status_code == 200 + assert versions() == dict(zip(run_ids[1:], (2, 3, 4))) + # the last file of a run, which takes its directory with it + assert ( + client.delete(f"/api/gallery/acorn/cut/{run_ids[1]}/film.png").status_code + == 200 + ) + assert not (identity / run_ids[1]).exists() + assert versions() == dict(zip(run_ids[2:], (3, 4))) + + def test_gallery_only_orphans_lists_media_less_run_directories(server, tmp_path): """#170: a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything, has no diff --git a/ui/src/lib/pages/GalleryPage.svelte b/ui/src/lib/pages/GalleryPage.svelte index 3c6be3e3..32507caf 100644 --- a/ui/src/lib/pages/GalleryPage.svelte +++ b/ui/src/lib/pages/GalleryPage.svelte @@ -326,6 +326,10 @@ title="version {file.version} of this workflow" >v{file.version} + + {' '} {/if}{file.label} @@ -547,11 +551,10 @@ across lines */ .version { display: inline-block; - margin-right: 0.35rem; padding: 0 0.3rem; border-radius: 0.2rem; - background: var(--chip, rgb(255 255 255 / 0.08)); - color: var(--fg); + background: var(--line); + color: var(--ink); font-weight: 600; word-break: keep-all; } diff --git a/ui/src/lib/pages/GalleryPage.test.ts b/ui/src/lib/pages/GalleryPage.test.ts index 68a859d6..3d9d6339 100644 --- a/ui/src/lib/pages/GalleryPage.test.ts +++ b/ui/src/lib/pages/GalleryPage.test.ts @@ -1,12 +1,12 @@ import { cleanup, + fireEvent, render, screen, waitFor, within, } from '@testing-library/svelte' import { afterEach, beforeEach, expect, it, vi } from 'vitest' -import { fireEvent } from '@testing-library/dom' // Hoisted above the imports so the static import of the component below - // itself hoisted - sees an initialized mock. Importing the component inside // the test instead would charge its (multi-second) compile to the test timeout @@ -398,6 +398,10 @@ it('marks each file with the version of the run that wrote it', async () => { await waitFor(() => expect(screen.getAllByText('film.mp4')).toHaveLength(2)) expect(screen.getByText('v1')).toBeTruthy() expect(screen.getByText('v4')).toBeTruthy() + // One space between chip and label, so a screen reader does not run + // 'v4' into the file name + const caption = screen.getByText('v4').closest('.caption') as HTMLElement + expect(caption.textContent?.replace(/\s+/g, ' ').trim()).toBe('v4 film.mp4') }) it('shows no version for a flat-layout file, which belongs to no run', async () => { From 667642e4e720d9a5da429b8ce7d2e57fa7456b3d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 16:12:40 -0500 Subject: [PATCH 003/181] feat: add versioning to output references and gallery listings - Enhanced the `list_gallery` function to support filtering by `folder` and `version`, allowing users to specify a particular run's output. - Introduced a `version` parameter in various API endpoints and functions to track the ordinal of runs, displayed as `v` in the gallery. - Updated the `Job` model to include `run_version`, reflecting the version of the run associated with each job. - Modified the realization process to correctly handle output references that specify a version, ensuring that the correct run is referenced. - Adjusted the UI components to display the run version alongside job details, improving user clarity on which version of a workflow is being referenced. - Added tests to verify the correct behavior of versioned output references and ensure that the system behaves as expected when handling versions. --- CLAUDE.md | 15 +++++++--- docs/MCP.md | 2 +- docs/SERVER.md | 3 +- docs/WORKFLOW_GUIDE.md | 13 ++++++--- dw/realize.py | 18 ++++++++---- dw/runs.py | 48 ++++++++++++++++++++++++++++---- dw/server/app.py | 29 ++++++++++++++++++- dw/server/exports.py | 12 ++++++++ dw/server/jobs.py | 24 ++++++++++++---- dw/workflow.py | 1 + dw_mcp/catalog.py | 18 ++++++++++-- dw_mcp/server.py | 9 ++++-- tests/test_events.py | 3 ++ tests/test_mcp_catalog.py | 13 +++++++++ tests/test_realize.py | 15 ++++++++++ tests/test_runs.py | 41 +++++++++++++++++++++++++++ tests/test_server.py | 12 ++++++++ tests/test_server_exports.py | 9 ++++++ tests/test_server_jobs.py | 8 ++++++ ui/CLAUDE.md | 6 ++-- ui/src/lib/pages/JobPage.svelte | 16 +++++++++++ ui/src/lib/pages/JobPage.test.ts | 18 ++++++++++++ ui/src/lib/pages/JobsPage.svelte | 7 +++++ ui/src/lib/types.ts | 6 ++++ 24 files changed, 313 insertions(+), 33 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index f6849a83..06e89aa8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -153,7 +153,9 @@ references) are documented above in *Workflow sources* and *Type System*. `"output:ltx2/Gyre/latest/still.png"`. The name is `//` under the output root, and `latest` in the run-id position picks the newest run that holds the file (run ids sort by their UTC timestamp; a failed or fully-cached run holds - only a manifest and is skipped). Resolved in `realize_args` beside `asset:` (`dw/runs.py`), + only a manifest and is skipped), and `v` there picks the run whose version is N + (below) - exactly that run, with no fallback to an older one. Either is a selector only + where run directories are, and the realized workflow pins both to the run id. Resolved in `realize_args` beside `asset:` (`dw/runs.py`), against the output root `Workflow.run` activates, and confined to it - A generated file becomes a stable input with `POST /api/assets/keep` (gallery "Keep as asset", MCP `keep_output`): it is hard-linked, else copied, from the workspace's outputs @@ -395,9 +397,14 @@ same reason - default setup cannot load a pack. not only a deletion: a failed run or a fully cached rerun takes a number and may have no media for the gallery to show under it. `GET /api/gallery` and the metadata route carry `version` and `run_id` - (`run_versions` read once per identity per listing, not per file), MCP - `list_gallery` teaches the vocabulary, and the web UI reads the field only - - a `v4` chip on the card, the run id in the detail pane. Nothing on disk is + (`run_versions` read once per identity per listing, not per file), and + `?folder=&version=` lists one run's files. The number is also a name: + `output:/v4/`. The `run_start` event carries it, the job + records it (`run_version`, a `jobs.sqlite` column) and the export README and + zip download name (`-v4-.zip`) carry it too. MCP + `list_gallery` teaches the vocabulary and takes `folder`/`version`, and the + web UI reads the field only - a `v4` chip on the gallery card, the jobs list + and the job page, the run id in the gallery's detail pane. Nothing on disk is renamed, so `output:` references, the step cache and `keep_output` are untouched. Two limits taken deliberately: deleting the *newest* run frees its number for reuse (the high-water mark lived in the manifest that went diff --git a/docs/MCP.md b/docs/MCP.md index f06956cb..b3ed1c69 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -226,7 +226,7 @@ when no single workflow covers it. | `get_health()` | — | Check that the server is alive, and which machine answered: `version`, `device`, whether a model process is currently resident (`worker_alive`), the job running now and the queue depth. `worker_alive: false` is the normal idle state on a server that has not run a job since startup or the last memory clear - not a fault - the on-demand worker starts with the next job (#206) | | `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in | | `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds | -| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. Tools take `name`, never `version`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None, folder=None, version=None)` | `limit`, `subfolder`, `only_orphans`, `workspace`, `folder`, `version` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. `folder` with `version` lists that one run's files, and `output:/v5/` names one in a workflow; every other tool takes `name`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | ### Media diff --git a/docs/SERVER.md b/docs/SERVER.md index f357d87c..30afe061 100644 --- a/docs/SERVER.md +++ b/docs/SERVER.md @@ -441,7 +441,8 @@ The editor's forms come from these; they are just as usable from scripts: gallery entry carries `folder` (the workflow identity, the run id dropped) and `subfolder` (what followed the run id - the `final`/`intermediate` a step's `result.subfolder` chose, `''` when it chose none); `?folder=` and - `?subfolder=` filter independently, and the reply's `folders` and + `?subfolder=` filter independently (`?version=` too - with `?folder=`, + the one run the gallery labels `v4`), and the reply's `folders` and `subfolders` list every distinct value over the whole tree, `''` always a member of each so root-level files stay selectable - `GET /api/gallery/{name:path}/download` — download an output file diff --git a/docs/WORKFLOW_GUIDE.md b/docs/WORKFLOW_GUIDE.md index 53fbeac6..895b34cd 100644 --- a/docs/WORKFLOW_GUIDE.md +++ b/docs/WORKFLOW_GUIDE.md @@ -289,7 +289,8 @@ for existence. to the same file. - `output:` — `output://` is a file an earlier run wrote, under the output root and confined to it. `latest` in the run-id position - picks the newest run that holds that file. A run id is not stable against + picks the newest run that holds that file; `v` picks the run the gallery labels + `v` (`list_gallery`'s `version`), and only that run. A run id is not stable against pruning: to depend on a generated file, promote it with `keep_output` and reference the `asset:` name instead. - `prompt:` — `prompt:name` or `prompt:folder/name` is a stored prompt's @@ -1459,7 +1460,7 @@ Beside that manifest the run also writes `workflow.json` — the *realized* workflow, meaning the one that actually ran. Every mutable input is pinned into it: the caller's `arguments` folded into the `variables` defaults, the seed the run used, each `prompt:` reference replaced by the stored text, and each -`output:/latest/` rewritten to the run id it resolved to. +`output:/latest/` (or `/v/`) rewritten to the run id it resolved to. `asset:`, `constant:`, `previous_result:` and `builtin:` are kept as written — each already names something pinned by the asset library or by the manifest's `dw_version` — and a sub-workflow named by local path is kept with its file's @@ -1657,8 +1658,12 @@ second-stage workflow name the first stage's product without being edited after run - and keeps working when the newest run failed part way, or reused every step from the cache and so wrote nothing of its own but a manifest. Runs sort by their id, which starts with a UTC timestamp, so "newest" needs no file timestamps and survives a -directory being copied. `latest` only selects a run where run directories are; a -workflow or file that happens to be called `latest` is still named as itself. +directory being copied. `v` in the same position names the run whose version is N - +the `v4` the gallery labels its files with - so the number a person was told is a name +a workflow can take. Unlike `latest` it picks exactly one run: `v4` not holding the file +is an error, not a reason to try `v3`. `latest` and `v` only select a run where run +directories are; a workflow or file that happens to be called either is still named as +itself. Like `asset:`, a reference resolves to a path and then whatever loads paths loads it, so it works under `image`, `video`, a `from_file`, or a list of them. The audio tasks take diff --git a/dw/realize.py b/dw/realize.py index c970aad2..f00f402d 100644 --- a/dw/realize.py +++ b/dw/realize.py @@ -29,6 +29,7 @@ is_output_reference, output_root as default_output_root, resolve_output_reference, + version_selector, ) from .security import SecurityError, validate_workflow_path from .workflow_sources import resolve_sub_workflow, SubWorkflowNotFound @@ -170,14 +171,19 @@ def _inline_prompt(reference, annotations, prompt_dir, base_dir): def _pin_output(reference, output_root): - """'output:/latest/' rewritten to the run it resolved to. - - An explicit run id is already pinned, so it is returned untouched without - touching the disk - realizing must not fail on a reference the run has - not reached yet. + """'output:/latest/' - or '/v4/' - rewritten to the run it + resolved to. + + A version is stable, but deleting the newest run frees its number for + reuse, so the realized copy names the run id either way. An explicit run + id is already pinned, so it is returned untouched without touching the + disk - realizing must not fail on a reference the run has not reached + yet. """ name = reference.removeprefix(OUTPUT_PREFIX).strip() - if LATEST not in name.split("/"): + if not any( + part == LATEST or version_selector(part) is not None for part in name.split("/") + ): return reference root = output_root or default_output_root() try: diff --git a/dw/runs.py b/dw/runs.py index 797369ca..919a3437 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -57,6 +57,17 @@ # newest run directory LATEST = "latest" +# 'v4' in the run-id position of an 'output:' reference: the run whose +# ordinal is 4 - the number the gallery shows and an agent quotes +_VERSION_SELECTOR = re.compile(r"^v([1-9][0-9]*)$") + + +def version_selector(segment): + """The ordinal a 'v' segment names, or None for any other segment.""" + match = _VERSION_SELECTOR.match(segment) + return int(match.group(1)) if match else None + + # What a run id looks like: a UTC timestamp and a short digest of the spec. # The pattern is not only documentation - the gallery reads it to group a # workflow's runs under one folder rather than listing every run separately @@ -126,8 +137,8 @@ def _runs_newest_first(directory): def _resolve_segments(directory, parts, reference, root): - """Build the path a name stands for, expanding 'latest' where it names - a run. + """Build the path a name stands for, expanding 'latest' or 'v' where + it names a run. 'latest' means the newest run *that has the file*, not the newest run directory: a run that failed part way, or one whose every step was a @@ -136,9 +147,14 @@ def _resolve_segments(directory, parts, reference, root): stage before it plainly produced something. So the runs are tried newest first and the first one holding the rest of the name wins. + 'v' means the run whose recorded ordinal is N - the 'v4' the gallery + shows - so the number quoted to a person is also a name a workflow can + take. Unlike 'latest' it picks exactly one run: a v4 that did not write + the file is an error, not a reason to try v3. + Only a segment standing where run directories are is a run selector. A - 'latest' segment in a directory that holds no runs is a name like any - other, so a workflow or a file called 'latest' stays reachable. + 'latest' or 'v4' segment in a directory that holds no runs is a name + like any other, so a workflow or a file called either stays reachable. Returns the path, or None when runs were found and none of them holds the file. @@ -162,6 +178,27 @@ def _resolve_segments(directory, parts, reference, root): f"'{reference}' names the newest run of a workflow that " f"has not produced one" ) + wanted = version_selector(part) + if wanted is not None and _runs_newest_first(directory): + versions = run_versions(directory) + matching = [run for run, version in versions.items() if version == wanted] + if not matching: + held = ", ".join(f"v{v}" for v in sorted(set(versions.values()))) + raise ValueError( + f"No run v{wanted} under {os.path.relpath(directory, root)} - " + f"'{reference}' names a run by its version, and the runs " + f"there are {held}" + ) + # Normally one; two only where history predating versions could + # not be ranked beneath the first recorded number. Newest first, + # as 'latest' would try them + for run in sorted(matching, key=run_id_sort_key, reverse=True): + candidate = _resolve_segments( + os.path.join(directory, run), rest, reference, root + ) + if candidate and os.path.isfile(candidate): + return candidate + return None return _resolve_segments(os.path.join(directory, part), rest, reference, root) @@ -170,7 +207,8 @@ def resolve_output_reference(reference, root=None): The name is a path under the output directory - '//' - and the run id may be written as 'latest', which resolves - to the newest run of that workflow that holds the file. That is what + to the newest run of that workflow that holds the file, or as 'v', + the run whose version is N. 'latest' is what lets a second-stage workflow name the first stage's product without being edited after every run, and without breaking when the newest run failed or reused cached files and so wrote none of its own. diff --git a/dw/server/app.py b/dw/server/app.py index 28ee5929..cb98172f 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -14,6 +14,7 @@ import tempfile import copy import json +import re import uuid import asyncio import logging @@ -2824,6 +2825,7 @@ def gallery( folder: Optional[str] = None, subfolder: Optional[str] = None, only_orphans: bool = False, + version: Optional[int] = None, ws: Workspace = Depends(selected_workspace), ): """A page of media files in the output directory, newest first. @@ -2838,6 +2840,8 @@ def gallery( way: the in-run subfolders steps wrote into ('final', 'intermediate'), '' for files at a run's root. `folder` and `subfolder` filter independently and intersect when both are given. + `version` narrows to the runs holding that ordinal - with `folder`, + the one run "v4" names; without it, that run of every workflow. `only_orphans=true` inverts the whole call: instead of media files, it returns run directories holding nothing but their own @@ -2868,6 +2872,8 @@ def gallery( entries = [e for e in entries if e["folder"] == folder] if subfolder is not None: entries = [e for e in entries if e["subfolder"] == subfolder] + if version is not None: + entries = [e for e in entries if e["version"] == version] offset = max(0, offset) limit = max(0, limit) page = entries[offset : offset + limit] @@ -4165,6 +4171,27 @@ async def input_file( files = _static_files_for(roots[0]) return await files.get_response(name, request.scope) + def _export_download_name(directory, job_id): + """'-v4-.zip' when the exported manifest says which + run it was, else '.zip'. Only the saved file's name: the + URL and the entries inside keep the job id, so nothing that already + names an export changes.""" + try: + with open(os.path.join(directory, MANIFEST_FILE_NAME)) as file: + manifest = json.load(file) + except (OSError, ValueError): + return f"{job_id}.zip" + if not isinstance(manifest, dict): + return f"{job_id}.zip" + version = manifest.get("version") + identity = (manifest.get("workflow") or {}).get("identity") + if not isinstance(version, int) or isinstance(version, bool): + return f"{job_id}.zip" + if not isinstance(identity, str) or not identity: + return f"v{version}-{job_id}.zip" + slug = re.sub(r"[^A-Za-z0-9_.-]+", "-", identity).strip("-.") + return f"{slug}-v{version}-{job_id}.zip" if slug else f"v{version}-{job_id}.zip" + # Ungated for the same reason the two above are: a download link cannot # attach an Authorization header either @app.get("/exports/{job_id}.zip") @@ -4186,7 +4213,7 @@ def export_zip(job_id: str, ws: Workspace = Depends(selected_workspace)): path = os.path.join(current, name) entry = os.path.relpath(path, directory).replace(os.sep, "/") entries.append((f"{job_id}/{entry}", path)) - return _zip_download(entries, f"{job_id}.zip") + return _zip_download(entries, _export_download_name(directory, job_id)) # ---------------------------------------------------------------- the UI diff --git a/dw/server/exports.py b/dw/server/exports.py index 7953f7a9..bf573a45 100644 --- a/dw/server/exports.py +++ b/dw/server/exports.py @@ -77,6 +77,7 @@ "error", "run_id", "run_dir", + "run_version", ) README_TEMPLATE = """# {workflow_name} - job {job_id} @@ -90,6 +91,7 @@ | Job | `{job_id}` | | Workflow | `{workflow_name}` | | Catalog entry | {catalog_name} | +| Run | {run} | | Status | {status} | | Started | {started_at} | | Finished | {finished_at} | @@ -394,6 +396,16 @@ def _readme(job_id, detail, manifest, workflow, summary, realized): "predates run tracking, so its arguments and prompts are not " "pinned into it." ), + run=( + f"`{detail['run_id']}`" + + ( + f" - version {manifest['version']}" + if isinstance(manifest.get("version"), int) + else "" + ) + if detail.get("run_id") + else "not recorded" + ), status=detail.get("status"), started_at=detail.get("started_at"), finished_at=detail.get("finished_at"), diff --git a/dw/server/jobs.py b/dw/server/jobs.py index c01b6950..69fee683 100644 --- a/dw/server/jobs.py +++ b/dw/server/jobs.py @@ -134,6 +134,11 @@ def __init__(self, db_path): connection.execute("ALTER TABLE jobs ADD COLUMN run_id TEXT") if "run_dir" not in columns: connection.execute("ALTER TABLE jobs ADD COLUMN run_dir TEXT") + # That run's ordinal among the workflow's runs - the 'v4' the + # gallery shows. NULL before the column, and for a job that + # never opened a run + if "run_version" not in columns: + connection.execute("ALTER TABLE jobs ADD COLUMN run_version INTEGER") # Which form of cost acknowledgement queued the job. Rows before # the column are 'none' - nothing recorded is nothing recorded if "acknowledged" not in columns: @@ -178,8 +183,8 @@ def record(self, job): " started_at, finished_at, arguments, spec, manifest, warnings," " error, events, workspace, workflow_name, run_id, run_dir," " acknowledged, host_memory_peak_rss_mb," - " host_memory_job_peak_rss_mb) VALUES" - " (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)", + " host_memory_job_peak_rss_mb, run_version) VALUES" + " (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)", ( job.id, job.workflow_name, @@ -204,6 +209,7 @@ def record(self, job): # itself allows getattr(job, "host_memory_peak_rss_mb", None), getattr(job, "host_memory_job_peak_rss_mb", None), + getattr(job, "run_version", None), ), ) @@ -219,7 +225,8 @@ def recent_summaries(self, limit=200, workspace=None, statuses=None): """ query = ( "SELECT id, workflow, status, created_at, started_at, finished_at," - " workspace, workflow_name, run_id, acknowledged FROM jobs" + " workspace, workflow_name, run_id, acknowledged, run_version" + " FROM jobs" ) params = [] clauses = [] @@ -249,6 +256,7 @@ def recent_summaries(self, limit=200, workspace=None, statuses=None): "workflow_name": row[7], "run_id": row[8], "acknowledged": row[9] or ACK_NONE, + "run_version": row[10], "historical": True, } for row in rows @@ -259,8 +267,8 @@ def get(self, job_id): row = connection.execute( "SELECT id, workflow, status, created_at, started_at, finished_at," " arguments, spec, manifest, warnings, error, workspace," - " workflow_name, run_id, run_dir, acknowledged, events FROM jobs" - " WHERE id = ?", + " workflow_name, run_id, run_dir, acknowledged, events," + " run_version FROM jobs WHERE id = ?", (job_id,), ).fetchone() return self._to_detail(row) if row else None @@ -479,6 +487,7 @@ def parse(text, fallback): "workflow_name": row[12], "run_id": row[13], "run_dir": row[14], + "run_version": row[17], "acknowledged": row[15] or ACK_NONE, "acknowledged_cost": (spec or {}).get("acknowledged_cost"), "traceback": None, @@ -510,6 +519,7 @@ def __init__(self, spec): # never got that far self.run_id = None self.run_dir = None + self.run_version = None # Which form of cost acknowledgement queued this job (#85) self.acknowledged = spec.get("acknowledged") or ACK_NONE # The worker's own high-water mark for this run, from its final @@ -685,6 +695,9 @@ def summary(self): # so it defaults the same way history's column does "workspace": self.spec.get("workspace") or DEFAULT_WORKSPACE_NAME, "run_id": self.run_id, + # The run's ordinal - 'v4' - so the job that just ran can be + # named the way the gallery will name it + "run_version": self.run_version, "acknowledged": self.acknowledged, } @@ -1303,6 +1316,7 @@ def _consume_results(self, job): if event.get("event") == "run_start": job.run_id = event.get("run_id") job.run_dir = event.get("run_dir") + job.run_version = event.get("version") job.add_event(event) elif message_type in ("output", "workflow_loaded"): text = message.get("message") or message.get("workflow_name", "") diff --git a/dw/workflow.py b/dw/workflow.py index 0845237e..e65e173b 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -1172,6 +1172,7 @@ def run( run_context.emit( "run_start", run_id=run_id, + version=self._run_version, identity=workflow_identity(self.file_spec, workflow_id), run_dir=os.path.relpath(self._run_dir, self.output_dir).replace( os.sep, "/" diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 4e1c70bb..669fd502 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -213,10 +213,20 @@ def list_jobs(client, limit=20, status=None, workspace=None): return answer -def list_gallery(client, limit=50, subfolder=None, only_orphans=False, workspace=None): +def list_gallery( + client, + limit=50, + subfolder=None, + only_orphans=False, + workspace=None, + folder=None, + version=None, +): """Generated media in the output directory, newest first. `subfolder` narrows to one in-run subfolder ('final', 'intermediate', '' for files - at a run's root); None means every file. + at a run's root); None means every file. `folder` narrows to one + workflow and `version` to one run's ordinal, so the two together list + the run a person calls "v4". `only_orphans=True` inverts the call: instead of files, it returns run directories holding nothing but their own bookkeeping (manifest.json, @@ -242,6 +252,10 @@ def list_gallery(client, limit=50, subfolder=None, only_orphans=False, workspace params = {"limit": limit} if subfolder is not None: params["subfolder"] = subfolder + if folder is not None: + params["folder"] = folder + if version is not None: + params["version"] = version if only_orphans: params["only_orphans"] = "true" return client.get_json("/api/gallery", params=params, workspace=workspace) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 44e02089..4272c9ba 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -395,6 +395,8 @@ def list_gallery( subfolder: str | None = None, only_orphans: bool = False, workspace: str | None = None, + folder: str | None = None, + version: int | None = None, ) -> dict: """List generated output files, newest first. A name is //, where may itself sit in a @@ -412,8 +414,9 @@ def list_gallery( Entries also carry `run_id` and `version`, that run's ordinal among the workflow's runs - stable, never renumbered. Quote the version - to a person: the web UI labels the same file `v5`. Tools still take - `name`. + to a person: the web UI labels the same file `v5`. `folder=` with + `version=` lists that one run; "output:/v5/" names it + in a workflow. Other tools still take `name`. `only_orphans=True` inverts the call: instead of files, it returns run directories holding nothing but their own bookkeeping @@ -438,6 +441,8 @@ def list_gallery( subfolder=subfolder, only_orphans=only_orphans, workspace=workspace, + folder=folder, + version=version, ) def get_gallery_metadata( diff --git a/tests/test_events.py b/tests/test_events.py index 1a60baa8..b97f7e3e 100644 --- a/tests/test_events.py +++ b/tests/test_events.py @@ -89,6 +89,9 @@ def test_progress_event_sequence(): names = [event["event"] for event in events] assert names[0] == "run_start" + # the run's ordinal, so a job can name its run the way the gallery will + # (the output root here is shared, so only its shape is fixed) + assert isinstance(events[0]["version"], int) and events[0]["version"] >= 1 assert names[1] == "workflow_start" assert names[-1] == "workflow_end" assert "step_start" in names and "step_end" in names diff --git a/tests/test_mcp_catalog.py b/tests/test_mcp_catalog.py index 2d28cb48..b7b698cb 100644 --- a/tests/test_mcp_catalog.py +++ b/tests/test_mcp_catalog.py @@ -117,6 +117,19 @@ def test_list_gallery_sends_a_subfolder_only_when_given(): assert seen["params"]["subfolder"] == "" +def test_list_gallery_sends_folder_and_version_only_when_given(): + client, seen = recording_client() + catalog.list_gallery(client, limit=7) + assert "folder" not in seen["params"] + assert "version" not in seen["params"] + + # together they name one run - what a person calls "v4" + client, seen = recording_client() + catalog.list_gallery(client, folder="acorn/cut", version=4) + assert seen["params"]["folder"] == "acorn/cut" + assert seen["params"]["version"] == "4" + + def test_list_gallery_sends_only_orphans_only_when_true(): client, seen = recording_client() catalog.list_gallery(client, limit=7) diff --git a/tests/test_realize.py b/tests/test_realize.py index cec60148..656399b6 100644 --- a/tests/test_realize.py +++ b/tests/test_realize.py @@ -136,6 +136,21 @@ def test_latest_is_pinned_to_the_run_it_resolved_to(self, output_root): f"output:ltx2/Gyre/{run_id}/still.png" ) + def test_a_version_is_pinned_to_the_run_it_named(self, output_root): + # A version is stable, but deleting the newest run frees its number, + # so the realized copy names the run id as it does for 'latest' + root, run_id = output_root + source = definition() + source["steps"][0]["pipeline"]["arguments"]["image"] = ( + "output:ltx2/Gyre/v1/still.png" + ) + + realized, _ = realize_workflow(source, {}, 7, output_root=root) + + assert realized["steps"][0]["pipeline"]["arguments"]["image"] == ( + f"output:ltx2/Gyre/{run_id}/still.png" + ) + def test_an_explicit_run_id_is_kept_as_written(self, output_root): root, run_id = output_root written = f"output:ltx2/Gyre/{run_id}/still.png" diff --git a/tests/test_runs.py b/tests/test_runs.py index 8fe7f208..360355a7 100644 --- a/tests/test_runs.py +++ b/tests/test_runs.py @@ -905,6 +905,47 @@ def test_opening_a_run_pins_the_numbers_of_older_runs(self, tmp_path): "20260903-120000-aaaaaaaa": 3, } + def test_an_output_reference_can_name_a_run_by_its_version(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + first = self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + self._run(str(identity), "20260903-120000-cccccccc", version=3) + with open(os.path.join(first, "still.png"), "wb") as file: + file.write(b"png") + + resolved = resolve_output_reference( + "output:ltx2/Gyre/v1/still.png", str(tmp_path) + ) + assert resolved == os.path.realpath(os.path.join(first, "still.png")) + + def test_a_version_that_did_not_write_the_file_does_not_fall_back(self, tmp_path): + # Unlike 'latest', 'v3' picks one run: v3 lacking the file is an + # error, not a reason to hand back v1's + identity = tmp_path / "ltx2" / "Gyre" + first = self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + self._run(str(identity), "20260903-120000-cccccccc", version=3) + with open(os.path.join(first, "still.png"), "wb") as file: + file.write(b"png") + + with pytest.raises(ValueError, match="not found"): + resolve_output_reference("output:ltx2/Gyre/v3/still.png", str(tmp_path)) + + def test_a_missing_version_names_the_ones_there_are(self, tmp_path): + identity = tmp_path / "ltx2" / "Gyre" + self._run(str(identity), "20260901-120000-aaaaaaaa", version=1) + self._run(str(identity), "20260903-120000-cccccccc", version=3) + + with pytest.raises(ValueError, match="No run v2 .* v1, v3"): + resolve_output_reference("output:ltx2/Gyre/v2/still.png", str(tmp_path)) + + def test_a_v_segment_where_no_runs_are_is_an_ordinary_name(self, tmp_path): + # A workflow folder called 'v2' stays reachable, as 'latest' does + target = tmp_path / "flat" / "v2" + target.mkdir(parents=True) + (target / "still.png").write_bytes(b"png") + + resolved = resolve_output_reference("output:flat/v2/still.png", str(tmp_path)) + assert resolved == os.path.realpath(str(target / "still.png")) + def test_pinning_leaves_a_run_with_no_manifest_alone(self, tmp_path): # Writing one would invent a record of a run nobody recorded identity = tmp_path / "ltx2" / "Gyre" diff --git a/tests/test_server.py b/tests/test_server.py index cdccac7c..b808d253 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -2043,6 +2043,18 @@ def _run(identity, run_id, version=None): assert by_name["ltx/flat.png"]["version"] is None assert by_name["ltx/flat.png"]["run_id"] == "" + # "show me v3": folder and version together list exactly that run + names = [ + f["name"] + for f in client.get( + "/api/gallery", params={"folder": "acorn/cut", "version": 3} + ).json()["files"] + ] + assert names == ["acorn/cut/20260903-120000-cccccccc/final/film.7-0.0.png"] + # version alone spans workflows: each one's v1 + v1 = client.get("/api/gallery", params={"version": 1}).json()["files"] + assert {f["folder"] for f in v1} == {"acorn/cut", "acorn/score"} + def test_gallery_metadata_names_the_run_and_its_version(server, tmp_path): """After "look at version 3", the next call is usually this one - so it diff --git a/tests/test_server_exports.py b/tests/test_server_exports.py index 95163a52..35f11140 100644 --- a/tests/test_server_exports.py +++ b/tests/test_server_exports.py @@ -56,8 +56,10 @@ def exporting_script(command): json.dump( { "run_id": RUN_ID, + "version": 4, "status": "completed", "seed": 7, + "workflow": {"identity": "server_test"}, "steps": [{"step": "gen", "files": ["still.png"]}], }, file, @@ -66,6 +68,7 @@ def exporting_script(command): "type": "progress", "event": "run_start", "run_id": RUN_ID, + "version": 4, "identity": "server_test", "run_dir": RUN_DIR, } @@ -282,6 +285,8 @@ def test_the_readme_names_the_job_and_says_how_to_run_it(self, server): assert job_id in readme assert "python -m dw.run workflow.json" in readme assert "Git LFS" in readme + # which run, in the form the gallery labels it + assert f"`{RUN_ID}` - version 4" in readme def test_the_job_s_own_asset_dir_is_used_not_the_export_s_workspace( self, server, workspace_root @@ -376,6 +381,10 @@ def test_it_lists_the_same_entries_as_the_directory(self, server): response = client.get(f"/exports/{job_id}.zip") assert response.status_code == 200 + # the saved file says which run it is; the URL and the entries + # inside keep the job id + disposition = response.headers["content-disposition"] + assert f"server_test-v4-{job_id}.zip" in disposition archive = zipfile.ZipFile(io.BytesIO(response.content)) assert sorted(archive.namelist()) == sorted( f"{job_id}/{entry['path']}" for entry in body["files"] diff --git a/tests/test_server_jobs.py b/tests/test_server_jobs.py index d5f4d2ef..0ab4347d 100644 --- a/tests/test_server_jobs.py +++ b/tests/test_server_jobs.py @@ -28,6 +28,7 @@ def tracked_script(command): "run_id": RUN_ID, "identity": "server_test", "run_dir": RUN_DIR, + "version": 4, } yield {"type": "success", "message": "ok", "run_count": 1, "manifest": []} @@ -59,6 +60,9 @@ def test_run_start_populates_the_job(manager): assert job.run_dir == RUN_DIR assert job.summary()["run_id"] == RUN_ID assert job.detail()["run_dir"] == RUN_DIR + # the ordinal the gallery shows for this run's files + assert job.run_version == 4 + assert job.summary()["run_version"] == 4 def test_both_persist_and_read_back(manager): @@ -66,6 +70,10 @@ def test_both_persist_and_read_back(manager): historical = manager.history.get(job.id) assert historical["run_id"] == RUN_ID assert historical["run_dir"] == RUN_DIR + assert historical["run_version"] == 4 + # and in the polled list, not only the detail + (summary,) = manager.history.recent_summaries() + assert summary["run_version"] == 4 def test_realized_reads_the_file_the_run_wrote(manager, tmp_path): diff --git a/ui/CLAUDE.md b/ui/CLAUDE.md index 979067f3..47005503 100644 --- a/ui/CLAUDE.md +++ b/ui/CLAUDE.md @@ -65,8 +65,10 @@ Stripping the run id is also what makes four runs of one workflow four identical captions, since a file's name is per step rather than per run. The gallery entry carries `version` - the run's ordinal, assigned by the engine and never renumbered - and the grid draws it as a `v4` chip ahead -of the label, with the run id in the detail pane beside it. The UI reads -the field only; nothing here computes or orders a version. +of the label, with the run id in the detail pane beside it. A job carries +the same number as `run_version`, drawn as `v4` in the jobs list and on the +job page (from the `run_start` event while the job runs). The UI reads the +field only; nothing here computes or orders a version. Every picture in the app sits in the global `.frame` (app.css): the media fills it edge to edge, with no inner padding and no rounding of its own, diff --git a/ui/src/lib/pages/JobPage.svelte b/ui/src/lib/pages/JobPage.svelte index f8a7c07d..2cf92e1f 100644 --- a/ui/src/lib/pages/JobPage.svelte +++ b/ui/src/lib/pages/JobPage.svelte @@ -215,6 +215,13 @@ events.find((e) => e.event === 'workflow_start')?.seed as number | undefined, ) + // The record's number once the job has one; while it runs, the run_start + // event says it first - so the page names the run the moment it opens + const runVersion = $derived( + job?.run_version ?? + (events.find((e) => e.event === 'run_start')?.version as + number | undefined), + ) const etaSeconds = $derived.by(() => { if (!denoise?.total_steps || stepTimes.length < 3) return null const window = stepTimes.slice(-6) @@ -345,6 +352,15 @@ >seed {seed} {/if} + {#if runVersion} + + v{runVersion} + {/if} {#if job.acknowledged === 'bound'} diff --git a/ui/src/lib/pages/JobPage.test.ts b/ui/src/lib/pages/JobPage.test.ts index 6e0f90c7..94fb8e4d 100644 --- a/ui/src/lib/pages/JobPage.test.ts +++ b/ui/src/lib/pages/JobPage.test.ts @@ -407,3 +407,21 @@ it("corrects the URL to the job's own workspace", async () => { render(JobPage, { jobId: 'j1' }) await waitFor(() => expect(location.hash).toBe('#/ws/studio/jobs/j1')) }) + +it('names the run by the version the gallery labels its files with', async () => { + detail.job = { + ...job([]), + run_id: '20260922-120000-aaaaaaaa', + run_version: 4, + } + render(JobPage, { jobId: 'j1' }) + const chip = await waitFor(() => screen.getByText('v4')) + expect(chip.getAttribute('title')).toContain('20260922-120000-aaaaaaaa') +}) + +it('shows no version for a job that never opened a run', async () => { + detail.job = { ...job([]), run_version: null } + render(JobPage, { jobId: 'j1' }) + await waitFor(() => expect(screen.getByText('j1')).toBeTruthy()) + expect(screen.queryByText(/^v\d+$/)).toBeNull() +}) diff --git a/ui/src/lib/pages/JobsPage.svelte b/ui/src/lib/pages/JobsPage.svelte index 3f494da2..9c94b050 100644 --- a/ui/src/lib/pages/JobsPage.svelte +++ b/ui/src/lib/pages/JobsPage.svelte @@ -131,6 +131,13 @@ {job.status} {job.workflow} + {#if job.run_version} + v{job.run_version} + {/if} {#if scope === 'all' && (workspace.names?.length ?? 0) > 1} {job.workspace} {/if} diff --git a/ui/src/lib/types.ts b/ui/src/lib/types.ts index 357fcbaf..a1e586cf 100644 --- a/ui/src/lib/types.ts +++ b/ui/src/lib/types.ts @@ -13,6 +13,12 @@ export interface JobSummary { * and any caller that sent nothing), a bare boolean, or one bound to the * plan a validate answered with. Absent on rows from older servers. */ acknowledged?: 'none' | 'boolean' | 'bound' + /** The run this job opened - null until it opens one, and for a job + * recorded before runs were tracked. */ + run_id?: string | null + /** That run's ordinal among the workflow's runs - the `v4` the gallery + * shows for its files. Null until the run opens, and for older rows. */ + run_version?: number | null } export interface ManifestEntry { From 7132590c1bb578561f057dc03705f8258927b4cf Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 16:17:03 -0500 Subject: [PATCH 004/181] Pay for list_gallery's new parameters; find npm on a non-interactive PATH The folder/version parameters pushed the MCP surface 69 tokens over budget. The docstring now costs less than it did before the change (descriptions 9_078 -> 9_068); the two schema entries (+49) are taken by the budget deliberately, 13_850 -> 13_890, since no docstring can pay for a schema. deploy.sh ran `npm run build` on whatever PATH the caller's shell gave it. `ssh lem deploy.sh` is non-interactive, and lem's ~/.bashrc puts ~/.local/node/bin on PATH only after its interactive-only guard, so npm that works at a prompt was missing. find_npm now looks in DW_NODE_BIN and the usual per-user install locations (and nvm) before the UI build, and fails naming DW_NODE_BIN rather than with a bare command-not-found. Co-Authored-By: Claude Opus 5.5 (1M context) --- dw_mcp/server.py | 16 +++++++--------- scripts/deploy.sh | 30 +++++++++++++++++++++++++++++- tests/test_mcp_server.py | 8 +++++++- 3 files changed, 43 insertions(+), 11 deletions(-) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 4272c9ba..093e7bb0 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -408,15 +408,13 @@ def list_gallery( by convention `final` is the deliverable and `intermediate` the scratch work, '' when the step chose none); `subfolder=` filters on the latter, so `subfolder="final"` is "what did these runs - deliver". Each entry also carries a ready-made `url` for viewing the - file over HTTP, already scoped to the right workspace; use it as - given rather than composing one from the name. - - Entries also carry `run_id` and `version`, that run's ordinal among - the workflow's runs - stable, never renumbered. Quote the version - to a person: the web UI labels the same file `v5`. `folder=` with - `version=` lists that one run; "output:/v5/" names it - in a workflow. Other tools still take `name`. + deliver". Each entry's `url` is already scoped to its workspace; + use it as given rather than composing one from the name. + + Entries also carry `run_id` and `version`, the run's stable ordinal + (the web UI shows `v5`) - quote the version to a person. `folder=` + plus `version=` lists that run; "output:/v5/" names + it. Other tools take `name`. `only_orphans=True` inverts the call: instead of files, it returns run directories holding nothing but their own bookkeeping diff --git a/scripts/deploy.sh b/scripts/deploy.sh index e90ac842..289cf005 100755 --- a/scripts/deploy.sh +++ b/scripts/deploy.sh @@ -30,7 +30,9 @@ # # Environment overrides, all optional: # DW_DIR (checkout, default ~/diffusers-workflow), DW_TOKEN (default xyz), -# DW_PORT (8765), DW_WORKSPACE (~/diffusers-workspace), DW_HOST (0.0.0.0). +# DW_PORT (8765), DW_WORKSPACE (~/diffusers-workspace), DW_HOST (0.0.0.0), +# DW_NODE_BIN (the directory holding npm, when it is not on a +# non-interactive PATH and not in one of the places find_npm looks). set -euo pipefail DW_DIR="${DW_DIR:-$HOME/diffusers-workflow}" @@ -57,6 +59,31 @@ say() { echo "[deploy $(ts)] $*"; } health() { curl -s -m 5 -H "Authorization: Bearer $DW_TOKEN" "$HEALTH" 2>/dev/null; } server_pids() { pgrep -f 'python -m dw\.serve' || true; } +# `ssh lem deploy.sh` runs a non-interactive shell, and a per-user node +# install is usually put on PATH by ~/.bashrc *after* its "not interactive, +# stop here" guard - so npm that works at a prompt is missing here. Look in +# the usual per-user places rather than depend on the caller's shell +find_npm() { + command -v npm >/dev/null 2>&1 && return 0 + local dir + for dir in "${DW_NODE_BIN:-}" "$HOME/.local/node/bin" "$HOME/.volta/bin" \ + "$HOME/.local/share/fnm/aliases/default/bin" "$HOME/.local/bin" /usr/local/bin; do + if [ -n "$dir" ] && [ -x "$dir/npm" ]; then + PATH="$dir:$PATH"; export PATH + say "npm not on PATH; using $dir" + return 0 + fi + done + # nvm is a shell function, not a directory on PATH + if [ -s "${NVM_DIR:-$HOME/.nvm}/nvm.sh" ]; then + # shellcheck disable=SC1091 + . "${NVM_DIR:-$HOME/.nvm}/nvm.sh" >/dev/null 2>&1 && command -v npm >/dev/null 2>&1 \ + && { say "npm from nvm: $(command -v npm)"; return 0; } + fi + say "npm not found (PATH=$PATH); set DW_NODE_BIN to the directory holding npm" + exit 1 +} + cd "$DW_DIR" [ -z "$branch" ] && branch="$(git branch --show-current)" @@ -80,6 +107,7 @@ if echo "$changed" | grep -qx 'pyproject.toml'; then venv/bin/pip install -q -e . fi if echo "$changed" | grep -q '^ui/' || [ ! -d ui/dist ]; then + find_npm if echo "$changed" | grep -qx 'ui/package-lock.json' || [ ! -d ui/node_modules ]; then say "ui/package-lock.json changed; npm ci" (cd ui && npm ci --silent --no-audit --no-fund) diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index b4a45d5e..4248fedb 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -1441,7 +1441,13 @@ def test_the_stated_tool_count_is_the_registered_one(): # as tightly as it can be said and still 44 tokens over, so the budget takes # them deliberately rather than the sentence being cut to nothing. Measured # 2026-09-22 at 13_844.0 (9_078.0 / 3_752.0 / 1_014.0). 6 tokens of headroom. -SURFACE_BUDGET = 13_850 +# Then list_gallery took `folder` and `version`, so "show me v5" is one call +# rather than a scan of the listing. The docstring paid for its own new +# sentence and then some (descriptions 9_078 -> 9_068, the `url` sentence +# said in fewer words); the two schema entries (+49) are what the budget +# takes, since no docstring can pay for a parameter's schema. Measured +# 2026-09-22 at 13_883.0 (9_068.0 / 3_801.0 / 1_014.0). 7 tokens of headroom. +SURFACE_BUDGET = 13_890 @pytest.mark.asyncio From b2fbf6e7825f050abb91bb6e3556b6e090a499fc Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 16:28:33 -0500 Subject: [PATCH 005/181] wait_for_job's slim job carries run_version The slim projection named run_id but not run_version, so an agent that had just waited on a run could not say "that was v5" - the label the gallery puts on its files - without a second call to get_job. One key in _SLIM_KEYS; no docstring change, so the MCP surface budget is untouched. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/MCP.md | 2 +- dw_mcp/diagnose.py | 3 +++ tests/test_mcp_diagnose.py | 18 ++++++++++++++++++ 3 files changed, 22 insertions(+), 1 deletion(-) diff --git a/docs/MCP.md b/docs/MCP.md index b3ed1c69..1ddac9ee 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -306,7 +306,7 @@ references written in the same session. | `get_job_workflow(job_id)` | `job_id` | The REST equivalent is `GET /api/jobs/{id}/workflow` (see [SERVER.md](SERVER.md#jobs-api)). The workflow the job actually ran. `realized: true` means every mutable input is pinned (arguments, seed, prompts, `output:latest`); `false` means the job predates run tracking and this is the definition as submitted. Pass it to `save_workflow` to keep it under a name | | `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `/exports//` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination - fetch the zip URL and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file); the archive already unpacks into one folder named after the job id | | `get_job_events(job_id, after=-1, limit=200)` | `job_id`, `after`, `limit` | Get a page of a job's progress events | -| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) | +| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, `run_id` and `run_version` (the run's `v5`, as the gallery labels it), and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) | | `cancel_job(job_id)` | `job_id` | Ask a queued or running job to stop | | `clear_memory()` | — | Drop every loaded pipeline and the step cache, freeing VRAM/RAM immediately instead of waiting for the next job to evict one model for another. Also drops the step cache, so a seeded workflow that would otherwise reuse cached results regenerates on its next run. Refused with a 409 while a job is running or queued - the queue is FIFO, so wait for it to finish and retry rather than expecting this call to block until it does (#221) | | `rerun_job(job_id, acknowledged_cost=False, new_seed=False)` | `job_id`, `acknowledged_cost`, `new_seed` | Queue a fresh job from a previous job's stored specification. Costs GPU time, so it passes the same gate as `run_workflow`. `new_seed=true` draws a fresh seed into the workflow's seed variable — without it a seeded workflow's rerun repeats its arguments exactly and the step cache serves the whole run from the earlier one's files (`reused: true`), generating nothing. `get_job_workflow`'s `seed_variable` says whether there is one - `acknowledged_cost` is `true` or the bound `{fingerprint, minutes, downloads}` from the validate plan; a 409 means the plan changed and the message carries the new estimate | diff --git a/dw_mcp/diagnose.py b/dw_mcp/diagnose.py index 66d7bd79..ade6ca36 100644 --- a/dw_mcp/diagnose.py +++ b/dw_mcp/diagnose.py @@ -198,6 +198,9 @@ def get_job_events(client, job_id, after=-1, limit=200): "finished_at", "workspace", "run_id", + # The run's ordinal - the 'v5' the gallery labels its files with - so + # the caller can name the run it just waited on without another call + "run_version", "queue_position", "warnings", "error", diff --git a/tests/test_mcp_diagnose.py b/tests/test_mcp_diagnose.py index 01192f1c..98e2474a 100644 --- a/tests/test_mcp_diagnose.py +++ b/tests/test_mcp_diagnose.py @@ -417,6 +417,24 @@ def test_wait_for_job_keeps_the_manifest_and_error_once_terminal(monkeypatch): assert "get_job" in result["next"] +def test_wait_for_job_names_the_run_the_way_the_gallery_will(monkeypatch): + """The job that just finished is the one a person asks about next, and + the gallery labels its files 'v5' - so the slim job carries the number + beside the run id rather than sending the caller to get_job for it.""" + done = { + **FAT_JOB, + "status": "succeeded", + "run_id": "20260922-212005-cd189c68", + "run_version": 5, + } + client, _ = scripted({("GET", "/api/jobs/job-1"): (200, done)}) + + result = diagnose.wait_for_job(client, "job-1") + + assert result["job"]["run_id"] == "20260922-212005-cd189c68" + assert result["job"]["run_version"] == 5 + + def test_wait_for_job_reports_queue_position_for_a_still_queued_job(monkeypatch): monkeypatch.setattr(diagnose, "MAX_WAIT_SECONDS", 0) queued = {**FAT_JOB, "status": "queued", "queue_position": 2} From e79f3913ddff2e6591e42e7002262c8ea431575f Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 17:17:38 -0500 Subject: [PATCH 006/181] fix(variables): #338 - refuse fractional args on int-typed variables get_value coerced a caller's argument with type(default), so a variable whose default happened to be a JSON integer (e.g. score_gain: 1) silently truncated a fractional override (0.3 -> 0) with no warning or error, from both run_workflow and validate_workflow (argument_errors shares the same path). Now a non-integral float against an int-typed variable raises, naming the variable, the value, and the truncated result. Float-ify the catalog defaults that invited this: guidance_scale in upscale-diffusion, flux2-dev and multi-image-reference, and audio_bleed_gain_db in assemble-and-score, minimax/music-video and minimax/dialogue-short. Co-Authored-By: Claude Sonnet 5 --- dw/variables.py | 15 +++++++++++++++ workflows/models/flux2-dev.json | 2 +- workflows/templates/assemble-and-score.json | 2 +- workflows/templates/minimax/dialogue-short.json | 2 +- workflows/templates/minimax/music-video.json | 2 +- workflows/templates/multi-image-reference.json | 2 +- workflows/templates/upscale-diffusion.json | 2 +- 7 files changed, 21 insertions(+), 6 deletions(-) diff --git a/dw/variables.py b/dw/variables.py index c9f8bed2..b5d99e34 100644 --- a/dw/variables.py +++ b/dw/variables.py @@ -327,6 +327,21 @@ def get_value(v, desired_type, name=None): if isinstance(v, PIL.Image.Image): return v + # A variable typed int by its default (e.g. `"score_gain": 1`) silently + # truncates a fractional override - int(0.3) == 0, with no error - which + # reads as a valid, if small, argument rather than the wrong type. Refuse + # it instead of truncating; the fix is to declare the default as a float + # (`1.0`) if the variable is meant to accept fractions + if desired_type is int and isinstance(v, float) and not v.is_integer(): + var_label = name if name is not None else "" + message = ( + f"{var_label} {v!r} would be realized as {int(v)}: this variable " + f"is typed integer by its default; declare its default as a float " + f"(e.g. {float(int(v))!r}) to accept fractional values" + ) + logger.error(message) + raise ValueError(message) + # A string cannot be coerced into a dict or a None - dict('/a/b.png') is # nonsense, NoneType('x') a TypeError. Those defaults are how media # variables ({'location': ...}) and optional inputs (null) are declared, diff --git a/workflows/models/flux2-dev.json b/workflows/models/flux2-dev.json index 4b49bb37..80ac952b 100644 --- a/workflows/models/flux2-dev.json +++ b/workflows/models/flux2-dev.json @@ -4,7 +4,7 @@ "prompt": "A realistic photograph of a mouse wearing a skirt playing volleyball against a team of professional volleyball players.", "num_images_per_prompt": 1, "num_inference_steps": 40, - "guidance_scale": 4 + "guidance_scale": 4.0 }, "id": "Flux2Dev", "description": "Text-to-image with FLUX.2 dev, pre-quantized to 4-bit BitsAndBytes. The 4-bit text encoder is loaded first as its own component and rested on the CPU, since BitsAndBytes materializes on the accelerator and the two together do not fit at load; model offloading then hands the accelerator to one component at a time, so both run on a 24 GB card.", diff --git a/workflows/templates/assemble-and-score.json b/workflows/templates/assemble-and-score.json index c57a3ab8..d14a3fe8 100644 --- a/workflows/templates/assemble-and-score.json +++ b/workflows/templates/assemble-and-score.json @@ -22,7 +22,7 @@ "sample_rate": 44100, "fps": 24, "audio_bleed_ms": 0, - "audio_bleed_gain_db": 0, + "audio_bleed_gain_db": 0.0, "seam_fade_ms": null, "match_levels": null, "match_levels_dbfs": null, diff --git a/workflows/templates/minimax/dialogue-short.json b/workflows/templates/minimax/dialogue-short.json index 482b7138..b7683ee1 100644 --- a/workflows/templates/minimax/dialogue-short.json +++ b/workflows/templates/minimax/dialogue-short.json @@ -117,7 +117,7 @@ "voice_reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3AudioReference", "subject_reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3ImageReference", "audio_bleed_ms": 1800, - "audio_bleed_gain_db": 0, + "audio_bleed_gain_db": 0.0, "seam_fade_ms": null, "match_levels": null, "width": 960, diff --git a/workflows/templates/minimax/music-video.json b/workflows/templates/minimax/music-video.json index 7e41634f..8e70cfb8 100644 --- a/workflows/templates/minimax/music-video.json +++ b/workflows/templates/minimax/music-video.json @@ -54,7 +54,7 @@ "num_frames": 124, "fps": 24, "audio_bleed_ms": 0, - "audio_bleed_gain_db": 0, + "audio_bleed_gain_db": 0.0, "seam_fade_ms": null, "match_levels": null, "width": 960, diff --git a/workflows/templates/multi-image-reference.json b/workflows/templates/multi-image-reference.json index cab352db..3ee0d637 100644 --- a/workflows/templates/multi-image-reference.json +++ b/workflows/templates/multi-image-reference.json @@ -4,7 +4,7 @@ "prompt": "combine the two images into a Photorealistic photograph of two friends posing together in front of the Eiffel Tower in Paris. On the left: a realistic bird creature with full anthropomorphic body, standing upright, wearing casual clothes and multi colored reflective sunglasses. On the right: a realistic Cthulhu-inspired humanoid with tentacled face, full body visible, wearing modern casual attire. Both are smiling/friendly, arms around each other's shoulders in buddy pose. Natural daylight, tourist photo style, full body shot, realistic textures and lighting, high detail, DSLR quality.", "num_images_per_prompt": 1, "num_inference_steps": 40, - "guidance_scale": 4, + "guidance_scale": 4.0, "first_image": { "location": "https://pbs.twimg.com/profile_images/1982336432430571520/CchTWvKW_400x400.jpg" }, diff --git a/workflows/templates/upscale-diffusion.json b/workflows/templates/upscale-diffusion.json index 6fe915ff..624e4556 100644 --- a/workflows/templates/upscale-diffusion.json +++ b/workflows/templates/upscale-diffusion.json @@ -13,7 +13,7 @@ "negative_prompt": null, "mode": "x2", "num_inference_steps": 25, - "guidance_scale": 9, + "guidance_scale": 9.0, "noise_level": 20, "device": "cuda" }, From d591404b9c860131a790f0f6ec9120c1c31eaa5c Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 17:24:02 -0500 Subject: [PATCH 007/181] fix(audio): #339, #342 - world_fade_out_ms and slice_audio tail warning #339: assemble-and-score and dissolve-between-shots mixed the world track against the score at a constant level, so a score's written decay sat 20+ dB under the world in its final seconds with no way to fix it short of destroying world_gain everywhere. Adds world_fade_out_ms (default 0, a no-op) and a world_faded fade_audio step between world and mixed. #342: slice_audio silently dropped a source's tail when the requested slice landed a few seconds short of its natural end (e.g. a frame-lattice total that cannot land exactly on a score's length), including whatever was loudest there. _warn_on_slice_trims_tail warns only when the dropped remainder is both short in absolute terms (<10s) and small next to the slice (<5%) and not silence - a deliberate excerpt out of a long recording still passes without warning. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/audio_utils.py | 52 +++++++++++++++++++ workflows/templates/assemble-and-score.json | 18 +++++-- .../templates/dissolve-between-shots.json | 18 +++++-- 3 files changed, 82 insertions(+), 6 deletions(-) diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index a2fb2160..da868015 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -31,6 +31,13 @@ # frame-aligned slicing produces, not a slice that overran its source SLICE_PAD_WARN_MS = 10.0 +# A dropped tail is only the "almost reached the end" signature this warning +# exists for when it is both short next to the slice and short in absolute +# terms - a deliberate excerpt out of a long recording drops most of the +# source and should not warn +SLICE_TRIM_WARN_SECONDS = 10.0 +SLICE_TRIM_WARN_FRACTION = 0.05 + def as_channels_samples(audio): """Normalize a waveform to a (channels, samples) float32 numpy array. @@ -500,6 +507,7 @@ def slice_audio( ) _warn_on_slice_past_end(total, start, length, sample_rate) + _warn_on_slice_trims_tail(waveform, total, start, length, sample_rate) # #309: a cut out of a source that was already near-silent (room tone, # a deliberate quiet bed) is not a defect the slice introduced - measure # the source before cutting it down, so save can tell the two apart from @@ -677,6 +685,50 @@ def _warn_on_slice_past_end(total, start, length, sample_rate): ) +def _warn_on_slice_trims_tail(waveform, total, start, length, sample_rate): + """Say when a slice left material behind that the caller likely wanted. + + slice_audio is a slice, so most unused remainders are deliberate excerpts + and warning on every one would be noise. What #342 found is a narrower + signature: a cut landing a few seconds short of a source's natural end + (a frame-lattice total that cannot land exactly on the score's length) + silently drops the source's tail, including whatever is loudest there. + Only fires when the dropped remainder is both short in absolute terms + and small next to the slice itself, and only when that remainder is not + already silence - a track that legitimately ends in a fade should not + warn just because its last seconds are quiet. + """ + if not sample_rate or length <= 0: + return + slice_end = start + length + remainder = total - slice_end + if remainder <= 0: + return + remainder_seconds = remainder / float(sample_rate) + if remainder_seconds >= SLICE_TRIM_WARN_SECONDS: + return + if remainder_seconds / (length / float(sample_rate)) >= SLICE_TRIM_WARN_FRACTION: + return + dropped = waveform[:, slice_end:total] + peak_dbfs = level_dbfs(dropped, "peak") + if peak_dbfs is None: + # No level at all is silence - nothing was lost + return + emit_warning( + f"slice_audio: the slice ends {remainder_seconds:.2f} s before the " + f"{total / float(sample_rate):.2f} s source does, dropping its tail " + f"(peak {peak_dbfs:.1f} dBFS in the dropped {remainder_seconds:.2f} s) " + f"- if the slice was meant to reach the source's end, adjust " + f"start/length to land there, or fade the source's own tail first", + kind="slice_trimmed_tail", + command="slice_audio", + source_seconds=round(total / float(sample_rate), 3), + dropped_seconds=round(remainder_seconds, 3), + dropped_peak_dbfs=round(peak_dbfs, 1), + sample_rate=sample_rate, + ) + + def resample_waveform(waveform, sample_rate, target_sample_rate): """A waveform at a different rate, as a plain (channels, samples) array. diff --git a/workflows/templates/assemble-and-score.json b/workflows/templates/assemble-and-score.json index d14a3fe8..6e8eac19 100644 --- a/workflows/templates/assemble-and-score.json +++ b/workflows/templates/assemble-and-score.json @@ -1,6 +1,6 @@ { "id": "assemble-and-score", - "description": "Cuts existing shots into one scored film. It is the pass you re-run while cutting, when re-generating the footage each time would cost GPU hours to change a fade. The shots go straight into concat_videos as hard cuts - nothing is stabilized or rescaled on the way, so a deliberate camera move survives the edit; they must already share one size and frame rate, and the task refuses a set that does not. 'shots' is a list, so a diptych is two entries and a reel is however many the cut needs. The shots' own recorded sound is carried up to the score's sample rate and mixed underneath it, so the room tone of each world survives the edit instead of being replaced by music. Supply the shots and the score as asset: references - 'shot_1.mp4' and friends, uploaded with upload_asset or promoted from a generated run with keep_output. Shots generated independently also drift in loudness - 10 dB between two shots of one scene is ordinary - and no seam control can hide a level jump, because it is either side of the cut rather than at it; 'match_levels' ('rms' for perceived level, 'peak' for the loudest sample) evens the shots out before they are joined, and left null, as it is by default, a wide spread is warned about in the log rather than passing in silence. 'match_levels_dbfs' sets the level match_levels moves every shot to, below its own default; a shot whose gain to that level would clip is held short of it instead, and the run says so rather than leaving the shots still apart after a match. 'audio_bleed_gain_db' ducks 'audio_bleed_ms'' full-scale bleed, left at 0 for no change. 'total_frames' is the length of the cut in frames and 'score' is a separate asset with its own length: the score is sliced to 'total_frames' from 'score_start_frame', and a score that does not reach that far is padded to it with digital silence, which leaves the rest of the film unscored under the shots' own sound (the run says so as a 'slice_past_end' warning, but the film itself sounds plausible). A score must therefore be at least as long as the cut; to stretch a short bed to reach, make a longer one with the 'loop_audio' task first and pass that as 'score'.", + "description": "Cuts existing shots into one scored film. It is the pass you re-run while cutting, when re-generating the footage each time would cost GPU hours to change a fade. The shots go straight into concat_videos as hard cuts - nothing is stabilized or rescaled on the way, so a deliberate camera move survives the edit; they must already share one size and frame rate, and the task refuses a set that does not. 'shots' is a list, so a diptych is two entries and a reel is however many the cut needs. The shots' own recorded sound is carried up to the score's sample rate and mixed underneath it, so the room tone of each world survives the edit instead of being replaced by music. Supply the shots and the score as asset: references - 'shot_1.mp4' and friends, uploaded with upload_asset or promoted from a generated run with keep_output. Shots generated independently also drift in loudness - 10 dB between two shots of one scene is ordinary - and no seam control can hide a level jump, because it is either side of the cut rather than at it; 'match_levels' ('rms' for perceived level, 'peak' for the loudest sample) evens the shots out before they are joined, and left null, as it is by default, a wide spread is warned about in the log rather than passing in silence. 'match_levels_dbfs' sets the level match_levels moves every shot to, below its own default; a shot whose gain to that level would clip is held short of it instead, and the run says so rather than leaving the shots still apart after a match. 'audio_bleed_gain_db' ducks 'audio_bleed_ms'' full-scale bleed, left at 0 for no change. 'total_frames' is the length of the cut in frames and 'score' is a separate asset with its own length: the score is sliced to 'total_frames' from 'score_start_frame', and a score that does not reach that far is padded to it with digital silence, which leaves the rest of the film unscored under the shots' own sound (the run says so as a 'slice_past_end' warning, but the film itself sounds plausible). A score must therefore be at least as long as the cut; to stretch a short bed to reach, make a longer one with the 'loop_audio' task first and pass that as 'score'. The world track runs at a constant level through the score's own ending, so a score that decays into silence can sit 20+ dB under it there and be inaudible even though the mix's peak looks healthy; 'world_fade_out_ms', 0 by default, fades the world track out over its last N ms so the score's written ending is what is heard.", "cost": [ { "device": "cuda", @@ -29,7 +29,8 @@ "score_start_frame": 0, "total_frames": 360, "score_gain": 1.0, - "world_gain": 1.8 + "world_gain": 1.8, + "world_fade_out_ms": 0 }, "steps": [ { @@ -57,6 +58,17 @@ } } }, + { + "name": "world_faded", + "task": { + "command": "fade_audio", + "arguments": { + "audio": "previous_result:world", + "fade_out_ms": "variable:world_fade_out_ms", + "sample_rate": "variable:sample_rate" + } + } + }, { "name": "soundtrack", "task": { @@ -86,7 +98,7 @@ "arguments": { "audios": [ "previous_result:soundtrack_resampled", - "previous_result:world" + "previous_result:world_faded" ], "gains": [ "variable:score_gain", diff --git a/workflows/templates/dissolve-between-shots.json b/workflows/templates/dissolve-between-shots.json index 0c3d18b8..43d0a794 100644 --- a/workflows/templates/dissolve-between-shots.json +++ b/workflows/templates/dissolve-between-shots.json @@ -1,6 +1,6 @@ { "id": "dissolve-between-shots", - "description": "assemble-and-score.json with the cuts softened into cross-dissolves. dissolve_videos overlaps each pair by 'dissolve_frames' rather than butting them together, which suits shots that share a composition - registered on the same centre at the same size, an overlap reads as one thing becoming another rather than as a fade between two pictures. Shots whose framing disagrees will read as a plain cross-fade, so use hard cuts there. As in assemble-and-score the shots are a list and go in untouched - a shot that drifts wants stabilize_video run on it first, as its own step, not on every shot by default. The audio bed is mixed the same way as in assemble-and-score. Shots generated independently drift in loudness and no overlap can hide a level jump, because it is either side of the seam rather than at it; 'match_levels' ('rms' for perceived level, 'peak' for the loudest sample) evens the shots out before they are joined, and left null, as it is by default, a wide spread is warned about in the log rather than passing in silence. 'total_frames' is the length of the joined cut, which a dissolve shortens: every seam eats one 'dissolve_frames' overlap, so n shots of f frames joined with d-frame dissolves run n*f - (n-1)*d frames, not n*f. 'score' is a separate asset with its own length: it is sliced to 'total_frames' from 'score_start_frame', and a score that does not reach that far is padded to it with digital silence, which leaves the rest of the film unscored under the shots' own sound (the run says so as a 'slice_past_end' warning, but the film itself sounds plausible). A score must therefore be at least as long as the cut; to stretch a short bed to reach, make a longer one with the 'loop_audio' task first ('target_frames' and 'fps') and pass that as 'score'.", + "description": "assemble-and-score.json with the cuts softened into cross-dissolves. dissolve_videos overlaps each pair by 'dissolve_frames' rather than butting them together, which suits shots that share a composition - registered on the same centre at the same size, an overlap reads as one thing becoming another rather than as a fade between two pictures. Shots whose framing disagrees will read as a plain cross-fade, so use hard cuts there. As in assemble-and-score the shots are a list and go in untouched - a shot that drifts wants stabilize_video run on it first, as its own step, not on every shot by default. The audio bed is mixed the same way as in assemble-and-score. Shots generated independently drift in loudness and no overlap can hide a level jump, because it is either side of the seam rather than at it; 'match_levels' ('rms' for perceived level, 'peak' for the loudest sample) evens the shots out before they are joined, and left null, as it is by default, a wide spread is warned about in the log rather than passing in silence. 'total_frames' is the length of the joined cut, which a dissolve shortens: every seam eats one 'dissolve_frames' overlap, so n shots of f frames joined with d-frame dissolves run n*f - (n-1)*d frames, not n*f. 'score' is a separate asset with its own length: it is sliced to 'total_frames' from 'score_start_frame', and a score that does not reach that far is padded to it with digital silence, which leaves the rest of the film unscored under the shots' own sound (the run says so as a 'slice_past_end' warning, but the film itself sounds plausible). A score must therefore be at least as long as the cut; to stretch a short bed to reach, make a longer one with the 'loop_audio' task first ('target_frames' and 'fps') and pass that as 'score'. The world track runs at a constant level through the score's own ending, so a score that decays into silence can sit 20+ dB under it there and be inaudible even though the mix's peak looks healthy; 'world_fade_out_ms', 0 by default, fades the world track out over its last N ms so the score's written ending is what is heard.", "cost": [ { "device": "cuda", @@ -27,7 +27,8 @@ "score_start_frame": 0, "total_frames": 360, "score_gain": 1.0, - "world_gain": 1.8 + "world_gain": 1.8, + "world_fade_out_ms": 0 }, "steps": [ { @@ -54,6 +55,17 @@ } } }, + { + "name": "world_faded", + "task": { + "command": "fade_audio", + "arguments": { + "audio": "previous_result:world", + "fade_out_ms": "variable:world_fade_out_ms", + "sample_rate": "variable:sample_rate" + } + } + }, { "name": "soundtrack", "task": { @@ -83,7 +95,7 @@ "arguments": { "audios": [ "previous_result:soundtrack_resampled", - "previous_result:world" + "previous_result:world_faded" ], "gains": [ "variable:score_gain", From f735b416e7c98864a81936779f14b41dec2fe51f Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 17:32:36 -0500 Subject: [PATCH 008/181] fix(plan): #341 - price a composed child by the arguments its step passes _sub_workflow_paths now yields each composing step's `arguments` beside its path; estimate() folds them over the child's own declared variable defaults and reuses _scalar_driver_shifted (#267) and list re-pricing against that folded view instead of the child's bare defaults, so a child run at a different num_frames/width/height/num_inference_steps than its catalog cost was measured for is marked unpriced (partial: true) rather than quoted at the wrong figure. Co-Authored-By: Claude Sonnet 5 --- dw/plan.py | 46 +++++++++++++++++++++++++++++++++++++++++----- tests/test_plan.py | 36 ++++++++++++++++++++++++++++++++++++ 2 files changed, 77 insertions(+), 5 deletions(-) diff --git a/dw/plan.py b/dw/plan.py index ba0d96be..aa69cf69 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -175,6 +175,12 @@ def list_entries(definition, realized): return entries +# `estimate`'s own `list_entries` parameter (the parent's) shadows this +# function's name in its scope - a child's entries are computed under this +# alias instead (#341) +_list_entries = list_entries + + def cached_steps(definition, realized, arguments, cache_probe): """How many steps the step cache would answer for this run: 0 without asking when the workflow is unseeded (the cache is off then), None @@ -468,7 +474,7 @@ def estimate( children_all_observed = True child_runs = [] child_measured_on = set() - for path in _sub_workflow_paths(expanded): + for path, step_arguments in _sub_workflow_paths(expanded): had_child = True # A builtin is the parent's to price; a local child prices itself raw = read_sub_workflow(path, base_dir, workflow_dir) @@ -499,7 +505,33 @@ def estimate( child = _observed(child_observed, device) if child is None: children_all_observed = False - child = _price(child_cost, device, {}, {}) + child_list_entries = {} + child_measured_entries = {} + child_expanded = {"variables": {}} + if child_definition is not None: + child_measured_entries = _list_entries(child_definition, child_definition) + # The composing step's own `arguments` are what the child + # actually runs with - folded over its declared defaults the + # same way a caller's arguments are, since `expanded` has + # already substituted them to concrete values (#341) + effective_variables = dict(child_definition.get("variables") or {}) + effective_variables.update(step_arguments) + child_expanded = {"variables": effective_variables} + child_list_entries = _list_entries(child_definition, child_expanded) + child = _price( + child_cost, device, child_list_entries, child_measured_entries + ) + if child["basis"] == CATALOG and child_definition is not None and ( + _scalar_driver_shifted( + child_definition, child_expanded, child_list_entries + ) + ): + # A scalar cost_driver the composing step overrode (H3's + # num_frames at 345 against a default of 124, say) is the + # same #267 failure one level down - the child's own + # catalog figure was never measured for the value this + # step actually passes it (#341) + child = {"minutes": None, "basis": UNKNOWN, "measured_on": None} else: child_runs.append(child["runs"]) child_measured_on.add(child["measured_on"]) @@ -588,12 +620,16 @@ def _observed(observed, device): def _sub_workflow_paths(expanded): - """The local (non-builtin) sub-workflow path of every composing step.""" + """The local (non-builtin) sub-workflow path and composing arguments of + every composing step, as (path, arguments) - `expanded` has already + substituted and expanded `for_each`, so each occurrence carries the + concrete arguments that step actually passes the child (#341).""" for step in expanded.get("steps") or []: reference = step.get("workflow") if isinstance(step, dict) else None path = reference.get("path") if isinstance(reference, dict) else None if isinstance(path, str) and not path.startswith(BUILTIN_PREFIX): - yield path + arguments = reference.get("arguments") + yield path, arguments if isinstance(arguments, dict) else {} def _only_composes_children(definition): @@ -702,7 +738,7 @@ def downloads_required(expanded, base_dir, workflow_dir, cache_dir, lookup_sizes names = [] urls = [] _collect_sources(expanded, names, urls) - for path in _sub_workflow_paths(expanded): + for path, _arguments in _sub_workflow_paths(expanded): raw = read_sub_workflow(path, base_dir, workflow_dir) if raw is None: continue diff --git a/tests/test_plan.py b/tests/test_plan.py index cbe53d4f..663e975a 100644 --- a/tests/test_plan.py +++ b/tests/test_plan.py @@ -630,6 +630,42 @@ def test_a_child_measured_on_another_device_is_still_added(self, plan, tmp_path) # the parent's own basis is what is reported assert answer["basis"] == "catalog" + def test_a_composed_childs_shifted_scalar_driver_is_unpriced(self, plan, tmp_path): + """The composing step's own `arguments` are what the child actually + runs with, not its declared defaults - a scalar cost_driver moved + away from the value the child's catalog cost was measured against + is the same #267 failure one level down (#341).""" + child = { + "id": "child", + "cost": [cost("cuda", 5)], + "cost_drivers": ["num_frames"], + "variables": {"num_frames": 124}, + "steps": [], + } + (tmp_path / "child.json").write_text(json.dumps(child)) + parent = composing("child.json") + parent["steps"][1]["workflow"]["arguments"] = {"num_frames": 345} + answer = plan(parent)["estimate"] + assert (answer["minutes"], answer["partial"]) == (2.0, True) + assert answer["unpriced"] == ["child.json"] + + def test_a_composed_child_at_its_default_is_still_priced(self, plan, tmp_path): + """The companion case: a composing step that passes the child's own + default for a declared scalar driver is priced normally (#341).""" + child = { + "id": "child", + "cost": [cost("cuda", 5)], + "cost_drivers": ["num_frames"], + "variables": {"num_frames": 124}, + "steps": [], + } + (tmp_path / "child.json").write_text(json.dumps(child)) + parent = composing("child.json") + parent["steps"][1]["workflow"]["arguments"] = {"num_frames": 124} + answer = plan(parent)["estimate"] + assert (answer["minutes"], answer["partial"]) == (7.0, False) + assert answer["unpriced"] == [] + def test_a_childs_observed_history_rolls_up_when_the_parent_has_none( self, plan, tmp_path ): From aa67bf2a78c7ebf1f7cf5f26061da2a5b1d2fcc4 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 17:47:31 -0500 Subject: [PATCH 009/181] fix(events): #343 - make a from_pretrained download visible on the run A job stuck downloading a model looked identical to a hang: the phase-stall watchdog had nothing but silence to report for as long as the pull took, and neither list_downloads nor the job's own progress said a download was even happening. Per the issue's disposition, this does not intercept or pre-fetch the download (a snapshot_download call could pull more than from_pretrained would for the same repo, and get it wrong for a variant or an alternate weight file). Instead, dw/download_watch.py watches the Hugging Face cache directory the download writes into for the duration of the loading phase - the same *.incomplete blobs hub_cache.py already scans for the explicit model-download route - and emits a throttled download_progress event while bytes are growing. That growth resets the watchdog's silence clock the same as any other event, so a real, ongoing download no longer reads as a stall; a download that genuinely stops still does. dw/pipeline_processors/pipeline.py wraps the one from_pretrained call site (load_component) in the watch; from_single_file and already-cached loads are untouched. dw/server/jobs.py folds download_progress into the job's existing phase_detail field (" downloading : 12.3 GB, 38 MB/s"), so get_job/wait_for_job surface it with no new field for a consumer to learn. Listing a job's downloads in list_downloads is issue #357's territory, not this one. Co-Authored-By: Claude Sonnet 5 --- dw/download_watch.py | 162 +++++++++++++++++++++++++++++ dw/pipeline_processors/pipeline.py | 10 +- dw/server/jobs.py | 11 ++ tests/test_download_watch.py | 78 ++++++++++++++ 4 files changed, 258 insertions(+), 3 deletions(-) create mode 100644 dw/download_watch.py create mode 100644 tests/test_download_watch.py diff --git a/dw/download_watch.py b/dw/download_watch.py new file mode 100644 index 00000000..e20b493a --- /dev/null +++ b/dw/download_watch.py @@ -0,0 +1,162 @@ +"""Makes a `from_pretrained` download visible on the run's event stream (#343). + +`from_pretrained` downloads with no hook of its own, and a pull that runs +long looks identical to a hang: the phase-stall watchdog (dw/events.py) has +nothing but silence to report. Rather than intercepting the download - which +would mean second-guessing what `from_pretrained` fetches, and getting it +wrong for a variant or an alternate weight file it would not have pulled - +this only watches the Hugging Face cache directory the download writes into +while a `loading` phase is in progress, the same directory `hub_cache.py` +scans for `list_downloads`. Byte growth there is real progress whoever +triggered it, and is reported as such. +""" + +import logging +import os +import threading +import time + +from huggingface_hub.constants import HF_HUB_CACHE +from huggingface_hub.file_download import repo_folder_name +from huggingface_hub.utils import HFValidationError, validate_repo_id + +logger = logging.getLogger("dw") + +# How often the watcher re-measures the cache directory, and the minimum gap +# between emitted events - well under the phase-stall threshold (30s) so a +# real, ongoing download never trips it, but not so tight that a run emits a +# progress event on every tick of a fast-growing directory. +CHECK_INTERVAL_SECONDS = 1.0 +EMIT_INTERVAL_SECONDS = 5.0 + + +def is_watchable_repo_id(name): + """Whether name is shaped like a Hugging Face repo id - a local path or + checkout is not something a cache directory watch means anything for.""" + try: + validate_repo_id(name) + return True + except HFValidationError: + return False + + +def _blob_dir_size(blob_dir): + total = 0 + try: + with os.scandir(blob_dir) as entries: + for entry in entries: + try: + if entry.is_file(follow_symlinks=False): + total += entry.stat(follow_symlinks=False).st_size + except OSError: + # A blob renamed or removed mid-scan (a completed + # .incomplete file, for instance) is not a fault + continue + except FileNotFoundError: + return 0 + return total + + +class DownloadWatch: + """Context manager that watches one repo's cache directory for growth + for as long as it is open, emitting a throttled `download_progress` + event on the active run context while bytes are arriving. + + Not a download itself - `from_pretrained` runs unmodified inside the + `with` block. Growth with nothing watching it (a cache miss the caller + already had, or a repo id from_pretrained resolves differently than + expected) simply emits nothing, same as no watch at all. + """ + + def __init__(self, repo_id, context, cache_dir=None, repo_type="model"): + self._repo_id = repo_id + self._context = context + resolved_cache_dir = cache_dir or HF_HUB_CACHE + self._blob_dir = os.path.join( + resolved_cache_dir, + repo_folder_name(repo_id=repo_id, repo_type=repo_type), + "blobs", + ) + self._stop = threading.Event() + self._thread = None + + def __enter__(self): + self._thread = threading.Thread( + target=self._run, daemon=True, name="dw-download-watch" + ) + self._thread.start() + return self + + def __exit__(self, *exc_info): + self._stop.set() + if self._thread is not None: + self._thread.join(timeout=CHECK_INTERVAL_SECONDS * 2) + return False + + def _run(self): + baseline = _blob_dir_size(self._blob_dir) + last_size = baseline + last_sample_at = time.monotonic() + last_emit_at = 0.0 + while not self._stop.wait(CHECK_INTERVAL_SECONDS): + try: + size = _blob_dir_size(self._blob_dir) + now = time.monotonic() + if size <= last_size: + last_size = size + last_sample_at = now + continue + elapsed = now - last_sample_at + rate = (size - last_size) / elapsed if elapsed > 0 else None + last_size = size + last_sample_at = now + if now - last_emit_at < EMIT_INTERVAL_SECONDS: + continue + last_emit_at = now + self._context.emit( + "download_progress", + repo_id=self._repo_id, + downloaded_bytes=max(0, size - baseline), + bytes_per_second=rate, + ) + except Exception as e: + # A watch that breaks must not take the download down with it + logger.debug(f"Download watch for '{self._repo_id}' failed: {e}") + return + + +def _format_bytes(n): + for unit in ("B", "KB", "MB", "GB", "TB"): + if n < 1024 or unit == "TB": + return f"{n:.1f} {unit}" if unit != "B" else f"{n:.0f} {unit}" + n /= 1024 + + +def format_progress(repo_id, downloaded_bytes, bytes_per_second): + """The `phase_detail` text a job's progress reports while this repo is + downloading - `downloading : 12.3 GB, 38 MB/s`, per #343.""" + text = f"downloading {repo_id}: {_format_bytes(downloaded_bytes)}" + if bytes_per_second: + text += f", {_format_bytes(bytes_per_second)}/s" + return text + + +def watch(repo_id, cache_dir=None, repo_type="model"): + """A DownloadWatch on the active run context, or a no-op context manager + when repo_id is not shaped like a hub repo (a local path, for instance).""" + from .events import get_context + + if not is_watchable_repo_id(repo_id): + return _NULL_WATCH + return DownloadWatch(repo_id, get_context(), cache_dir=cache_dir, repo_type=repo_type) + + +class _NullWatch: + def __enter__(self): + return self + + def __exit__(self, *exc_info): + return False + + +_NULL_WATCH = _NullWatch() diff --git a/dw/pipeline_processors/pipeline.py b/dw/pipeline_processors/pipeline.py index f8972457..c7295983 100644 --- a/dw/pipeline_processors/pipeline.py +++ b/dw/pipeline_processors/pipeline.py @@ -27,6 +27,7 @@ # imported where they are used - at module scope they add seconds to every startup from ..events import WorkflowCancelled, emit_phase, emit_warning, get_context +from .. import download_watch from huggingface_hub.errors import HfHubHTTPError logger = logging.getLogger("dw") @@ -1820,9 +1821,12 @@ def load_component( model_name = from_pretrained_arguments.pop("model_name") logger.info(f"Loading {component_name} from model: {model_name}") emit_phase("loading", detail=f"{component_name}: {model_name}") - component = component_type.from_pretrained( - model_name, **from_pretrained_arguments - ) + with download_watch.watch( + model_name, cache_dir=from_pretrained_arguments.get("cache_dir") + ): + component = component_type.from_pretrained( + model_name, **from_pretrained_arguments + ) # Load from single file elif "from_single_file" in from_pretrained_arguments: diff --git a/dw/server/jobs.py b/dw/server/jobs.py index 69fee683..2b923671 100644 --- a/dw/server/jobs.py +++ b/dw/server/jobs.py @@ -17,6 +17,7 @@ import logging import threading +from ..download_watch import format_progress from ..repl_worker import WorkerManager from ..workflow import SEED_BITS, workflow_from_file, workflow_from_definition from ..introspection import workflow_argument_warnings @@ -581,6 +582,16 @@ def _note_progress(self, event): elif kind == "pipeline_step": self.denoise_step = event.get("step") self.denoise_total_steps = event.get("total_steps") + elif kind == "download_progress": + # Folded into phase_detail rather than a field of its own - a + # poller already reads phase_detail for what the loading phase + # is waiting on, and the next "phase" event (loading ending) + # overwrites it same as any other detail (#343) + self.phase_detail = format_progress( + event.get("repo_id"), + event.get("downloaded_bytes"), + event.get("bytes_per_second"), + ) elif kind == "warning": # Both channels, on purpose: the event log keeps the moment it # happened, `warnings` keeps it where a caller who polled the diff --git a/tests/test_download_watch.py b/tests/test_download_watch.py new file mode 100644 index 00000000..03cb6bf9 --- /dev/null +++ b/tests/test_download_watch.py @@ -0,0 +1,78 @@ +"""A download from_pretrained triggers is watched, not intercepted (#343): +byte growth in the Hugging Face cache counts as progress, so a real download +does not trip the phase-stall watchdog, and a download that genuinely stops +still does.""" + +import time +from unittest.mock import patch + +from dw import download_watch +from dw import events as events_module +from dw.events import RunContext + + +def _fast_watchdog(): + return patch.multiple( + events_module, + PHASE_STALL_THRESHOLD_SECONDS=0.1, + PHASE_STALL_CHECK_INTERVAL_SECONDS=0.02, + ) + + +def _fast_watch(monkeypatch): + monkeypatch.setattr(download_watch, "CHECK_INTERVAL_SECONDS", 0.02) + monkeypatch.setattr(download_watch, "EMIT_INTERVAL_SECONDS", 0.05) + + +def test_growing_download_emits_progress_and_suppresses_stall(tmp_path, monkeypatch): + _fast_watch(monkeypatch) + repo_id = "some-org/some-model" + blob_dir = tmp_path / "models--some-org--some-model" / "blobs" + blob_dir.mkdir(parents=True) + blob_file = blob_dir / "abc123.incomplete" + + events = [] + context = RunContext(on_event=events.append) + with _fast_watchdog(): + context.enter_run() + try: + context.note_phase("loading") + with download_watch.DownloadWatch(repo_id, context, cache_dir=str(tmp_path)): + for _ in range(6): + with open(blob_file, "ab") as f: + f.write(b"x" * 4096) + time.sleep(0.06) + finally: + context.exit_run() + + progress = [e for e in events if e["event"] == "download_progress"] + assert progress, "growing bytes must be reported as download_progress" + assert progress[0]["repo_id"] == repo_id + assert all(e["downloaded_bytes"] > 0 for e in progress) + + stalls = [e for e in events if e.get("kind") == "phase_stall"] + assert not stalls, "a real, ongoing download must not read as a stall" + + +def test_stalled_download_still_stalls(tmp_path, monkeypatch): + _fast_watch(monkeypatch) + repo_id = "some-org/some-model" + blob_dir = tmp_path / "models--some-org--some-model" / "blobs" + blob_dir.mkdir(parents=True) + blob_file = blob_dir / "abc123.incomplete" + blob_file.write_bytes(b"x" * 4096) # written once, never grows again + + events = [] + context = RunContext(on_event=events.append) + with _fast_watchdog(): + context.enter_run() + try: + context.note_phase("loading") + with download_watch.DownloadWatch(repo_id, context, cache_dir=str(tmp_path)): + time.sleep(0.3) + finally: + context.exit_run() + + assert not [e for e in events if e["event"] == "download_progress"] + stalls = [e for e in events if e.get("kind") == "phase_stall"] + assert stalls, "bytes that stopped growing must still read as a stall" From b0965b3443bd9937a85a84c74453f39594e71898 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 17:58:23 -0500 Subject: [PATCH 010/181] fix(minimax): #344 - release H3 pipeline after the shot list before assembly dialogue-short and music-video never set release_pipeline on their shot for_each step, so H3 (and its group-offloaded weights) stayed resident through concat_videos/pair_audio/normalize_audio - a composed multi-shot run could SIGKILL at or after assembly with the pipeline still loaded. for_each carries release_pipeline/release_models onto only the last expanded member, so setting it once on the template step is enough. Updates the minimax-h3 skill's existing release_pipeline guidance to name the shot-before-assembly case explicitly. Co-Authored-By: Claude Sonnet 5 --- plugins/dw/skills/minimax-h3/SKILL.md | 7 ++++--- workflows/templates/minimax/dialogue-short.json | 1 + workflows/templates/minimax/music-video.json | 1 + 3 files changed, 6 insertions(+), 3 deletions(-) diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 9e81013e..145eb795 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -117,9 +117,10 @@ read the `workflows` guide's authoring section first. cuts is laid under the concat. - H3 is guidance-distilled: no `guidance_scale`, no negative prompt. Say what is there, not what is not. -- When deriving a variant, keep `release_pipeline` where the template puts it: - it frees the Z-Image boards before H3 loads. A run SIGKILLed near the end in - a warm worker but fine in a fresh one is host memory, not the prompt. +- Keep `release_pipeline` where the template puts it: it frees Z-Image before + H3 loads, and frees H3 itself on `shot` before a concat runs. A run + SIGKILLed near the end in a warm worker but fine in a fresh one is host + memory, not the prompt. - Ref2VA limits: at most 9 images, 3 videos, 3 audio clips, 12 files; audio can never be the only reference. References are labelled in order. - Music3 reads `audio_duration` as a ceiling, not a target: ask for more than diff --git a/workflows/templates/minimax/dialogue-short.json b/workflows/templates/minimax/dialogue-short.json index b7683ee1..930862aa 100644 --- a/workflows/templates/minimax/dialogue-short.json +++ b/workflows/templates/minimax/dialogue-short.json @@ -202,6 +202,7 @@ { "name": "shot", "for_each": "variable:shots", + "release_pipeline": true, "pipeline": { "configuration": { "component_type": "ModularPipeline", diff --git a/workflows/templates/minimax/music-video.json b/workflows/templates/minimax/music-video.json index 8e70cfb8..a30e8810 100644 --- a/workflows/templates/minimax/music-video.json +++ b/workflows/templates/minimax/music-video.json @@ -155,6 +155,7 @@ { "name": "shot", "for_each": "variable:shots", + "release_pipeline": true, "pipeline": { "configuration": { "component_type": "ModularPipeline", From cf96edf1fc18951336696483b89518a6feef837b Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 18:21:21 -0500 Subject: [PATCH 011/181] fix(introspection): #345 - refuse a component_type/scheduler_type/config_type that would die 3s into the run validate_workflow answered valid: true for a pipeline step naming a class diffusers does not export - the field is resolved lazily deep inside pipeline construction, so the job queued, the worker loaded a checkpoint the plan had already quoted a download for, and only then died on diffusers' own AttributeError. component_type_errors (dw/introspection.py) walks a step's pipeline for every component_type/scheduler_type/config_type and resolves it the same way the run itself does for a '*_type' value - type_helpers.load_type_from_name, gated by security.py's TRUSTED_TOP_LEVEL_PACKAGES - rather than load_allowed_class's narrower ALLOWED_MODULES (sdnq only), which would refuse real catalog entries the run already loads today (transformers.*, dw.community_pipelines.*). A name outside the trusted ecosystem entirely reports the UntrustedWorkflowError message; a name that is merely wrong gets "does not exist" plus up to three difflib suggestions - distinct messages for the two cases. No change to dw/plan.py: component_type never fed downloads_required, and POST /api/validate skips build_plan whenever validation_errors is non-empty, so a refused step already quoted nothing. Wired into Workflow.validation_errors beside task_signature_errors (#285), the same shape. Swept every workflows/**/*.json and dw/workflows/*.json entry clean. Co-Authored-By: Claude Sonnet 5 --- dw/introspection.py | 120 ++++++++++++++++++++ dw/workflow.py | 8 +- tests/test_component_type_errors.py | 167 ++++++++++++++++++++++++++++ 3 files changed, 294 insertions(+), 1 deletion(-) create mode 100644 tests/test_component_type_errors.py diff --git a/dw/introspection.py b/dw/introspection.py index 32ef3fb9..49ae4782 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -14,6 +14,7 @@ import re import inspect import logging +import difflib from .variables import undeclared_variable_references logger = logging.getLogger("dw") @@ -635,6 +636,125 @@ def report(key, message): return errors +_TYPE_REFERENCE_KEYS = ("component_type", "scheduler_type", "config_type") + +# A class-name-shaped string, bare or dotted - excludes a {}-escaped literal +# and a variable:/constant:/asset:/... reference, which use ':' or braces +# and are checked elsewhere +_DOTTED_NAME_PATTERN = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)*$") + + +def _type_reference_candidates(key): + """Names to suggest a close match from, keyed by which field was wrong.""" + if key == "component_type": + return list_pipelines() + list_classes("models") + if key == "scheduler_type": + return list_classes("schedulers") + return list_classes("quantization") + + +def _type_reference_error(key, value, path): + """One component_type/scheduler_type/config_type value, checked against + the resolver the run itself uses for a '*_type' value + (type_helpers.load_type_from_name) - not load_allowed_class's narrower + ALLOWED_MODULES, which would refuse names the catalog already relies on + (e.g. 'transformers.AutoProcessor', 'dw.community_pipelines...') that + TRUSTED_TOP_LEVEL_PACKAGES lets the run itself load. Using the real + resolver is what makes #345's own invariant hold: this can never refuse + a name that would in fact have run. + + Returns an error dict ({path, message}), or None if `value` would resolve. + """ + if not isinstance(value, str) or not _DOTTED_NAME_PATTERN.match(value): + return None + + from .type_helpers import load_type_from_name + from .security import UntrustedWorkflowError + + try: + load_type_from_name(value) + except UntrustedWorkflowError as e: + return {"path": path, "message": str(e)} + except (ImportError, AttributeError, ValueError): + class_name = value.rsplit(".", 1)[-1] + suggestions = difflib.get_close_matches( + class_name, _type_reference_candidates(key), n=3, cutoff=0.6 + ) + message = f"{key} {value!r} does not exist" + if suggestions: + message += f" (closest matches: {', '.join(suggestions)})" + return {"path": path, "message": message} + return None + + +def _walk_type_references(node, path, errors): + if isinstance(node, dict): + for key in _TYPE_REFERENCE_KEYS: + if key in node: + error = _type_reference_error(key, node[key], path + (key,)) + if error is not None: + errors.append(error) + for k, v in node.items(): + _walk_type_references(v, path + (k,), errors) + elif isinstance(node, list): + for i, item in enumerate(node): + _walk_type_references(item, path + (i,), errors) + + +def component_type_errors(workflow_definition, source_indices=None): + """Every component_type/scheduler_type/config_type in a step's pipeline + naming a class the run itself could not load, as [{path, message}] - a + misspelled class used to validate clean and only die ~3s into the run, + after the worker had already loaded a checkpoint the plan's + downloads_required quoted for a pipeline that could never exist (#345). + + A class outside the trusted ecosystem entirely (UntrustedWorkflowError, + see _type_reference_error) is reported with a distinct message from one + that is merely spelled wrong - "not allowed" is not "does not exist". + + The definition handed here has already been substituted and expanded, + matching task_signature_errors; source_indices maps each expanded step + back to the step the author wrote. + """ + from .for_each import MEMBER_SEPARATOR, render_path + + steps = workflow_definition.get("steps") + if not isinstance(steps, list): + return [] + + errors = [] + for index, step in enumerate(steps): + if not isinstance(step, dict): + continue + pipeline = step.get("pipeline") + if not isinstance(pipeline, dict): + continue + source = ( + source_indices[index] + if source_indices is not None and index < len(source_indices) + else index + ) + name = step.get("name") + where = ( + f" in member '{name}'" + if isinstance(name, str) and MEMBER_SEPARATOR in name + else "" + ) + found = [] + _walk_type_references(pipeline, ("pipeline",), found) + for error in found: + full_message = f"{error['message']}{where}" + if not full_message.endswith("."): + full_message += "." + errors.append( + { + "path": render_path(("steps", source) + error["path"]), + "message": full_message, + } + ) + return errors + + def _resolved_value(arguments, key, values): """`arguments[key]` as a number, resolving a `variable:name` reference against `values` (declared defaults merged with the caller's own diff --git a/dw/workflow.py b/dw/workflow.py index e65e173b..d9685f30 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -32,7 +32,7 @@ from .reference_limits import reference_limit_errors from .adapter_compatibility import adapter_errors, warn_adapters from .elision import elide_definition, warn_elided -from .introspection import task_signature_errors +from .introspection import task_signature_errors, component_type_errors from .task_domains import task_argument_errors from .select_validation import select_errors from .variable_constraints import ( @@ -689,6 +689,12 @@ def validation_errors(self, arguments=None, composing=None): # is the one mistake a free pre-flight most obviously exists for # (dw/introspection.py, #141) + task_signature_errors(expanded, source_indices) + # A step's pipeline names a component_type/scheduler_type/ + # config_type that does not exist (or is outside the trusted + # ecosystem entirely) - validated clean and died 3s into the run + # after a checkpoint the plan had already quoted for downloading + # (dw/introspection.py, #345) + + component_type_errors(expanded, source_indices) # A value outside a rule the workflow declares - the bound that # cost 138 s of loading to discover, refused for free at the # path the value sits at (dw/variable_constraints.py, #96) diff --git a/tests/test_component_type_errors.py b/tests/test_component_type_errors.py new file mode 100644 index 00000000..3aed385d --- /dev/null +++ b/tests/test_component_type_errors.py @@ -0,0 +1,167 @@ +"""A pipeline step whose component_type/scheduler_type/config_type diffusers +does not export. + +`validate_workflow` used to answer `valid: true` for a misspelled class name - +`configuration.component_type` is resolved lazily, deep inside pipeline +construction, so the job queued, the worker loaded a checkpoint the plan had +already quoted a download for, and only then died on diffusers' own +AttributeError, ~3s in (#345). This mirrors #285's task_signature_errors: +checked against the same resolver the run itself uses for a '*_type' value +(`type_helpers.load_type_from_name`), so the rule can never refuse a name +that would in fact have run. +""" + +import json +import pathlib + +import pytest + +from dw.introspection import component_type_errors + + +def pipeline_step(component_type, name="a", extra=None): + step = { + "name": name, + "pipeline": {"configuration": {"component_type": component_type}}, + "result": {"content_type": "image/png"}, + } + if extra: + step["pipeline"].update(extra) + return step + + +def errors_for(component_type): + return component_type_errors({"id": "ct", "steps": [pipeline_step(component_type)]}) + + +class TestAMisspelledClass: + def test_it_is_refused_with_a_suggestion(self): + errors = errors_for("FluxPipelin") + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].pipeline.configuration.component_type" + assert "does not exist" in errors[0]["message"] + assert "FluxPipeline" in errors[0]["message"] + + def test_a_real_class_is_accepted(self): + assert errors_for("FluxPipeline") == [] + + def test_a_real_dotted_class_outside_the_narrow_allowlist_is_accepted(self): + """These name modules that are not on introspection's own + ALLOWED_MODULES (sdnq only) but are on the runtime's broader + TRUSTED_TOP_LEVEL_PACKAGES, and are real, shipped catalog entries - + the check must resolve them the way the run itself does, not refuse + something that would have worked.""" + assert errors_for("transformers.AutoProcessor") == [] + assert ( + errors_for( + "dw.community_pipelines.pipeline_flux_rf_inversion." + "RFInversionFluxPipeline" + ) + == [] + ) + + +class TestAllowlistedAbsentVsPresentButDisallowed: + def test_an_absent_class_in_a_trusted_module_says_does_not_exist(self): + errors = component_type_errors( + { + "id": "ct", + "steps": [ + pipeline_step( + None, + extra={"quantization_config": {"config_type": "sdnq.NoSuchConfig"}}, + ) + ], + } + ) + assert len(errors) == 1 + assert "does not exist" in errors[0]["message"] + + def test_a_class_in_a_disallowed_module_says_not_allowed(self, monkeypatch): + # This suite trusts workflows by default (conftest.py); the trust + # gate itself is only exercised untrusted, matching + # tests/test_workflow_trust.py's own pattern. + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "0") + errors = errors_for("os.system") + assert len(errors) == 1 + assert "does not exist" not in errors[0]["message"] + assert "outside the ecosystem" in errors[0]["message"] + + def test_the_two_messages_differ(self, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "0") + absent = component_type_errors( + { + "id": "ct", + "steps": [ + pipeline_step( + None, + extra={"quantization_config": {"config_type": "sdnq.NoSuchConfig"}}, + ) + ], + } + )[0]["message"] + disallowed = errors_for("os.system")[0]["message"] + assert absent != disallowed + + +class TestSchedulerAndQuantizationFields: + def test_a_misspelled_scheduler_type_is_refused(self): + definition = { + "id": "ct", + "steps": [ + { + "name": "a", + "pipeline": { + "scheduler": {"configuration": {"scheduler_type": "DDIMSchedulr"}} + }, + "result": {"content_type": "image/png"}, + } + ], + } + errors = component_type_errors(definition) + assert len(errors) == 1 + assert errors[0]["path"] == ( + "steps[0].pipeline.scheduler.configuration.scheduler_type" + ) + + def test_a_real_config_type_is_accepted(self): + definition = { + "id": "ct", + "steps": [pipeline_step(None, extra={"quantization_config": {"config_type": "sdnq.SDNQConfig"}})], + } + assert component_type_errors(definition) == [] + + +class TestNoDownloadIsQuotedForARefusedStep: + def test_component_type_plays_no_part_in_collecting_sources(self): + """POST /api/validate builds `plan` (and its downloads_required) only + when validation_errors is empty - see dw/server/app.py's validate + route - and component_type never feeds `_collect_sources` at all, so + a step that fails this check was never going to contribute a + download either way.""" + from dw.plan import _collect_sources + + definition = {"id": "ct", "steps": [pipeline_step("FluxPipelin")]} + assert errors_for("FluxPipelin") != [] + names, urls = [], [] + _collect_sources(definition, names, urls) + assert names == [] + assert urls == [] + + +class TestTheCatalogItself: + """Every workflow shipped in the repo passes the new check.""" + + @pytest.mark.parametrize( + "path", + sorted( + str(p) + for p in list(pathlib.Path("workflows").rglob("*.json")) + + list(pathlib.Path("dw/workflows").glob("*.json")) + ), + ) + def test_workflow_has_no_component_type_error(self, path): + definition = json.loads(pathlib.Path(path).read_text()) + if not isinstance(definition, dict) or "steps" not in definition: + pytest.skip("not a workflow") + assert component_type_errors(definition) == [] From 09ce39744db58a8af1739fdfbf7e0e73051472eb Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 18:22:53 -0500 Subject: [PATCH 012/181] fix(templates): #346 - drop low_cpu_mem_usage:false from restore-faces diffusers refuses low_cpu_mem_usage:false whenever parallel loading is enabled (dw's default), so the generate step died ~6s in. The key's absence defaults to true, matching every other template. Co-Authored-By: Claude Sonnet 5 --- workflows/templates/restore-faces.json | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/workflows/templates/restore-faces.json b/workflows/templates/restore-faces.json index 22d3afa0..9d1e8761 100644 --- a/workflows/templates/restore-faces.json +++ b/workflows/templates/restore-faces.json @@ -20,8 +20,7 @@ }, "from_pretrained_arguments": { "model_name": "Tongyi-MAI/Z-Image-Turbo", - "torch_dtype": "torch.bfloat16", - "low_cpu_mem_usage": false + "torch_dtype": "torch.bfloat16" }, "arguments": { "prompt": "variable:prompt", From 604882264ab4eba6089a6ff923ef788b0d9d5cf4 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 18:34:51 -0500 Subject: [PATCH 013/181] fix(validate): #347 - refuse a still image's extension on a video argument loop_frames's video argument advertised taking a still image, but a path or asset:/output: reference with an image extension (asset:sheet.png) validated clean and then died inside fetch_video's extension gate seconds into the run. dw/video_extensions.py adds video_extension_errors, wired into validation_errors: for every video-named key or {"media_type": "video", ...} reference whose location's extension is already knowable at validate time (a literal path, or an asset:/output: reference - not a URL, and not a previous_result:/variable:/item:/gather: reference still unresolved), checks it against ALLOWED_VIDEO_EXTENSIONS and names the {"media_type": "image", "location": ...} form when the extension is an image one. Never refuses anything fetch_video's own gate would accept (URLs are ungated there and stay ungated here). loop_frames's docstring now says how to pass a still: the media_type form, or previous_result: of an image step. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/video_utils.py | 7 +- dw/video_extensions.py | 141 +++++++++++++++++++++++++++++++++ dw/workflow.py | 7 ++ tests/test_video_extensions.py | 104 ++++++++++++++++++++++++ 4 files changed, 258 insertions(+), 1 deletion(-) create mode 100644 dw/video_extensions.py create mode 100644 tests/test_video_extensions.py diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index ba92b02f..99aebd9c 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -142,7 +142,12 @@ def loop_frames(video, num_frames): not a defect and blending two frames of a reference sheet would be. Args: - video: A still image, or frames in any shape a result carries + video: Frames in any shape a result carries, or a still image - but + the `video` argument loads *video files* by convention (#347), so + a still on disk has to be passed as + `{"media_type": "image", "location": "asset:x.png"}` rather than + a bare path or `asset:`/`output:` reference; a still made earlier + in the same workflow is `previous_result:` num_frames: How many frames to hand back, one or more Returns: diff --git a/dw/video_extensions.py b/dw/video_extensions.py new file mode 100644 index 00000000..609b7e66 --- /dev/null +++ b/dw/video_extensions.py @@ -0,0 +1,141 @@ +"""The extension of a `video` argument, checked before the run when it can be. + +`fetch_video` (`dw/arguments.py`) loads every argument named `video` or +`*_video`, and every `{"media_type": "video", "location": ...}` reference +whatever it is named, as a video file - and refuses one whose extension is +not `ALLOWED_VIDEO_EXTENSIONS` there, at run time. A still image handed to +one of those arguments (`"video": "asset:sheet.png"`) validated clean and +then failed the job in seconds with "Video file extension not allowed: +.png", after `get_task("loop_frames")` had described the same argument as +taking "a still image, or frames in any shape a result carries" (#347). + +Moved here, into `validation_errors`, for exactly the cases whose extension +is knowable without running anything: a literal file path, or an +`asset:`/`output:` reference whose name carries one - `resolve_path_references` +turns either into a real path before `fetch_video` ever sees it, and the +extension survives that unchanged. A `previous_result:` (or any reference +`expand_for_each` left unresolved) names no file yet and is left to the +existing run-time check, and a URL is left alone too - `fetch_video` never +gates a URL's extension, so refusing one here would refuse something the run +itself accepts. +""" + +from .arguments import ( + CONSTANT_PREFIX, + PROMPT_PREFIX, + is_media_reference, +) +from .for_each import MEMBER_SEPARATOR, render_path +from .security import ALLOWED_IMAGE_EXTENSIONS, ALLOWED_VIDEO_EXTENSIONS + +# Left to the run-time check: not yet resolved to anything an extension can +# be read off, at the point validation walks the expanded definition +_UNRESOLVED_PREFIXES = ("previous_result:", "variable:", "item:", "gather:") + + +def _is_video_key(key): + return isinstance(key, str) and (key == "video" or key.endswith("_video")) + + +def _extension_problem(value): + """Why this string's extension is not one `fetch_video` will accept, or + None - including None for anything whose extension is not yet knowable.""" + if not isinstance(value, str): + return None + if value.startswith(("http://", "https://")): + return None + if value.startswith(_UNRESOLVED_PREFIXES): + return None + if value.startswith(CONSTANT_PREFIX) or value.startswith(PROMPT_PREFIX): + return None + ext = value.rsplit(".", 1) + if len(ext) != 2 or not ext[1] or "/" in ext[1]: + return None + ext = f".{ext[1].lower()}" + if ext in ALLOWED_VIDEO_EXTENSIONS: + return None + if ext in ALLOWED_IMAGE_EXTENSIONS: + return ( + f"'{value}' is a still image, and a video argument loads video " + f"files - pass it as " + f'{{"media_type": "image", "location": "{value}"}} to load it as ' + f"a still, or reference a prior image step with previous_result:" + ) + return f"Video file extension not allowed: {ext}" + + +def _video_key_values(value, path): + """Strings a video-named key hands to `fetch_video` - itself, or each + element of a list of them. A dict here is a `media_type` reference and + is walked by `_video_values` instead, same as `fetch_video` handles it.""" + if isinstance(value, str): + yield path, value + elif isinstance(value, list): + for index, item in enumerate(value): + yield from _video_key_values(item, path + (index,)) + elif isinstance(value, dict): + yield from _video_values(value, path) + + +def _video_values(value, path): + """Every string this walk can attribute to a location `fetch_video` + will load, paired with the path it sits at.""" + if isinstance(value, dict): + if is_media_reference(value) and value.get("media_type") == "video": + location = value.get("location") + if isinstance(location, str): + yield path + ("location",), location + return + for key, item in value.items(): + if _is_video_key(key): + yield from _video_key_values(item, path + (key,)) + else: + yield from _video_values(item, path + (key,)) + elif isinstance(value, list): + for index, item in enumerate(value): + yield from _video_values(item, path + (index,)) + + +def video_extension_errors(workflow_definition, source_indices=None): + """Every video argument whose extension `fetch_video` will refuse, for + every case that extension is knowable before the run, as + [{path, message}]. + + Walks the substituted, expanded definition, the same convention + `reference_name_errors` and `task_argument_errors` follow: `source_indices` + maps an expanded step back to the one the author wrote, and a path inside + a `for_each` member names the member. + """ + steps = workflow_definition.get("steps") + if not isinstance(steps, list): + return [] + + errors = [] + for index, step in enumerate(steps): + if not isinstance(step, dict): + continue + source = ( + source_indices[index] + if source_indices is not None and index < len(source_indices) + else index + ) + name = step.get("name") + where = ( + f" in member '{name}'" + if isinstance(name, str) and MEMBER_SEPARATOR in name + else "" + ) + for path, value in _video_values(step, ()): + problem = _extension_problem(value) + if problem is None: + continue + errors.append( + { + "path": render_path(("steps", source) + path), + "message": f"{problem}{where}", + } + ) + return errors + + +__all__ = ["video_extension_errors"] diff --git a/dw/workflow.py b/dw/workflow.py index d9685f30..0640c55e 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -43,6 +43,7 @@ ) from .subfolders import step_subfolder, subfolder_errors from .reference_names import reference_name_errors +from .video_extensions import video_extension_errors from .content_types import content_type_errors from .scalar_result_validation import scalar_result_errors from .kernel_availability import kernel_availability_errors @@ -645,6 +646,12 @@ def validation_errors(self, arguments=None, composing=None): # by a message that named a valid form and not the objection # (dw/reference_names.py, #162) + reference_name_errors(expanded, source_indices) + # A still image handed to a 'video' argument by path or asset:/ + # output: reference validated clean and then died inside + # fetch_video's extension gate in the first seconds of the run - + # refused here for the cases the extension is already knowable + # (dw/video_extensions.py, #347) + + video_extension_errors(expanded, source_indices) # A result content_type no writer will accept - a bare word like # "video" validated clean and then died inside the writer with a # traceback naming neither the field nor the value diff --git a/tests/test_video_extensions.py b/tests/test_video_extensions.py new file mode 100644 index 00000000..cbc731bd --- /dev/null +++ b/tests/test_video_extensions.py @@ -0,0 +1,104 @@ +"""A video argument's extension, refused for free when it is already knowable. + +#347: `loop_frames`'s `video` argument took a still image by its own +docstring, but `asset:sheet.png` validated clean and then died inside +`fetch_video`'s extension gate in the first seconds of the run. +""" + +import tempfile + +from dw.video_extensions import _extension_problem, video_extension_errors +from dw.workflow import workflow_from_definition + + +def workflow_holding(video): + return { + "id": "holding", + "steps": [ + { + "name": "hold", + "task": { + "command": "loop_frames", + "arguments": {"video": video, "num_frames": 121}, + }, + "result": {"content_type": "video/mp4", "fps": 24}, + } + ], + } + + +class TestTheFault: + def test_a_still_asset_names_the_media_type_form(self): + problem = _extension_problem("asset:sheet.png") + + assert "media_type" in problem + assert "asset:sheet.png" in problem + + def test_a_video_asset_is_fine(self): + assert _extension_problem("asset:clip.mp4") is None + + def test_a_video_output_reference_is_fine(self): + assert _extension_problem("output:t/20260914-171601-adeee23c/i/x.mp4") is None + + def test_an_extension_the_run_would_also_refuse_is_named(self): + assert _extension_problem("asset:notes.txt") is not None + + def test_a_url_is_left_to_the_run(self): + """`fetch_video` never gates a URL's extension, so this pass must not + refuse one either - #347's binding scope.""" + assert _extension_problem("https://example.com/sheet.png") is None + + def test_a_deferred_reference_is_not_yet_knowable(self): + assert _extension_problem("previous_result:make_image") is None + assert _extension_problem("variable:video_path") is None + + def test_a_path_with_no_extension_is_not_this_passs_complaint(self): + assert _extension_problem("asset:sheet") is None + + def test_a_non_string_is_not_this_passs_complaint(self): + assert _extension_problem(None) is None + assert _extension_problem({"media_type": "video", "location": "asset:x.mp4"}) is None + + +class TestTheValidationPass: + def test_a_still_by_key_convention_is_refused_before_the_queue(self): + workflow = workflow_from_definition( + workflow_holding("asset:sheet.png"), tempfile.mkdtemp() + ) + + problems = workflow.validation_errors() + + assert any( + problem["path"] == "steps[0].task.arguments.video" for problem in problems + ) + + def test_the_media_type_form_validates(self): + workflow = workflow_from_definition( + workflow_holding({"media_type": "image", "location": "asset:sheet.png"}), + tempfile.mkdtemp(), + ) + + assert workflow.validation_errors() == [] + + def test_a_media_type_video_reference_checks_its_location(self): + workflow = workflow_from_definition( + workflow_holding({"media_type": "video", "location": "asset:sheet.png"}), + tempfile.mkdtemp(), + ) + + problems = workflow.validation_errors() + + assert any( + problem["path"] == "steps[0].task.arguments.video.location" + for problem in problems + ) + + def test_a_real_video_asset_validates(self): + workflow = workflow_from_definition( + workflow_holding("asset:clip.mp4"), tempfile.mkdtemp() + ) + + assert workflow.validation_errors() == [] + + def test_nothing_is_reported_for_a_definition_with_no_video_arguments(self): + assert video_extension_errors({"steps": [{"name": "a", "task": {}}]}) == [] From f617c194acbf987cd90942570077312395214087 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 18:41:31 -0500 Subject: [PATCH 014/181] fix(memory): #348 - fit host memory projection as fixed+marginal, not per-entry linear A resident for_each's projection divided the whole observed peak (dominated by the one-time pipeline load) by the entry count and multiplied back out, projecting a 12-entry batch at roughly five times its measured peak. With history at two or more distinct list lengths, base and marginal growth are now fit through the smallest and largest observed counts; with history at only one length, the measured peak is projected flat rather than scaled. Also drops the "held resident together" wording for text naming what was actually fit, and suppresses the projection entirely for a task-only workflow with no pipeline step to accumulate. Co-Authored-By: Claude Sonnet 5 --- dw/host_memory_projection.py | 89 ++++++++++++++++++++++------ tests/test_host_memory_projection.py | 51 ++++++++++++---- 2 files changed, 113 insertions(+), 27 deletions(-) diff --git a/dw/host_memory_projection.py b/dw/host_memory_projection.py index 0c29d027..10180d80 100644 --- a/dw/host_memory_projection.py +++ b/dw/host_memory_projection.py @@ -19,6 +19,7 @@ import logging from .for_each import FOR_EACH_KEY +from .plan import _has_seedable_step logger = logging.getLogger("dw") @@ -101,6 +102,52 @@ def _row_count(row, list_entries, variables): return max(counts) if counts else None +def _peak_by_count(rows, list_entries, variables): + """This workflow's history grouped by list length, each count's peak + reduced to its median - the input a fixed+marginal fit reads.""" + by_count = {} + for row in rows: + peak = row.get("host_memory_job_peak_rss_mb") + count = _row_count(row, list_entries, variables) + if isinstance(peak, (int, float)) and count: + by_count.setdefault(count, []).append(peak) + medians = {} + for count, peaks in by_count.items(): + peaks.sort() + medians[count] = peaks[len(peaks) // 2] + return medians + + +def _fit_peak_model(medians): + """A run's peak as fixed (one model load) plus marginal (per extra + entry), fit from this workflow's own history rather than assumed. + + A single model load dominates a resident run's peak - dividing the + whole observed figure by the entry count and multiplying back out + (the old per-entry model) counted that load once per entry instead of + once per run, projecting a 12-entry batch at roughly five times its + measured peak (#348). With history at two or more distinct list + lengths, the base and the marginal are fit through the smallest and + largest observed counts - the same "two extremes" a least-squares fit + would reduce to with only two points, and cheaper than one with more. + With history at only one list length, there is nothing to fit a slope + from: the measured peak is projected flat, which is the conservative + reading of "the measured behavior" the run has actually shown, rather + than guessing a growth rate with a single data point. + + Returns `(base_mb, slope_mb_per_entry, low_count, high_count)`. + """ + counts = sorted(medians) + if len(counts) == 1: + only = counts[0] + return medians[only], 0.0, only, only + low, high = counts[0], counts[-1] + slope = (medians[high] - medians[low]) / (high - low) + slope = max(slope, 0.0) + base = medians[low] - slope * low + return base, slope, low, high + + def _already_survived(rows, list_entries, requested, projected_mb, variables): """Whether a run at least as large as this request already finished on this box at or above the projected peak. @@ -139,6 +186,10 @@ def host_memory_warnings(definition, list_entries, rows, ceiling_mb): requested = _requested_count(list_entries) if requested is None or not rows or not ceiling_mb: return [] + if not _has_seedable_step(definition): + # A task-only workflow loads no model to accumulate across entries - + # the failure mode this module projects for doesn't apply to it + return [] variables = definition.get("variables") or {} peaks = [row["host_memory_job_peak_rss_mb"] for row in rows] peaks = [value for value in peaks if isinstance(value, (int, float))] @@ -148,28 +199,32 @@ def host_memory_warnings(definition, list_entries, rows, ceiling_mb): return [] if releases_between_iterations(definition): projected_mb = max(peaks) - shape = "the largest single iteration observed" + shape = f"~{round(projected_mb)} MB, the largest single iteration observed" else: - per_entry = [] - for row in rows: - peak = row["host_memory_job_peak_rss_mb"] - count = _row_count(row, list_entries, variables) - if isinstance(peak, (int, float)) and count: - per_entry.append(peak / count) - if not per_entry: + medians = _peak_by_count(rows, list_entries, variables) + if not medians: return [] - per_entry.sort() - median_per_entry = per_entry[len(per_entry) // 2] - projected_mb = median_per_entry * requested - shape = f"{requested} entries held resident together" + base, slope, low_count, high_count = _fit_peak_model(medians) + projected_mb = base + slope * requested + if low_count == high_count: + shape = ( + f"~{round(projected_mb)} MB at {requested} entries, based on " + f"this workflow's own {low_count}-entry runs with no larger " + "history yet to project growth from" + ) + else: + shape = ( + f"~{round(projected_mb)} MB at {requested} entries, " + f"extrapolated from runs of {low_count} and {high_count} entries" + ) if projected_mb <= ceiling_mb: return [] if _already_survived(rows, list_entries, requested, projected_mb, variables): return [] return [ - "Projected host memory for this run (~" - f"{round(projected_mb)} MB, {shape}) exceeds this machine's usable RAM " - f"(~{round(ceiling_mb)} MB) - based on {len(rows)} run(s) of this " - "workflow's own history on this machine, not a curated figure. The " - "run is not blocked, but it may be killed by the OS partway through." + f"Projected host memory for this run ({shape}) exceeds this " + f"machine's usable RAM (~{round(ceiling_mb)} MB) - based on " + f"{len(rows)} run(s) of this workflow's own history on this " + "machine, not a curated figure. The run is not blocked, but it " + "may be killed by the OS partway through." ] diff --git a/tests/test_host_memory_projection.py b/tests/test_host_memory_projection.py index 5cf77898..52a31303 100644 --- a/tests/test_host_memory_projection.py +++ b/tests/test_host_memory_projection.py @@ -1,4 +1,4 @@ -"""Host-RAM projection for a list-driven or composed run (#243). +"""Host-RAM projection for a list-driven or composed run (#243, #348). Warn, not refuse: this module never fails a run, only says the caller's own history projects past what this machine's RAM can hold. Each rule below is a @@ -24,13 +24,21 @@ def for_each_workflow(release=False): step = { "name": "shot", "for_each": "variable:shots", - "task": "noop", + "pipeline": {"pipeline_type": "TestPipeline"}, } if release: step["release_pipeline"] = True return {"id": "w", "variables": {"shots": []}, "steps": [step]} +def task_only_workflow(): + return { + "id": "w", + "variables": {"shots": []}, + "steps": [{"name": "shot", "for_each": "variable:shots", "task": "noop"}], + } + + class TestReleaseDetection: def test_no_for_each_step_is_treated_as_released(self): definition = {"id": "w", "variables": {}, "steps": [{"name": "a"}]} @@ -60,19 +68,41 @@ def test_history_with_no_memory_reading_warns_nothing(self): warnings = host_memory_warnings(definition, {"shots": 12}, rows, 10_000 * MB) assert warnings == [] - def test_resident_pipeline_projects_per_entry_times_requested_count(self): + def test_task_only_workflow_warns_nothing(self): + # #348: a task-only utility loads no model to accumulate across + # entries - the failure mode this module projects for doesn't apply + definition = task_only_workflow() + rows = [row(9000 * MB, {"shots": [1, 2, 3]})] + warnings = host_memory_warnings(definition, {"shots": 12}, rows, 8_000 * MB) + assert warnings == [] + + def test_single_list_length_in_history_projects_flat(self): + # #348: with history at only one list length, there is nothing to + # fit a growth rate from - the measured peak is projected flat + # rather than divided per-entry and multiplied back out, which used + # to inflate a 12-entry projection to twelve times a 3-entry peak definition = for_each_workflow(release=False) - # 3-entry runs each peaked at 3000 MB -> 1000 MB/entry; 12 entries - # projects to 12000 MB, over an 8000 MB ceiling rows = [row(3000 * MB, {"shots": [1, 2, 3]}) for _ in range(3)] warnings = host_memory_warnings(definition, {"shots": 12}, rows, 8_000 * MB) + assert warnings == [] + + def test_two_list_lengths_fit_base_plus_marginal_growth(self): + # #348 (DW-17): a fixed model load plus a small per-entry marginal, + # fit through the smallest and largest observed counts. 1 entry + # peaked at 31000 MB, 5 entries at 62000 MB -> a 6-entry request + # should project roughly 70000 MB, not 31000 * 6 = 186000 MB + definition = for_each_workflow(release=False) + rows = [row(31000 * MB, {"shots": [1]}), row(62000 * MB, {"shots": [1, 2, 3, 4, 5]})] + warnings = host_memory_warnings(definition, {"shots": 6}, rows, 8_000 * MB) assert len(warnings) == 1 - assert "12000" in warnings[0] or "12,000" in warnings[0] + assert "69750" in warnings[0] or "69,750" in warnings[0] + assert "186000" not in warnings[0] + assert "held resident together" not in warnings[0] - def test_resident_pipeline_under_the_ceiling_warns_nothing(self): + def test_growth_fit_under_the_ceiling_warns_nothing(self): definition = for_each_workflow(release=False) - rows = [row(3000 * MB, {"shots": [1, 2, 3]}) for _ in range(3)] - warnings = host_memory_warnings(definition, {"shots": 4}, rows, 8_000 * MB) + rows = [row(31000 * MB, {"shots": [1]}), row(62000 * MB, {"shots": [1, 2, 3, 4, 5]})] + warnings = host_memory_warnings(definition, {"shots": 2}, rows, 100_000 * MB) assert warnings == [] def test_released_pipeline_projects_the_largest_single_iteration(self): @@ -118,5 +148,6 @@ def test_default_arguments_row_still_counts_toward_the_projection(self): definition = for_each_workflow(release=False) definition["variables"]["shots"] = [1, 2] rows = [row(2000 * MB, arguments={}) for _ in range(3)] - warnings = host_memory_warnings(definition, {"shots": 32}, rows, 8_000 * MB) + rows.append(row(3200 * MB, {"shots": list(range(16))})) + warnings = host_memory_warnings(definition, {"shots": 32}, rows, 3_000 * MB) assert len(warnings) == 1 From e9e44d5e2c39948922d1461f54e85089200a776b Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 18:51:56 -0500 Subject: [PATCH 015/181] docs: 2026-09-22 quality review and remediation plan Independent read-only review of develop @ 09ce397 (issues #17-#367, git history since the web UI/MCP additions) and a phased plan built from it: security fixes, CI on develop + durable jobs, one prepare pipeline for run/validate/plan, audio value model, fitted estimator, server router split, explicit workflow-language forms, and loop process changes. Co-Authored-By: Claude Opus 5.5 (1M context) --- .../audits/2026-09-22-quality-review.md | 567 ++++++++++++++++++ .../plans/2026-09-22-quality-remediation.md | 155 +++++ 2 files changed, 722 insertions(+) create mode 100644 docs/proposals/audits/2026-09-22-quality-review.md create mode 100644 docs/superpowers/plans/2026-09-22-quality-remediation.md diff --git a/docs/proposals/audits/2026-09-22-quality-review.md b/docs/proposals/audits/2026-09-22-quality-review.md new file mode 100644 index 00000000..9b1a9514 --- /dev/null +++ b/docs/proposals/audits/2026-09-22-quality-review.md @@ -0,0 +1,567 @@ +# diffusers-workflow — independent quality review + +**Reviewed:** `develop` @ `09ce397` (2026-09-22 18:22 CDT), a fresh clone. I also read the GitHub issues (#17–#367: 280 issues, 87 PRs) and the full git history (1,169 commits). +**Reviewer stance:** read-only. I did not call the live server, did not ssh, and did not run the test suite (it needs torch and diffusers-from-git). All paths below are relative to the repo root. + +--- + +## 1. Executive summary + +**Verdict.** The project is ambitious, and at the unit level it is carefully engineered. Many things are done properly: + +- containment and trust gating +- a bound cost acknowledgement +- a graceful worker-crash path +- about 3,500 tests +- a real attempt to keep numbers in the docs honest by pinning them to diffusers symbols + +The growth since the web UI (2026-08-29) and the MCP server (2026-09-01) has still outpaced the architecture: + +- Engine, server and MCP code tripled in 3.5 weeks (13.9k → 42.9k LOC). +- The main server module grew **10×** (`dw/server/app.py` 431 → 4,226 lines). +- The share of commits that are fixes climbed every week: + + | Weeks | Fix share | + |---|---| + | Aug 29–Sep 11 | 17–18% | + | Sep 12–16 | 39% | + | Sep 17–22 | 57% | + +The issue stream is not converging. The loop closes issues quickly (median 2.75 h from filing to verified close), but inflow keeps pace: 50 and 49 new issues on the last two days, and the open backlog is at its highest yet (34). + +Most of that inflow belongs to a few **bug classes produced by the structure**, not to one-off defects: + +- validation that is a hand-kept copy of execution +- an audio subsystem with no single value model +- a cost estimator that grows one special case per issue +- copy-pasted templates + +Each fix closes one parameter combination, and the next session finds the neighbouring one. The code is decent. The main risk is that the process is very good at producing point fixes, which lets the project put off the few structural changes that would stop the flow. + +**Top five recommendations, ranked by leverage:** + +1. **Make validation and execution share one preparation pipeline** (substitute → constrain → expand → elide → realize). Put each static check next to the task, pipeline or writer it guards. + - Today there are 18 validator passes in `Workflow.validation_errors` (`dw/workflow.py:603-720`), plus 2 more error sources and 9 warning sources inside the HTTP route (`dw/server/app.py:1658-1876`). + - They check a separately prepared definition (`expanded_definition`, `workflow.py:439`, a hand copy of `_prepare_definition`, `workflow.py:798`). + - `JobManager.submit` runs a **different, weaker** check that ignores the caller's arguments (`dw/server/jobs.py:818,831` → `loaded.validate()`). + - This is the root of the largest code-bug class ("validates clean, dies at run time"): 50 of 272 loop-era issue titles mention `validate`. +2. **Fix the SSRF and `from_file` holes, and centralise containment** (security, cheap, urgent). + - `requests.get` follows redirects after a one-time host check, and DNS is resolved twice (`dw/tasks/audio_utils.py:361`, `dw/tasks/video_utils.py:443`, diffusers `load_image` via `dw/arguments.py:943`). + - Run-time `from_file` checks only the scheme and an unrooted path (`dw/arguments.py:848-876` → `security.validate_url`). + - There are at least five separate "is inside root" predicates. +3. **Give audio one value type and one meaning for `sample_rate`, at the dispatch boundary.** + - Today `sample_rate` means *relabel* in some commands, *resample* in others, and *silently relabel* in `pair_audio` (`dw/tasks/pair_audio.py:171-173`). + - Clipping is warned about after the fact instead of being prevented. + - About 75 of 272 issues are audio, level or mix issues. + - Separately, decide whether a mastering toolkit (compressor, gate, EQ, spectral analysis in `audio_utils.py`) belongs in this engine at all, rather than in ffmpeg filters. +4. **Put the loop's output under CI, and add a durable job ledger.** + - CI runs on push to `master` and on PRs only (`.github/workflows/ci.yml:3-8`). The loop merges to `develop` and deploys `develop` to lem, so nothing the agents ship is CI-checked until a release PR. `develop` → `master` PRs failed CI 6 times on 09-18/19, and one failing merge landed on `master` (#241). + - Jobs are written to SQLite only when they finish, so a restart erases queued and running jobs (#300; `jobs.py:757-760, 1158-1163`). +5. **Stop growing the cost estimator by special case; redesign it as one fitted model with priors.** + - 15+ issues in 10 days touched it (#91 → #93 → #154 → #242 → #252 → #255 → #267 → #268 → #275 → #301 → #312 → #315 → #319 → #330 → #341). + - It now has six `basis` values, plus tempered, low_confidence, partial and unpriced overlays, plus parent/child roll-up rules (`dw/plan.py:282-402`). + - It duplicates driver logic from `dw/server/observed_cost.py`, and the two copies already diverge. + + While you are at it, split `create_app` (a single 3,524-line closure) into routers. + +--- + +## 2. Quality assessment by area + +| Area | Grade | One-line reason | +|---|---|---| +| Core engine (pipeline loading, chains, for_each, step cache) | B | Capable and heavily featured. Long functions (`Workflow.run` 490 lines, `workflow.py:1010`); model-specific names leak into generic code. | +| Workflow JSON language | C+ | Powerful, but at least 11 string prefixes resolved in different phases, key-name magic, and variable types inferred from defaults. Sharp edges show up as issues (#338, #363, #365). | +| Static validation | C | 20+ hand-maintained validator passes that copy run-time logic. Submit runs a weaker check. Several validators fail open (`except Exception` → `[]`). | +| Cost/memory planning | C | Accreting heuristics, a duplicated driver logic, and a stream of follow-up issues. Well unit-tested, but the model itself is the problem. | +| HTTP server | C+ | Good per-request workspace handling and containment intent. One 3.5k-line closure; no response models; 400-for-everything with `str(e)`; small concurrency hazards. | +| Job manager | C | Clean crash handling, but in-memory queue and running jobs, no hang timeout, ad-hoc SQLite migrations, and the thinnest tests in the codebase (test/code 0.2). | +| MCP surface | B- | Mostly thin over HTTP, which is good. 58 tools and about 14k tokens of prose kept under a ratcheting budget test; tool prose kept in 3 places; parameter aliases instead of one name (#179); mounted mode shares one workspace across sessions (#298). | +| Tasks: audio | C- | No shared audio model; `sample_rate` semantics vary by command; a mastering stack with pure-Python per-sample loops on the single GPU worker. | +| Tasks: image/video | B- | Main task metadata comes from signatures (good), with three hand-kept side tables that drift (#185, #350, #366). Image processors test/code 0.2. | +| Security | B- (design) / C (current) | Serious and thoughtful, with the trust gate and path policies. Containment is not centralised, and there are two URL validators of different strength with a redirect bypass. | +| Tests | B- | Large and fast. Real behaviour tests for DSP, heavily mocked elsewhere. 17% of asserts pin wording; only 2 real generation tests; no property-based tests; not run on `develop`. | +| UI | B- | Reasonable size (20k LOC, 40 vitest files, 6 Playwright specs). Hand-written `types.ts` against untyped server dicts; e2e not in CI. | +| Docs | C+ | Plentiful and mostly accurate, but sprawling (14k lines), mixed with plans and proposals, and restated in code comments and tool docstrings. 395 inline `#NNN` references in 56 of 115 Python files. | +| Catalog/templates | C+ | Useful, but copy-paste siblings (`assemble-and-score` vs `dissolve-between-shots` are 95% identical), so fixes miss siblings. Descriptions are 2–4 KB of prose acting as agent instructions. | + +--- + +## 3. Weaknesses, ranked by leverage + +### W1. Validation is a parallel re-implementation of execution (redesign) + +**What.** Static validation has three layers: + +1. **Engine passes.** `Workflow.validation_errors` (`dw/workflow.py:603-720`) runs the JSON schema, then concatenates 18 error passes from 15 modules: previous_results, subfolders, reference_names, content_types, scalar_result, locations, reference_limits, adapter_compatibility, task_domains, select_validation, introspection ×2, variable_constraints ×2, vram_estimate, kernel_availability, sub_workflow and undeclared variables. +2. **Server-only passes.** The HTTP route adds `argument_errors`, a 107-line `_argument_reference_errors` closure (`app.py:1531`, which the CLI and REPL never run), 9 warning sources and 4 `_inert_*` helpers. +3. **A separate definition.** All of this runs on a definition prepared by `expanded_definition` (`workflow.py:439`), a hand copy of the run path's `_prepare_definition` (`workflow.py:798`). The copy skips `apply_constraints`, constraint-reference resolution, `realize_args` and elision. + +**Evidence.** + +- Nearly every validator's comment names the incident that created it: #89, #96, #136, #139, #140, #141, #155, #162, #168, #178, #212, #265, #345. +- The rules are duplicated, not shared. `select_validation.py:15` hard-codes `_RULES = {"argmax", …}` in parallel with `tasks/select.py:82-106`. +- `_UNRESOLVED_PREFIXES` is redefined 6 times in 3 different variants (`subfolders.py:30`, `content_types.py:29`, `kernel_availability.py:35`, `reference_names.py:34`, `reference_limits.py:48`, `adapter_compatibility.py:57`). +- `JobManager.submit` (`jobs.py:818,831`) calls `loaded.validate()`, which is `validation_errors()` **with no arguments**. So a `run_workflow` that skipped `validate_workflow` is queued without argument-aware checks. +- Validators fail open: `workflow.py:571` and `:733`, and the plan at `app.py:1858`. +- The issue record: #89, #96, #118, #123, #136, #141, #145, #155, #166, #168, #208, #213, #285, #287, #293, and today's #345, #347, #364, #365. + +**Why it matters.** Every new task parameter, pipeline or template creates an unguarded gap until someone runs into it. The agent loop is good at finding these, so they turn into a steady flow of issues. Fail-open validators make this worse: a bug in the validator hides a bug in the workflow. + +**What I'd do (redesign).** + +1. Write one `prepare(definition, arguments) -> PreparedWorkflow` used by `run`, `validate` and `plan`. +2. Replace per-incident modules with a `static_check` hook declared next to each implementation. `register_command` already does this for argument signatures; extend it. +3. Make `submit` run the same argument-aware validation that `validate_workflow` runs. +4. Make validator crashes errors, not silence. +5. Consider validating the schema *after* substitution. Today 55 typed schema fields reject `variable:` and only 16 accept it, which is the whole #363 class. + +### W2. Security: SSRF redirect and rebinding bypass, `from_file` policy gap, containment in five places (fix now, then refactor) + +**Evidence.** + +- **Redirect bypass.** `validate_media_url` (`dw/locations.py:287-321`) resolves the host and refuses internal addresses. The fetches that follow are plain `requests.get(validated_url, timeout=…)` (`audio_utils.py:361`, `video_utils.py:443`) and diffusers `load_image` (`arguments.py:943`, `tasks/gather.py:64`). + - Redirects are followed by default, so a public URL that 302s to `169.254.169.254` or loopback gets through. + - The host is resolved twice, so DNS rebinding works. + - Responses have no size cap. +- **`from_file` at run time.** It goes through `validate_media_location` (`arguments.py:848-876`), which calls `security.validate_url` (scheme only, `security.py:327-351`) and `validate_path` with no roots. + - The static pass does check literal `from_file` values (`locations.py:403-450`). + - A value that arrives by variable override most likely reaches the loader under the weaker policy. That contradicts the module's own promise that "the loaders call the same functions at run time" (`locations.py:30-35`). +- **Containment predicates.** At least five: `security.validate_path` (prefix after realpath), `locations._within`, `workspace._is_within` (commonpath), `dw_mcp/media._confine`, and `dw_mcp/assets.py:72-100`. That is why #113, #114, #124 and #138 each needed a separate fix. +- **Error text.** `POST /api/jobs` maps *every* exception to 400 with `str(e)` (`app.py:1157-1160`; also `:1275`, `:1294`, `:2086`, `:2467`), which can leak absolute server paths despite the #247 and #310 hardening. + +**What I'd do.** + +1. One fetch helper for all remote media: + - `allow_redirects=False`, or re-validate the host on every hop + - pin the resolved IP for the connection + - a byte cap +2. One `contain(path, roots)` function. +3. One URL validator. Delete `security.validate_url`'s direct uses (`plan.py:830`, `arguments.py:871`). +4. Map internal exceptions to 500 with a generic message. +5. Add each case to the security regression suite. + +### W3. The audio subsystem has no single model, and has grown into a mastering suite (rethink scope, refactor core) + +**Evidence.** + +- `AudioTrack` exists only as a return type (`_as_track`, `audio_utils.py:378`). Inputs are `str | AudioVideo | ndarray`, normalised separately per command (`_waveform_and_rate` `:1344`, `_load_tracks_matching_rate` `:837`) and again in `pair_audio`, `concat_videos` and `dissolve_videos`. +- `sample_rate` means different things depending on the command: + + | Where | What `sample_rate` does | + |---|---| + | Single-track commands (`audio_utils.py:1355-1383`) | Relabels, with a warning | + | `concat_videos` (`concat_videos.py:146-174`) | Resample target | + | Mix/crossfade when pinned (`:858`) | Relabels every track | + | `pair_audio` (`pair_audio.py:171-173`) | Relabels silently: the #180 bug class, still present | + +- The issue chains show the same thing: #108 → #180 → #196 → #205 → #287 → #293 for sample rate. For headroom: #158 → #159 → #161 → #174 → #194 → #286 → #295 → #305 → #306 → #323 → #362, and #362 is still open today ("still clips after mp3 encode"). +- Levels are warned about after encode (`result.py:116-160`, `:190-235`) rather than enforced at the write boundary. +- `audio_utils.py` grew 266 → 1,710 lines since 08-29 and has had 36 commits since 09-01. It now contains: + - a compressor, limiter and gate with a pure-Python per-sample envelope loop (`:1494-1513`, about 8M iterations for a 3-minute track, run on the single FIFO GPU worker) + - biquad EQ that depends on scipy only transitively, via controlnet-aux (`:1601-1631`) + - spectral flatness and harmonicity analysis + +**Why it matters.** About 75 of 272 loop-era issues fall in the audio/level/mix area, the largest category. The pattern is always "silently wrong, add a warning", which fixes one combination per issue. + +**What I'd do.** + +1. Coerce every audio input to one `AudioTrack(waveform float32 [C,N], rate, source)` at task dispatch. +2. Make `sample_rate` always mean "resample to"; the relabel escape hatch gets its own name. +3. Enforce a true-peak ceiling at the single write boundary. Measure *after* encode and correct, rather than warning. +4. Push DSP (EQ, dynamics, loudness, LUFS per #361) down to ffmpeg filter graphs through PyAV, which is already a dependency. +5. Before adding more of it (#349 CPU grade, #361 LUFS), decide deliberately how much post-production this project should own. + +### W4. The loop's quality gate is the implementer's local pytest; `develop` isn't CI-checked, and "verified" is repro-shaped (process, high leverage) + +**Evidence.** + +- `ci.yml:3-8` triggers on `push: [master]` and on `pull_request` only. The loop merges straight to `develop` and deploys it. +- CI history: 6 failed `develop` → `master` PR runs on 09-18/19, and a failed push to `master` on 09-19 (#241). +- Playwright e2e is not in CI. +- Only 2 tests run a real generation, gated on an accelerator (`tests/test_worker.py:37-40`). +- #197 is the cautionary tale: + 1. The first fix called `_fit_audio_to_frames` on a CUDA tensor and broke **every** in-memory LTX-2 audio+video save. + 2. The tester caught that and bounced it. + 3. The second fix passed verification. + 4. The next day the original repro came back: the fit had been applied to a local variable, not to the artifact the concat reads. + 5. A fourth hand-off closed it. + + The tester verifies the filed repro. It does not verify the invariant. +- Bounce rate is low (11 of 184 handed-off issues bounced, 6%), and only 4 issues were formally reopened. But 63 issues cite an earlier issue in their title, 55 of them a `status:verified` one. By my hand classification, about 30 of those are "the fix was incomplete or missed a sibling" and about 20 are "the fix broke a suite expectation". + +**What I'd do.** + +1. Run CI on push to `develop`, and gate lem deploys on it: the deploy script checks the commit's status. +2. Add a nightly GPU smoke on lem: one real generation per family, driven by pytest rather than an agent. +3. For "silently wrong" classes, add property or fuzz tests over parameter combinations (Hypothesis on the audio tasks and `validate` ≡ `run` agreement). Point tests are the only kind the suite has today. +4. Ask the tester prompt to verify *the stated root-cause invariant* plus one sibling, not only the repro. + +### W5. The cost/memory estimator grows by special case (redesign) + +**Evidence.** + +- `estimate` (`plan.py:402`) is 189 lines. +- Six `basis` values (`plan.py:282-287`), plus these rules and overlays: + - tempered and low_confidence blending (#301, #319) + - partial and unpriced (#242, #252) + - parent roll-up with minimum runs (#268, #275) + - a shifted driver forces unknown, at parent level (#267) and child level (#341) + - child-observed only when the parent is unpriced (#315) +- `_declared_drivers` and `_driver_comparable` (`plan.py:354-377`) are self-described copies of `observed_cost.py:57/95`, and they already diverge: one buckets a list by length, the other JSON-dumps it. +- The host-memory projection has its own chain: #243 → #254 → #264 → #272 → #274 → #334 → #348. #348 is open: a shared model load is scaled linearly with list length. +- The regression suites needed at least 8 wording changes just to track estimator changes (#172, #276, #304, #307, #312, #320, #327, #330, #331). + +**Why it matters.** This is the most-churned feature in the loop era, and each change breaks suite expectations downstream. The loop spends a lot of its budget here for a figure that is inherently approximate. + +**What I'd do.** + +1. Replace it with one explicit model: `minutes = f(drivers)`, fitted per workflow and device from history, with curated values as a Bayesian prior and a confidence interval. +2. Report `{estimate, low, high, n}` and drop the basis taxonomy. +3. Memory: model `fixed + per_entry × n`, not `per_entry × n`. +4. Put driver comparison in one module that both the engine and the server import. + +### W6. `dw/server/app.py` is one 3,524-line closure; responses are untyped (refactor) + +**Evidence.** + +- `create_app` spans `app.py:703-4226`: 64 routes, about 25 helpers and 3 middlewares, all closures. There is no `APIRouter` and 0 `response_model`s. +- Business logic lives in route bodies: + - validation orchestration (`validate_workflow` 219 lines, `:1658-1876`) + - cost acknowledgement (`:985-1045`) + - gallery indexing and orphan scanning (`:2678-2830`) + - zip building (`:3275-3360`) + - frame tiling and budgeting (`gallery_frames` 139 lines) +- A run request validates three times: the route, then `submit`, then the worker. +- Concurrency hazards: + - `_workflow_detail_cache` and `_prompt_detail_cache` are module-global dicts mutated from threadpool routes without a lock (`:212`, `:226-236`, `:527`). + - Workspace delete checks for queued jobs under `manager._lock`, releases it, then deletes (`:1973-1991`), which is a check-then-act race. + - Chunked uploads bypass the size precheck and are read into memory whole (`:3517-3530`). +- The UI's `types.ts` (363 lines) and `api.ts` (673 lines) are hand-written against these untyped dicts. + +**What I'd do.** + +1. Split into routers (jobs, catalog, validate, workspaces, gallery, assets, models, system) and service modules the routers call. +2. Add pydantic response models, and generate the UI's TypeScript types from OpenAPI. +3. A mechanical refactor, well covered by the 8k lines of server tests. It is a good job for the implementer agent *if* it is scoped as "no behaviour change". + +### W7. Job durability and worker hangs (inspect, then fix; small) + +**Evidence.** + +- The queue is only in memory (`jobs.py:757-760`). History rows are written on terminal state only (`:81-88`, `:1180-1186`). +- `shutdown()` (`:1158-1163`) never marks pending or running jobs, so #300 covers the running job too, not just queued ones. +- There is no timeout on a hung worker (`repl_worker.py:115-120`), and cancel is a message a hung worker can't read. +- The step cache lives only in memory, per process (`step_cache.py:39`; #244, #281). +- SQLite migrations are ten try/except `ALTER TABLE`s with no schema version (`jobs.py:94-164`). +- `jobs.py` has test/code 0.2, the lowest of any major module. + +**What I'd do.** + +1. Insert the job row at submit. +2. On startup, mark leftover `queued`/`running` rows as `interrupted`. +3. Add a watchdog: no event for N minutes while a phase is expected to report → kill and restart the worker, and fail the job. +4. Add a `user_version`-based migration. +5. Write the job-lifecycle tests. + +### W8. The workflow language's sharp edges (rethink, incrementally) + +**Evidence.** + +- At least 11 string prefixes resolved in different phases: `variable:`, `constant:`, `item:`, `gather:`, `asset:`, `output:`, `prompt:`, `previous_result:`, `builtin:`, `constraint:`, `frame:`. +- On top of those, dict forms (`from_file`, `from_previous_result`, `{media_type, location}`) and key-name magic: + - `_image`/`_video` suffixes auto-load media + - `_type`/`_dtype`/`dtype` import Python objects, with exceptions `content_type` and `offload_type` (`arguments.py:26,136-171`) + - a `{}` escape + - an `EscapedString` class that exists only so a second realize pass doesn't undo the first (`arguments.py:29`) +- Because `realize_args` also runs over the variables block (`workflow.py:846`), a *variable's* name triggers media loading (#365). +- Variable types come from `type(default)` (`variables.py:277`), which produces #338 (silent int truncation) and #364 (null default). +- `null` means "drop the key" except under `MEDIA_SOURCE_KEYS` (`variables.py:98-112`). + +**Why it matters.** Agents author this language. Every implicit rule is a trap that turns into an issue, then a warning, then docs prose, then a suite case. + +**What I'd do.** + +1. Add declared variable types: `{"type": "float", "default": 1, "enum": […]}`, with the inferred form kept as legacy. +2. Make type import and media loading explicit, e.g. `{"$type": "torch.bfloat16"}` and `{"$media": "asset:x.png"}`, instead of key-name inference. Deprecate the suffix magic behind a schema version. +3. Resolve all references in one documented phase order. + +### W9. Model-family knowledge in generic engine code (refactor) + +**Evidence.** + +- `dw/adapter_compatibility.py` is entirely MiniMax-H3 (`H3_WORKFLOWS = {"ref2va","t2va","fl2va"}`, `:39-47`) and runs on every validation (`workflow.py:672`). +- `pipeline.py:1985` hard-codes `("transformer","transformer_ref")`. +- `prompt_weighting.py:293-315` checks Flux by name, and `teacache.py:291` keys on `FluxTransformer2DModel`. +- `workflow_schema.json:361,367` quotes H3 numbers. + +**Why it matters.** The project's own rule (root `CLAUDE.md`) is "model knowledge lives [in skills and the catalog], never in engine code". Each new family will add more of these. + +**What I'd do.** Add a per-family plugin hook (validators and component-partition hints) registered from a `families/` package, so the engine core stays generic. + +### W10. Template duplication and prose-as-configuration (refactor) + +**Evidence.** + +- `workflows/templates/assemble-and-score.json` and `dissolve-between-shots.json` have identical step lists and are 95% line-similar. +- The sibling-missed-fix issues are the direct result: #129, #196, #205, #215, #227, #302, and #339/#342 open now for both. +- Template `description` fields reach 4 KB (`minimax/dialogue-short.json` 4,068 chars; 43 KB across 83 files). The five skills add 48 KB, and MCP tool descriptions add about 36 KB. +- The same guidance is restated in `docs/*.md`, tool docstrings, handler docstrings, template descriptions and SKILL.md. The MCP prose already drifts: the `run_workflow` handler says to poll `get_job_events` while the server instructions say to use `wait_for_job`. +- `tests/test_mcp_server.py:1450` keeps `SURFACE_BUDGET = 13_890` with 7 tokens of headroom under about 100 lines of changelog comments. Every docstring edit becomes a trimming exercise. + +**What I'd do.** + +1. Compose shared assembly through `builtin:` sub-workflows (the engine already supports them) instead of copying. +2. Make one source for each piece of guidance (the guides served by `get_guide`), and have tool docstrings point to it. +3. Replace the character budget with a real token count and per-tool caps, so the budget stops acting as a ratchet. + +### W11. Comment and docstring archaeology (hygiene) + +**Evidence.** + +- 395 inline `#NNN` references across 56 of 115 Python files (app.py 47, result.py 35, workflow.py 34, audio_utils.py 33, plan.py 33), plus 293 in tests. +- app.py is about 30% comments and docstrings. In `validation_errors`, comments outnumber code. +- Docstrings tell incident stories ("reported succeeded with no warnings (#139)…") instead of stating the invariant. + +**Why it matters.** This is the fingerprint of a fix-per-issue process. It makes the code harder to read and hides the actual contracts. The history belongs in git and the issues; comments should state *what must stay true*. + +**What I'd do.** Leave existing comments alone, but change the implementer prompt's convention: a comment states the invariant, and the commit message carries the history. + +### W12. Tests: numerous, but skewed (inspect) + +**Evidence.** + +- 3,522 tests and 55k LOC, about 1.29 lines of test per line of code. +- 716 mock/patch uses across 70 files. +- About 17% of asserts (984 of 5,787) are `"literal" in message` checks, plus 272 `match=` checks. +- Hard token budgets are pinned (`test_catalog_structure.py:402-451`, `test_mcp_server.py:1450`). +- Test/code ratio for the weakest modules: + + | Module | Test/code | + |---|---| + | `server/jobs.py` | 0.2 | + | `image_utils.py` | 0.2 | + | `pipeline.py` | 0.7 | + | `workflow.py` | 0.7 | + +- The DSP tests are real behaviour tests on synthetic waveforms, which is good. + +**What I'd do.** Move wording pins to error *codes* (structured `{code, path, message}` errors are also better for agents). Add property tests (W4). Test the job lifecycle. Add a small, real-codec end-to-end encode test (PyAV on CPU is cheap). + +--- + +## 4. Quality trajectory + +### 4.1 Where the inflection is + +- **Web UI and HTTP server:** first commit `d50ddce`, 2026-08-29, "feat(server,ui): HTTP server, introspection service, and Svelte SPA" (31 files, +3,314). +- **MCP:** design and spec on 2026-09-01 (`5459435`, `9e9df55`); package moved out of `dw` in `7d64a88` the same day. +- **Plugin skills:** 2026-09-08 (`4ace771`). +- **Agent loop on GitHub Issues:** from 2026-09-12 (issues #69+). + +Before 2026-08 the repo had about 260 commits over 21 months. **August 1–28 was already a big engine push** (dw 7.3k → 13.9k LOC, tests 4.2k → 14.0k), so "before" is not a quiet baseline. + +### 4.2 Size snapshots + +I took the last `develop` commit before each date and counted lines with `git cat-file`. `dw` excludes `dw/community_pipelines` and `dw/workflows`. + +| Date | dw | dw_mcp | ui/src (non-test) | ui tests | tests/ LOC | test fns | docs & *.md lines | +|---|---|---|---|---|---|---|---| +| 2026-08-01 | 7,264 | 0 | 0 | 0 | 4,201 | 236 | 4,008 | +| 2026-08-29 | 13,913 | 0 | 2,223 | 68 | 13,954 | 991 | 5,539 | +| 2026-09-05 | 18,953 | 1,478 | 10,240 | 1,639 | 24,385 | 1,627 | 19,440 | +| 2026-09-12 | 26,684 | 2,613 | 12,328 | 2,647 | 37,614 | 2,456 | 40,722 | +| 2026-09-18 | 35,162 | 3,484 | 13,879 | 4,285 | 49,640 | 3,226 | 49,403 | +| 2026-09-22 (HEAD) | 38,791 | 4,133 | 14,999 | 5,283 | 55,148 | 3,519 | 15,829 | + +- **Code (dw + dw_mcp) grew 3.1×** in 24 days; Python tests grew 4.0×. The test:code LOC ratio rose from 1.00 to 1.29, and the UI went from 0.03 to 0.35. +- **Test volume kept up with code.** Test *effectiveness* did not keep up with the bug classes (W4, W12). +- **Docs** peaked around 49k lines on 09-18, including plans and proposals. They were cut to about 16k on 09-20 (`c5d4ac5` "cleanup docs"), which was a healthy correction. + +Hotspot file growth, in lines: + +| File | 08-29 | 09-05 | 09-12 | 09-17 | 09-22 | +|---|---|---|---|---|---| +| dw/server/app.py | 431 | 1,476 | 2,860 | 3,655 | **4,226** | +| dw/workflow.py | 507 | 733 | 1,209 | 1,705 | 1,820 | +| dw/pipeline_processors/pipeline.py | 1,678 | 1,797 | 2,007 | 2,289 | 2,293 | +| dw/result.py | 850 | 907 | 972 | 1,290 | 1,585 | +| dw/server/jobs.py | 296 | 700 | 1,174 | 1,360 | 1,551 | +| dw_mcp/server.py | — | 439 | 906 | 1,180 | 1,340 | +| dw/tasks/audio_utils.py | 266 | 579 | 804 | 1,463 | **1,710** | +| dw/plan.py | — | — | — | 587 | 889 | + +`pipeline.py` has stabilised (+4 lines in the last week). Growth has moved to the **server, result writer, audio and planner**, which are the surfaces the agents drive. + +**Most-touched files since 08-29:** + +| File | Commits touching it | +|---|---| +| app.py | 122 | +| test_server.py | 102 | +| dw_mcp/server.py | 87 | +| docs/MCP.md | 79 | +| docs/SERVER.md | 72 | +| workflow.py | 68 | +| root CLAUDE.md | 67 | +| jobs.py | 53 | + +Documentation files are among the hottest, because every surface change is restated in 2–4 prose locations. + +### 4.3 Commit mix + +Non-merge commits on `develop`, bucketed by committer date. "Fix" is a case-insensitive `^fix`, `^correct` or `^repair` subject match. + +| Window | Commits | Fix | Fix % | feat/add | Code +/− | Test lines + | test/code added | +|---|---|---|---|---|---|---|---| +| ≤ 2026-07 | 253 | 12 | 4% | 31 | +7,320 / −2,613 | 4,093 | 0.56 | +| Aug 1–28 | 64 | 3 | 4% | 9 | +6,105 / −1,326 | 8,946 | 1.47 | +| Aug 29–Sep 4 | 180 | 32 | 17% | 67 | +21,806 / −3,972 | 12,934 | 0.59 | +| Sep 5–11 | 92 | 17 | 18% | 19 | +13,786 / −1,954 | 15,471 | 1.12 | +| Sep 12–16 | 189 | 75 | 39% | 48 | +12,094 / −1,989 | 13,850 | 1.15 | +| Sep 17–22 | 211 | 122 | **57%** | 34 | +11,314 / −1,927 | 11,589 | 1.02 | + +Caveats: + +- The loop names every issue-driven change `fix(...)`, including some enhancements, so the jump at 09-12 partly reflects that labelling convention. +- Even so, 57% fixes on a stable ~11–12k lines of code added per window means **new surface is still being added at a steady rate while the fix share climbs**. That is not a stabilising codebase. +- Deletions stay at 15–18% of additions throughout, so little is being consolidated. +- 484 of 665 commits since 08-29 carry a Claude co-author trailer. + +### 4.4 Issue flow (loop era, #69+) + +| Day | Created | Closed | Open at end of day | +|---|---|---|---| +| 09-12 | 18 | 16 | 2 | +| 09-13 | 39 | 32 | 9 | +| 09-14 | 31 | 34 | 6 | +| 09-16 | 13 | 14 | 4 | +| 09-17 | 24 | 9 | 19 | +| 09-18 | 23 | 35 | 7 | +| 09-19 | 16 | 11 | 12 | +| 09-20 | 9 | 13 | 8 | +| 09-21 | 50 | 37 | 21 | +| 09-22 | 49 | 36 | **34** | + +- Median time from filing to a completed close is **2.75 h** (p90 16 h). The process is fast. +- There is no decline in inflow. Some of the last two days' spike is the external findings ledger (#338–#367) and a heavier regression cadence. +- Closed outcomes for #69+: 216 completed (176 with `status:verified`), 22 not planned (18 `wontfix`, 2 `duplicate`), 34 open. + +**Rough categories by title keyword** (272 loop-era issues, first match wins; approximate): + +| Category | Issues | +|---|---| +| Audio / level / mix | 75 | +| Regression-suite drift or process | 66 | +| Cost/memory estimation | 34 | +| MCP/API surface | 29 | +| Templates/catalog/skills | 25 | +| Validation gap (not caught above) | 17 | +| Other | 24 | + +Separately, 50 titles mention `validate`, and 12 issues carry the `security` label. + +**Process noise versus code signal.** About a quarter of the stream (the 66 drift/process issues) is the loop maintaining its own regression suites: expectations made stale by the loop's own fixes, and approval requests to edit them (for example #172, #173, #260, #263, #276, #304, #307, #320–#323, #327–#331). That is real cost, but it is not a code defect. The `regression` label (55 issues) mostly means "found by the regression agent", not "a verified fix regressed". Most real code regressions came from dependency drift: + +- #169 and #232: transformers 5.17 broke all TTS +- #133: safety-checker black images + +A handful came from fix-on-fix: #197, #150 (after #98), #161 (after #159), #330 (after #312), #334 (after #272). + +### 4.5 Regressions and fix-on-fix + +- **Hand-offs to verify** (timeline `labeled status:fixed-pending-verify`): 184 issues had at least one. 173 passed first time, 10 needed 2 hand-offs, and #197 needed 4. +- **Formally reopened:** 4 (#107, #184, #197, #214). +- **Follow-on issues:** 63 issues cite an earlier issue number in their *title*; 55 of those cite a `status:verified` one. By my hand classification, about 30 are incomplete fixes or missed siblings, and about 20 are suite drift caused by a fix. Examples of incomplete fixes: + - #123 after #118, #129 after #126, #138 after #113, #145 after #96, #150 after #98 + - #161 after #159, #196 and #205 after #180, #227 after #199, #254 after #243 + - #275 after #268, #287/#293 after #108, #290/#291 after #288, #302 after #235 + - #319 after #301, #334 after #272, #341 after #267, #342 after #246, #345 after #285 +- **Fix commits per issue:** 152 issues were referenced by fix commits since 09-12. 23 had 2 or more; #193 had 18 (a feature built through "fix" commits), #85 had 5, and #90/#150/#170/#197 had 4 each. + +**Reading.** The first-pass verify rate (94%) overstates quality, because verification checks the filed repro. The real rework rate is better measured by follow-on issues: about 30 incomplete fixes out of roughly 180 fixed issues is **about 1 in 6**. Those follow-ons concentrate in exactly the structural areas above: validation, audio sample rate and headroom, the estimator, and sibling templates. + +--- + +## 5. Broader perspective + +**Product scope is drifting from "declarative diffusers workflow engine" toward "agent-operated video post-production studio".** + +- Recent and open work: audio mastering (EQ, dynamics, LUFS), shot assembly, series/episode tooling, cost estimation, a CPU colour grade (#349), transcription, a contact-sheet vision tool (#193, 18 commits). +- Each of these is individually reasonable. Together, they put a DAW and an NLE on a single-GPU FIFO worker, with CPU-heavy DSP running in the process that owns the GPU. +- My suggestion is to draw a line on purpose: generation and the minimum assembly needed to deliver a shot stay in dw, and post-production goes to ffmpeg graphs or a separate tool. Make that call before the loop's discovery process makes it for you. The standing tester task explores the post-production workflow, so it will keep finding post-production gaps. + +**Many implicit behaviours make the system hard for agents.** + +- The MCP interface is heavily documented to compensate for implicit semantics: key-name magic, phase-dependent prefixes, warnings in place of errors. +- The tool surface costs about 14k tokens resident, and the loop spends effort keeping prose within budget. +- Making the semantics explicit (typed variables, explicit `$media`/`$type`, structured error codes, one verb per concept rather than the #179 aliases) would let much of that prose go away. Structured errors would also let the MCP client stop flattening 409 bodies into prose (`dw_mcp/client.py:364-451`). + +**Warnings have become a substitute for errors or corrections.** + +- There are about 15 `emit_warning` sites in audio and result writing alone, and many issues read "silent no-op, add a warning" (#288, #290, #291, #292, #294). +- Warnings accumulate, collide with suite assertions (#194, #258, #305, #309), and do not change outcomes. +- Prefer, in order: refuse at validation, then correct automatically, then warn. Use the last only when the first two are impossible. + +**One process, many owners of state.** The worker owns models, step cache and memory stats; the server owns the queue, history and workspaces; and in mounted mode, the MCP client owns a workspace pin shared by every session (#298). + +- A single-user posture is a legitimate decision (Don's note on #298). +- If so, make it explicit in code: refuse a second MCP session, or lock the pin. The current posture is to warn when it happens. + +**The agent loop itself.** + +- *Worked well:* + - It found and closed ~220 real issues in 10 days. + - The security suite found genuine holes (#112–#117, #124, #138). + - The consumer-only tester is a genuinely independent check. +- *What it lacks:* something that works against the structural debt. Each session's scope is one issue, and the budgets reward the narrowest patch. Two options: + - Periodically inject structural work, for example "W1 phase 1: single prepare pipeline, no behaviour change", scoped by you and verified by the full pytest suite plus the regression suites. + - Have triage detect *clusters*: three or more issues in the same subsystem within a week trigger an `owner:don` design note instead of another point fix. +- *Suite-maintenance overhead* (about 66 issues) suggests the suites pin too much incidental output (warning counts, exact basis strings, exact wording). Pinning outcomes and error codes would make them sturdier. + +**Operability.** + +- Good: the deploy script (`scripts/deploy.sh`), a systemd unit, health checks, and `get_server_info` with environment details (#222). +- Missing: + - job durability (W7) + - a hang watchdog, since the phase-stall events are only advisory (#357) + - persisting the step cache across restarts (#244) + - visibility into downloads a job triggers (#343) + - a schema version on `jobs.sqlite` + +**Dependency risk.** The project tracks diffusers git HEAD and new transformers releases closely; #169, #232, #178 and the CI break in #224/#228 all came from upstream. A pinned, known-good lock for lem (with the updater as an explicit upgrade path) would separate "upstream moved" from "we broke it". + +--- + +## 6. Method and caveats + +**Code.** + +- I read the hotspots directly and spot-checked the key claims: + - `jobs.py:818/831` validates without arguments + - `result.fps` is `"type": "integer"` in the schema + - `app.py:1157-1160` maps every error to 400 with `str(e)` + - `requests.get` has no redirect control + - `validate_media_location` calls the scheme-only `validate_url` + - the per-sample Python loop in `audio_utils.py:1494-1513` +- Three read-only review sub-agents covered: server/MCP; engine/validation/planning; tasks, security, tests, UI and docs. Function lengths and nesting came from `ast` scans. +- One sub-agent built the MCP server object from the existing dw-agent venv to count tool schema sizes. It imported code from there but modified nothing; I note it because you asked that those checkouts not be used. + +**Issues.** + +- `gh issue list` (all 280 issues) and GraphQL timelines (labeled, reopened and closed events, comment counts), plus full reads of #193, #197, #298 and #300 and targeted reads of others. +- Categories are keyword heuristics, and the fix-on-fix split is my hand classification of 63 title-linked follow-ons. Treat both as ±20%. + +**History.** + +- Commit mix and size numbers come from `git log --numstat` and `git cat-file` over `develop`. +- Some history was rewritten: many 2026-08-15 commits share a committer timestamp. I bucketed by committer date. +- "Fix %" depends on commit-message conventions, which changed when the loop started. + +**Not done.** + +- I did not run pytest (needs torch and diffusers-git) or the UI tests. +- I made no live server calls, so every run-time claim is from reading the code, not reproducing it. +- In particular, I infer the SSRF redirect bypass and the run-time `from_file` gap from the code. They should be confirmed with a test against a redirecting URL and a variable-supplied `from_file` in the security suite before and after fixing. + +**Issues #338–#367** are mid-flight. I used them only as signal about the current state: they repeat the same classes (validation gaps #345, #347, #364, #365; audio levels #358, #361, #362; estimator #341, #348; get_task metadata #350, #366; language #338, #363). diff --git a/docs/superpowers/plans/2026-09-22-quality-remediation.md b/docs/superpowers/plans/2026-09-22-quality-remediation.md new file mode 100644 index 00000000..d5a97411 --- /dev/null +++ b/docs/superpowers/plans/2026-09-22-quality-remediation.md @@ -0,0 +1,155 @@ +# Quality Remediation Plan + +**Source:** `docs/proposals/audits/2026-09-22-quality-review.md`. It is an independent, read-only review of `develop` @ `09ce397`, covering 280 issues and the git history since 2026-08-29. Weakness IDs (W1–W12) below refer to that document's section 3. + +**Goal:** Stop the issue stream from re-growing out of the same structural bug classes. The loop fixes issues fast (median 2.75 h to a verified close), but about 1 in 6 fixes needs a follow-on issue. Fix-commit share rose from 17% to 57% in three weeks while new surface kept landing at a steady rate. Each phase below removes a *class* of bug rather than one instance. + +**Principles** + +- **No behaviour change first, behaviour change second.** Every redesign starts with a phase that the full pytest suite plus the regression suites prove is behaviour-preserving. Behaviour changes land after that, one per issue, each with a `breaking-change` label if it changes the MCP/HTTP interface. +- **Structural work is scoped by Don, executed by the loop.** Each phase becomes one GitHub issue, parked `owner:don` + `status:needs-approval` until Don approves its scope. Once approved, it hands to `owner:implementer` like any other issue. The tester verifies it against the *invariant* stated in the issue, not only a repro. +- **Freeze point fixes in a class once its redesign is approved.** New issues in that class are linked to the phase issue instead of patched individually, unless they are security or data-loss issues. +- **Order of preference for a wrong outcome:** refuse at validation, then correct automatically, then warn. Use a warning only when the first two are impossible. + +--- + +## Phase 0 — Security fixes (now; small) + +W2. Confirm each finding with a failing case in `regression-suite-security.md` (in the iterate repo) before the fix, and see it pass after. + +- [ ] **Remote media fetch helper.** One function used by `audio_utils.py:361`, `video_utils.py:443`, and the `load_image` paths (`arguments.py:943`, `tasks/gather.py:64`). It should: + - refuse redirects, or re-validate the host on every hop; + - pin the resolved IP for the connection, so DNS rebinding can't swap it; + - cap the response size. +- [ ] **Run-time `from_file`** goes through the same URL and root policy as the static pass (`locations.py:403-450`), not `security.validate_url` (scheme only) with an unrooted `validate_path` (`arguments.py:848-876`). +- [ ] **Error text.** Internal exceptions on `POST /api/jobs` and the other `str(e)` → 400 sites (`app.py:1157-1160, 1275, 1294, 2086, 2467`) become 500 with a generic message. Only known validation errors stay 400 with detail. +- [ ] **One `contain(path, roots)`** replaces `security.validate_path`, `locations._within`, `workspace._is_within`, `dw_mcp/media._confine` and `dw_mcp/assets.py:72-100`. This one is a refactor with no behaviour change; the four fixes above change behaviour. + +**Done when:** the security suite has cases for a redirect to an internal address, a variable-supplied `from_file` outside the roots, and path leakage in 400 bodies. All of them pass. + +## Phase 1 — Gates and durability (small; unblocks everything after it) + +W4 and W7. + +- [ ] **CI on `develop`.** Add `develop` to `push.branches` in `.github/workflows/ci.yml`. Also add the Playwright e2e job, or a smoke subset of it. +- [ ] **Deploy gate.** `scripts/deploy.sh` refuses a commit whose CI status on GitHub isn't `success`. It waits briefly if the status is pending, and `--force` exists for emergencies. +- [ ] **Nightly GPU smoke on lem.** Pytest-driven, not agent-driven: one real generation per model family, run from a systemd timer, with failures filed as issues. +- [ ] **Durable job ledger (#300):** + - insert the job row at submit; + - on startup, mark leftover `queued`/`running` rows as `interrupted`; + - have `shutdown()` mark in-flight jobs; + - replace the try/except `ALTER TABLE` chain (`jobs.py:94-164`) with a `PRAGMA user_version` migration. +- [ ] **Hung-worker watchdog (#357):** + - if there has been no worker event for N minutes while a phase is expected to report, kill and restart the worker and fail the job with a clear reason; + - cancel must work against a hung worker. +- [ ] **Job-lifecycle tests.** `server/jobs.py` has the lowest test/code ratio in the codebase (0.2). + +**Done when:** a push to `develop` runs CI, and lem refuses to deploy a red commit. Restarting `dw-serve` with a queued job and a running job leaves both visible as `interrupted`. + +## Phase 2 — One preparation pipeline for run, validate and plan (redesign; highest leverage) + +W1. This is the root of the "validates clean, dies at run time" class: 50 issue titles mention `validate`. Currently open examples: #345, #347, #363, #364, #365. + +- [ ] **2a (no behaviour change).** Extract `prepare(definition, arguments) -> PreparedWorkflow` from `_prepare_definition` (`workflow.py:798`). It runs substitute → constrain → expand → elide → realize. + - Delete the hand copy `expanded_definition` (`workflow.py:439`). + - `run`, `validate_workflow` and `plan` all consume `PreparedWorkflow`. + - Proof: the full pytest suite, plus all four regression suites, pass unchanged. +- [ ] **2b.** `JobManager.submit` (`jobs.py:818,831`) runs the same argument-aware validation as `validate_workflow`, instead of `loaded.validate()` with no arguments. Remove the resulting triple validation on the run path (route → submit → worker). +- [ ] **2c.** Validators fail closed. A validator crash (`workflow.py:571, 733`, `app.py:1858`) becomes a validation error naming the validator, not an empty list. +- [ ] **2d.** Move checks next to what they guard. + - Add a `static_check` hook to `register_command` (which already owns argument signatures), and to pipelines and writers. + - Migrate the 18 per-incident validator modules onto it one at a time, deleting the duplicated rule tables as you go. Examples: `select_validation._RULES` vs `tasks/select.py`, and the six `_UNRESOLVED_PREFIXES` copies. + - Move the server-only passes (`_argument_reference_errors`, `app.py:1531`) into the engine, so the CLI and REPL get them too. +- [ ] **2e.** Validate the JSON schema *after* variable substitution, so typed fields accept `variable:` references. That removes the #363 class: 55 typed fields reject `variable:` today. +- [ ] **2f.** Add a Hypothesis property test: for generated argument combinations over the catalog templates, `validate` is clean ⇒ `prepare` succeeds. This test is the invariant the tester verifies against. + +## Phase 3 — Audio value model and scope decision + +W3. Audio is the largest issue category (about 75 issues). Open: #358, #361, #362; #349 is adjacent. + +- [ ] **Decision (Don): how much post-production dw owns.** The review recommends that generation, plus the minimum assembly needed to deliver a shot, stays in dw. Mastering-grade DSP should go to ffmpeg filter graphs through PyAV (already a dependency), or out of dw entirely. #349 (CPU grade) and #361 (LUFS) should wait for this decision. +- [ ] **3a (no behaviour change).** Coerce every audio input to one `AudioTrack(waveform: float32 [C,N], rate, source)` at task dispatch. Delete the per-command normalisers: + - `_waveform_and_rate` (`audio_utils.py:1344`); + - `_load_tracks_matching_rate` (`:837`); + - the copies in `pair_audio`, `concat_videos` and `dissolve_videos`. +- [ ] **3b (`breaking-change`).** `sample_rate` always means *resample to*. The relabel behaviour gets its own explicit argument. Fix `pair_audio.py:171-173`, which relabels silently today. +- [ ] **3c.** Enforce a true-peak ceiling at the single audio write boundary: measure after encode and correct, instead of warning after the fact (`result.py:116-160, 190-235`). That closes the headroom chain (#158 … #362). +- [ ] **3d.** Depending on the scope decision, move EQ, dynamics and loudness to ffmpeg filter graphs. This replaces the pure-Python per-sample envelope loop (`audio_utils.py:1494-1513`) that runs on the GPU worker, and removes the transitive-only scipy dependency. + +## Phase 4 — Cost and memory estimator as one fitted model + +W5. The estimator is the most-churned feature: 15+ issues since #91, plus about 8 suite-wording issues that followed its changes. + +- [ ] **4a.** Create one driver-comparison module that both `plan.py` and `server/observed_cost.py` import. It replaces `_declared_drivers`/`_driver_comparable` (`plan.py:354-377`), whose two copies already diverge on list bucketing. +- [ ] **4b.** + - Replace the six `basis` values and their overlays with one model: `minutes = f(drivers)`, fitted per workflow and device from history, with curated catalog values as the prior. + - Report `{estimate, low, high, n}`. + - This is `breaking-change` for the `plan` payload. Update the regression suites to pin the *shape* and the bounds, not basis strings. +- [ ] **4c.** The host-memory model is `fixed + per_entry × n`. #348 (`f617c19`, landed after the review snapshot) moved to exactly this. Fold it into the unified model instead of keeping a separate chain. + +## Phase 5 — Server decomposition and typed responses (mechanical refactor) + +W6. Well suited to the implementer, because 8k lines of server tests cover it. + +- [ ] **5a (no behaviour change).** Split `create_app` (`app.py:703-4226`) into `APIRouter`s: jobs, catalog, validate, workspaces, gallery, assets, models and system. + - Move business logic out of route bodies into service modules: validation orchestration, cost acknowledgement, gallery indexing and orphan scan, zip building, frame tiling. +- [ ] **5b.** Add pydantic response models for every route, and generate the UI's `types.ts` from the OpenAPI schema in place of the 363 hand-written lines. +- [ ] **5c.** Fix the concurrency hazards found in the review: + - the module-global detail caches mutated without a lock (`app.py:212, 226-236, 527`); + - the check-then-act race in workspace delete (`:1973-1991`); + - chunked uploads that bypass the size precheck (`:3517-3530`). +- [ ] **5d.** Structured errors, `{code, path, message}`, from validation and the API. Move test and suite assertions from wording to codes; about 17% of assertions pin message text. `dw_mcp/client.py:364-451` then stops flattening 409 bodies into prose. + +## Phase 6 — Workflow language: explicit over implicit + +W8. Every change here is `breaking-change` and gated on a workflow schema version, so existing workflows keep their meaning. Open: #338, #363 (via Phase 2e), #364, #365. + +- [ ] **6a.** Declared variable types, `{"type": "float", "default": 1, "enum": [...]}`. The inferred-from-default form (`variables.py:277`) stays as a legacy mode. Fixes the #338 (truncation) and #364 (null default) class. +- [ ] **6b.** Explicit media and type references, `{"$media": "asset:x.png"}` and `{"$type": "torch.bfloat16"}`, replace inference from key names: + - the `_image`/`_video` suffixes, and `_type`/`_dtype` with their exceptions (`arguments.py:26, 136-171`); + - `realize_args` stops running over the variables block (`workflow.py:846`; #365). + - Deprecate the inferred forms behind the schema version. +- [ ] **6c.** Document one resolution-phase order for all 11 prefixes, and enforce it in `prepare` (Phase 2). + +## Phase 7 — Hygiene and consolidation (continuous, low risk) + +- [ ] **W9. Model families out of generic code.** Add a `families/` package with a per-family hook for validators and component hints. Move MiniMax-H3 (`adapter_compatibility.py`, `pipeline.py:1985`), Flux (`prompt_weighting.py:293-315`, `teacache.py:291`) and the H3 numbers in `workflow_schema.json` into it. +- [ ] **W10. Templates.** + - Compose shared assembly through `builtin:` sub-workflows, starting with `assemble-and-score` / `dissolve-between-shots` (95% identical; #339/#342 are the latest missed-sibling pair). + - Make one source for each piece of guidance (`get_guide`), with tool docstrings pointing at it. + - Replace `SURFACE_BUDGET`'s character ratchet (`test_mcp_server.py:1450`) with a token count and per-tool caps. +- [ ] **W10. Task metadata.** Derive the hand-kept side tables from signatures, or from one registry (#185, #350, #366). +- [ ] **W11. Comments.** New comments state the invariant; the history goes in the commit message. No sweeping rewrite of existing comments. Edit a comment when its code is next touched. +- [ ] **W12. Tests.** Add a real-codec CPU encode end-to-end test (PyAV). Add property tests for the audio tasks alongside 2f. +- [ ] **Dependencies.** A pinned, known-good lock on lem, with `update_diffusers` as the explicit upgrade path. That separates "upstream moved" (#169, #232) from "we broke it". + +--- + +## Loop process changes (iterate repo, not this repo) + +These go in the role prompts and drivers in `dkackman/iterate`, alongside this plan. + +- [ ] **Cluster escalation in triage.** When three or more open or recently closed issues fall in one subsystem within seven days, triage parks the newest one `owner:don` + `status:needs-approval` with a `triage: cluster` comment naming the others and the phase above they belong to, instead of dispatching another point fix. +- [ ] **Tester verifies the invariant.** The implementer's hand-off states the root-cause invariant. The tester checks it plus one sibling combination, not only the filed repro (the #197 lesson). +- [ ] **Comment convention** (W11) goes into the implementer prompt. +- [ ] **Suites pin outcomes, not incidentals.** New cases assert outcomes and error codes, not warning counts, basis strings or exact wording. This targets the ~66 suite-maintenance issues. +- [ ] **Structural sessions.** Phase items run as ordinary issues with a "no behaviour change" scope line. Their verify step is the full pytest suite plus the regression suites, run by the regression agent, rather than one repro. + +## Decisions needed from Don + +1. **Post-production scope** (Phase 3), before #349 and #361 proceed. +2. **Single-user posture** (#298). Either make it explicit in code (refuse a second MCP session, or lock the workspace pin) or design for several sessions. Today it only warns. +3. **Phase order.** Recommended order: 0 → 1 → 2 → 5a → 3 → 4 → 6, with 7 running continuously. 5a is placed early because the router split makes 2b/2d and 4 easier to land. +4. **Freeze policy.** Whether new issues in a class under redesign are linked and held, or still point-fixed. + +## Open issues mapped to phases (as of 2026-09-22) + +| Phase | Issues | +|---|---| +| 1 | #300, #357, #244 | +| 2 | #345, #347, #363, #364, #365 | +| 3 | #349, #358, #361, #362 | +| 4 | #341, #348 | +| 5 | #298 (posture decision) | +| 6 | #338, #364, #365 | +| 7 | #339, #342, #350, #366 | From 5ca3d40acbe676cb9532abc238617be9d7b28140 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:04:27 -0500 Subject: [PATCH 016/181] fix(introspection): #350 - describe real image-processor parameters, not the generic (image, device) shape get_task/list_tasks reported every image processor as a generic ControlNet preprocessor with accepts_kwargs: true, hiding real keyword arguments (e.g. recenter_crop's center_x, center_y, crop, width, height, fill). image_processor_target maps a processor name to its plain backing function when one exists, and describe_task introspects it directly; a detector-backed processor (controlnet_aux, transformers, dw_pose, sam) still falls back to the old generic shape. Co-Authored-By: Claude Sonnet 5 --- dw/introspection.py | 47 +++++++++++++++++++++++++++---------- dw/tasks/image_utils.py | 38 ++++++++++++++++++++++++++++++ tests/test_introspection.py | 39 ++++++++++++++++++++++++++++++ 3 files changed, 111 insertions(+), 13 deletions(-) diff --git a/dw/introspection.py b/dw/introspection.py index 49ae4782..4c834e37 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -413,20 +413,39 @@ def describe_task(command): } if info["kind"] == "image_processor": + from .tasks.image_utils import image_processor_target + + target = image_processor_target(command) + image_parameter = { + "name": "image", + "required": True, + "default": None, + "annotation": None, + "description": "The image to process", + } + if target is None: + return { + "name": command, + "summary": f"'{command}' image processor (ControlNet preprocessor)", + "accepts_kwargs": True, + "parameters": [image_parameter, device_parameter], + } + + # target is a plain (image, **kwargs) function - introspect it directly + # rather than reporting the generic (image, device) shape every other + # image processor shares (#350). Its first positional parameter is + # the image (named "image" or "img" across these functions), dropped + # in favor of the uniform image_parameter above. + parameters, accepts_kwargs = _callable_parameters(target) + parameters = [image_parameter] + parameters[1:] + if not any(p["name"] == "device" for p in parameters): + parameters.append(device_parameter) + summary = _first_paragraph(inspect.getdoc(target)) return { "name": command, - "summary": f"'{command}' image processor (ControlNet preprocessor)", - "accepts_kwargs": True, - "parameters": [ - { - "name": "image", - "required": True, - "default": None, - "annotation": None, - "description": "The image to process", - }, - device_parameter, - ], + "summary": summary, + "accepts_kwargs": accepts_kwargs, + "parameters": parameters, } if info["implementation"] is None: @@ -460,7 +479,9 @@ def describe_task(command): if domain is not None: parameter["domain"] = domain - summary = _first_paragraph(inspect.getdoc(implementation)) + summary = info.get("summary") + if not summary: + summary = _first_paragraph(inspect.getdoc(implementation)) if not summary: from .tasks.task import _COMMAND_REGISTRY diff --git a/dw/tasks/image_utils.py b/dw/tasks/image_utils.py index e31647b7..9bb2392b 100644 --- a/dw/tasks/image_utils.py +++ b/dw/tasks/image_utils.py @@ -150,6 +150,7 @@ def get_zoe_depth_map(image, device): def image_to_canny(image, low_threshold=100, high_threshold=200): + """Raw cv2 Canny edge map at the image's native resolution.""" # Raw cv2.Canny at the image's native resolution - intentionally kept # separate from the "canny" controlnet_aux CannyDetector above, which # resizes to 512px first. See comment on _ZERO_ARG_DETECTOR_SPECS["canny"]. @@ -202,6 +203,7 @@ def image_to_depth(image, device, height=1024, width=1024): def image_to_segmentation(image): + """Semantic segmentation map from the UperNet ConvNeXt model, colored by class.""" from transformers import AutoImageProcessor, UperNetForSemanticSegmentation image_processor = AutoImageProcessor.from_pretrained( @@ -226,10 +228,12 @@ def image_to_segmentation(image): def get_image_size(image): + """Return the image's width and height in pixels.""" return {"width": image.width, "height": image.height} def crop_square(img: Image) -> Image: + """Crop the image to a centered square of its shorter side.""" # Determine the shortest side min_side = min(img.width, img.height) @@ -248,6 +252,7 @@ def crop_square(img: Image) -> Image: def resize_center_crop(img, height=768, width=768): + """Crop the image to its centered square and resize to width x height.""" output_size = (width, height) W, H = img.size @@ -266,11 +271,14 @@ def resize_center_crop(img, height=768, width=768): def resize_rescale(image, height=768, width=768): + """Resize the image to width x height, ignoring its original aspect ratio.""" input_image = image.convert("RGB") return input_image.resize((width, height)) def resize_resample(image, resolution=1024): + """Resize the image so its shorter side is `resolution`, rounded to a + multiple of 64, preserving aspect ratio.""" input_image = image.convert("RGB") W, H = input_image.size k = float(resolution) / min(H, W) @@ -559,6 +567,36 @@ def available_processors(): return sorted(_PROCESSORS) +# Processors whose handler is a plain (image, device, kwargs) -> function(image, **kwargs) +# forward - i.e. every argument beyond `image` is the named function's own, so +# introspection can read them straight off its signature and docstring instead of +# reporting the generic (image, device) shape every processor otherwise shares +# (#350). Detector-backed processors (controlnet_aux, transformers, dw_pose, sam) +# are left out on purpose: their real argument shape is the detector's __call__, +# not a Python function get_task can point at. +_PROCESSOR_TARGETS = { + "get_image_size": get_image_size, + "add_border_and_mask": add_border_and_mask, + "add_border_and_mask_with_size": add_border_and_mask_with_size, + "canny_cv": image_to_canny, + "segmentation": image_to_segmentation, + "resize_center_crop": resize_center_crop, + "resize_resample": resize_resample, + "crop_square": crop_square, + "recenter_crop": recenter_crop, + "resize_rescale": resize_rescale, + "resize_bucket": resize_bucket, + "strip_exif": strip_exif, + "add_watermark": add_watermark, +} + + +def image_processor_target(processor): + """The plain function backing `processor`'s handler, or None when the + processor is detector-backed and has no such function to introspect.""" + return _PROCESSOR_TARGETS.get(processor) + + def process_image(image, processor, device, kwargs): processor = processor.lower() diff --git a/tests/test_introspection.py b/tests/test_introspection.py index e40568a1..8ca4a4f3 100644 --- a/tests/test_introspection.py +++ b/tests/test_introspection.py @@ -2,6 +2,7 @@ from dw.introspection import ( describe_pipeline, + describe_task, unknown_call_arguments, workflow_argument_warnings, load_pipeline_class, @@ -21,6 +22,44 @@ def test_describe_merges_signature_and_docstring(): assert "self" not in parameters +def test_describe_task_reports_an_image_processors_real_parameters(): + # #350 - recenter_crop has real keyword arguments, not the generic + # (image, device) shape every other image processor used to report. + description = describe_task("recenter_crop") + parameters = {p["name"]: p for p in description["parameters"]} + assert description["accepts_kwargs"] is False + assert "image" in parameters + assert "device" in parameters + for name in ("center_x", "center_y", "crop", "width", "height", "fill"): + assert name in parameters, name + assert description["summary"] + + +def test_describe_task_falls_back_for_a_detector_backed_processor(): + # canny has no plain backing function to introspect (controlnet_aux) - + # it keeps the old generic (image, device) shape. + description = describe_task("canny") + assert description["accepts_kwargs"] is True + names = {p["name"] for p in description["parameters"]} + assert names == {"image", "device"} + + +def test_describe_task_gives_get_first_and_last_frame_their_own_summary(): + # #366 - get_first_frame/get_last_frame share get_frame's implementation + # and used to share its generic docstring summary too. + first = describe_task("get_first_frame") + last = describe_task("get_last_frame") + frame = describe_task("get_frame") + assert first["summary"] != frame["summary"] + assert last["summary"] != frame["summary"] + assert first["summary"] != last["summary"] + assert "first" in first["summary"].lower() + assert "last" in last["summary"].lower() + # frame_index is pinned by these two, not a caller-supplied argument + assert "frame_index" not in {p["name"] for p in first["parameters"]} + assert "frame_index" not in {p["name"] for p in last["parameters"]} + + def test_load_pipeline_class_rejects_non_bare_names(): for bad in ("os.path", "../etc", "", "diffusers.ZImagePipeline", "no_such_thing"): with pytest.raises(ValueError): From 1fd7c0457c46f5e373de1e24d2d03d1cf64cf6f2 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:04:35 -0500 Subject: [PATCH 017/181] fix(video): #367 - get_frame seeks one frame instead of decoding the whole clip get_frame/get_first_frame/get_last_frame only ever need one frame, but reached it through fetch_video/load_video's full decode, which SIGKILL'd on large clips. VideoFileReference wraps a validated file path; realize_args routes a get_frame-family task's 'video' argument to it instead of the eager loader when the source is a file/URL reference, and get_frame seeks to the requested frame with PyAV's frames_at instead of decoding every frame. The existing past-the-end error wording is preserved. Adds tests/test_video_utils.py::TestVideoFileReference, including a monkeypatch proof that load_video is never called for these commands. Co-Authored-By: Claude Sonnet 5 --- dw/arguments.py | 51 ++++++++++++++++ dw/tasks/video_utils.py | 31 ++++++++++ tests/test_video_utils.py | 120 ++++++++++++++++++++++++++++++++++++++ 3 files changed, 202 insertions(+) diff --git a/dw/arguments.py b/dw/arguments.py index b6111251..2cfbb06a 100644 --- a/dw/arguments.py +++ b/dw/arguments.py @@ -135,6 +135,14 @@ def realize_args(arg, base_dir=None): elif k.endswith("_image") or k == "image": logger.debug(f"Loading image for key: {k}") arg[k] = fetch_image(v, base_dir) + # get_frame/get_first_frame/get_last_frame only ever need one frame + # out of their 'video' - loading the ordinary way decodes the whole + # clip to throw all but one frame away, which is what OOM-killed a + # long clip (#367). Recognized by the sibling 'command' on this same + # task object, since that is the only place the command name and + # this argument meet before a task handler runs + elif k == "arguments" and arg.get("command") in _LAZY_FRAME_COMMANDS: + _realize_lazy_frame_arguments(v, base_dir) # Handle video loading for keys ending in '_video' or exactly 'video' elif k.endswith("_video") or k == "video": logger.debug(f"Loading video for key: {k}") @@ -1056,3 +1064,46 @@ def fetch_video(video_spec, base_dir=None): except Exception as e: logger.error(f"Failed to load video {video_spec}: {e}") raise + + +# get_frame and its two fixed-index siblings - see _realize_lazy_frame_arguments +_LAZY_FRAME_COMMANDS = frozenset({"get_frame", "get_first_frame", "get_last_frame"}) + + +def _realize_lazy_frame_arguments(arguments, base_dir): + """Realize a get_frame/get_first_frame/get_last_frame step's arguments, + reading a file-based 'video' by reference rather than decoding it (#367). + + Everything but 'video' is realized the ordinary way. A 'video' naming a + real file or an asset/output path becomes a VideoFileReference the task + reads one frame out of by seeking; a 'previous_result:'/'variable:' + reference is still deferred, and a URL still goes through the ordinary + eager fetch_video, since a seek needs a local, seekable file. + """ + from .tasks.video_utils import VideoFileReference + + if "video" in arguments: + video = arguments["video"] + if is_path_reference(video) or isinstance(video, (list, dict)): + video = resolve_path_references(video, base_dir) + deferred = isinstance(video, str) and ( + video.startswith("previous_result:") or video.startswith("variable:") + ) + url = isinstance(video, str) and ( + video.startswith("http://") or video.startswith("https://") + ) + if isinstance(video, str) and not deferred and not url: + validated_path = validate_media_path(video, base_dir, "a video argument") + ext = os.path.splitext(validated_path)[1].lower() + if ext not in ALLOWED_VIDEO_EXTENSIONS: + raise SecurityError(f"Video file extension not allowed: {ext}") + arguments["video"] = VideoFileReference(validated_path) + elif video is not None: + arguments["video"] = fetch_video(video, base_dir) + + for k, v in list(arguments.items()): + if k == "video": + continue + single = {k: v} + realize_args(single, base_dir) + arguments[k] = single[k] diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index 99aebd9c..361f3ffa 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -38,7 +38,38 @@ def process_video(video, processor, device, kwargs): raise Exception(f"Unknown video processor type: {processor}") +class VideoFileReference: + """A 'video' argument realized to a file on disk rather than an in-memory + clip - built by dw/arguments.py's _realize_lazy_frame_arguments so + get_frame can seek to the one frame it needs instead of decoding the + whole file (#367). Not a public shape; nothing else constructs or + consumes one.""" + + __slots__ = ("path",) + + def __init__(self, path): + self.path = path + + def get_frame(video, frame_index=0): + """Pull one frame out of a video as a PIL image. + + Args: + video: List of PIL images, numpy array or torch tensor of frames, an + AudioVideo, a one-video batch wrapping any of those, or a + VideoFileReference naming a file this call reads by seeking + rather than decoding in full + frame_index: Frame to extract, 0-based; negative indexes count from + the end (-1 is the last frame). Past either end of the clip + raises an error naming the clip's frame count + + Returns: + The frame as a PIL image + """ + if isinstance(video, VideoFileReference): + from ..media_frames import frames_at + + return frames_at(video.path, [f"frame:{frame_index}"])[0]["image"] return extract_frame(video, frame_index) diff --git a/tests/test_video_utils.py b/tests/test_video_utils.py index 114ff3fe..024b3338 100644 --- a/tests/test_video_utils.py +++ b/tests/test_video_utils.py @@ -399,6 +399,126 @@ def test_a_disallowed_extension_is_refused(self, tmp_path): load_audio_video(str(payload)) +class TestVideoFileReference: + """#367. get_frame/get_first_frame/get_last_frame only need one frame; a + VideoFileReference lets get_frame seek to it with PyAV instead of + decoding the whole clip through fetch_video/load_video.""" + + def write_long_clip(self, path, num_frames=300, fps=30, marked=()): + """A clip whose frames are black except the given indexes, which are + pure red - a marker robust to a lossy codec's compression noise, + unlike a unique near-black shade per frame.""" + from diffusers.utils.export_utils import encode_video + + marked = set(marked) + frames = [ + Image.new("RGB", (8, 8), (255, 0, 0) if index in marked else (0, 0, 0)) + for index in range(num_frames) + ] + encode_video(frames, fps=fps, output_path=str(path)) + return str(path) + + def assert_is_red(self, frame): + r, g, b = frame.getpixel((0, 0)) + assert r > 128 and r > g + 64 and r > b + 64 + + def assert_is_black(self, frame): + r, g, b = frame.getpixel((0, 0)) + assert r < 96 + + def test_get_frame_seeks_rather_than_decoding_the_whole_clip(self, tmp_path): + from dw.tasks.video_utils import VideoFileReference + + path = self.write_long_clip(tmp_path / "long.mp4", marked=[250]) + ref = VideoFileReference(path) + + self.assert_is_red(get_frame(ref, 250)) + self.assert_is_black(get_frame(ref, 100)) + + def test_negative_indexes_count_from_the_end(self, tmp_path): + from dw.tasks.video_utils import VideoFileReference + + path = self.write_long_clip(tmp_path / "long.mp4", marked=[299]) + ref = VideoFileReference(path) + + self.assert_is_red(get_frame(ref, -1)) + + def test_an_out_of_range_index_names_the_frame_count(self, tmp_path): + from dw.tasks.video_utils import VideoFileReference + + path = self.write_long_clip(tmp_path / "long.mp4") + ref = VideoFileReference(path) + + with pytest.raises(ValueError, match="past the end of a 300-frame clip"): + get_frame(ref, 999999) + + def test_process_video_dispatches_first_and_last_through_the_reference( + self, tmp_path + ): + from dw.tasks.video_utils import VideoFileReference + + path = self.write_long_clip(tmp_path / "long.mp4", marked=[0, 299]) + ref = VideoFileReference(path) + + first = process_video(ref, "get_first_frame", "cpu", {}) + last = process_video(ref, "get_last_frame", "cpu", {}) + + self.assert_is_red(first) + self.assert_is_red(last) + + def test_realize_args_builds_a_reference_without_calling_load_video( + self, tmp_path, monkeypatch + ): + """The whole point of #367: a get_frame step's 'video' must not go + through the eager, whole-clip fetch_video/load_video path.""" + import dw.arguments as arguments_module + from dw.tasks.video_utils import VideoFileReference + + path = self.write_long_clip(tmp_path / "long.mp4", marked=[250]) + + def _boom(*args, **kwargs): + raise AssertionError("load_video must not be called for get_frame (#367)") + + monkeypatch.setattr(arguments_module, "load_video", _boom) + + task = { + "command": "get_frame", + "arguments": {"video": path, "frame_index": 250}, + } + arguments_module.realize_args(task, base_dir=str(tmp_path)) + + video = task["arguments"]["video"] + assert isinstance(video, VideoFileReference) + self.assert_is_red(get_frame(video, 250)) + + def test_a_deferred_previous_result_reference_is_left_unchanged(self, tmp_path): + import dw.arguments as arguments_module + + task = { + "command": "get_frame", + "arguments": {"video": "previous_result:shot", "frame_index": 0}, + } + arguments_module.realize_args(task, base_dir=str(tmp_path)) + + assert task["arguments"]["video"] == "previous_result:shot" + + def test_a_variable_reference_is_left_unchanged(self, tmp_path): + import dw.arguments as arguments_module + + task = { + "command": "get_frame", + "arguments": {"video": "variable:my_video"}, + } + arguments_module.realize_args(task, base_dir=str(tmp_path)) + + assert task["arguments"]["video"] == "variable:my_video" + + def test_an_in_memory_frame_list_is_unaffected(self, video): + """A step whose 'video' is an earlier step's in-memory result still + goes through the ordinary extract_frame path.""" + assert get_frame(video, 2) is video[2] + + class TestIsVideo: def test_the_shapes_that_are_videos(self): import numpy From dc47a3db06aaff94e744cb0b453d6831020cef35 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:04:40 -0500 Subject: [PATCH 018/181] fix(introspection): #366 - give get_first_frame/get_last_frame their own summary get_frame's docstring gave get_task("get_frame") a real summary, but get_first_frame/get_last_frame share its implementation and so shared that same generic summary, with no mention of which end of the clip each pulls from. _VIDEO_PROCESSOR_INFO entries can now carry a 'summary' override, which describe_task prefers over the shared implementation's docstring. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/task.py | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/dw/tasks/task.py b/dw/tasks/task.py index 5ea4ec73..61a6ea7a 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -591,7 +591,9 @@ def _handle_video_processing(task, arguments, previous_pipelines): # Command names process_video (video_utils.py) accepts, with the function # whose signature carries their arguments. video_utils dispatches via a plain # if-chain, so keep this in sync with the branches in process_video(). -# get_first/last_frame pin frame_index themselves, so it is 'provided'. +# get_first/last_frame pin frame_index themselves, so it is 'provided'; they +# share get_frame's implementation and so would share its generic docstring +# summary too (#366) - 'summary' overrides that per command. _VIDEO_PROCESSOR_INFO = { "get_frame": { "kind": "video_processor", @@ -602,11 +604,13 @@ def _handle_video_processing(task, arguments, previous_pipelines): "kind": "video_processor", "implementation": "dw.tasks.video_utils.get_frame", "provided": ("frame_index",), + "summary": "The first frame of a video, as a PIL image.", }, "get_last_frame": { "kind": "video_processor", "implementation": "dw.tasks.video_utils.get_frame", "provided": ("frame_index",), + "summary": "The last frame of a video, as a PIL image.", }, } _VIDEO_PROCESSOR_COMMANDS = sorted(_VIDEO_PROCESSOR_INFO) From 8a418834dca667d20e8cc6ea2f133df4bbabc3f8 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:26:30 -0500 Subject: [PATCH 019/181] fix(mcp): #353 - configured absolute URLs, auth-aware export handoff, mounted download_output refusal export_job and list_gallery return only server-relative URLs, which are useless handed to a person who isn't a browser already pointed at the box; export_job's instructions also told an MCP-only agent to fetch its own zip even when the endpoint requires auth it cannot attach. And a mounted download_output with no destination silently wrote into the workspace root, where nothing could find or delete it later. - Add a DW_PUBLIC_URL env var / public_url setting. When configured, export_job and list_gallery additionally return absolute_url / absolute_zip_url built from it; unconfigured, those fields are omitted rather than guessed from request/forwarded headers. - export_job's `next` field, and the export paragraphs in the minimax-h3, ltx-2.5 and minimax-music3 plugin skills, now say: when auth_required is true, hand the URL to the person instead of fetching it, and keep using get_output_image/_audio/_frames for inline content. - download_output on a mounted endpoint now refuses an omitted destination with a message pointing at keep_output and the gallery URL, instead of defaulting into the workspace root. An explicit destination is unaffected. Breaking change: the download_output refusal turns a previously-succeeding mounted call (omitted destination) into an error for existing callers. Co-Authored-By: Claude Sonnet 5 --- docs/REMOTE.md | 23 +++- dw/server/app.py | 133 ++++++++++++++-------- dw/settings.py | 9 ++ dw_mcp/exports.py | 65 ++++++++--- dw_mcp/media.py | 19 +++- plugins/dw/skills/ltx-2.5/SKILL.md | 11 +- plugins/dw/skills/minimax-h3/SKILL.md | 12 +- plugins/dw/skills/minimax-music3/SKILL.md | 12 +- tests/test_mcp_exports.py | 40 +++++++ tests/test_mcp_media.py | 15 +++ tests/test_server.py | 17 +++ tests/test_server_exports.py | 62 ++++++++++ 12 files changed, 337 insertions(+), 81 deletions(-) diff --git a/docs/REMOTE.md b/docs/REMOTE.md index d8b3d507..2e26a1b0 100644 --- a/docs/REMOTE.md +++ b/docs/REMOTE.md @@ -33,6 +33,21 @@ Check it from the laptop: `hostname` and `device` are there so you can tell which machine answered. +If an agent on this box will ever export a job or list gallery/asset URLs to +hand to a person who isn't at a terminal on the box itself, set +`DW_API_TOKEN` and also set `DW_PUBLIC_URL` to this server's origin (for +example `https://dw.example.com`, or `http://:8765` with no proxy): + + DW_API_TOKEN= DW_PUBLIC_URL=https://dw.example.com dw-serve --host 0.0.0.0 --mcp + +Without it, `export_job`, `list_gallery` and the upload/asset routes only +return paths relative to this server (`/exports/job-1.zip`) - correct for a +browser already pointed at the box, useless handed to someone who isn't. +With `DW_PUBLIC_URL` set (or the equivalent `public_url` setting), those +responses add an `absolute_url` / `absolute_zip_url` built from it; nothing +guesses this from request headers, so an unconfigured server omits the +field rather than composing a wrong origin. + ## Browser Open `http://:8765`. Click the key icon next to the theme toggle, @@ -55,8 +70,12 @@ both require the token in an `Authorization: Bearer` header - the Two things differ from the local stdio setup: - `download_output` writes on the GPU box (where the MCP server runs), not - on your laptop. Use `get_output_image` / `get_output_text` to see a - result, or open `http://:8765/outputs/` in the browser. + on your laptop, so an omitted destination is refused rather than dropped + loose in the workspace root - nothing on your laptop would find or + delete it there (#353). Pass an explicit destination inside the + workspace to save one anyway, or use `get_output_image` / + `get_output_text` to see a result, or open + `http://:8765/outputs/` in the browser. - The connection is a plain HTTP call per tool invocation; there is no subprocess to restart. - `use_workspace`/`create_workspace` pin *this box's* one MCP client, shared diff --git a/dw/server/app.py b/dw/server/app.py index cb98172f..2cdaefff 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -151,6 +151,7 @@ from .catalog_shape import derive_catalog_metadata, project_listing from . import guides from .guides import GuideError +from .. import settings logger = logging.getLogger("dw") @@ -1328,7 +1329,16 @@ def export_job_route( raise HTTPException(status_code=404, detail=message) raise HTTPException(status_code=409, detail=message) body = summary.as_dict() - body["zip_url"] = _served_url(f"/exports/{quote(job_id)}.zip", ws) + zip_path = f"/exports/{quote(job_id)}.zip" + body["zip_url"] = _served_url(zip_path, ws) + absolute_zip_url = _absolute_served_url(zip_path, ws) + if absolute_zip_url is not None: + body["absolute_zip_url"] = absolute_zip_url + # Same rule get_server_info's field states (#353): whether the zip + # URL above needs a bearer token an MCP-only agent has no way to + # attach itself, which is what tells the caller whether to fetch it + # or hand it to the person. + body["auth_required"] = bool(token) for key, name in ( ("workflow", "workflow.json"), ("manifest", "manifest.json"), @@ -2675,6 +2685,19 @@ def _served_url(path, ws, version=None): separator = "&" if "?" in url else "?" return f"{url}{separator}v={version}" + def _absolute_served_url(path, ws, version=None): + """The same URL, made openable by a client with no other way to + learn this server's origin (#353) - an MCP-only agent, which is + never told a request's Host and must not guess one. `None` unless + an operator has configured `public_url` (or `DW_PUBLIC_URL`): + deriving an origin from request/forwarded headers would trust + whatever the caller claims to be, so a caller gets nothing rather + than a guess.""" + origin = os.environ.get("DW_PUBLIC_URL") or settings.public_url + if not origin: + return None + return f"{origin.rstrip('/')}{_served_url(path, ws, version)}" + def _iter_gallery_files(root, group_runs=True): """Every media file under a directory tree. Yields (relative_name, folder, subfolder, kind, path) - relative_name always uses '/' so @@ -2747,35 +2770,36 @@ def _version(folder, run_id): # and a step that writes more than one kind of file (e.g. a still # plus a video) would lose the extension that tells them apart label = os.path.basename(relative_name) - entries.append( - { - "name": relative_name, - "folder": folder, - "subfolder": subfolder, - # Which run wrote it, and that run's ordinal among this - # workflow's runs - the 'v4' a person sees in the grid - # and an agent says out loud. Two runs write the same - # basename, so `label` cannot tell them apart and - # `name` is too long to quote. None under the flat - # layout, which has no runs to number - "run_id": run_id, - "version": _version(folder, run_id), - # Quoted (slashes kept literal): a name carrying '#', '?' - # or '%' would otherwise break the src the gallery - # renders it into. The mtime still rides along for cache - # busting when a file's content changes without its name - # changing (e.g. a manual overwrite outside the engine) - - # normal reruns get a fresh name instead, see - # dw/result.py's output_file_path - "url": _served_url( - f"/outputs/{quote(relative_name)}", ws, int(stat.st_mtime) - ), - "kind": kind, - "size": stat.st_size, - "mtime": stat.st_mtime, - "label": label, - } - ) + output_path = f"/outputs/{quote(relative_name)}" + entry = { + "name": relative_name, + "folder": folder, + "subfolder": subfolder, + # Which run wrote it, and that run's ordinal among this + # workflow's runs - the 'v4' a person sees in the grid + # and an agent says out loud. Two runs write the same + # basename, so `label` cannot tell them apart and + # `name` is too long to quote. None under the flat + # layout, which has no runs to number + "run_id": run_id, + "version": _version(folder, run_id), + # Quoted (slashes kept literal): a name carrying '#', '?' + # or '%' would otherwise break the src the gallery + # renders it into. The mtime still rides along for cache + # busting when a file's content changes without its name + # changing (e.g. a manual overwrite outside the engine) - + # normal reruns get a fresh name instead, see + # dw/result.py's output_file_path + "url": _served_url(output_path, ws, int(stat.st_mtime)), + "kind": kind, + "size": stat.st_size, + "mtime": stat.st_mtime, + "label": label, + } + absolute_url = _absolute_served_url(output_path, ws, int(stat.st_mtime)) + if absolute_url is not None: + entry["absolute_url"] = absolute_url + entries.append(entry) entries.sort(key=lambda e: e["mtime"], reverse=True) return entries @@ -3572,15 +3596,25 @@ async def upload_media( await run_in_threadpool(_write_bytes, dest, body) logger.info(f"Saved upload {filename!r} -> {dest}") if shared or ws.assets: - return { + path = f"/inputs/{UPLOADS_SUBDIR}/{quote(name)}" + result = { "path": f"asset:{UPLOADS_SUBDIR}/{name}", - "url": _served_url(f"/inputs/{UPLOADS_SUBDIR}/{quote(name)}", ws), + "url": _served_url(path, ws), "shared": shared, } - return { + absolute_url = _absolute_served_url(path, ws) + if absolute_url is not None: + result["absolute_url"] = absolute_url + return result + path = f"/outputs/{UPLOADS_SUBDIR}/{quote(name)}" + result = { "path": dest, - "url": _served_url(f"/outputs/{UPLOADS_SUBDIR}/{quote(name)}", ws), + "url": _served_url(path, ws), } + absolute_url = _absolute_served_url(path, ws) + if absolute_url is not None: + result["absolute_url"] = absolute_url + return result def _asset_origin(ws, root): """Which library an asset came from: this workspace's own, the one @@ -3664,20 +3698,23 @@ def list_assets(ws: Workspace = Depends(selected_workspace)): ) continue seen[relative] = origin - assets.append( - { - "name": relative, - "reference": f"asset:{relative}", - "folder": folder, - "kind": kind, - "size": stat.st_size, - "mtime": stat.st_mtime, - "origin": origin, - # For the editor's own preview - fetchable the same - # way an upload's URL is - "url": _served_url(f"/inputs/{quote(relative)}", ws), - } - ) + asset_path = f"/inputs/{quote(relative)}" + asset_entry = { + "name": relative, + "reference": f"asset:{relative}", + "folder": folder, + "kind": kind, + "size": stat.st_size, + "mtime": stat.st_mtime, + "origin": origin, + # For the editor's own preview - fetchable the same + # way an upload's URL is + "url": _served_url(asset_path, ws), + } + absolute_url = _absolute_served_url(asset_path, ws) + if absolute_url is not None: + asset_entry["absolute_url"] = absolute_url + assets.append(asset_entry) assets.sort(key=lambda entry: entry["mtime"], reverse=True) return { # The workspace's own library, unchanged: where an upload lands diff --git a/dw/settings.py b/dw/settings.py index b6fa36cd..70c0c1d5 100644 --- a/dw/settings.py +++ b/dw/settings.py @@ -28,6 +28,13 @@ class Settings: cudnn_benchmark: bool = True # cuDNN autotuner (faster for fixed sizes) cudnn_deterministic: bool = False # Set True for reproducibility + # This server's public origin (e.g. "https://dw.example.com"), for a + # client that can't otherwise turn a served path into a URL it can open + # itself. None (the default) means no such origin is configured, so + # nothing composes one - see dw/server/app.py's `_served_url`. The + # DW_PUBLIC_URL environment variable overrides this for a single run. + public_url: str = None + def load_settings(): settings = Settings() @@ -53,6 +60,8 @@ def load_settings(): settings.cudnn_benchmark = settings_dict.get("cudnn_benchmark", True) settings.cudnn_deterministic = settings_dict.get("cudnn_deterministic", False) + settings.public_url = settings_dict.get("public_url", None) + return settings diff --git a/dw_mcp/exports.py b/dw_mcp/exports.py index e0897071..94ae71a3 100644 --- a/dw_mcp/exports.py +++ b/dw_mcp/exports.py @@ -3,7 +3,9 @@ The one thing this module has to keep saying: the directory it makes is on the machine running dw.serve, which over a `dw.serve --mcp` endpoint is the GPU box and not where the agent is. The zip URL is the way to it from -anywhere else. +anywhere else - but when the server requires a bearer token (#353), that URL +is for the person to open, not for this agent to fetch on their behalf; see +`export_job`'s `auth_required` / `open_url`. """ from dw_mcp.client import api_path @@ -15,29 +17,66 @@ def export_job(client, job_id, overwrite=False): assets/, inputs/, outputs/. The export copies every output and input file rather than linking them, so a video job's export costs its size again on the server's disk; `total_bytes` in the result reports what - was copied. Returns the directory, the zip URL, the file list with + was copied. Returns the directory, the zip URL(s), the file list with sizes and the total. The three JSON files are in the zip, not repeated - here. The directory is on the server machine, not this one - use the - zip URL to fetch it elsewhere.""" + here. The directory is on the server machine, not this one. + + `auth_required` says whether the zip needs this server's bearer token + to open - a token this agent has no way to attach to a browser or hand + to someone else's tooling. When it is true, `open_url` is for the + *person* to open, not for this agent to fetch: hand it to them (see + `next`). When it is false, `open_url` may be fetched directly. It is + `absolute_zip_url` when the server has one configured (`DW_PUBLIC_URL` + / the `public_url` setting), else the relative `zip_url`.""" body = client.post_json( api_path("api", "jobs", job_id, "export"), params={"overwrite": "true" if overwrite else "false"}, ) directory = body.get("directory") + zip_url = body.get("zip_url") + absolute_zip_url = body.get("absolute_zip_url") + auth_required = bool(body.get("auth_required")) + if auth_required: + open_url = absolute_zip_url or zip_url + next_text = ( + "The directory is on the server, and the zip is behind this " + "server's bearer token - hand open_url to the person and let " + "them open it themselves; do not fetch it. " + + ( + "It is already absolute." + if absolute_zip_url + else "It is relative - tell the person the server's own " + "address, since none is configured (DW_PUBLIC_URL/public_url)." + ) + + " Individual results stay reachable inline via " + "get_output_image/get_output_audio/get_output_frames without " + "opening the zip at all. workflow.json, manifest.json and " + "job.json are inside it - they are not repeated here; " + "get_job_workflow and get_job serve them individually." + ) + else: + open_url = absolute_zip_url or zip_url + next_text = ( + "The directory is on the server. To give the user the files, " + "fetch open_url and unpack it into exports/ under the session's " + "working directory - it is the user's deliverable, not a " + "temporary file, so not a scratch or temp directory. The archive " + "already unpacks into one folder named after the job id; do not " + "create that folder first or the id is doubled in the path. " + "workflow.json, manifest.json and job.json are inside it - they " + "are not repeated here; get_job_workflow and get_job serve them " + "individually." + ) return { "job_id": job_id, "where": f"{directory} on the machine running the MCP server", "directory": directory, - "zip_url": body.get("zip_url"), + "zip_url": zip_url, + "absolute_zip_url": absolute_zip_url, + "auth_required": auth_required, + "open_url": open_url, "files": body.get("files") or [], "total_bytes": body.get("total_bytes"), "missing": body.get("missing") or [], - "next": "The directory is on the server. To give the user the files, " - "fetch zip_url and unpack it into exports/ under the session's " - "working directory - it is the user's deliverable, not a temporary " - "file, so not a scratch or temp directory. The archive already " - "unpacks into one folder named after the job id; do not create " - "that folder first or the id is doubled in the path. workflow.json, " - "manifest.json and job.json are inside it - they are not repeated " - "here; get_job_workflow and get_job serve them individually.", + "next": next_text, } diff --git a/dw_mcp/media.py b/dw_mcp/media.py index a45b360b..ca28c75a 100644 --- a/dw_mcp/media.py +++ b/dw_mcp/media.py @@ -516,10 +516,26 @@ def download_output(client, name, destination=None, overwrite=False, workspace=N Over a `dw.serve --mcp` endpoint the file lands on the *server*, not on the calling agent's machine, so there the destination is confined to that workspace: an absolute or '~' path outside it is refused rather than - written (#113). A stdio `dw-mcp` keeps writing anywhere the user can, + written (#113). An omitted `destination` is refused outright there + rather than defaulting into the workspace root - a file dropped loose in + the root has no run to delete it with and nothing names it back as an + output (#353); pass an explicit destination inside the workspace to save + one anyway. A stdio `dw-mcp` keeps writing anywhere the user can, and an + omitted `destination` keeps defaulting to the current working directory, because there "local disk" is genuinely their own. """ + root = _remote_root(client) if destination is None: + if root: + raise DwApiError( + "destination is required over a dw.serve --mcp endpoint - " + "omitting it would drop the file loose in the workspace " + "root, where nothing can find or delete it later. Pass an " + "explicit destination inside the workspace, or use the url " + "list_gallery reports, get_output_image / get_output_audio / " + "get_output_frames for inline content, or keep_output to " + "make it a named asset instead." + ) destination = os.path.basename(name) destination = os.path.expanduser(destination) if ".." in pathlib.PurePath(destination).parts: @@ -533,7 +549,6 @@ def download_output(client, name, destination=None, overwrite=False, workspace=N # this transport: the caller's own working directory for stdio, the # server's workspace when the tool runs inside dw.serve - where the # process's cwd is an implementation detail the caller never chose - root = _remote_root(client) destination = ( os.path.abspath(os.path.join(root, destination)) if root and not os.path.isabs(destination) diff --git a/plugins/dw/skills/ltx-2.5/SKILL.md b/plugins/dw/skills/ltx-2.5/SKILL.md index 05ec88be..d18714bf 100644 --- a/plugins/dw/skills/ltx-2.5/SKILL.md +++ b/plugins/dw/skills/ltx-2.5/SKILL.md @@ -156,11 +156,12 @@ AESTHETIC QUALITY (in addition to the above, without breaking the objective capt `get_job` for the manifest and its warnings, `get_gallery_metadata` for duration, size and whether audio is present, and hand the user the gallery `url` (`list_gallery`, or the manifest's file name). -6. After an inline run worth keeping, `get_job_workflow` and `save_workflow` it, - so the next run is by name rather than pasted JSON; `export_job` bundles the - run for git. Fetch its zip URL and unpack it into `exports/` under the - session's working directory, never a temp dir - the archive already - unpacks into a job-id folder. +6. After a run worth keeping, `get_job_workflow` and `save_workflow` it, so + the next run is by name not pasted JSON; `export_job` bundles it on the + server. `auth_required: false` - fetch `open_url` into `exports/` under + the working directory (never a temp dir; unpacks into a job-id folder). + `true` - hand `open_url` to the person instead, keep using + `get_output_image`/`_audio`/`_frames` ## Sources diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 145eb795..160e08cc 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -187,12 +187,12 @@ portrait's composition. boards of `dialogue-short`, `storyboard`, `generated-subject-reference` and `music-video`. Also look for a storyboard skipped, every shot the same length, one look word on every board softening all of them. -5. After an inline run worth keeping, `get_job_workflow` and `save_workflow` - it, so the next run is by name rather than pasted JSON; `export_job` bundles - the run — workflow, manifest, job row and media — for git. It is on the - server: fetch its zip URL and unpack it into `exports/` under the session's - working directory, never a temp dir; the archive holds a job-id folder, so - do not make one first. +5. After a run worth keeping, `get_job_workflow` and `save_workflow` it, so + the next run is by name not pasted JSON; `export_job` bundles it on the + server. `auth_required: false` - fetch `open_url` into `exports/` under + the working directory (never a temp dir; it unpacks into a job-id + folder). `true` - hand `open_url` to the person instead and keep using + `get_output_image`/`_audio`/`_frames`. ## Sources diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 30c75e53..2efde10c 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -160,11 +160,13 @@ Control" section. to trim it in the same run, chain `templates/audio-trim-fade` on the output. 6. After an inline run worth keeping, `get_job_workflow` and `save_workflow` it, so the next run is by name rather than by pasting JSON; `export_job` bundles - the run — workflow, manifest, job row and media — for git. The bundle is on - the server: fetch its zip URL and unpack it into `exports/` under the - session's working directory, never a temp directory, and do not make a - folder named after the job id first, since the archive already unpacks - into one. + the run — workflow, manifest, job row and media — for git, on the server. + If `auth_required` is false, fetch `open_url` and unpack it into + `exports/` under the session's working directory, never a temp directory + (the archive already unpacks into a job-id folder, don't make one first). + If true, this agent can't attach the token itself - hand `open_url` to + the person, and keep working via + `get_output_image`/`get_output_audio`/`get_output_frames`. ## Sources diff --git a/tests/test_mcp_exports.py b/tests/test_mcp_exports.py index 0c5db7ce..d509e448 100644 --- a/tests/test_mcp_exports.py +++ b/tests/test_mcp_exports.py @@ -115,3 +115,43 @@ def test_the_next_hint_sends_the_zip_to_the_working_directory(): assert "working directory" in hint assert "temp" in hint assert "do not create that folder" in hint + + +def test_auth_required_tells_the_agent_to_hand_the_zip_to_a_person(): + """#353: when the server gates the zip with a bearer token, this agent + has no way to attach one to a fetch made on the person's behalf - the + hint has to say hand it over, not fetch it.""" + client, _ = exporting(body={**SUMMARY, "auth_required": True}) + + result = exports.export_job(client, "job-1") + + assert result["auth_required"] is True + assert result["open_url"] == "/exports/job-1.zip" + assert "hand open_url to the person" in result["next"] + assert "do not fetch it" in result["next"] + assert "get_output_image" in result["next"] + + +def test_an_absolute_open_url_is_preferred_and_said_to_be_absolute(): + client, _ = exporting( + body={ + **SUMMARY, + "auth_required": True, + "absolute_zip_url": "https://dw.example.com/exports/job-1.zip", + } + ) + + result = exports.export_job(client, "job-1") + + assert result["open_url"] == "https://dw.example.com/exports/job-1.zip" + assert "already absolute" in result["next"] + + +def test_no_auth_required_still_fetches_the_zip_itself(): + client, _ = exporting(body={**SUMMARY, "auth_required": False}) + + result = exports.export_job(client, "job-1") + + assert result["auth_required"] is False + assert result["open_url"] == "/exports/job-1.zip" + assert "fetch open_url" in result["next"] diff --git a/tests/test_mcp_media.py b/tests/test_mcp_media.py index f7b86633..63ce0324 100644 --- a/tests/test_mcp_media.py +++ b/tests/test_mcp_media.py @@ -956,6 +956,21 @@ def test_a_mounted_server_writes_a_relative_destination_into_its_workspace(tmp_p assert (workspace / "kept" / "probe.jpg").read_bytes() == png_bytes(4, 4) +def test_a_mounted_server_refuses_an_omitted_destination(tmp_path): + """#353: defaulting into the workspace root stranded a file nothing could + later find or delete - so a mounted endpoint now requires an explicit + destination rather than picking one.""" + workspace = tmp_path / "workspace" + workspace.mkdir() + client = mounted(png_bytes(4, 4), "image/png", workspace) + + with pytest.raises(DwApiError) as refusal: + download_output(client, "run/probe.jpg") + + assert "destination is required" in str(refusal.value) + assert list(workspace.iterdir()) == [] + + def test_a_stdio_client_still_writes_wherever_the_user_can(tmp_path): """Unmounted, 'local disk' is genuinely the caller's own machine.""" client = serving(png_bytes(4, 4), "image/png") diff --git a/tests/test_server.py b/tests/test_server.py index b808d253..3c1ee989 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -2356,6 +2356,23 @@ def test_upload_media_saves_file_and_returns_absolute_path(server, tmp_path): assert fetched.content == b"not-really-png-bytes" +def test_upload_media_adds_an_absolute_url_when_a_public_url_is_configured( + server, monkeypatch +): + # #353: a client with no way to learn this server's origin otherwise - + # an MCP-only agent - gets an absolute_url only when an operator + # configured one; nothing derives an origin from request headers. + monkeypatch.setenv("DW_PUBLIC_URL", "https://dw.example.com") + with server(success_script) as client: + body = client.post( + "/api/uploads", + params={"filename": "source-image.png"}, + content=b"not-really-png-bytes", + ).json() + + assert body["absolute_url"] == f"https://dw.example.com{body['url']}" + + @pytest.fixture def asset_server(tmp_path): """A server with an asset library configured, which is where uploads go.""" diff --git a/tests/test_server_exports.py b/tests/test_server_exports.py index 35f11140..cfa24bc4 100644 --- a/tests/test_server_exports.py +++ b/tests/test_server_exports.py @@ -4,6 +4,7 @@ import io import json import os +import time import zipfile import pytest @@ -352,6 +353,67 @@ def test_a_second_export_without_overwrite_is_409(self, server): forced = client.post(f"/api/jobs/{job_id}/export?overwrite=true") assert forced.status_code == 201 + def test_auth_required_reflects_whether_a_token_is_configured( + self, workspace_root, tmp_path + ): + # #353: an MCP-only agent has no way to attach a bearer token to a + # fetch on the person's behalf, so export_job's `next` hint branches + # on this field rather than assuming the zip is open to fetch. + manager = JobManager( + workspace_root.outputs, + worker_manager=ScriptedWorkerManager(exporting_script), + history_path=str(tmp_path / "jobs.sqlite"), + workflow_dir=workspace_root.workflows, + ) + app = create_app( + workflow_dir=workspace_root.workflows, + output_dir=workspace_root.outputs, + job_manager=manager, + prompt_dir=workspace_root.prompts, + asset_dir=workspace_root.assets, + workspace=workspace_root.root, + token="s3cr3t", + ) + with TestClient(app, base_url="http://localhost") as client: + headers = {"Authorization": "Bearer s3cr3t"} + submitted = client.post( + "/api/jobs", + json={"workflow": valid_workflow(), "arguments": {}}, + headers=headers, + ).json() + deadline = time.time() + 5.0 + detail = None + while time.time() < deadline: + detail = client.get( + f"/api/jobs/{submitted['id']}", headers=headers + ).json() + if detail["status"] in TERMINAL_STATES: + break + time.sleep(0.02) + assert detail["status"] in TERMINAL_STATES + body = client.post( + f"/api/jobs/{submitted['id']}/export", headers=headers + ).json() + + assert body["auth_required"] is True + + def test_an_absolute_zip_url_is_added_when_a_public_url_is_configured( + self, server, monkeypatch + ): + monkeypatch.setenv("DW_PUBLIC_URL", "https://dw.example.com") + with server() as client: + job_id = finished(client) + body = client.post(f"/api/jobs/{job_id}/export").json() + + assert body["absolute_zip_url"] == f"https://dw.example.com/exports/{job_id}.zip" + + def test_no_absolute_zip_url_when_no_public_url_is_configured(self, server): + with server() as client: + job_id = finished(client) + body = client.post(f"/api/jobs/{job_id}/export").json() + + assert "absolute_zip_url" not in body + class TestExportWithoutAWorkspace: def test_a_server_with_no_workspace_root_answers_409_not_a_crash( From 3dc697adc189766afffc02833fd65c178e0dc1e2 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:28:07 -0500 Subject: [PATCH 020/181] fix(plugin): #355 - skills tell agent to use clear_memory instead of claiming nothing clears VRAM minimax-music3 and ltx-2.5 SKILL.md step 4 said "Nothing over MCP clears it ... ask the operator to restart the worker," but clear_memory does clear resident VRAM (and drops the step cache) while the server is idle. Point the agent at clear_memory + a re-read of get_memory instead, keeping the don't-retry-into-a-failed-attempt reasoning. Co-Authored-By: Claude Sonnet 5 --- plugins/dw/skills/ltx-2.5/SKILL.md | 9 ++++++--- plugins/dw/skills/minimax-music3/SKILL.md | 8 +++++--- 2 files changed, 11 insertions(+), 6 deletions(-) diff --git a/plugins/dw/skills/ltx-2.5/SKILL.md b/plugins/dw/skills/ltx-2.5/SKILL.md index d18714bf..b363cd9c 100644 --- a/plugins/dw/skills/ltx-2.5/SKILL.md +++ b/plugins/dw/skills/ltx-2.5/SKILL.md @@ -24,9 +24,12 @@ Lightricks' own caption spec, quoted below from diffusers. reading's `gpu_memory_allocated_mb`. Only those are the worker's own: `info: null` means nothing is resident, and a `live: false` reading is cached from another moment. A non-trivial idle figure is what an earlier - run left behind and comes off what this one has. Nothing over MCP clears - it - ask the operator to restart the worker rather than retrying into it, - since a failed attempt is itself what leaves weight resident. + run left behind and comes off what this one has. With the server idle, + `clear_memory` clears it (refused while a job is queued or running) - it + also drops the step cache, so the next run, including a seeded rerun, is + cold and regenerates. Re-read `get_memory` afterwards to confirm. Don't + retry into a failed attempt without clearing first - a failed attempt is + itself what leaves weight resident. ## Which shape is the request diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 2efde10c..981da119 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -25,9 +25,11 @@ shapes; do not author a new workflow until the shape decision below fails. is cached from another moment - it reads low while a job is loading a model, so ask again once the server is idle rather than trusting it. A non-trivial idle figure is what an earlier run left behind, and it comes off the ~22 GB - these templates need. Nothing over - MCP clears it, so say so and ask the operator to restart the worker rather - than retrying into it - a failed attempt is itself what leaves weight + these templates need. With the server idle, `clear_memory` clears it + (refused while a job is queued or running) - it also drops the step cache, + so the next run, including a seeded rerun, is cold and regenerates. Re-read + `get_memory` afterwards to confirm. Don't retry into a failed attempt + without clearing first - a failed attempt is itself what leaves weight resident, so an immediate retry starts from less than the attempt that just failed had. From abc9b61589f467ff01bf8aa274e6f1cda84b49e2 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:35:07 -0500 Subject: [PATCH 021/181] fix(server): #356 - accept output: prefix on gallery reads, add media duration to list_gallery Every gallery read that resolves a name through _output_file (metadata, audio, frames, thumbnail, download, delete, keep, archive, and the static /outputs/{name} route) rejected an "output:"-prefixed name with a misleading "path does not exist", even though that is exactly how a workflow argument references the same file. A shared _strip_output_prefix helper strips one leading "output:" before the name is used for path resolution, job lookup, or run-path parsing - containment is still validate_path's, applied to the stripped remainder. GET /api/gallery also gains an opt-in media=true that adds duration_seconds to audio/video entries (via the same probe_media get_gallery_metadata already uses), bounded to the page actually returned so distinguishing two takes of one workflow no longer needs one metadata call per candidate. Default listing is unchanged. Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 55 ++++++++++++++++++++++++++++++++---- tests/test_server.py | 67 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 116 insertions(+), 6 deletions(-) diff --git a/dw/server/app.py b/dw/server/app.py index 2cdaefff..3930c9fa 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -96,6 +96,7 @@ from ..plan import build_plan, gate_warnings, unseeded_cache_warnings from ..runs import ( MANIFEST_FILE_NAME, + OUTPUT_PREFIX, REALIZED_FILE_NAME, is_output_reference, is_run_id, @@ -2511,10 +2512,30 @@ def enhance(request: EnhanceRequest, ws: Workspace = Depends(selected_workspace) # Longest side of an on-demand gallery thumbnail, in pixels GALLERY_THUMBNAIL_MAX_DIM = 320 + def _strip_output_prefix(name): + """A gallery name, accepting the way a workflow argument would + reference it ('output:', #356) as well as the bare form + every gallery listing reports. `asset:` already gets this courtesy + on this same endpoint family (`is_asset_reference` below); a caller + who spelled a name by copying an `output:` reference used to be met + with a wrong-looking "path does not exist" instead, because the + prefix was joined straight into the path rather than stripped first. + + Applied once, at the top of every route that takes a gallery + `name`, so the rest of that route - job lookups, run-path parsing, + the file it echoes back - sees the same bare name `_output_file` + resolves, rather than resolving the file correctly while a sibling + lookup keyed on the untouched string quietly misses. + """ + if is_output_reference(name): + return name.removeprefix(OUTPUT_PREFIX).strip() + return name + def _output_file(name, root=None): """A file inside a workspace's output directory, or a 404 - never outside it.""" root = root or manager.output_dir + name = _strip_output_prefix(name) try: path = validate_path( os.path.join(root, name), @@ -2850,6 +2871,7 @@ def gallery( subfolder: Optional[str] = None, only_orphans: bool = False, version: Optional[int] = None, + media: bool = False, ws: Workspace = Depends(selected_workspace), ): """A page of media files in the output directory, newest first. @@ -2876,7 +2898,15 @@ def gallery( do not apply in this mode, since an orphan run has no file to carry either. `name` is exactly what `DELETE /api/gallery/{name}` accepts, so listing and deleting an orphan is a two-call round - trip (#170).""" + trip (#170). + + `media=true` adds `duration_seconds` to each audio/video entry, + probed the same way `get_gallery_metadata` reports it - which two + takes of the same workflow otherwise have no way to be told apart + by, since size and mtime are misleading proxies for length (#356). + Off by default and bounded by `limit`: only the page actually + returned is probed, not the whole listing, so the cost of asking + stays proportional to the page size rather than the library size.""" if only_orphans: entries = _orphan_entries(ws.outputs) offset = max(0, offset) @@ -2901,6 +2931,13 @@ def gallery( offset = max(0, offset) limit = max(0, limit) page = entries[offset : offset + limit] + if media: + for entry in page: + if entry["kind"] not in ("audio", "video"): + continue + probed = probe_media(os.path.join(ws.outputs, entry["name"])) + if probed is not None: + entry["duration_seconds"] = probed.get("duration_seconds") return { "files": page, "total": len(entries), @@ -2938,6 +2975,7 @@ def gallery_metadata( into the output directory. `job` is null for an asset (nothing here produced it) and `source` says which of the two roots answered.""" run_id, version = "", None + name = _strip_output_prefix(name) if is_asset_reference(name): path = _asset_file(name, ws) source, job = "asset", None @@ -2993,6 +3031,7 @@ def gallery_audio( An audio-only file asked for whole is served as its own bytes in its own encoding - there is nothing to extract, and a transcode would change what the agent hears.""" + name = _strip_output_prefix(name) if is_asset_reference(name): path = _asset_file(name, ws) else: @@ -3099,6 +3138,7 @@ def gallery_frames( every sampled frame before any stamping, fitting or composing, so it names the same region whatever `max_dimension` downscales the result to.""" + name = _strip_output_prefix(name) if is_asset_reference(name): path = _asset_file(name, ws) else: @@ -3372,7 +3412,8 @@ def archive_outputs( # selection fails the request instead of yielding a partial zip # the gallery-relative name is the entry name, so a workflow's output # subfolders stay intact inside the download - paths = [(name, _output_file(name, ws.outputs)) for name in request.names] + names = [_strip_output_prefix(name) for name in request.names] + paths = [(name, _output_file(name, ws.outputs)) for name in names] return _archive_selection(paths, "output") @@ -3461,6 +3502,7 @@ def delete_output(name: str, ws: Workspace = Depends(selected_workspace)): (`/`), which removes the whole run - the only handle on a run that failed before it wrote any media (#134). """ + name = _strip_output_prefix(name) run_dir = _run_directory(name, ws.outputs) if run_dir is not None: # As in _prune_empty_run_directory: pin the siblings' numbers @@ -3774,15 +3816,16 @@ def keep_output_as_asset( ), ) - source = _output_file(request.name, ws.outputs) - asset_name = request.asset_name or os.path.basename(request.name) + kept_name = _strip_output_prefix(request.name) + source = _output_file(kept_name, ws.outputs) + asset_name = request.asset_name or os.path.basename(kept_name) # The kept file's own extension when the name carries none, and a # refusal when it carries a contradicting one - exactly what the # upload route does with its `asset_name`. Without this a kept asset # could be written under an extensionless name, which the library # listing (which reads by kind) never shows again: the call reported # success and the asset was invisible (T014) - extension = os.path.splitext(os.path.basename(request.name))[1].lower() + extension = os.path.splitext(os.path.basename(kept_name))[1].lower() if not os.path.splitext(asset_name)[1]: asset_name = f"{asset_name}{extension}" elif os.path.splitext(asset_name)[1].lower() != extension: @@ -4182,7 +4225,7 @@ async def output_file( ): """One generated file, from the workspace that made it.""" files = _static_files_for(ws.outputs) - return await files.get_response(name, request.scope) + return await files.get_response(_strip_output_prefix(name), request.scope) @app.get("/inputs/{name:path}") async def input_file( diff --git a/tests/test_server.py b/tests/test_server.py index 3c1ee989..202cf560 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -1305,6 +1305,59 @@ def test_gallery_metadata_describes_audio_and_video(server, tmp_path): assert still["media"] is None +def test_gallery_metadata_accepts_an_output_reference(server, tmp_path): + """A name copied from an 'output:' reference used to 404 with "path does + not exist" instead of resolving - the prefix was joined straight into the + path rather than stripped first (#356).""" + from tests.test_media_info import write_wav + + with server(success_script) as client: + outputs = tmp_path / "outputs" + write_wav(outputs / "score-gen.0-0.0.wav", seconds=2.0) + + plain = client.get("/api/gallery/score-gen.0-0.0.wav/metadata") + prefixed = client.get("/api/gallery/output:score-gen.0-0.0.wav/metadata") + + assert prefixed.status_code == plain.status_code == 200 + assert prefixed.json()["media"]["duration_seconds"] == pytest.approx( + 2.0, abs=0.01 + ) + + +def test_gallery_output_reference_traversal_is_still_refused(server, tmp_path): + """Stripping the 'output:' prefix must not open a new escape - the + stripped remainder still goes through validate_path (#356).""" + with server(success_script) as client: + response = client.get("/api/gallery/output:../jobs.sqlite/metadata") + assert response.status_code == 404 + assert (tmp_path / "jobs.sqlite").exists() + + +def test_gallery_lists_media_duration_when_asked(server, tmp_path): + """size and mtime are misleading proxies for a take's length - a + bitrate difference can make a shorter file the bigger one - so + ?media=true adds duration_seconds per entry, bounded by the page + returned rather than the whole library (#356).""" + from tests.test_media_info import write_mp4, write_wav + + with server(success_script) as client: + outputs = tmp_path / "outputs" + write_wav(outputs / "score-gen.0-0.0.wav", seconds=2.0) + write_mp4(outputs / "shot-gen.0-0.0.mp4", frames=12, fps=6) + + default = client.get("/api/gallery").json() + assert all("duration_seconds" not in e for e in default["files"]) + + with_media = client.get("/api/gallery", params={"media": "true"}).json() + by_name = {e["name"]: e for e in with_media["files"]} + assert by_name["score-gen.0-0.0.wav"]["duration_seconds"] == pytest.approx( + 2.0, abs=0.01 + ) + assert by_name["shot-gen.0-0.0.mp4"]["duration_seconds"] == pytest.approx( + 2.0, abs=0.1 + ) + + def test_gallery_audio_extracts_a_videos_soundtrack(server, tmp_path): """get_output_audio refused video/mp4 outright, so a generated clip's soundtrack could only be heard by fetching the file and demuxing it by @@ -2331,6 +2384,20 @@ def test_gallery_delete_and_job_linkage(server, tmp_path): assert (tmp_path / "jobs.sqlite").exists() +def test_gallery_delete_accepts_an_output_reference(server, tmp_path): + """The same 'output:' prefix delete_output rejected before #356.""" + from PIL import Image + + with server(success_script) as client: + outputs = tmp_path / "outputs" + Image.new("RGB", (4, 4)).save(outputs / "victim.png") + + response = client.delete("/api/gallery/output:victim.png") + + assert response.status_code == 200 + assert not (outputs / "victim.png").exists() + + def test_upload_media_saves_file_and_returns_absolute_path(server, tmp_path): with server(success_script) as client: response = client.post( From b7b52272ba74930cb06433256017e7b7b68c77f1 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:42:38 -0500 Subject: [PATCH 022/181] fix(events): #357 - reword phase_stall watchdog as informational, not a fault The watchdog's repeat report bumped its own last-event clock, so a caller reading get_job_events during a long healthy silence (H3's reference encode, block-cache gaps) saw fault-shaped warnings with a climbing event_count and mistook them for liveness or a stuck run. RunContext now tracks the last *real* progress event separately from the watchdog's own emits, and the phase_stall message and a new seconds_since_last_progress field report time since that last real event rather than resetting each time the watchdog itself speaks. get_job_events/wait_for_job docstrings and the minimax-h3 skill are updated to say phase_stall narrates silence rather than signaling a fault. Co-Authored-By: Claude Sonnet 5 --- dw/events.py | 34 ++++++++++++++++++++++----- dw_mcp/server.py | 21 ++++++++++++++--- plugins/dw/skills/minimax-h3/SKILL.md | 7 +++--- tests/test_events.py | 24 +++++++++++++++++++ 4 files changed, 74 insertions(+), 12 deletions(-) diff --git a/dw/events.py b/dw/events.py index dc9038b9..d516f347 100644 --- a/dw/events.py +++ b/dw/events.py @@ -58,6 +58,13 @@ def __init__(self, on_event=None): # measured in elapsed time, not affected by clock adjustments. self._last_event_at = time.monotonic() self._phase_started_at = time.monotonic() + # Last *real* event - anything but the watchdog's own phase_stall + # report - and what kind it was. Tracked separately from + # _last_event_at (which the stall report itself also bumps, to drive + # its own repeat cadence) so "how long has it actually been quiet" + # doesn't reset every time the watchdog speaks. + self._last_progress_at = self._last_event_at + self._last_progress_kind = None # Reference-counted: a sub-workflow runs inside its parent's # RunContext (Workflow.run reuses the ambient one), so the watchdog # starts on the outermost run() and stops on the outermost's exit, @@ -67,7 +74,11 @@ def __init__(self, on_event=None): self._watchdog_stop = threading.Event() def emit(self, event_type, **data): - self._last_event_at = time.monotonic() + now = time.monotonic() + self._last_event_at = now + if not (event_type == "warning" and data.get("kind") == "phase_stall"): + self._last_progress_at = now + self._last_progress_kind = event_type if self._on_event is None: return try: @@ -108,9 +119,15 @@ def _watchdog_loop(self): since any event' and 'what phase are we in', never at what a particular pipeline does inside a phase (see docs/proposals/step-callback-lead-in-instrumentation.md, Option A). - A stall report is itself an event, so it naturally repeats on - PHASE_STALL_THRESHOLD_SECONDS while the silence continues and stops - the moment a real progress event arrives. + A stall report is itself an event, so it bumps _last_event_at and + naturally repeats on PHASE_STALL_THRESHOLD_SECONDS while the silence + continues and stops the moment a real progress event arrives - + but it deliberately does not bump _last_progress_at, so the + seconds_since_last_progress it reports keeps climbing across + repeats rather than resetting itself every 30s (#357). It is a + watchdog notice, not progress: a caller polling events should not + read a run of these, or a climbing event_count made of them, as + liveness. """ while not self._watchdog_stop.wait(PHASE_STALL_CHECK_INTERVAL_SECONDS): phase = self._current_phase @@ -120,9 +137,13 @@ def _watchdog_loop(self): if now - self._last_event_at < PHASE_STALL_THRESHOLD_SECONDS: continue seconds_since_phase_start = round(now - self._phase_started_at, 1) + seconds_since_last_progress = round(now - self._last_progress_at, 1) + last_kind = self._last_progress_kind or "phase start" + last_offset = round(max(0.0, self._last_progress_at - self._phase_started_at), 1) message = ( - f"still in phase '{phase}', {seconds_since_phase_start:.1f}s " - "since it started with no progress event" + f"no progress event for {seconds_since_last_progress:.1f}s in phase " + f"'{phase}' (last: {last_kind} at +{last_offset:.1f}s) - informational; " + "long silent stretches are normal for some models, see the model's skill" ) logger.warning(message) self.emit( @@ -131,6 +152,7 @@ def _watchdog_loop(self): kind="phase_stall", phase=phase, seconds_since_phase_start=seconds_since_phase_start, + seconds_since_last_progress=seconds_since_last_progress, ) def cancel(self): diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 093e7bb0..682b12b1 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -1189,7 +1189,17 @@ def get_job_events(job_id: str, after: int = -1, limit: int = 200) -> dict: seconds since the job started, so where a step's time went is the difference between two events. For 'is it still moving?' the `progress` block on get_job/wait_for_job is cheaper than a page of - events.""" + events. + + A `kind: "phase_stall"` entry is a watchdog notice, not progress - + it fires every ~30s a phase goes quiet and repeats on that cadence + for as long as the silence continues, so it is not evidence of a + hang by itself and a climbing event_count made of nothing else is + not liveness either. It carries `seconds_since_last_progress` + (climbs across repeats) and `seconds_since_phase_start`; some + models are silent for minutes at a time in normal operation (a + video reference encode, block-cache gaps) - check the model's + skill/guide for what's expected before treating one as a fault.""" return diagnose.get_job_events(client, job_id, after=after, limit=limit) def wait_for_job(job_id: str, timeout_seconds: int = 20) -> dict: @@ -1215,8 +1225,13 @@ def wait_for_job(job_id: str, timeout_seconds: int = 20) -> dict: `denoise_step` has moved since a poll minutes ago, not by silence: a video reference's lead-in can run many minutes emitting nothing, and denoise gaps are uneven under a transformer block cache - both - normal. Full diagnosis, and why `denoise_total_steps` can read one - less than asked, in WORKFLOW_GUIDE's "The loop", step 5.""" + normal. If you're also reading get_job_events, a `phase_stall` + entry there is the same silence being narrated, not a fault or a + sign of progress - it repeats every ~30s the phase stays quiet, so + neither seeing one nor watching its event_count climb tells you + anything `denoise_step` doesn't already say better. Full diagnosis, + and why `denoise_total_steps` can read one less than asked, in + WORKFLOW_GUIDE's "The loop", step 5.""" return diagnose.wait_for_job(client, job_id, timeout_seconds=timeout_seconds) # The cap is a number a caller paces against, so the description states diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 160e08cc..3dfa30a2 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -171,7 +171,9 @@ portrait's composition. on to its next step boundary, minutes on this model. Silence is no hang: `denoise_step` is null through the reference encode (~90 s; 629 s for one 5 s 960x544 video reference on a 3090), and the block cache makes later - steps uneven - two-minute gaps are healthy. Each entry carries + steps uneven - two-minute gaps are healthy. `phase_stall` in + `get_job_events` narrates it, not a fault; judge by `denoise_step`. + Each entry carries `subfolder`: `final` is the deliverable (`episode`, `music_video`, `voyage`), `intermediate` the scratch; keep that split in anything you compose. @@ -185,8 +187,7 @@ portrait's composition. audio is present, and hand the user the gallery `url` (`list_gallery`). `get_output_image` works only on image steps - the Z-Image portraits and boards of `dialogue-short`, `storyboard`, `generated-subject-reference` and - `music-video`. Also look for a storyboard skipped, every shot the same - length, one look word on every board softening all of them. + `music-video`. 5. After a run worth keeping, `get_job_workflow` and `save_workflow` it, so the next run is by name not pasted JSON; `export_job` bundles it on the server. `auth_required: false` - fetch `open_url` into `exports/` under diff --git a/tests/test_events.py b/tests/test_events.py index b97f7e3e..aa11a24c 100644 --- a/tests/test_events.py +++ b/tests/test_events.py @@ -463,9 +463,33 @@ def test_watchdog_event_carries_the_required_fields(): assert stall["event"] == "warning" assert stall["phase"] == "saving" assert isinstance(stall["seconds_since_phase_start"], (int, float)) + assert isinstance(stall["seconds_since_last_progress"], (int, float)) assert "message" in stall +def test_watchdog_reports_last_progress_kind_and_does_not_reset_on_repeat(): + events = [] + context = RunContext(on_event=events.append) + with _fast_watchdog(): + context.enter_run() + try: + context.note_phase("generating") + context.emit("pipeline_step", step=1) + time.sleep(0.3) + finally: + context.exit_run() + + stalls = [e for e in events if e.get("kind") == "phase_stall"] + assert len(stalls) >= 2 + assert "pipeline_step" in stalls[0]["message"] + # seconds_since_last_progress climbs across repeats rather than + # resetting each time the watchdog itself emits (#357) + assert ( + stalls[-1]["seconds_since_last_progress"] + > stalls[0]["seconds_since_last_progress"] + ) + + def test_each_run_records_its_own_version(tmp_path): """Consecutive runs of one workflow number themselves 1, 2, 3 - the ordinal the gallery shows as 'v2' and an agent quotes.""" From 9e07b5c93c77944c43d8cd04da5f57e88289599a Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:50:17 -0500 Subject: [PATCH 023/181] fix(mcp): #359 - clearer messages for three misdiagnosed errors Three MCP tool call sites steered agents to the wrong fix: - validate_workflow, run_workflow and save_prompt rejected a JSON-encoded string for a dict-typed param with a raw pydantic dict_type error. They now widen the annotation to dict | str and run the value through a shared coerce_json_object() (moved from authoring.py's local _coerce_json_object, now in dw_mcp/client.py so authoring.py, diagnose.py and prompts.py all use one implementation), matching the coercion save_workflow already had from #247. A string that fails to parse is now reported as invalid JSON, not a type mismatch. - get_output_frames's _moment_to_index (dw/media_frames.py) gave an ambiguous "130.0 is past the end of a 141-frame clip" for an out-of-range bare number. The out-of-range message for a bare (seconds) value now says the clip's duration in seconds, spells out that bare numbers are seconds, and suggests the frame:N form with the caller's own numeral. - concat_videos/dissolve_videos's level_spread warning (warn_on_level_spread, dw/tasks/audio_utils.py) always told a caller to pass match_levels, even when a level difference between shots is intentional (e.g. a shot deliberately silent against a score hole). The message now names that possibility before recommending match_levels. kind="level_spread" and its other fields are unchanged so existing regression-suite matches on kind still hold. Co-Authored-By: Claude Sonnet 5 --- dw/media_frames.py | 36 +++++++++++++++++++++++++++--------- dw/tasks/audio_utils.py | 6 ++++-- dw_mcp/authoring.py | 35 +++++------------------------------ dw_mcp/client.py | 25 +++++++++++++++++++++++++ dw_mcp/diagnose.py | 4 +++- dw_mcp/prompts.py | 3 ++- dw_mcp/server.py | 22 ++++++++++++++-------- 7 files changed, 80 insertions(+), 51 deletions(-) diff --git a/dw/media_frames.py b/dw/media_frames.py index 5c5a72ac..b348ba9e 100644 --- a/dw/media_frames.py +++ b/dw/media_frames.py @@ -216,21 +216,39 @@ def _moment_to_index(moment, shape): total = shape["frame_count"] if isinstance(moment, str) and moment.startswith("frame:"): index = int(moment[len("frame:") :]) - else: - fps = shape["fps"] - if fps is None: - raise ValueError("this clip has no frame rate, so name a frame: 'frame:N'") - seconds = float(moment) - if not math.isfinite(seconds): - raise ValueError(f"{moment!r} is not a moment in seconds") - index = int(round(seconds * fps)) + if index < 0: + index += total + if not 0 <= index < total: + raise ValueError( + f"frame {index} is past the end of a {total}-frame clip " + f"(frames 0-{total - 1})" + ) + return index + fps = shape["fps"] + if fps is None: + raise ValueError("this clip has no frame rate, so name a frame: 'frame:N'") + seconds = float(moment) + if not math.isfinite(seconds): + raise ValueError(f"{moment!r} is not a moment in seconds") + index = int(round(seconds * fps)) if index < 0: index += total if not 0 <= index < total: - raise ValueError(f"{moment!r} is past the end of a {total}-frame clip") + duration = total / fps + raise ValueError( + f"{moment!r} s is past the end of a {duration:.2f} s " + f"({total}-frame) clip - bare numbers in 'at' are seconds; " + f'use "frame:{_format_moment(moment)}" for a frame index' + ) return index +def _format_moment(moment): + if isinstance(moment, float) and moment.is_integer(): + return str(int(moment)) + return str(moment) + + def _seconds(index, shape): return index / shape["fps"] if shape["fps"] else float(index) diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index da868015..52a57ed6 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -1216,8 +1216,10 @@ def warn_on_level_spread(waveforms, command="concat_videos", measure="rms"): # the one who can act on it (#82) emit_warning( f"{command}: the tracks being joined span {spread:.1f} dB " - f"({measure} {min(levels):.1f} to {max(levels):.1f} dBFS) - the cut " - f"will be audible as a level jump. Pass match_levels to even them out", + f"({measure} {min(levels):.1f} to {max(levels):.1f} dBFS) - " + "audible as a level jump unless the difference is intended (a " + "shot written silent against the score). If it is not, pass " + "match_levels to even them out", kind="level_spread", command=command, spread_db=round(spread, 1), diff --git a/dw_mcp/authoring.py b/dw_mcp/authoring.py index 17408361..c7a970f0 100644 --- a/dw_mcp/authoring.py +++ b/dw_mcp/authoring.py @@ -6,9 +6,7 @@ re-implements it - a second, subtly different check is how the two drift. """ -import json - -from dw_mcp.client import DwApiError, api_path +from dw_mcp.client import DwApiError, api_path, coerce_json_object def validate_workflow( @@ -38,6 +36,8 @@ def validate_workflow( `outputs/` - a run sitting in a different workspace, however identical its arguments and seed, does not count as a hit. Pin `workspace` to the one an earlier run actually used if you want to see it credited.""" + workflow = coerce_json_object(workflow, "workflow") + inline_workflow = coerce_json_object(inline_workflow, "inline_workflow") if workflow is not None and inline_workflow is not None: raise DwApiError( "`workflow` and `inline_workflow` are the same thing - provide only one." @@ -82,39 +82,14 @@ def save_workflow(client, name, workflow=None, patch=None): "Provide exactly one of `workflow` (a full replacement) or " "`patch` (a JSON merge patch onto the stored version)." ) - workflow = _coerce_json_object(workflow, "workflow") - patch = _coerce_json_object(patch, "patch") + workflow = coerce_json_object(workflow, "workflow") + patch = coerce_json_object(patch, "patch") if patch is not None: current = client.get_json(api_path("api", "workflows", name)) workflow = _merge_patch(current, patch) return client.put_json(api_path("api", "workflows", name), {"workflow": workflow}) -def _coerce_json_object(value, param_name): - """A tool argument typed as an object can still arrive as a JSON-encoded - string (a caller that serialized it before handing it over, or a client - that couldn't parse a malformed document and passed the raw text - through). Accept that case rather than letting it reach the server as a - string, where the schema rejection names the wrong problem - not "this - isn't an object" but a bare pydantic `type=dict_type, input_type=str`, - which reads as if the field itself were misdeclared.""" - if value is None or isinstance(value, dict): - return value - if not isinstance(value, str): - raise DwApiError( - f"`{param_name}` must be a JSON object, not {type(value).__name__}." - ) - try: - parsed = json.loads(value) - except json.JSONDecodeError as e: - raise DwApiError(f"`{param_name}` is not valid JSON: {e}") from e - if not isinstance(parsed, dict): - raise DwApiError( - f"`{param_name}` must be a JSON object, not {type(parsed).__name__}." - ) - return parsed - - def _merge_patch(target, patch): """RFC 7396 JSON Merge Patch: each dict key in `patch` merges recursively into `target`; any other value replaces `target` outright; diff --git a/dw_mcp/client.py b/dw_mcp/client.py index bbee5b0e..a8de49b4 100644 --- a/dw_mcp/client.py +++ b/dw_mcp/client.py @@ -25,6 +25,31 @@ def is_loopback_url(url): return (urlparse(url).hostname or "").lower() in LOOPBACK_HOSTS +def coerce_json_object(value, param_name): + """A tool argument typed as an object can still arrive as a JSON-encoded + string (a caller that serialized it before handing it over, or a client + that couldn't parse a malformed document and passed the raw text + through). Accept that case rather than letting it reach the server as a + string, where the schema rejection names the wrong problem - not "this + isn't an object" but a bare pydantic `type=dict_type, input_type=str`, + which reads as if the field itself were misdeclared.""" + if value is None or isinstance(value, dict): + return value + if not isinstance(value, str): + raise DwApiError( + f"`{param_name}` must be a JSON object, not {type(value).__name__}." + ) + try: + parsed = json.loads(value) + except json.JSONDecodeError as e: + raise DwApiError(f"`{param_name}` is not valid JSON: {e}") from e + if not isinstance(parsed, dict): + raise DwApiError( + f"`{param_name}` must be a JSON object, not {type(parsed).__name__}." + ) + return parsed + + def path_segment(name): """Percent-encode a name for interpolation into a request path, including its '/' characters. diff --git a/dw_mcp/diagnose.py b/dw_mcp/diagnose.py index ade6ca36..1a40ea43 100644 --- a/dw_mcp/diagnose.py +++ b/dw_mcp/diagnose.py @@ -10,7 +10,7 @@ import os import time -from dw_mcp.client import DwApiError, api_path +from dw_mcp.client import DwApiError, api_path, coerce_json_object TERMINAL_STATUSES = {"succeeded", "failed", "cancelled"} @@ -105,6 +105,8 @@ def run_workflow( raise DwApiError( "`workflow_path` and `name` are the same thing - provide only one." ) + inline_workflow = coerce_json_object(inline_workflow, "inline_workflow") + workflow = coerce_json_object(workflow, "workflow") if inline_workflow is not None and workflow is not None: raise DwApiError( "`inline_workflow` and `workflow` are the same thing - provide only one." diff --git a/dw_mcp/prompts.py b/dw_mcp/prompts.py index f77d3c91..54fcab26 100644 --- a/dw_mcp/prompts.py +++ b/dw_mcp/prompts.py @@ -11,7 +11,7 @@ disagree. """ -from dw_mcp.client import DwApiError, api_path +from dw_mcp.client import DwApiError, api_path, coerce_json_object ENHANCE_COST_REFUSAL = ( "Enhancing a prompt loads a language model and queues a real job on the " @@ -56,6 +56,7 @@ def get_prompt_schema(client): def save_prompt(client, name, prompt): """Write a prompt into the library, overwriting any prompt of that name. The server validates before it writes.""" + prompt = coerce_json_object(prompt, "prompt") return client.put_json(api_path("api", "prompts", name), {"prompt": prompt}) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 682b12b1..4bb9e659 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -907,9 +907,9 @@ def delete_workspace(name: str, acknowledged_cost: bool = False) -> dict: # ----------------------------------------------------------- authoring def validate_workflow( - workflow: dict | None = None, + workflow: dict | str | None = None, name: str | None = None, - inline_workflow: dict | None = None, + inline_workflow: dict | str | None = None, workflow_path: str | None = None, workspace: str | None = None, arguments: dict | None = None, @@ -922,6 +922,8 @@ def validate_workflow( accept both spellings, so a definition or a name checked here can be handed straight to `run_workflow` without renaming a key. Every schema error comes back at once, each with its JSON path. + `workflow`/`inline_workflow` may also be a JSON-encoded string; a + parse failure is reported as invalid JSON, not a type mismatch. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's. @@ -1051,12 +1053,14 @@ def get_prompt_schema() -> dict: before writing one, as you would get_schema before a workflow.""" return prompts.get_prompt_schema(client) - def save_prompt(name: str, prompt: dict) -> dict: + def save_prompt(name: str, prompt: dict | str) -> dict: """Save a prompt to the library, overwriting any prompt of that name. Its `text` may not itself begin with a reference prefix (variable:, previous_result:, constant:, asset:, output:, prompt:) - the server refuses that to prevent a reference resolving twice. - The library is shared by every workspace on this server.""" + The library is shared by every workspace on this server. `prompt` + may also be a JSON-encoded string; a parse failure is reported as + invalid JSON, not a type mismatch.""" return prompts.save_prompt(client, name, prompt) def delete_prompt(name: str) -> dict: @@ -1102,8 +1106,8 @@ def enhance_prompt( def run_workflow( workflow_path: str | None = None, - inline_workflow: dict | None = None, - workflow: dict | None = None, + inline_workflow: dict | str | None = None, + workflow: dict | str | None = None, name: str | None = None, arguments: dict | None = None, acknowledged_cost: bool | dict = False, @@ -1125,8 +1129,10 @@ def run_workflow( catalog name from `list_workflows`, with or without .json, or a path on the server - or `inline_workflow`, a full definition nothing stored covers; `validate_workflow` calls these `name` and - `workflow`, and both tools accept both spellings. `arguments` - overrides the workflow's variables by name. `workspace` pins this + `workflow`, and both tools accept both spellings. + `inline_workflow`/`workflow` may also be a JSON-encoded string; a + parse failure is reported as invalid JSON, not a type mismatch. + `arguments` overrides the workflow's variables by name. `workspace` pins this call to another workspace without switching the session (where its `output:`/`asset:` references live). From 86d37c891cf0b1813745fe7c2b47c33f44ad1274 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:52:20 -0500 Subject: [PATCH 024/181] fix(mcp): #360 - run_workflow's queued answer names its workspace Per #298 the mounted server shares one pin across all clients; a client whose pin was moved by another client's use_workspace/ create_workspace only found out by going looking for its outputs. run_workflow's queued (non-wait) answer now carries job["workspace"], matching what the wait_seconds>0 path already returns via _SLIM_KEYS. Co-Authored-By: Claude Sonnet 5 --- dw_mcp/diagnose.py | 1 + 1 file changed, 1 insertion(+) diff --git a/dw_mcp/diagnose.py b/dw_mcp/diagnose.py index 1a40ea43..00ab0420 100644 --- a/dw_mcp/diagnose.py +++ b/dw_mcp/diagnose.py @@ -138,6 +138,7 @@ def run_workflow( "job_id": job.get("id"), "status": job.get("status"), "queue_position": job.get("queue_position"), + "workspace": job.get("workspace"), "next": "Poll get_job_events(job_id) for progress, then get_job(job_id) " "for the manifest or the error.", } From dbe6d26b1cede5274626b2fce35b8a346512cbb7 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 19:54:03 -0500 Subject: [PATCH 025/181] fix(templates): #362 - minimax/music balanced step normalizes to -3 dBFS The default mp3 deliverable normalized to -1 dBFS still clipped after encode (+0.56 dBFS measured), contradicting the engine's own audio_clipped warning, which prescribes -3. Every other audio-deliverable template already uses -3; bring music.json in line and update its description and the pinning test in tests/test_result.py. Co-Authored-By: Claude Sonnet 5 --- tests/test_result.py | 8 ++++++-- workflows/templates/minimax/music.json | 4 ++-- 2 files changed, 8 insertions(+), 4 deletions(-) diff --git a/tests/test_result.py b/tests/test_result.py index 665c5fb6..973b30e6 100644 --- a/tests/test_result.py +++ b/tests/test_result.py @@ -1943,7 +1943,11 @@ class TestTheMusicTemplatesLeaveHeadroom: #161: -1 dBFS was enough for an mp3 (measured at -0.07 and -0.41) and not for the mux - the AAC encode overshoots by around 1.9 dB on this material, so `music-video`'s finished mp4 still decoded at +0.94 dBFS. - A deliverable that ends in a video mux normalizes to -3.""" + A deliverable that ends in a video mux normalizes to -3. + + #362: -1 dBFS was not reliably enough for the mp3 either - a run + measured +0.56 dBFS after the encode, so `music.json`'s own deliverable + takes -3 too.""" def steps_of(self, path): with open(path) as definition_file: @@ -1953,7 +1957,7 @@ def steps_of(self, path): @pytest.mark.parametrize( "path,source,target", [ - ("workflows/templates/minimax/music.json", "generate_music", -1.0), + ("workflows/templates/minimax/music.json", "generate_music", -3.0), ("workflows/templates/minimax/music-video.json", "write_song", -3.0), ], ) diff --git a/workflows/templates/minimax/music.json b/workflows/templates/minimax/music.json index fbab14cc..ac67944e 100644 --- a/workflows/templates/minimax/music.json +++ b/workflows/templates/minimax/music.json @@ -1,6 +1,6 @@ { "id": "MiniMaxMusic", - "description": "Audio-only generation: a music track from a prompt, showing that a pipeline's output need not be visual. The step declares an 'audios' output and an audio result type, and the engine encodes it at the given sample rate. Also the minimal way to run a modular pipeline: rather than placing components by hand, a 'components_manager' with auto CPU offload takes ownership of device placement and keeps only the running components resident, so no quantization or per-component offload configuration is needed. 'audio_duration' is a ceiling, not a target: the model stops where the music stops, so ask for more time than the song needs - a value trimmed to the intended length cuts the outro off mid-decay - and finish the tail with 'templates/audio-trim-fade.json'. The result declares 'sample_rate' because a modular pipeline's dict output carries none of its own. The track is peak-normalized to -1 dBFS before it is written: Music 3 lands wherever it lands, which on this box is at or over full scale every time, and a deliverable with no headroom clips in the encoder (#158, #159). Normalization is a gain change and nothing else, so the dynamics are the model's.", + "description": "Audio-only generation: a music track from a prompt, showing that a pipeline's output need not be visual. The step declares an 'audios' output and an audio result type, and the engine encodes it at the given sample rate. Also the minimal way to run a modular pipeline: rather than placing components by hand, a 'components_manager' with auto CPU offload takes ownership of device placement and keeps only the running components resident, so no quantization or per-component offload configuration is needed. 'audio_duration' is a ceiling, not a target: the model stops where the music stops, so ask for more time than the song needs - a value trimmed to the intended length cuts the outro off mid-decay - and finish the tail with 'templates/audio-trim-fade.json'. The result declares 'sample_rate' because a modular pipeline's dict output carries none of its own. The track is peak-normalized to -3 dBFS before it is written: Music 3 lands wherever it lands, which on this box is at or over full scale every time, and a deliverable with no headroom clips in the encoder (#158, #159, #362) - the mp3 encode can add back over a dB of overshoot on top of the peak-normalized level, so -1 dBFS left no margin. Normalization is a gain change and nothing else, so the dynamics are the model's.", "cost": [ { "device": "cuda", @@ -52,7 +52,7 @@ "command": "normalize_audio", "arguments": { "audio": "previous_result:generate_music", - "peak_dbfs": -1.0, + "peak_dbfs": -3.0, "sample_rate": 44100 } }, From 277265ebd0603f137894e0f0075f558e554ea14e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 20:06:47 -0500 Subject: [PATCH 026/181] fix(introspection): #364 - null-defaulted variable feeding a required task argument is a warning, not a refusal A required task argument written as `variable:name`, where `name`'s declared default is null, was reported as "the step does not supply" it and refused by save_workflow and by validate_workflow with no arguments - even though the step does supply it, just not yet a value. task_signature_errors now compares the expanded definition against the author-written one: when a missing required key was written as `variable:name` and name resolves to null, the error carries a `variable` key and a message naming the variable instead of the generic #141 wording. Workflow.validation_errors downgrades that specific case to a warning when called with no arguments (save, or validate with no arguments), via the new null_variable_argument_warnings(), while keeping it a hard error whenever the run/validate call's own arguments leave the variable null. The schema's null description now explains the caveat. Co-Authored-By: Claude Sonnet 5 --- dw/introspection.py | 93 ++++++++++++++++++++++++++--- dw/server/app.py | 6 ++ dw/workflow.py | 53 +++++++++++++++- dw/workflow_schema.json | 2 +- tests/test_task_signature_errors.py | 71 ++++++++++++++++++++++ 5 files changed, 213 insertions(+), 12 deletions(-) diff --git a/dw/introspection.py b/dw/introspection.py index 4c834e37..0f6b6563 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -559,7 +559,49 @@ def missing_task_argument_message(command, missing): ) -def task_signature_errors(workflow_definition, source_indices=None): +def null_variable_task_argument_message(command, missing_arg, variable_name): + """The wording for a required argument the step *does* supply, by + `variable:`, but the variable's value is null (#364). + + `missing_task_argument_message` says "the step does not supply" it, + which is false here - the step names the variable, the variable just + hasn't been given a real value yet. That is a caller's job to do at + run time, not a defect in the document. + """ + return ( + f"'{missing_arg}' is fed by variable '{variable_name}', which is " + f"null - task '{command}' requires a real value for it. Pass " + f"arguments={{'{variable_name}': ...}} when running or validating, " + f"or give '{variable_name}' a non-null default" + ) + + +def _null_fed_variable(written_steps, source_index, key, declared_variables): + """The variable name, if the argument at `key` was written as + `variable:` naming a declared variable - the shape that makes a + "missing" required argument actually a null-variable one (#364). None + otherwise, including when `written_steps` can't be indexed (a for_each + template step, whose members are checked by `item:`/`gather:` instead). + """ + if not isinstance(source_index, int) or source_index >= len(written_steps): + return None + step = written_steps[source_index] + if not isinstance(step, dict): + return None + task = step.get("task") + if not isinstance(task, dict): + return None + arguments = task.get("arguments") + if not isinstance(arguments, dict): + return None + value = arguments.get(key) + if not isinstance(value, str) or not value.startswith("variable:"): + return None + name = value[len("variable:") :] + return name if name in declared_variables else None + + +def task_signature_errors(workflow_definition, source_indices=None, written_definition=None): """Every task step whose arguments its command's signature refuses, as [{path, message}] - a required argument left unset, and an argument the command does not take - plus a step naming a command that is not @@ -587,6 +629,16 @@ def task_signature_errors(workflow_definition, source_indices=None): The definition handed here has already been substituted and expanded, so a for_each member is checked as it will run; `source_indices` maps each expanded step back to the step the author wrote. + + `written_definition`, when given, is that step *as the author wrote it* - + before substitution - plus the declared `variables` block. A required + argument reported missing whose written form is `variable:` naming + a declared variable is not a step that "does not supply" it (#364): the + step does name it, the variable's value just resolved to null (the only + way substitution drops a `variable:` reference, per #209). That error + carries a `variable` key naming it, so a caller checking a document with + no arguments of its own can treat it as caller input rather than a + defect in the document. """ from .for_each import MEMBER_SEPARATOR, render_path from .tasks.task import task_command_info @@ -595,6 +647,17 @@ def task_signature_errors(workflow_definition, source_indices=None): if not isinstance(steps, list): return [] + written_steps = ( + (written_definition or {}).get("steps") or [] + if isinstance(written_definition, dict) + else [] + ) + declared_variables = ( + (written_definition or {}).get("variables") or {} + if isinstance(written_definition, dict) + else {} + ) + errors = [] for index, step in enumerate(steps): if not isinstance(step, dict): @@ -642,16 +705,28 @@ def task_signature_errors(workflow_definition, source_indices=None): if not missing and not unknown: continue - def report(key, message): - errors.append( - { - "path": render_path(("steps", source, "task", "arguments", key)), - "message": f"{message}{where}.", - } - ) + def report(key, message, variable=None): + entry = { + "path": render_path(("steps", source, "task", "arguments", key)), + "message": f"{message}{where}.", + } + if variable is not None: + entry["variable"] = variable + errors.append(entry) if missing: - report(missing[0], missing_task_argument_message(command, missing)) + key = missing[0] + variable = _null_fed_variable( + written_steps, source, key, declared_variables + ) + if variable is not None: + report( + key, + null_variable_task_argument_message(command, key, variable), + variable=variable, + ) + else: + report(key, missing_task_argument_message(command, missing)) for key in unknown: report(key, unknown_task_argument_message(command, key)) return errors diff --git a/dw/server/app.py b/dw/server/app.py index 3930c9fa..fa37ac99 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -1793,6 +1793,11 @@ def validate_workflow( # future checkpoint cannot be predicted, but nothing at run time # would say it loaded onto the wrong one (#155) + candidate.adapter_warnings(request.arguments) + # A required task argument fed by variable:name where name's + # default is null - a fine document, but a run left as-is would + # fail; empty once request.arguments names anything, since that + # condition is a hard error above instead (#364) + + candidate.null_variable_argument_warnings(request.arguments) # An argument a sub-workflow step passes to a workflow that # declares no variable for it - dropped in silence at run time + candidate.sub_workflow_warnings(), @@ -2107,6 +2112,7 @@ def save_workflow( # to shape-first discovery metadata = derive_catalog_metadata(request.workflow) warnings = list(workflow_argument_warnings(request.workflow)) + warnings += candidate.null_variable_argument_warnings() if not metadata["summary"]: warnings.append( "No summary: add a 'description' (its first sentence becomes " diff --git a/dw/workflow.py b/dw/workflow.py index 0640c55e..ebbc7277 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -638,6 +638,19 @@ def validation_errors(self, arguments=None, composing=None): base_dir = ( os.path.dirname(os.path.abspath(self.file_spec)) if self.file_spec else None ) + task_errors = task_signature_errors(expanded, source_indices, self.workflow_definition) + if arguments is None: + # A step that feeds a required argument from `variable:name` and + # a variable whose default is null is a fine document - the + # variable just hasn't been given a value yet, which is exactly + # what no-arguments means here (save_workflow, or + # validate_workflow called to check the document rather than a + # specific run). Downgraded to a warning + # (null_variable_argument_warnings) rather than dropped outright, + # since it is still true a run left as-is would fail (#364). + # Anything else task_signature_errors reports - a genuinely + # missing or unknown argument - stays a hard error regardless + task_errors = [e for e in task_errors if "variable" not in e] return ( previous_result_reference_errors(expanded, source_indices) + subfolder_errors(expanded, source_indices) @@ -694,8 +707,13 @@ def validation_errors(self, arguments=None, composing=None): # A required task argument left unset validated as `valid: true` # and then failed the job on Python's own signature error, which # is the one mistake a free pre-flight most obviously exists for - # (dw/introspection.py, #141) - + task_signature_errors(expanded, source_indices) + # (dw/introspection.py, #141). When the step supplies it by + # `variable:name` and only the variable's value is null, the + # error carries a `variable` key (#364) so a caller checking the + # document itself - no arguments of its own - can tell "the + # variable needs a value at run time" apart from "the step is + # broken", and downgrade the former below + + task_errors # A step's pipeline names a component_type/scheduler_type/ # config_type that does not exist (or is outside the trusted # ecosystem entirely) - validated clean and died 3s into the run @@ -749,6 +767,37 @@ def adapter_warnings(self, arguments=None): supplied=set(arguments or {}), ) + def null_variable_argument_warnings(self, arguments=None): + """Every required task argument fed by `variable:name` where name's + value is null - downgraded out of `validation_errors` when + `arguments` is None (#364), surfaced here so a caller checking the + document without arguments of its own (save_workflow, + validate_workflow with no `arguments`) still sees it, just not as a + reason the document is invalid. + + Empty once `arguments` is given: at that point the same condition is + a hard error in `validation_errors`, since a real run or a validate + call naming its own arguments needed the variable to hold something. + + Best effort: a definition the schema or the expander refuses has its + own errors to report and none of them are this one. + """ + if arguments is not None: + return [] + try: + source_indices = [] + expanded = self.expanded_definition(arguments, source_indices) + except Exception: + logger.debug("No null-variable-argument warnings available", exc_info=True) + return [] + return [ + f"{entry['path']}: {entry['message']}" + for entry in task_signature_errors( + expanded, source_indices, self.workflow_definition + ) + if "variable" in entry + ] + def _undeclared_variable_errors(self, arguments=None): """Every 'variable:' reference naming nothing the workflow declares. diff --git a/dw/workflow_schema.json b/dw/workflow_schema.json index f7b893c1..a6a83397 100644 --- a/dw/workflow_schema.json +++ b/dw/workflow_schema.json @@ -264,7 +264,7 @@ "arguments": { "type": "object", "additionalProperties": { - "description": "null declares an argument that is optional - a workflow can expose a variable a caller may pass without inventing a sentinel value for its absence.", + "description": "null declares an argument that is optional - a workflow can expose a variable a caller may pass without inventing a sentinel value for its absence. If a step feeds a required task argument from a variable whose value is null, the document itself still saves and validates - but a run (or a validate/save call that supplies its own arguments) that leaves the variable null fails, since the argument the step needs was never actually given a value.", "type": [ "string", "integer", diff --git a/tests/test_task_signature_errors.py b/tests/test_task_signature_errors.py index c92d9c6f..4636b37a 100644 --- a/tests/test_task_signature_errors.py +++ b/tests/test_task_signature_errors.py @@ -23,6 +23,7 @@ class of mistake a free pre-flight most obviously exists for (#141). An workflow_argument_warnings, ) from dw.tasks.task import Task +from dw.workflow import Workflow def task_step(command, arguments, name="a"): @@ -111,6 +112,76 @@ def test_a_free_form_command_accepts_anything(self): assert errors_for("gather_inputs", {"whatever": 1}) == [] +class TestARequiredArgumentFedByANullVariable: + """A step that names the argument by `variable:name`, where `name`'s + value is null, is not a step that "does not supply" it (#364) - the + error carries a `variable` key so a caller with no arguments of its own + can tell the two apart.""" + + def test_the_repro_carries_the_variable_key_and_a_clearer_message(self): + written = { + "id": "sig", + "variables": {"audio": None}, + "steps": [ + task_step( + "resample_audio", + {"audio": "variable:audio", "target_sample_rate": 16000}, + ) + ], + } + # replace_variables drops a variable: reference resolved to null + # from its containing dict (#209) - this is what the expanded + # definition looks like once that has happened + expanded = { + "id": "sig", + "steps": [task_step("resample_audio", {"target_sample_rate": 16000})], + } + errors = task_signature_errors(expanded, written_definition=written) + assert len(errors) == 1 + assert errors[0]["variable"] == "audio" + assert errors[0]["path"] == "steps[0].task.arguments.audio" + assert "variable 'audio'" in errors[0]["message"] + assert "does not supply" not in errors[0]["message"] + + def test_no_written_definition_keeps_the_original_wording(self): + """Without the author-written form to compare against - the #141 + call sites already in the codebase before #364 - nothing changes.""" + errors = errors_for("resample_audio", {"target_sample_rate": 16000}) + assert "variable" not in errors[0] + assert "does not supply" in errors[0]["message"] + + def test_a_genuinely_missing_argument_is_unaffected(self): + """No `variable:` reference at all in the written step - still the + plain #141 message, even with a written_definition available.""" + written = { + "id": "sig", + "steps": [task_step("resample_audio", {"target_sample_rate": 16000})], + } + errors = task_signature_errors(written, written_definition=written) + assert "variable" not in errors[0] + assert "does not supply" in errors[0]["message"] + + def test_a_variable_not_declared_is_unaffected(self): + """`variable:audio` written but nothing declares `audio` - not the + shape #364 covers, so the original wording stands.""" + written = { + "id": "sig", + "steps": [ + task_step( + "resample_audio", + {"audio": "variable:audio", "target_sample_rate": 16000}, + ) + ], + } + expanded = { + "id": "sig", + "steps": [task_step("resample_audio", {"target_sample_rate": 16000})], + } + errors = task_signature_errors(expanded, written_definition=written) + assert "variable" not in errors[0] + assert "does not supply" in errors[0]["message"] + + class TestTheRunTimeBackstop: """The static pass sees literals. A required argument that arrived from a variable or an earlier step and resolved to nothing reaches the command, From e5bfb9ecceab311355658b6c85a3efb08265d8b3 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 20:23:48 -0500 Subject: [PATCH 027/181] fix(engine): #365 - media name-conventions no longer pre-load the variables dict by variable name realize_args applied its image/video/type key-name conventions when realizing the variables dict, keying off the variable's own name rather than the argument it would end up filling - a variable named "image" fed into a step's "video" argument was pre-loaded as a PIL Image before substitution ever ran, and fetch_video then raised a raw "got " naming neither the argument nor the step. Add apply_key_conventions to realize_args (default True, unchanged for every step-level call) and pass apply_key_conventions=False only at the variables-dict call site in workflow.py's _prepare_definition - explicit references (asset:/output:/constant:/prompt:/media_type objects) still resolve there, but a plain value now loads only once substituted into a real argument key. Any type mismatch that still reaches a loader now names the argument, the step, and what the value already was (_fetch_image_with_context/_fetch_video_with_context, _describe_value_source). Catalog sweep of workflows/** found no variable whose name and consuming argument key disagree on media type, so no template needed remediation. Co-Authored-By: Claude Sonnet 5 --- dw/arguments.py | 78 +++++++++++++++++++++++++++---- dw/workflow.py | 11 ++++- tests/test_arguments.py | 74 +++++++++++++++++++++++++++++ tests/test_workflow_step_cache.py | 6 ++- 4 files changed, 157 insertions(+), 12 deletions(-) diff --git a/dw/arguments.py b/dw/arguments.py index 2cfbb06a..874d16e4 100644 --- a/dw/arguments.py +++ b/dw/arguments.py @@ -96,7 +96,7 @@ def __repr__(self): # Helper functions for processing and loading workflow arguments -def realize_args(arg, base_dir=None): +def realize_args(arg, base_dir=None, apply_key_conventions=True): """ Recursively processes workflow arguments to: 1. Convert type references into actual Python types @@ -108,6 +108,17 @@ def realize_args(arg, base_dir=None): arg: The arguments to process, modified in place base_dir: Directory relative file paths are resolved against - the workflow file's directory. Defaults to the process working directory + apply_key_conventions: Whether a bare value loads by what its key looks + like (an 'image'/'video'/'_type' name). Explicit references + (asset:, output:, constant:, prompt:, a {media_type, location} + dict) always resolve regardless of this flag - only the fallback + that guesses from the key name is gated. The top-level variables + dict is realized with this off (dw/workflow.py): a variable's own + name is not the argument it will end up filling, so 'image' guessed + a variable named that way into a PIL Image before the step that + actually names its argument 'video' ever saw the value (#365). + Nested structures still recurse with this at its default, since by + then a dict key is a real argument name again """ if isinstance(arg, dict): logger.debug(f"Processing dictionary arguments: {list(arg.keys())}") @@ -132,23 +143,29 @@ def realize_args(arg, base_dir=None): elif is_media_reference(v): arg[k] = fetch_media(v, base_dir) # Handle image loading for keys ending in '_image' or exactly 'image' - elif k.endswith("_image") or k == "image": + elif apply_key_conventions and (k.endswith("_image") or k == "image"): logger.debug(f"Loading image for key: {k}") - arg[k] = fetch_image(v, base_dir) + arg[k] = _fetch_image_with_context(v, base_dir, k) # get_frame/get_first_frame/get_last_frame only ever need one frame # out of their 'video' - loading the ordinary way decodes the whole # clip to throw all but one frame away, which is what OOM-killed a # long clip (#367). Recognized by the sibling 'command' on this same # task object, since that is the only place the command name and # this argument meet before a task handler runs - elif k == "arguments" and arg.get("command") in _LAZY_FRAME_COMMANDS: + elif ( + apply_key_conventions + and k == "arguments" + and arg.get("command") in _LAZY_FRAME_COMMANDS + ): _realize_lazy_frame_arguments(v, base_dir) # Handle video loading for keys ending in '_video' or exactly 'video' - elif k.endswith("_video") or k == "video": + elif apply_key_conventions and (k.endswith("_video") or k == "video"): logger.debug(f"Loading video for key: {k}") - arg[k] = fetch_video(v, base_dir) + arg[k] = _fetch_video_with_context(v, base_dir, k) # Handle type references, and the keys that only look like one - elif k.endswith("_type") or k.endswith("_dtype") or k == "dtype": + elif apply_key_conventions and ( + k.endswith("_type") or k.endswith("_dtype") or k == "dtype" + ): if isinstance(v, EscapedString): # An earlier pass already consumed this value's escape continue @@ -199,7 +216,14 @@ def realize_args(arg, base_dir=None): if is_media_reference(item): arg[i] = fetch_media(item, base_dir) continue - realize_args(item, base_dir) + try: + realize_args(item, base_dir) + except ValueError as error: + if isinstance(item, dict) and "name" in item: + raise ValueError( + f"{error} (step '{item['name']}')" + ) from error + raise arg[i] = realize_object(item, base_dir) # An optional entry whose media is null leaves the list rather than # reaching the pipeline as a reference with nothing in it @@ -895,6 +919,44 @@ def resolve_relative_path(path, base_dir): return path +def _describe_value_source(value): + """A short, human phrase for what a mistyped value already is - the + 'source' half of an argument-mismatch error, since the type name alone + (PIL.Image.Image) doesn't say *how* it got there.""" + if hasattr(value, "mode") and hasattr(value, "size"): + return "an already-loaded image" + if isinstance(value, tuple) and value and hasattr(value[0], "size"): + return "already-loaded video frames" + return f"a {type(value).__name__}" + + +def _fetch_image_with_context(v, base_dir, key): + """fetch_image, with the argument key folded into a type-mismatch error - + a bare 'got ' names neither the argument nor what the value + already was (#365).""" + try: + return fetch_image(v, base_dir) + except ValueError as error: + raise ValueError( + f"{error} (argument '{key}' expected an image, got " + f"{_describe_value_source(v)} - check what variable or previous " + f"result feeds it)" + ) from error + + +def _fetch_video_with_context(v, base_dir, key): + """fetch_video, with the same argument-key context as + _fetch_image_with_context.""" + try: + return fetch_video(v, base_dir) + except ValueError as error: + raise ValueError( + f"{error} (argument '{key}' expected a video, got " + f"{_describe_value_source(v)} - check what variable or previous " + f"result feeds it)" + ) from error + + def fetch_image(img_spec, base_dir=None): """ Load image from file path or URL with security validation. diff --git a/dw/workflow.py b/dw/workflow.py index ebbc7277..8ec14a98 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -896,8 +896,15 @@ def _prepare_definition(self, workflow_def, arguments, base_dir): # than starting a job the decode step was always going to OOM # on (dw/vram_estimate.py, #265) apply_vram_estimate(workflow_def, variables) - # realize the variables, initializing downloads of images etc - realize_args(variables, base_dir) + # realize the variables - explicit references only (asset:, + # output:, constant:, prompt:, a {media_type, location} dict). + # Key-name conventions (an 'image'/'video'/'_type' argument) are + # left off here: a variable's own name is not the argument it + # will end up filling, so a variable named 'image' fed to a step's + # 'video' argument was pre-loaded as a PIL Image before that step + # was ever substituted in (#365). The step-level realize_args + # passes below apply the conventions under the real argument key + realize_args(variables, base_dir, apply_key_conventions=False) ## then replace any variable references in the workflow definition with the actual values # replace_variables returns a new structure rather than mutating in # place, so the result must be captured here diff --git a/tests/test_arguments.py b/tests/test_arguments.py index 84afbc55..7f454c94 100644 --- a/tests/test_arguments.py +++ b/tests/test_arguments.py @@ -420,6 +420,80 @@ def test_realize_already_loaded_type(self): assert args["scheduler_type"] == mock_type + def test_variables_dict_skips_key_conventions(self): + # A variable named 'image', with apply_key_conventions off, is left + # as the plain path it was declared with - the value only loads once + # the argument that actually consumes it is realized (#365) + with tempfile.TemporaryDirectory() as temp_dir: + video_path = os.path.join(temp_dir, "clip.mp4") + with open(video_path, "wb") as f: + f.write(b"not a real video, just a placeholder") + + variables = {"image": video_path} + realize_args(variables, apply_key_conventions=False) + + assert variables["image"] == video_path + + def test_variable_named_image_feeds_video_argument(self): + # The exact #365 shape: a variable named 'image' is substituted into + # a step's 'video' argument. Realizing the variables dict without key + # conventions, then the step with them, loads it as a video rather + # than pre-loading it as an image and handing fetch_video a PIL Image + with patch("dw.arguments.load_video") as mock_load: + with patch("dw.arguments.validate_media_url") as mock_validate: + mock_validate.return_value = "https://example.com/clip.mp4" + mock_load.return_value = ["frame1", "frame2"] + + variables = {"image": "https://example.com/clip.mp4"} + realize_args(variables, apply_key_conventions=False) + + # simulate substitution of the variable into the step's argument + steps = {"video": variables["image"]} + realize_args(steps) + + assert steps["video"] == ["frame1", "frame2"] + + def test_star_image_variable_still_loads_as_image_under_image_argument(self): + # Regression check: a variable named like a media convention (e.g. + # 'input_image') still loads correctly once substituted into a + # matching argument - only the variable-stage guess is removed + with tempfile.TemporaryDirectory() as temp_dir: + test_image = Image.new("RGB", (50, 50), color="purple") + image_path = os.path.join(temp_dir, "subject.png") + test_image.save(image_path) + + variables = {"input_image": image_path} + realize_args(variables, apply_key_conventions=False) + assert variables["input_image"] == image_path + + steps = {"image": variables["input_image"]} + realize_args(steps, base_dir=temp_dir) + + assert isinstance(steps["image"], Image.Image) + + def test_reference_type_variable_stays_a_string(self): + # A '_type'-suffixed variable holding a plain category string must + # not be run through load_type_from_name at the variable stage + variables = {"reference_type": "character"} + realize_args(variables, apply_key_conventions=False) + + assert variables["reference_type"] == "character" + + def test_image_variable_fed_to_video_argument_error_names_argument(self): + # If a value still reaches the wrong loader, the error names the + # argument and describes the value's source rather than a raw + # '' message + img = Image.new("RGB", (10, 10)) + steps = {"video": img} + + with pytest.raises(ValueError) as exc_info: + realize_args(steps) + + message = str(exc_info.value) + assert "must be a string" in message + assert "'video'" in message + assert "already-loaded image" in message + class Reference: """Stands in for a pipeline argument built from a file, e.g. MiniMaxH3ImageReference""" diff --git a/tests/test_workflow_step_cache.py b/tests/test_workflow_step_cache.py index d36f2fcd..fe545965 100644 --- a/tests/test_workflow_step_cache.py +++ b/tests/test_workflow_step_cache.py @@ -126,8 +126,10 @@ def __deepcopy__(self, memo): original_realize_args = workflow_module.realize_args - def realize_and_poison(target, base_dir): - original_realize_args(target, base_dir) + def realize_and_poison(target, base_dir, apply_key_conventions=True): + original_realize_args( + target, base_dir, apply_key_conventions=apply_key_conventions + ) if isinstance(target, list): # the steps list, not the variables dict target[0]["pipeline"]["arguments"]["image"] = NotCopyable() From fed4ec588ec6f37631dab85c5a97e4e67187d225 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 23:02:20 -0500 Subject: [PATCH 028/181] fix(deploy): update SSH command syntax for clarity --- scripts/deploy.sh | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scripts/deploy.sh b/scripts/deploy.sh index 289cf005..8df6d313 100755 --- a/scripts/deploy.sh +++ b/scripts/deploy.sh @@ -4,7 +4,7 @@ # scripts/deploy.sh [branch] [--force] # # Run ON the server box (lem), from anywhere: -# ssh lem ~/diffusers-workflow/scripts/deploy.sh develop +# ssh lem '~/diffusers-workflow/scripts/deploy.sh develop' # # What it does, in order, stopping at the first failure: # 1. fetch; check out (default: the current branch); fast-forward From d10c8582fe5077a5cceab696f8a3512fc0723d81 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 23:18:09 -0500 Subject: [PATCH 029/181] fix(mcp): fit server instructions and tool descriptions under the client's 2,048-char cut Claude Code truncates an MCP server's instructions and each tool description at 2,048 characters. The instructions (4,056) lost the run loop, acknowledged_cost and the reference prefixes; validate_workflow (2,803) lost its plan/estimate paragraph; list_workflows (2,129) its tail. Rewrite all three under the limit, put the plan first in validate_workflow, and pin every text with CLIENT_TEXT_LIMIT. Also drop issue numbers from agent-facing docstrings and skills, and correct minimax-music3's "you cannot listen" now that get_output_audio exists. ltx-2.5 keeps its DFR fps note, trimmed elsewhere to stay under the 12 KiB skill cap. Co-Authored-By: Claude Opus 5.5 (1M context) --- dw_mcp/CLAUDE.md | 10 + dw_mcp/server.py | 212 ++++++++------------- plugins/dw/skills/ltx-2.5/SKILL.md | 27 ++- plugins/dw/skills/minimax-music3/SKILL.md | 19 +- plugins/dw/skills/script-to-video/SKILL.md | 5 +- plugins/dw/skills/series-episodes/SKILL.md | 4 +- tests/test_mcp_server.py | 22 ++- 7 files changed, 130 insertions(+), 169 deletions(-) diff --git a/dw_mcp/CLAUDE.md b/dw_mcp/CLAUDE.md index 50a53f2d..46e87ccc 100644 --- a/dw_mcp/CLAUDE.md +++ b/dw_mcp/CLAUDE.md @@ -80,3 +80,13 @@ run-directory delete already accepts, so a whole run goes in one call without a gallery listing to find its name. Authoring has two halves: `get_schema` describes a workflow and `get_prompt_schema` a stored prompt, which a workflow reaches by `"prompt:name"`. See docs/MCP.md. + +Claude Code shows an MCP server's `instructions` and each tool description +only up to 2,048 characters, then appends "[truncated]": nothing past that +reaches the agent. The instructions ran to 4,056 and cut off before the run +loop, `acknowledged_cost` and the reference prefixes; `validate_workflow`'s +`plan` paragraph sat past the cut. `CLIENT_TEXT_LIMIT` in +`tests/test_mcp_server.py` pins every one under it, alongside +`SURFACE_BUDGET`'s total. Put what an agent must act on first, and point at a +guide section rather than restating it - the instructions and `list_workflows` +sit within a few dozen characters of the limit. diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 4bb9e659..bd7396f8 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -67,88 +67,47 @@ def build_server(client): server = MCPServer( "diffusers-workflow", instructions=( - "Generate images and video on a real GPU: author, run and " - "diagnose diffusers-workflow jobs against a running dw.serve. " - "A workflow is a JSON document of named steps, each a " + "Generate images, video and audio on a real GPU: author, run " + "and diagnose diffusers-workflow jobs against a running " + "dw.serve. A workflow is a JSON document of named steps, each a " "diffusers pipeline or a utility task; the engine runs one job " "at a time.\n" "\n" - "Start from `list_workflows(shape=...)`: the server keeps a " - "large catalog, and its compact listing carries each " - "workflow's summary, shape, traits, cost, variable names and, " - "for a list-driven workflow, `lists` - what an entry of each " - "list carries; a list-driven workflow's `cost` may also carry " - "a measured `per_entry` (neither template does yet) - " - "run what is already there, with `arguments` overriding its " - "variables, rather than authoring a new workflow for a " - "request an existing one covers. Shapes: image, image-set, " - "image-edit, shot, sequence, audio, text, utility. Traits: " - "has-audio, chained, image-conditioned, identity-referenced, " - "needs-input-media, composes-workflows.\n" + "Start from `list_workflows(shape=...)` and run what the catalog " + "already holds, with `arguments` overriding its variables. " + "Shapes: image, image-set, image-edit, shot, sequence, audio, " + "text, utility. Traits: has-audio, chained, image-conditioned, " + "identity-referenced, needs-input-media, composes-workflows. An " + "open-ended request names a subject, not a shape - decide the " + "deliverable's shape first. `list_guides` indexes the docs by " + "section and `list_tasks` is what a shape is composed from; " + "author new JSON only when neither the catalog nor a " + "composition covers the request.\n" "\n" - "When a request is open-ended - a subject rather than a shape " - '("a lego movie trailer set in the marvel universe") - no ' - "catalog entry will name it, because entries are written in " - "shapes: a single image, an image set, one shot, a multi-shot " - "cut sequence, video with speech. Decide which shape the " - "deliverable is first, then call `list_workflows` with it; " - "`list_guides` indexes the documentation by section so a shape " - "can be looked up rather than guessed at, and `list_tasks` is " - "what a shape is composed from when no single workflow covers " - "it. Author new JSON only once neither does, and say what it " - "will cost before spending it.\n" - "\n" - "The engine that answers is one machine: `get_server_info` " - "reports its accelerator, its directories and which workspace " - "this session works in, and what a workflow can ask for " - "follows from that - a CUDA-only choice is not available on an " + "`get_server_info` reports the accelerator, directories and this " + "session's workspace; a CUDA-only choice is unavailable on an " "mps or cpu server.\n" "\n" - "The loop for anything that generates: `get_guide` " - '("workflows", section "Authoring a workflow from an agent") ' - "before writing or repairing any JSON, since the reference " - "conventions below are engine-specific and a draft that guesses " - "them validates and then fails at run time -> `validate_workflow` " - "(free, catches schema errors and arguments the pipeline does " - "not accept; repeat until it is clean, since fixing one layer " - "exposes the next) -> `run_workflow` -> `wait_for_job` rather than a " - "polling loop -> `get_job` for the manifest -> " - "`get_output_image`, `get_output_frames` and `get_output_audio` to " - "actually look at and listen to what was made and say whether " - "it answers the request. Tools that cost GPU minutes, " - "disk or unrecoverable deletion refuse until " - "`acknowledged_cost=true`: tell the user what it will cost (a " - "workflow you wrote or copied has no `cost`; quote the figure " - "from the `models/` entry that loads the same pipeline, found " - "with `list_workflows(include_models=true)`, times the number " - "of images), get their go-ahead, then call again.\n" - "\n" - "Workflow arguments carry references rather than literals, " - "which is what makes multi-stage work composable: " - '"variable:name" (an override), ' - '"previous_result:step" (an earlier step in the same run), ' - '"prompt:name" or "prompt:folder/name" (the stored prompt ' - "library - `list_prompts`, `get_prompt_schema`), " - '"asset:name.ext" (input media on the server - `list_assets`, ' - "`upload_asset` to push a local file, `keep_output` to promote " - "a generated file into a stable input), and " - '"output:workflow/run-id/file.png" (a file an earlier run ' - 'wrote, with "latest" in the run-id position picking the ' - "newest run holding it). Prefer an asset: or output: reference " - "over a filesystem path: a path on this machine usually means " - "nothing to the server.\n" + "The loop: `get_guide(\"workflows\", section=\"Authoring a " + "workflow from an agent\")` before writing or repairing JSON -> " + "`validate_workflow` (free; repeat until clean) -> quote its " + "`plan.estimate` and get the user's go-ahead -> `run_workflow` " + "-> `wait_for_job` -> `get_job` -> `get_output_image`, " + "`get_output_frames`, `get_output_audio` to look at and listen " + "to the result and judge it against the request. Tools that " + "spend GPU time or disk, or delete for good, refuse until " + "`acknowledged_cost` is set. A workflow you wrote has no " + "measured cost: quote the `models/` entry that loads the same " + "pipeline (`list_workflows(include_models=true)`) times the " + "number of images.\n" "\n" - "Each run writes its own directory, " - "//, with a manifest beside its files; " - "`list_gallery` and a job's manifest name files the way " - "`get_output_image`, `download_output` and `keep_output` " - "expect them.\n" - "\n" - "The server can hold several workspaces - separate workflows, " - "assets and outputs, one shared prompt library: " - "`list_workspaces` shows them and `use_workspace` picks one " - "for the rest of the session, which is how to keep your work " - "out of another agent's namespace." + "Arguments carry references rather than literals: `variable:`, " + "`previous_result:`, `prompt:` (the stored prompt library), " + "`asset:` (input media on the server - `upload_asset`, " + "`keep_output`) and `output:` (an earlier run's file). The " + "guide's References section defines each; prefer them to a " + "local path, which means nothing to the server. " + "`use_workspace` picks the workspace this session works in." ), ) @@ -180,10 +139,9 @@ def list_workflows( image-conditioned, identity-referenced, needs-input-media, composes-workflows. Each entry carries a one-line `summary`, its `shape` and `traits` (what it needs supplied), `cost` (curated: - figures a maintainer measured once on the devices named and wrote - into the workflow, never derived from this server's job history - - so null means nobody wrote one down, not that the run is cheap; - the answer's `cost_basis` says as much. It is also the mark of an + measured once by a maintainer on the devices named, never derived + from this server's history - null means nobody wrote one down, not + that the run is cheap. It is also the mark of an entry that has been run through on a real device: one without a `cost` has only been authored, and its first run is the one that finds what the description could not verify. `observed_minutes` and @@ -191,9 +149,8 @@ def list_workflows( of that workflow - the cold median, model load included, and how many runs are behind it. Prefer it when quoting a price for this machine, fall back to `cost`, and say "unknown" only when neither - is there; say which one you used, since a figure a maintainer - measured on their card and one this box averaged last week are - different claims), output kinds and variable names. `lists`, present for a + is there; say which one you used - a maintainer's card and this + box's average are different claims), output kinds and variable names. `lists`, present for a list-driven workflow, names per list variable the fields an entry takes, the steps run over it and the default's length; there `cost[].per_entry`, when present, is the measured cost of one @@ -424,7 +381,7 @@ def list_gallery( writing anything, invisible to a normal listing because it has no file to show. `subfolder` does not apply in this mode. `name` is exactly what `delete_output` accepts, so clearing the backlog is - list, then delete each name (#170). A run that wrote any file at + list, then delete each name. A run that wrote any file at all - a text-shape prompt, a utility's side output - is not listed; this call only lists, so deciding whether a listed entry is actually junk before calling `delete_output` on it is still yours to make. @@ -432,7 +389,7 @@ def list_gallery( `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace is reachable from - here without leaving this one (#99).""" + here without leaving this one.""" return catalog.list_gallery( client, limit=limit, @@ -472,7 +429,7 @@ def get_gallery_metadata( and rates are arguments the caller supplies, and a wrong one is a failed job or, worse, silence padded onto the end of a track. - `workspace` pins this call to another workspace (#99).""" + `workspace` pins this call to another workspace.""" return catalog.get_gallery_metadata( client, name, envelope=envelope, workspace=workspace ) @@ -538,7 +495,7 @@ def get_output_image( `crop` is `[x, y, width, height]` in the original's pixels, cut before the downscale. - `workspace` pins this call to another workspace (#99).""" + `workspace` pins this call to another workspace.""" result = media.get_output_image( client, name, max_dimension=max_dimension, workspace=workspace, crop=crop ) @@ -571,7 +528,7 @@ def get_output_audio( seconds, per `get_gallery_metadata`'s envelope. The text part says what was cut. To *see* a video, `get_output_frames`. - `workspace` pins this call to another workspace (#99).""" + `workspace` pins this call to another workspace.""" result = media.get_output_audio( client, name, start=start, duration=duration, workspace=workspace ) @@ -610,7 +567,7 @@ def get_output_frames( pixels, cut from every frame before any downscale, like `get_output_image`'s. - `workspace` pins this call to another workspace (#99).""" + `workspace` pins this call to another workspace.""" result = media.get_output_frames( client, name, @@ -687,7 +644,7 @@ def get_output_text( `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace is reachable from - here without leaving this one (#99).""" + here without leaving this one.""" return media.get_output_text( client, name, max_characters=max_characters, workspace=workspace ) @@ -710,7 +667,7 @@ def delete_output( directory, or unknown, is an error. `workspace` pins this call to another workspace without switching - the session (#99); a `job_id` delete with no `workspace` goes to + the session; a `job_id` delete with no `workspace` goes to the workspace the job ran in.""" return media.delete_output(client, name, workspace=workspace, job_id=job_id) @@ -743,7 +700,7 @@ def download_output( `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace is reachable from - here without leaving this one (#99).""" + here without leaving this one.""" return media.download_output( client, name, @@ -837,7 +794,7 @@ def keep_output( `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace is reachable from - here without leaving this one (#99).""" + here without leaving this one.""" return assets.keep_output( client, name, @@ -916,55 +873,36 @@ def validate_workflow( ) -> dict: """Check a workflow against the schema and against real pipeline signatures. Free and instant - always run this before run_workflow. - Give exactly one of `workflow` or `name` - `name` being a stored - workflow as `list_workflows` reports it. `run_workflow` calls these - same two concepts `inline_workflow` and `workflow_path`; both tools - accept both spellings, so a definition or a name checked here can be - handed straight to `run_workflow` without renaming a key. Every - schema error comes back at once, each with its JSON path. - `workflow`/`inline_workflow` may also be a JSON-encoded string; a - parse failure is reported as invalid JSON, not a type mismatch. - `workspace` names the workspace for this one call without switching - the session to it - use it to pin a job whose `output:` or `asset:` - references live in a workspace other than the session's. - - Pass the same `arguments` you will pass to `run_workflow` and they - are checked too: a name the workflow no longer declares, a value - that will not coerce to the declared type, and an `asset:`, - `prompt:` or `output:` reference that names nothing this workspace - can reach - each with `arguments.` as its path. - `checked_arguments` lists what was checked, so a valid answer says - whether it covered your values or only the stored defaults. - - A value outside a bound the workflow declares for that variable is - an error here rather than a failed run: H3's frame count has to be - `17 * n + 5` between 124 and 345, and 61 used to validate and then - fail 138 s into the run, after the weights were loaded. A value the - workflow rounds up instead of refusing comes back as a warning - naming what it becomes, so a frame count the run changes is known - before the run. `get_workflow(variables_only=true)` and - `list_workflows` report the bound beside the default. - - A `result.subfolder` or `file_base_name` that cannot be written (a - `..`, a backslash, a separator in `file_base_name`) is reported here - at its JSON path, after `for_each` expansion. - - A sub-workflow step is resolved too: a `workflow.path` that names - nothing this server can reach is an error at - `steps[N].workflow.path` (the message lists where it looked), the - workflow it names is validated in turn under that path, a - composition cycle is refused, and an argument passed down that the - composed workflow declares no variable for comes back as a - warning. + Give exactly one of `workflow` or `name` (a stored workflow as + `list_workflows` reports it); `run_workflow`'s `inline_workflow` + and `workflow_path` spellings are accepted here too. `workflow` may + be a JSON-encoded string. Every error comes back at once, each with + its JSON path. `workspace` pins this one call to another workspace + without switching the session. A valid answer carries `plan`: what will execute for these arguments - `estimate.minutes` and its `basis` (`observed`, `per_entry`, `catalog`, `derived`, `other_device` or `unknown` - - what each means and how to quote it is WORKFLOW_GUIDE's "The loop", - step 4), each `downloads_required` entry as its own cost line, and - `steps`/`list_entries` for how many members the list actually - produced. `plan` is null when it could not be built; the verdict - stands.""" + how to quote each is WORKFLOW_GUIDE's "The loop", step 4), each + `downloads_required` entry as its own cost line, and + `steps`/`list_entries` for how many members the list produced. + `plan` is null when it could not be built; the verdict stands. + + Pass the same `arguments` you will pass to `run_workflow` and they + are checked too: an undeclared name, a value that will not coerce, + and an `asset:`/`prompt:`/`output:` reference naming nothing this + workspace can reach, each at `arguments.`. + `checked_arguments` says whether your values or only the stored + defaults were checked. A value outside a bound the workflow + declares (H3's frame count is `17 * n + 5` from 124 to 345) is an + error here rather than a failed run; one the workflow rounds up + comes back as a warning naming what it becomes. + + Also checked: an unwritable `result.subfolder` or `file_base_name`, + after `for_each` expansion; and each sub-workflow step - an + unreachable `workflow.path`, the composed workflow in turn, a + composition cycle, and (as a warning) an argument passed down that + it declares no variable for.""" return authoring.validate_workflow( client, workflow=workflow, diff --git a/plugins/dw/skills/ltx-2.5/SKILL.md b/plugins/dw/skills/ltx-2.5/SKILL.md index b363cd9c..99991f41 100644 --- a/plugins/dw/skills/ltx-2.5/SKILL.md +++ b/plugins/dw/skills/ltx-2.5/SKILL.md @@ -7,8 +7,8 @@ description: Use when a dw MCP server is connected and the user wants LTX-2.5 vi LTX-2.5 generates video and a soundtrack together, 24 fps, on a distilled schedule that is not a knob. Every template here fits a 24 GB card. This -skill chooses the template and the arguments; the prompt is written to -Lightricks' own caption spec, quoted below from diffusers. +skill chooses the template and the arguments; the prompt follows the +caption spec below. ## Before anything @@ -23,11 +23,11 @@ Lightricks' own caption spec, quoted below from diffusers. 2x upscale - `get_memory` on an idle server and read a `live: true` reading's `gpu_memory_allocated_mb`. Only those are the worker's own: `info: null` means nothing is resident, and a `live: false` reading is - cached from another moment. A non-trivial idle figure is what an earlier - run left behind and comes off what this one has. With the server idle, + cached from another moment. A non-trivial idle figure is an earlier run's + leftover and comes off what this one has. With the server idle, `clear_memory` clears it (refused while a job is queued or running) - it also drops the step cache, so the next run, including a seeded rerun, is - cold and regenerates. Re-read `get_memory` afterwards to confirm. Don't + cold and regenerates. Re-read `get_memory` to confirm. Don't retry into a failed attempt without clearing first - a failed attempt is itself what leaves weight resident. @@ -43,9 +43,8 @@ Lightricks' own caption spec, quoted below from diffusers. 768x448, a 2x latent upsample, then renoise and three stage-two sigmas at 1536x896 carrying the audio latents through. The upsample alone is soft; the refine pass is where the detail comes from. -- **Comparing decoders**: `templates/ltx2/diffusion-decode`. Do not offer it: - without a `shi-labs/natten` build for the installed torch the FlexAttention - fallback OOMs on 24GB at any size (#153). +- `templates/ltx2/diffusion-decode` compares decoders. Do not offer it: without + a `shi-labs/natten` build its fallback OOMs on 24GB at any size. - **A generative 2x render**: `templates/ltx2/generative-upscale` draws its own low-resolution pass and has an IC-LoRA re-render it twice the size, inventing detail. `base_width`/`base_height` are that first render's size, @@ -61,8 +60,8 @@ Lightricks' own caption spec, quoted below from diffusers. make. Each inverts one defect and no other - neither is an upscale, neither removes motion blur or grain - so say which defect you think it is and let the user correct you. -- **Longer**: `templates/ltx2/extend-clip` generates an opening and continues it - conditioned on the whole opening, not one frame; +- **Longer**: `templates/ltx2/extend-clip` continues an opening conditioned + on all of it, not one frame; `templates/ltx2/chained-segments` re-runs per segment on the previous last frame and stitches. Neither is a Lightricks recipe; both are dw's, and a single 481-frame pass reaches 20 seconds before either is needed. @@ -76,12 +75,12 @@ read the `workflows` guide's authoring section first. The templates generate at 24 fps. RoPE time is `frame / fps` and the model is trained around 24, 25, 30 and 60, so for a higher-fps request generate at 24 (or condition at 60 at most - `MAX_CONDITIONING_FPS`, never 120) and let - playback set the rate. The temporal-upscaling path that renders 48 and 96 - fps belongs to the DFR pipelines, which no template here uses yet. + playback set the rate. 48 and 96 fps are the DFR pipelines', which no + template uses. - The distilled transformer runs its eight trained sigmas (`DISTILLED_SIGMA_VALUES`) with `guidance_scale` 1.0 and STG and modality guidance off. No - `num_inference_steps`. Those knobs mean something only against the dev - transformer, which no 24 GB template ships. + `num_inference_steps`. Those knobs belong to the dev transformer, which + no 24 GB template ships. - Stage two of the two-stage flow: renoise at 0.909375 (the first `STAGE_2_DISTILLED_SIGMA_VALUES` entry), three sigmas at full size. - An image condition is re-compressed at CRF 18 to match training and needs a diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 981da119..840973af 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -148,16 +148,15 @@ Control" section. deliverables. Keep the convention in anything you compose from a template: the step whose output the user will be shown is `final`, every other saving step `intermediate`. -4. You cannot listen: no tool returns audio inline. Hand the user the gallery - `url` (`list_gallery`, or the manifest's file name) and check what you can - yourself - `get_gallery_metadata` for the file's duration against the - ceiling: `media.duration_seconds` within 0.2 s of `audio_duration` means - the ceiling cut the track (raise it and rerun); well short of it means the - song finished on its own. Also check the sample rate. Ask the user to - listen for the family's failure - modes: a song that went instrumental (name the vocals in the caption), an - ending cut mid-note (raise the ceiling, then trim), a structure that ignored - the tags (fewer sections, plainer directions). +4. Judge it yourself. `get_gallery_metadata` for duration and sample rate: + `media.duration_seconds` within 0.2 s of `audio_duration` means the + ceiling cut the track (raise it and rerun); well short of it means the + song finished on its own. Then listen with `get_output_audio` (a long + track in `start`/`duration` excerpts) for the family's failure modes: a + song that went instrumental (name the vocals in the caption), an ending + cut mid-note (raise the ceiling, then trim), a structure that ignored the + tags (fewer sections, plainer directions). Hand the user the gallery + `url` (`list_gallery`, or the manifest's file name). 5. To use the track in a later workflow, `keep_output` makes it an `asset:`; to trim it in the same run, chain `templates/audio-trim-fade` on the output. 6. After an inline run worth keeping, `get_job_workflow` and `save_workflow` it, diff --git a/plugins/dw/skills/script-to-video/SKILL.md b/plugins/dw/skills/script-to-video/SKILL.md index 9a8a53a9..0a28e1f5 100644 --- a/plugins/dw/skills/script-to-video/SKILL.md +++ b/plugins/dw/skills/script-to-video/SKILL.md @@ -24,10 +24,7 @@ does not parse screenplay formats. The unit that matters is *shots*, not lines or scenes: one scene may be several shots, and `dw:minimax-h3`'s sweet spot is 4-6 second shots (the `17*n+5` frame grid). For each shot, name it, note which character(s) -appear, what happens, and roughly how long it runs. This is genuine -reasoning work with no existing mechanism to lean on - the one step where -"an LLM does this well" is actually true here, unlike auto-tuning a slow -model. +appear, what happens, and roughly how long it runs. ## 3. Cast recurring characters once, before any shot generates diff --git a/plugins/dw/skills/series-episodes/SKILL.md b/plugins/dw/skills/series-episodes/SKILL.md index f53944ea..1b8cb21e 100644 --- a/plugins/dw/skills/series-episodes/SKILL.md +++ b/plugins/dw/skills/series-episodes/SKILL.md @@ -123,6 +123,4 @@ mismatch between episodes shows up before a viewer notices it. `workflows/templates/assemble-and-score.json`, `dw/tasks/audio_utils.py` (`loop_audio`, `match_levels`, `normalize_audio`), `dw/tasks/pair_audio.py`, `docs/WORKSPACES.md` (the shared asset library), the `minimax-h3` skill -this composes into. Written from issue #217 (dkackman/diffusers-workflow), -which named the drift as a recurring, undocumented failure across two -hand-built episodes. +this composes into. diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 4248fedb..ce0a5eaa 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -1304,13 +1304,33 @@ async def test_the_instructions_name_the_vocabulary(): assert word in server.instructions +# Claude Code shows an MCP server's instructions and each tool description +# only up to this many characters and appends "[truncated]"; everything past +# it never reaches the agent, however carefully it was written. +CLIENT_TEXT_LIMIT = 2048 + + +@pytest.mark.asyncio +async def test_no_text_the_agent_reads_is_cut_off_by_the_client(): + server = server_over(ok({})) + tools = await tools_of(server) + + assert len(server.instructions) <= CLIENT_TEXT_LIMIT + too_long = { + name: len(tool.description or "") + for name, tool in tools.items() + if len(tool.description or "") > CLIENT_TEXT_LIMIT + } + assert not too_long, too_long + + @pytest.mark.asyncio async def test_validate_workflow_teaches_quoting_from_the_plan(): """The number an agent says out loud is the plan's - priced for the arguments it will run with, naming the weights this box lacks - not the listing's defaults-only cost (#85).""" tools = await tools_of(server_over(ok({}))) - doc = tools["validate_workflow"].description + doc = tools["validate_workflow"].description[:CLIENT_TEXT_LIMIT] assert "plan" in doc assert "downloads_required" in doc assert "estimate" in doc From c1e103b333168f148056f26120a0f4d1b15cc939 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Tue, 22 Sep 2026 23:22:53 -0500 Subject: [PATCH 030/181] format --- dw/arguments.py | 4 +--- dw/download_watch.py | 4 +++- dw/events.py | 4 +++- dw/introspection.py | 8 ++++++-- dw/plan.py | 14 ++++++++++---- dw/workflow.py | 4 +++- dw_mcp/server.py | 4 ++-- tests/test_component_type_errors.py | 19 +++++++++++++++---- tests/test_download_watch.py | 8 ++++++-- tests/test_host_memory_projection.py | 10 ++++++++-- tests/test_server_exports.py | 4 +++- tests/test_task_signature_errors.py | 1 - tests/test_video_extensions.py | 5 ++++- ui/src/lib/pages/GalleryPage.svelte | 1 + 14 files changed, 65 insertions(+), 25 deletions(-) diff --git a/dw/arguments.py b/dw/arguments.py index 874d16e4..70867641 100644 --- a/dw/arguments.py +++ b/dw/arguments.py @@ -220,9 +220,7 @@ def realize_args(arg, base_dir=None, apply_key_conventions=True): realize_args(item, base_dir) except ValueError as error: if isinstance(item, dict) and "name" in item: - raise ValueError( - f"{error} (step '{item['name']}')" - ) from error + raise ValueError(f"{error} (step '{item['name']}')") from error raise arg[i] = realize_object(item, base_dir) # An optional entry whose media is null leaves the list rather than diff --git a/dw/download_watch.py b/dw/download_watch.py index e20b493a..c40ced60 100644 --- a/dw/download_watch.py +++ b/dw/download_watch.py @@ -148,7 +148,9 @@ def watch(repo_id, cache_dir=None, repo_type="model"): if not is_watchable_repo_id(repo_id): return _NULL_WATCH - return DownloadWatch(repo_id, get_context(), cache_dir=cache_dir, repo_type=repo_type) + return DownloadWatch( + repo_id, get_context(), cache_dir=cache_dir, repo_type=repo_type + ) class _NullWatch: diff --git a/dw/events.py b/dw/events.py index d516f347..8ad10de2 100644 --- a/dw/events.py +++ b/dw/events.py @@ -139,7 +139,9 @@ def _watchdog_loop(self): seconds_since_phase_start = round(now - self._phase_started_at, 1) seconds_since_last_progress = round(now - self._last_progress_at, 1) last_kind = self._last_progress_kind or "phase start" - last_offset = round(max(0.0, self._last_progress_at - self._phase_started_at), 1) + last_offset = round( + max(0.0, self._last_progress_at - self._phase_started_at), 1 + ) message = ( f"no progress event for {seconds_since_last_progress:.1f}s in phase " f"'{phase}' (last: {last_kind} at +{last_offset:.1f}s) - informational; " diff --git a/dw/introspection.py b/dw/introspection.py index 0f6b6563..e32b3634 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -601,7 +601,9 @@ def _null_fed_variable(written_steps, source_index, key, declared_variables): return name if name in declared_variables else None -def task_signature_errors(workflow_definition, source_indices=None, written_definition=None): +def task_signature_errors( + workflow_definition, source_indices=None, written_definition=None +): """Every task step whose arguments its command's signature refuses, as [{path, message}] - a required argument left unset, and an argument the command does not take - plus a step naming a command that is not @@ -737,7 +739,9 @@ def report(key, message, variable=None): # A class-name-shaped string, bare or dotted - excludes a {}-escaped literal # and a variable:/constant:/asset:/... reference, which use ':' or braces # and are checked elsewhere -_DOTTED_NAME_PATTERN = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)*$") +_DOTTED_NAME_PATTERN = re.compile( + r"^[A-Za-z_][A-Za-z0-9_]*(\.[A-Za-z_][A-Za-z0-9_]*)*$" +) def _type_reference_candidates(key): diff --git a/dw/plan.py b/dw/plan.py index aa69cf69..971f5414 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -509,7 +509,9 @@ def estimate( child_measured_entries = {} child_expanded = {"variables": {}} if child_definition is not None: - child_measured_entries = _list_entries(child_definition, child_definition) + child_measured_entries = _list_entries( + child_definition, child_definition + ) # The composing step's own `arguments` are what the child # actually runs with - folded over its declared defaults the # same way a caller's arguments are, since `expanded` has @@ -521,9 +523,13 @@ def estimate( child = _price( child_cost, device, child_list_entries, child_measured_entries ) - if child["basis"] == CATALOG and child_definition is not None and ( - _scalar_driver_shifted( - child_definition, child_expanded, child_list_entries + if ( + child["basis"] == CATALOG + and child_definition is not None + and ( + _scalar_driver_shifted( + child_definition, child_expanded, child_list_entries + ) ) ): # A scalar cost_driver the composing step overrode (H3's diff --git a/dw/workflow.py b/dw/workflow.py index 8ec14a98..ddf14099 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -638,7 +638,9 @@ def validation_errors(self, arguments=None, composing=None): base_dir = ( os.path.dirname(os.path.abspath(self.file_spec)) if self.file_spec else None ) - task_errors = task_signature_errors(expanded, source_indices, self.workflow_definition) + task_errors = task_signature_errors( + expanded, source_indices, self.workflow_definition + ) if arguments is None: # A step that feeds a required argument from `variable:name` and # a variable whose default is null is a fine document - the diff --git a/dw_mcp/server.py b/dw_mcp/server.py index bd7396f8..447a6f4c 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -88,8 +88,8 @@ def build_server(client): "session's workspace; a CUDA-only choice is unavailable on an " "mps or cpu server.\n" "\n" - "The loop: `get_guide(\"workflows\", section=\"Authoring a " - "workflow from an agent\")` before writing or repairing JSON -> " + 'The loop: `get_guide("workflows", section="Authoring a ' + 'workflow from an agent")` before writing or repairing JSON -> ' "`validate_workflow` (free; repeat until clean) -> quote its " "`plan.estimate` and get the user's go-ahead -> `run_workflow` " "-> `wait_for_job` -> `get_job` -> `get_output_image`, " diff --git a/tests/test_component_type_errors.py b/tests/test_component_type_errors.py index 3aed385d..9a00741c 100644 --- a/tests/test_component_type_errors.py +++ b/tests/test_component_type_errors.py @@ -69,7 +69,9 @@ def test_an_absent_class_in_a_trusted_module_says_does_not_exist(self): "steps": [ pipeline_step( None, - extra={"quantization_config": {"config_type": "sdnq.NoSuchConfig"}}, + extra={ + "quantization_config": {"config_type": "sdnq.NoSuchConfig"} + }, ) ], } @@ -95,7 +97,9 @@ def test_the_two_messages_differ(self, monkeypatch): "steps": [ pipeline_step( None, - extra={"quantization_config": {"config_type": "sdnq.NoSuchConfig"}}, + extra={ + "quantization_config": {"config_type": "sdnq.NoSuchConfig"} + }, ) ], } @@ -112,7 +116,9 @@ def test_a_misspelled_scheduler_type_is_refused(self): { "name": "a", "pipeline": { - "scheduler": {"configuration": {"scheduler_type": "DDIMSchedulr"}} + "scheduler": { + "configuration": {"scheduler_type": "DDIMSchedulr"} + } }, "result": {"content_type": "image/png"}, } @@ -127,7 +133,12 @@ def test_a_misspelled_scheduler_type_is_refused(self): def test_a_real_config_type_is_accepted(self): definition = { "id": "ct", - "steps": [pipeline_step(None, extra={"quantization_config": {"config_type": "sdnq.SDNQConfig"}})], + "steps": [ + pipeline_step( + None, + extra={"quantization_config": {"config_type": "sdnq.SDNQConfig"}}, + ) + ], } assert component_type_errors(definition) == [] diff --git a/tests/test_download_watch.py b/tests/test_download_watch.py index 03cb6bf9..401a8a06 100644 --- a/tests/test_download_watch.py +++ b/tests/test_download_watch.py @@ -37,7 +37,9 @@ def test_growing_download_emits_progress_and_suppresses_stall(tmp_path, monkeypa context.enter_run() try: context.note_phase("loading") - with download_watch.DownloadWatch(repo_id, context, cache_dir=str(tmp_path)): + with download_watch.DownloadWatch( + repo_id, context, cache_dir=str(tmp_path) + ): for _ in range(6): with open(blob_file, "ab") as f: f.write(b"x" * 4096) @@ -68,7 +70,9 @@ def test_stalled_download_still_stalls(tmp_path, monkeypatch): context.enter_run() try: context.note_phase("loading") - with download_watch.DownloadWatch(repo_id, context, cache_dir=str(tmp_path)): + with download_watch.DownloadWatch( + repo_id, context, cache_dir=str(tmp_path) + ): time.sleep(0.3) finally: context.exit_run() diff --git a/tests/test_host_memory_projection.py b/tests/test_host_memory_projection.py index 52a31303..f7906d86 100644 --- a/tests/test_host_memory_projection.py +++ b/tests/test_host_memory_projection.py @@ -92,7 +92,10 @@ def test_two_list_lengths_fit_base_plus_marginal_growth(self): # peaked at 31000 MB, 5 entries at 62000 MB -> a 6-entry request # should project roughly 70000 MB, not 31000 * 6 = 186000 MB definition = for_each_workflow(release=False) - rows = [row(31000 * MB, {"shots": [1]}), row(62000 * MB, {"shots": [1, 2, 3, 4, 5]})] + rows = [ + row(31000 * MB, {"shots": [1]}), + row(62000 * MB, {"shots": [1, 2, 3, 4, 5]}), + ] warnings = host_memory_warnings(definition, {"shots": 6}, rows, 8_000 * MB) assert len(warnings) == 1 assert "69750" in warnings[0] or "69,750" in warnings[0] @@ -101,7 +104,10 @@ def test_two_list_lengths_fit_base_plus_marginal_growth(self): def test_growth_fit_under_the_ceiling_warns_nothing(self): definition = for_each_workflow(release=False) - rows = [row(31000 * MB, {"shots": [1]}), row(62000 * MB, {"shots": [1, 2, 3, 4, 5]})] + rows = [ + row(31000 * MB, {"shots": [1]}), + row(62000 * MB, {"shots": [1, 2, 3, 4, 5]}), + ] warnings = host_memory_warnings(definition, {"shots": 2}, rows, 100_000 * MB) assert warnings == [] diff --git a/tests/test_server_exports.py b/tests/test_server_exports.py index cfa24bc4..b73544c4 100644 --- a/tests/test_server_exports.py +++ b/tests/test_server_exports.py @@ -405,7 +405,9 @@ def test_an_absolute_zip_url_is_added_when_a_public_url_is_configured( job_id = finished(client) body = client.post(f"/api/jobs/{job_id}/export").json() - assert body["absolute_zip_url"] == f"https://dw.example.com/exports/{job_id}.zip" + assert ( + body["absolute_zip_url"] == f"https://dw.example.com/exports/{job_id}.zip" + ) def test_no_absolute_zip_url_when_no_public_url_is_configured(self, server): with server() as client: diff --git a/tests/test_task_signature_errors.py b/tests/test_task_signature_errors.py index 4636b37a..02cc4d7e 100644 --- a/tests/test_task_signature_errors.py +++ b/tests/test_task_signature_errors.py @@ -23,7 +23,6 @@ class of mistake a free pre-flight most obviously exists for (#141). An workflow_argument_warnings, ) from dw.tasks.task import Task -from dw.workflow import Workflow def task_step(command, arguments, name="a"): diff --git a/tests/test_video_extensions.py b/tests/test_video_extensions.py index cbc731bd..3ad7a6b0 100644 --- a/tests/test_video_extensions.py +++ b/tests/test_video_extensions.py @@ -57,7 +57,10 @@ def test_a_path_with_no_extension_is_not_this_passs_complaint(self): def test_a_non_string_is_not_this_passs_complaint(self): assert _extension_problem(None) is None - assert _extension_problem({"media_type": "video", "location": "asset:x.mp4"}) is None + assert ( + _extension_problem({"media_type": "video", "location": "asset:x.mp4"}) + is None + ) class TestTheValidationPass: diff --git a/ui/src/lib/pages/GalleryPage.svelte b/ui/src/lib/pages/GalleryPage.svelte index 32507caf..29895125 100644 --- a/ui/src/lib/pages/GalleryPage.svelte +++ b/ui/src/lib/pages/GalleryPage.svelte @@ -329,6 +329,7 @@ + {' '} {/if}{file.label} From a737bbdde38eacc5da3f8969c90d0704378c6a0f Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 00:53:20 -0500 Subject: [PATCH 031/181] docs(proposals): retire the two shipped proposals - h3-video-mux-headroom-warning-partial.md: fixes (1)-(4) shipped in bbe4adb and the suite wording landed via #323. The resolution and the still-deferred fix (2), with its undershoot evidence, move into the complete doc. - live-memory-during-run.md: design (a) shipped in fe824c9 (#273). The unshipped fallback (b) is recorded on #273. - todo.md: Tier 1 is empty. Open items are now `feature` issues (#374-#380, #244). Co-Authored-By: Claude Opus 5.5 (1M context) --- .../h3-video-mux-headroom-warning-complete.md | 30 ++++++ .../h3-video-mux-headroom-warning-partial.md | 59 ------------ docs/proposals/live-memory-during-run.md | 96 ------------------- docs/proposals/todo.md | 12 ++- 4 files changed, 37 insertions(+), 160 deletions(-) delete mode 100644 docs/proposals/h3-video-mux-headroom-warning-partial.md delete mode 100644 docs/proposals/live-memory-during-run.md diff --git a/docs/proposals/complete/h3-video-mux-headroom-warning-complete.md b/docs/proposals/complete/h3-video-mux-headroom-warning-complete.md index cf04a4be..a7aaff0d 100644 --- a/docs/proposals/complete/h3-video-mux-headroom-warning-complete.md +++ b/docs/proposals/complete/h3-video-mux-headroom-warning-complete.md @@ -205,3 +205,33 @@ emitting nothing when it's clean — scoped to `content_type.startswith("video") only, audio-only saves unaffected. Fix (2) (a `normalize_audio` gain step on the fifteen H3 video templates) remains held, unchanged from the original decision — no run on record has an H3 video mux actually clipping. + +## Resolution (2026-09-20) + +Approved and implemented in `bbe4adb`, fixes (1)-(4): + +1. For a video mux (`content_type.startswith("video")`), the pre-encode + `warn_without_headroom` result is held rather than emitted immediately. +2. The post-encode `warn_if_written_above_full_scale` probe runs as before. +3. The held warning is dropped if the written file measures clean, and + upgraded to `audio_clipped` with the measured value if the file is genuinely + over. +4. Audio-only saves are unaffected (unconditional, unheld). + +The regression cases were updated to match: S-F031 and M-F008 in the external +suite (#323). A video mux now carries at most one headroom warning, and it +is the post-encode one. + +**Fix (2) is still deferred, and no decision is needed yet.** It would add a +`normalize_audio` gain stage to the fifteen H3 video templates. No run on +record has an H3 video mux actually clipping. Wait for a real +clipped-in-practice case to size the gain against, the way #158 gave #159 one. +The undershoot record this rests on: + +| Run | Template / mux | Predicted (pre-encode) | Written (post-encode) | Delta | +|---|---|---|---|---| +| #161's `music-video` case | AAC, different song | — | — | +1.94 dB (the one positive measurement on record, a different template family) | +| `M-F012.jsonl` #1 | H3 video mux | target | measured | -0.92 dB | +| `M-F012.jsonl` #2 | H3 video mux | target | measured | -1.11 dB | +| M-F008 original report | `video-with-audio-768p` | -0.14 dBFS | -1.1158 dBFS | -0.98 dB | +| Tester's 2026-09-16 re-verify | `video-with-audio-768p` | -0.14 dBFS | -1.1158844… dBFS | -0.97 dB | diff --git a/docs/proposals/h3-video-mux-headroom-warning-partial.md b/docs/proposals/h3-video-mux-headroom-warning-partial.md deleted file mode 100644 index 5a04f7b0..00000000 --- a/docs/proposals/h3-video-mux-headroom-warning-partial.md +++ /dev/null @@ -1,59 +0,0 @@ -# Remaining work: H3 video mux headroom warning (#174) - pre-encode/post-encode reconciliation - -Split from `h3-video-mux-headroom-warning-complete.md` on 2026-09-20. Fix (1) -— stop suppressing the post-encode ground-truth probe for a video mux — is -implemented (`dw/result.py`, shipped `ce82f06`/merged `aab4ef5`, deployed to -`lem`). The tester's 2026-09-16 verification run showed it is insufficient on -its own: the pre-encode `audio_no_headroom` warning still fires unconditionally -before the write, independent of what the post-encode probe finds, so -`warnings: []` still fails on `video-with-audio-768p` stock defaults. What -remains is the amendment's open decision and its implementation, unresolved: - -## The remaining decision, "Asking" (2026-09-16 amendment) - -Approve or decline holding the pre-encode `audio_no_headroom` warning for a -video mux until the post-encode probe has run, replacing it with -`audio_clipped` (measured value) when the file is genuinely over, and -emitting nothing when it's clean — scoped to `content_type.startswith("video")` -only, audio-only saves unaffected. - -This is a real behavior change beyond fix (1): it changes *when* a warning -reaches the caller (after the write completes, not at the point the risk is -detected) and *which* checks a video mux can ever surface (post-encode ground -truth only, never the prediction) — for every content type that already goes -through `warn_without_headroom`, not just H3's video family, since the -suppression logic in `dw/result.py` is generic. - -## Remaining steps, if approved - -1. Hold the pre-encode `warn_without_headroom` result (`dw/result.py:77`) - rather than emitting it immediately via `emit_warning`, for - `content_type.startswith("video")`. -2. Run the post-encode `warn_if_written_above_full_scale` probe as today. -3. Drop the held warning if the file measures clean; upgrade it to - `audio_clipped` with the measured value if the file is genuinely over. -4. Leave audio-only saves unaffected (unconditional, unheld, as today). -5. Update the M-F008 regression case wording once this ships (it should then - pass `warnings: []` on `video-with-audio-768p` stock defaults for real, - not just as originally predicted by fix (1) alone). - - Fixes (1)-(4) implemented 2026-09-20 (commit bbe4adb); item 5 (updating - the M-F008 regression case wording) is still owed against the external - suite. - -## Fix (2) — remains explicitly deferred, no decision needed yet - -Whether to add a `normalize_audio` gain stage to the 15 H3 video templates. -Held per the original recommendation: no run on record has an H3 video mux -actually clipping (four consecutive measurements all undershoot by roughly -1 dB); wait for a real clipped-in-practice case to size the gain against, -the way #158 gave #159 one. - -## Evidence this decision rests on (undershoot record, all four points negative) - -| Run | Template / mux | Predicted (pre-encode) | Written (post-encode) | Delta | -|---|---|---|---|---| -| #161's `music-video` case | AAC, different song | — | — | +1.94 dB (the one positive measurement on record, a different template family) | -| `M-F012.jsonl` #1 | H3 video mux | target | measured | -0.92 dB | -| `M-F012.jsonl` #2 | H3 video mux | target | measured | -1.11 dB | -| M-F008 original report | `video-with-audio-768p` | -0.14 dBFS | -1.1158 dBFS | -0.98 dB | -| Tester's 2026-09-16 re-verify | `video-with-audio-768p` | -0.14 dBFS | -1.1158844… dBFS | -0.97 dB | diff --git a/docs/proposals/live-memory-during-run.md b/docs/proposals/live-memory-during-run.md deleted file mode 100644 index 4c034753..00000000 --- a/docs/proposals/live-memory-during-run.md +++ /dev/null @@ -1,96 +0,0 @@ -# Live `get_memory` during a run (#269, gap 2) - -Split from #269 on 2026-09-21. Gap 1 of that issue (a failed job's `progress` -going null instead of keeping the phase it died in) is fixed directly -(`dw/server/jobs.py`, `Job.progress()`) and does not need this decision. Gap 2 -does: `get_memory` refusing with `reason: "job_running"` for the whole -duration of a run, which is exactly when a caller wants it most - to tell a -merely slow run from one that is thrashing toward an OOM, in time to act -(`cancel_job`) rather than reading the answer in a post-mortem after #265/#266 -had already failed. - -## Why this isn't a mechanical fix - -`memory_status()` (`dw/server/jobs.py`) answers live only when the worker is -idle; while a job runs it returns `_cached_memory("job_running")` without -asking the worker anything. That refusal isn't a missing branch - it reflects -two real constraints: - -1. `_worker_lock` is held by the runner thread for a job's entire duration. - `memory_status()` cannot take it to send an ordinary command and block on - the reply without serializing behind the run it's trying to inspect. -2. `_consume_results()` is the sole reader of `result_queue` while a job is - in flight. It expects a fixed sequence of messages for *that* job - (`progress`, `error`/`success`, etc.) and has no mechanism today to - recognize and route an out-of-band reply to a concurrent caller. The - closest existing pattern, `probe_cache`'s `probe_id` correlation, exists - for exactly this kind of request/response-while-something-else-runs - problem - but `probe_cache` itself explicitly refuses when a job is - active (`self._current_job_id is not None`), so it's a model to extend, - not a mechanism already fit for this. - -Either fix means widening what can happen while the worker is mid-run: -new command handling in `dw/worker.py`'s `_watch_commands()` (already -mid-run-only, currently limited to `cancel`/`ping`/`shutdown`), and new -correlation state in `JobManager` to route a reply back to the right caller -instead of the runner thread. That's new engine surface, not a one-line -guard change - hence the escalation. - -## Two designs - -**(a) Proactive: emit `memory_info` at phase boundaries.** -The worker already emits one `memory_info` message, once, right after a run -finishes. Extend that to fire at each phase transition (`loading` -> -`generating` -> `decoding` -> `saving`) instead of only at the end. -`JobManager` folds each into the job's cached memory reading the same way it -already does for the post-run one (`_record_memory`), so `get_memory` while -a job is running is instantly answerable - it is not live in the sense of -"queried right now," but it's fresh as of the last phase boundary, which is -already enough resolution to catch "still climbing" vs. "flat" over the -handful of phases a run has. - -- Pro: no new correlation machinery, no new command type, no contention with - `_worker_lock` - it rides the same one-directional message flow that - progress events already use. - Con: resolution is coarse (phase boundaries, not on-demand); a phase that - runs long (a slow `generating` on a big denoise loop) still leaves a stale - reading for its whole duration, which is close to today's failure mode for - exactly the case #265/#266 cared about most. - -**(b) On-demand: a correlated request/response over the worker's command -queue, mid-run.** -Add a `memory_status` command that `_watch_commands()` answers even while a -job is active (it can - `_get_memory_info()` is pure stat reads and already -thread-safe against the run loop), tagged with a request id. `JobManager` -sends it without taking `_worker_lock` (a new, PID/queue-safe path, not -the same lock the runner holds), and `_consume_results()` recognizes a -tagged `memory_info` reply and hands it to the waiting caller instead of -folding it into the job's own event stream - the same shape as -`probe_cache`'s `probe_id`, extended to run concurrently with an active job -rather than refusing when one exists. - -- Pro: genuinely live, matches what `get_memory`'s docstring already - promises for the idle case. - Con: new correlation state in `JobManager` shared between the runner - thread and whatever thread services `get_memory` calls; needs care that a - slow or wedged worker (already mid-OOM) doesn't leave a `get_memory` call - hanging on a queue nobody is servicing - probably wants its own short - timeout, distinct from a job's. - -## Recommendation - -(a) first: it is the smaller change, reuses machinery this codebase already -trusts (the existing post-run `memory_info` message, `_record_memory`), and -directly serves the motivating case (distinguishing "slow" from "climbing") -without touching `_worker_lock` or `_consume_results`'s message routing. If -phase-boundary resolution turns out to be too coarse in practice - a single -long `generating` phase hiding a mid-phase spike - (b) is the fallback, and -(a)'s phase-boundary readings remain useful as a baseline even if (b) is -later added on top. - -Filed as its own issue against #269 rather than answered here because it -adds new mid-run worker protocol either way and the tradeoff above is a -product decision (how fresh does "live" need to be), not an implementation -detail. - -Model: sonnet. Provider: anthropic. diff --git a/docs/proposals/todo.md b/docs/proposals/todo.md index 69dd5a63..6f0a1eb0 100644 --- a/docs/proposals/todo.md +++ b/docs/proposals/todo.md @@ -9,13 +9,15 @@ benefit vs. added complexity/risk, highest ROI first. Updated 2026-09-20 original Tier 1 items and both fully-finished proposals (`score-and-select`, `script-to-video-agent-skill`) were removed. +Since 2026-09-23 every open item below is also a GitHub issue labeled +`feature` (#374–#380, and #244 for `resume.md`), parked with Don. Its +`priority:N` label mirrors the tier here. Work starts from the issue. + ## Tier 1 — do these first (small, scoped, clear payoff) -1. **h3-video-mux-headroom-warning-partial.md** — fixes (1)-(4) shipped - 2026-09-20 (`bbe4adb`); item 5 (updating the M-F008 regression case - wording against the external suite) is still owed, and fix (2) (a - `normalize_audio` gain stage on the H3 video templates) stays deferred - pending a real clipped-in-practice case. +None open. The last item, the H3 video mux headroom warning, shipped. Its +remaining deferred fix is recorded in +`complete/h3-video-mux-headroom-warning-complete.md`. ## Tier 2 — solid ROI, moderate scope From a9181050d3f937e3aca1eecdfbd021b29ad483ab Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 06:47:24 -0500 Subject: [PATCH 032/181] test: fix tests that checked nothing, drop ones that gated nothing A review of the suite for tests that are not quality gates. Fixed rather than removed: - test_download_watch: stall threshold was 2x the emit interval, so scheduler jitter tripped it (failed 4/5 runs); now ~8x like production - test_text_sections: H3 prompt glob matched nothing since the templates moved; repointed, plus a guard that the sweep is non-empty - catalog sweeps in test_elision, test_component_type_errors, test_task_signature_errors, test_pair_audio_fit: anchored to the repo root so they don't collect nothing when run from elsewhere - test_workflow_trust: set_trust_workflows(True) could not fail under conftest's DW_TRUST_WORKFLOWS=1 - keep-as-asset: link test asserted only inside `if linked`; now required, and the copy fallback gets its own test - several can't-fail or tautological tests rewritten to assert real behaviour (wait_for_job cap, REPL commands, gather, cached pipeline arguments/seed, watermark fallback, format.ts mtime, and others) Removed: test_worker_with_simple_workflow (strict subset of the cache-hit test, ~16s), assertion-free REPL scripts, dead __main__ blocks, and verified duplicates/trivial tests across ~25 files. Suite: 5720 passed, 14 skipped, 107s (was 5772/17, 150s). Co-Authored-By: Claude Opus 5.5 (1M context) --- tests/test_argument_updates.py | 298 +++++-------------------- tests/test_arguments.py | 11 +- tests/test_best_of_n_template.py | 10 - tests/test_borders.py | 6 - tests/test_component_type_errors.py | 38 ++-- tests/test_diffusion_upscale.py | 12 +- tests/test_download_watch.py | 12 +- tests/test_elision.py | 25 +-- tests/test_gather.py | 81 ++----- tests/test_list_images.py | 17 -- tests/test_mcp_diagnose.py | 17 +- tests/test_mcp_exports.py | 8 +- tests/test_mcp_media.py | 7 - tests/test_mcp_server.py | 10 - tests/test_model_cache.py | 13 -- tests/test_modular_progress.py | 20 +- tests/test_pair_audio_fit.py | 15 +- tests/test_pipeline_caching.py | 28 --- tests/test_plugin_skills.py | 7 +- tests/test_previous_results.py | 13 -- tests/test_realize.py | 12 - tests/test_repl_commands.py | 102 ++++----- tests/test_repl_hierarchical.py | 119 ---------- tests/test_repl_reorganization.py | 144 ------------ tests/test_result.py | 43 +--- tests/test_schema.py | 8 - tests/test_security.py | 28 +-- tests/test_select.py | 5 - tests/test_server.py | 46 ---- tests/test_server_workspaces.py | 40 +++- tests/test_step_cache.py | 11 - tests/test_strip_exif_and_watermark.py | 9 +- tests/test_task.py | 35 ++- tests/test_task_signature_errors.py | 10 +- tests/test_tensor_image.py | 5 - tests/test_text_sections.py | 8 +- tests/test_type_helpers.py | 6 - tests/test_video_utils.py | 5 - tests/test_worker.py | 68 ------ tests/test_workflow_trust.py | 3 + ui/src/lib/format.test.ts | 8 +- 41 files changed, 295 insertions(+), 1068 deletions(-) delete mode 100644 tests/test_repl_hierarchical.py delete mode 100644 tests/test_repl_reorganization.py diff --git a/tests/test_argument_updates.py b/tests/test_argument_updates.py index c6055914..2e9aedb2 100644 --- a/tests/test_argument_updates.py +++ b/tests/test_argument_updates.py @@ -1,250 +1,64 @@ -#!/usr/bin/env python3 -""" -Test to verify that cached pipelines get fresh arguments on each run. -This addresses the bug where changing arguments between runs didn't work. -""" +"""A cached pipeline reuses its loaded model but takes each run's own +arguments and seed. Guards the bug where changing arguments between runs of +a cached pipeline did nothing.""" -import os -import sys -import logging -from unittest.mock import patch, MagicMock +import copy +from unittest.mock import MagicMock, patch -# Add parent directory to path -sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) - -from dw.workflow import Workflow from dw.pipeline_processors.pipeline import Pipeline -from dw import get_device - -# Setup logging -logging.basicConfig( - level=logging.DEBUG, format="%(name)s - %(levelname)s - %(message)s" -) -logger = logging.getLogger(__name__) - - -def test_cached_pipeline_uses_new_arguments(): - """Test that cached pipelines receive fresh arguments on each run.""" - - # Create a workflow definition - workflow_def = { - "id": "test_args", - "steps": [ - { - "name": "generate", - "pipeline": { - "configuration": { - "component_type": "MockPipeline", - }, - "from_pretrained_arguments": {"model_name": "test-model"}, - "arguments": { - "prompt": "INITIAL_PROMPT", - "num_inference_steps": 10, - }, - }, - } - ], +from dw.workflow import Workflow + + +def pipeline_definition(**overrides): + definition = { + "configuration": {"component_type": "MockPipeline"}, + "from_pretrained_arguments": {"model_name": "test-model"}, + "arguments": {"prompt": "a cat", "num_inference_steps": 20}, } + definition.update(overrides) + return definition - workflow = Workflow(workflow_def, "/tmp/test_output", "test.json") - - # Track what arguments are passed to pipeline.run() - captured_arguments = [] - - original_pipeline_init = Pipeline.__init__ - - def mock_pipeline_init(self, *args, **kwargs): - original_pipeline_init(self, *args, **kwargs) - self.pipeline = MagicMock() - def mock_pipeline_load(self, *args, **kwargs): - logger.info("Pipeline.load() called") - self.pipeline = MagicMock() +def fake_load(self, shared_components): + self.pipeline = MagicMock() - def mock_pipeline_run(self, arguments, *args, **kwargs): - prompt = arguments.get("prompt", "NO_PROMPT") - steps = arguments.get("num_inference_steps", "NO_STEPS") - logger.info(f"🔵 Pipeline.run() called with: prompt='{prompt}', steps={steps}") - captured_arguments.append(arguments.copy()) - return MagicMock() - - with patch.object(Pipeline, "__init__", mock_pipeline_init): - with patch.object(Pipeline, "load", mock_pipeline_load): - with patch.object(Pipeline, "run", mock_pipeline_run): - pipeline_cache = {} - - # First run - should create and cache pipeline - logger.info("\n" + "=" * 60) - logger.info("RUN 1: Initial prompt") - logger.info("=" * 60) - - action1 = workflow.create_step_action( - workflow_def["steps"][0], {}, pipeline_cache, 42, get_device() - ) - - # Simulate step.run() calling action.run() - action1.run({"prompt": "a cat", "num_inference_steps": 20}, {}) - - logger.info("✅ Run 1 complete") - logger.info(f" Prompt passed: '{captured_arguments[-1]['prompt']}'") - logger.info( - f" Steps passed: {captured_arguments[-1]['num_inference_steps']}" - ) - - # Second run - should reuse cached model but with NEW arguments - logger.info("\n" + "=" * 60) - logger.info("RUN 2: Changed prompt (should use NEW prompt)") - logger.info("=" * 60) - - # Modify the workflow definition to simulate new arguments - workflow_def["steps"][0]["pipeline"]["arguments"]["prompt"] = "a dog" - workflow_def["steps"][0]["pipeline"]["arguments"][ - "num_inference_steps" - ] = 30 - - action2 = workflow.create_step_action( - workflow_def["steps"][0], {}, pipeline_cache, 42, get_device() - ) - - # Simulate step.run() calling action.run() - action2.run({"prompt": "a dog", "num_inference_steps": 30}, {}) - - logger.info("✅ Run 2 complete") - logger.info(f" Prompt passed: '{captured_arguments[-1]['prompt']}'") - logger.info( - f" Steps passed: {captured_arguments[-1]['num_inference_steps']}" - ) - - # Verify results - logger.info("\n" + "=" * 60) - logger.info("VERIFICATION") - logger.info("=" * 60) - - assert len(captured_arguments) == 2, ( - f"Expected 2 runs, got {len(captured_arguments)}" - ) - logger.info("✅ Both runs executed") - - run1_prompt = captured_arguments[0].get("prompt") - run2_prompt = captured_arguments[1].get("prompt") - - assert run1_prompt == "a cat", ( - f"Run 1 should have 'a cat', got '{run1_prompt}'" - ) - logger.info(f"✅ Run 1 used correct prompt: '{run1_prompt}'") - - assert run2_prompt == "a dog", ( - f"Run 2 should have 'a dog', got '{run2_prompt}'" - ) - logger.info(f"✅ Run 2 used NEW prompt: '{run2_prompt}'") - - assert run1_prompt != run2_prompt, ( - "Arguments should be different between runs!" - ) - logger.info("✅ Arguments changed between runs") - - run1_steps = captured_arguments[0].get("num_inference_steps") - run2_steps = captured_arguments[1].get("num_inference_steps") - - assert run1_steps == 20, f"Run 1 should have 20 steps, got {run1_steps}" - logger.info(f"✅ Run 1 used correct steps: {run1_steps}") - - assert run2_steps == 30, f"Run 2 should have 30 steps, got {run2_steps}" - logger.info(f"✅ Run 2 used NEW steps: {run2_steps}") - - logger.info("\n" + "=" * 60) - logger.info("🎉 TEST PASSED - Cached pipelines use fresh arguments!") - logger.info("=" * 60) - - -def test_generator_seed_updates(): - """Test that generator seed is updated on cached pipeline runs.""" - - workflow_def = { - "id": "test_seed", - "steps": [ - { - "name": "generate", - "pipeline": { - "configuration": { - "component_type": "MockPipeline", - }, - "from_pretrained_arguments": {}, - "arguments": {"prompt": "test"}, - "seed": 100, - }, - } - ], - } - workflow = Workflow(workflow_def, "/tmp/test_output", "test.json") - pipeline_cache = {} - - original_pipeline_init = Pipeline.__init__ - - def mock_pipeline_init(self, *args, **kwargs): - original_pipeline_init(self, *args, **kwargs) - self.pipeline = MagicMock() - - def mock_pipeline_load(self, *args, **kwargs): - self.pipeline = MagicMock() - - with patch.object(Pipeline, "__init__", mock_pipeline_init): - with patch.object(Pipeline, "load", mock_pipeline_load): - logger.info("\n" + "=" * 60) - logger.info("SEED TEST") - logger.info("=" * 60) - - # Run 1 with seed 100 - workflow.create_step_action( - workflow_def["steps"][0], {}, pipeline_cache, 42, get_device() - ) - seed1 = workflow_def["steps"][0]["pipeline"].get("seed", 42) - logger.info(f"Run 1: seed={seed1}") - - # Run 2 with seed 200 - workflow_def["steps"][0]["pipeline"]["seed"] = 200 - action2 = workflow.create_step_action( - workflow_def["steps"][0], {}, pipeline_cache, 42, get_device() - ) - seed2 = workflow_def["steps"][0]["pipeline"].get("seed", 42) - logger.info(f"Run 2: seed={seed2}") - - # Check that action2 has the new pipeline definition with seed 200 - assert action2.pipeline_definition["seed"] == 200, ( - f"Expected seed 200, got {action2.pipeline_definition.get('seed')}" - ) - - logger.info("✅ Pipeline wrapper gets updated seed") - - logger.info("\n" + "=" * 60) - logger.info("🎉 SEED TEST PASSED!") - logger.info("=" * 60) - - -if __name__ == "__main__": - print("\n" + "=" * 60) - print("Testing Cached Pipeline Argument Updates") - print("=" * 60 + "\n") - - try: - test_cached_pipeline_uses_new_arguments() - test_generator_seed_updates() - - print("\n" + "=" * 60) - print("✅ ALL TESTS PASSED!") - print("=" * 60) - print("\nFix verified:") - print(" • Cached pipelines receive fresh arguments on each run") - print(" • Generator seeds update correctly") - print(" • Arguments don't get stuck with old values") - - except AssertionError as e: - print(f"\n❌ TEST FAILED: {e}") - sys.exit(1) - except Exception as e: - print(f"\n❌ ERROR: {e}") - import traceback - - traceback.print_exc() - sys.exit(1) +def create_both(tmp_path, first, second): + """Create the step twice against one pipeline cache, as two runs do.""" + workflow = Workflow({"id": "test_args", "steps": []}, str(tmp_path), "test.json") + cache = {} + with patch.object(Pipeline, "load", autospec=True, side_effect=fake_load) as load: + action1 = workflow.create_step_action( + {"name": "generate", "pipeline": first}, {}, cache, 42, "cpu" + ) + action2 = workflow.create_step_action( + {"name": "generate", "pipeline": second}, {}, cache, 42, "cpu" + ) + return load, action1, action2 + + +def test_cached_pipeline_uses_new_arguments(tmp_path): + first = pipeline_definition() + second = copy.deepcopy(first) + second["arguments"] = {"prompt": "a dog", "num_inference_steps": 30} + + load, action1, action2 = create_both(tmp_path, first, second) + + # One load; the second run reuses the model under a fresh wrapper + assert load.call_count == 1 + assert action2 is not action1 + assert action2.pipeline is action1.pipeline + assert action2.argument_template["prompt"] == "a dog" + assert action2.argument_template["num_inference_steps"] == 30 + + +def test_generator_seed_updates(tmp_path): + first = pipeline_definition(seed=100) + second = copy.deepcopy(first) + second["seed"] = 200 + + load, action1, action2 = create_both(tmp_path, first, second) + + assert load.call_count == 1 + assert action2.pipeline is action1.pipeline + assert action2.argument_template["generator"].initial_seed() == 200 diff --git a/tests/test_arguments.py b/tests/test_arguments.py index 7f454c94..4a9b8b63 100644 --- a/tests/test_arguments.py +++ b/tests/test_arguments.py @@ -314,8 +314,11 @@ def test_realize_escaped_variable_substituted_into_type_key(self): assert steps["arguments"]["weights_dtype"] == "int4" def test_realize_escaped_offload_type_survives_second_pass(self): + # The previously mandatory {} escape keeps working after the key + # was excluded from type conversion, and stays unescaped on a rerealize args = {"group_offload": {"offload_type": "{leaf_level}"}} realize_args(args) + assert args["group_offload"]["offload_type"] == "leaf_level" realize_args(args) assert args["group_offload"]["offload_type"] == "leaf_level" @@ -334,14 +337,6 @@ def test_realize_offload_type_not_converted(self): assert args["group_offload"]["offload_type"] == "leaf_level" - def test_realize_escaped_offload_type_is_unescaped(self): - # The previously mandatory {} escape keeps working after the key - # was excluded from type conversion - args = {"group_offload": {"offload_type": "{leaf_level}"}} - realize_args(args) - - assert args["group_offload"]["offload_type"] == "leaf_level" - def test_explicit_media_reference_loads_under_any_key(self): # A "mask" argument names no media in its key - the explicit form # says what it is instead diff --git a/tests/test_best_of_n_template.py b/tests/test_best_of_n_template.py index 9d2023a7..52b928f3 100644 --- a/tests/test_best_of_n_template.py +++ b/tests/test_best_of_n_template.py @@ -54,13 +54,3 @@ def test_the_template_declares_no_cost_yet(): definition = load() assert "cost" not in definition - - -def test_saving_steps_are_marked_final_or_intermediate(): - steps = load()["steps"] - saving_steps = [s for s in steps if "result" in s] - - for step in saving_steps: - assert step["result"].get("subfolder") in ("final", "intermediate"), ( - f"step '{step['name']}' saves without a subfolder marking" - ) diff --git a/tests/test_borders.py b/tests/test_borders.py index 3bd19996..a0bed7e7 100644 --- a/tests/test_borders.py +++ b/tests/test_borders.py @@ -31,12 +31,6 @@ def test_returns_a_bordered_image_and_a_matching_mask(self, image): assert result["bordered_image"].mode == "RGB" assert result["mask"].mode == "L" - def test_padding_extends_only_the_requested_side(self, image): - result = add_border_and_mask(image, zoom_left=0.5) - - # 100 + 50 left pad = 150, snapped up to the nearest multiple of 32 - assert result["bordered_image"].size == (160, 96) - @pytest.mark.parametrize( "kwargs, expected", [ diff --git a/tests/test_component_type_errors.py b/tests/test_component_type_errors.py index 9a00741c..9de07505 100644 --- a/tests/test_component_type_errors.py +++ b/tests/test_component_type_errors.py @@ -18,6 +18,8 @@ from dw.introspection import component_type_errors +REPO_ROOT = pathlib.Path(__file__).resolve().parent.parent + def pipeline_step(component_type, name="a", extra=None): step = { @@ -144,20 +146,26 @@ def test_a_real_config_type_is_accepted(self): class TestNoDownloadIsQuotedForARefusedStep: - def test_component_type_plays_no_part_in_collecting_sources(self): + def test_the_refusal_reaches_validation_errors(self, tmp_path): """POST /api/validate builds `plan` (and its downloads_required) only when validation_errors is empty - see dw/server/app.py's validate - route - and component_type never feeds `_collect_sources` at all, so - a step that fails this check was never going to contribute a - download either way.""" - from dw.plan import _collect_sources + route - so a misspelled class on a step that names a checkpoint must + surface there, or the plan quotes a download for a step that cannot + run.""" + from dw.workflow import Workflow + + step = pipeline_step( + "FluxPipelin", + extra={ + "from_pretrained_arguments": {"model_name": "org/model"}, + "arguments": {"prompt": "a cat"}, + }, + ) + definition = {"id": "ct", "steps": [step]} + workflow = Workflow(definition, str(tmp_path), str(tmp_path / "ct.json")) - definition = {"id": "ct", "steps": [pipeline_step("FluxPipelin")]} - assert errors_for("FluxPipelin") != [] - names, urls = [], [] - _collect_sources(definition, names, urls) - assert names == [] - assert urls == [] + paths = [error["path"] for error in workflow.validation_errors()] + assert "steps[0].pipeline.configuration.component_type" in paths class TestTheCatalogItself: @@ -166,13 +174,13 @@ class TestTheCatalogItself: @pytest.mark.parametrize( "path", sorted( - str(p) - for p in list(pathlib.Path("workflows").rglob("*.json")) - + list(pathlib.Path("dw/workflows").glob("*.json")) + str(p.relative_to(REPO_ROOT)) + for p in list((REPO_ROOT / "workflows").rglob("*.json")) + + list((REPO_ROOT / "dw" / "workflows").glob("*.json")) ), ) def test_workflow_has_no_component_type_error(self, path): - definition = json.loads(pathlib.Path(path).read_text()) + definition = json.loads((REPO_ROOT / path).read_text()) if not isinstance(definition, dict) or "steps" not in definition: pytest.skip("not a workflow") assert component_type_errors(definition) == [] diff --git a/tests/test_diffusion_upscale.py b/tests/test_diffusion_upscale.py index 591c1896..fabc8c92 100644 --- a/tests/test_diffusion_upscale.py +++ b/tests/test_diffusion_upscale.py @@ -4,7 +4,7 @@ from unittest.mock import patch, MagicMock from PIL import Image -from dw.tasks.diffusion_upscale import diffusion_upscale, _MODELS +from dw.tasks.diffusion_upscale import diffusion_upscale class TestDiffusionUpscale(unittest.TestCase): @@ -160,16 +160,6 @@ def test_invalid_mode_raises(self): diffusion_upscale(self._make_image(), device="cpu", mode="x8") self.assertIn("x8", str(ctx.exception)) - def test_models_config_has_expected_modes(self): - self.assertIn("x4", _MODELS) - self.assertIn("x2", _MODELS) - self.assertEqual( - _MODELS["x4"]["pipeline_class"], "StableDiffusionUpscalePipeline" - ) - self.assertEqual( - _MODELS["x2"]["pipeline_class"], "StableDiffusionLatentUpscalePipeline" - ) - class TestDiffusionUpscaleRegistration(unittest.TestCase): """Test that diffusion_upscale is registered as a task command.""" diff --git a/tests/test_download_watch.py b/tests/test_download_watch.py index 401a8a06..66391453 100644 --- a/tests/test_download_watch.py +++ b/tests/test_download_watch.py @@ -11,10 +11,14 @@ from dw.events import RunContext +# Scaled down from production's 5s emit / 30s stall, keeping the margin +# between them wide enough that scheduler jitter cannot open a stall-sized +# gap between two progress events. Each test still runs longer than the +# threshold, so silence would trip the watchdog. def _fast_watchdog(): return patch.multiple( events_module, - PHASE_STALL_THRESHOLD_SECONDS=0.1, + PHASE_STALL_THRESHOLD_SECONDS=0.4, PHASE_STALL_CHECK_INTERVAL_SECONDS=0.02, ) @@ -40,10 +44,10 @@ def test_growing_download_emits_progress_and_suppresses_stall(tmp_path, monkeypa with download_watch.DownloadWatch( repo_id, context, cache_dir=str(tmp_path) ): - for _ in range(6): + for _ in range(20): with open(blob_file, "ab") as f: f.write(b"x" * 4096) - time.sleep(0.06) + time.sleep(0.04) finally: context.exit_run() @@ -73,7 +77,7 @@ def test_stalled_download_still_stalls(tmp_path, monkeypatch): with download_watch.DownloadWatch( repo_id, context, cache_dir=str(tmp_path) ): - time.sleep(0.3) + time.sleep(0.8) finally: context.exit_run() diff --git a/tests/test_elision.py b/tests/test_elision.py index 9c95a36a..18e196bf 100644 --- a/tests/test_elision.py +++ b/tests/test_elision.py @@ -18,6 +18,8 @@ from dw.elision import elide_definition, elide_unreferenced_steps from dw.workflow import Workflow +REPO_ROOT = pathlib.Path(__file__).resolve().parent.parent + def step(name, **extra): return {"name": name, **extra} @@ -53,16 +55,6 @@ def test_elision_is_transitive(self): assert names(kept) == ["kept"] assert [e["step"] for e in elided] == ["first", "second"] - def test_the_records_are_in_written_order(self): - _, elided = elide_unreferenced_steps( - [ - task("a"), - task("b", reads="a"), - task("last", result={"content_type": "audio/wav"}), - ] - ) - assert [e["step"] for e in elided] == ["a", "b"] - class TestTheGuardrails: def test_a_step_that_saves_is_kept(self): @@ -362,7 +354,7 @@ def test_an_orphan_still_gets_the_diagnosis(self): class TestTheMusicVideoSinger: """#146's happy path, end to end through the template itself.""" - PATH = "workflows/templates/minimax/music-video.json" + PATH = str(REPO_ROOT / "workflows/templates/minimax/music-video.json") def definition(self): return json.loads(pathlib.Path(self.PATH).read_text()) @@ -385,7 +377,7 @@ def test_a_supplied_portrait_elides_without_suggesting_a_mistake(self): class TestDialogueShort: """The case that raised it.""" - PATH = "workflows/templates/minimax/dialogue-short.json" + PATH = str(REPO_ROOT / "workflows/templates/minimax/dialogue-short.json") def definition(self): return json.loads(pathlib.Path(self.PATH).read_text()) @@ -439,9 +431,14 @@ class TestTheCatalogIsUnchanged: step to it on its stored defaults.""" @pytest.mark.parametrize( - "path", sorted(str(p) for p in pathlib.Path("workflows").rglob("*.json")) + "path", + sorted( + str(p.relative_to(REPO_ROOT)) + for p in (REPO_ROOT / "workflows").rglob("*.json") + ), ) def test_no_step_is_elided_on_the_defaults(self, path): + path = str(REPO_ROOT / path) definition = json.loads(pathlib.Path(path).read_text()) if not isinstance(definition, dict) or "steps" not in definition: pytest.skip("not a workflow") @@ -452,7 +449,7 @@ def test_no_step_is_elided_on_the_defaults(self, path): class TestMusicVideo: """#146: a standing cast member sings, without copying the template.""" - PATH = "workflows/templates/minimax/music-video.json" + PATH = str(REPO_ROOT / "workflows/templates/minimax/music-video.json") def definition(self): return json.loads(pathlib.Path(self.PATH).read_text()) diff --git a/tests/test_gather.py b/tests/test_gather.py index a084a630..5d169336 100644 --- a/tests/test_gather.py +++ b/tests/test_gather.py @@ -93,30 +93,26 @@ def test_gather_images_from_urls(self, mock_validate_url, mock_load_image): assert mock_validate_url.call_count == 2 assert mock_load_image.call_count == 2 - def test_gather_images_mixed_sources(self): - """Test gathering images from both files and URLs""" - with tempfile.TemporaryDirectory() as temp_dir: - # Create test image file - img1 = Image.new("RGB", (50, 50)) - path1 = os.path.join(temp_dir, "local.jpg") - img1.save(path1) - - glob_pattern = os.path.join(temp_dir, "*.jpg") + def test_gather_images_mixed_sources(self, tmp_path): + """Local matches and URLs are both gathered - the files first, then + the URLs in the order given.""" + Image.new("RGB", (50, 50)).save(tmp_path / "local.jpg") + remote = Image.new("RGB", (100, 100)) - with patch("dw.tasks.gather.load_image") as mock_load: - with patch("dw.tasks.gather.validate_media_url") as mock_validate: - mock_validate.return_value = "https://example.com/remote.jpg" - mock_load.side_effect = [ - Image.new("RGB", (50, 50)), # For file - Image.new("RGB", (100, 100)), # For URL - ] - - images = gather_images( - glob=glob_pattern, urls=["https://example.com/remote.jpg"] - ) + with ( + patch("dw.tasks.gather.load_image", return_value=remote) as mock_load, + patch( + "dw.tasks.gather.validate_media_url", + side_effect=lambda url, what=None: url, + ), + ): + images = gather_images( + glob=os.path.join(str(tmp_path), "*.jpg"), + urls=["https://example.com/remote.jpg"], + ) - # Should have images from both sources - assert len(images) >= 1 + assert [img.size for img in images] == [(50, 50), (100, 100)] + mock_load.assert_called_once_with("https://example.com/remote.jpg") def test_gather_images_no_results_raises_error(self): """Test that gathering no images raises ValueError""" @@ -158,19 +154,6 @@ def test_gather_images_traversal_glob_raises(self, tmp_path): with pytest.raises(SecurityError): gather_images(glob=traversal_pattern) - def test_gather_images_none_defaults(self): - """Test that None URLs parameter works (fixed mutable default)""" - # This tests the fix for mutable default arguments - with tempfile.TemporaryDirectory() as temp_dir: - img = Image.new("RGB", (50, 50)) - path = os.path.join(temp_dir, "test.jpg") - img.save(path) - - glob_pattern = os.path.join(temp_dir, "*.jpg") - images = gather_images(glob=glob_pattern) # urls=None - - assert len(images) == 1 - def test_gather_images_happy_path_tmp_path(self, tmp_path): """Allowed-extension local files under a glob still load fine after routing through fetch_image's validation.""" @@ -233,16 +216,6 @@ def test_gather_videos_no_results_raises_error(self): assert "No videos found" in str(exc_info.value) - def test_gather_videos_none_defaults(self): - """Test that None URLs parameter works""" - with tempfile.TemporaryDirectory() as temp_dir: - # Since we can't easily create real video files in tests, - # we'll just test that the function handles None properly - with pytest.raises(ValueError) as exc_info: - gather_videos(glob=os.path.join(temp_dir, "*.mp4")) - - assert "No videos found" in str(exc_info.value) - def test_gather_videos_sorted_not_filesystem_order(self, tmp_path): """Videos gathered for a concat come back in sorted order. @@ -305,21 +278,3 @@ def test_gather_inputs_returns_unchanged(self): result = gather_inputs(inputs) assert result == inputs - - def test_gather_inputs_with_list(self): - """Test gather_inputs with list input""" - inputs = ["item1", "item2", "item3"] - result = gather_inputs(inputs) - - assert result == inputs - - def test_gather_inputs_with_nested_structure(self): - """Test gather_inputs with nested structures""" - inputs = {"outer": {"inner": ["value1", "value2"]}} - result = gather_inputs(inputs) - - assert result == inputs - - -if __name__ == "__main__": - pytest.main([__file__, "-v"]) diff --git a/tests/test_list_images.py b/tests/test_list_images.py index 2dc9ac6d..68ec6396 100644 --- a/tests/test_list_images.py +++ b/tests/test_list_images.py @@ -59,20 +59,3 @@ def test_fetch_video_list(): assert isinstance(result, list), "Result should be a list" assert len(result) == 2, "Should have 2 frames" print("✓ List of already loaded frames works") - - -if __name__ == "__main__": - try: - test_fetch_image_list() - test_fetch_video_list() - print("\n✅ All tests passed!") - sys.exit(0) - except AssertionError as e: - print(f"\n❌ Test failed: {e}") - sys.exit(1) - except Exception as e: - print(f"\n❌ Error: {e}") - import traceback - - traceback.print_exc() - sys.exit(1) diff --git a/tests/test_mcp_diagnose.py b/tests/test_mcp_diagnose.py index 98e2474a..99293206 100644 --- a/tests/test_mcp_diagnose.py +++ b/tests/test_mcp_diagnose.py @@ -310,16 +310,23 @@ def test_wait_for_job_reports_still_running_at_timeout(monkeypatch): assert elapsed < 1, "must return once timeout_seconds elapses, not hang" -def test_wait_for_job_caps_the_timeout_it_is_given(): +def test_wait_for_job_caps_the_timeout_it_is_given(monkeypatch): """A caller asking for an absurd timeout does not get an absurd wait - - the value is clamped before it ever reaches the poll loop.""" + the value is clamped before it ever reaches the poll loop. The job never + finishes, so only the cap can end the call.""" + monkeypatch.setattr(diagnose, "WAIT_POLL_SECONDS", 0.01) + monkeypatch.setattr(diagnose, "MAX_WAIT_SECONDS", 0.05) client, seen = sequenced( - ("GET", "/api/jobs/job-1"), [{"id": "job-1", "status": "succeeded"}] + ("GET", "/api/jobs/job-1"), [{"id": "job-1", "status": "running"}] ) - diagnose.wait_for_job(client, "job-1", timeout_seconds=10_000) + started = time.monotonic() + result = diagnose.wait_for_job(client, "job-1", timeout_seconds=10_000) + elapsed = time.monotonic() - started - assert len(seen) == 1, "a terminal status on the first poll returns immediately" + assert result["still_running"] is True + assert elapsed < 2, "the cap, not the requested 10,000s, bounds the wait" + assert len(seen) >= 2, "it still polls inside the capped budget" def test_wait_for_job_says_when_it_capped_the_timeout(monkeypatch): diff --git a/tests/test_mcp_exports.py b/tests/test_mcp_exports.py index d509e448..bea5242c 100644 --- a/tests/test_mcp_exports.py +++ b/tests/test_mcp_exports.py @@ -99,11 +99,9 @@ def test_a_409_reaches_the_model_as_a_readable_refusal(): assert "already exists" in str(caught.value) -def test_the_docstring_says_copying_costs_disk_and_names_total_bytes(): - # A caller reading only the handler's docstring has to learn this before - # exporting a video job fills the server's disk a second time - the - # export copies files rather than linking them. - assert "copies every output and input" in exports.export_job.__doc__ +def test_the_docstring_names_total_bytes(): + # The field a caller reads to see what an export cost on disk has to be + # named where the caller looks - the handler's docstring. assert "total_bytes" in exports.export_job.__doc__ diff --git a/tests/test_mcp_media.py b/tests/test_mcp_media.py index 63ce0324..0e6d1a1b 100644 --- a/tests/test_mcp_media.py +++ b/tests/test_mcp_media.py @@ -333,13 +333,6 @@ def handler(request): assert stream.iterated is False -def test_an_undecodable_body_is_refused_clearly(): - client = serving(b"not an image at all", "image/png") - - with pytest.raises(DwApiError, match="could not be decoded"): - get_output_image(client, "broken.png") - - def test_a_missing_file_propagates_the_api_error(): def handler(request): return httpx.Response(404, json={"detail": "Unknown file"}) diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index ce0a5eaa..7c61660f 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -278,16 +278,6 @@ async def test_a_read_only_tool_round_trips_to_the_api(): assert "workflows" in json.dumps(_text_of(result)) -@pytest.mark.asyncio -async def test_run_workflow_refuses_without_acknowledgement(): - server = server_over(ok({"id": "job-1", "status": "queued"})) - - with pytest.raises(Exception) as caught: - await server.call_tool("run_workflow", {"workflow_path": "w.json"}) - - assert "acknowledged_cost" in str(caught.value) - - @pytest.mark.asyncio async def test_an_unreachable_server_reports_how_to_start_it(): def refusing(request): diff --git a/tests/test_model_cache.py b/tests/test_model_cache.py index 4d79fac4..b16627e9 100644 --- a/tests/test_model_cache.py +++ b/tests/test_model_cache.py @@ -35,19 +35,6 @@ def test_distinct_keys_get_distinct_loads(self): factory_a.assert_called_once() factory_b.assert_called_once() - def test_device_is_part_of_the_key(self): - """Same model name on two different devices must load independently.""" - factory_cpu = MagicMock(return_value="on-cpu") - factory_cuda = MagicMock(return_value="on-cuda") - - result_cpu = cached_model(("task", "model-a", "cpu"), factory_cpu) - result_cuda = cached_model(("task", "model-a", "cuda"), factory_cuda) - - assert result_cpu == "on-cpu" - assert result_cuda == "on-cuda" - factory_cpu.assert_called_once() - factory_cuda.assert_called_once() - def test_returns_the_same_object_instance(self): """Callers get the identical cached object back, not a copy.""" loaded = object() diff --git a/tests/test_modular_progress.py b/tests/test_modular_progress.py index f60dd949..2a7e9d5c 100644 --- a/tests/test_modular_progress.py +++ b/tests/test_modular_progress.py @@ -274,9 +274,25 @@ def test_the_block_dispatch_is_handed_back_unpatched(): assert FakeSequentialBlocks.__call__ is original +class FakePipelineWithoutBlocks: + """No step callback and no `_blocks` tree - the bar is the pipeline's + own, the way a classic pipeline that predates the callback holds it.""" + + def progress_bar(self, iterable=None, total=None): + return tqdm(total=total, disable=True) + + def __call__(self, prompt=None, num_inference_steps=None, generator=None): + with self.progress_bar(total=3) as bar: + for _ in range(3): + bar.update() + return FakeOutput() + + def test_a_pipeline_with_no_blocks_still_runs(): - """Not every pipeline without a step callback is modular.""" - assert _steps(_events(FakeModularPipeline())) == [(1, 3), (2, 3), (3, 3)] + """Not every pipeline without a step callback is modular: with no + `_blocks` there is nothing to narrate, and its own bar still reports.""" + assert not hasattr(FakePipelineWithoutBlocks(), "_blocks") + assert _steps(_events(FakePipelineWithoutBlocks())) == [(1, 3), (2, 3), (3, 3)] class FakeConditionalBlocks: diff --git a/tests/test_pair_audio_fit.py b/tests/test_pair_audio_fit.py index be933198..d24a782b 100644 --- a/tests/test_pair_audio_fit.py +++ b/tests/test_pair_audio_fit.py @@ -17,6 +17,8 @@ from dw.result import AudioVideo from dw.tasks.pair_audio import pair_audio +REPO_ROOT = pathlib.Path(__file__).resolve().parent.parent + SAMPLE_RATE = 44100 FPS = 24 @@ -53,11 +55,6 @@ def record(event): class TestFitToTheVideo: - def test_the_reported_case_is_cut_to_the_two_shot_edit(self): - """248 frames at 24 fps is 10.33 s, from a 30 s song.""" - result = pair_audio(cut(248), song(30), sample_rate=SAMPLE_RATE, fit="video") - assert samples(result) == round(248 / FPS * SAMPLE_RATE) - def test_the_default_four_shot_length_is_unchanged(self): """496 frames - what the hardcoded slice used to produce.""" result = pair_audio(cut(496), song(30), sample_rate=SAMPLE_RATE, fit="video") @@ -72,7 +69,9 @@ def test_the_other_direction_pads_and_warns(self, warnings_emitted): def test_trimming_the_track_warns_with_the_seconds_cut(self, warnings_emitted): """The pad direction cannot lose content; the trim direction always - can, so it is the one that most needs saying out loud (#246).""" + can, so it is the one that most needs saying out loud (#246). This is + also #142's reported case: 248 frames at 24 fps is 10.33 s, from a + 30 s song.""" result = pair_audio(cut(248), song(30), sample_rate=SAMPLE_RATE, fit="video") assert samples(result) == round(248 / FPS * SAMPLE_RATE) trimmed = [w for w in warnings_emitted if "trimmed" in w] @@ -115,7 +114,7 @@ def test_unknown_fps_is_not_a_mismatch(self, warnings_emitted): class TestTheTemplateItself: def test_music_video_derives_its_soundtrack(self): definition = json.loads( - pathlib.Path("workflows/templates/minimax/music-video.json").read_text() + (REPO_ROOT / "workflows/templates/minimax/music-video.json").read_text() ) steps = {s["name"]: s for s in definition["steps"]} assert "soundtrack" not in steps, "the hardcoded 496-frame slice is gone" @@ -131,7 +130,7 @@ def test_no_template_hardcodes_a_soundtrack_length(self): """Every `slice_audio` in the catalog whose count is a literal is one a list cannot resize out from under - so the literal must not be the length of a whole cut.""" - for path in pathlib.Path("workflows").rglob("*.json"): + for path in (REPO_ROOT / "workflows").rglob("*.json"): definition = json.loads(path.read_text()) if not isinstance(definition, dict): continue diff --git a/tests/test_pipeline_caching.py b/tests/test_pipeline_caching.py index 3170d4ab..dae119d1 100644 --- a/tests/test_pipeline_caching.py +++ b/tests/test_pipeline_caching.py @@ -362,34 +362,6 @@ def mock_pipeline_load(self, shared_components): clear_model_cache() -if __name__ == "__main__": - print("\n" + "=" * 60) - print("Testing Pipeline Caching Implementation") - print("=" * 60 + "\n") - - try: - test_pipeline_caching() - test_pipeline_caching_different_steps() - - print("\n" + "=" * 60) - print("✅ ALL TESTS PASSED!") - print("=" * 60) - print("\nModels will now persist in GPU memory across workflow runs!") - print( - "This significantly improves performance by avoiding repeated model loading." - ) - - except AssertionError as e: - print(f"\n❌ TEST FAILED: {e}") - sys.exit(1) - except Exception as e: - print(f"\n❌ ERROR: {e}") - import traceback - - traceback.print_exc() - sys.exit(1) - - def test_cache_hit_republishes_shared_components(): """A warm sharing step must refill the fresh shared_components dict, or a later reusing step that missed the cache finds nothing.""" diff --git a/tests/test_plugin_skills.py b/tests/test_plugin_skills.py index f6cc497d..96c5a920 100644 --- a/tests/test_plugin_skills.py +++ b/tests/test_plugin_skills.py @@ -133,7 +133,6 @@ def test_the_frame_rule_and_bounds_are_the_pipeline_s(self): assert modular_pipeline.MINIMAX_H3_FPS == 24 and "24 fps" in text # 124 and 345 are the smallest and largest 17n + 5 inside 5 to 15 seconds at 24 fps assert "124" in text and "345" in text - assert 124 == 17 * 7 + 5 and 345 == 17 * 20 + 5 # min_duration/max_duration are instance properties on MiniMaxH3ModularPipeline, so # the window is pinned by the check that reads them rather than by their values assert ( @@ -184,10 +183,7 @@ def test_the_canvas_rules_are_the_pipeline_s(self): def test_the_reference_and_audio_limits_are_the_pipeline_s(self): import inspect - from diffusers.modular_pipelines.minimax_h3 import ( - before_encoder, - modular_pipeline, - ) + from diffusers.modular_pipelines.minimax_h3 import modular_pipeline from diffusers.modular_pipelines.minimax_h3.before_encoder import ( MiniMaxH3Ref2VASetupStep, ) @@ -204,7 +200,6 @@ def test_the_reference_and_audio_limits_are_the_pipeline_s(self): modular_pipeline.MINIMAX_H3_AUDIO_CHANNELS == 2 and "32 kHz stereo" in text ) assert "audio can" in text and "never be the only reference" in text - assert before_encoder is not None def test_the_checkpoint_coupling_is_stated(self): """A checkpoint comes with its canvas, its shift and its alpha. diff --git a/tests/test_previous_results.py b/tests/test_previous_results.py index 1f6fe6f9..d8a29b7e 100644 --- a/tests/test_previous_results.py +++ b/tests/test_previous_results.py @@ -422,19 +422,6 @@ def test_the_template_is_left_intact_for_the_next_iteration(self): assert template["references"][0] is description assert description["from_previous_result"] == "draw" - def test_siblings_are_shared_rather_than_copied(self): - result = Result({}) - result.add_result(["first", "second"]) - - frames = [object()] - template = {"video": frames, "prompt": "previous_result:write"} - iterations = get_iterations(template, {"write": result}) - - # Only the containers on the path to the substitution are copied - a - # deep copy would duplicate the media the iterations mean to share - assert iterations[0]["video"] is frames - assert iterations[1]["video"] is frames - if __name__ == "__main__": pytest.main([__file__, "-v"]) diff --git a/tests/test_realize.py b/tests/test_realize.py index 656399b6..993055a8 100644 --- a/tests/test_realize.py +++ b/tests/test_realize.py @@ -280,18 +280,6 @@ def test_pin_outputs_false_still_inlines_prompts(self, prompt_library): assert realized["variables"]["prompt"] == "a harbour at dusk" assert annotations["prompts"] == ["scenic/dusk"] - def test_the_default_still_pins(self, output_root): - root, run_id = output_root - spec = definition() - spec["steps"][0]["pipeline"]["arguments"]["image"] = ( - "output:ltx2/Gyre/latest/still.png" - ) - realized, _ = realize_workflow(spec, {}, 7, output_root=root) - assert ( - realized["steps"][0]["pipeline"]["arguments"]["image"] - == f"output:ltx2/Gyre/{run_id}/still.png" - ) - class TestReadSubWorkflow: def test_reads_a_child_beside_the_parent(self, tmp_path): diff --git a/tests/test_repl_commands.py b/tests/test_repl_commands.py index 429b6038..ade4dd3b 100644 --- a/tests/test_repl_commands.py +++ b/tests/test_repl_commands.py @@ -1,60 +1,56 @@ -#!/usr/bin/env python -""" -Test the reorganized REPL commands interactively. -""" +"""The REPL's command groups: help routing, and what each group answers +before any workflow is loaded or any worker has started.""" -import sys -import os - -# Add parent to path -sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) +import pytest from dw.repl import DiffusersWorkflowREPL -def test_repl_commands(): - """Test the reorganized command structure""" +def output_of(repl, capsys, line): + capsys.readouterr() + repl.onecmd(line) + return capsys.readouterr().out + + +def test_help_lists_every_command_group(capsys): + out = output_of(DiffusersWorkflowREPL(), capsys, "help") + for group in DiffusersWorkflowREPL.COMMAND_GROUPS: + assert f" {group}" in out + + +@pytest.mark.parametrize("group", DiffusersWorkflowREPL.COMMAND_GROUPS) +def test_group_help_and_help_group_tell_the_same_story(capsys, group): + repl = DiffusersWorkflowREPL() + via_group = output_of(repl, capsys, f"{group} ?") + via_help = output_of(repl, capsys, f"help {group}") + assert f"{group.capitalize()} commands:" in via_group + assert via_help == via_group + + +@pytest.mark.parametrize( + "line, expected", + [ + ("workflow status", "No workflow currently loaded"), + ("arg show", "No workflow loaded"), + ("memory show", "No worker process running"), + ("memory clear", "No worker process running"), + ("workflow bogus", "Unknown workflow subcommand: bogus"), + ("memory bogus", "Unknown memory subcommand: bogus"), + ("config bogus", "Unknown config subcommand: bogus"), + ], +) +def test_commands_answer_before_a_workflow_or_worker_exists(capsys, line, expected): + assert expected in output_of(DiffusersWorkflowREPL(), capsys, line) + + +def test_config_show_lists_the_session_settings(capsys): repl = DiffusersWorkflowREPL() + out = output_of(repl, capsys, "config show") + for name, value in repl.globals.items(): + assert f" {name}={value}" in out + - test_commands = [ - ("help", "Main help"), - ("workflow ?", "Workflow help"), - ("arg ?", "Arg help"), - ("model ?", "Model help"), - ("memory ?", "Memory help"), - ("config ?", "Config help"), - ("config show", "Show config"), - ("workflow status", "Workflow status (no workflow loaded)"), - ("arg show", "Show args (no workflow loaded)"), - # Test backward compatibility - ("status", "Old status command"), - ("load", "Old load command (no args)"), - ] - - print("=" * 70) - print("Testing REPL Command Reorganization") - print("=" * 70) - - for cmd, description in test_commands: - print(f"\n{'=' * 70}") - print(f"Test: {description}") - print(f"Command: {cmd}") - print(f"{'=' * 70}") - repl.onecmd(cmd) - - print("\n" + "=" * 70) - print("✅ All command tests completed successfully!") - print("=" * 70) - print("\nCommand hierarchy implemented:") - print(" • workflow - Load and manage workflows") - print(" • arg - Manage workflow arguments") - print(" • model - Control model execution") - print(" • memory - Monitor and manage GPU memory") - print(" • config - Configure global settings") - print("\nBackward compatibility maintained for:") - print(" load, reload, status, run, restart, clear, set, clear_args") - print("\nUse ' ?' to explore any command group!") - - -if __name__ == "__main__": - test_repl_commands() +def test_unknown_command_suggests_the_close_match(capsys): + out = output_of(DiffusersWorkflowREPL(), capsys, "worklfow status") + assert "Unknown command: worklfow status" in out + assert "Did you mean: workflow?" in out diff --git a/tests/test_repl_hierarchical.py b/tests/test_repl_hierarchical.py deleted file mode 100644 index e8c5a7cb..00000000 --- a/tests/test_repl_hierarchical.py +++ /dev/null @@ -1,119 +0,0 @@ -#!/usr/bin/env python -""" -Test the hierarchical REPL command structure. -Verifies all commands work correctly without backward compatibility. -""" - -import sys -import os - -sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) - -from dw.repl import DiffusersWorkflowREPL - - -def test_hierarchical_commands(): - """Test all hierarchical command groups""" - repl = DiffusersWorkflowREPL() - - print("=" * 70) - print("Testing Hierarchical REPL Commands") - print("=" * 70) - - # Test help system - print("\n▶ Testing help system") - repl.onecmd("?") - repl.onecmd("help") - - # Test each command group's help - groups = ["workflow", "arg", "model", "memory", "config"] - for group in groups: - print(f"\n▶ Testing {group} ?") - repl.onecmd(f"{group} ?") - - # Test workflow commands - print("\n▶ Testing workflow commands") - repl.onecmd("workflow status") - repl.onecmd("workflow load FluxDev") - repl.onecmd("workflow status") - - # Test arg commands - print("\n▶ Testing arg commands") - repl.onecmd("arg show") - repl.onecmd('arg set prompt="test"') - repl.onecmd("arg show") - repl.onecmd("arg clear") - repl.onecmd("arg show") - - # Test config commands - print("\n▶ Testing config commands") - repl.onecmd("config show") - repl.onecmd("config set output_dir=./outputs") - repl.onecmd("config show") - - # Test memory commands - print("\n▶ Testing memory commands") - repl.onecmd("memory show") - - print("\n" + "=" * 70) - print("✅ All hierarchical commands tested successfully!") - print("=" * 70) - - -def test_command_structure(): - """Verify command structure""" - print("\n" + "=" * 70) - print("Command Structure Verification") - print("=" * 70) - - expected_structure = { - "workflow": ["load", "reload", "status"], - "arg": ["show", "set", "clear"], - "model": ["run", "restart"], - "memory": ["show", "clear"], - "config": ["show", "set"], - } - - print("\nExpected command groups and subcommands:") - for group, subcommands in expected_structure.items(): - print(f"\n {group}:") - for subcmd in subcommands: - print(f" - {subcmd}") - - print("\n✅ Command structure is clean and hierarchical") - print("✅ No backward compatibility aliases") - print("✅ All commands use '?' for help") - - # Assertions for pytest - assert expected_structure is not None - assert len(expected_structure) > 0 - - -def main(): - try: - test_hierarchical_commands() - test_command_structure() - - print("\n" + "=" * 70) - print("🎉 ALL TESTS PASSED!") - print("=" * 70) - print("\nREPL Command Structure:") - print(" • workflow - Load and manage workflows") - print(" • arg - Manage workflow arguments") - print(" • model - Control model execution") - print(" • memory - Monitor GPU memory") - print(" • config - Configure settings") - print("\nUsage: [args]") - print("Discovery: Use '?' with any command to see subcommands") - - return 0 - except Exception as e: - print(f"\n❌ TEST FAILED: {e}") - import traceback - - traceback.print_exc() - return 1 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/tests/test_repl_reorganization.py b/tests/test_repl_reorganization.py deleted file mode 100644 index f50408da..00000000 --- a/tests/test_repl_reorganization.py +++ /dev/null @@ -1,144 +0,0 @@ -#!/usr/bin/env python -""" -Comprehensive test of the reorganized REPL command structure. -Tests both new hierarchical commands and backward compatibility. -""" - -import sys -import os - -sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) - -from dw.repl import DiffusersWorkflowREPL - - -def test_command_equivalence(): - """Test that old and new commands produce the same results""" - - print("=" * 70) - print("Testing Command Equivalence (Old vs New)") - print("=" * 70) - - tests = [ - { - "old": "status", - "new": "workflow status", - "description": "Check workflow status", - }, - {"old": "set", "new": "config set", "description": "Config set (no args)"}, - {"old": "clear_args", "new": "arg clear", "description": "Clear arguments"}, - ] - - for test in tests: - repl = DiffusersWorkflowREPL() # Fresh REPL for each test - - print(f"\n{'-' * 70}") - print(f"Test: {test['description']}") - print(f"Old command: {test['old']}") - print(f"New command: {test['new']}") - print(f"{'-' * 70}") - - # The output should be the same - print("Old command output:") - repl.onecmd(test["old"]) - - print("\nNew command output:") - repl.onecmd(test["new"]) - - print("✅ Both commands work") - - print("\n" + "=" * 70) - print("✅ All equivalence tests passed!") - print("=" * 70) - - -def test_help_system(): - """Test the hierarchical help system""" - - print("\n" + "=" * 70) - print("Testing Help System") - print("=" * 70) - - repl = DiffusersWorkflowREPL() - - commands = ["workflow", "arg", "model", "memory", "config"] - - for cmd in commands: - print(f"\n{'-' * 70}") - print(f"Testing: {cmd} ?") - print(f"{'-' * 70}") - repl.onecmd(f"{cmd} ?") - print(f"✅ Help for '{cmd}' works") - - print("\n" + "=" * 70) - print("✅ All help tests passed!") - print("=" * 70) - - -def test_command_flow(): - """Test a typical command flow""" - - print("\n" + "=" * 70) - print("Testing Typical Command Flow") - print("=" * 70) - - repl = DiffusersWorkflowREPL() - - flow = [ - ("config show", "Show configuration"), - ("workflow status", "Check workflow (none loaded)"), - ("arg show", "Try to show args (no workflow)"), - ("config set output_dir=./outputs", "Set output directory"), - ("memory show", "Try to show memory (no worker)"), - ] - - for cmd, description in flow: - print(f"\n{'-' * 70}") - print(f"Step: {description}") - print(f"Command: {cmd}") - print(f"{'-' * 70}") - repl.onecmd(cmd) - print(f"✅ {description} - OK") - - print("\n" + "=" * 70) - print("✅ Command flow test passed!") - print("=" * 70) - - -def main(): - """Run all tests""" - try: - test_command_equivalence() - test_help_system() - test_command_flow() - - print("\n" + "=" * 70) - print("🎉 ALL TESTS PASSED!") - print("=" * 70) - print("\nREPL Command Reorganization Summary:") - print(" ✅ Hierarchical command structure implemented") - print(" ✅ Help system with '?' support") - print(" ✅ Backward compatibility maintained") - print(" ✅ All command groups working") - print("\nCommand Groups:") - print(" • workflow - Load and manage workflows") - print(" • arg - Manage workflow arguments") - print(" • model - Control model execution") - print(" • memory - Monitor GPU memory") - print(" • config - Configure settings") - print("\nDocumentation:") - print(" • docs/REPL_COMMANDS.md - Full command reference") - print(" • docs/REPL_WORKER_GUIDE.md - Architecture guide") - - return 0 - - except Exception as e: - print(f"\n❌ TEST FAILED: {e}") - import traceback - - traceback.print_exc() - return 1 - - -if __name__ == "__main__": - sys.exit(main()) diff --git a/tests/test_result.py b/tests/test_result.py index 973b30e6..cfd349ae 100644 --- a/tests/test_result.py +++ b/tests/test_result.py @@ -371,16 +371,6 @@ class MockResult: # channels-first, the layout AudioTrack documents assert artifacts[0].audio.shape == (2, 100) - def test_audios_without_a_rate_stay_bare_waveforms(self): - # Nothing to carry: the shape every existing consumer already handles - class MockResult: - audios = numpy.zeros((1, 2, 100), dtype=numpy.float32) - - artifacts = get_artifact_list(MockResult()) - - assert isinstance(artifacts[0], numpy.ndarray) - assert artifacts[0].shape == (100, 2) - def test_images_take_precedence_over_frames(self): # The output-field registry is consulted in order - a result exposing both # (which nothing real does, but the registry order must still be deterministic) @@ -1141,15 +1131,6 @@ def test_no_metadata_when_not_enabled(self): saved_img = Image.open(output_file) assert "parameters" not in saved_img.info - def test_set_metadata_method(self): - """set_metadata should store metadata on the Result instance.""" - result = Result({}) - assert result.metadata is None - - metadata = {"workflow_id": "test", "step_name": "step1"} - result.set_metadata(metadata) - assert result.metadata == metadata - def test_metadata_with_embed_false(self): """When embed_metadata is explicitly false, no metadata embedded even if set.""" with tempfile.TemporaryDirectory() as temp_dir: @@ -1624,22 +1605,6 @@ def test_the_muxed_deliverable_is_measured_too(self): def test_a_video_with_a_quiet_track_is_not_warned_about(self): assert self.events_from(lambda: self.save_muxed(torch.zeros((2, 100)))) == [] - def test_a_video_is_probed_even_though_the_waveform_already_warned(self): - """#174: suppressing the post-encode probe whenever the pre-encode - check already fired assumed the encoder only ever adds overshoot - - true for the mp3s #159/#161 measured, backwards for an H3 video mux, - whose AAC mux can land under full scale after starting over it. A - video always gets the ground-truth post-encode read, and once that - read is in, it - not the pre-encode guess - is what the caller sees.""" - with patch( - "dw.media_info.probe_media", - return_value={"peak_dbfs": 0.94, "kind": "video"}, - ): - warnings = self.events_from(lambda: self.save_muxed(torch.ones((2, 100)))) - - kinds = {w["kind"] for w in warnings} - assert kinds == {"audio_clipped"} - def test_a_clean_video_mux_drops_the_stale_prediction(self): """#174 amendment: the pre-encode prediction fires on H3's own soundtrack every run, and the post-encode probe already proved the @@ -1669,7 +1634,13 @@ def test_a_video_mux_is_decoded_only_once_for_both_level_checks(self): def test_a_dirty_video_mux_reports_only_the_measured_clip(self): """The post-encode probe found a real clip - report that, not the - pre-encode guess, so the caller gets one answer with a real number.""" + pre-encode guess, so the caller gets one answer with a real number. + + #174: suppressing the post-encode probe whenever the pre-encode check + already fired assumed the encoder only ever adds overshoot - true for + the mp3s #159/#161 measured, backwards for an H3 video mux, whose AAC + mux can land under full scale after starting over it. A video always + gets the ground-truth post-encode read.""" with patch( "dw.media_info.probe_media", return_value={"peak_dbfs": 0.94, "kind": "video"}, diff --git a/tests/test_schema.py b/tests/test_schema.py index ff0c3967..9337482b 100644 --- a/tests/test_schema.py +++ b/tests/test_schema.py @@ -11,14 +11,6 @@ ) -def test_load_schema(): - # Test that we can load the workflow schema - schema = load_schema("workflow") - assert schema is not None - assert "$schema" in schema - assert "properties" in schema - - def test_validate_data_valid(valid_workflow_json): # Test validation with valid workflow schema = load_schema("workflow") diff --git a/tests/test_security.py b/tests/test_security.py index ef1c923d..8952c01e 100644 --- a/tests/test_security.py +++ b/tests/test_security.py @@ -120,29 +120,6 @@ def test_string_input_validation(): validate_string_input("hello\x01world") -def test_command_sanitization(): - """Test command argument sanitization""" - # Normal arguments should work - args = ["python", "-m", "dw.run", "workflow.json"] - sanitized = sanitize_command_args(args) - assert len(sanitized) == len(args) - assert ( - sanitized == args - ) # With shell=False, arguments pass through after validation - - # Arguments with semicolons should fail - with pytest.raises(InvalidInputError): - sanitize_command_args(["rm", "-rf", "; rm -rf /"]) - - # Arguments with $ should fail - with pytest.raises(InvalidInputError): - sanitize_command_args(["echo", "$(malicious_command)"]) - - # Arguments with pipes should fail - with pytest.raises(InvalidInputError): - sanitize_command_args(["cat", "/etc/passwd | grep root"]) - - class TestValidatePathRejections: """Inputs validate_path must refuse outright""" @@ -432,6 +409,11 @@ def test_a_component_with_a_null_byte_is_rejected(self): class TestSanitizeCommandArgs: + def test_ordinary_arguments_pass_through_unchanged(self): + # shell=False does the quoting, so a clean argument is returned verbatim + args = ["python", "-m", "dw.run", "workflow.json"] + assert sanitize_command_args(args) == args + def test_non_string_arguments_are_coerced(self): assert sanitize_command_args(["--steps", 25, 1.5]) == ["--steps", "25", "1.5"] diff --git a/tests/test_select.py b/tests/test_select.py index ac5148a9..98907922 100644 --- a/tests/test_select.py +++ b/tests/test_select.py @@ -112,8 +112,3 @@ def test_it_runs_through_the_task_dispatch(self): ) assert result == "b" - - def test_command_registered(self): - from dw.tasks.task import _COMMAND_REGISTRY - - assert "select" in _COMMAND_REGISTRY diff --git a/tests/test_server.py b/tests/test_server.py index 202cf560..87cbd40b 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -1189,13 +1189,6 @@ def test_bearer_token_gates_the_api_when_configured(tmp_path): assert client.get(f"/api/jobs/{job['id']}/events").status_code == 401 -def test_no_token_configured_means_no_auth(server): - """The default, unconfigured behavior is unchanged: no token means no - Authorization check at all.""" - with server(success_script) as client: - assert client.get("/api/health").status_code == 200 - - def test_save_workflow_roundtrip_and_confinement(server, tmp_path): with server(success_script) as client: workflow = valid_workflow("saved") @@ -2700,17 +2693,6 @@ def test_the_asset_library_is_reported(asset_server, tmp_path): assert directories["assets"] == str(tmp_path / "assets") -def test_without_an_asset_library_uploads_keep_the_old_shape(server, tmp_path): - with server(success_script) as client: - body = client.post( - "/api/uploads", - params={"filename": "source-image.png"}, - content=b"bytes", - ).json() - assert os.path.isabs(body["path"]) - assert body["url"].startswith("/outputs/uploads/") - - def test_upload_media_rejects_disallowed_extension(server): with server(success_script) as client: response = client.post( @@ -2778,34 +2760,6 @@ def test_workflow_listing_carries_details(server): assert listing["details"]["Basic"]["kinds"] == [] -def test_workflow_details_name_the_template_a_model_config_configures(server): - """A model config is a tuned instance of a template, and a client that - cannot see which is which shows it as just another catalog entry - the - thing the two-tree layout exists to stop.""" - with server(success_script) as client: - client.put( - "/api/workflows/templates/text-to-image", - json={"workflow": valid_workflow("tti")}, - ) - - workflow = valid_workflow("tuned") - workflow["configures"] = "templates/text-to-image" - client.put("/api/workflows/models/Tuned", json={"workflow": workflow}) - - listing = client.get("/api/workflows").json() - - assert listing["details"]["models/Tuned"]["configures"] == ( - "templates/text-to-image" - ) - - -def test_a_workflow_that_configures_nothing_says_so(server): - with server(success_script) as client: - listing = client.get("/api/workflows").json() - - assert listing["details"]["Basic"]["configures"] == "" - - def test_the_listing_filters_and_compacts(server): with server(success_script) as client: client.put( diff --git a/tests/test_server_workspaces.py b/tests/test_server_workspaces.py index 9fd381a7..e73d1ca0 100644 --- a/tests/test_server_workspaces.py +++ b/tests/test_server_workspaces.py @@ -251,10 +251,6 @@ def test_saving_a_workflow_with_a_relative_sub_workflow_path( ) assert response.status_code == 200 - def test_an_unknown_workspace_query_param_is_a_404(self, server): - with server() as client: - assert client.get("/api/workflows?workspace=nope").status_code == 404 - def test_a_traversal_attempt_as_a_workspace_name_is_a_400(self, server): with server() as client: response = client.get("/api/workflows", params={"workspace": "../x"}) @@ -393,14 +389,16 @@ def test_asset_listing_entries_carry_a_fetchable_url(self, server, workspace_roo assert fetched.status_code == 200 assert fetched.content == b"iris" - def test_get_workflow_reports_its_origin_and_writability( - self, server, workspace_root - ): + def test_get_workflow_reports_its_origin_and_writability(self, server): + # a named workspace's own library is as writable as the default's - + # the headers follow the workspace the route was scoped to with server() as client: + client.post("/api/workspaces", json={"name": "shots"}) client.put( - "/api/workflows/Basic", json={"workflow": valid_workflow("mine")} + "/api/workflows/Basic?workspace=shots", + json={"workflow": valid_workflow("mine")}, ) - response = client.get("/api/workflows/Basic") + response = client.get("/api/workflows/Basic?workspace=shots") assert response.status_code == 200 assert response.headers["x-workflow-origin"] == "workspace" assert response.headers["x-workflow-writable"] == "true" @@ -448,8 +446,28 @@ def test_it_links_rather_than_copying_when_it_can(self, server, workspace_root): "/api/assets/keep", json={"name": "Gyre/run/clip.mp4"} ).json() assert body["reference"] == "asset:clip.mp4" - if body["linked"]: - assert os.stat(source).st_ino == os.stat(body["path"]).st_ino + # outputs and assets share one temporary filesystem, so a link is + # always possible here + assert body["linked"] is True + assert os.stat(source).st_ino == os.stat(body["path"]).st_ino + + def test_it_copies_when_it_cannot_link(self, server, workspace_root, monkeypatch): + """A different filesystem, or one with no links, still keeps the + output - as a separate copy of the same bytes.""" + source = self.written(workspace_root.outputs, "Gyre/run/clip.mp4", b"clip") + + def no_links(*args, **kwargs): + raise OSError("cross-device link") + + monkeypatch.setattr(os, "link", no_links) + with server() as client: + body = client.post( + "/api/assets/keep", json={"name": "Gyre/run/clip.mp4"} + ).json() + assert body["linked"] is False + assert os.stat(source).st_ino != os.stat(body["path"]).st_ino + with open(body["path"], "rb") as kept: + assert kept.read() == b"clip" def test_the_name_defaults_to_the_files_own(self, server, workspace_root): self.written(workspace_root.outputs, "Gyre/run/still.png") diff --git a/tests/test_step_cache.py b/tests/test_step_cache.py index 2c20d390..aa289aee 100644 --- a/tests/test_step_cache.py +++ b/tests/test_step_cache.py @@ -110,17 +110,6 @@ def test_step_cache_miss_on_first_run(): assert cache.get("w", step_data, 42, set(), "/out", True) is None -def test_step_cache_hit_when_output_dir_unchanged(): - cache = StepCache() - step_data = {"name": "gen", "pipeline": {"arguments": {"prompt": "a cat"}}} - result = FakeResult("first") - cache.put("w", step_data, 42, result, "/out/a", True) - - hit = cache.get("w", step_data, 42, set(), "/out/a", True) - - assert hit is result - - def test_step_cache_miss_when_output_dir_changes(): """A hit reuses the entry's saved_files/manifest paths verbatim, so a changed effective output dir must force a miss rather than silently diff --git a/tests/test_strip_exif_and_watermark.py b/tests/test_strip_exif_and_watermark.py index ec56964e..26425e60 100644 --- a/tests/test_strip_exif_and_watermark.py +++ b/tests/test_strip_exif_and_watermark.py @@ -79,9 +79,12 @@ def test_all_positions(self): def test_invalid_position_falls_back(self): img = Image.new("RGB", (400, 200)) - # Unknown position should fall back to bottom-right - result = add_watermark(img, position="nonsense") - self.assertIsInstance(result, Image.Image) + # Unknown position should fall back to bottom-right, pixel for pixel + result = add_watermark(img, position="nonsense", opacity=255) + expected = add_watermark(img, position="bottom-right", opacity=255) + elsewhere = add_watermark(img, position="top-left", opacity=255) + self.assertEqual(result.tobytes(), expected.tobytes()) + self.assertNotEqual(result.tobytes(), elsewhere.tobytes()) def test_custom_color(self): img = Image.new("RGB", (400, 200)) diff --git a/tests/test_task.py b/tests/test_task.py index 74184497..fe084a5d 100644 --- a/tests/test_task.py +++ b/tests/test_task.py @@ -65,20 +65,22 @@ def test_image_processor_command_dispatches_without_registry_entry(): assert result.size == (4, 4) -@pytest.mark.skip(reason="Requires network access to external URLs which may be flaky") -def test_gather_images_task(): - task_def = { - "command": "gather_images", - "arguments": { - "urls": [ - "https://pbs.twimg.com/media/Gf5iaDGXsAA0R30?format=jpg&name=small", - "https://pbs.twimg.com/media/Gf7vNQJXoAAY5Cm?format=jpg&name=small", - ] - }, - } +def test_gather_images_task_dispatches_to_gather(): + urls = ["https://example.com/a.jpg", "https://example.com/b.jpg"] + images = [Image.new("RGB", (4, 4)), Image.new("RGB", (8, 8))] + task_def = {"command": "gather_images", "arguments": {"urls": urls}} task = Task(task_def, "cpu") - result = task.run(task_def["arguments"]) - assert isinstance(result, list), "Expected a list of images from gather_images" + + with ( + patch( + "dw.tasks.gather.validate_media_url", side_effect=lambda url, what=None: url + ), + patch("dw.tasks.gather.load_image", side_effect=images) as load_image, + ): + result = task.run(task_def["arguments"]) + + assert result == images + assert [call.args[0] for call in load_image.call_args_list] == urls def test_gather_inputs_task(): @@ -119,7 +121,6 @@ def test_format_chat_message_task(): assert text_inputs[1]["content"] == "unit_test", "User message content mismatch" -@pytest.mark.skip(reason="Test not fully implemented yet") def test_batch_decode_post_process_task(): # We use a mock pipeline to simulate previous_pipelines behavior. class MockPipeline: @@ -145,10 +146,8 @@ def post_process_generation(self, generated_text, task): } task = Task(task_def, "cpu") result = task.run(task_def["arguments"], previous_pipelines=mock_previous_pipelines) - assert result == [ - "decoded-foo", - "decoded-bar", - ], "Should return batch-decoded strings" + # The first decoded sequence, post-processed and read back under the task key + assert result == "decoded-foo" class TestTaskDevice: diff --git a/tests/test_task_signature_errors.py b/tests/test_task_signature_errors.py index 02cc4d7e..1d449663 100644 --- a/tests/test_task_signature_errors.py +++ b/tests/test_task_signature_errors.py @@ -24,6 +24,8 @@ class of mistake a free pre-flight most obviously exists for (#141). An ) from dw.tasks.task import Task +REPO_ROOT = pathlib.Path(__file__).resolve().parent.parent + def task_step(command, arguments, name="a"): return { @@ -208,13 +210,13 @@ class TestTheCatalogItself: @pytest.mark.parametrize( "path", sorted( - str(p) - for p in list(pathlib.Path("workflows").rglob("*.json")) - + list(pathlib.Path("dw/workflows").glob("*.json")) + str(p.relative_to(REPO_ROOT)) + for p in list((REPO_ROOT / "workflows").rglob("*.json")) + + list((REPO_ROOT / "dw" / "workflows").glob("*.json")) ), ) def test_workflow_has_no_task_signature_error(self, path): - definition = json.loads(pathlib.Path(path).read_text()) + definition = json.loads((REPO_ROOT / path).read_text()) if not isinstance(definition, dict) or "steps" not in definition: pytest.skip("not a workflow") assert task_signature_errors(definition) == [] diff --git a/tests/test_tensor_image.py b/tests/test_tensor_image.py index 20ad2a60..fb6bfd5d 100644 --- a/tests/test_tensor_image.py +++ b/tests/test_tensor_image.py @@ -43,11 +43,6 @@ def test_dtype_cast(self): tensor = pil_to_float_tensor(image, "cpu", dtype=torch.float64) assert tensor.dtype == torch.float64 - def test_dtype_defaults_to_float32(self): - image = Image.new("RGB", (4, 4), color=(1, 2, 3)) - tensor = pil_to_float_tensor(image, "cpu") - assert tensor.dtype == torch.float32 - class TestFloatTensorToPil: def test_accepts_batched_and_unbatched(self): diff --git a/tests/test_text_sections.py b/tests/test_text_sections.py index 33e732c1..ac2f7d44 100644 --- a/tests/test_text_sections.py +++ b/tests/test_text_sections.py @@ -90,7 +90,7 @@ def test_no_sections_requested_is_a_passthrough(): def h3_prompts(): """Every hand-written H3 prompt in the examples, as (file, key, text).""" found = [] - pattern = os.path.join(REPO_ROOT, "workflows", "minimax", "MiniMaxH3*.json") + pattern = os.path.join(REPO_ROOT, "workflows", "templates", "minimax", "*.json") for path in sorted(glob.glob(pattern)): with open(path, encoding="utf-8") as handle: workflow = json.load(handle) @@ -100,6 +100,12 @@ def h3_prompts(): return found +def test_the_sweep_finds_the_shipped_prompts(): + """The parametrized test below collects nothing, and so passes, if the + templates move again.""" + assert h3_prompts() + + @pytest.mark.parametrize("name,key,prompt", h3_prompts()) def test_a_hand_written_prompt_passes_through_unchanged(name, key, prompt): """The trim must be a no-op on a prompt that is already well formed. diff --git a/tests/test_type_helpers.py b/tests/test_type_helpers.py index c8208024..1ba05ee9 100644 --- a/tests/test_type_helpers.py +++ b/tests/test_type_helpers.py @@ -33,12 +33,6 @@ def test_get_type_invalid_attribute(self): class TestLoadTypeFromName: """Test loading type by name from diffusers""" - def test_load_type_with_full_path(self): - # Test with fully qualified name - result = load_type_from_full_name("os.path.join") - assert callable(result) - assert result.__name__ == "join" - def test_load_type_invalid_full_path(self): with pytest.raises(ModuleNotFoundError): load_type_from_full_name("fake.module.Type") diff --git a/tests/test_video_utils.py b/tests/test_video_utils.py index 024b3338..456049fe 100644 --- a/tests/test_video_utils.py +++ b/tests/test_video_utils.py @@ -513,11 +513,6 @@ def test_a_variable_reference_is_left_unchanged(self, tmp_path): assert task["arguments"]["video"] == "variable:my_video" - def test_an_in_memory_frame_list_is_unaffected(self, video): - """A step whose 'video' is an earlier step's in-memory result still - goes through the ordinary extract_frame path.""" - assert get_frame(video, 2) is video[2] - class TestIsVideo: def test_the_shapes_that_are_videos(self): diff --git a/tests/test_worker.py b/tests/test_worker.py index 75f652b1..5a98a612 100644 --- a/tests/test_worker.py +++ b/tests/test_worker.py @@ -149,74 +149,6 @@ def test_worker_clear_memory(worker_process): assert "gpu_available" in info -@pytest.mark.skipif( - not os.path.exists(TEST_WORKFLOW_PATH), - reason=f"test workflow not found: {TEST_WORKFLOW_PATH}", -) -@requires_accelerator -def test_worker_with_simple_workflow(worker_process, tmp_path): - """Worker executes a real workflow and reuses cached models on a second run.""" - cmd_queue, res_queue, worker = worker_process - - output_dir = str(tmp_path / "test_outputs") - os.makedirs(output_dir, exist_ok=True) - - cmd_queue.put( - { - "type": "execute", - "workflow_path": TEST_WORKFLOW_PATH, - "arguments": {}, - "output_dir": output_dir, - "log_level": "INFO", - } - ) - - # First message after spawn (or after a fresh command with no prior - # traffic) must tolerate child import cost; workflow execution itself - # (model load + inference) is also slow, so keep the generous timeout - # for every message in this loop rather than switching to the short one. - success = False - saw_workflow_loaded = False - while True: - result = res_queue.get(timeout=WORKER_READY_TIMEOUT) - result_type = result.get("type") - - if result_type == "workflow_loaded": - saw_workflow_loaded = True - elif result_type == "success": - success = True - break - elif result_type == "error": - pytest.fail(f"Workflow execution error: {result['message']}") - - assert success - assert saw_workflow_loaded - - # Run again to exercise the model-reuse/caching path. - cmd_queue.put( - { - "type": "execute", - "workflow_path": TEST_WORKFLOW_PATH, - "arguments": {}, - "output_dir": output_dir, - "log_level": "INFO", - } - ) - - second_run_count = None - while True: - result = res_queue.get(timeout=WORKER_READY_TIMEOUT) - result_type = result.get("type") - - if result_type == "success": - second_run_count = result["run_count"] - break - elif result_type == "error": - pytest.fail(f"Workflow execution error on second run: {result['message']}") - - assert second_run_count == 2 - - @pytest.mark.skipif( not os.path.exists(TEST_WORKFLOW_PATH), reason=f"test workflow not found: {TEST_WORKFLOW_PATH}", diff --git a/tests/test_workflow_trust.py b/tests/test_workflow_trust.py index 3285decd..98fe4c09 100644 --- a/tests/test_workflow_trust.py +++ b/tests/test_workflow_trust.py @@ -51,6 +51,9 @@ def _untrust(monkeypatch): class TestTrustFlag: def test_set_trust_workflows_true(self, monkeypatch): + # conftest already trusts; start untrusted so this can fail + _untrust(monkeypatch) + assert workflows_are_trusted() is False set_trust_workflows(True) assert workflows_are_trusted() is True diff --git a/ui/src/lib/format.test.ts b/ui/src/lib/format.test.ts index 4c05b5e9..de14db25 100644 --- a/ui/src/lib/format.test.ts +++ b/ui/src/lib/format.test.ts @@ -17,7 +17,9 @@ it('formats gigabyte sizes to two decimals at and above 1 GB', () => { expect(formatBytes(1024 ** 3 * 2.5)).toBe('2.50 GB') }) -it('renders a unix timestamp as a locale date/time string', () => { - const mtime = 1700000000 - expect(formatMtime(mtime)).toBe(new Date(mtime * 1000).toLocaleString()) +it('reads mtime as unix seconds, not milliseconds', () => { + // 1700000000 s is 2023-11-14T22:13:20Z; read as ms it would land in January 1970 + const rendered = formatMtime(1700000000) + expect(rendered).toBe(new Date('2023-11-14T22:13:20Z').toLocaleString()) + expect(rendered).not.toBe(new Date(1700000000).toLocaleString()) }) From 4493fa8b83f1f5fc81b2ab9d8121348f6617dd5d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 06:54:28 -0500 Subject: [PATCH 033/181] fix(ui): enhance directory hint styling for better responsiveness --- ui/src/lib/pages/PromptEditorPage.svelte | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/ui/src/lib/pages/PromptEditorPage.svelte b/ui/src/lib/pages/PromptEditorPage.svelte index 70523454..8c98a5bc 100644 --- a/ui/src/lib/pages/PromptEditorPage.svelte +++ b/ui/src/lib/pages/PromptEditorPage.svelte @@ -546,7 +546,7 @@ / {/if} - .json in {promptDir} + .json in {promptDir} {#if savePath()} Date: Wed, 23 Sep 2026 07:05:56 -0500 Subject: [PATCH 034/181] test: second pass - drop duplicate MCP/server tests, fix weak ones Policy: delete verified duplicates, keep prose pins on tool descriptions and parameter-forwarding tests, fix rather than delete tests that were flaky or asserted nothing. - MCP: per-tool pass-through and error-passthrough tests removed where TOOL_WIRING and test_mcp_client already cover the tool; added a client test for the streamed get_media_if error path they were the only cover for - server/result/elision/model_cache/etc: duplicates removed, unique assertions moved into the surviving test; prompt-reference sweep now also covers dw/workflows - getters and restated constants removed; vocabularies, schema closure and API fields (cost_basis, common_assets) kept - fixed: test_events watchdog timing (6x ratio, mutation-checked), integration tests now assert decoded outputs (multi-step fixture was eliding step 1), watermark/recenter/netinfo/host-memory/type-helper tests assert real effects Suite: 5636 passed, 14 skipped. Co-Authored-By: Claude Opus 5.5 (1M context) --- tests/test_component_type_errors.py | 26 ++--- tests/test_dialogue_short_voices.py | 10 -- tests/test_elision.py | 20 ---- tests/test_events.py | 60 +++++++---- tests/test_gather.py | 16 --- tests/test_host_memory.py | 14 ++- tests/test_image_utils.py | 8 -- tests/test_integration.py | 63 ++++++++++-- tests/test_mcp_authoring.py | 23 ----- tests/test_mcp_catalog.py | 18 +--- tests/test_mcp_client.py | 30 +++--- tests/test_mcp_diagnose.py | 56 +--------- tests/test_mcp_exports.py | 12 +-- tests/test_mcp_guides.py | 16 --- tests/test_mcp_media.py | 39 ------- tests/test_mcp_models.py | 58 ----------- tests/test_mcp_prompts.py | 137 +------------------------ tests/test_mcp_server.py | 23 ----- tests/test_model_cache.py | 29 +----- tests/test_output_references.py | 4 - tests/test_pipeline_components.py | 4 - tests/test_prompt_references.py | 14 ++- tests/test_prompts.py | 51 ++------- tests/test_recenter_crop.py | 8 +- tests/test_resize_bucket.py | 11 +- tests/test_result.py | 19 +--- tests/test_security.py | 5 +- tests/test_server.py | 47 +-------- tests/test_server_info.py | 33 +++--- tests/test_server_workspaces.py | 27 ++--- tests/test_speech_generation.py | 17 --- tests/test_strip_exif_and_watermark.py | 41 ++++++-- tests/test_task.py | 8 -- tests/test_teacache.py | 10 -- tests/test_tensor_image.py | 6 -- tests/test_text_generation.py | 9 -- tests/test_type_helpers.py | 6 +- tests/test_workflow.py | 27 +---- tests/test_workspace.py | 5 - 39 files changed, 241 insertions(+), 769 deletions(-) diff --git a/tests/test_component_type_errors.py b/tests/test_component_type_errors.py index 9de07505..eb14f7a8 100644 --- a/tests/test_component_type_errors.py +++ b/tests/test_component_type_errors.py @@ -64,7 +64,13 @@ def test_a_real_dotted_class_outside_the_narrow_allowlist_is_accepted(self): class TestAllowlistedAbsentVsPresentButDisallowed: - def test_an_absent_class_in_a_trusted_module_says_does_not_exist(self): + @pytest.mark.parametrize("trust", ["0", "1"]) + def test_an_absent_class_in_a_trusted_module_says_does_not_exist( + self, monkeypatch, trust + ): + """Untrusted too: an allowlisted module's missing class is a + misspelling, not the disallowed-module refusal below.""" + monkeypatch.setenv("DW_TRUST_WORKFLOWS", trust) errors = component_type_errors( { "id": "ct", @@ -91,24 +97,6 @@ def test_a_class_in_a_disallowed_module_says_not_allowed(self, monkeypatch): assert "does not exist" not in errors[0]["message"] assert "outside the ecosystem" in errors[0]["message"] - def test_the_two_messages_differ(self, monkeypatch): - monkeypatch.setenv("DW_TRUST_WORKFLOWS", "0") - absent = component_type_errors( - { - "id": "ct", - "steps": [ - pipeline_step( - None, - extra={ - "quantization_config": {"config_type": "sdnq.NoSuchConfig"} - }, - ) - ], - } - )[0]["message"] - disallowed = errors_for("os.system")[0]["message"] - assert absent != disallowed - class TestSchedulerAndQuantizationFields: def test_a_misspelled_scheduler_type_is_refused(self): diff --git a/tests/test_dialogue_short_voices.py b/tests/test_dialogue_short_voices.py index 5a52bf18..208be10d 100644 --- a/tests/test_dialogue_short_voices.py +++ b/tests/test_dialogue_short_voices.py @@ -100,16 +100,6 @@ def test_a_named_voice_is_referenced_in_the_shots_it_speaks_in( built = [r for r in references if not isinstance(r, dict)] assert len(built) == (1 if "a" in SPEAKERS[name] else 0), name - def test_the_tag_runs_longer(self, definition): - frames = {e["name"]: e["num_frames"] for e in definition["variables"]["shots"]} - assert frames == { - "cold_open": 124, - "deflect": 124, - "react": 124, - "button": 124, - "tag": 141, - } - def test_the_variable_names_are_roles_rather_than_a_cast(self, definition): """Every run carried howie_portrait_prompt and shot_3_howie_incredulous through its arguments, manifest and export whatever the cast was.""" diff --git a/tests/test_elision.py b/tests/test_elision.py index 18e196bf..98dd8146 100644 --- a/tests/test_elision.py +++ b/tests/test_elision.py @@ -401,12 +401,6 @@ def test_the_draw_steps_save_nothing(self): for name in ("draw_character_a", "draw_character_b"): assert steps[name]["result"]["save"] is False - def test_the_default_run_is_unchanged(self): - expanded = self.expanded(self.definition()) - elided = elide_definition(expanded) - assert elided == [] - assert "draw_character_a" in names(expanded["steps"]) - def test_a_cast_episode_draws_nothing(self): expanded = self.expanded(self.cast_from_files(self.definition())) elided = elide_definition(expanded) @@ -480,17 +474,3 @@ def test_the_portrait_saves_nothing(self): otherwise keep the step a cast episode has no use for.""" steps = {s["name"]: s for s in self.definition()["steps"]} assert steps["draw_singer"]["result"]["save"] is False - - def test_a_cast_singer_draws_nothing(self): - definition = self.definition() - definition["variables"]["singer_reference"] = { - "reference_type": "variable:image_reference_type", - "from_file": "asset:qa-cast/priya-portrait.jpg", - } - expanded = self.expanded(definition) - assert [e["step"] for e in elide_definition(expanded)] == ["draw_singer"] - - def test_the_default_run_still_draws(self): - expanded = self.expanded(self.definition()) - assert elide_definition(expanded) == [] - assert "draw_singer" in names(expanded["steps"]) diff --git a/tests/test_events.py b/tests/test_events.py index aa11a24c..71e1dafc 100644 --- a/tests/test_events.py +++ b/tests/test_events.py @@ -325,16 +325,36 @@ def mock_load(self, shared_components): assert current_context() is None, "context must deactivate after the run" -def _fast_watchdog(): +def _fast_watchdog(threshold=0.05, interval=0.02): """Patches the watchdog's timing constants down to something a test can wait out in real time, without touching the production defaults.""" return patch.multiple( events_module, - PHASE_STALL_THRESHOLD_SECONDS=0.05, - PHASE_STALL_CHECK_INTERVAL_SECONDS=0.02, + PHASE_STALL_THRESHOLD_SECONDS=threshold, + PHASE_STALL_CHECK_INTERVAL_SECONDS=interval, ) +# Production checks six times per threshold (30s / 5s). A test that asserts +# something does *not* happen within a window keeps that ratio, so the margin +# between the window and the threshold is several check intervals wide rather +# than a scheduler hiccup wide. +_RATIO_THRESHOLD = 0.6 +_RATIO_INTERVAL = 0.1 + + +def _stalls(events): + return [e for e in events if e.get("kind") == "phase_stall"] + + +def _wait_for_stall(events, deadline_seconds=5.0): + """Poll until the watchdog has reported at least one stall.""" + deadline = time.monotonic() + deadline_seconds + while not _stalls(events): + assert time.monotonic() < deadline, "watchdog never reported a stall" + time.sleep(0.005) + + def test_watchdog_fires_after_threshold_with_no_events(): events = [] context = RunContext(on_event=events.append) @@ -355,16 +375,18 @@ def test_watchdog_fires_after_threshold_with_no_events(): def test_watchdog_does_not_fire_before_threshold(): events = [] context = RunContext(on_event=events.append) - with _fast_watchdog(): + with _fast_watchdog(_RATIO_THRESHOLD, _RATIO_INTERVAL): context.enter_run() try: context.note_phase("generating") - time.sleep(0.03) + # Two check intervals, so the watchdog has really looked (a + # watchdog that ignored the threshold would fire here), and well + # short of the threshold + time.sleep(2.5 * _RATIO_INTERVAL) finally: context.exit_run() - stalls = [e for e in events if e.get("kind") == "phase_stall"] - assert not stalls + assert not _stalls(events) def test_watchdog_repeats_while_the_stall_continues(): @@ -390,26 +412,28 @@ def test_watchdog_repeats_while_the_stall_continues(): def test_watchdog_stops_once_a_new_event_arrives(): + # The stall report bumps the silence clock itself, so it repeats one + # threshold after the last report. A progress event part-way through that + # wait must restart the clock: the window below runs past when the repeat + # would have come without the reset, and ends well before one threshold + # after the progress event. + threshold = _RATIO_THRESHOLD events = [] context = RunContext(on_event=events.append) - with _fast_watchdog(): + with _fast_watchdog(threshold, _RATIO_INTERVAL): context.enter_run() try: context.note_phase("generating") - time.sleep(0.06) + _wait_for_stall(events) + stalls_before_progress = len(_stalls(events)) + time.sleep(0.6 * threshold) context.emit("pipeline_step", step=1) - time.sleep(0.01) - count_after_progress = len( - [e for e in events if e.get("kind") == "phase_stall"] - ) - time.sleep(0.01) - count_soon_after = len( - [e for e in events if e.get("kind") == "phase_stall"] - ) + time.sleep(0.65 * threshold) + stalls_after_progress = _stalls(events)[stalls_before_progress:] finally: context.exit_run() - assert count_soon_after == count_after_progress, ( + assert stalls_after_progress == [], ( "a fresh event must reset the silence clock, not just a fresh phase" ) diff --git a/tests/test_gather.py b/tests/test_gather.py index 5d169336..9b091c17 100644 --- a/tests/test_gather.py +++ b/tests/test_gather.py @@ -77,22 +77,6 @@ def test_gather_images_sorted_not_filesystem_order(self, tmp_path): (0, 0, 255), ] - @patch("dw.tasks.gather.load_image") - @patch("dw.tasks.gather.validate_media_url") - def test_gather_images_from_urls(self, mock_validate_url, mock_load_image): - """Test gathering images from URLs""" - mock_validate_url.side_effect = lambda url, what=None: url - mock_image = Image.new("RGB", (100, 100)) - mock_load_image.return_value = mock_image - - urls = ["https://example.com/img1.jpg", "https://example.com/img2.jpg"] - - images = gather_images(urls=urls) - - assert len(images) == 2 - assert mock_validate_url.call_count == 2 - assert mock_load_image.call_count == 2 - def test_gather_images_mixed_sources(self, tmp_path): """Local matches and URLs are both gathered - the files first, then the URLs in the order given.""" diff --git a/tests/test_host_memory.py b/tests/test_host_memory.py index 9b69fd26..cb303dc2 100644 --- a/tests/test_host_memory.py +++ b/tests/test_host_memory.py @@ -6,6 +6,8 @@ """ import builtins +import os +import sys import pytest @@ -50,17 +52,27 @@ def boom(): def test_the_stdlib_fallback_covers_a_machine_without_psutil(monkeypatch): """psutil is opportunistic here, not a declared dependency.""" real_import = builtins.__import__ + refused = [] def no_psutil(name, *args, **kwargs): if name == "psutil": + refused.append(name) raise ImportError("no psutil") return real_import(name, *args, **kwargs) monkeypatch.setattr(builtins, "__import__", no_psutil) stats = host_memory.host_memory_stats() monkeypatch.undo() + # the psutil path was taken and refused, and the call still answered + assert refused + assert set(stats) == {"rss_mb", "peak_rss_mb", "total_mb", "available_mb"} # The peak comes from resource(2), which needs neither psutil nor /proc - assert stats["peak_rss_mb"] is None or stats["peak_rss_mb"] > 1.0 + if os.name == "posix": + assert stats["peak_rss_mb"] > 1.0 + # /proc answers the rest on Linux + if sys.platform.startswith("linux"): + assert stats["rss_mb"] > 1.0 + assert stats["total_mb"] > 1.0 def test_the_worker_reports_host_fields_beside_the_gpu_ones(): diff --git a/tests/test_image_utils.py b/tests/test_image_utils.py index 47c08cf3..b026702c 100644 --- a/tests/test_image_utils.py +++ b/tests/test_image_utils.py @@ -18,14 +18,6 @@ class TestUnknownProcessor(unittest.TestCase): - def test_unknown_processor_raises_with_expected_message(self): - img = Image.new("RGB", (10, 10)) - with self.assertRaises(Exception) as ctx: - process_image(img, "not_a_real_processor", "cpu", {}) - self.assertEqual( - str(ctx.exception), "Unknown image processor type: not_a_real_processor" - ) - def test_unknown_processor_message_uses_lowered_name(self): img = Image.new("RGB", (10, 10)) with self.assertRaises(Exception) as ctx: diff --git a/tests/test_integration.py b/tests/test_integration.py index 3cf9a262..9fd81760 100644 --- a/tests/test_integration.py +++ b/tests/test_integration.py @@ -7,9 +7,33 @@ import os import json import tempfile +from PIL import Image from dw.workflow import Workflow, workflow_from_file +def decode_qr(image): + """The text a QR code image carries, read back with OpenCV. + + OpenCV's default detector misses some codes at 768px that it reads at + another size (it cannot read "Overridden Content" at 768), so this tries + the ArUco-based detector and a second size before giving up. Decoding is + deterministic, so the fallbacks add no flakiness.""" + cv2 = pytest.importorskip("cv2") + import numpy as np + + gray = image.convert("L") + attempts = ( + (cv2.QRCodeDetectorAruco, gray), + (cv2.QRCodeDetector, gray), + (cv2.QRCodeDetector, gray.resize((256, 256))), + ) + for detector, candidate in attempts: + text, _, _ = detector().detectAndDecode(np.array(candidate)) + if text: + return text + return "" + + @pytest.fixture def temp_workflow_dir(): """Create temporary directory for test workflows""" @@ -57,7 +81,7 @@ def multi_step_workflow(): "command": "format_chat_message", "arguments": { "system_prompt": "System", - "user_message": "variable:text1", + "user_message": "previous_result:gather_inputs", }, }, "result": {"content_type": "application/json", "save": False}, @@ -89,7 +113,8 @@ def test_workflow_with_variable_override( # Override the content variable result = workflow.run({"content": "Overridden Content"}) - assert result is not None + assert len(result) == 1 + assert decode_qr(result[0]) == "Overridden Content" def test_multi_step_workflow(self, multi_step_workflow, temp_workflow_dir): """Test workflow with multiple steps""" @@ -98,8 +123,22 @@ def test_multi_step_workflow(self, multi_step_workflow, temp_workflow_dir): result = workflow.run({}) - # Should return the last step's results - assert result is not None + # The last step's results: one chat message per value the first step + # gathered, each carrying that step's substituted variable + assert result == [ + { + "text_inputs": [ + {"role": "system", "content": "System"}, + {"role": "user", "content": "First"}, + ] + }, + { + "text_inputs": [ + {"role": "system", "content": "System"}, + {"role": "user", "content": "Second"}, + ] + }, + ] def test_workflow_from_file_execution(self, simple_qr_workflow, temp_workflow_dir): """Test loading and executing workflow from file""" @@ -113,7 +152,15 @@ def test_workflow_from_file_execution(self, simple_qr_workflow, temp_workflow_di workflow.validate() result = workflow.run({}) - assert result is not None + assert decode_qr(result[0]) == "Hello World" + # The file's name is the run's identity, and the saved image is the + # one the step returned + [entry] = workflow.manifest + [saved] = entry["files"] + relative = os.path.relpath(saved, os.path.realpath(temp_workflow_dir)) + assert relative.split(os.sep)[0] == "test_workflow" + with Image.open(saved) as image: + assert decode_qr(image) == "Hello World" def test_workflow_result_saving(self, simple_qr_workflow, temp_workflow_dir): """Test that workflow results are saved to output directory""" @@ -218,7 +265,7 @@ def test_simple_dependency(self, temp_workflow_dir): "name": "step2", "task": { "command": "gather_inputs", - "inputs": ["previous_result:step1"], + "arguments": {"value": "previous_result:step1"}, }, "result": {"content_type": "application/json", "save": False}, }, @@ -229,8 +276,8 @@ def test_simple_dependency(self, temp_workflow_dir): workflow.validate() result = workflow.run({}) - # step2 should receive the results from step1 - assert result is not None + # step2 runs once per result step1 produced, receiving each one + assert result == [{"value": "value1"}, {"value": "value2"}] if __name__ == "__main__": diff --git a/tests/test_mcp_authoring.py b/tests/test_mcp_authoring.py index 0ef41f48..72cd7bfb 100644 --- a/tests/test_mcp_authoring.py +++ b/tests/test_mcp_authoring.py @@ -152,20 +152,6 @@ def test_save_surfaces_a_rejected_definition(): authoring.save_workflow(client, "mine", WORKFLOW) -def test_save_surfaces_a_path_the_server_refuses(): - client, _seen = scripted( - { - ("PUT", "/api/workflows/../escape"): ( - 400, - {"detail": "Path traversal is not allowed"}, - ) - } - ) - - with pytest.raises(DwApiError, match="traversal"): - authoring.save_workflow(client, "../escape", WORKFLOW) - - def test_save_with_patch_merges_onto_the_stored_definition(): """A small edit shouldn't require resending the whole document (#202).""" stored = { @@ -234,15 +220,6 @@ def test_delete_calls_delete(): assert seen == [("DELETE", "/api/workflows/mine")] -def test_delete_surfaces_a_missing_workflow(): - client, _seen = scripted( - {("DELETE", "/api/workflows/ghost"): (404, {"detail": "No such workflow"})} - ) - - with pytest.raises(DwApiError, match="No such workflow"): - authoring.delete_workflow(client, "ghost") - - def body_recording_client(): """A client that keeps the request body, for the parts of a payload the (method, path) recorders above cannot see.""" diff --git a/tests/test_mcp_catalog.py b/tests/test_mcp_catalog.py index b7b698cb..a6aef5a9 100644 --- a/tests/test_mcp_catalog.py +++ b/tests/test_mcp_catalog.py @@ -5,7 +5,7 @@ import pytest from dw_mcp import catalog -from dw_mcp.client import DwApiError, DwClient +from dw_mcp.client import DwClient def recording_client(body=None, status=200): @@ -140,12 +140,6 @@ def test_list_gallery_sends_only_orphans_only_when_true(): assert seen["params"]["only_orphans"] == "true" -def test_a_pass_through_tool_returns_the_body_unchanged(): - client, _seen = recording_client({"workflows": ["a"], "details": {}}) - - assert catalog.list_workflows(client)["workflows"] == ["a"] - - FULL_ENTRY = { "summary": "a cut sequence", "shape": "sequence", @@ -212,16 +206,6 @@ def test_list_workflows_passes_its_filters_through(): } -def test_a_missing_workflow_propagates_the_api_error(): - def handler(request): - return httpx.Response(404, json={"detail": "No such workflow: ghost"}) - - client = DwClient(transport=httpx.MockTransport(handler)) - - with pytest.raises(DwApiError, match="ghost"): - catalog.get_workflow(client, "ghost") - - def test_get_workflow_sends_the_name_percent_encoded_on_the_wire(): """httpx.URL.path decodes escapes back for display, so an unquoted and a quoted request can look identical on `.path` - only the wire bytes diff --git a/tests/test_mcp_client.py b/tests/test_mcp_client.py index 68239700..ba3ed338 100644 --- a/tests/test_mcp_client.py +++ b/tests/test_mcp_client.py @@ -142,6 +142,22 @@ def handler(request): assert "/outputs/a.png" in str(caught.value) +def test_get_media_if_surfaces_an_error_status_verbatim(): + """The streamed path has to read the body before it can report an + error status - the media tools (get_output_image/_audio/_text) all + reach the server this way rather than through get_json.""" + + def handler(request): + return httpx.Response(404, json={"detail": "Unknown file"}) + + with pytest.raises(DwApiError) as caught: + client_with(handler).get_media_if( + "/api/gallery/ghost.wav/audio", lambda ct: True, max_bytes=1024 + ) + + assert str(caught.value) == "Unknown file" + + def test_a_refused_connection_says_how_to_start_the_server(): def handler(request): raise httpx.ConnectError("refused", request=request) @@ -241,24 +257,10 @@ def test_api_path_joins_literal_and_encoded_segments(): assert api_path("api", "jobs", "j1", "cancel") == "/api/jobs/j1/cancel" -def test_api_path_quotes_a_slash_in_a_segment(): - assert ( - api_path("api", "workflows", "flux/FluxDev") == "/api/workflows/flux%2FFluxDev" - ) - - def test_api_path_quotes_dot_segments_so_they_cannot_traverse(): assert api_path("api", "workflows", "../escape") == "/api/workflows/..%2Fescape" -def test_api_path_quotes_a_hash(): - assert api_path("api", "gallery", "a#1") == "/api/gallery/a%231" - - -def test_api_path_quotes_a_space(): - assert api_path("outputs", "a b.png") == "/outputs/a%20b.png" - - # Every dw_mcp handler module that talks to the API must build request paths # through `api_path`, not by hand - `path_segment` alone is easy to forget on # one call site among many, and a bare f-string skips quoting entirely. This diff --git a/tests/test_mcp_diagnose.py b/tests/test_mcp_diagnose.py index 99293206..6399e6da 100644 --- a/tests/test_mcp_diagnose.py +++ b/tests/test_mcp_diagnose.py @@ -55,18 +55,9 @@ def test_run_submits_once_the_cost_is_acknowledged(): assert result["job_id"] == "job-1" assert result["status"] == "queued" assert result["queue_position"] == 2 - assert len(seen) == 1 - - -def test_run_returns_immediately_rather_than_waiting_for_the_job(): - """A generation takes minutes; no MCP client will hold a call open. The - contract is submit-then-poll, so exactly one request goes out.""" - client, seen = submitting() - - result = diagnose.run_workflow( - client, workflow_path="w.json", acknowledged_cost=True - ) - + # A generation takes minutes and no MCP client holds a call open: the + # contract is submit-then-poll, so one request goes out and the answer + # names the tool to poll with assert [entry["key"] for entry in seen] == [("POST", "/api/jobs")] assert "get_job_events" in result["next"] @@ -123,15 +114,6 @@ def test_run_passes_variable_overrides(): assert b"a cat" in seen[0]["body"] -def test_run_surfaces_a_rejected_workflow(): - client, _seen = scripted( - {("POST", "/api/jobs"): (400, {"detail": "steps must not be empty"})} - ) - - with pytest.raises(DwApiError, match="steps must not be empty"): - diagnose.run_workflow(client, inline_workflow=WORKFLOW, acknowledged_cost=True) - - def test_get_job_returns_the_detail_payload(): client, _seen = scripted( { @@ -206,20 +188,6 @@ def test_cancel_rerun_and_move_call_their_routes(): assert b"front" in seen[2]["body"] -def test_move_surfaces_a_job_that_has_left_the_queue(): - client, _seen = scripted( - { - ("POST", "/api/jobs/job-1/move"): ( - 409, - {"detail": "Job is not queued - only queued jobs move"}, - ) - } - ) - - with pytest.raises(DwApiError, match="only queued jobs move"): - diagnose.move_job(client, "job-1", "up") - - def test_rerun_refuses_without_an_acknowledged_cost(): """A rerun queues the same generation from a stored spec - the same GPU minutes on the same one-job-at-a-time engine. The gate on run_workflow @@ -365,18 +333,6 @@ def test_wait_for_job_reports_an_uncapped_budget_honestly(monkeypatch): assert "waited_seconds" in result -def test_wait_for_job_does_not_require_acknowledged_cost(): - """It reads an already-queued job rather than starting anything, so the - cost gate other job-queuing tools carry does not apply here.""" - client, _seen = sequenced( - ("GET", "/api/jobs/job-1"), [{"id": "job-1", "status": "succeeded"}] - ) - - result = diagnose.wait_for_job(client, "job-1") - - assert result["status"] == "succeeded" - - FAT_JOB = { "id": "job-1", "workflow_name": "minimax/dialogue-short", @@ -626,12 +582,6 @@ def test_next_names_the_two_tools_that_use_it(self): assert "save_workflow" in result["next"] assert "run_workflow" in result["next"] - def test_an_unknown_job_raises_the_client_error(self): - client, _ = scripted({}) - - with pytest.raises(DwApiError): - diagnose.get_job_workflow(client, "nope") - def test_run_pins_a_job_to_a_named_workspace_without_switching(): """A session restart resets the session workspace to default, and a diff --git a/tests/test_mcp_exports.py b/tests/test_mcp_exports.py index bea5242c..20085258 100644 --- a/tests/test_mcp_exports.py +++ b/tests/test_mcp_exports.py @@ -2,10 +2,9 @@ says so - the lesson download_output taught.""" import httpx -import pytest from dw_mcp import exports -from dw_mcp.client import DwApiError, DwClient +from dw_mcp.client import DwClient SUMMARY = { "job_id": "job-1", @@ -90,15 +89,6 @@ def test_overwrite_travels_as_a_query_parameter(): assert seen[0]["params"]["overwrite"] == "true" -def test_a_409_reaches_the_model_as_a_readable_refusal(): - client, _ = exporting(status=409, body={"detail": "An export already exists"}) - - with pytest.raises(DwApiError) as caught: - exports.export_job(client, "job-1") - - assert "already exists" in str(caught.value) - - def test_the_docstring_names_total_bytes(): # The field a caller reads to see what an export cost on disk has to be # named where the caller looks - the handler's docstring. diff --git a/tests/test_mcp_guides.py b/tests/test_mcp_guides.py index 4070b1b8..00ca74f4 100644 --- a/tests/test_mcp_guides.py +++ b/tests/test_mcp_guides.py @@ -64,22 +64,6 @@ def test_get_guide_section_is_a_query_parameter(): assert seen[0]["params"] == {"section": "speech-generation"} -def test_a_404_detail_reaches_the_model_as_the_message(): - # The server writes "No guide named 'x'. The guides are: ..." - that text - # is the answer, and it must not be replaced by an HTTP status - client, _seen = scripted( - { - ("GET", "/api/guides/nonexistent"): ( - 404, - {"detail": "No guide named 'nonexistent'. The guides are: tasks."}, - ) - } - ) - - with pytest.raises(DwApiError, match="The guides are: tasks"): - guides.get_guide(client, "nonexistent") - - def test_a_guide_name_is_path_encoded(): """A name with '..' must reach the server intact, so its own validation - not httpx's dot-segment normalisation - decides what it means. diff --git a/tests/test_mcp_media.py b/tests/test_mcp_media.py index 0e6d1a1b..5839bf0a 100644 --- a/tests/test_mcp_media.py +++ b/tests/test_mcp_media.py @@ -230,15 +230,6 @@ def test_the_budget_is_checked_against_the_base64_size_not_the_raw_bytes( with pytest.raises(DwApiError, match="byte limit"): get_output_audio(client, "clip.wav") - def test_a_missing_file_propagates_the_api_error(self): - def handler(request): - return httpx.Response(404, json={"detail": "Unknown file"}) - - client = DwClient(transport=httpx.MockTransport(handler)) - - with pytest.raises(DwApiError, match="Unknown file"): - get_output_audio(client, "ghost.wav") - def test_audio_is_fetched_from_the_gallery_audio_route(): seen = [] @@ -333,16 +324,6 @@ def handler(request): assert stream.iterated is False -def test_a_missing_file_propagates_the_api_error(): - def handler(request): - return httpx.Response(404, json={"detail": "Unknown file"}) - - client = DwClient(transport=httpx.MockTransport(handler)) - - with pytest.raises(DwApiError, match="Unknown file"): - get_output_image(client, "ghost.png") - - def test_the_name_is_url_quoted_in_the_request(): # "#" starts a URL fragment when left unescaped - an unquoted name would # arrive at the server truncated ("a b", with "1.png" silently dropped as @@ -509,16 +490,6 @@ def test_the_text_tool_names_the_tool_that_can_read_an_image(): media.get_output_text(client, "out.png") -def test_a_missing_text_output_surfaces_the_error(): - def handler(request): - return httpx.Response(404, json={"detail": "Unknown file"}) - - client = DwClient(transport=httpx.MockTransport(handler)) - - with pytest.raises(DwApiError): - media.get_output_text(client, "ghost.txt") - - def test_undecodable_bytes_do_not_crash_the_tool(): """A file the server labels text but that is not valid UTF-8 should read as damaged output, not as a tool that blew up.""" @@ -545,16 +516,6 @@ def handler(request): assert seen == [("DELETE", "/api/gallery/out.png")] -def test_delete_output_surfaces_a_missing_file(): - def handler(request): - return httpx.Response(404, json={"detail": "Unknown file"}) - - client = DwClient(transport=httpx.MockTransport(handler)) - - with pytest.raises(DwApiError, match="Unknown file"): - media.delete_output(client, "ghost.png") - - def deleting_by_job(job): """GET /api/jobs/job-1 answers `job`; DELETE on the gallery answers as the server's run-directory form does.""" diff --git a/tests/test_mcp_models.py b/tests/test_mcp_models.py index c3c36631..258ed6a8 100644 --- a/tests/test_mcp_models.py +++ b/tests/test_mcp_models.py @@ -58,35 +58,6 @@ def test_it_refuses_without_an_acknowledgement(self, call): assert "acknowledged_cost=true" in str(excinfo.value) - @pytest.mark.parametrize( - "call, method, path", - [ - ( - lambda c: models.download_model(c, "org/model", acknowledged_cost=True), - "POST", - "/api/models/download", - ), - ( - lambda c: models.delete_model(c, "org/model", acknowledged_cost=True), - "DELETE", - "/api/models", - ), - ( - lambda c: models.update_diffusers(c, acknowledged_cost=True), - "POST", - "/api/system/diffusers/update", - ), - ], - ids=["download_model", "delete_model", "update_diffusers"], - ) - def test_an_acknowledgement_lets_it_through(self, call, method, path): - client, seen = recording_client() - - call(client) - - assert seen["method"] == method - assert seen["path"] == path - def test_each_refusal_says_what_that_particular_tool_costs(self): # One shared message would tell the user "this occupies the GPU" for # a download, which is wrong and trains them to wave the gate through @@ -119,12 +90,6 @@ def test_download_model_reports_that_it_did_not_wait(self): assert result["id"] == "d1" assert "list_downloads" in result["next"] - def test_list_downloads_is_a_plain_read(self): - client, seen = recording_client({"downloads": []}) - - assert models.list_downloads(client) == {"downloads": []} - assert (seen["method"], seen["path"]) == ("GET", "/api/models/downloads") - def test_cancel_download_needs_no_acknowledgement(self): # Cancelling stops a cost rather than starting one - gating it would # make the safe direction the harder one @@ -145,12 +110,6 @@ def test_a_download_id_is_quoted_into_the_path(self): assert seen["raw_path"] == "/api/models/downloads/..%2Fescape/cancel" - def test_an_unknown_download_reports_the_servers_message(self): - client, _ = recording_client({"detail": "Unknown download"}, status=404) - - with pytest.raises(DwApiError, match="Unknown download"): - models.cancel_download(client, "nope") - class TestDeletion: def test_delete_model_sends_the_repo_as_a_query_parameter(self): @@ -160,25 +119,8 @@ def test_delete_model_sends_the_repo_as_a_query_parameter(self): assert seen["params"] == {"repo": "org/model"} - def test_a_busy_server_refusal_reaches_the_caller(self): - # The server refuses a delete while a job is queued or running; that - # reason is the actionable part, so it must not be flattened - client, _ = recording_client( - {"detail": "A job is running or queued - deleting model files ..."}, - status=409, - ) - - with pytest.raises(DwApiError, match="A job is running or queued"): - models.delete_model(client, "org/model", acknowledged_cost=True) - class TestDiffusersVersion: - def test_get_diffusers_state_is_a_plain_read(self): - client, seen = recording_client({"version": "0.31.0"}) - - assert models.get_diffusers_state(client) == {"version": "0.31.0"} - assert (seen["method"], seen["path"]) == ("GET", "/api/system/diffusers") - def test_update_diffusers_warns_that_it_can_break_the_install(self): with pytest.raises(DwApiError) as excinfo: models.update_diffusers(refusing_client()) diff --git a/tests/test_mcp_prompts.py b/tests/test_mcp_prompts.py index 8b0563e0..88e64ea1 100644 --- a/tests/test_mcp_prompts.py +++ b/tests/test_mcp_prompts.py @@ -13,21 +13,6 @@ PROMPT = {"text": "a duke on a sofa", "description": "a duke"} -def scripted(routes): - """routes: {(method, path): (status, json_body)}""" - seen = [] - - def handler(request): - key = (request.method, request.url.path) - seen.append(key) - if key not in routes: - return httpx.Response(404, json={"detail": f"unrouted {key}"}) - status, body = routes[key] - return httpx.Response(status, json=body) - - return DwClient(transport=httpx.MockTransport(handler)), seen - - def recording(response): """A client that records the one request it is given.""" seen = {} @@ -45,24 +30,9 @@ def handler(request): # ----------------------------------------------------------------- reading -def test_list_prompts_returns_the_library(): - client, seen = scripted( - { - ("GET", "/api/prompts"): ( - 200, - {"prompt_dir": "/p", "prompts": ["duke"], "details": {}}, - ) - } - ) - - result = prompts.list_prompts(client) - - assert result["prompts"] == ["duke"] - assert seen == [("GET", "/api/prompts")] - - def scripted_with_params(routes): - """`scripted`, keeping each request's query parameters.""" + """routes: {(method, path): (status, json_body)}; records each request's + (method, path) and query parameters.""" seen = [] def handler(request): @@ -106,13 +76,6 @@ def test_list_prompts_forwards_the_filters_and_can_ask_for_the_text(): } -def test_get_prompt_reads_one_by_name(): - client, seen = scripted({("GET", "/api/prompts/duke"): (200, PROMPT)}) - - assert prompts.get_prompt(client, "duke")["text"] == "a duke on a sofa" - assert seen == [("GET", "/api/prompts/duke")] - - def test_get_prompt_encodes_a_foldered_name(): """A prompt lives at `folder/name`, and the slash has to survive as a path segment the server validates rather than one httpx normalizes.""" @@ -130,25 +93,6 @@ def handler(request): assert seen["raw_path"] == b"/api/prompts/sitcom%2Fduke" -def test_get_prompt_surfaces_a_missing_prompt(): - client, _seen = scripted( - {("GET", "/api/prompts/ghost"): (404, {"detail": "No such prompt"})} - ) - - with pytest.raises(DwApiError, match="No such prompt"): - prompts.get_prompt(client, "ghost") - - -def test_get_prompt_schema_has_its_own_path(): - """Not /api/prompts/schema - a prompt named 'schema' would shadow it.""" - client, seen = scripted( - {("GET", "/api/prompt-schema"): (200, {"type": "object"})}, - ) - - assert prompts.get_prompt_schema(client)["type"] == "object" - assert seen == [("GET", "/api/prompt-schema")] - - # ----------------------------------------------------------------- writing @@ -165,77 +109,9 @@ def test_save_prompt_puts_the_definition_under_its_name(): assert result["name"] == "duke" -def test_save_prompt_surfaces_a_rejected_definition(): - client, _seen = scripted( - {("PUT", "/api/prompts/duke"): (400, {"detail": "'text' is a required"})} - ) - - with pytest.raises(DwApiError, match="required"): - prompts.save_prompt(client, "duke", {"description": "no text"}) - - -def test_save_prompt_surfaces_a_reserved_text_prefix(): - """The engine refuses a prompt whose text is itself a reference - the - server says so and the message has to reach the model unaltered.""" - client, _seen = scripted( - { - ("PUT", "/api/prompts/duke"): ( - 400, - { - "detail": "A prompt's text may not itself begin with a " - "reference prefix (variable:, previous_result:)" - }, - ) - } - ) - - with pytest.raises(DwApiError, match="reference prefix"): - prompts.save_prompt(client, "duke", {"text": "variable:x"}) - - -def test_delete_prompt_calls_delete(): - client, seen = scripted( - {("DELETE", "/api/prompts/duke"): (200, {"name": "duke", "deleted": True})} - ) - - assert prompts.delete_prompt(client, "duke")["deleted"] is True - assert seen == [("DELETE", "/api/prompts/duke")] - - -def test_delete_prompt_surfaces_a_path_the_server_refuses(): - client, _seen = scripted( - { - ("DELETE", "/api/prompts/../escape"): ( - 400, - {"detail": "Path traversal is not allowed"}, - ) - } - ) - - with pytest.raises(DwApiError, match="traversal"): - prompts.delete_prompt(client, "../escape") - - # ---------------------------------------------------------------- enhancer -def test_list_enhancers_returns_the_presets(): - # The shape the live server sends: a list of preset records, not a map - client, seen = scripted( - { - ("GET", "/api/enhancers"): ( - 200, - {"presets": [{"key": "h3", "default_model": "Qwen/Qwen3-4B"}]}, - ) - } - ) - - presets = prompts.list_enhancers(client)["presets"] - - assert [preset["key"] for preset in presets] == ["h3"] - assert seen == [("GET", "/api/enhancers")] - - def test_enhance_prompt_refuses_without_acknowledgement(): """Enhancing loads a language model and queues a real job on the one-at-a-time engine, so it is gated exactly like a run.""" @@ -301,12 +177,3 @@ def test_enhance_prompt_omits_overrides_it_was_not_given(): assert "model_name" not in seen["body"] assert "device" not in seen["body"] - - -def test_enhance_prompt_surfaces_an_unknown_preset(): - client, _seen = scripted( - {("POST", "/api/enhance"): (400, {"detail": "Unknown preset 'nope'"})} - ) - - with pytest.raises(DwApiError, match="Unknown preset"): - prompts.enhance_prompt(client, "an idea", preset="nope", acknowledged_cost=True) diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 7c61660f..3a367d8e 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -232,13 +232,6 @@ async def test_no_tool_exposes_base_dir(): assert "base_dir" not in json.dumps(tool.input_schema), name -@pytest.mark.asyncio -async def test_run_workflow_takes_an_acknowledged_cost_flag(): - tools = await tools_of(server_over(ok({}))) - - assert "acknowledged_cost" in tools["run_workflow"].input_schema["properties"] - - @pytest.mark.asyncio async def test_run_workflow_and_delete_output_take_the_turn_saving_parameters(): """Almost every run is followed by a wait, and most deletes are of the @@ -269,15 +262,6 @@ async def test_run_workflow_advertises_its_cost(): assert "acknowledged_cost" in description -@pytest.mark.asyncio -async def test_a_read_only_tool_round_trips_to_the_api(): - server = server_over(ok({"workflows": ["a"], "details": {}})) - - result = await server.call_tool("list_workflows", {}) - - assert "workflows" in json.dumps(_text_of(result)) - - @pytest.mark.asyncio async def test_an_unreachable_server_reports_how_to_start_it(): def refusing(request): @@ -926,13 +910,6 @@ async def test_optional_parameters_are_declared_nullable(): ) -@pytest.mark.asyncio -async def test_rerun_job_takes_an_acknowledged_cost_flag(): - tools = await tools_of(server_over(ok({}))) - - assert "acknowledged_cost" in tools["rerun_job"].input_schema["properties"] - - @pytest.mark.asyncio async def test_rerun_job_advertises_its_cost(): tools = await tools_of(server_over(ok({}))) diff --git a/tests/test_model_cache.py b/tests/test_model_cache.py index b16627e9..0307467d 100644 --- a/tests/test_model_cache.py +++ b/tests/test_model_cache.py @@ -14,13 +14,15 @@ class TestCachedModel: def test_factory_called_once_for_same_key(self): - factory = MagicMock(return_value="the-model") + """Callers get the identical cached object back, not a copy.""" + loaded = object() + factory = MagicMock(return_value=loaded) first = cached_model(("task", "model-a", "cpu"), factory) second = cached_model(("task", "model-a", "cpu"), factory) - assert first == "the-model" - assert second == "the-model" + assert first is loaded + assert second is loaded factory.assert_called_once() def test_distinct_keys_get_distinct_loads(self): @@ -35,27 +37,6 @@ def test_distinct_keys_get_distinct_loads(self): factory_a.assert_called_once() factory_b.assert_called_once() - def test_returns_the_same_object_instance(self): - """Callers get the identical cached object back, not a copy.""" - loaded = object() - factory = MagicMock(return_value=loaded) - - first = cached_model(("task", "model-a", "cpu"), factory) - second = cached_model(("task", "model-a", "cpu"), factory) - - assert first is loaded - assert second is loaded - - def test_factory_not_called_until_needed(self): - """A cache hit must not invoke factory at all, even zero times isn't assumed.""" - cached_model(("task", "shared-key", "cpu"), MagicMock(return_value="v1")) - - never_called = MagicMock(return_value="v2") - result = cached_model(("task", "shared-key", "cpu"), never_called) - - assert result == "v1" - never_called.assert_not_called() - def test_clear_model_cache_releases_entries(self): factory = MagicMock(return_value="the-model") cached_model(("task", "model-a", "cpu"), factory) diff --git a/tests/test_output_references.py b/tests/test_output_references.py index 441e266f..a77d9b9d 100644 --- a/tests/test_output_references.py +++ b/tests/test_output_references.py @@ -210,7 +210,3 @@ def test_a_refusal_names_the_character_it_objected_to(self): assert "' '" in str(caught.value) assert "position" in str(caught.value) - - def test_traversal_is_still_refused(self): - with pytest.raises((InvalidInputError, SecurityError)): - resolve_output_reference("output:ltx2/../../etc/passwd") diff --git a/tests/test_pipeline_components.py b/tests/test_pipeline_components.py index fc702851..35e92bdb 100644 --- a/tests/test_pipeline_components.py +++ b/tests/test_pipeline_components.py @@ -23,10 +23,6 @@ class TestDeclaredComponentNames: """A component outside the known list is loaded, not silently dropped""" - def test_known_names_are_always_included(self): - names = declared_component_names({}) - assert names == optional_component_names - def test_component_shaped_keys_are_detected(self): definition = { "configuration": {"component_type": "SomePipeline"}, diff --git a/tests/test_prompt_references.py b/tests/test_prompt_references.py index b8b4cf9e..c8f31880 100644 --- a/tests/test_prompt_references.py +++ b/tests/test_prompt_references.py @@ -17,6 +17,18 @@ from tests.test_examples import REPO_ROOT, get_example_files PROMPT_DIR = os.path.join(REPO_ROOT, "prompts") +BUILTIN_DIR = os.path.join(REPO_ROOT, "dw", "workflows") + + +def get_builtin_files(): + """The packaged builtin workflows - what 'builtin:' steps name - which + resolve prompt: references against the same library.""" + return sorted( + os.path.relpath(os.path.join(root, name), REPO_ROOT) + for root, _, files in os.walk(BUILTIN_DIR) + for name in files + if name.endswith(".json") + ) def prompt_references(definition): @@ -27,7 +39,7 @@ def prompt_references(definition): ] -@pytest.mark.parametrize("example_file", get_example_files()) +@pytest.mark.parametrize("example_file", get_example_files() + get_builtin_files()) def test_every_prompt_reference_resolves(example_file): path = os.path.join(REPO_ROOT, example_file) with open(path, encoding="utf-8") as file: diff --git a/tests/test_prompts.py b/tests/test_prompts.py index 096e248d..61b19bd4 100644 --- a/tests/test_prompts.py +++ b/tests/test_prompts.py @@ -7,7 +7,12 @@ import pytest from dw.arguments import is_prompt_reference, realize_args -from dw.prompts import fetch_prompt, get_prompt_dir, load_prompt +from dw.prompts import ( + RESERVED_TEXT_PREFIXES, + fetch_prompt, + get_prompt_dir, + load_prompt, +) from dw.security import InvalidInputError, SecurityError @@ -180,48 +185,12 @@ def shipped_prompt_files(): ) -def shipped_prompt_references(): - """Every distinct 'prompt:' string in the shipped and builtin workflows.""" - references = set() - - def collect(value): - if isinstance(value, str) and is_prompt_reference(value): - references.add(value) - elif isinstance(value, dict): - for item in value.values(): - collect(item) - elif isinstance(value, list): - for item in value: - collect(item) - - for tree in ( - os.path.join(REPO_ROOT, "workflows"), - os.path.join(REPO_ROOT, "dw", "workflows"), - ): - for root, _, files in os.walk(tree): - for name in files: - if not name.endswith(".json"): - continue - try: - with open(os.path.join(root, name), encoding="utf-8") as file: - collect(json.load(file)) - except json.JSONDecodeError: - continue - return sorted(references) - - class TestShippedLibrary: - """The prompt library the repo ships stays loadable, and the workflows - that lean on it keep pointing at prompts that exist - a rename under - prompts/ should fail here, not at run time.""" + """The prompt library the repo ships stays loadable. That the workflows + leaning on it point at prompts that exist is + tests/test_prompt_references.py's sweep.""" @pytest.mark.parametrize("path", shipped_prompt_files()) def test_shipped_prompt_is_valid(self, path): prompt = load_prompt(os.path.join(REPO_ROOT, path)) - assert not prompt["text"].startswith( - ("previous_result:", "variable:", "constant:", "prompt:") - ) - - @pytest.mark.parametrize("reference", shipped_prompt_references()) - def test_shipped_reference_resolves(self, reference): - assert fetch_prompt(reference, prompt_dir=LIBRARY_DIR) + assert not prompt["text"].startswith(RESERVED_TEXT_PREFIXES) diff --git a/tests/test_recenter_crop.py b/tests/test_recenter_crop.py index 90d33a66..75b228b9 100644 --- a/tests/test_recenter_crop.py +++ b/tests/test_recenter_crop.py @@ -15,10 +15,16 @@ def _marked(size=200, mark=(20, 180), colour="red"): class TestRecenterCrop: def test_the_chosen_point_lands_at_the_centre(self): + # A window wholly inside the source, so no fill is involved + image = _marked(mark=(70, 130)) out = recenter_crop( - _marked(), center_x=0.5, center_y=0.5, crop=0.5, width=100, height=100 + image, center_x=0.35, center_y=0.65, crop=0.5, width=100, height=100 ) assert out.size == (100, 100) + assert out.getpixel((50, 50)) == (255, 0, 0) + # and only there - the rest of the window is the navy field + assert out.getpixel((5, 5)) == (0, 0, 128) + assert out.getpixel((95, 95)) == (0, 0, 128) def test_a_corner_feature_is_brought_to_the_centre(self): image = _marked(mark=(20, 180)) diff --git a/tests/test_resize_bucket.py b/tests/test_resize_bucket.py index d6128595..87a94c51 100644 --- a/tests/test_resize_bucket.py +++ b/tests/test_resize_bucket.py @@ -3,7 +3,7 @@ import unittest from PIL import Image -from dw.tasks.image_utils import resize_bucket, _DEFAULT_RATIOS +from dw.tasks.image_utils import resize_bucket class TestResizeBucket(unittest.TestCase): @@ -71,15 +71,6 @@ def test_converts_to_rgb(self): result = resize_bucket(img, resolution=512) self.assertEqual(result.mode, "RGB") - def test_default_ratios_has_expected_entries(self): - # Sanity check that we have the standard ratios - ratio_values = {(r[0], r[1]) for r in _DEFAULT_RATIOS} - self.assertIn((1, 1), ratio_values) - self.assertIn((16, 9), ratio_values) - self.assertIn((9, 16), ratio_values) - self.assertIn((4, 3), ratio_values) - self.assertIn((3, 4), ratio_values) - class TestResizeBucketRegistration(unittest.TestCase): """Test that resize_bucket is accessible via process_image.""" diff --git a/tests/test_result.py b/tests/test_result.py index cfd349ae..e7fa9e71 100644 --- a/tests/test_result.py +++ b/tests/test_result.py @@ -40,6 +40,7 @@ def test_add_list_result(self): result = Result({}) result.add_result(["item1", "item2", "item3"]) assert result.result_list == ["item1", "item2", "item3"] + assert result.get_artifacts() == ["item1", "item2", "item3"] def test_add_string_strips_quotes(self): result = Result({}) @@ -55,12 +56,6 @@ def test_add_selected_unwraps_value_and_records_metadata(self): assert result.result_list == ["b"] assert result.selected == {"position": 1, "score": 0.9} - def test_get_artifacts_from_simple_list(self): - result = Result({}) - result.add_result(["item1", "item2"]) - artifacts = result.get_artifacts() - assert artifacts == ["item1", "item2"] - def test_get_artifact_properties(self): result = Result({}) result.add_result( @@ -1754,18 +1749,6 @@ def test_the_audio_and_video_writes_measure_what_they_wrote(self, tmp_path): measured.assert_called_once() assert measured.call_args.args[0].endswith(".wav") - def test_it_says_both_the_prediction_and_the_written_clip_for_a_wav(self, tmp_path): - """A wav's write is itself the clip (#295): unlike a lossy re-encode, - there is no later encode step for the pre-write warning to describe - as a future risk, so the pre-write prediction and the post-write - ground truth are two different facts about this file and both fire.""" - kinds = [ - warning["kind"] - for warning in self.warnings_from(lambda: self.save_wav(1.5, str(tmp_path))) - ] - - assert kinds == ["audio_no_headroom", "audio_clipped"] - def test_an_image_is_never_probed(self, tmp_path): """Only a file that can carry a soundtrack pays for the read-back.""" from dw.result import Result diff --git a/tests/test_security.py b/tests/test_security.py index 8952c01e..42c083af 100644 --- a/tests/test_security.py +++ b/tests/test_security.py @@ -63,10 +63,7 @@ def test_url_validation(): ) assert validate_url("http://localhost:8080/api") == "http://localhost:8080/api" - # Invalid schemes should fail - with pytest.raises(InvalidInputError): - validate_url("file:///etc/passwd") - + # Invalid schemes should fail (file:// is in TestValidateUrl) with pytest.raises(InvalidInputError): validate_url("ftp://example.com/file") diff --git a/tests/test_server.py b/tests/test_server.py index 87cbd40b..a8608cdc 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -2675,18 +2675,6 @@ def test_listing_assets_without_a_library_is_empty_not_an_error(server): assert body["shadowed"] == [] -def test_an_audio_file_can_be_uploaded(asset_server, tmp_path): - """A workflow's audio reference is built from a .wav - refusing it would - leave one input kind with no way onto the machine.""" - with asset_server(success_script) as client: - response = client.post( - "/api/uploads", params={"filename": "voice.wav"}, content=b"riff" - ) - assert response.status_code == 201 - assert response.json()["path"].startswith("asset:uploads/") - assert response.json()["path"].endswith(".wav") - - def test_the_asset_library_is_reported(asset_server, tmp_path): with asset_server(success_script) as client: directories = client.get("/api/server").json()["directories"] @@ -4000,25 +3988,6 @@ def test_event_log_clamps_a_negative_after(server): assert body["events"][0]["seq"] == 0 -def test_event_log_serves_a_historical_jobs_persisted_events(server): - """A job recovered from sqlite is a plain dict, but its event tail was - persisted with it - that is what makes last night's failure explainable.""" - with server(success_script) as client: - manager = client.app.state.job_manager - manager.get = lambda job_id: {"id": job_id, "status": "failed"} - manager.history.events_for = lambda job_id: [ - {"seq": 0, "event": "phase", "phase": "loading"}, - {"seq": 1, "event": "job_status", "status": "failed"}, - ] - - body = client.get("/api/jobs/historical/event-log").json() - - assert [event["seq"] for event in body["events"]] == [0, 1] - assert body["last_seq"] == 1 - assert body["truncated"] is False - assert body["note"] is None - - def test_event_log_pages_a_historical_jobs_events(server): with server(success_script) as client: manager = client.app.state.job_manager @@ -4111,6 +4080,7 @@ def test_a_recorded_job_reads_back_through_the_event_log_route(server): body = client.get("/api/jobs/recorded/event-log").json() assert [event["seq"] for event in body["events"]] == [0, 1] + assert body["last_seq"] == 1 assert body["events"][0]["phase"] == "loading" assert body["status"] == "complete" assert body["truncated"] is False @@ -4136,21 +4106,6 @@ def test_a_recorded_job_whose_log_was_dropped_says_so_through_the_route(server): assert f"last {MAX_PERSISTED_EVENTS}" in body["note"] -def test_workflow_details_name_their_variables(server): - """The listing says which knobs a workflow takes, so an agent picking a - workflow to run knows what to pass without fetching each candidate's - full definition. Names only - the defaults of every workflow on disk - are an order of magnitude more payload on a listing the UI reloads.""" - with server(success_script) as client: - workflow = valid_workflow("knobby") - workflow["variables"] = {"prompt": "a cat", "steps": 25} - client.put("/api/workflows/Knobby", json={"workflow": workflow}) - - details = client.get("/api/workflows").json()["details"] - assert details["Knobby"]["variable_names"] == ["prompt", "steps"] - assert details["Knobby"]["variables"] == 2 - - def test_workflow_details_describe_their_lists(server): """A list-driven workflow's listing says what an entry carries, so an agent can write the list without opening the definition.""" diff --git a/tests/test_server_info.py b/tests/test_server_info.py index 68e28f00..7aaf3e93 100644 --- a/tests/test_server_info.py +++ b/tests/test_server_info.py @@ -102,6 +102,8 @@ def test_prompt_dir_may_be_absent(tmp_path): def test_auth_required_and_token_never_disclosed(tmp_path): token = "s3cr3t-token-value" with client(tmp_path, token=token) as c: + # gated like every other API route + assert c.get("/api/server").status_code == 401 response = c.get("/api/server", headers={"Authorization": f"Bearer {token}"}) assert response.status_code == 200 body = response.json() @@ -127,15 +129,6 @@ def numbers(value): assert len(token) not in numbers(body) -def test_requires_the_token_like_every_other_api_route(tmp_path): - with client(tmp_path, token="abc123") as c: - assert c.get("/api/server").status_code == 401 - assert ( - c.get("/api/server", headers={"Authorization": "Bearer abc123"}).status_code - == 200 - ) - - def test_mcp_mounted_reported(tmp_path): pytest.importorskip("mcp", reason="the mcp extra is not installed") with client(tmp_path, mcp=True) as c: @@ -175,12 +168,24 @@ def test_netinfo_falls_back_to_stdlib_without_psutil(monkeypatch): def no_psutil(): raise ImportError("no psutil") + import socket + + def fake_getaddrinfo(host, port): + return [ + (socket.AF_INET6, socket.SOCK_STREAM, 0, "", ("2001:db8::5%eth0", 0, 0, 0)), + (socket.AF_INET, socket.SOCK_STREAM, 0, "", ("127.0.0.1", 0)), + (socket.AF_INET, socket.SOCK_STREAM, 0, "", ("192.168.1.50", 0)), + ] + monkeypatch.setattr(netinfo, "_psutil_addresses", no_psutil) - entries = netinfo.local_addresses() - assert isinstance(entries, list) - for entry in entries: - assert entry["interface"] is None - assert entry["family"] in ("IPv4", "IPv6") + monkeypatch.setattr(netinfo.socket, "getaddrinfo", fake_getaddrinfo) + # the outbound probe finds an address getaddrinfo already had: reported once + monkeypatch.setattr(netinfo, "_outbound_address", lambda: "192.168.1.50") + + assert netinfo.local_addresses() == [ + {"address": "192.168.1.50", "family": "IPv4", "interface": None}, + {"address": "2001:db8::5", "family": "IPv6", "interface": None}, + ] def test_usable_filters(monkeypatch): diff --git a/tests/test_server_workspaces.py b/tests/test_server_workspaces.py index e73d1ca0..420e43a0 100644 --- a/tests/test_server_workspaces.py +++ b/tests/test_server_workspaces.py @@ -660,23 +660,6 @@ def test_a_name_cannot_leave_the_library(self, server, workspace_root, name): class TestRunning: - def test_a_job_runs_in_the_workspace_it_named(self, server, workspace_root): - with server() as client: - client.post("/api/workspaces", json={"name": "shots"}) - client.put( - "/api/workflows/Mine?workspace=shots", - json={"workflow": valid_workflow("mine")}, - ) - response = client.post( - "/api/jobs", json={"workflow_path": "Mine", "workspace": "shots"} - ) - assert response.status_code == 201 - detail = wait_for_status( - client, response.json()["id"], {"succeeded", "failed"} - ) - - assert detail["status"] == "succeeded" - def test_enhance_runs_in_the_selected_workspace(self, server): """The enhance job used to be submitted unscoped, so its text landed in the default workspace's outputs while the editor read it back @@ -728,10 +711,14 @@ def test_rerun_stays_in_the_workspace_it_ran_in(self, server, workspace_root): "/api/workflows/Mine?workspace=shots", json={"workflow": valid_workflow("mine")}, ) - original = client.post( + submitted = client.post( "/api/jobs", json={"workflow_path": "Mine", "workspace": "shots"} - ).json() - wait_for_status(client, original["id"], {"succeeded", "failed"}) + ) + # a stored workflow found only in the named workspace runs there + assert submitted.status_code == 201 + original = submitted.json() + detail = wait_for_status(client, original["id"], {"succeeded", "failed"}) + assert detail["status"] == "succeeded" rerun = client.post(f"/api/jobs/{original['id']}/rerun") assert rerun.status_code == 201 diff --git a/tests/test_speech_generation.py b/tests/test_speech_generation.py index dc8f2a77..88233a76 100644 --- a/tests/test_speech_generation.py +++ b/tests/test_speech_generation.py @@ -301,23 +301,6 @@ def test_messages_that_pass_the_shape_guard_still_reach_the_pipeline( device="cpu", messages=[{"role": "narrator", "content": "hi"}] ) - @patch("dw.tasks.speech_generation.hf_pipeline") - def test_messages_on_a_model_with_no_chat_template_propagates_the_error( - self, mock_pipeline - ): - # Bark and other non-chat-templated models have no apply_chat_template - # to call; transformers' own failure surfaces rather than dw silently - # falling back to treating messages as text - pipe = MagicMock(side_effect=ValueError("no chat template is set")) - mock_pipeline.return_value = pipe - - with self.assertRaisesRegex(ValueError, "no chat template is set"): - generate_speech( - device="cpu", - model_name=_DEFAULT_MODEL, - messages=[{"role": "user", "content": "hi"}], - ) - @patch("dw.tasks.speech_generation.hf_pipeline") def test_a_seed_seeds_the_global_rng_before_generating(self, mock_pipeline): # transformers' generate() takes no generator= kwarg, unlike a diffusers diff --git a/tests/test_strip_exif_and_watermark.py b/tests/test_strip_exif_and_watermark.py index 26425e60..4ead8b66 100644 --- a/tests/test_strip_exif_and_watermark.py +++ b/tests/test_strip_exif_and_watermark.py @@ -1,6 +1,8 @@ """Tests for strip_exif and add_watermark image processing commands.""" import unittest + +import numpy as np from PIL import Image from PIL.PngImagePlugin import PngInfo @@ -61,15 +63,20 @@ def test_modifies_pixels(self): self.assertNotEqual(img.tobytes(), result.tobytes()) def test_default_text(self): - # Should not raise with defaults img = Image.new("RGB", (400, 200)) result = add_watermark(img) - self.assertIsInstance(result, Image.Image) + # The defaults draw "AI Generated", pixel for pixel + expected = add_watermark(img, text="AI Generated") + other = add_watermark(img, text="SOMETHING ELSE") + self.assertNotEqual(result.tobytes(), img.tobytes()) + self.assertEqual(result.tobytes(), expected.tobytes()) + self.assertNotEqual(result.tobytes(), other.tobytes()) def test_custom_text(self): img = Image.new("RGB", (400, 200)) result = add_watermark(img, text="DO NOT DISTRIBUTE") - self.assertIsInstance(result, Image.Image) + self.assertNotEqual(result.tobytes(), img.tobytes()) + self.assertNotEqual(result.tobytes(), add_watermark(img).tobytes()) def test_all_positions(self): img = Image.new("RGB", (400, 200)) @@ -88,13 +95,24 @@ def test_invalid_position_falls_back(self): def test_custom_color(self): img = Image.new("RGB", (400, 200)) - result = add_watermark(img, color=(255, 0, 0)) - self.assertIsInstance(result, Image.Image) + result = add_watermark(img, color=(255, 0, 0), opacity=255) + # Red text on black: every drawn pixel is some shade of pure red + pixels = np.asarray(result).reshape(-1, 3) + drawn = pixels[pixels.any(axis=1)] + self.assertTrue(len(drawn)) + self.assertFalse(drawn[:, 1:].any()) + self.assertEqual(drawn[:, 0].max(), 255) def test_custom_font_size(self): img = Image.new("RGB", (400, 200)) - result = add_watermark(img, font_size=24) - self.assertIsInstance(result, Image.Image) + + def inked(font_size): + result = add_watermark(img, text="W", font_size=font_size, opacity=255) + return int(np.asarray(result).any(axis=2).sum()) + + # the auto size here is max(12, 200 // 30) = 12; a larger size draws more + self.assertGreater(inked(48), inked(24)) + self.assertGreater(inked(24), inked(0)) def test_rgba_input_converted(self): img = Image.new("RGBA", (200, 100)) @@ -103,8 +121,13 @@ def test_rgba_input_converted(self): def test_dispatch_via_process_image(self): img = Image.new("RGB", (200, 100)) - result = process_image(img, "add_watermark", "cpu", {"text": "TEST"}) - self.assertIsInstance(result, Image.Image) + result = process_image( + img, "add_watermark", "cpu", {"text": "TEST", "opacity": 255} + ) + # the kwargs reach add_watermark: same pixels as a direct call + expected = add_watermark(img, text="TEST", opacity=255) + self.assertEqual(result.tobytes(), expected.tobytes()) + self.assertNotEqual(result.tobytes(), img.tobytes()) if __name__ == "__main__": diff --git a/tests/test_task.py b/tests/test_task.py index fe084a5d..89d34a53 100644 --- a/tests/test_task.py +++ b/tests/test_task.py @@ -83,14 +83,6 @@ def test_gather_images_task_dispatches_to_gather(): assert [call.args[0] for call in load_image.call_args_list] == urls -def test_gather_inputs_task(): - task_def = {"command": "gather_inputs", "inputs": ["value1", "value2"]} - task = Task(task_def, "cpu") - result = task.run(task_def["inputs"]) - assert isinstance(result, list), "Expected a list of inputs from gather_inputs" - assert "value1" in result and "value2" in result, "Should gather all passed inputs" - - def test_format_chat_message_task(): task_def = { "command": "format_chat_message", diff --git a/tests/test_teacache.py b/tests/test_teacache.py index 644ab154..d7faf668 100644 --- a/tests/test_teacache.py +++ b/tests/test_teacache.py @@ -101,16 +101,6 @@ def _call(bound_forward, timestep_value): ) -def test_duplicate_timestep_raises_runtime_error(): - """Two forward calls with the identical timestep (true CFG) must raise.""" - _, bound_forward = _make_bound_forward() - - _call(bound_forward, 0.9) # first call: no prior timestep, always allowed - - with pytest.raises(RuntimeError, match="true classifier-free guidance"): - _call(bound_forward, 0.9) # duplicate timestep: simulates the uncond pass - - def test_duplicate_timestep_error_names_both_features(): """The guard's error message must name both TeaCache and true CFG.""" _, bound_forward = _make_bound_forward() diff --git a/tests/test_tensor_image.py b/tests/test_tensor_image.py index fb6bfd5d..c145742c 100644 --- a/tests/test_tensor_image.py +++ b/tests/test_tensor_image.py @@ -123,12 +123,6 @@ def test_near_half_maps_by_rounding_not_floor(self): assert truncated == 127 assert pixel[0] != truncated - def test_half_value_rounds_up_not_down(self): - """0.5/255-scaled exact half (127.5) rounds to nearest even (128).""" - tensor = torch.full((1, 3, 1, 1), 127.5 / 255.0) - image = float_tensor_to_pil(tensor) - assert image.getpixel((0, 0)) == (128, 128, 128) - class TestModuleAdoption: """Verify upscale.py and interpolate_frames.py actually route through the diff --git a/tests/test_text_generation.py b/tests/test_text_generation.py index e8a5173f..e7fe2327 100644 --- a/tests/test_text_generation.py +++ b/tests/test_text_generation.py @@ -336,14 +336,5 @@ def test_other_devices_still_use_device_map(self, mock_pipeline): self.assertNotIn("device", kwargs) -class TestTextGenerationRegistration(unittest.TestCase): - """Test that text_generation is registered as a task command.""" - - def test_command_registered(self): - from dw.tasks.task import _COMMAND_REGISTRY - - self.assertIn("text_generation", _COMMAND_REGISTRY) - - if __name__ == "__main__": unittest.main() diff --git a/tests/test_type_helpers.py b/tests/test_type_helpers.py index 1ba05ee9..8a36f030 100644 --- a/tests/test_type_helpers.py +++ b/tests/test_type_helpers.py @@ -15,11 +15,9 @@ class TestGetType: """Test getting type from module""" def test_get_type_from_diffusers(self): - # This would work if diffusers is installed - # For testing, we'll use a built-in type + import diffusers - result = get_type("sys", "version") - assert result is not None + assert get_type("diffusers", "DiffusionPipeline") is diffusers.DiffusionPipeline def test_get_type_invalid_module(self): with pytest.raises(ModuleNotFoundError): diff --git a/tests/test_workflow.py b/tests/test_workflow.py index 35f7feac..76f58892 100644 --- a/tests/test_workflow.py +++ b/tests/test_workflow.py @@ -32,29 +32,12 @@ def test_workflow_validation_invalid(invalid_workflow_json, tmp_path): assert "Validation error" in str(exc_info.value) -def test_workflow_name(valid_workflow_json, tmp_path): - workflow = Workflow(valid_workflow_json, str(tmp_path), "") - assert workflow.name == "test_workflow" - - def test_workflow_from_file(test_data_dir, tmp_path): workflow_path = os.path.join(test_data_dir, "workflows", "valid_workflow.json") workflow = workflow_from_file(workflow_path, str(tmp_path)) assert isinstance(workflow, Workflow) -def test_workflow_variables_property(valid_workflow_json, tmp_path): - workflow = Workflow(valid_workflow_json, str(tmp_path), "") - assert "prompt" in workflow.variables - assert workflow.variables["prompt"] == "test prompt" - - -def test_workflow_argument_template(valid_workflow_json, tmp_path): - workflow = Workflow(valid_workflow_json, str(tmp_path), "") - # Should return empty dict if no argument_template - assert workflow.argument_template == {} - - def test_workflow_security_validation(tmp_path): from dw.security import SecurityError @@ -409,7 +392,7 @@ def test_a_parent_directory_step_inside_the_root_is_allowed(self, tmp_path): device="cpu", ) - assert action is not None + assert action.name == "child" def test_a_parent_directory_step_escaping_the_root_is_refused(self, tmp_path): import json @@ -523,7 +506,7 @@ def test_an_unconfined_run_may_climb_to_a_sibling_catalog_folder(self, tmp_path) device="cpu", ) - assert action is not None + assert action.name == "child" def test_a_file_outside_any_catalog_is_confined_to_its_own_directory( self, tmp_path @@ -1041,12 +1024,6 @@ def test_a_name_that_resolves_nowhere_says_where_it_looked(self, tmp_path): assert "does-not-exist.json" in message assert "outside the root" not in message - def test_a_relative_path_beside_the_file_still_wins(self, tmp_path): - """The '../models/x.json' form every template uses is unchanged.""" - action = self._resolve(tmp_path, "minimax/ref2va.json") - - assert action.name == "child" - class TestComposedStepSavesOnce: """A sub-workflow step that declares a result owns the file: the child's diff --git a/tests/test_workspace.py b/tests/test_workspace.py index 2e79e9b7..6c45ea87 100644 --- a/tests/test_workspace.py +++ b/tests/test_workspace.py @@ -189,11 +189,6 @@ class TestSharedAssetLibrary: the prompt library's treatment applied to a recurring cast, which belongs to no one workspace.""" - def test_it_hangs_off_the_root(self, tmp_path): - root = Workspace(tmp_path / "studio", FLAG) - - assert root.common_assets == os.path.join(root.root, "common", "assets") - def test_a_named_workspace_points_back_at_the_root_s(self, tmp_path): from dw.workspace import named_workspace From ea8a0b6a8d283eea063c6dbef72880a595558859 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 07:09:34 -0500 Subject: [PATCH 035/181] fix(engine): resolve previous_result: references inside a task's inputs list get_iterations handed an 'inputs' list template back untouched, so "inputs": ["previous_result:step1"] reached the task as the literal string - while previous_result_reference_errors already validated the same reference as real, so nothing warned. A reference entry now expands to one iteration per result it names, and an object entry expands the way an 'arguments' template does; other entries pass through as before. The list is capped at MAX_ITERATIONS like the cartesian product. Found by the test-suite review: test_simple_dependency passed while its second step received the reference string. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/WORKFLOW_GUIDE.md | 4 +++- dw/previous_results.py | 26 +++++++++++++++++++++++--- tests/test_integration.py | 5 +++-- tests/test_previous_results.py | 22 ++++++++++++++++++++++ 4 files changed, 51 insertions(+), 6 deletions(-) diff --git a/docs/WORKFLOW_GUIDE.md b/docs/WORKFLOW_GUIDE.md index 895b34cd..4a82ef6d 100644 --- a/docs/WORKFLOW_GUIDE.md +++ b/docs/WORKFLOW_GUIDE.md @@ -124,7 +124,9 @@ Run utility operations (image processing, QR codes, data gathering): ``` A task can take `inputs` (a plain array) instead of `arguments`. Each array item becomes -its own iteration, the same way multiple `previous_result` values do: +its own iteration, the same way multiple `previous_result` values do. An item that is a +`previous_result:` reference becomes one iteration per result it names, and an object item +expands the way an `arguments` object would: ```json { diff --git a/dw/previous_results.py b/dw/previous_results.py index 3a093f37..2dc97c7f 100644 --- a/dw/previous_results.py +++ b/dw/previous_results.py @@ -28,10 +28,30 @@ def get_iterations(argument_template, previous_results): Returns: List of argument dictionaries, one for each possible combination """ - # Special case: if template is a list, use it directly without processing + # An 'inputs' list is already one iteration per entry. The static check + # (previous_result_reference_errors) accepts a reference inside one, so the + # run must resolve it too rather than hand the task the literal string: + # a reference entry is one iteration per result it names, and an object + # entry expands exactly as an 'arguments' template would if isinstance(argument_template, list): - logger.debug("Using list argument template directly") - return argument_template + iterations = [] + for entry in argument_template: + if isinstance(entry, str) and entry.startswith(PREVIOUS_RESULT_PREFIX): + iterations.extend( + get_previous_results( + previous_results, entry[len(PREVIOUS_RESULT_PREFIX) :] + ) + ) + elif isinstance(entry, dict): + iterations.extend(get_iterations(entry, previous_results)) + else: + iterations.append(entry) + if len(iterations) > MAX_ITERATIONS: + raise ValueError( + f"Too many iterations generated: more than {MAX_ITERATIONS} " + f"from an 'inputs' list. Consider splitting it across steps." + ) + return iterations # Find any references to previous results in the template # Returns dict of {arg_key: result_reference} diff --git a/tests/test_integration.py b/tests/test_integration.py index 9fd81760..db9cb7a2 100644 --- a/tests/test_integration.py +++ b/tests/test_integration.py @@ -265,7 +265,7 @@ def test_simple_dependency(self, temp_workflow_dir): "name": "step2", "task": { "command": "gather_inputs", - "arguments": {"value": "previous_result:step1"}, + "inputs": ["previous_result:step1"], }, "result": {"content_type": "application/json", "save": False}, }, @@ -277,7 +277,8 @@ def test_simple_dependency(self, temp_workflow_dir): result = workflow.run({}) # step2 runs once per result step1 produced, receiving each one - assert result == [{"value": "value1"}, {"value": "value2"}] + # rather than the literal reference string + assert result == ["value1", "value2"] if __name__ == "__main__": diff --git a/tests/test_previous_results.py b/tests/test_previous_results.py index d8a29b7e..e91b67cd 100644 --- a/tests/test_previous_results.py +++ b/tests/test_previous_results.py @@ -269,6 +269,28 @@ def test_list_template_returns_as_is(self): iterations = get_iterations(template, previous_results) assert iterations == template + def test_a_reference_entry_in_a_list_template_expands_to_its_results(self): + """An 'inputs' list is one iteration per entry, so a reference entry + is one iteration per result it names - not the literal string the + static check already accepts as a reference.""" + step1 = Result({}) + step1.add_result(["a", "b"]) + + iterations = get_iterations( + ["first", "previous_result:step1", "last"], {"step1": step1} + ) + assert iterations == ["first", "a", "b", "last"] + + def test_a_reference_inside_a_list_templates_object_resolves(self): + step1 = Result({}) + step1.add_result(["a", "b"]) + + iterations = get_iterations( + [{"value": "previous_result:step1"}, {"value": "plain"}], + {"step1": step1}, + ) + assert iterations == [{"value": "a"}, {"value": "b"}, {"value": "plain"}] + def test_three_way_cartesian_product(self): """Test with 3 dimensions: 2x2x2 = 8 combinations""" result1 = Result({}) From 96e2ea2209dbc59f3ca597727d6b4b7176f0c19e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 07:33:26 -0500 Subject: [PATCH 036/181] fix(plugin): #340, #354 - score ducking recipe + transcription route in skills #340: minimax-h3, minimax-music3 and series-episodes skills now teach ducking the *score* under a voice-over shot with gain_audio regions (one per voice-over shot, start_frame = running sum of preceding shots' num_frames, applied before the track is passed as `score`), rather than raising `world_gain`, which lifts narration and action sound together. Names the measured symptom: narration ~24dB over its own ambience, score 6-15dB over that. #354: get_output_audio's docstring, and the minimax-h3/series-episodes "Run and judge" sections, now point at the two-call transcription route (run_workflow(name="templates/transcribe-audio", ...) then get_output_text) instead of implying a new read tool. Verified live against dw.serve on both a plain wav and an H3 mp4 with a muxed soundtrack - both transcribed correctly with no code changes needed. minimax-h3/SKILL.md trimmed for the 12KB skill size cap while adding the ducking recipe and transcription pointer. Co-Authored-By: Claude Sonnet 5 --- dw_mcp/server.py | 10 +- plugins/dw/skills/minimax-h3/SKILL.md | 149 ++++++++++----------- plugins/dw/skills/minimax-music3/SKILL.md | 5 +- plugins/dw/skills/series-episodes/SKILL.md | 13 +- 4 files changed, 98 insertions(+), 79 deletions(-) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 447a6f4c..12a6267e 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -526,7 +526,15 @@ def get_output_audio( excerpted. No downscale exists for audio - a whole clip too large is refused; ask for a part with `start`/`duration` in seconds, per `get_gallery_metadata`'s envelope. The text part - says what was cut. To *see* a video, `get_output_frames`. + says what was cut. To *see* a video, `get_output_frames`. To + confirm the *words* an output speaks rather than hear it - a + text-only client can't consume the `AudioContent` block this + returns - run `run_workflow(name="templates/transcribe-audio", + arguments={"input_audio": "output:"}, wait_seconds=55)` (it + takes an audio file or a video's muxed soundtrack directly) then + `get_output_text` on the result, and `delete_output` the scratch + run afterward. Two calls and a short wait, not a GPU-spending read + tool - keep the normal queue rather than adding one. `workspace` pins this call to another workspace.""" result = media.get_output_audio( diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 3dfa30a2..7f00e92d 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -1,22 +1,21 @@ --- name: minimax-h3 -description: Use when a dw MCP server is connected and the user wants MiniMax H3 video - a clip with its own sound, dialogue or a speaking character, a music video, a multi-shot short with cuts, a longer take, or a subject or voice kept consistent from a reference. Picks the template for the shape, states the rules that bite, quotes cost, and points at MiniMax's own prompt guides. +description: Use when a dw MCP server is connected and the user wants MiniMax H3 video - a clip with its own sound, dialogue or a speaking character, a music video, a multi-shot short with cuts, a longer take, or a subject/voice kept consistent from a reference. Picks the template, states the rules that bite, quotes cost, points at MiniMax's own prompt guides. --- # MiniMax H3 on a dw server H3 generates video and audio together: speech with lip sync, ambient sound, -score. Every template here fits a 24 GB card. This skill chooses the template -and the arguments; the prompt format is MiniMax's, from their text not here. +score. Every template fits a 24 GB card. This skill chooses the template and +arguments; the prompt format is MiniMax's, from their text not here. ## Before anything -1. `get_server_info`: the device (H3 templates are CUDA; the quantized - configurations do not run on mps) and which workspace this session is in. +1. `get_server_info`: the device (H3 is CUDA; quantized configurations do + not run on mps) and which workspace this session is in. 2. `list_workflows(shape="shot")`, `list_workflows(shape="sequence")` and - `list_workflows(shape="audio")`: the family's templates by their current - names, with `summary`, `traits` and `cost`. Trust the listing over the - names quoted below. + `list_workflows(shape="audio")`: the family's templates by current name, + `summary`, `traits` and `cost`. Trust the listing over names quoted below. 3. `get_workflow` on the one chosen, for its variables and their defaults. ## Which shape is the request @@ -38,20 +37,19 @@ and the arguments; the prompt format is MiniMax's, from their text not here. `templates/minimax/voice-timbre-reference` fixes a voice from a Bark line. - **Several boards in one generation, one unbroken score**: `templates/minimax/storyboard` - H3 cuts between the boards inside a single - generation, which no concat of separate clips can match for continuous audio. - It is one beat with fixed cut points, not a building block; past one beat - with a recurring cast, use the cuts pattern below. -- **Longer than 14.4 seconds**: decide first whether the seam is a cut or a - continuation. Chain when one action or line of speech crosses the seam; cut - when the scene changes, and treat each cut as its own generation. Six - distinct scenes are a cuts piece, not a chain. -- **Longer than 14.4 seconds as one take**: a chain. `templates/minimax/chained-segments` + generation, which no concat of separate clips can match. One beat with fixed + cut points, not a building block; past one beat with a recurring cast, use + the cuts pattern below. +- **Longer than 14.4 seconds**: chain when one action or line of speech + crosses the seam, cut when the scene changes (each cut its own generation). + Six distinct scenes are a cuts piece, not a chain. +- **As one take (a chain)**: `templates/minimax/chained-segments` (last-frame continuity), `templates/minimax/chain-video-continuity` (the - previous segment's tail rides along as a video reference - motion, camera and + previous segment's tail rides as a video reference - motion, camera and voice carry across the seam), `templates/minimax/chain-matched-to-audio` (a - supplied track sets the length and is muxed back seamless), + supplied track sets the length, muxed back seamless), `templates/minimax/chain-matched-and-aligned` (all of it, per-segment - prompts). Drift compounds per seam: reference the subject picture in every + prompts). Drift compounds per seam: reference the subject picture every segment, prefer `last_segment` continuity, use the longest segments memory allows. - **A piece with cuts**: fresh shots from shared portraits, then a concat. @@ -71,17 +69,22 @@ and the arguments; the prompt format is MiniMax's, from their text not here. for N entries. Without `per_entry`, quote the total and say it is the default list's. A cut erases drift: the last shot is as clean as the first. Each shot makes - its own audio, so write - `non_diegetic_music: N/A` in every shot and lay one score under the concat - afterwards: `templates/minimax/music` writes the track and - `templates/assemble-and-score` shows the `pair_audio` step that mixes it - under the world sound (it takes three shots; for more, author the concat + its own audio, so write `non_diegetic_music: N/A` in every shot and lay one + score under the concat afterwards: `templates/minimax/music` writes the + track and `templates/assemble-and-score` shows the `pair_audio` step that + mixes it under the world sound (three shots; for more, author the concat and score steps the same way). A character speaking in several shots keeps one voice by passing the same clip as an audio reference in each (the - `voice-timbre-reference` pattern); a repeated description alone drifts. - That clip carries delivery as well as timbre, and Bark's presets are - conversational - for gravitas, `upload_asset` a read in that register. + `voice-timbre-reference` pattern) - a repeated description alone drifts. Each entry's `num_frames` is its own, so pace the cut. + A score burying voice-over is not a `world_gain` fix: the world track + carries narration and action together (narration ~24 dB over its own + ambience, score 6-15 dB over that), so raising it lifts both. Duck the + *score* instead, before passing it as `score`: one `gain_audio` (#187) + step per voice-over shot, chained, `start_frame` = running sum of + preceding shots' `num_frames`, `num_frames` that shot's length, `fps` + the cut's rate, negative `gain_db` (-6 to -10). `dissolve-between-shots` + eats a `dissolve_frames` per seam, so its sum isn't plain. - **Music alone**: `templates/minimax/music` (Music3); the `minimax-music3` skill. If none fits, compose from `list_tasks` before authoring a new workflow, and @@ -108,23 +111,20 @@ read the `workflows` guide's authoring section first. so it only degrades the output. `validate_workflow` refuses it and warns on a `weight_name` naming neither path. - Nine steps for an eight-step LoRA: the scheduler counts sigma grid points, - terminal zero included, so `denoise_total_steps` reports 8. Expected; do not - read upstream's `--inference-steps 8` literally. -- Nothing carries between generations except what is passed as a reference: - no latent memory and no extension mode, in the checkpoint, the API or - diffusers. Identity rides on a picture, voice on an audio clip, motion and - camera on a video tail (what a chain passes forward), and a score across - cuts is laid under the concat. -- H3 is guidance-distilled: no `guidance_scale`, no negative prompt. Say what - is there, not what is not. -- Keep `release_pipeline` where the template puts it: it frees Z-Image before - H3 loads, and frees H3 itself on `shot` before a concat runs. A run - SIGKILLed near the end in a warm worker but fine in a fresh one is host - memory, not the prompt. -- Ref2VA limits: at most 9 images, 3 videos, 3 audio clips, 12 files; - audio can never be the only reference. References are labelled in order. + terminal zero included, so `denoise_total_steps` reports 8 - expected, not + upstream's `--inference-steps 8` read literally. +- Nothing carries between generations except a passed reference: no latent + memory, no extension mode. Identity rides on a picture, voice on an audio + clip, motion/camera on a video tail (a chain's), score across cuts under + the concat. +- H3 is guidance-distilled: no `guidance_scale`, no negative prompt. +- Keep `release_pipeline` where the template puts it: frees Z-Image before H3 + loads, and H3 itself on `shot` before a concat runs, or a warm worker can + SIGKILL near the end where a fresh one would not. +- Ref2VA limits: at most 9 images, 3 videos, 3 audio clips, 12 files; audio can + never be the only reference. References are labelled in order. - Music3 reads `audio_duration` as a ceiling, not a target: ask for more than - the song needs and trim with `templates/audio-trim-fade`. + needed and trim with `templates/audio-trim-fade`. - Write the prompt for the length generated: timestamps should span the duration, or a five-second script tells a five-second story whatever the frame count. @@ -149,12 +149,10 @@ the rules come from MiniMax, not from paraphrasing one: built-in enhancer writes the format from those guides. Its `idea` is framed as `Task: T2VA. Duration: 5.17 seconds. Idea: ...`. -Whichever route: write the whole script before the first shot - the lines in -order, read once, should carry the piece - then place them. Repeat a speaker's -voice description verbatim across shots, and when a reference picture should -fix identity but not framing, say so in the prompt itself, in the lines that -define the subject and what each reference keeps - or every shot inherits the -portrait's composition. +Whichever route: write the whole script before the first shot, then place the +lines. Repeat a speaker's voice description verbatim across shots, and when a +reference picture should fix identity but not framing, say so in the prompt +itself - or every shot inherits the portrait's composition. ## Run and judge @@ -168,36 +166,37 @@ portrait's composition. Get the go-ahead, then `run_workflow` with `acknowledged_cost` = the plan's `{fingerprint, minutes, downloads}`. 3. `wait_for_job`, then `get_job` for the manifest. A cancelled H3 job runs - on to its next step boundary, minutes on this model. Silence is no hang: - `denoise_step` is null through the reference encode (~90 s; 629 s for one - 5 s 960x544 video reference on a 3090), and the block cache makes later - steps uneven - two-minute gaps are healthy. `phase_stall` in - `get_job_events` narrates it, not a fault; judge by `denoise_step`. - Each entry carries - `subfolder`: `final` is the deliverable (`episode`, `music_video`, - `voyage`), `intermediate` the scratch; keep that split in anything you - compose. + on to its next step boundary. Silence is no hang: `denoise_step` is null + through the reference encode (~90 s; 629 s for one 5 s 960x544 video + reference on a 3090), and the block cache makes later steps uneven - + two-minute gaps are healthy. `phase_stall` in `get_job_events` narrates + it, not a fault; judge by `denoise_step`. Each entry carries `subfolder`: + `final` is the deliverable (`episode`, `music_video`, `voyage`), + `intermediate` the scratch; keep that split in anything you compose. 4. Judge it yourself: `get_output_frames(count=12)` for a clip's shape, - `seams=true` (with each later shot's start frame) for a cut's joins - a character - that changes between shots (reference the same portraits everywhere), a - portrait imposing its framing on every shot - `at` late in a chain for drift - sharpening into noise, and `get_output_audio` for a voice-over without - affect (the reference's delivery came through). Then `get_job` for the - manifest and its warnings, `get_gallery_metadata` for duration and whether - audio is present, and hand the user the gallery `url` (`list_gallery`). - `get_output_image` works only on image steps - the Z-Image portraits and - boards of `dialogue-short`, `storyboard`, `generated-subject-reference` and - `music-video`. -5. After a run worth keeping, `get_job_workflow` and `save_workflow` it, so - the next run is by name not pasted JSON; `export_job` bundles it on the - server. `auth_required: false` - fetch `open_url` into `exports/` under - the working directory (never a temp dir; it unpacks into a job-id + `seams=true` (each later shot's start frame) for a cut's joins - a + character that changes between shots (reference the same portraits + everywhere), a portrait imposing its framing on every shot - `at` late in + a chain for drift sharpening into noise, and `get_output_audio` for a + voice-over without affect. `get_output_audio` returns sound, not text; to + confirm a line rendered rather than judge its delivery, + `run_workflow(name="templates/transcribe-audio", + arguments={"input_audio": "output:"}, wait_seconds=55)` (takes the + muxed soundtrack directly) then `get_output_text` on the result, and + `delete_output` the scratch run. Then `get_gallery_metadata` for duration + and whether audio is present, and hand the user the gallery `url` + (`list_gallery`). `get_output_image` works only on image steps - the + Z-Image portraits and boards of `dialogue-short`, `storyboard`, + `generated-subject-reference` and `music-video`. +5. After a run worth keeping, `get_job_workflow` and `save_workflow` it; + `export_job` bundles it on the server. `auth_required: false` - fetch + `open_url` into `exports/` (never a temp dir; unpacks into a job-id folder). `true` - hand `open_url` to the person instead and keep using `get_output_image`/`_audio`/`_frames`. ## Sources MiniMax-H3 model card and prompt guides (huggingface.co/MiniMaxAI/MiniMax-H3), -the `h3-prompt-writing` skill (github.com/MiniMax-AI/MiniMax-H3), the diffusers -MiniMax-H3 modular pipeline, the `lightx2v/Minimax-h3-Turbo` LoRA notes; read -2026-09-07; the audit is `docs/proposals/audits/2026-09-07-minimax-h3-audit.md`. +`h3-prompt-writing` (github.com/MiniMax-AI/MiniMax-H3), the diffusers +MiniMax-H3 modular pipeline, `lightx2v/Minimax-h3-Turbo` LoRA notes; read +2026-09-07; audit `docs/proposals/audits/2026-09-07-minimax-h3-audit.md`. diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 840973af..3cb6f0c8 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -46,7 +46,10 @@ shapes; do not author a new workflow until the shape decision below fails. ceiling comfortably longer than the cut, trimmed and faded with `templates/audio-trim-fade`, then mixed under the picture the way `templates/assemble-and-score` does with `pair_audio`. Each H3 shot should - have written `non_diegetic_music: N/A` so the two scores do not fight. + have written `non_diegetic_music: N/A` so the two scores do not fight. A + score that buries a shot's voice-over is not a `world_gain` fix - see the + `minimax-h3` skill's ducking recipe: `gain_audio` regions on the score + itself, one per voice-over shot, applied before it is passed as `score`. - **A music video**: `templates/minimax/music-video`. The song is written first, `slice_audio` deals frame-exact pieces to lip-synced H3 shots, and `pair_audio` lays the unbroken track back over the edit. The ceiling must diff --git a/plugins/dw/skills/series-episodes/SKILL.md b/plugins/dw/skills/series-episodes/SKILL.md index 1b8cb21e..b51ab7bb 100644 --- a/plugins/dw/skills/series-episodes/SKILL.md +++ b/plugins/dw/skills/series-episodes/SKILL.md @@ -86,7 +86,11 @@ composes into, not something to re-derive: - **match_levels**: shots generated independently drift in loudness - `assemble-and-score`'s `match_levels` (`"rms"` or `"peak"`) evens them before the cut; leaving it null only warns on a wide spread instead of - fixing it. + fixing it. A shot whose voice-over the score buries is a different + problem, not fixed by `match_levels` or `world_gain` (which lifts the + whole world track, action sound included) - see `minimax-h3`'s ducking + recipe: `gain_audio` regions on the score, one per voice-over shot, + applied before the score is passed in. - **normalize**: the mixed world sound and score are normalized together (`assemble-and-score`'s `balanced` step, -3 dBFS) so one episode is not louder than the next. @@ -116,7 +120,12 @@ between shots, a portrait imposing its framing, a voice without affect): watch two episodes back to back and check the cast reads as the same people - the failure mode step 0 exists to prevent. `get_gallery_metadata` on each episode's final file for duration and loudness, so a level -mismatch between episodes shows up before a viewer notices it. +mismatch between episodes shows up before a viewer notices it. To confirm a +line actually rendered rather than judging it by ear, `get_output_audio` +returns sound, not text: `run_workflow(name="templates/transcribe-audio", +arguments={"input_audio": "output:"}, wait_seconds=55)` then +`get_output_text` on the result (it takes the episode's muxed soundtrack +directly), and delete the scratch run afterward. ## Sources From ea460a8f0a8ba7cb2518f7836e7e03ae77c9f9ed Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 07:45:32 -0500 Subject: [PATCH 037/181] fix(engine): #341 - bucket a composed child's observed cost by the composing step's arguments observed_for_child was called with only (path, child_definition), so a composed child's observed-history rollup always answered the *default* bucket's figure regardless of what the composing step actually passed for a declared scalar cost_driver - bypassing the #267/#341 shifted-driver check entirely whenever any observed history existed. Thread the composing step's own arguments through dw/plan.py's estimate() into dw/server/app.py's observed_for_child, matching the pattern the top-level observed callback already uses, so the lookup buckets against the value the step runs with and correctly falls through to the catalog + _scalar_driver_shifted unpriced check when no history exists for that bucket. Co-Authored-By: Claude Sonnet 5 --- dw/plan.py | 23 ++++++++---- dw/server/app.py | 14 ++++++-- tests/test_plan.py | 88 ++++++++++++++++++++++++++++++++++++++++++---- 3 files changed, 109 insertions(+), 16 deletions(-) diff --git a/dw/plan.py b/dw/plan.py index 971f5414..77453f91 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -75,12 +75,15 @@ def build_plan( run's arguments and answering one - what the estimate quotes in preference to a curated figure (#154). None on a caller that has no history to offer, which is every caller but the server. - observed_for_child: A callable taking a composed child's local path - and its parsed definition, answering that child's own `observed` - block or None - so a composing workflow's estimate can roll up a - child's history instead of resetting to `unknown` when the - parent has no figure of its own (#268). None on a caller that - cannot resolve a child's catalog name to look history up by. + observed_for_child: A callable taking a composed child's local path, + its parsed definition, and the composing step's own `arguments`, + answering that child's own `observed` block or None - so a + composing workflow's estimate can roll up a child's history + instead of resetting to `unknown` when the parent has no figure + of its own (#268), bucketed against the value the composing step + actually passes rather than always the child's stored defaults + (#341). None on a caller that cannot resolve a child's catalog + name to look history up by. """ definition = candidate.workflow_definition base_dir = ( @@ -499,7 +502,13 @@ def estimate( and child_definition is not None ): try: - child_observed = observed_for_child(path, child_definition) + # `step_arguments` are the composing step's own overrides - + # passed through so a child observed lookup buckets against + # the value this step actually runs with rather than always + # the child's stored defaults (#341) + child_observed = observed_for_child( + path, child_definition, step_arguments + ) except Exception: child_observed = None child = _observed(child_observed, device) diff --git a/dw/server/app.py b/dw/server/app.py index fa37ac99..c7a07899 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -1814,11 +1814,19 @@ def validate_workflow( command = _probe_command_for(candidate, request, workspace, source_root) - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): """A composed child's own observed figure, keyed by the catalog name it resolves to - so a parent with no figure of its own can quote what this box's runs of the *child* took - rather than falling back to unknown (#268).""" + rather than falling back to unknown (#268). + + `arguments` are the composing step's own overrides - the + same role `arguments` plays for the top-level `observed` + callback - so a child whose composing step shifted a + declared scalar `cost_driver` (#341) is bucketed against + *that* value rather than always the child's stored + defaults, which silently answered the default bucket's + history for every override.""" base_dir = ( os.path.dirname(os.path.abspath(candidate.file_spec)) if candidate.file_spec @@ -1841,7 +1849,7 @@ def observed_for_child(path, child_definition): workspace.name if child_root == workspace.workflows else None ) return _observed_for_name( - child_name, child_definition, workspace=child_workspace + child_name, child_definition, arguments, workspace=child_workspace ) answer["plan"] = build_plan( diff --git a/tests/test_plan.py b/tests/test_plan.py index 663e975a..ebe3f7c3 100644 --- a/tests/test_plan.py +++ b/tests/test_plan.py @@ -551,7 +551,7 @@ def test_a_thin_rolled_up_child_is_flagged(self, plan, tmp_path): ], } - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): return observed(minutes=6, runs=1) answer = plan(parent, observed_for_child=observed_for_child)["estimate"] @@ -682,7 +682,7 @@ def test_a_childs_observed_history_rolls_up_when_the_parent_has_none( ], } - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): return observed(minutes=6) answer = plan(parent, observed_for_child=observed_for_child)["estimate"] @@ -706,7 +706,7 @@ def test_the_rolled_up_estimate_carries_the_childs_runs_and_measured_on( ], } - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): return observed(minutes=6, runs=3, name="RTX 3090") answer = plan(parent, observed_for_child=observed_for_child)["estimate"] @@ -727,7 +727,7 @@ def test_the_rolled_up_runs_is_the_weakest_childs(self, plan, tmp_path): ], } - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): runs = 3 if path == "a.json" else 9 return observed(minutes=6, runs=runs, name="RTX 3090") @@ -751,7 +751,7 @@ def test_the_rolled_up_measured_on_is_null_when_children_disagree( ], } - def observed_for_child(path, child_definition): + def observed_for_child(path, child_definition, arguments=None): name = "RTX 3090" if path == "a.json" else "RTX 4090" return observed(minutes=6, runs=3, name=name) @@ -774,11 +774,87 @@ def test_an_unpriced_childs_history_leaves_the_estimate_partial( {"name": "child", "workflow": {"path": "child.json", "arguments": {}}}, ], } - answer = plan(parent, observed_for_child=lambda path, defn: None)["estimate"] + answer = plan( + parent, observed_for_child=lambda path, defn, arguments=None: None + )["estimate"] assert answer["minutes"] is None assert answer["basis"] == "unknown" assert answer["partial"] is False + def test_a_composed_childs_shifted_scalar_driver_stays_unpriced_despite_default_bucket_history( + self, plan, tmp_path + ): + """The observed-history rollup must not paper over a shifted scalar + driver by quoting the *default* bucket's figure just because it + exists - `observed_for_child` is called with the composing step's + own arguments, and answering None for the shifted bucket falls + through to the catalog path's `_scalar_driver_shifted` unpriced + check rather than silently reusing the default bucket's history + (#341).""" + child = { + "id": "child", + "cost": [cost("cuda", 5)], + "cost_drivers": ["num_frames"], + "variables": {"num_frames": 124}, + "steps": [], + } + (tmp_path / "child.json").write_text(json.dumps(child)) + parent = { + "id": "parent", + "steps": [ + { + "name": "child", + "workflow": { + "path": "child.json", + "arguments": {"num_frames": 345}, + }, + }, + ], + } + + calls = [] + + def observed_for_child(path, child_definition, arguments=None): + calls.append(arguments) + if (arguments or {}).get("num_frames") == 124: + return observed(minutes=6, runs=3) + return None + + answer = plan(parent, observed_for_child=observed_for_child)["estimate"] + assert calls == [{"num_frames": 345}] + assert answer["basis"] == "unknown" + assert (answer["minutes"], answer["partial"]) == (None, False) + + def test_a_composed_childs_observed_history_is_bucketed_by_the_composing_steps_arguments( + self, plan, tmp_path + ): + """The companion case: when the composing step's arguments do match + a bucket this box has history for, that bucket's real figure is + quoted rather than always the child's stored-default bucket + (#341).""" + (tmp_path / "child.json").write_text(json.dumps({"id": "child", "steps": []})) + parent = { + "id": "parent", + "steps": [ + { + "name": "child", + "workflow": { + "path": "child.json", + "arguments": {"num_frames": 345}, + }, + }, + ], + } + + def observed_for_child(path, child_definition, arguments=None): + if (arguments or {}).get("num_frames") == 345: + return observed(minutes=9, runs=2) + return observed(minutes=6, runs=5) + + answer = plan(parent, observed_for_child=observed_for_child)["estimate"] + assert answer["basis"] == "observed" + assert answer["minutes"] == 9.0 + class TestDownloadsRequired: def test_a_repo_not_in_the_cache_is_required(self, plan): From fc8c1ea0f2ce53cc1aa50e78357fcf1814ca35a0 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 08:03:44 -0500 Subject: [PATCH 038/181] fix(engine): #343 - heartbeat DownloadWatch on a timer to close the xet reconstruction blind window The prior fix (aa67bf2) only emitted download_progress on detected growth in the hub cache's blob directory, which goes blind during an xet-backed transfer: reconstruction against the local CAS cache can leave the on-disk size unchanged for 20+ seconds at a stretch even while genuinely still running, reproducing the exact phase_stall silence #343 reported. No on-disk location (blob dir, HF_XET_CACHE, the tracked .incomplete file itself) closes that gap by polling alone. DownloadWatch._run() now emits on a fixed cadence regardless of growth, reporting bytes_per_second=0 honestly during a quiet interval instead of staying silent - a heartbeat, not a rate guess. This still only watches; it does not intercept or otherwise change what from_pretrained fetches. Traded away deliberately: a watched download can no longer trip phase_stall for as long as DownloadWatch's thread is alive, including in the case of a true from_pretrained hang with no eventual timeout. Co-Authored-By: Claude Sonnet 5 --- dw/download_watch.py | 34 +++++++++++------- tests/test_download_watch.py | 70 +++++++++++++++++++++++++++++++++--- 2 files changed, 87 insertions(+), 17 deletions(-) diff --git a/dw/download_watch.py b/dw/download_watch.py index c40ced60..a426401a 100644 --- a/dw/download_watch.py +++ b/dw/download_watch.py @@ -9,6 +9,14 @@ while a `loading` phase is in progress, the same directory `hub_cache.py` scans for `list_downloads`. Byte growth there is real progress whoever triggered it, and is reported as such. + +A `download_progress` event fires on a fixed cadence, not only when the +watched size has grown (#343 follow-up): an xet-backed file reconstructs +against its local CAS cache in bursts, and can hold an unchanged size on +disk for well past the phase-stall threshold while a transfer is genuinely +still running underneath. Ticking on a timer reports that honestly - a +quiet interval is `bytes_per_second=0`, not silence - which is what keeps +the watchdog from mistaking it for a hang. """ import logging @@ -95,23 +103,25 @@ def __exit__(self, *exc_info): def _run(self): baseline = _blob_dir_size(self._blob_dir) - last_size = baseline - last_sample_at = time.monotonic() - last_emit_at = 0.0 + last_emitted_size = baseline + last_emit_at = time.monotonic() while not self._stop.wait(CHECK_INTERVAL_SECONDS): try: size = _blob_dir_size(self._blob_dir) now = time.monotonic() - if size <= last_size: - last_size = size - last_sample_at = now - continue - elapsed = now - last_sample_at - rate = (size - last_size) / elapsed if elapsed > 0 else None - last_size = size - last_sample_at = now - if now - last_emit_at < EMIT_INTERVAL_SECONDS: + elapsed = now - last_emit_at + if elapsed < EMIT_INTERVAL_SECONDS: continue + # Emitted on a timer, not on growth: a large xet-backed file + # reconstructs in bursts (dedup against its CAS cache) and + # can sit with an unchanged size on disk for well over + # EMIT_INTERVAL_SECONDS while genuinely still transferring + # (#343) - waiting for growth before emitting reproduces the + # exact silence the phase-stall watchdog is meant to catch. + # A tick with no growth still reports honestly: 0 B/s, not a + # guessed or carried-over rate. + rate = (size - last_emitted_size) / elapsed if elapsed > 0 else None + last_emitted_size = size last_emit_at = now self._context.emit( "download_progress", diff --git a/tests/test_download_watch.py b/tests/test_download_watch.py index 66391453..c208d471 100644 --- a/tests/test_download_watch.py +++ b/tests/test_download_watch.py @@ -1,7 +1,21 @@ """A download from_pretrained triggers is watched, not intercepted (#343): byte growth in the Hugging Face cache counts as progress, so a real download -does not trip the phase-stall watchdog, and a download that genuinely stops -still does.""" +does not trip the phase-stall watchdog. + +Progress ticks on a timer, not only on growth (#343 follow-up): an +xet-backed transfer reconstructs against its local CAS cache in bursts and +can leave the watched size unchanged on disk for well past the phase-stall +threshold while genuinely still running underneath - confirmed by polling a +real large download (stabilityai/sd-turbo's non-fp16 unet, ~3.5GB) directly, +which sat with an unchanged blob size for 20+ seconds more than once before +jumping hundreds of MB at a time. There is no on-disk signal, in the blob +directory or in HF_XET_CACHE, that fills that gap, so the trade this file +makes deliberately: a watched download cannot read as a stall for as long as +DownloadWatch's thread is alive, in exchange for never mistaking a +reconstruction pause for one. What that costs is real: an `from_pretrained` +call that hangs with no eventual timeout or exception, for the length of the +hang, does not raise a phase-stall warning either - orthogonal to what #343 +reported, and a smaller loss than 17.8 minutes of false ones.""" import time from unittest.mock import patch @@ -60,7 +74,11 @@ def test_growing_download_emits_progress_and_suppresses_stall(tmp_path, monkeypa assert not stalls, "a real, ongoing download must not read as a stall" -def test_stalled_download_still_stalls(tmp_path, monkeypatch): +def test_unchanging_size_still_heartbeats_and_does_not_stall(tmp_path, monkeypatch): + # A size that never changes for the life of the watch is exactly what a + # healthy xet reconstruction pause looks like on disk (#343 follow-up) - + # there is no growth-based signal available to tell it apart from a + # genuine hang, so the watch must still speak up on its own cadence. _fast_watch(monkeypatch) repo_id = "some-org/some-model" blob_dir = tmp_path / "models--some-org--some-model" / "blobs" @@ -81,6 +99,48 @@ def test_stalled_download_still_stalls(tmp_path, monkeypatch): finally: context.exit_run() - assert not [e for e in events if e["event"] == "download_progress"] + progress = [e for e in events if e["event"] == "download_progress"] + assert progress, "an active watch must heartbeat even with no growth" + assert all(e["bytes_per_second"] == 0 for e in progress) + + stalls = [e for e in events if e.get("kind") == "phase_stall"] + assert not stalls, "a watched download must not read as a stall" + + +def test_large_file_freeze_then_burst_does_not_stall(tmp_path, monkeypatch): + # Reproduces the shape measured against a real ~3.5GB xet-backed download + # (stabilityai/sd-turbo unet): the tracked size holds flat for stretches + # well past one emit interval, then jumps hundreds of MB at once when + # reconstruction catches up - not a mock of DownloadWatch's own size + # function, a real directory written on that timing. + _fast_watch(monkeypatch) + repo_id = "some-org/some-model" + blob_dir = tmp_path / "models--some-org--some-model" / "blobs" + blob_dir.mkdir(parents=True) + blob_file = blob_dir / "abc123.incomplete" + blob_file.write_bytes(b"x" * 1024) + + events = [] + context = RunContext(on_event=events.append) + with _fast_watchdog(): + context.enter_run() + try: + context.note_phase("loading") + with download_watch.DownloadWatch( + repo_id, context, cache_dir=str(tmp_path) + ): + # Freeze well past one emit interval (0.05s) ... + time.sleep(0.3) + # ... then a burst, then freeze again. + with open(blob_file, "ab") as f: + f.write(b"x" * (16 * 1024 * 1024)) + time.sleep(0.3) + finally: + context.exit_run() + + progress = [e for e in events if e["event"] == "download_progress"] + assert progress, "growth and quiet stretches must both be reported" + assert progress[-1]["downloaded_bytes"] >= 16 * 1024 * 1024 + stalls = [e for e in events if e.get("kind") == "phase_stall"] - assert stalls, "bytes that stopped growing must still read as a stall" + assert not stalls, "a freeze-then-burst download must not read as a stall" From 998a739e8341d5691ce61f3e1fb91bdb30532f60 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 08:15:47 -0500 Subject: [PATCH 039/181] fix(engine): #349 - add a CPU grade task for images and video New grade task command (exposure, contrast, saturation, temperature, tint), CPU-only via numpy/PIL, per Don's narrow v1 disposition. Video is graded per frame through the existing _per_frame helper, which carries audio and fps through untouched. contrast/saturation are declared NON_NEGATIVE in task_domains.py so validate_workflow refuses a negative value before a run; exposure/temperature/tint have no natural non-negative-only domain and are left to the command itself, per that module's own rule that only undebatable domains are declared there. temperature/tint use a documented linear scale, +-1.0 shifting the R/B or G channel by up to 15% of the channel range. Co-Authored-By: Claude Sonnet 5 --- dw/task_domains.py | 4 + dw/tasks/grade.py | 95 ++++++++++++++++++++++ dw/tasks/task.py | 10 +++ tests/test_grade.py | 188 ++++++++++++++++++++++++++++++++++++++++++++ 4 files changed, 297 insertions(+) create mode 100644 dw/tasks/grade.py create mode 100644 tests/test_grade.py diff --git a/dw/task_domains.py b/dw/task_domains.py index 9a22da9b..ccf90cf7 100644 --- a/dw/task_domains.py +++ b/dw/task_domains.py @@ -114,6 +114,10 @@ "sample_rate": POSITIVE, }, "analyze_audio": {"sample_rate": POSITIVE}, + "grade": { + "contrast": NON_NEGATIVE, + "saturation": NON_NEGATIVE, + }, } diff --git a/dw/tasks/grade.py b/dw/tasks/grade.py new file mode 100644 index 00000000..136b8802 --- /dev/null +++ b/dw/tasks/grade.py @@ -0,0 +1,95 @@ +""" +CPU colour grading: exposure, contrast, saturation, temperature/tint. + +Pure numpy/PIL - no model, no GPU. A video is graded per frame by the +command handler (task.py's _per_frame), so this module only ever sees a +single PIL Image. +""" + +import numpy as np +from PIL import Image + +# Luma weights (Rec. 709), used as the saturation pivot +_LUMA_WEIGHTS = np.array([0.2126, 0.7152, 0.0722], dtype=np.float32) + +# Fraction of the full [0, 1] channel range shifted at |temperature| == 1.0 or +# |tint| == 1.0. Temperature moves the red and blue channels apart (warmer is +# more red, less blue); tint moves the green channel against them (more +# magenta is less green). Both are linear in their argument and independent +# of each other - this is the "documented scale" v1 promises rather than a +# physical colour-temperature-in-Kelvin model, which is more expensive to +# compute and no more useful to a generation pipeline's output. +_WHITE_BALANCE_STRENGTH = 0.15 + + +def grade_image( + image, + exposure=0.0, + contrast=1.0, + saturation=1.0, + temperature=0.0, + tint=0.0, +): + """Adjust exposure, contrast, saturation and white balance of an image. + + Every parameter is optional; omitting one leaves that adjustment at its + identity value, so calling with no arguments returns the input pixels + unchanged (within rounding). Adjustments apply in this order: exposure, + then contrast, then temperature/tint, then saturation. + + Args: + image: PIL Image to grade. + exposure: Stops to brighten (positive) or darken (negative) by, + applied as a multiply of 2**exposure. 0.0 (default) is identity. + contrast: Multiplier applied around the mid grey point (0.5). + 1.0 (default) is identity; above 1 increases contrast, below 1 + (down to 0) flattens it. + saturation: Multiplier applied around each pixel's own luma + (Rec. 709 weights). 1.0 (default) is identity; 0.0 is greyscale. + temperature: Warm/cool white-balance shift from -1.0 (coolest, shifts + toward blue) to 1.0 (warmest, shifts toward red), linear, moving + the red and blue channels apart by up to 15% of the channel + range at |1.0|. 0.0 (default) is identity. + tint: Green/magenta white-balance shift from -1.0 (green) to 1.0 + (magenta), linear, moving the green channel by up to 15% of the + channel range at |1.0| in the opposite direction to the shift's + sign. 0.0 (default) is identity. + + Returns: + PIL Image, same size and mode as the input (graded). An alpha + channel, if the input has one, passes through untouched. + """ + alpha = None + if image.mode in ("RGBA", "LA"): + alpha = image.getchannel("A") + + rgb = image.convert("RGB") + array = np.asarray(rgb, dtype=np.float32) / 255.0 + + if exposure != 0.0: + array = array * (2.0**exposure) + + if contrast != 1.0: + array = (array - 0.5) * contrast + 0.5 + + if temperature != 0.0: + shift = temperature * _WHITE_BALANCE_STRENGTH + array[..., 0] += shift + array[..., 2] -= shift + + if tint != 0.0: + shift = tint * _WHITE_BALANCE_STRENGTH + array[..., 1] -= shift + + if saturation != 1.0: + luma = np.tensordot(array, _LUMA_WEIGHTS, axes=([-1], [0])) + array = luma[..., None] + (array - luma[..., None]) * saturation + + array = np.clip(array * 255.0, 0, 255).astype(np.uint8) + graded = Image.fromarray(array, mode="RGB") + + if alpha is not None: + graded = graded.convert("RGBA") + graded.putalpha(alpha) + + return graded diff --git a/dw/tasks/task.py b/dw/tasks/task.py index 61a6ea7a..b2d1b1e1 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -426,6 +426,16 @@ def _handle_restore_faces(task, arguments, previous_pipelines): ) +@register_command("grade", implementation="dw.tasks.grade.grade_image") +def _handle_grade(task, arguments, previous_pipelines): + """Adjust exposure, contrast, saturation and white balance of an image or video""" + logger.debug("Grading image") + image = arguments.pop("image") + from .grade import grade_image + + return _per_frame(image, lambda frame: grade_image(frame, **arguments)) + + @register_command( "segment", implementation="dw.tasks.segment.segment_image", consumes_device=True ) diff --git a/tests/test_grade.py b/tests/test_grade.py new file mode 100644 index 00000000..8cf688c1 --- /dev/null +++ b/tests/test_grade.py @@ -0,0 +1,188 @@ +"""Tests for the grade task command (#349): CPU-only exposure, contrast, +saturation and white-balance adjustment for an image or a video. +""" + +import numpy +import pytest +from PIL import Image + +from dw.result import AudioVideo +from dw.task_domains import task_argument_errors +from dw.tasks.grade import grade_image +from dw.tasks.task import Task +from dw.workflow import Workflow + + +def _swatch(rgb=(128, 96, 160), size=16): + return Image.new("RGB", (size, size), rgb) + + +def _mean_channels(image): + array = numpy.asarray(image, dtype=numpy.float32) + return array[..., 0].mean(), array[..., 1].mean(), array[..., 2].mean() + + +class TestIdentity: + def test_no_arguments_returns_pixels_unchanged(self): + source = _swatch() + graded = grade_image(source) + assert numpy.array_equal(numpy.asarray(source), numpy.asarray(graded)) + + def test_identity_values_return_pixels_unchanged(self): + source = _swatch() + graded = grade_image( + source, + exposure=0.0, + contrast=1.0, + saturation=1.0, + temperature=0.0, + tint=0.0, + ) + array = numpy.abs( + numpy.asarray(source, dtype=numpy.int16) + - numpy.asarray(graded, dtype=numpy.int16) + ) + assert array.max() <= 1 # rounding only + + +class TestEachParameterMovesTheDirectionItDocuments: + def test_positive_exposure_brightens(self): + source = _swatch((100, 100, 100)) + graded = grade_image(source, exposure=1.0) + assert _mean_channels(graded)[0] > _mean_channels(source)[0] + + def test_negative_exposure_darkens(self): + source = _swatch((100, 100, 100)) + graded = grade_image(source, exposure=-1.0) + assert _mean_channels(graded)[0] < _mean_channels(source)[0] + + def test_contrast_above_one_pushes_a_bright_pixel_brighter(self): + source = _swatch((200, 200, 200)) + graded = grade_image(source, contrast=1.5) + assert _mean_channels(graded)[0] > _mean_channels(source)[0] + + def test_contrast_below_one_pulls_a_bright_pixel_toward_grey(self): + source = _swatch((200, 200, 200)) + graded = grade_image(source, contrast=0.5) + assert _mean_channels(graded)[0] < _mean_channels(source)[0] + + def test_saturation_zero_greys_out_a_colour_swatch(self): + source = _swatch((200, 50, 50)) + graded = grade_image(source, saturation=0.0) + r, g, b = _mean_channels(graded) + assert r == pytest.approx(g, abs=1.0) + assert g == pytest.approx(b, abs=1.0) + + def test_positive_temperature_moves_red_up_and_blue_down(self): + source = _swatch((128, 128, 128)) + graded = grade_image(source, temperature=1.0) + r, _, b = _mean_channels(graded) + sr, _, sb = _mean_channels(source) + assert r > sr + assert b < sb + + def test_negative_temperature_moves_blue_up_and_red_down(self): + source = _swatch((128, 128, 128)) + graded = grade_image(source, temperature=-1.0) + r, _, b = _mean_channels(graded) + sr, _, sb = _mean_channels(source) + assert r < sr + assert b > sb + + def test_positive_tint_reduces_green(self): + source = _swatch((128, 128, 128)) + graded = grade_image(source, tint=1.0) + assert _mean_channels(graded)[1] < _mean_channels(source)[1] + + def test_negative_tint_increases_green(self): + source = _swatch((128, 128, 128)) + graded = grade_image(source, tint=-1.0) + assert _mean_channels(graded)[1] > _mean_channels(source)[1] + + +class TestAlphaPassesThrough: + def test_an_alpha_channel_is_untouched(self): + source = Image.new("RGBA", (4, 4), (200, 50, 50, 77)) + graded = grade_image(source, exposure=1.0) + assert graded.mode == "RGBA" + assert graded.getpixel((0, 0))[3] == 77 + + +class TestVideoIsGradedPerFrame: + def video(self): + frames = [Image.new("RGB", (4, 4), (i * 40, 100, 100)) for i in range(3)] + return AudioVideo(frames, numpy.zeros((2, 50), dtype=numpy.float32), 100, fps=24) + + def test_grade_dispatches_over_every_frame_and_keeps_audio_and_fps(self): + task = Task({"command": "grade", "arguments": {}}, "cpu") + result = task.run({"image": self.video(), "exposure": 1.0}) + + assert isinstance(result, AudioVideo) + assert len(result.frames) == 3 + assert result.fps == 24 + assert result.sample_rate == 100 + assert result.audio.shape == (2, 50) + # each frame was actually graded, not passed through untouched + for source_frame, graded_frame in zip(self.video().frames, result.frames): + assert not numpy.array_equal( + numpy.asarray(source_frame), numpy.asarray(graded_frame) + ) + + +class TestDomains: + def _errors(self, arguments): + definition = { + "id": "grade-domains", + "steps": [ + { + "name": "grade", + "task": {"command": "grade", "arguments": arguments}, + "result": {"content_type": "image/png"}, + } + ], + } + return task_argument_errors(definition) + + def test_a_negative_contrast_is_refused_at_its_path(self): + errors = self._errors({"image": "asset:a.png", "contrast": -0.5}) + assert [e["path"] for e in errors] == ["steps[0].task.arguments.contrast"] + + def test_a_negative_saturation_is_refused(self): + errors = self._errors({"image": "asset:a.png", "saturation": -1.0}) + assert [e["path"] for e in errors] == ["steps[0].task.arguments.saturation"] + + def test_zero_saturation_is_a_legitimate_request(self): + assert self._errors({"image": "asset:a.png", "saturation": 0.0}) == [] + + def test_validate_workflow_refuses_it_before_a_run(self): + workflow = Workflow( + { + "id": "grade-domains", + "steps": [ + { + "name": "grade", + "task": { + "command": "grade", + "arguments": {"image": "asset:a.png", "contrast": -1.0}, + }, + "result": {"content_type": "image/png"}, + } + ], + }, + "outputs", + None, + ) + errors = workflow.validation_errors() + assert [e["path"] for e in errors] == ["steps[0].task.arguments.contrast"] + + def test_exposure_and_temperature_are_unconstrained(self): + from dw.introspection import describe_task + + parameters = {p["name"]: p for p in describe_task("grade")["parameters"]} + assert "domain" not in parameters["exposure"] + assert "domain" not in parameters["temperature"] + assert "domain" not in parameters["tint"] + + +if __name__ == "__main__": + pytest.main([__file__, "-v"]) From 9368b204d541b29974b737fa301f6ad0cbab49ad Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 08:18:32 -0500 Subject: [PATCH 040/181] fix(engine): #349 - ruff format test_grade.py Co-Authored-By: Claude Sonnet 5 --- tests/test_grade.py | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/tests/test_grade.py b/tests/test_grade.py index 8cf688c1..cedade35 100644 --- a/tests/test_grade.py +++ b/tests/test_grade.py @@ -111,7 +111,9 @@ def test_an_alpha_channel_is_untouched(self): class TestVideoIsGradedPerFrame: def video(self): frames = [Image.new("RGB", (4, 4), (i * 40, 100, 100)) for i in range(3)] - return AudioVideo(frames, numpy.zeros((2, 50), dtype=numpy.float32), 100, fps=24) + return AudioVideo( + frames, numpy.zeros((2, 50), dtype=numpy.float32), 100, fps=24 + ) def test_grade_dispatches_over_every_frame_and_keeps_audio_and_fps(self): task = Task({"command": "grade", "arguments": {}}, "cpu") From 2f1c6c86279be252fd543c27ad601f2f2a3a5431 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 08:32:43 -0500 Subject: [PATCH 041/181] fix(catalog): #351 - every generative entry declares seed: variable:seed Converts the catalog's two seed conventions into one: the 18 entries that pinned a literal seed now expose it as "seed": "variable:seed" with the same value declared in variables.seed (default run realizes the identical seed as before, no behavior change), and the 51 unseeded entries gain "seed": "variable:seed" with variables.seed: null (a null-valued variable: seed is already treated as unseeded by _is_seeded/the step cache, so a default run still draws a fresh seed and the cache stays off). This gives rerun_job(new_seed=true) and the realized-workflow seed-readback a uniform variable: seed to work with on every entry - no changes needed in dw/server/jobs.py or dw/realize.py, both already handle a variable-sourced seed whether its default is an int or null. Rewords the no-seed warning in unseeded_cache_warnings (dw/plan.py) to recommend the variable:seed + declared-default convention rather than a bare top-level literal. Bumps the compact-listing token budget (tests/test_catalog_structure.py) to 8_450: every entry now declaring seed in variables adds one name to variable_names, a compact field, following the budget's existing measured-and-documented growth pattern. Co-Authored-By: Claude Sonnet 5 --- dw/plan.py | 5 +++-- tests/test_catalog_structure.py | 5 ++++- workflows/models/flux-dev-compile.json | 4 +++- workflows/models/flux-dev.json | 4 +++- workflows/models/flux-gguf.json | 4 +++- workflows/models/flux-krea.json | 4 +++- workflows/models/flux-torchao.json | 4 +++- workflows/models/flux2-dev.json | 4 +++- workflows/models/krea2.json | 4 +++- workflows/models/qwen-image-2.1.json | 4 +++- workflows/models/z-image-sdnq.json | 4 +++- workflows/templates/attention-processor.json | 4 +++- workflows/templates/base-and-refiner.json | 4 ++++ workflows/templates/best-of-n-to-video.json | 4 +++- workflows/templates/community-pipeline.json | 4 +++- workflows/templates/compose-workflows.json | 4 +++- workflows/templates/consistent-set.json | 4 +++- workflows/templates/controlnet-component.json | 4 +++- workflows/templates/controlnet.json | 4 ++++ workflows/templates/depth-marigold.json | 4 +++- workflows/templates/describe-and-regenerate.json | 4 +++- workflows/templates/embed-metadata.json | 4 +++- workflows/templates/image-edit.json | 4 +++- workflows/templates/image-to-image.json | 4 +++- workflows/templates/image-variation.json | 4 +++- workflows/templates/inpaint.json | 4 +++- workflows/templates/interpolate-frames.json | 4 +++- workflows/templates/ip-adapter.json | 4 +++- workflows/templates/lora-styles.json | 4 +++- workflows/templates/lora.json | 4 +++- workflows/templates/ltx2/chained-segments.json | 4 +++- workflows/templates/ltx2/diffusion-decode.json | 5 +++-- workflows/templates/ltx2/enhance-prompt.json | 4 +++- workflows/templates/ltx2/extend-clip.json | 5 +++-- workflows/templates/ltx2/generative-upscale.json | 5 +++-- workflows/templates/ltx2/image-to-video.json | 4 +++- workflows/templates/ltx2/keyframes.json | 5 +++-- workflows/templates/ltx2/reference-sheet.json | 5 +++-- workflows/templates/ltx2/restore-deblur.json | 5 +++-- workflows/templates/ltx2/restore-decompression.json | 5 +++-- workflows/templates/ltx2/text-to-video.json | 5 +++-- workflows/templates/ltx2/two-stage.json | 5 +++-- workflows/templates/minimax/chain-matched-and-aligned.json | 5 +++-- workflows/templates/minimax/chain-matched-to-audio.json | 4 +++- workflows/templates/minimax/chain-video-continuity.json | 5 +++-- workflows/templates/minimax/chained-segments.json | 4 +++- workflows/templates/minimax/composable-references.json | 5 +++-- workflows/templates/minimax/dialogue-short.json | 5 +++-- workflows/templates/minimax/enhance-prompt-with-image.json | 4 +++- workflows/templates/minimax/enhance-prompt.json | 5 +++-- workflows/templates/minimax/first-and-last-frame.json | 4 +++- workflows/templates/minimax/generated-subject-reference.json | 4 +++- workflows/templates/minimax/image-to-video.json | 4 +++- workflows/templates/minimax/last-frame-only.json | 4 +++- workflows/templates/minimax/music-video.json | 5 +++-- workflows/templates/minimax/music.json | 4 +++- workflows/templates/minimax/reference-to-video.json | 4 +++- workflows/templates/minimax/storyboard.json | 5 +++-- workflows/templates/minimax/video-with-audio-768p.json | 5 +++-- workflows/templates/minimax/video-with-audio.json | 5 +++-- workflows/templates/minimax/voice-timbre-reference.json | 4 +++- workflows/templates/multi-image-reference.json | 4 +++- workflows/templates/outpaint.json | 4 +++- workflows/templates/prompt-weighting.json | 4 +++- workflows/templates/qr-code.json | 4 +++- workflows/templates/restore-faces.json | 4 +++- workflows/templates/segment-and-inpaint.json | 4 +++- workflows/templates/step-caching.json | 4 +++- workflows/templates/sub-workflow.json | 4 +++- workflows/templates/surface-normals.json | 4 +++- workflows/templates/text-to-image.json | 4 +++- 71 files changed, 216 insertions(+), 88 deletions(-) diff --git a/dw/plan.py b/dw/plan.py index 77453f91..b2096e54 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -215,8 +215,9 @@ def unseeded_cache_warnings(definition, arguments=None): return [ "This workflow sets no 'seed', so the step cache is disabled and " "'cached_steps' is 0 without being probed - every step regenerates " - "on every run. Set a top-level 'seed' to make a repeat run reuse " - "what it already produced" + "on every run. Set a top-level 'seed': 'variable:seed' with a " + "declared default in 'variables' to make a repeat run reuse what it " + "already produced" ] diff --git a/tests/test_catalog_structure.py b/tests/test_catalog_structure.py index 425cc4ba..6e2f6b43 100644 --- a/tests/test_catalog_structure.py +++ b/tests/test_catalog_structure.py @@ -399,7 +399,10 @@ def test_no_stale_entry_in_the_allowlist(): # storyboard) each add a `[{device, name, vram_gb, minutes}]` entry to the # listing; `vram_estimate` beside it is not a compact field and costs nothing # here. -COMPACT_BUDGET = 8_350 +# Then to 8_450 for the catalog-wide seed convention (#351): every generative +# entry now declares `seed` in `variables`, which is a compact field +# (`variable_names`), measured at 8_424. +COMPACT_BUDGET = 8_450 FILTERED_BUDGET = 1_500 diff --git a/workflows/models/flux-dev-compile.json b/workflows/models/flux-dev-compile.json index 5c5473ea..5f846731 100644 --- a/workflows/models/flux-dev-compile.json +++ b/workflows/models/flux-dev-compile.json @@ -13,8 +13,10 @@ "guidance_scale": 3.5, "width": 1024, "height": 1024, - "model_name": "black-forest-labs/FLUX.1-dev" + "model_name": "black-forest-labs/FLUX.1-dev", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/models/flux-dev.json b/workflows/models/flux-dev.json index 43cb5e09..af634b8a 100644 --- a/workflows/models/flux-dev.json +++ b/workflows/models/flux-dev.json @@ -6,7 +6,8 @@ "num_inference_steps": 25, "guidance_scale": 3.5, "width": 768, - "height": 768 + "height": 768, + "seed": null }, "id": "FluxDev", "description": "Text-to-image with FLUX.1 dev - the reference FLUX workflow.", @@ -14,6 +15,7 @@ {"device": "cuda", "name": "RTX 3090", "vram_gb": 24, "minutes": 1.0} ], "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/models/flux-gguf.json b/workflows/models/flux-gguf.json index 91104dd6..f0eff64f 100644 --- a/workflows/models/flux-gguf.json +++ b/workflows/models/flux-gguf.json @@ -1,11 +1,13 @@ { "variables": { "checkpoint_path": "https://huggingface.co/city96/FLUX.1-dev-gguf/blob/main/flux1-dev-Q2_K.gguf", - "prompt": "a marmot wearing a tophat, in a forest scene." + "prompt": "a marmot wearing a tophat, in a forest scene.", + "seed": null }, "id": "FluxGGUF", "description": "FLUX.1 dev loaded from a GGUF-quantized transformer checkpoint.", "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/models/flux-krea.json b/workflows/models/flux-krea.json index 72f65e24..9ab6780c 100644 --- a/workflows/models/flux-krea.json +++ b/workflows/models/flux-krea.json @@ -5,11 +5,13 @@ "num_inference_steps": 25, "guidance_scale": 4.5, "width": 768, - "height": 768 + "height": 768, + "seed": null }, "id": "FluxDevKrea", "description": "Text-to-image with FLUX.1 Krea dev, tuned for photographic aesthetics.", "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/models/flux-torchao.json b/workflows/models/flux-torchao.json index ffa2b510..ff55f7f1 100644 --- a/workflows/models/flux-torchao.json +++ b/workflows/models/flux-torchao.json @@ -1,10 +1,12 @@ { "variables": { - "prompt": "a marmot wearing a tophat, in a forest scene." + "prompt": "a marmot wearing a tophat, in a forest scene.", + "seed": null }, "id": "FluxTorchAO", "description": "FLUX.1 dev with the transformer quantized to int4 via TorchAO, compiled. Unmeasured; the int8 variant ran at over a minute per denoising step on an RTX 3090, so measure before relying on it - models/flux-dev-compile records what was measured there.", "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/models/flux2-dev.json b/workflows/models/flux2-dev.json index 80ac952b..9d689512 100644 --- a/workflows/models/flux2-dev.json +++ b/workflows/models/flux2-dev.json @@ -4,7 +4,8 @@ "prompt": "A realistic photograph of a mouse wearing a skirt playing volleyball against a team of professional volleyball players.", "num_images_per_prompt": 1, "num_inference_steps": 40, - "guidance_scale": 4.0 + "guidance_scale": 4.0, + "seed": null }, "id": "Flux2Dev", "description": "Text-to-image with FLUX.2 dev, pre-quantized to 4-bit BitsAndBytes. The 4-bit text encoder is loaded first as its own component and rested on the CPU, since BitsAndBytes materializes on the accelerator and the two together do not fit at load; model offloading then hands the accelerator to one component at a time, so both run on a 24 GB card.", @@ -17,6 +18,7 @@ } ], "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/models/krea2.json b/workflows/models/krea2.json index b76b649a..f3937d7e 100644 --- a/workflows/models/krea2.json +++ b/workflows/models/krea2.json @@ -1,10 +1,12 @@ { "variables": { - "prompt": "prompt:zimage/kidney_trade_in" + "prompt": "prompt:zimage/kidney_trade_in", + "seed": null }, "id": "Krea2", "description": "Text-to-image with Krea 2 Turbo - fast, few-step photorealistic generation.", "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/models/qwen-image-2.1.json b/workflows/models/qwen-image-2.1.json index 89f633b1..d2db51c3 100644 --- a/workflows/models/qwen-image-2.1.json +++ b/workflows/models/qwen-image-2.1.json @@ -18,7 +18,8 @@ "num_images_per_prompt": 1, "num_inference_steps": 40, "width": 2048, - "height": 2048 + "height": 2048, + "seed": null }, "id": "QwenImage2.1", "description": "Text-to-image with Qwen-Image-2.1: a unified 7B model (32 single-stream DiT layers, Qwen3-VL 8B text encoder) that also does instruction editing and native RGBA output; this entry exercises plain text-to-image and models/qwen-image-2.1-edit exercises editing. No guidance_scale - the pipeline takes only true_cfg_scale, which the vendor recommends leaving at its default of 1.0 (off). width/height are set explicitly because the pipeline falls back to a 1024x1024 output_resolution default otherwise, well under the model's native 2K. The result is image/png because the pipeline returns RGBA: a jpeg result would flatten the alpha over white (and warn). The VAE is tiled and sliced because at 2048x2048 on a 24GB card the decode - not the denoise - runs out of memory, with all 40 steps already done. Offload is a property of the mode, not the model: 'model' offload fits plain text-to-image at 2048x2048, but an edit prepends the condition image's latents to the sequence and needs 'sequential' at half those pixels - which is why the edit entry differs. Qwen Research License: non-commercial use only, unlike every other model in this catalog, which are Apache-2.0 - see https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE.", @@ -35,6 +36,7 @@ "reason": "Qwen-Image-2.1 rounds width/height down to a multiple of 32 (vae_scale_factor * 2), so an off-grid value is not refused by the pipeline but silently shrunk - refused here instead, with no 'snap', since rounding the other way would be a second silent change to the size asked for" } }, + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/models/z-image-sdnq.json b/workflows/models/z-image-sdnq.json index b0a52306..7d36930f 100644 --- a/workflows/models/z-image-sdnq.json +++ b/workflows/models/z-image-sdnq.json @@ -5,11 +5,13 @@ "num_inference_steps": 9, "guidance_scale": 0.0, "width": 1024, - "height": 1024 + "height": 1024, + "seed": null }, "id": "ZImageSDNQ", "description": "Z-Image Turbo pre-quantized to SDNQ uint4 - the same nine-step generation in a fraction of the VRAM.", "configures": "templates/text-to-image", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/templates/attention-processor.json b/workflows/templates/attention-processor.json index d3893669..b5f5a256 100644 --- a/workflows/templates/attention-processor.json +++ b/workflows/templates/attention-processor.json @@ -6,8 +6,10 @@ "num_inference_steps": 25, "guidance_scale": 7.5, "height": 512, - "width": 512 + "width": 512, + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "generate_with_sdpa", diff --git a/workflows/templates/base-and-refiner.json b/workflows/templates/base-and-refiner.json index 17ddf6fd..b0f4502a 100644 --- a/workflows/templates/base-and-refiner.json +++ b/workflows/templates/base-and-refiner.json @@ -1,6 +1,10 @@ { "id": "sdxl", "description": "SDXL base plus refiner, two-stage.", + "variables": { + "seed": null + }, + "seed": "variable:seed", "steps": [ { "name": "sdxl_base", diff --git a/workflows/templates/best-of-n-to-video.json b/workflows/templates/best-of-n-to-video.json index aed9a1aa..f2c390da 100644 --- a/workflows/templates/best-of-n-to-video.json +++ b/workflows/templates/best-of-n-to-video.json @@ -9,8 +9,10 @@ "rubric": "How sharp and well-composed is this photograph? Penalize blur, extra or malformed limbs, and flat or blown-out lighting.", "scale": [0, 10], "num_inference_steps": 25, - "video_prompt": "prompt:ltx2/best-of-n-video-motion" + "video_prompt": "prompt:ltx2/best-of-n-video-motion", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "still", diff --git a/workflows/templates/community-pipeline.json b/workflows/templates/community-pipeline.json index ac28c2cf..ba09d813 100644 --- a/workflows/templates/community-pipeline.json +++ b/workflows/templates/community-pipeline.json @@ -1,10 +1,12 @@ { "variables": { "prompt": "a marmot", - "num_images_per_prompt": 4 + "num_images_per_prompt": 4, + "seed": null }, "id": "FluxRfInversion", "description": "Inverts an input image into FLUX.1 latents for editing, via the community RF-inversion pipeline (4-bit quantized).", + "seed": "variable:seed", "steps": [ { "name": "invert", diff --git a/workflows/templates/compose-workflows.json b/workflows/templates/compose-workflows.json index ff2b4e3d..8a633ba2 100644 --- a/workflows/templates/compose-workflows.json +++ b/workflows/templates/compose-workflows.json @@ -7,10 +7,12 @@ "image_prompt": "portrait | wide angle shot of eyes off to one side of frame, lucid dream-like 3d model of a marmot, game asset, blender, looking off in distance ::8 style | glowing ::8 background | forest, vivid neon wonderland, particles, blue, green, orange ::7 parameters | rule of thirds, golden ratio, asymmetric composition, hyper- maximalist, octane render, photorealism, cinematic realism, unreal engine, 8k ::7 --ar 16:9 --s 1000", "video_generation_workflow": "./minimax/image-to-video.json", "video_prompt": "the marmot blinks and looks around", - "video_num_inference_steps": 20 + "video_num_inference_steps": 20, + "seed": null }, "id": "txt2img2vid", "description": "Composes any text-to-image workflow with any image-to-video workflow into one still-then-animate run. Both are named as variables, so the pair can be swapped without touching this file. The still the first step generates is handed to the second as its 'image'. Paths resolve against this file's own directory, so a workflow elsewhere in the catalog is named relatively: '../models/z-image.json' here for the still, './minimax/image-to-video.json' for the animation.", + "seed": "variable:seed", "steps": [ { "name": "image_generation", diff --git a/workflows/templates/consistent-set.json b/workflows/templates/consistent-set.json index 11d1abaa..7968b3e0 100644 --- a/workflows/templates/consistent-set.json +++ b/workflows/templates/consistent-set.json @@ -3,7 +3,8 @@ "prompt": "a single ceramic coffee mug on a plain wooden table, soft window light, photograph", "first_edit": "make the mug red", "second_edit": "make the mug blue", - "third_edit": "make the mug green" + "third_edit": "make the mug green", + "seed": null }, "id": "ConsistentSet", "shape": "image-set", @@ -16,6 +17,7 @@ "minutes": 11.5 } ], + "seed": "variable:seed", "steps": [ { "name": "base", diff --git a/workflows/templates/controlnet-component.json b/workflows/templates/controlnet-component.json index e533ddb3..b5209bed 100644 --- a/workflows/templates/controlnet-component.json +++ b/workflows/templates/controlnet-component.json @@ -1,9 +1,11 @@ { "variables": { - "mask_image_uri": "https://nftfactory.blob.core.windows.net/images/chia256.jpg" + "mask_image_uri": "https://nftfactory.blob.core.windows.net/images/chia256.jpg", + "seed": null }, "id": "sd15-controlnet", "description": "Estimates a depth map from an input image, then generates with SD 1.5 guided by a depth ControlNet.", + "seed": "variable:seed", "steps": [ { "name": "depth", diff --git a/workflows/templates/controlnet.json b/workflows/templates/controlnet.json index 98a3c9d6..82567537 100644 --- a/workflows/templates/controlnet.json +++ b/workflows/templates/controlnet.json @@ -1,6 +1,10 @@ { "id": "FluxCanny", "description": "Extracts Canny edges from an input image and generates with FLUX.1 Canny dev guided by them. The preprocessor is the only thing that changes between control types: swap the 'canny' task for 'depth' and the model for 'black-forest-labs/FLUX.1-Depth-dev' (guidance_scale 10 rather than 30) for depth control. image-processors.json runs every preprocessor side by side.", + "variables": { + "seed": null + }, + "seed": "variable:seed", "steps": [ { "name": "canny", diff --git a/workflows/templates/depth-marigold.json b/workflows/templates/depth-marigold.json index f8fc08a7..6be12396 100644 --- a/workflows/templates/depth-marigold.json +++ b/workflows/templates/depth-marigold.json @@ -2,8 +2,10 @@ "id": "MarigoldDepth", "description": "Estimates depth maps for input images with Marigold.", "variables": { - "image_url": "https://marigoldmonodepth.github.io/images/einstein.jpg" + "image_url": "https://marigoldmonodepth.github.io/images/einstein.jpg", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "load_image", diff --git a/workflows/templates/describe-and-regenerate.json b/workflows/templates/describe-and-regenerate.json index a1f3a7f8..12cd2d6a 100644 --- a/workflows/templates/describe-and-regenerate.json +++ b/workflows/templates/describe-and-regenerate.json @@ -2,10 +2,12 @@ "variables": { "image_url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg?download:true", "vision_model": "Qwen/Qwen3-VL-4B-Instruct", - "language_model": "Qwen/Qwen2.5-1.5B-Instruct" + "language_model": "Qwen/Qwen2.5-1.5B-Instruct", + "seed": null }, "id": "img2txt2img", "description": "Describes an input image with a VLM, expands the caption with an LLM, then regenerates it with FLUX.1 dev. Both text steps are the 'text_generation' task - with an image it runs the vision model named in 'vision_model', without one the language model in 'language_model' - and 'release_models' drops each before the next model loads. Every model is one transformers ships, so nothing needs --trust-workflows.", + "seed": "variable:seed", "steps": [ { "name": "input_image", diff --git a/workflows/templates/embed-metadata.json b/workflows/templates/embed-metadata.json index 5b262de8..572f3fed 100644 --- a/workflows/templates/embed-metadata.json +++ b/workflows/templates/embed-metadata.json @@ -3,8 +3,10 @@ "description": "FLUX.1 dev generation with full generation metadata embedded in the saved PNG.", "variables": { "prompt": "A serene mountain landscape at golden hour, photorealistic", - "steps": 25 + "steps": 25, + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "generate", diff --git a/workflows/templates/image-edit.json b/workflows/templates/image-edit.json index 9cf5bd6e..a6caca8a 100644 --- a/workflows/templates/image-edit.json +++ b/workflows/templates/image-edit.json @@ -3,10 +3,12 @@ "prompt": "recolor this person's hair blonde", "input_image": { "location": "https://pbs.twimg.com/media/HNi-N_eWIAAJfTZ?format=jpg&name=900x900" - } + }, + "seed": null }, "id": "FluxKontextEdit", "description": "Edits an input picture toward a written instruction rather than regenerating it from scratch. FLUX.1 Kontext holds everything the instruction does not name - subject, framing, light - and changes only what it does. A diffusers pipeline, so it runs without --trust-workflows.", + "seed": "variable:seed", "steps": [ { "name": "edit", diff --git a/workflows/templates/image-to-image.json b/workflows/templates/image-to-image.json index 82a3aac5..cfe83a6f 100644 --- a/workflows/templates/image-to-image.json +++ b/workflows/templates/image-to-image.json @@ -9,10 +9,12 @@ "guidance_scale": 2.0, "strength": 0.75, "width": 96, - "height": 96 + "height": 96, + "seed": null }, "id": "FluxImg2Img", "description": "Image-to-image with FLUX.1 dev - restyles an input image toward the prompt.", + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/templates/image-variation.json b/workflows/templates/image-variation.json index a6b8ec50..8c7c9734 100644 --- a/workflows/templates/image-variation.json +++ b/workflows/templates/image-variation.json @@ -1,11 +1,13 @@ { "variables": { "input_image": "https://pbs.twimg.com/media/GgpbmxGWkAATZtp?format=jpg&name=small", - "num_images_per_prompt": 4 + "num_images_per_prompt": 4, + "seed": null }, "id": "FluxRedux", "description": "Image variation with FLUX.1 Redux - encodes a reference image and regenerates it with FLUX.1 dev.", "shape": "image-edit", + "seed": "variable:seed", "steps": [ { "name": "prior", diff --git a/workflows/templates/inpaint.json b/workflows/templates/inpaint.json index 2e01f671..691681eb 100644 --- a/workflows/templates/inpaint.json +++ b/workflows/templates/inpaint.json @@ -1,10 +1,12 @@ { "variables": { "image": "https://huggingface.co/datasets/diffusers/diffusers-images-docs/resolve/main/cup.png", - "mask_image": "https://huggingface.co/datasets/diffusers/diffusers-images-docs/resolve/main/cup_mask.png" + "mask_image": "https://huggingface.co/datasets/diffusers/diffusers-images-docs/resolve/main/cup_mask.png", + "seed": null }, "id": "FluxFill", "description": "Inpaints the masked region of an image with FLUX.1 Fill dev.", + "seed": "variable:seed", "steps": [ { "name": "fill", diff --git a/workflows/templates/interpolate-frames.json b/workflows/templates/interpolate-frames.json index 25fefb0b..de0eb767 100644 --- a/workflows/templates/interpolate-frames.json +++ b/workflows/templates/interpolate-frames.json @@ -3,8 +3,10 @@ "description": "Generates a short Mochi video, then smooths it by interpolating extra frames between the generated ones.", "variables": { "prompt": "A cat walking across a sunlit room", - "multiplier": 2 + "multiplier": 2, + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "generate_video", diff --git a/workflows/templates/ip-adapter.json b/workflows/templates/ip-adapter.json index 8562bd82..27b9eae9 100644 --- a/workflows/templates/ip-adapter.json +++ b/workflows/templates/ip-adapter.json @@ -1,9 +1,11 @@ { "variables": { - "prompt": "A marmot sits at the counter and drinks a milkshake" + "prompt": "A marmot sits at the counter and drinks a milkshake", + "seed": null }, "id": "FluxIP", "description": "FLUX.1 dev with an IP-Adapter conditioning generation on a reference image. The same 'ip_adapter' block takes a face model, which keeps a person's face across a new scene: 'h94/IP-Adapter' with subfolder 'models', weight_name 'ip-adapter-full-face_sd15.bin' and a scale of 0.5, on a StableDiffusionPipeline.", + "seed": "variable:seed", "steps": [ { "name": "main", diff --git a/workflows/templates/lora-styles.json b/workflows/templates/lora-styles.json index 8bb2f158..d928a390 100644 --- a/workflows/templates/lora-styles.json +++ b/workflows/templates/lora-styles.json @@ -6,10 +6,12 @@ "guidance_scale": 3.5, "width": 2048, "height": 1024, - "weight_name": "couple-profile.safetensors" + "weight_name": "couple-profile.safetensors", + "seed": null }, "id": "FluxInContext", "description": "Runs the same prompt through ten FLUX.1 in-context LoRA styles - one image per adapter.", + "seed": "variable:seed", "steps": [ { "name": "couple", diff --git a/workflows/templates/lora.json b/workflows/templates/lora.json index b30240b1..57705276 100644 --- a/workflows/templates/lora.json +++ b/workflows/templates/lora.json @@ -4,10 +4,12 @@ "lora": "XLabs-AI/flux-RealismLora", "num_images_per_prompt": 1, "num_inference_steps": 25, - "guidance_scale": 3.5 + "guidance_scale": 3.5, + "seed": null }, "id": "FluxLora", "description": "FLUX.1 dev with a LoRA adapter. The same 'loras' list works on any pipeline that supports adapters - on StableDiffusion3Pipeline with 'stabilityai/stable-diffusion-3.5-large' and the 'linoyts/yart_art_sd3-5_lora' adapter, for instance.", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/templates/ltx2/chained-segments.json b/workflows/templates/ltx2/chained-segments.json index 5350f8d7..0a9c44b0 100644 --- a/workflows/templates/ltx2/chained-segments.json +++ b/workflows/templates/ltx2/chained-segments.json @@ -15,7 +15,8 @@ "segments": 3, "image": { "location": "https://pbs.twimg.com/media/HPxFgsIXIAAjtUu?format=jpg&name=medium" - } + }, + "seed": null }, "variable_constraints": { "num_frames": { @@ -25,6 +26,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, + "seed": "variable:seed", "steps": [ { "name": "chained_image_to_video", diff --git a/workflows/templates/ltx2/diffusion-decode.json b/workflows/templates/ltx2/diffusion-decode.json index e9e4d254..54752d93 100644 --- a/workflows/templates/ltx2/diffusion-decode.json +++ b/workflows/templates/ltx2/diffusion-decode.json @@ -14,7 +14,8 @@ "width": 960, "height": 544, "num_frames": 121, - "frame_rate": 24.0 + "frame_rate": 24.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -24,7 +25,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "latents", diff --git a/workflows/templates/ltx2/enhance-prompt.json b/workflows/templates/ltx2/enhance-prompt.json index 51172f6c..83d43e4b 100644 --- a/workflows/templates/ltx2/enhance-prompt.json +++ b/workflows/templates/ltx2/enhance-prompt.json @@ -16,8 +16,10 @@ "image": { "location": "https://pbs.twimg.com/media/HPxFgsIXIAAjtUu?format=jpg&name=medium" }, - "enhancer_weights_dtype": "{uint4}" + "enhancer_weights_dtype": "{uint4}", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "enhanced_image_to_video", diff --git a/workflows/templates/ltx2/extend-clip.json b/workflows/templates/ltx2/extend-clip.json index f4567f30..d2cdbbf1 100644 --- a/workflows/templates/ltx2/extend-clip.json +++ b/workflows/templates/ltx2/extend-clip.json @@ -13,7 +13,8 @@ "height": 448, "clip_frames": 121, "num_frames": 241, - "frame_rate": 24.0 + "frame_rate": 24.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -23,7 +24,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "opening", diff --git a/workflows/templates/ltx2/generative-upscale.json b/workflows/templates/ltx2/generative-upscale.json index b46ca15c..6cbfdbec 100644 --- a/workflows/templates/ltx2/generative-upscale.json +++ b/workflows/templates/ltx2/generative-upscale.json @@ -12,7 +12,8 @@ "width": 960, "height": 576, "num_frames": 121, - "frame_rate": 24.0 + "frame_rate": 24.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -22,7 +23,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "low_resolution", diff --git a/workflows/templates/ltx2/image-to-video.json b/workflows/templates/ltx2/image-to-video.json index 167cc67b..ac73056a 100644 --- a/workflows/templates/ltx2/image-to-video.json +++ b/workflows/templates/ltx2/image-to-video.json @@ -13,7 +13,8 @@ "num_frames": 481, "image": { "location": "https://pbs.twimg.com/media/HPxFgsIXIAAjtUu?format=jpg&name=medium" - } + }, + "seed": null }, "variable_constraints": { "num_frames": { @@ -23,6 +24,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, + "seed": "variable:seed", "steps": [ { "name": "image_to_video", diff --git a/workflows/templates/ltx2/keyframes.json b/workflows/templates/ltx2/keyframes.json index c83bf262..5a26bcda 100644 --- a/workflows/templates/ltx2/keyframes.json +++ b/workflows/templates/ltx2/keyframes.json @@ -18,7 +18,8 @@ }, "last_image": { "location": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/flf2v_input_last_frame.png" - } + }, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -28,7 +29,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "keyframes_to_video", diff --git a/workflows/templates/ltx2/reference-sheet.json b/workflows/templates/ltx2/reference-sheet.json index 7cd39d3a..18968de2 100644 --- a/workflows/templates/ltx2/reference-sheet.json +++ b/workflows/templates/ltx2/reference-sheet.json @@ -18,7 +18,8 @@ "num_frames": 121, "reference_frames": 121, "frame_rate": 24.0, - "lora_scale": 1.0 + "lora_scale": 1.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -32,7 +33,7 @@ "reason": "the Ingredients IC-LoRA reads the reference through a 121-frame bucket, and every target it was trained on was at least that long - a shorter static sheet breaks the reference encoding rather than shortening it" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "static_sheet", diff --git a/workflows/templates/ltx2/restore-deblur.json b/workflows/templates/ltx2/restore-deblur.json index 9d6e0d3c..8fb779f8 100644 --- a/workflows/templates/ltx2/restore-deblur.json +++ b/workflows/templates/ltx2/restore-deblur.json @@ -17,7 +17,8 @@ "height": 544, "num_frames": 121, "frame_rate": 24.0, - "lora_scale": 1.0 + "lora_scale": 1.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -27,7 +28,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "restored", diff --git a/workflows/templates/ltx2/restore-decompression.json b/workflows/templates/ltx2/restore-decompression.json index d0392d67..64cb023a 100644 --- a/workflows/templates/ltx2/restore-decompression.json +++ b/workflows/templates/ltx2/restore-decompression.json @@ -17,7 +17,8 @@ "height": 544, "num_frames": 121, "frame_rate": 24.0, - "lora_scale": 1.0 + "lora_scale": 1.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -27,7 +28,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "restored", diff --git a/workflows/templates/ltx2/text-to-video.json b/workflows/templates/ltx2/text-to-video.json index 0d239701..b7bae28b 100644 --- a/workflows/templates/ltx2/text-to-video.json +++ b/workflows/templates/ltx2/text-to-video.json @@ -19,7 +19,8 @@ "width": 960, "height": 544, "num_frames": 121, - "frame_rate": 24.0 + "frame_rate": 24.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -29,7 +30,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "text_to_video", diff --git a/workflows/templates/ltx2/two-stage.json b/workflows/templates/ltx2/two-stage.json index 54633859..d6f34ebd 100644 --- a/workflows/templates/ltx2/two-stage.json +++ b/workflows/templates/ltx2/two-stage.json @@ -21,7 +21,8 @@ "full_width": 1536, "full_height": 896, "num_frames": 121, - "frame_rate": 24.0 + "frame_rate": 24.0, + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -31,7 +32,7 @@ "reason": "LTX-2.5's video VAE encodes 8 * n + 1 frames, so an off-grid count is not refused by the pipeline but floored to the grid below - a shorter clip than the one asked for. Refused here instead: no 'snap', because rounding a length the other way from the pipeline's own would be a second silent change" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "base", diff --git a/workflows/templates/minimax/chain-matched-and-aligned.json b/workflows/templates/minimax/chain-matched-and-aligned.json index 11230539..a3248930 100644 --- a/workflows/templates/minimax/chain-matched-and-aligned.json +++ b/workflows/templates/minimax/chain-matched-and-aligned.json @@ -2,7 +2,7 @@ "id": "MiniMaxH3Ref2VAChainedAligned", "description": "The chaining features used together, for a long clip driven by a supplied soundtrack. 'match_audio' sizes the chain to the audio while 'last_segment' continuity carries motion and framing across the seams, with 'carry_audio' off because the segments already take their audio from the supplied track rather than from each other. Also shows 'prompts', which lets a chain say something different on its opening segment than on its continuations, and 'save_segments', which writes each finished segment to disk so memory stays bounded to one segment and a crashed run keeps its progress.", "summary": "A long H3 clip sized to a supplied soundtrack, with video-tail continuity across every seam.", - "seed": 12345, + "seed": "variable:seed", "cost_drivers": [ "num_frames", "num_inference_steps", @@ -27,7 +27,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 12345 }, "variable_constraints": { "num_frames": { diff --git a/workflows/templates/minimax/chain-matched-to-audio.json b/workflows/templates/minimax/chain-matched-to-audio.json index 1888470e..64acdb27 100644 --- a/workflows/templates/minimax/chain-matched-to-audio.json +++ b/workflows/templates/minimax/chain-matched-to-audio.json @@ -23,7 +23,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": null }, "variable_constraints": { "num_frames": { @@ -35,6 +36,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "chained_reference_to_video_audio", diff --git a/workflows/templates/minimax/chain-video-continuity.json b/workflows/templates/minimax/chain-video-continuity.json index a87e1989..4fdce8d3 100644 --- a/workflows/templates/minimax/chain-video-continuity.json +++ b/workflows/templates/minimax/chain-video-continuity.json @@ -2,7 +2,7 @@ "id": "MiniMaxH3Ref2VAChainedVideo", "description": "The stronger form of chain continuity. Instead of handing the next segment a single frame, 'last_segment' hands it the tail of the previous segment as a video reference, complete with the audio generated alongside it - so motion, camera and voice survive the seam, not just appearance. 'carry_frames' bounds how much of that tail is carried, which is the lever between continuity and the sequence length (and memory) each segment costs.", "summary": "A long H3 clip chained on each segment's video tail, so motion, camera and voice carry across the seams.", - "seed": 12345, + "seed": "variable:seed", "cost_drivers": [ "segments", "num_frames", @@ -27,7 +27,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": 12345 }, "variable_constraints": { "num_frames": { diff --git a/workflows/templates/minimax/chained-segments.json b/workflows/templates/minimax/chained-segments.json index 1c7c6ad5..345af458 100644 --- a/workflows/templates/minimax/chained-segments.json +++ b/workflows/templates/minimax/chained-segments.json @@ -26,7 +26,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": null }, "variable_constraints": { "num_frames": { @@ -38,6 +39,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "chained_keyframe_to_video_audio", diff --git a/workflows/templates/minimax/composable-references.json b/workflows/templates/minimax/composable-references.json index b5977a11..364e9484 100644 --- a/workflows/templates/minimax/composable-references.json +++ b/workflows/templates/minimax/composable-references.json @@ -2,7 +2,7 @@ "id": "MiniMaxH3Ref2VAVideo", "description": "Reference types are composable, and each one contributes a different attribute. Here an image reference supplies the subject and a video reference supplies the shot - its framing, lighting and camera movement - so the two are recombined into a clip that copies neither wholesale. The prompt's retention analysis is what assigns those roles; the reference list only declares what is available.", "summary": "H3 video: an image reference supplies the subject, a video reference the framing and camera.", - "seed": 12345, + "seed": "variable:seed", "cost_drivers": [ "num_frames", "num_inference_steps", @@ -24,7 +24,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": 12345 }, "variable_constraints": { "num_frames": { diff --git a/workflows/templates/minimax/dialogue-short.json b/workflows/templates/minimax/dialogue-short.json index 930862aa..66097a66 100644 --- a/workflows/templates/minimax/dialogue-short.json +++ b/workflows/templates/minimax/dialogue-short.json @@ -130,7 +130,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -142,7 +143,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "draw_character_a", diff --git a/workflows/templates/minimax/enhance-prompt-with-image.json b/workflows/templates/minimax/enhance-prompt-with-image.json index a7c87bff..ae07384d 100644 --- a/workflows/templates/minimax/enhance-prompt-with-image.json +++ b/workflows/templates/minimax/enhance-prompt-with-image.json @@ -26,7 +26,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": null }, "variable_constraints": { "num_frames": { @@ -38,6 +39,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "prompt_enhancer", diff --git a/workflows/templates/minimax/enhance-prompt.json b/workflows/templates/minimax/enhance-prompt.json index f04be4b6..73275c54 100644 --- a/workflows/templates/minimax/enhance-prompt.json +++ b/workflows/templates/minimax/enhance-prompt.json @@ -41,7 +41,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -53,7 +54,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "prompt_enhancer", diff --git a/workflows/templates/minimax/first-and-last-frame.json b/workflows/templates/minimax/first-and-last-frame.json index 2eb6b89c..064f8e91 100644 --- a/workflows/templates/minimax/first-and-last-frame.json +++ b/workflows/templates/minimax/first-and-last-frame.json @@ -27,7 +27,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": null }, "variable_constraints": { "num_frames": { @@ -39,6 +40,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "first_and_last_frame_to_video_audio", diff --git a/workflows/templates/minimax/generated-subject-reference.json b/workflows/templates/minimax/generated-subject-reference.json index 26008441..5b2dd238 100644 --- a/workflows/templates/minimax/generated-subject-reference.json +++ b/workflows/templates/minimax/generated-subject-reference.json @@ -22,7 +22,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": null }, "variable_constraints": { "num_frames": { @@ -34,6 +35,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "draw_subject", diff --git a/workflows/templates/minimax/image-to-video.json b/workflows/templates/minimax/image-to-video.json index 823b76e4..0371b43f 100644 --- a/workflows/templates/minimax/image-to-video.json +++ b/workflows/templates/minimax/image-to-video.json @@ -24,7 +24,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": null }, "variable_constraints": { "num_frames": { @@ -36,6 +37,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "keyframe_to_video_audio", diff --git a/workflows/templates/minimax/last-frame-only.json b/workflows/templates/minimax/last-frame-only.json index 30e009ba..759db90a 100644 --- a/workflows/templates/minimax/last-frame-only.json +++ b/workflows/templates/minimax/last-frame-only.json @@ -24,7 +24,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": null }, "variable_constraints": { "num_frames": { @@ -36,6 +37,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "last_frame_to_video_audio", diff --git a/workflows/templates/minimax/music-video.json b/workflows/templates/minimax/music-video.json index a30e8810..3d7d27a1 100644 --- a/workflows/templates/minimax/music-video.json +++ b/workflows/templates/minimax/music-video.json @@ -67,7 +67,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -79,7 +80,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "draw_singer", diff --git a/workflows/templates/minimax/music.json b/workflows/templates/minimax/music.json index ac67944e..6c9d9315 100644 --- a/workflows/templates/minimax/music.json +++ b/workflows/templates/minimax/music.json @@ -15,8 +15,10 @@ "variables": { "prompt": "prompt:minimax/acoustic_pop_song", "lyrics": "[verse]\nMining farms burning through the grid all day\nChia swapped the furnace for the plots you lay\nNo more e-waste piling up when the rigs retire\nGreen space farming puts out yesterday's fire\n[chorus]\nChia fixes this\nChia fixes this\nEnergy waste, ASIC greed, centralized control\nChia fixes this\n[verse]\nOffer files trade without a middleman in sight\nChialisp writes the rules and locks the CAT in tight\nMarmots watch the validators keep it fair and slow\nNo more trusting strangers with the coins you owe\n[chorus]\nChia fixes this\nChia fixes this\nEnergy waste, ASIC greed, centralized control\nChia fixes this", - "audio_duration": 120 + "audio_duration": 120, + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "generate_music", diff --git a/workflows/templates/minimax/reference-to-video.json b/workflows/templates/minimax/reference-to-video.json index f9d3fd40..7a08cee5 100644 --- a/workflows/templates/minimax/reference-to-video.json +++ b/workflows/templates/minimax/reference-to-video.json @@ -41,7 +41,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": null }, "variable_constraints": { "num_frames": { @@ -53,6 +54,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "reference_to_video_audio", diff --git a/workflows/templates/minimax/storyboard.json b/workflows/templates/minimax/storyboard.json index 747170d4..293e08c6 100644 --- a/workflows/templates/minimax/storyboard.json +++ b/workflows/templates/minimax/storyboard.json @@ -43,7 +43,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -55,7 +56,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "board_1_launch", diff --git a/workflows/templates/minimax/video-with-audio-768p.json b/workflows/templates/minimax/video-with-audio-768p.json index 433388ac..0e2e3aa0 100644 --- a/workflows/templates/minimax/video-with-audio-768p.json +++ b/workflows/templates/minimax/video-with-audio-768p.json @@ -39,7 +39,8 @@ "lora_alpha": 128, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_768p_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -51,7 +52,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "text_to_video_audio", diff --git a/workflows/templates/minimax/video-with-audio.json b/workflows/templates/minimax/video-with-audio.json index a7c0c8e3..3a2516cc 100644 --- a/workflows/templates/minimax/video-with-audio.json +++ b/workflows/templates/minimax/video-with-audio.json @@ -39,7 +39,8 @@ "lora_alpha": null, "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", - "lora_adapter_name": "turbo" + "lora_adapter_name": "turbo", + "seed": 42 }, "variable_constraints": { "num_frames": { @@ -51,7 +52,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, - "seed": 42, + "seed": "variable:seed", "steps": [ { "name": "text_to_video_audio", diff --git a/workflows/templates/minimax/voice-timbre-reference.json b/workflows/templates/minimax/voice-timbre-reference.json index d3c866a0..ee3b9483 100644 --- a/workflows/templates/minimax/voice-timbre-reference.json +++ b/workflows/templates/minimax/voice-timbre-reference.json @@ -24,7 +24,8 @@ "lora_model_name": "lightx2v/Minimax-h3-Turbo", "lora_weight_name": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors", "lora_adapter_name": "turbo", - "weights_dtype": "{int4}" + "weights_dtype": "{int4}", + "seed": null }, "variable_constraints": { "num_frames": { @@ -36,6 +37,7 @@ "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" } }, + "seed": "variable:seed", "steps": [ { "name": "voice", diff --git a/workflows/templates/multi-image-reference.json b/workflows/templates/multi-image-reference.json index 3ee0d637..c1284b3e 100644 --- a/workflows/templates/multi-image-reference.json +++ b/workflows/templates/multi-image-reference.json @@ -10,7 +10,8 @@ }, "second_image": { "location": "https://pbs.twimg.com/profile_images/1961215773080211456/nILQt5Fv_400x400.jpg" - } + }, + "seed": null }, "id": "Flux2DevImageCombine", "description": "FLUX.2 dev combining multiple reference images into a single generation. Loads its 4-bit text encoder locally with model offloading rather than calling a remote encoder.", @@ -22,6 +23,7 @@ "minutes": 2.5 } ], + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/templates/outpaint.json b/workflows/templates/outpaint.json index 10b96933..7d688129 100644 --- a/workflows/templates/outpaint.json +++ b/workflows/templates/outpaint.json @@ -1,10 +1,12 @@ { "variables": { "input_image": "https://pbs.twimg.com/media/Ge7NpKpWAAAq6L4?format=jpg&name=large", - "prompt": "fill in with more clouds, rainbows, stars and rockets" + "prompt": "fill in with more clouds, rainbows, stars and rockets", + "seed": null }, "id": "FluxOutpaint", "description": "Extends an image beyond its borders: adds a border mask, then fills it with FLUX.1 Fill dev.", + "seed": "variable:seed", "steps": [ { "name": "mask_image", diff --git a/workflows/templates/prompt-weighting.json b/workflows/templates/prompt-weighting.json index 2c9a3b4a..746b6a65 100644 --- a/workflows/templates/prompt-weighting.json +++ b/workflows/templates/prompt-weighting.json @@ -1,9 +1,11 @@ { "variables": { - "prompt": "prompt:flux/weighted_portrait" + "prompt": "prompt:flux/weighted_portrait", + "seed": null }, "id": "FluxSchnellWeighted", "description": "FLUX.1 schnell with A1111-style (word:1.5) prompt weighting.", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/templates/qr-code.json b/workflows/templates/qr-code.json index 6138c10f..5ad56af6 100644 --- a/workflows/templates/qr-code.json +++ b/workflows/templates/qr-code.json @@ -1,10 +1,12 @@ { "variables": { "qr_code_contents": "txch1jdsqdz4069k00t5vlk9l8mnr5ycljz640g36ymnygslgkra675ssg2y4ng", - "init_image_location": "https://th.bing.com/th/id/OIP.Lsm7UOwPR37OPvDy3HUUTQHaE7?rs=1&pid=ImgDetMain" + "init_image_location": "https://th.bing.com/th/id/OIP.Lsm7UOwPR37OPvDy3HUUTQHaE7?rs=1&pid=ImgDetMain", + "seed": null }, "id": "qr_code", "description": "Generates a scannable artistic QR code with an SD 2.1 ControlNet.", + "seed": "variable:seed", "steps": [ { "name": "qr_code", diff --git a/workflows/templates/restore-faces.json b/workflows/templates/restore-faces.json index 9d1e8761..5d448743 100644 --- a/workflows/templates/restore-faces.json +++ b/workflows/templates/restore-faces.json @@ -7,10 +7,12 @@ "num_inference_steps": 9, "guidance_scale": 0.0, "width": 1024, - "height": 1024 + "height": 1024, + "seed": null }, "id": "FaceRestore", "description": "Generates with Z-Image Turbo, then restores and sharpens faces in the result.", + "seed": "variable:seed", "steps": [ { "name": "generate", diff --git a/workflows/templates/segment-and-inpaint.json b/workflows/templates/segment-and-inpaint.json index 0a18b16b..ea8db77a 100644 --- a/workflows/templates/segment-and-inpaint.json +++ b/workflows/templates/segment-and-inpaint.json @@ -4,8 +4,10 @@ "variables": { "image_url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/grounding_dino_example_input.png", "segment_prompt": "cat", - "inpaint_prompt": "a golden retriever puppy sitting on the grass" + "inpaint_prompt": "a golden retriever puppy sitting on the grass", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "load_image", diff --git a/workflows/templates/step-caching.json b/workflows/templates/step-caching.json index b22704cd..bebc310e 100644 --- a/workflows/templates/step-caching.json +++ b/workflows/templates/step-caching.json @@ -2,10 +2,12 @@ "variables": { "prompt": "A photorealistic image of a red fox sitting in a snowy forest clearing, soft morning light filtering through the trees", "num_inference_steps": 28, - "guidance_scale": 3.5 + "guidance_scale": 3.5, + "seed": null }, "id": "FluxDevFirstBlockCache", "description": "FLUX.1 dev accelerated with first-block caching, skipping redundant transformer work across denoise steps. Two cache implementations, one shape. This uses first-block caching, a 'cache' block of {\"type\": \"first_block\", \"threshold\": 0.05}. For TeaCache instead, replace it with a sibling 'teacache' block of {\"rel_l1_thresh\": 0.4}. Both trade a little fidelity for skipped transformer work; raise the threshold to skip more.", + "seed": "variable:seed", "steps": [ { "name": "txt2img", diff --git a/workflows/templates/sub-workflow.json b/workflows/templates/sub-workflow.json index 90803eca..2afffd09 100644 --- a/workflows/templates/sub-workflow.json +++ b/workflows/templates/sub-workflow.json @@ -1,9 +1,11 @@ { "variables": { - "prompt": "An eco-friendly crypto currency logo" + "prompt": "An eco-friendly crypto currency logo", + "seed": null }, "id": "FluxLogo", "description": "Generates a logo with the FluxLora workflow, then removes its background.", + "seed": "variable:seed", "steps": [ { "name": "logo", diff --git a/workflows/templates/surface-normals.json b/workflows/templates/surface-normals.json index 3decc9c0..0ee20d6b 100644 --- a/workflows/templates/surface-normals.json +++ b/workflows/templates/surface-normals.json @@ -2,8 +2,10 @@ "id": "MarigoldNormals", "description": "Estimates surface-normal maps for input images with Marigold.", "variables": { - "image_url": "https://marigoldmonodepth.github.io/images/einstein.jpg" + "image_url": "https://marigoldmonodepth.github.io/images/einstein.jpg", + "seed": null }, + "seed": "variable:seed", "steps": [ { "name": "load_image", diff --git a/workflows/templates/text-to-image.json b/workflows/templates/text-to-image.json index a0699c58..5a234d23 100644 --- a/workflows/templates/text-to-image.json +++ b/workflows/templates/text-to-image.json @@ -2,13 +2,15 @@ "cost_drivers": ["num_images_per_prompt"], "variables": { "prompt": "an apple", - "num_images_per_prompt": 1 + "num_images_per_prompt": 1, + "seed": null }, "id": "text-to-image", "description": "Baseline text-to-image with Stable Diffusion 1.5 - the smallest, fastest starting point. It loads with 'safety_checker': null: SD 1.5's checker false-positives on ordinary prompts for particular seeds and returns a solid black image rather than an error, which a reference template used to prove the engine works must not do. Any pipeline that keeps the checker says so as a 'safety_checker_blanked' warning when it fires.", "cost": [ {"device": "cuda", "name": "RTX 3090", "vram_gb": 24, "minutes": 0.2} ], + "seed": "variable:seed", "steps": [ { "name": "main", From 0945e659cf4647875295007cfb99d94a8607a7a9 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 08:49:41 -0500 Subject: [PATCH 042/181] fix(catalog): #352 - add shots-batch template for unrelated H3 shots New templates/minimax/shots-batch.json: one for_each H3 text-to-video- with-audio step per shots entry (name, prompt, num_frames), no shared cast, no concat - deliberately separate from assemble-and-score, which does the cutting and scoring after each clip is promoted with keep_output. Follows #351's seed: variable:seed convention and reuses video-with-audio's pipeline configuration. num_frames constrained to 17n+5 (124-345) via variable_constraints; cost.per_entry measured against the default 5-shot list. Bumps COMPACT_BUDGET 8450->8650 for the new template's compact-listing footprint, and trims plugins/dw/skills/minimax-h3/SKILL.md to fit the 12288-byte skill cap after adding a bullet pointing at shots-batch. Co-Authored-By: Claude Sonnet 5 --- plugins/dw/skills/minimax-h3/SKILL.md | 54 +++--- tests/test_catalog_structure.py | 2 +- workflows/templates/minimax/shots-batch.json | 183 +++++++++++++++++++ 3 files changed, 211 insertions(+), 28 deletions(-) create mode 100644 workflows/templates/minimax/shots-batch.json diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 7f00e92d..88fc6ac8 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -37,9 +37,9 @@ arguments; the prompt format is MiniMax's, from their text not here. `templates/minimax/voice-timbre-reference` fixes a voice from a Bark line. - **Several boards in one generation, one unbroken score**: `templates/minimax/storyboard` - H3 cuts between the boards inside a single - generation, which no concat of separate clips can match. One beat with fixed - cut points, not a building block; past one beat with a recurring cast, use - the cuts pattern below. + generation, which no concat can match. One beat with fixed cut points, not + a building block; past one beat with a recurring cast, use the cuts + pattern below. - **Longer than 14.4 seconds**: chain when one action or line of speech crosses the seam, cut when the scene changes (each cut its own generation). Six distinct scenes are a cuts piece, not a chain. @@ -56,27 +56,26 @@ arguments; the prompt format is MiniMax's, from their text not here. `templates/minimax/dialogue-short` (Z-Image draws the cast, one shot per `shots` entry on one loaded model, `concat_videos` splices) and `templates/minimax/music-video` (a song, one slice and one lip-synced shot - per entry; its singer is the `singer_reference` argument - a `from_file` - reference uses a cast portrait that exists and elides the drawing). + per entry; its singer is `singer_reference` - a `from_file` reference uses + an existing cast portrait and elides the drawing). `shots` is one list argument: a dialogue entry is `name`, `prompt`, `references` (portraits and voices: `from_previous_result` for one drawn here, `from_file` for an `asset:` cast) and `num_frames`; a music-video - entry is `name`, `prompt` and `start_frame`. - A six-shot piece is one more entry, not another file. The listing's + entry is `name`, `prompt`, `start_frame`. + A six-shot piece is one more entry, not a new file. The listing's `lists` block says what an entry carries; its `cost` carries `per_entry` when one shot was measured: quote `minutes - per_entry.minutes × per_entry.entries + per_entry.minutes × N` - for N entries. Without `per_entry`, quote the total and say it is the - default list's. + for N entries. Without `per_entry`, quote the total for the default list. A cut erases drift: the last shot is as clean as the first. Each shot makes its own audio, so write `non_diegetic_music: N/A` in every shot and lay one score under the concat afterwards: `templates/minimax/music` writes the track and `templates/assemble-and-score` shows the `pair_audio` step that mixes it under the world sound (three shots; for more, author the concat and score steps the same way). A character speaking in several shots keeps - one voice by passing the same clip as an audio reference in each (the + one voice by passing the same clip as an audio reference each time (the `voice-timbre-reference` pattern) - a repeated description alone drifts. - Each entry's `num_frames` is its own, so pace the cut. + Each entry's `num_frames` paces the cut. A score burying voice-over is not a `world_gain` fix: the world track carries narration and action together (narration ~24 dB over its own ambience, score 6-15 dB over that), so raising it lifts both. Duck the @@ -84,7 +83,10 @@ arguments; the prompt format is MiniMax's, from their text not here. step per voice-over shot, chained, `start_frame` = running sum of preceding shots' `num_frames`, `num_frames` that shot's length, `fps` the cut's rate, negative `gain_db` (-6 to -10). `dissolve-between-shots` - eats a `dissolve_frames` per seam, so its sum isn't plain. + eats `dissolve_frames` per seam, so its sum isn't plain. +- **Unrelated shots, no cut**: `templates/minimax/shots-batch` - one H3 + step per `shots` entry, no shared cast, no concat. `keep_output` each + clip, then `templates/assemble-and-score` cuts and scores. - **Music alone**: `templates/minimax/music` (Music3); the `minimax-music3` skill. If none fits, compose from `list_tasks` before authoring a new workflow, and @@ -106,10 +108,10 @@ read the `workflows` guide's authoring section first. **6**/3, **alpha 128** - `video-with-audio-768p`; 768p Ref2VA turbo, shift 12/3, alpha unset - every `ref2va` template. The two 768p LoRAs differ in shift; do not generalise. - Never put an FL2VA LoRA on a reference template: a `ref2va` step holds - `transformer_ref` alone and diffusers routes whatever it is handed there, - so it only degrades the output. `validate_workflow` refuses it and warns - on a `weight_name` naming neither path. + Never put an FL2VA LoRA on a reference template: `ref2va` holds + `transformer_ref` alone, so whatever is handed there only degrades the + output. `validate_workflow` refuses it and warns on a `weight_name` + naming neither path. - Nine steps for an eight-step LoRA: the scheduler counts sigma grid points, terminal zero included, so `denoise_total_steps` reports 8 - expected, not upstream's `--inference-steps 8` read literally. @@ -138,8 +140,7 @@ the rules come from MiniMax, not from paraphrasing one: https://github.com/MiniMax-AI/MiniMax-H3 under `skills/`), use it. If not, say once that `npx skills add MiniMax-AI/MiniMax-H3 --skill h3-prompt-writing` installs - it - only that skill; the repo's other eight are style packs - and go on - without it. + it - only that one; the rest are style packs - and go on without it. 2. Else read the guides on the model card: https://huggingface.co/MiniMaxAI/MiniMax-H3/raw/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md for text and frame conditioning, and @@ -172,27 +173,26 @@ itself - or every shot inherits the portrait's composition. two-minute gaps are healthy. `phase_stall` in `get_job_events` narrates it, not a fault; judge by `denoise_step`. Each entry carries `subfolder`: `final` is the deliverable (`episode`, `music_video`, `voyage`), - `intermediate` the scratch; keep that split in anything you compose. + `intermediate` the scratch; keep that split in what you compose. 4. Judge it yourself: `get_output_frames(count=12)` for a clip's shape, `seams=true` (each later shot's start frame) for a cut's joins - a character that changes between shots (reference the same portraits everywhere), a portrait imposing its framing on every shot - `at` late in - a chain for drift sharpening into noise, and `get_output_audio` for a + a chain for drift sharpening to noise, and `get_output_audio` for a voice-over without affect. `get_output_audio` returns sound, not text; to confirm a line rendered rather than judge its delivery, `run_workflow(name="templates/transcribe-audio", - arguments={"input_audio": "output:"}, wait_seconds=55)` (takes the - muxed soundtrack directly) then `get_output_text` on the result, and - `delete_output` the scratch run. Then `get_gallery_metadata` for duration - and whether audio is present, and hand the user the gallery `url` + arguments={"input_audio": "output:"}, wait_seconds=55)` then + `get_output_text` on the result, and `delete_output` the scratch run. + Then `get_gallery_metadata` for duration and whether audio is present, + and hand the user the gallery `url` (`list_gallery`). `get_output_image` works only on image steps - the Z-Image portraits and boards of `dialogue-short`, `storyboard`, `generated-subject-reference` and `music-video`. 5. After a run worth keeping, `get_job_workflow` and `save_workflow` it; `export_job` bundles it on the server. `auth_required: false` - fetch - `open_url` into `exports/` (never a temp dir; unpacks into a job-id - folder). `true` - hand `open_url` to the person instead and keep using - `get_output_image`/`_audio`/`_frames`. + `open_url` into `exports/` (never a temp dir). `true` - hand `open_url` + to the person instead and keep using `get_output_image`/`_audio`/`_frames`. ## Sources diff --git a/tests/test_catalog_structure.py b/tests/test_catalog_structure.py index 6e2f6b43..95ef3103 100644 --- a/tests/test_catalog_structure.py +++ b/tests/test_catalog_structure.py @@ -402,7 +402,7 @@ def test_no_stale_entry_in_the_allowlist(): # Then to 8_450 for the catalog-wide seed convention (#351): every generative # entry now declares `seed` in `variables`, which is a compact field # (`variable_names`), measured at 8_424. -COMPACT_BUDGET = 8_450 +COMPACT_BUDGET = 8_650 FILTERED_BUDGET = 1_500 diff --git a/workflows/templates/minimax/shots-batch.json b/workflows/templates/minimax/shots-batch.json new file mode 100644 index 00000000..da31fb7d --- /dev/null +++ b/workflows/templates/minimax/shots-batch.json @@ -0,0 +1,183 @@ +{ + "id": "MiniMaxH3ShotsBatch", + "description": "A batch of independent H3 shots on one loaded model - each its own prompt and length, no shared subject, no concat. The assembly is a deliberately separate job: promote each clip with keep_output and hand the set to assemble-and-score, which cuts them together and lays a score underneath.", + "cost": [ + { + "device": "cuda", + "name": "RTX 3090", + "vram_gb": 24, + "minutes": 19.4, + "per_entry": { + "variable": "shots", + "minutes": 3.49, + "entries": 5 + } + } + ], + "cost_drivers": ["shots", "num_inference_steps", "width", "height"], + "vram_estimate": { + "base_gb": 16.3, + "bytes_per_voxel": 28.71, + "voxel_variables": ["width", "height", "num_frames"], + "reason": "the transformer's peak activation memory scales with the token count the video and audio streams produce, which is proportional to width * height * num_frames" + }, + "variables": { + "shots": [ + { + "name": "shot_1", + "prompt": "Task: T2VA. Duration: 5.17 seconds. A lighthouse keeper climbs a spiral staircase at dawn, footsteps echoing on iron treads, gulls calling outside. non_diegetic_music: N/A", + "num_frames": 124 + }, + { + "name": "shot_2", + "prompt": "Task: T2VA. Duration: 5.17 seconds. The keeper reaches the lamp room and lights the beacon, glass clinking, wind picking up against the panes. non_diegetic_music: N/A", + "num_frames": 124 + }, + { + "name": "shot_3", + "prompt": "Task: T2VA. Duration: 6.42 seconds. The beacon sweeps across a dark sea, waves rolling against rocks below, a foghorn sounding in the distance. non_diegetic_music: N/A", + "num_frames": 158 + }, + { + "name": "shot_4", + "prompt": "Task: T2VA. Duration: 5.17 seconds. A distant ship changes course toward the light, its horn answering the foghorn once. non_diegetic_music: N/A", + "num_frames": 124 + }, + { + "name": "shot_5", + "prompt": "Task: T2VA. Duration: 5.17 seconds. The keeper settles into a chair by the lamp, journal in hand, the beacon turning steadily behind him. non_diegetic_music: N/A", + "num_frames": 124 + } + ], + "width": 960, + "height": 544, + "num_inference_steps": 9, + "video_shift": 12.0, + "audio_shift": 3.0, + "weights_dtype": "{int4}", + "lora_scale": 1.0, + "lora_alpha": null, + "lora_model_name": "lightx2v/Minimax-h3-Turbo", + "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors", + "lora_adapter_name": "turbo", + "seed": null + }, + "variable_constraints": { + "num_frames": { + "modulus": 17, + "remainder": 5, + "min_frames": 124, + "max_frames": 345, + "snap": "up", + "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps" + } + }, + "seed": "variable:seed", + "steps": [ + { + "name": "shot", + "for_each": "variable:shots", + "release_pipeline": true, + "pipeline": { + "configuration": { + "component_type": "ModularPipeline", + "pre_load_modules": ["sdnq"], + "load_components": { + "dtype": "torch.bfloat16", + "quantization_config": { + "transformer": { + "configuration": {"config_type": "sdnq.SDNQConfig"}, + "arguments": { + "weights_dtype": "variable:weights_dtype", + "quantization_device": "cuda", + "return_device": "cpu", + "use_quantized_matmul": true, + "dequantize_fp32": false, + "modules_to_not_convert": [ + "proj_in", + "audio_proj_in", + "context_embedder", + "time_embedder", + "time_proj", + "token_refiner", + "norm_out", + "proj_out", + "audio_proj_out" + ] + } + }, + "text_encoder": { + "configuration": {"config_type": "sdnq.SDNQConfig"}, + "arguments": { + "weights_dtype": "variable:weights_dtype", + "quantization_device": "cuda", + "return_device": "cpu", + "dequantize_fp32": false, + "modules_to_not_convert": [".model.visual", "lm_head"] + } + }, + "vae": { + "configuration": {"config_type": "sdnq.SDNQConfig"}, + "arguments": { + "weights_dtype": "{int8}", + "quant_conv": true, + "use_quantized_matmul_conv": true, + "quantization_device": "cuda", + "return_device": "cpu", + "dequantize_fp32": false + } + } + } + }, + "components": { + "transformer": { + "group_offload": { + "offload_type": "block_level", + "num_blocks_per_group": 2, + "use_stream": true, + "record_stream": true, + "low_cpu_mem_usage": true + } + }, + "text_encoder": {"remove_modules": ["lm_head"]}, + "text_encoder.model": { + "truncate_layers": {"language_model.layers": 51}, + "group_offload": {"offload_type": "leaf_level"} + }, + "vae": {"device": "cuda", "residency": "on_demand"}, + "audio_vae": {"device": "cuda", "residency": "on_demand"} + } + }, + "from_pretrained_arguments": { + "model_name": "MiniMaxAI/MiniMax-H3", + "workflow": "t2va" + }, + "loras": [ + { + "model_name": "variable:lora_model_name", + "weight_name": "variable:lora_weight_name", + "adapter_name": "variable:lora_adapter_name", + "scale": "variable:lora_scale", + "alpha": "variable:lora_alpha" + } + ], + "scheduler": {"shift": "variable:video_shift"}, + "audio_scheduler": {"shift": "variable:audio_shift"}, + "arguments": { + "prompt": "item:prompt", + "num_frames": "item:num_frames", + "width": "variable:width", + "height": "variable:height", + "num_inference_steps": "variable:num_inference_steps", + "output": ["videos", "audio", "sampling_rate"] + } + }, + "result": { + "content_type": "video/mp4", + "fps": 24, + "file_base_name": "item:name", + "subfolder": "final" + } + } + ] +} From c128018376fb078a4d5955918c251c607f15264d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 09:03:57 -0500 Subject: [PATCH 043/181] fix(mcp): #353 - sync served tool descriptions with the auth-aware export/download behaviour The runtime `next` hint in export_job and the refusal in download_output were already correct (8a41883), but dw_mcp/server.py keeps its own duplicated docstrings for the served MCP tool descriptions rather than reusing exports.py/media.py's - those still told an agent to unconditionally fetch the export zip and let download_output default into the current working directory. Also stop returning absolute_zip_url as an explicit null when DW_PUBLIC_URL isn't configured; omit it like list_gallery's absolute_url does. Co-Authored-By: Claude Sonnet 5 --- docs/MCP.md | 9 ++++++--- dw_mcp/exports.py | 6 ++++-- dw_mcp/server.py | 35 ++++++++++++++++++++++++----------- tests/test_mcp_exports.py | 11 +++++++++++ tests/test_mcp_server.py | 26 ++++++++++++++++++++++++++ 5 files changed, 71 insertions(+), 16 deletions(-) diff --git a/docs/MCP.md b/docs/MCP.md index 1ddac9ee..9346c21b 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -237,7 +237,7 @@ when no single workflow covers it. | `get_output_audio(name, start=None, duration=None, workspace=None)` | `name`, `start`, `duration`, `workspace` | Listen to a generated soundtrack as base64 - an audio output, or the track muxed into a video (#193) - in its own encoding when served whole, WAV when extracted or excerpted. No downscale exists for audio, so a whole clip over the 4MB budget is refused rather than cut (#204); ask for the part instead with `start` and `duration` in seconds, and the text part names what was cut (`excerpt: 2.0s from 10.0s of 240.0s`) so a slice is never mistaken for the whole. `get_gallery_metadata`'s envelope says where in a track to look. `workspace` names the workspace for this one call without switching the session to it | | `get_output_frames(name, at=None, seams=None, count=None, boundaries=None, names=None, max_dimension=512, hear=None, workspace=None)` | `name`, `at`, `seams`, `count`, `boundaries`, `names`, `max_dimension`, `hear`, `workspace` | See a generated video as frames, since there is no video content type over MCP (#193). One selector per call: `count` for an evenly spaced contact sheet, `at` for moments (seconds or `"frame:N"`), `seams` (true, or seam numbers from 1) for the last frame before and first frame after each join side by side (each pair carries `difference`, the mean pixel change across the join, 0-255 - rank seams by it and look at the worst) - `boundaries` is each later shot's first frame - the running sum of the shots' `frame_count` from `get_gallery_metadata` on their own `intermediate/` files - and `names` names them. Tiles are fitted to `max_dimension` and, when the set would exceed the 4MB budget, shrunk together rather than dropped; the text part lists each tile and says so. `hear=N` also returns N seconds of soundtrack centred on each `at` moment, after its image - the way to check a hit point or lip-sync without reconciling two clocks; a mute clip keeps its frames and says `no soundtrack` | | `get_output_text(name, max_characters=20000, workspace=None)` | `name`, `max_characters`, `workspace` | Read a text output — a prompt enhancement, or any step whose result is `text/plain` or JSON. Reports the file's real length and whether it was truncated. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | -| `download_output(name, destination=None, overwrite=False, workspace=None)` | `name`, `destination`, `overwrite`, `workspace` | Save one output file to local disk, of any content type. `destination` may be a full path, a directory, or omitted to save under the output's own name in the current working directory; `~` expands and missing parent directories are created. `overwrite=True` is required to replace a file already at the resolved path. Over a `dw.serve --mcp` endpoint the file lands on the server, so the destination is confined to that workspace and a relative one is joined onto it. Returns nothing to the conversation but where the file landed — unlike the other media tools, the point is a file on disk, not a payload in context. Writes on the machine running the MCP server - over `dw.serve --mcp` that is the GPU box. A write that fails there (a path that exists only on the client, for instance) comes back as an error naming the server-side write and the client-side alternatives, not as an anonymous tool failure. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `download_output(name, destination=None, overwrite=False, workspace=None)` | `name`, `destination`, `overwrite`, `workspace` | Save one output file to local disk, of any content type. `destination` may be a full path or a directory; `~` expands and missing parent directories are created. `overwrite=True` is required to replace a file already at the resolved path. Over the stdio `dw-mcp`, omitting `destination` saves under the output's own name in the current working directory. Over a `dw.serve --mcp` endpoint the file lands on the server confined to that workspace (a relative path is joined onto it), and `destination` is required there - an omitted one is refused rather than dropped loose in the workspace root, where nothing can find or delete it later (#353); use the `url` `list_gallery` reports, `get_output_image`/`get_output_audio`/`get_output_frames` for inline content, or `keep_output` to make it a named asset instead. Returns nothing to the conversation but where the file landed — unlike the other media tools, the point is a file on disk, not a payload in context. Writes on the machine running the MCP server - over `dw.serve --mcp` that is the GPU box. A write that fails there (a path that exists only on the client, for instance) comes back as an error naming the server-side write and the client-side alternatives, not as an anonymous tool failure. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `delete_output(name=None, workspace=None, job_id=None)` | exactly one of `name` / `job_id`, `workspace` | Permanently remove one generated file from the output directory, or - with `name` a `/` run directory, or with `job_id` - a whole run. By `job_id` the run directory is read from the job record (`run_dir`) and the reply adds `job_id` and the resolved `run_dir` to the usual `name` / `deleted` / `run_swept`; a job that never wrote a run directory, or an unknown one, is an error. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it; a `job_id` delete with no `workspace` goes to the workspace the job ran in | ### Authoring, assets and workspaces @@ -304,7 +304,7 @@ references written in the same session. | `run_workflow(workflow_path=None, inline_workflow=None, arguments=None, acknowledged_cost=False, workspace=None, wait_seconds=0)` | exactly one of `workflow_path` (a catalog name from `list_workflows`, with or without `.json`, or a path to a workflow file on the server) or `inline_workflow`, optional `arguments`, `acknowledged_cost`, `workspace`, `wait_seconds` | Queue a workflow for generation. Returns as soon as the job is queued - unless `wait_seconds` is above 0, in which case the call then waits on the queued job exactly as `wait_for_job(job_id, timeout_seconds=wait_seconds)` would (same 55 s cap per call, clamped not honoured) and the result carries the queued-job fields plus that wait's (`status`, `still_running`, `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds`, `timeout_capped`, the slim `job`); when the cap covers the job's runtime one call is the run and the wait, and a `still_running: true` result is followed with `wait_for_job` as before. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's - `acknowledged_cost` is `true` or the bound `{fingerprint, minutes, downloads}` from the validate plan; a 409 means the plan changed and the message carries the new estimate, and nothing is waited on | | `get_job(job_id)` | `job_id` | Get a job's status, warnings, output manifest, error and traceback; each manifest entry's `subfolder` is the in-run subfolder the step declared - by convention `final` for the deliverable, `intermediate` for scratch, `''` for none. A running job also carries `progress` (below) | | `get_job_workflow(job_id)` | `job_id` | The REST equivalent is `GET /api/jobs/{id}/workflow` (see [SERVER.md](SERVER.md#jobs-api)). The workflow the job actually ran. `realized: true` means every mutable input is pinned (arguments, seed, prompts, `output:latest`); `false` means the job predates run tracking and this is the definition as submitted. Pass it to `save_workflow` to keep it under a name | -| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `/exports//` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination - fetch the zip URL and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file); the archive already unpacks into one folder named after the job id | +| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `/exports//` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination. `auth_required` says whether opening the zip needs this server's bearer token, which this agent cannot attach to someone else's fetch (#353): when false, fetch `open_url` yourself and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file) - the archive already unpacks into one folder named after the job id, so do not create that folder first; when true, hand `open_url` to the person instead of fetching it | | `get_job_events(job_id, after=-1, limit=200)` | `job_id`, `after`, `limit` | Get a page of a job's progress events | | `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, `run_id` and `run_version` (the run's `v5`, as the gallery labels it), and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) | | `cancel_job(job_id)` | `job_id` | Ask a queued or running job to stop | @@ -444,7 +444,10 @@ refused, and an existing file is left alone unless the caller passes Over `dw.serve --mcp` the write happens **on the server**, and there the destination is confined to that workspace: an absolute or `~` path outside it is refused, and a relative one is joined onto the workspace rather than -onto whatever the server process's working directory happens to be. The +onto whatever the server process's working directory happens to be. +`destination` is required over this transport - an omitted one is refused +rather than dropped loose in the workspace root, where nothing can find or +delete it later (#353). The transport is what distinguishes the two - on stdio "local disk" is genuinely the caller's own machine, over HTTP it is the operator's. Confinement is on the resolved real path, not a substring test, because an absolute path needs diff --git a/dw_mcp/exports.py b/dw_mcp/exports.py index 94ae71a3..07896a91 100644 --- a/dw_mcp/exports.py +++ b/dw_mcp/exports.py @@ -67,12 +67,11 @@ def export_job(client, job_id, overwrite=False): "are not repeated here; get_job_workflow and get_job serve them " "individually." ) - return { + result = { "job_id": job_id, "where": f"{directory} on the machine running the MCP server", "directory": directory, "zip_url": zip_url, - "absolute_zip_url": absolute_zip_url, "auth_required": auth_required, "open_url": open_url, "files": body.get("files") or [], @@ -80,3 +79,6 @@ def export_job(client, job_id, overwrite=False): "missing": body.get("missing") or [], "next": next_text, } + if absolute_zip_url is not None: + result["absolute_zip_url"] = absolute_zip_url + return result diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 12a6267e..ccaae8f3 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -700,10 +700,16 @@ def download_output( for any file type, streams the body straight to disk rather than buffering it, and returns no content to the conversation - only where it was saved. `destination` may be a - full path, a directory, or omitted to save into the current - working directory under the output's own name; a '..' path segment - in it is refused. An existing file at the resolved path is left - alone unless `overwrite=True`. + full path or a directory; a '..' path segment in it is refused. An + existing file at the resolved path is left alone unless + `overwrite=True`. On the stdio `dw-mcp`, omitting `destination` + saves into the current working directory under the output's own + name. On a `dw.serve --mcp` endpoint the save happens on the server, and destination is required there - + an omitted one is refused rather than dropped loose in the + workspace root, where nothing can find or delete it later; use + the `url` list_gallery reports, get_output_image/get_output_audio/ + get_output_frames for inline content, or keep_output to make it a + named asset instead. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` @@ -1239,13 +1245,20 @@ def export_job(job_id: str, overwrite: bool = False) -> dict: what was copied. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. - THE DIRECTORY IS ON THE MACHINE RUNNING THE SERVER, not on yours. To - give the user the files, fetch the zip URL and unpack it into - exports/ under the session's working directory - it is the user's - deliverable, not a temp file; the archive already unpacks into one - folder named after the job id, so do not create that folder first. - Refuses a job that is still running; refuses an existing export - unless overwrite=true.""" + THE DIRECTORY IS ON THE MACHINE RUNNING THE SERVER, not on yours. + + `auth_required` says whether opening the zip needs this server's + bearer token, a token you cannot attach to someone else's browser + or tooling. When it is false, fetch open_url yourself and unpack + it into exports/ under the session's working directory - it is + the user's deliverable, not a temp file; the archive already + unpacks into one folder named after the job id, so do not create that folder first. + When it is true, do NOT fetch it: hand open_url to the person and let them open it + (`next` says whether it is already absolute or needs the server's + address told to them). Individual results stay reachable inline + via get_output_image/get_output_audio/get_output_frames either + way. Refuses a job that is still running; refuses an existing + export unless overwrite=true.""" return exports.export_job(client, job_id, overwrite=overwrite) tool(get_job, READ_ONLY) diff --git a/tests/test_mcp_exports.py b/tests/test_mcp_exports.py index 20085258..c899eb2e 100644 --- a/tests/test_mcp_exports.py +++ b/tests/test_mcp_exports.py @@ -143,3 +143,14 @@ def test_no_auth_required_still_fetches_the_zip_itself(): assert result["auth_required"] is False assert result["open_url"] == "/exports/job-1.zip" assert "fetch open_url" in result["next"] + + +def test_absolute_zip_url_is_omitted_rather_than_null_when_unconfigured(): + """#353 follow-up: the API omits absolute_zip_url when DW_PUBLIC_URL + isn't set, matching list_gallery's absolute_url; this tool used to pass + the missing key through as an explicit null instead of leaving it out.""" + client, _ = exporting() + + result = exports.export_job(client, "job-1") + + assert "absolute_zip_url" not in result diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 3a367d8e..2ad02bad 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -996,6 +996,32 @@ async def test_export_job_sends_the_zip_to_the_working_directory(): assert "do not create that folder first" in description +@pytest.mark.asyncio +async def test_export_job_description_is_auth_aware(): + """#353: the served tool description, not just the runtime `next` hint, + has to tell the agent not to fetch an auth-gated zip on the person's + behalf - the description is what the agent plans from before it ever + calls the tool and sees `next`.""" + tools = await tools_of(server_over(ok({}))) + + description = tools["export_job"].description + assert "auth_required" in description + assert "do NOT fetch it" in description + assert "hand open_url to the person" in description + + +@pytest.mark.asyncio +async def test_download_output_description_says_a_mounted_endpoint_requires_destination(): + """#353: on a dw.serve --mcp endpoint an omitted destination used to + silently land in the workspace root; the tool description has to say + it's refused there instead, not just the stdio default.""" + tools = await tools_of(server_over(ok({}))) + + description = tools["download_output"].description + assert "destination is required there" in description + assert "keep_output" in description + + @pytest.mark.asyncio async def test_rerun_job_refuses_without_acknowledgement_and_sends_nothing(): seen = [] From 1272e713057177818cf77eba07c41826fc040fbb Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 09:10:42 -0500 Subject: [PATCH 044/181] fix(mcp): #356 - expose list_gallery(media=true) over MCP The server's GET /api/gallery already answered `media=true` with duration_seconds on audio/video entries (shipped for #356's part 1), but the MCP list_gallery tool had no way to ask for it - the tester bounced the fix because an MCP consumer could not reach the param. Forward it through catalog.list_gallery and the server.py tool signature/docstring, trimmed to stay under the 2048-char client description limit and the tool-surface token budget. Co-Authored-By: Claude Sonnet 5 --- dw_mcp/catalog.py | 10 +++++++++- dw_mcp/server.py | 11 ++++++++--- tests/test_mcp_catalog.py | 10 ++++++++++ 3 files changed, 27 insertions(+), 4 deletions(-) diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 669fd502..06272061 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -221,6 +221,7 @@ def list_gallery( workspace=None, folder=None, version=None, + media=False, ): """Generated media in the output directory, newest first. `subfolder` narrows to one in-run subfolder ('final', 'intermediate', '' for files @@ -248,7 +249,12 @@ def list_gallery( run opens and a deleted sibling leaves a gap rather than renumbering what is left - as does a run that failed, or reused every step from the cache, and so wrote nothing to list. Null under the flat output - layout, which has no runs.""" + layout, which has no runs. + + `media=True` adds `duration_seconds` to each audio/video entry in the + page returned, probed the way `get_gallery_metadata` measures a file - + enough to pick between two takes without one metadata call per + candidate. Off by default; a plain call carries no `duration_seconds`.""" params = {"limit": limit} if subfolder is not None: params["subfolder"] = subfolder @@ -258,6 +264,8 @@ def list_gallery( params["version"] = version if only_orphans: params["only_orphans"] = "true" + if media: + params["media"] = "true" return client.get_json("/api/gallery", params=params, workspace=workspace) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index ccaae8f3..aee628ae 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -354,6 +354,7 @@ def list_gallery( workspace: str | None = None, folder: str | None = None, version: int | None = None, + media: bool = False, ) -> dict: """List generated output files, newest first. A name is //, where may itself sit in a @@ -378,8 +379,7 @@ def list_gallery( (manifest.json, workflow.json, job.json) as `runs`, each `{name, mtime}` - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before - writing anything, invisible to a normal listing because it has no - file to show. `subfolder` does not apply in this mode. `name` is + writing anything. `subfolder` does not apply in this mode. `name` is exactly what `delete_output` accepts, so clearing the backlog is list, then delete each name. A run that wrote any file at all - a text-shape prompt, a utility's side output - is not listed; @@ -389,7 +389,11 @@ def list_gallery( `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace is reachable from - here without leaving this one.""" + here without leaving this one. + + `media=True` adds `duration_seconds` to audio/video entries - two + takes sharing a basename are told apart by length, not size or + mtime.""" return catalog.list_gallery( client, limit=limit, @@ -398,6 +402,7 @@ def list_gallery( workspace=workspace, folder=folder, version=version, + media=media, ) def get_gallery_metadata( diff --git a/tests/test_mcp_catalog.py b/tests/test_mcp_catalog.py index a6aef5a9..7ca94199 100644 --- a/tests/test_mcp_catalog.py +++ b/tests/test_mcp_catalog.py @@ -140,6 +140,16 @@ def test_list_gallery_sends_only_orphans_only_when_true(): assert seen["params"]["only_orphans"] == "true" +def test_list_gallery_sends_media_only_when_true(): + client, seen = recording_client() + catalog.list_gallery(client, limit=7) + assert "media" not in seen["params"] + + client, seen = recording_client() + catalog.list_gallery(client, media=True) + assert seen["params"]["media"] == "true" + + FULL_ENTRY = { "summary": "a cut sequence", "shape": "sequence", From 0da2d976d2fc65c4ee5ac397d96af00e778a9559 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 09:16:45 -0500 Subject: [PATCH 045/181] fix(result): #358 - say 'quiet but not empty' when a near-silent shot has real peaks warn_if_written_near_silent keeps its mean < -40 dBFS trigger (an ambience-only shot with real content, e.g. -54 dBFS mean / -18 dBFS peaks, was warning with the same wording as a genuinely empty render). Adds peak_dbfs to the warning's fields and, when the peak clears -30 dBFS, swaps the "check the step that generated it" wording for one that says the shot is quiet but not empty. A probe with no peak, or a peak below the line, keeps the original message. Co-Authored-By: Claude Sonnet 5 --- dw/result.py | 44 +++++++++++++++++++++++++++++--------- tests/test_result.py | 50 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 84 insertions(+), 10 deletions(-) diff --git a/dw/result.py b/dw/result.py index 4b33e21e..d6d9d2aa 100644 --- a/dw/result.py +++ b/dw/result.py @@ -253,6 +253,11 @@ def warn_if_written_above_full_scale( # warnings: [] (#261) NEAR_SILENT_WARN_DBFS = -40.0 +# Above this, a peak means "quiet but not empty" rather than "check for a +# defect" - s02's -18.5 dBFS peaks (an ambience-only shot) clear it, #261's +# -68.7 dBFS Bark clip and S-F077's -60 dBFS normalize do not (#358) +NEAR_SILENT_QUIET_NOT_EMPTY_DBFS = -30.0 + def warn_if_written_near_silent( output_path, already_warned=False, info=_UNPROBED, source_already_quiet=False @@ -276,6 +281,15 @@ def warn_if_written_near_silent( ("check the step that generated it... an unintended near-zero gain upstream") is aimed at a step that could plausibly have caused the level, which a plain cut out of an already-quiet recording did not (#309). + + The trigger stays mean-only (#358): a wordless, ambience-only shot (paws, + husks scraping, water) reads a low mean with real peaks - -54 dBFS mean, + -18 to -31 dBFS peaks - and split perfectly with dialogue presence, not + with silence, costing an investigation every run. The fix is the message, + not the gate: a peak above `NEAR_SILENT_QUIET_NOT_EMPTY_DBFS` says so + plainly rather than reusing the "check the step that generated it" + wording aimed at a genuinely empty render (#261's -68.7 dBFS, S-F077's + -60 dBFS). """ if already_warned or source_already_quiet: return None @@ -287,16 +301,26 @@ def warn_if_written_near_silent( if mean is None or mean >= NEAR_SILENT_WARN_DBFS: return mean name = os.path.basename(output_path) - emit_warning( - f"{name} decodes at a mean level of {mean:+.2f} dBFS - near-silent " - f"for a deliverable meant to be heard. Check the step that " - f"generated it: an empty or malformed prompt, a source model that " - f"produced no meaningful audio for this input, or an unintended " - f"near-zero gain upstream ('normalize_audio' or 'match_levels').", - kind="audio_near_silent", - file=name, - mean_dbfs=round(mean, 2), - ) + peak = info.get("peak_dbfs") + if peak is not None and peak >= NEAR_SILENT_QUIET_NOT_EMPTY_DBFS: + message = ( + f"{name} decodes at a mean level of {mean:+.2f} dBFS but peaks " + f"at {peak:+.2f} dBFS: quiet overall, not empty. Expected for " + f"an ambience-only shot; a concern only if this was meant to " + f"carry speech or music." + ) + else: + message = ( + f"{name} decodes at a mean level of {mean:+.2f} dBFS - near-silent " + f"for a deliverable meant to be heard. Check the step that " + f"generated it: an empty or malformed prompt, a source model that " + f"produced no meaningful audio for this input, or an unintended " + f"near-zero gain upstream ('normalize_audio' or 'match_levels')." + ) + fields = {"kind": "audio_near_silent", "file": name, "mean_dbfs": round(mean, 2)} + if peak is not None: + fields["peak_dbfs"] = round(peak, 2) + emit_warning(message, **fields) return mean diff --git a/tests/test_result.py b/tests/test_result.py index e7fa9e71..dde757cf 100644 --- a/tests/test_result.py +++ b/tests/test_result.py @@ -1862,6 +1862,56 @@ def test_a_file_that_decodes_near_silent_warns(self): assert warning["mean_dbfs"] == pytest.approx(-74.8, abs=0.01) assert warning["file"] == "line.wav" + def test_a_quiet_shot_with_real_peaks_says_quiet_not_empty(self): + """#358: an ambience-only shot (paws, husks scraping, water) reads a + low mean with real peaks - s02's own -54.46 dBFS mean, -18.5 dBFS + peak - and the old wording ("check the step that generated it...") + sent every one of those to an investigation. The trigger is + unchanged (mean still below -40); only the message and the added + peak_dbfs field distinguish it from a genuinely empty render.""" + from dw.result import warn_if_written_near_silent + + with patch( + "dw.media_info.probe_media", + return_value={"mean_dbfs": -54.46, "peak_dbfs": -18.5}, + ): + (warning,) = self.warnings_from( + lambda: warn_if_written_near_silent("/runs/final/shot.wav") + ) + + assert warning["kind"] == "audio_near_silent" + assert warning["mean_dbfs"] == pytest.approx(-54.46, abs=0.01) + assert warning["peak_dbfs"] == pytest.approx(-18.5, abs=0.01) + assert "quiet overall, not empty" in warning["message"] + assert "check the step that generated it" not in warning["message"].lower() + + def test_a_genuinely_empty_render_keeps_the_old_wording(self): + """#261's -68.7 dBFS Bark clip and S-F077's -60 dBFS normalize both + have low peaks too - those still get the "check the step" message, + not the ambience one.""" + from dw.result import warn_if_written_near_silent + + with patch( + "dw.media_info.probe_media", + return_value={"mean_dbfs": -68.7, "peak_dbfs": -55.0}, + ): + (warning,) = self.warnings_from( + lambda: warn_if_written_near_silent("/runs/final/empty.wav") + ) + + assert "check the step that generated it" in warning["message"].lower() + assert "quiet overall, not empty" not in warning["message"] + assert warning["peak_dbfs"] == pytest.approx(-55.0, abs=0.01) + + def test_no_peak_available_keeps_the_old_wording(self): + """A probe that reports only mean_dbfs (no peak) can't distinguish + the two cases, so it falls back to the original message and carries + no peak_dbfs field.""" + (warning,) = self.measured_at(-74.8) + + assert "check the step that generated it" in warning["message"].lower() + assert "peak_dbfs" not in warning + def test_a_file_at_the_threshold_is_quiet(self): assert self.measured_at(-40.0) == [] From 2cae9d0ec676f6eb4cb7e06247f31a2c35f28e0d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 09:37:17 -0500 Subject: [PATCH 046/181] fix(mcp): #361 - report BS.1770 LUFS/true-peak, add target_lufs to normalize_audio probe_media now reports integrated_lufs and true_peak_dbfs (dw/loudness.py, via pyloudnorm) alongside the existing peak_dbfs/mean_dbfs, for any audio or video-with-audio - None for a track too short for the 400ms gating block or one that's silent throughout, never a raised exception. normalize_audio gains an optional target_lufs: when given, gains the track to that integrated loudness with peak_dbfs still holding as a ceiling, warning (kind=target_lufs_capped) when the ceiling caps the gain short of the target, or (kind=target_lufs_unmeasurable) when the track can't be LUFS-measured. Domain declared in dw/task_domains.py (NON_POSITIVE, already present) and checked both statically and at run time. Default behavior (target_lufs=None) is unchanged. series-episodes and minimax-music3 skills now note that peak level doesn't say how loud a track reads and point at integrated_lufs/target_lufs for matching loudness across episodes or a mix. Co-Authored-By: Claude Sonnet 5 --- dw/loudness.py | 82 +++++++++++++++++++++ dw/media_info.py | 57 +++++++++++---- dw/task_domains.py | 16 ++++- dw/tasks/audio_utils.py | 56 +++++++++++++-- plugins/dw/skills/minimax-music3/SKILL.md | 7 +- plugins/dw/skills/series-episodes/SKILL.md | 12 +++- pyproject.toml | 2 + tests/test_audio_utils.py | 83 ++++++++++++++++++++++ tests/test_media_info.py | 54 ++++++++++++++ tests/test_task_domains.py | 12 +++- 10 files changed, 356 insertions(+), 25 deletions(-) create mode 100644 dw/loudness.py diff --git a/dw/loudness.py b/dw/loudness.py new file mode 100644 index 00000000..405404c5 --- /dev/null +++ b/dw/loudness.py @@ -0,0 +1,82 @@ +"""BS.1770 loudness measurement, shared by the media probe and the audio +tasks so both read a level the same way. + +Peak and loudness are different quantities - a sparse voice-over and a dense +score can share a peak and still sit tens of dB apart in how loud they sound, +because a peak is one sample and loudness is measured over the whole track +(#361). `integrated_lufs` is the BS.1770 integrated measure pyloudnorm +implements; `true_peak_dbfs` is the inter-sample peak BS.1770 defines +alongside it - a sample-peak reading can miss a peak that only appears +between samples, which is what an encoder's reconstruction filter can ring +up past 0 dBFS even when every decoded sample was under it. +""" + +import logging +import math + +import numpy +import pyloudnorm +import scipy.signal + +logger = logging.getLogger("dw") + +# The floor a level is reported at rather than -inf, which JSON cannot carry +SILENCE_DBFS = -120.0 + +# BS.1770's gating block is 400 ms; pyloudnorm refuses anything shorter +MIN_LUFS_SECONDS = 0.4 + +# Minimum oversampling BS.1770 defines for a true-peak measurement +TRUE_PEAK_OVERSAMPLE = 4 + + +def _dbfs(value): + if value <= 0: + return SILENCE_DBFS + return max(SILENCE_DBFS, 20.0 * math.log10(float(value))) + + +def integrated_lufs(samples, rate): + """Integrated loudness in LUFS, or None when it cannot be measured. + + `samples` is (frames, channels) or (frames,). None for a clip shorter + than the 400 ms gating block, an all-silent clip (pyloudnorm's own + -inf, which JSON cannot carry either), or anything pyloudnorm refuses - + a measurement that failed reads as "unknown" rather than as a crashed + task or probe. + """ + if samples is None or samples.size == 0: + return None + frames = samples.shape[0] + if frames < int(round(MIN_LUFS_SECONDS * rate)): + return None + try: + meter = pyloudnorm.Meter(rate) + value = float(meter.integrated_loudness(samples)) + except Exception as e: + logger.debug(f"integrated_lufs: could not measure: {e}") + return None + if not math.isfinite(value): + return None + return value + + +def true_peak_dbfs(samples, oversample=TRUE_PEAK_OVERSAMPLE): + """The inter-sample (true) peak of a track, in dBFS. + + `samples` is (frames, channels) or (frames,). Oversamples with a + polyphase FIR (BS.1770's own reconstruction) and reads the peak of the + interpolated signal, which is what a lossy encoder's own reconstruction + filter can ring up past a sample-peak reading that stayed under 0 dBFS. + SILENCE_DBFS for an empty track, never None - unlike integrated + loudness, a true peak is defined for any track, including a silent one. + """ + if samples is None or samples.size == 0: + return SILENCE_DBFS + try: + oversampled = scipy.signal.resample_poly(samples, oversample, 1, axis=0) + except Exception as e: + logger.debug(f"true_peak_dbfs: could not oversample: {e}") + oversampled = samples + peak = float(numpy.abs(oversampled).max(initial=0.0)) + return _dbfs(peak) diff --git a/dw/media_info.py b/dw/media_info.py index 74576de2..ded978b7 100644 --- a/dw/media_info.py +++ b/dw/media_info.py @@ -11,25 +11,29 @@ import av import numpy -logger = logging.getLogger("dw") - -# The floor a level is reported at rather than -inf, which JSON cannot carry -SILENCE_DBFS = -120.0 - +from .loudness import _dbfs, integrated_lufs, true_peak_dbfs -def _dbfs(value): - if value <= 0: - return SILENCE_DBFS - return max(SILENCE_DBFS, 20.0 * math.log10(float(value))) +logger = logging.getLogger("dw") def probe_media(path, envelope=False): """Duration, format and level of an audio or video file, or None. Video answers fps, frame_count, width and height, plus the soundtrack's - sample_rate, channels, peak_dbfs and mean_dbfs when it carries one; - audio answers the soundtrack fields. Levels come from decoding the - whole track, which is cheap next to generating it. + sample_rate, channels, peak_dbfs, mean_dbfs, integrated_lufs and + true_peak_dbfs when it carries one; audio answers the soundtrack fields. + Levels come from decoding the whole track, which is cheap next to + generating it. + + peak_dbfs and mean_dbfs are a single sample's level; integrated_lufs is + the BS.1770 loudness of the whole track (#361) - a sparse voice-over and + a dense score can share a peak and still sit tens of dB apart in how + loud they sound. integrated_lufs is None for a track shorter than the + 400 ms gating block or one that is silent throughout - "unmeasurable", + not zero. true_peak_dbfs is the inter-sample (oversampled) peak BS.1770 + also defines, which can read higher than peak_dbfs when an encoder's + reconstruction filter rings a decoded peak up past what any single + sample showed. When a frame count still needs counting and/or a soundtrack still needs its levels measured, both are gathered from a single decode pass over @@ -107,6 +111,19 @@ def probe_media(path, envelope=False): if bins is not None and info.get("duration_seconds") is not None else None ) + # The whole soundtrack, accumulated the same way regardless of + # envelope - integrated loudness and true peak are measured over + # the full track, not per frame, so they need it assembled + # rather than the running peak/sum above. Trimmed past the + # file's reported duration for the same reason as the envelope + # bins (#277): priming/padding is not real content to measure. + lufs_chunks = [] if audio is not None else None + lufs_seen = 0 + max_lufs_samples = ( + int(round(info["duration_seconds"] * audio.rate)) + if lufs_chunks is not None and info.get("duration_seconds") is not None + else None + ) streams = [ s for s in ((video if need_frame_count else None), audio) @@ -140,6 +157,17 @@ def probe_media(path, envelope=False): int(audio.channels), max_envelope_samples, ) + if lufs_chunks is not None: + frame = _as_frame_samples(samples, int(audio.channels)) + full_length = frame.shape[0] + length = ( + full_length + if max_lufs_samples is None + else max(0, min(full_length, max_lufs_samples - lufs_seen)) + ) + if length > 0: + lufs_chunks.append(frame[:length]) + lufs_seen += full_length if bins is not None and max_envelope_samples is None: _merge_trailing_fragment(bins, audio.rate) except Exception as e: @@ -156,6 +184,11 @@ def probe_media(path, envelope=False): rms = math.sqrt(total / count) if count else 0.0 info["peak_dbfs"] = _dbfs(peak) info["mean_dbfs"] = _dbfs(rms) + full = ( + numpy.concatenate(lufs_chunks, axis=0) if lufs_chunks else None + ) + info["integrated_lufs"] = integrated_lufs(full, audio.rate) + info["true_peak_dbfs"] = true_peak_dbfs(full) if bins is not None: info["envelope"] = _as_envelope(bins) return info diff --git a/dw/task_domains.py b/dw/task_domains.py index ccf90cf7..3d79cb9e 100644 --- a/dw/task_domains.py +++ b/dw/task_domains.py @@ -38,10 +38,15 @@ # second, since the head of a track is a legitimate place to begin POSITIVE = "positive" NON_NEGATIVE = "non_negative" +# A level that cannot exceed full scale - peak_dbfs's own long-standing rule +# (0 is full scale, positive is not a level any of these commands can reach), +# shared here with normalize_audio's target_lufs +NON_POSITIVE = "non_positive" _DOMAIN_TEXT = { POSITIVE: "above zero", NON_NEGATIVE: "zero or above", + NON_POSITIVE: "at or below full scale (0)", } # command -> argument -> domain. Every entry here is pinned to a real command @@ -80,7 +85,10 @@ "fade_out_ms": NON_NEGATIVE, "sample_rate": POSITIVE, }, - "normalize_audio": {"sample_rate": POSITIVE}, + "normalize_audio": { + "sample_rate": POSITIVE, + "target_lufs": NON_POSITIVE, + }, "crossfade_audio": { "crossfade_ms": NON_NEGATIVE, "sample_rate": POSITIVE, @@ -127,7 +135,11 @@ def in_domain(value, domain): number = as_number(value) if number is None: return True - return number > 0 if domain == POSITIVE else number >= 0 + if domain == POSITIVE: + return number > 0 + if domain == NON_POSITIVE: + return number <= 0 + return number >= 0 def as_number(value): diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index 52a57ed6..9e436952 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -15,6 +15,7 @@ import torch from ..events import emit_log, emit_warning +from ..loudness import integrated_lufs from ..task_domains import as_number, check_arguments from ..security import ( validate_file_extension, @@ -1299,7 +1300,7 @@ def fade_audio(audio, fade_in_ms=0, fade_out_ms=0, sample_rate=None): return _as_track(faded, sample_rate, "fade_audio") -def normalize_audio(audio, peak_dbfs=-1.0, sample_rate=None): +def normalize_audio(audio, peak_dbfs=-1.0, target_lufs=None, sample_rate=None): """Task command: scale a track so its loudest sample sits at a level. Generated music comes out at whatever level the model happened to land @@ -1312,13 +1313,25 @@ def normalize_audio(audio, peak_dbfs=-1.0, sample_rate=None): soundtrack is taken), a video generated with a soundtrack, or a waveform (which needs sample_rate alongside it) peak_dbfs: The level the loudest sample is moved to, in dB below full - scale. 0 is full scale; -1 leaves a little headroom + scale. 0 is full scale; -1 leaves a little headroom. Still + applies as a ceiling when target_lufs is also given + target_lufs: Integrated loudness (BS.1770) to gain the track to, in + LUFS. Peak alone says nothing about how loud a track sounds - a + sparse voice-over and a dense score can share a peak and still + sit tens of dB apart to the ear (#361). When given, the gain + targets this loudness first; peak_dbfs still holds as a ceiling, + and if reaching target_lufs would cross it the gain stops at the + ceiling and a warning names the shortfall in LU. None (the + default) leaves behavior exactly as peak-only sample_rate: Sample rate of a waveform passed directly Returns: An AudioTrack holding the scaled waveform and its rate; a silent track is returned unchanged """ + check_arguments( + "normalize_audio", sample_rate=sample_rate, target_lufs=target_lufs + ) waveform, sample_rate = _waveform_and_rate(audio, sample_rate, "normalize_audio") if peak_dbfs > 0: raise ValueError("normalize_audio 'peak_dbfs' cannot be above full scale (0)") @@ -1326,10 +1339,41 @@ def normalize_audio(audio, peak_dbfs=-1.0, sample_rate=None): if peak == 0.0: logger.warning("normalize_audio: the track is silent - left unchanged") return _as_track(waveform, sample_rate, "normalize_audio") - gain = 10 ** (peak_dbfs / 20) / peak - logger.debug( - f"normalize_audio: peak {peak:.3f}, gain {20 * numpy.log10(gain):+.1f} dB" - ) + + if target_lufs is None: + gain_db = peak_dbfs - 20 * numpy.log10(peak) + else: + peak_db = 20 * numpy.log10(peak) + ceiling_gain_db = peak_dbfs - peak_db + current_lufs = integrated_lufs(waveform.T, sample_rate) + if current_lufs is None: + emit_warning( + f"normalize_audio: target_lufs={target_lufs} was given, but the " + "track's loudness could not be measured (shorter than the 400 ms " + "gating block, or silent throughout) - falling back to peak_dbfs " + "alone.", + kind="target_lufs_unmeasurable", + command="normalize_audio", + target_lufs=target_lufs, + ) + gain_db = ceiling_gain_db + else: + target_gain_db = target_lufs - current_lufs + gain_db = min(target_gain_db, ceiling_gain_db) + if gain_db < target_gain_db: + emit_warning( + f"normalize_audio: target_lufs={target_lufs} would need " + f"{target_gain_db:+.1f} dB of gain, but peak_dbfs={peak_dbfs} " + f"caps it at {gain_db:+.1f} dB - " + f"{target_gain_db - gain_db:.1f} LU short of the target.", + kind="target_lufs_capped", + command="normalize_audio", + target_lufs=target_lufs, + peak_dbfs=peak_dbfs, + shortfall_lu=target_gain_db - gain_db, + ) + gain = 10 ** (gain_db / 20) + logger.debug(f"normalize_audio: peak {peak:.3f}, gain {gain_db:+.1f} dB") return _as_track( (waveform * gain).astype(numpy.float32), sample_rate, "normalize_audio" ) diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 3cb6f0c8..fcdaa357 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -154,7 +154,12 @@ Control" section. 4. Judge it yourself. `get_gallery_metadata` for duration and sample rate: `media.duration_seconds` within 0.2 s of `audio_duration` means the ceiling cut the track (raise it and rerun); well short of it means the - song finished on its own. Then listen with `get_output_audio` (a long + song finished on its own. Its `peak_dbfs` is a single sample and does not + say how loud the song reads end to end - `integrated_lufs` (BS.1770, + whole-track) is the field for that, and what `normalize_audio`'s optional + `target_lufs` targets when a score or a music-video mix needs to match + another track by ear rather than by peak alone. Then listen with + `get_output_audio` (a long track in `start`/`duration` excerpts) for the family's failure modes: a song that went instrumental (name the vocals in the caption), an ending cut mid-note (raise the ceiling, then trim), a structure that ignored the diff --git a/plugins/dw/skills/series-episodes/SKILL.md b/plugins/dw/skills/series-episodes/SKILL.md index b51ab7bb..2f5732c8 100644 --- a/plugins/dw/skills/series-episodes/SKILL.md +++ b/plugins/dw/skills/series-episodes/SKILL.md @@ -93,7 +93,12 @@ composes into, not something to re-derive: applied before the score is passed in. - **normalize**: the mixed world sound and score are normalized together (`assemble-and-score`'s `balanced` step, -3 dBFS) so one episode is not - louder than the next. + louder than the next. A shared peak ceiling does not mean a shared + loudness - a sparse, dialogue-only episode and a dense, score-heavy one + can both sit at -3 dBFS peak and still read as very different volumes; + `normalize_audio`'s optional `target_lufs` gains to a measured loudness + first, with `peak_dbfs` still holding as a ceiling, when episodes need to + match by ear rather than by sample. - **pair**: the normalized track is muxed onto the cut - the episode's deliverable. @@ -120,7 +125,10 @@ between shots, a portrait imposing its framing, a voice without affect): watch two episodes back to back and check the cast reads as the same people - the failure mode step 0 exists to prevent. `get_gallery_metadata` on each episode's final file for duration and loudness, so a level -mismatch between episodes shows up before a viewer notices it. To confirm a +mismatch between episodes shows up before a viewer notices it - its +`peak_dbfs` is a single sample and does not say how loud the episode reads +as a whole; `integrated_lufs` (BS.1770, whole-track) is the field that +answers that, and is what to compare across episodes. To confirm a line actually rendered rather than judging it by ear, `get_output_audio` returns sound, not text: `run_workflow(name="templates/transcribe-audio", arguments={"input_audio": "output:"}, wait_seconds=55)` then diff --git a/pyproject.toml b/pyproject.toml index 90000e75..6a021ac3 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -85,6 +85,8 @@ dependencies = [ "av>=18.1.0", "soundfile>=0.14.0", "matplotlib>=3.11.1", + # BS.1770 integrated loudness / true-peak metering (dw/loudness.py, #361) + "pyloudnorm>=0.2.0", # speaker_embedding preprocessing arg for generate_speech: reduces a # reference clip to the x-vector a SpeechT5 model conditions its voice # on (spkrec-xvect-voxceleb). Imported lazily in speech_generation.py, diff --git a/tests/test_audio_utils.py b/tests/test_audio_utils.py index 3e80a889..4837dde6 100644 --- a/tests/test_audio_utils.py +++ b/tests/test_audio_utils.py @@ -619,6 +619,89 @@ def test_a_target_above_full_scale_is_refused(self): with pytest.raises(ValueError, match="full scale"): normalize_audio(numpy.ones((1, 10)), peak_dbfs=1.0, sample_rate=100) + def _tone(self, rate=48000, seconds=2.0, amplitude=0.1, density=1.0): + """A calibration tone, optionally sparse - `density` zeroes out all + but that fraction of the track so a "sparse" and a "dense" signal + can share a peak and still sit apart in integrated loudness.""" + n = int(rate * seconds) + t = numpy.arange(n) / rate + tone = amplitude * numpy.sin(2 * numpy.pi * 1000 * t) + if density < 1.0: + mask = numpy.zeros(n, dtype=bool) + mask[: int(n * density)] = True + tone = tone * mask + return tone[numpy.newaxis, :].astype(numpy.float32), rate + + def test_target_lufs_hits_its_target_on_a_sparse_signal(self): + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone(density=0.1) + + scaled = samples( + normalize_audio(track, peak_dbfs=0.0, target_lufs=-16.0, sample_rate=rate) + ) + + from dw.loudness import integrated_lufs + + assert integrated_lufs(scaled, rate) == pytest.approx(-16.0, abs=0.5) + + def test_target_lufs_hits_its_target_on_a_dense_signal(self): + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone(density=1.0) + + scaled = samples( + normalize_audio(track, peak_dbfs=0.0, target_lufs=-16.0, sample_rate=rate) + ) + + from dw.loudness import integrated_lufs + + assert integrated_lufs(scaled, rate) == pytest.approx(-16.0, abs=0.5) + + def test_the_peak_ceiling_holds_and_warns_when_target_lufs_would_exceed_it( + self, + ): + from dw.events import RunContext, activate_context, deactivate_context + from dw.tasks.audio_utils import normalize_audio + + # A loud, dense tone: reaching -1 LUFS would need to push the gain + # up past the -1 dBFS ceiling, so the ceiling has to win. + track, rate = self._tone(amplitude=0.5, density=1.0) + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + scaled = samples( + normalize_audio( + track, peak_dbfs=-1.0, target_lufs=-1.0, sample_rate=rate + ) + ) + finally: + deactivate_context(token) + + peak_dbfs = 20 * numpy.log10(numpy.abs(scaled).max()) + assert peak_dbfs == pytest.approx(-1.0, abs=0.01) + + warnings = [e for e in events if e.get("kind") == "target_lufs_capped"] + assert len(warnings) == 1 + assert warnings[0]["shortfall_lu"] > 0 + + def test_default_behavior_is_unchanged_without_target_lufs(self): + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone() + + with_default = samples( + normalize_audio(track.copy(), peak_dbfs=-3.0, sample_rate=rate) + ) + explicit_none = samples( + normalize_audio( + track.copy(), peak_dbfs=-3.0, target_lufs=None, sample_rate=rate + ) + ) + + assert numpy.array_equal(with_default, explicit_none) + class TestAudioTasksTakeAnAudioVideo: """Every audio task accepts the video an earlier step generated with its diff --git a/tests/test_media_info.py b/tests/test_media_info.py index 1102d29f..cea2e583 100644 --- a/tests/test_media_info.py +++ b/tests/test_media_info.py @@ -135,6 +135,60 @@ def test_silence_is_clamped_not_minus_infinity(tmp_path): assert not math.isinf(info["mean_dbfs"]) +class TestLoudness: + """integrated_lufs and true_peak_dbfs, #361 - a level measured over the + whole track rather than a single sample.""" + + def test_a_known_tone_reads_within_half_a_lu_of_its_reference(self, tmp_path): + # The known reference is pyloudnorm's own measurement of the exact + # waveform write_wav encodes - probe_media's answer, reached through + # a full decode of the file it wrote, must agree with a direct + # measurement of the source samples to within codec/quantization + # noise (the wav here is lossless, so this is tight). + import pyloudnorm + + seconds, rate, amplitude = 2.0, 8000, 0.5 + write_wav( + tmp_path / "tone.wav", seconds=seconds, sample_rate=rate, amplitude=amplitude + ) + t = numpy.arange(int(seconds * rate)) / rate + tone = numpy.sin(2 * numpy.pi * 220 * t) * amplitude + reference = pyloudnorm.Meter(rate).integrated_loudness( + numpy.stack([tone, tone], axis=1) + ) + + info = probe_media(str(tmp_path / "tone.wav")) + + assert info["integrated_lufs"] == pytest.approx(reference, abs=0.5) + + def test_silence_reports_lufs_as_none_not_minus_infinity(self, tmp_path): + write_wav(tmp_path / "quiet.wav", amplitude=0.0) + + info = probe_media(str(tmp_path / "quiet.wav")) + + assert info["integrated_lufs"] is None + assert info["true_peak_dbfs"] == -120.0 + + def test_a_clip_shorter_than_the_gating_block_reports_lufs_as_none( + self, tmp_path + ): + write_wav(tmp_path / "short.wav", seconds=0.1, sample_rate=8000, amplitude=0.5) + + info = probe_media(str(tmp_path / "short.wav")) + + assert info["integrated_lufs"] is None + # true peak is still a single-sample-independent measurement, and is + # defined for any nonempty track regardless of length + assert info["true_peak_dbfs"] < 0.0 + + def test_true_peak_is_reported_alongside_sample_peak(self, tmp_path): + write_wav(tmp_path / "score.wav", seconds=2.0, sample_rate=8000, amplitude=0.5) + + info = probe_media(str(tmp_path / "score.wav")) + + assert info["true_peak_dbfs"] == pytest.approx(-6.0, abs=0.5) + + def test_a_file_that_is_not_media_answers_none(tmp_path): (tmp_path / "notes.txt").write_text("not media") diff --git a/tests/test_task_domains.py b/tests/test_task_domains.py index a94c0dc4..15e00170 100644 --- a/tests/test_task_domains.py +++ b/tests/test_task_domains.py @@ -16,6 +16,7 @@ from dw.task_domains import ( TASK_ARGUMENT_DOMAINS, NON_NEGATIVE, + NON_POSITIVE, POSITIVE, as_number, task_argument_errors, @@ -59,9 +60,9 @@ def test_every_argument_is_a_parameter_of_its_command(self): parameters = {p["name"] for p in describe_task(command)["parameters"]} assert set(domains) <= parameters, command - def test_every_domain_is_one_of_the_two(self): + def test_every_domain_is_one_of_the_three(self): for domains in TASK_ARGUMENT_DOMAINS.values(): - assert set(domains.values()) <= {POSITIVE, NON_NEGATIVE} + assert set(domains.values()) <= {POSITIVE, NON_NEGATIVE, NON_POSITIVE} class TestAsNumber: @@ -100,6 +101,13 @@ def test_a_zero_target_rate_is_refused(self): assert len(errors) == 1 assert errors[0]["path"] == "steps[0].task.arguments.target_sample_rate" + def test_a_positive_target_lufs_is_refused(self): + errors = errors_for( + "normalize_audio", {"audio": "asset:bed.wav", "target_lufs": 3.0} + ) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].task.arguments.target_lufs" + def test_a_zero_offset_is_fine(self): assert ( errors_for( From b6db9a4c1c6d08f55b30e1afbfca610a3220fe94 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 09:39:53 -0500 Subject: [PATCH 047/181] style(mcp): #361 - ruff format LUFS/target_lufs changes Co-Authored-By: Claude Sonnet 5 --- dw/media_info.py | 8 ++++---- dw/tasks/audio_utils.py | 4 +--- tests/test_media_info.py | 9 +++++---- 3 files changed, 10 insertions(+), 11 deletions(-) diff --git a/dw/media_info.py b/dw/media_info.py index ded978b7..3b7c5260 100644 --- a/dw/media_info.py +++ b/dw/media_info.py @@ -163,7 +163,9 @@ def probe_media(path, envelope=False): length = ( full_length if max_lufs_samples is None - else max(0, min(full_length, max_lufs_samples - lufs_seen)) + else max( + 0, min(full_length, max_lufs_samples - lufs_seen) + ) ) if length > 0: lufs_chunks.append(frame[:length]) @@ -184,9 +186,7 @@ def probe_media(path, envelope=False): rms = math.sqrt(total / count) if count else 0.0 info["peak_dbfs"] = _dbfs(peak) info["mean_dbfs"] = _dbfs(rms) - full = ( - numpy.concatenate(lufs_chunks, axis=0) if lufs_chunks else None - ) + full = numpy.concatenate(lufs_chunks, axis=0) if lufs_chunks else None info["integrated_lufs"] = integrated_lufs(full, audio.rate) info["true_peak_dbfs"] = true_peak_dbfs(full) if bins is not None: diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index 9e436952..5584bc0d 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -1329,9 +1329,7 @@ def normalize_audio(audio, peak_dbfs=-1.0, target_lufs=None, sample_rate=None): An AudioTrack holding the scaled waveform and its rate; a silent track is returned unchanged """ - check_arguments( - "normalize_audio", sample_rate=sample_rate, target_lufs=target_lufs - ) + check_arguments("normalize_audio", sample_rate=sample_rate, target_lufs=target_lufs) waveform, sample_rate = _waveform_and_rate(audio, sample_rate, "normalize_audio") if peak_dbfs > 0: raise ValueError("normalize_audio 'peak_dbfs' cannot be above full scale (0)") diff --git a/tests/test_media_info.py b/tests/test_media_info.py index cea2e583..3a8d7783 100644 --- a/tests/test_media_info.py +++ b/tests/test_media_info.py @@ -149,7 +149,10 @@ def test_a_known_tone_reads_within_half_a_lu_of_its_reference(self, tmp_path): seconds, rate, amplitude = 2.0, 8000, 0.5 write_wav( - tmp_path / "tone.wav", seconds=seconds, sample_rate=rate, amplitude=amplitude + tmp_path / "tone.wav", + seconds=seconds, + sample_rate=rate, + amplitude=amplitude, ) t = numpy.arange(int(seconds * rate)) / rate tone = numpy.sin(2 * numpy.pi * 220 * t) * amplitude @@ -169,9 +172,7 @@ def test_silence_reports_lufs_as_none_not_minus_infinity(self, tmp_path): assert info["integrated_lufs"] is None assert info["true_peak_dbfs"] == -120.0 - def test_a_clip_shorter_than_the_gating_block_reports_lufs_as_none( - self, tmp_path - ): + def test_a_clip_shorter_than_the_gating_block_reports_lufs_as_none(self, tmp_path): write_wav(tmp_path / "short.wav", seconds=0.1, sample_rate=8000, amplitude=0.5) info = probe_media(str(tmp_path / "short.wav")) From fbb3f5817d95cf523f0e0240f641f0d9d51d7224 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 10:01:31 -0500 Subject: [PATCH 048/181] fix(mcp): #363 - let result.fps take a variable:/item: reference Widens result.fps in the schema (integer -> integer|number|string) so a "variable:"/"item:" reference passes the raw-document schema gate before substitution runs, the same way result.subfolder already does. dw/result_fps.py checks the resolved value at validate time: it must be a positive whole number - encode_video hands fps straight to PyAV as rate=int(fps), so a fractional rate like 23.976 would be silently truncated rather than honored, and is refused rather than accepted. A float with no fractional part (24.0, as the ltx2 templates already declare) still passes. Repoints every template whose result.fps literal duplicated a frame_rate/fps variable at that variable: the 12 ltx2 templates and assemble-and-score, dissolve-between-shots and minimax/music-video. Templates with a hard-coded fps and no frame_rate/fps variable to duplicate (best-of-n-to-video and the MiniMax H3 catalog) are unchanged. Co-Authored-By: Claude Sonnet 5 --- dw/result_fps.py | 82 ++++++++ dw/workflow.py | 2 + dw/workflow_schema.json | 4 +- tests/test_result_fps.py | 189 ++++++++++++++++++ workflows/templates/assemble-and-score.json | 2 +- .../templates/dissolve-between-shots.json | 2 +- .../templates/ltx2/chained-segments.json | 2 +- .../templates/ltx2/diffusion-decode.json | 2 +- workflows/templates/ltx2/enhance-prompt.json | 2 +- workflows/templates/ltx2/extend-clip.json | 6 +- .../templates/ltx2/generative-upscale.json | 4 +- workflows/templates/ltx2/image-to-video.json | 2 +- workflows/templates/ltx2/keyframes.json | 2 +- workflows/templates/ltx2/reference-sheet.json | 2 +- workflows/templates/ltx2/restore-deblur.json | 2 +- .../templates/ltx2/restore-decompression.json | 2 +- workflows/templates/ltx2/text-to-video.json | 2 +- workflows/templates/ltx2/two-stage.json | 6 +- workflows/templates/minimax/music-video.json | 4 +- 19 files changed, 296 insertions(+), 23 deletions(-) create mode 100644 dw/result_fps.py create mode 100644 tests/test_result_fps.py diff --git a/dw/result_fps.py b/dw/result_fps.py new file mode 100644 index 00000000..677db7dd --- /dev/null +++ b/dw/result_fps.py @@ -0,0 +1,82 @@ +"""A step's result 'fps': the frame rate a video is written at. + +Widened the same way `subfolder` was (dw/subfolders.py): the schema lets the +raw document hold a `variable:` or `item:` reference so a workflow can keep +one frame-rate variable rather than a literal duplicated between the step +that generates and the step that writes (#363). This module owns the +resolved-value check: once substitution has run, `result.fps` has to be a +real, positive frame rate, and it must be a whole number - `encode_video` +(dw/result.py, the writer used whenever a step's result carries audio) hands +the value straight to PyAV as `rate=int(fps)`, so a fractional rate like +23.976 would be silently floored to 23 rather than written as asked. A +frame_rate variable declared as a float (24.0) still passes, since it carries +no fractional part; only a genuine fraction is refused. + +Containment doesn't apply here - there's no path to join - so unlike +subfolder_errors this only ever checks value shape. +""" + +from .for_each import MEMBER_SEPARATOR, render_path + +FPS_KEY = "fps" + +# Reference prefixes substitution resolves before this pass runs. One still +# spelled out here is one nothing resolved, and that is the undeclared- +# variable pass's complaint rather than a shape error +_UNRESOLVED_PREFIXES = ("variable:", "item:") + + +def fps_errors(workflow_definition, source_indices=None): + """Every result 'fps' that cannot be written, as [{path, message}]. + + The definition handed here has already been substituted and expanded, + so every value in it is literal; a 'variable:' or 'item:' still spelled + out is left alone. `source_indices`, when given, is the source step + index of each step - a 'for_each' group turns one written step into + several, and the path an error carries has to be one the author can + find in the file they wrote; the member is named in the message. + """ + steps = workflow_definition.get("steps") + if not isinstance(steps, list): + return [] + + errors = [] + for index, step in enumerate(steps): + if not isinstance(step, dict): + continue + result = step.get("result") + if not isinstance(result, dict) or FPS_KEY not in result: + continue + value = result[FPS_KEY] + if isinstance(value, str) and value.startswith(_UNRESOLVED_PREFIXES): + continue + + source = ( + source_indices[index] + if source_indices is not None and index < len(source_indices) + else index + ) + name = step.get("name") + where = ( + f" in member '{name}'" + if isinstance(name, str) and MEMBER_SEPARATOR in name + else "" + ) + message = None + if isinstance(value, bool) or not isinstance(value, (int, float)): + message = f"Invalid fps: {value!r} - fps is a number of frames per second" + elif value <= 0: + message = f"Invalid fps: {value!r} - fps must be greater than zero" + elif float(value) != int(value): + message = ( + f"Invalid fps: {value!r} - fps must be a whole number; the video " + f"writer truncates a fractional rate rather than honoring it" + ) + if message is not None: + errors.append( + { + "path": render_path(("steps", source, "result", FPS_KEY)), + "message": f"{message}{where}", + } + ) + return errors diff --git a/dw/workflow.py b/dw/workflow.py index ddf14099..38c8df2a 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -41,6 +41,7 @@ constraint_reference_errors, resolve_constraint_references, ) +from .result_fps import fps_errors from .subfolders import step_subfolder, subfolder_errors from .reference_names import reference_name_errors from .video_extensions import video_extension_errors @@ -656,6 +657,7 @@ def validation_errors(self, arguments=None, composing=None): return ( previous_result_reference_errors(expanded, source_indices) + subfolder_errors(expanded, source_indices) + + fps_errors(expanded, source_indices) # A reference name no workspace could ever resolve - the '@' a # for_each member's own file carries, rejected after the queue # by a message that named a valid form and not the objection diff --git a/dw/workflow_schema.json b/dw/workflow_schema.json index a6a83397..2a388baf 100644 --- a/dw/workflow_schema.json +++ b/dw/workflow_schema.json @@ -1238,8 +1238,8 @@ "type": "string" }, "fps": { - "description": "Frames per second - only used when output is video. Defaults to the rate the frames themselves carry (a chain's 'fps', a join's own rate, the rate of the file a task read) and to 8 only when nothing knows better. Set it to write at a rate other than the source's - a deliberate slow motion - which the run then warns about.", - "type": "integer", + "description": "Frames per second - only used when output is video. Defaults to the rate the frames themselves carry (a chain's 'fps', a join's own rate, the rate of the file a task read) and to 8 only when nothing knows better. Set it to write at a rate other than the source's - a deliberate slow motion - which the run then warns about. May be a 'variable:' or, inside a for_each step, an 'item:' reference, resolved the same way 'subfolder' is - the resolved value must be a positive whole number; a fractional rate is refused, since the video writer would silently truncate rather than honor it.", + "type": ["integer", "number", "string"], "default": 8 }, "sample_rate": { diff --git a/tests/test_result_fps.py b/tests/test_result_fps.py new file mode 100644 index 00000000..86bd93a1 --- /dev/null +++ b/tests/test_result_fps.py @@ -0,0 +1,189 @@ +import json +import os + +from dw.result_fps import fps_errors + +REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +CHAINED_SEGMENTS = os.path.join( + REPO_ROOT, "workflows", "templates", "ltx2", "chained-segments.json" +) + + +def _step(name, result=None, **extra): + step = {"name": name, "task": {"command": "gather_images", "arguments": {}}} + if result is not None: + step["result"] = result + step.update(extra) + return step + + +class TestFpsErrors: + def test_a_clean_definition_has_none(self): + definition = { + "steps": [ + _step("a", {"content_type": "video/mp4", "fps": 24}), + _step("b", {"content_type": "video/mp4"}), + ] + } + assert fps_errors(definition) == [] + + def test_a_whole_number_float_is_fine(self): + # a rate variable declared as JSON 24.0 still carries no fraction + definition = {"steps": [_step("a", {"fps": 24.0})]} + assert fps_errors(definition) == [] + + def test_a_fractional_fps_is_reported_at_its_path(self): + definition = {"steps": [_step("a", {"fps": 23.976})]} + errors = fps_errors(definition) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].result.fps" + assert "whole number" in errors[0]["message"] + + def test_a_non_positive_fps_is_reported(self): + definition = {"steps": [_step("a", {"fps": 0})]} + errors = fps_errors(definition) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].result.fps" + assert "greater than zero" in errors[0]["message"] + + def test_a_non_numeric_fps_is_reported(self): + definition = {"steps": [_step("a", {"fps": "not-a-reference"})]} + errors = fps_errors(definition) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].result.fps" + assert "number of frames" in errors[0]["message"] + + def test_a_boolean_fps_is_reported(self): + definition = {"steps": [_step("a", {"fps": True})]} + errors = fps_errors(definition) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].result.fps" + + def test_an_expanded_member_reports_the_source_step_and_names_the_member(self): + definition = { + "steps": [ + _step("intro", {"fps": 24}), + _step("shot@a", {"fps": 24}), + _step("shot@b", {"fps": 0}), + ] + } + errors = fps_errors(definition, source_indices=[0, 1, 1]) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[1].result.fps" + assert "shot@b" in errors[0]["message"] + + def test_an_unsubstituted_reference_is_left_alone(self): + definition = {"steps": [_step("a", {"fps": "variable:frame_rate"})]} + assert fps_errors(definition) == [] + definition = {"steps": [_step("a", {"fps": "item:frame_rate"})]} + assert fps_errors(definition) == [] + + def test_no_steps_is_fine(self): + assert fps_errors({}) == [] + assert fps_errors({"steps": "nope"}) == [] + + +class TestValidationErrorsIntegration: + def _workflow(self, definition, tmp_path): + from dw.workflow import Workflow + + return Workflow(definition, str(tmp_path), "/w/workflows/Fps.json") + + def test_a_variable_driven_fps_resolves(self, tmp_path): + definition = { + "id": "fps_test", + "variables": {"frame_rate": 24}, + "steps": [ + { + "name": "a", + "task": {"command": "gather_images", "arguments": {}}, + "result": { + "content_type": "video/mp4", + "fps": "variable:frame_rate", + }, + } + ], + } + workflow = self._workflow(definition, tmp_path) + assert workflow.validation_errors() == [] + + def test_a_bad_resolved_fps_is_refused_at_validate_time(self, tmp_path): + definition = { + "id": "fps_test", + # float-typed, like the ltx2 templates' own frame_rate default - + # an integer-typed variable would itself truncate 23.976 to 23 + # before fps_errors ever saw a fractional value + "variables": {"frame_rate": 24.0}, + "steps": [ + { + "name": "a", + "task": {"command": "gather_images", "arguments": {}}, + "result": { + "content_type": "video/mp4", + "fps": "variable:frame_rate", + }, + } + ], + } + workflow = self._workflow(definition, tmp_path) + errors = workflow.validation_errors(arguments={"frame_rate": 23.976}) + assert [e["path"] for e in errors] == ["steps[0].result.fps"] + + def test_an_item_driven_fps_is_checked_per_member(self, tmp_path): + definition = { + "id": "fps_test", + "steps": [ + { + "name": "shot", + "for_each": [ + {"name": "a", "rate": 24}, + {"name": "b", "rate": -1}, + ], + "task": {"command": "gather_images", "arguments": {}}, + "result": {"content_type": "video/mp4", "fps": "item:rate"}, + } + ], + } + errors = self._workflow(definition, tmp_path).validation_errors() + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].result.fps" + assert "shot@b" in errors[0]["message"] + + def test_an_integer_literal_fps_is_unchanged(self, tmp_path): + definition = { + "id": "fps_test", + "steps": [ + { + "name": "a", + "task": {"command": "gather_images", "arguments": {}}, + "result": {"content_type": "video/mp4", "fps": 8}, + } + ], + } + workflow = self._workflow(definition, tmp_path) + assert workflow.validation_errors() == [] + + def test_chained_segments_honors_a_frame_rate_argument(self, tmp_path): + # the real #363 repro: chained-segments.json's result.fps is + # "variable:frame_rate" - a caller setting frame_rate must see that + # value reach the resolved step, not the template's own default + definition = json.load(open(CHAINED_SEGMENTS, encoding="utf-8")) + workflow = self._workflow(definition, tmp_path) + assert workflow.validation_errors(arguments={"frame_rate": 30}) == [] + + expanded = workflow.expanded_definition(arguments={"frame_rate": 30}) + step = next( + s for s in expanded["steps"] if s["name"] == "chained_image_to_video" + ) + assert step["result"]["fps"] == 30 + + def test_chained_segments_default_run_is_unchanged(self, tmp_path): + definition = json.load(open(CHAINED_SEGMENTS, encoding="utf-8")) + workflow = self._workflow(definition, tmp_path) + assert workflow.validation_errors() == [] + + expanded = workflow.expanded_definition() + step = next( + s for s in expanded["steps"] if s["name"] == "chained_image_to_video" + ) + assert step["result"]["fps"] == 24.0 diff --git a/workflows/templates/assemble-and-score.json b/workflows/templates/assemble-and-score.json index 6e8eac19..8c711285 100644 --- a/workflows/templates/assemble-and-score.json +++ b/workflows/templates/assemble-and-score.json @@ -131,7 +131,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:fps", "subfolder": "final" } } diff --git a/workflows/templates/dissolve-between-shots.json b/workflows/templates/dissolve-between-shots.json index 43d0a794..23d4cdda 100644 --- a/workflows/templates/dissolve-between-shots.json +++ b/workflows/templates/dissolve-between-shots.json @@ -128,7 +128,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:fps", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/chained-segments.json b/workflows/templates/ltx2/chained-segments.json index 0a9c44b0..6cf5f720 100644 --- a/workflows/templates/ltx2/chained-segments.json +++ b/workflows/templates/ltx2/chained-segments.json @@ -151,7 +151,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/diffusion-decode.json b/workflows/templates/ltx2/diffusion-decode.json index 54752d93..1e503739 100644 --- a/workflows/templates/ltx2/diffusion-decode.json +++ b/workflows/templates/ltx2/diffusion-decode.json @@ -165,7 +165,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/enhance-prompt.json b/workflows/templates/ltx2/enhance-prompt.json index 83d43e4b..77a15e08 100644 --- a/workflows/templates/ltx2/enhance-prompt.json +++ b/workflows/templates/ltx2/enhance-prompt.json @@ -176,7 +176,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/extend-clip.json b/workflows/templates/ltx2/extend-clip.json index d2cdbbf1..6b8b1d10 100644 --- a/workflows/templates/ltx2/extend-clip.json +++ b/workflows/templates/ltx2/extend-clip.json @@ -146,7 +146,7 @@ "result": { "content_type": "video/mp4", "save": false, - "fps": 24 + "fps": "variable:frame_rate" } }, { @@ -160,7 +160,7 @@ "result": { "content_type": "video/mp4", "save": false, - "fps": 24 + "fps": "variable:frame_rate" } }, { @@ -245,7 +245,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/generative-upscale.json b/workflows/templates/ltx2/generative-upscale.json index 6cbfdbec..3cbc13e2 100644 --- a/workflows/templates/ltx2/generative-upscale.json +++ b/workflows/templates/ltx2/generative-upscale.json @@ -144,7 +144,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "intermediate" } }, @@ -236,7 +236,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/image-to-video.json b/workflows/templates/ltx2/image-to-video.json index ac73056a..2904c847 100644 --- a/workflows/templates/ltx2/image-to-video.json +++ b/workflows/templates/ltx2/image-to-video.json @@ -141,7 +141,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/keyframes.json b/workflows/templates/ltx2/keyframes.json index 5a26bcda..977b703a 100644 --- a/workflows/templates/ltx2/keyframes.json +++ b/workflows/templates/ltx2/keyframes.json @@ -160,7 +160,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/reference-sheet.json b/workflows/templates/ltx2/reference-sheet.json index 18968de2..5a8cc1b2 100644 --- a/workflows/templates/ltx2/reference-sheet.json +++ b/workflows/templates/ltx2/reference-sheet.json @@ -185,7 +185,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/restore-deblur.json b/workflows/templates/ltx2/restore-deblur.json index 8fb779f8..3b2db4bd 100644 --- a/workflows/templates/ltx2/restore-deblur.json +++ b/workflows/templates/ltx2/restore-deblur.json @@ -166,7 +166,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/restore-decompression.json b/workflows/templates/ltx2/restore-decompression.json index 64cb023a..d294c2b9 100644 --- a/workflows/templates/ltx2/restore-decompression.json +++ b/workflows/templates/ltx2/restore-decompression.json @@ -166,7 +166,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/text-to-video.json b/workflows/templates/ltx2/text-to-video.json index b7bae28b..6e960216 100644 --- a/workflows/templates/ltx2/text-to-video.json +++ b/workflows/templates/ltx2/text-to-video.json @@ -139,7 +139,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/ltx2/two-stage.json b/workflows/templates/ltx2/two-stage.json index d6f34ebd..0c8c87fa 100644 --- a/workflows/templates/ltx2/two-stage.json +++ b/workflows/templates/ltx2/two-stage.json @@ -152,7 +152,7 @@ "result": { "content_type": "video/mp4", "save": false, - "fps": 24 + "fps": "variable:frame_rate" } }, { @@ -187,7 +187,7 @@ "result": { "content_type": "video/mp4", "save": false, - "fps": 24 + "fps": "variable:frame_rate" } }, { @@ -311,7 +311,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:frame_rate", "subfolder": "final" } } diff --git a/workflows/templates/minimax/music-video.json b/workflows/templates/minimax/music-video.json index 3d7d27a1..b219a06d 100644 --- a/workflows/templates/minimax/music-video.json +++ b/workflows/templates/minimax/music-video.json @@ -294,7 +294,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:fps", "subfolder": "intermediate" } }, @@ -337,7 +337,7 @@ }, "result": { "content_type": "video/mp4", - "fps": 24, + "fps": "variable:fps", "subfolder": "final" } } From 8b357ad571e92d27a09aeaf46673f47192736266 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 10:12:46 -0500 Subject: [PATCH 049/181] fix(server): #364 - validate_workflow with no arguments matches save_workflow JobRequest.arguments defaults to {} (it is shared with run_workflow, which needs a real dict), so the validate_workflow route could not tell "the caller sent no arguments field" from "the caller sent {}" - both arrived as request.arguments == {}, and validation_errors()'s `arguments is None` gate (added in 277265e for #364) never saw None. A document save_workflow accepted was reported invalid by validate_workflow in the same no-arguments case. validate_workflow now reads model_fields_set to recover the caller's actual intent and passes None through to validation_errors / null_variable_argument_warnings when arguments was omitted, {} when it was sent explicitly (a real run with nothing supplied, still a hard error). dw_mcp/authoring.py's payload builder had the same collapse one layer up - `if arguments:` dropped an explicit `{}` before it ever left the client - fixed to `if arguments is not None:`. Added route-level coverage in tests/test_validate_arguments.py (TestNullVariableArgument) exercising the actual POST /api/validate and PUT /api/workflows/{name} endpoints for all four cases: no arguments key, explicit {}, arguments that still leave the variable null, and arguments that supply a value. Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 16 ++++- dw_mcp/authoring.py | 6 +- tests/test_validate_arguments.py | 101 +++++++++++++++++++++++++++++++ 3 files changed, 119 insertions(+), 4 deletions(-) diff --git a/dw/server/app.py b/dw/server/app.py index c7a07899..3a3f1591 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -1731,10 +1731,20 @@ def validate_workflow( detail="Workflow could not be constructed - the server log " "has the detail", ) + # `arguments` defaults to `{}` on the model (JobRequest is shared + # with run_workflow, which needs a dict), so an omitted field and an + # explicit `{}` are otherwise indistinguishable here - and the two + # mean different things: omitted is "check the document", explicit + # is "check a run with these arguments" (#364). model_fields_set + # tells them apart without changing the field's default for every + # other caller of validate_workflow. + caller_arguments = ( + request.arguments if "arguments" in request.model_fields_set else None + ) try: # The caller's list is the one a for_each expands over, so the # pre-flight checks the step set that will actually run - errors = candidate.validation_errors(arguments=request.arguments) + errors = candidate.validation_errors(arguments=caller_arguments) except Exception: # An error here is not the schema's verdict on the workflow - # validation_errors() reports that by returning it. It is the @@ -1795,9 +1805,9 @@ def validate_workflow( + candidate.adapter_warnings(request.arguments) # A required task argument fed by variable:name where name's # default is null - a fine document, but a run left as-is would - # fail; empty once request.arguments names anything, since that + # fail; empty once the caller names any arguments, since that # condition is a hard error above instead (#364) - + candidate.null_variable_argument_warnings(request.arguments) + + candidate.null_variable_argument_warnings(caller_arguments) # An argument a sub-workflow step passes to a workflow that # declares no variable for it - dropped in silence at run time + candidate.sub_workflow_warnings(), diff --git a/dw_mcp/authoring.py b/dw_mcp/authoring.py index c7a970f0..9aed52b8 100644 --- a/dw_mcp/authoring.py +++ b/dw_mcp/authoring.py @@ -55,7 +55,11 @@ def validate_workflow( ) params = {"workspace": workspace} if workspace else None payload = {"workflow_path": stored} if inline is None else {"workflow": inline} - if arguments: + if arguments is not None: + # An explicit {} still means "check a run with no values supplied" - + # distinct from omitting arguments entirely, which means "check the + # document" (#364). A falsy-but-not-None check here would drop that + # distinction before it ever reaches the server. payload["arguments"] = arguments # The server resolves a name against its own workflow directory, so # validation sees the same base directory a run would diff --git a/tests/test_validate_arguments.py b/tests/test_validate_arguments.py index 0171ae98..87fa1954 100644 --- a/tests/test_validate_arguments.py +++ b/tests/test_validate_arguments.py @@ -30,6 +30,27 @@ def typed_workflow(): return workflow +def null_variable_workflow(): + """A required task argument fed by `variable:audio`, where `audio`'s + declared default is null - the #364 repro. A fine document (the step + does supply the argument, just not yet a value); a run left as-is + would fail.""" + return { + "id": "null_var", + "variables": {"audio": None}, + "steps": [ + { + "name": "n", + "task": { + "command": "normalize_audio", + "arguments": {"audio": "variable:audio"}, + }, + "result": {"content_type": "audio/wav"}, + } + ], + } + + def placeholder_workflow(): """A stored default that names no file in this workspace, the shape of `templates/ltx2/reference-sheet` and its siblings (#166): a bare call @@ -47,6 +68,8 @@ def server(tmp_path): json.dump(typed_workflow(), file) with open(os.path.join(root.workflows, "Placeholder.json"), "w") as file: json.dump(placeholder_workflow(), file) + with open(os.path.join(root.workflows, "NullVariable.json"), "w") as file: + json.dump(null_variable_workflow(), file) with open(os.path.join(root.assets, "iris.png"), "wb") as file: file.write(b"not really a png, but it is a file under that name") with open(os.path.join(root.prompts, "hero.json"), "w") as file: @@ -411,6 +434,84 @@ def test_submission_and_validation_give_the_same_message(self, server): assert validated["errors"][0]["message"] in refused.json()["detail"] +class TestNullVariableArgument: + """#364: a required task argument fed by `variable:name` where `name`'s + declared value is null is a fine document - `save_workflow` accepts it, + and `validate_workflow` called with no `arguments` at all must agree, + since that is the same "check the document" question. The moment the + caller names arguments of their own - even `{}` - it is a real run + being checked, and one that leaves the variable null is a hard error. + """ + + def test_no_arguments_key_at_all_is_a_warning_not_an_error(self, server): + """The literal repro: no `arguments` field in the request body.""" + with server() as client: + response = client.post( + "/api/validate", json={"workflow_path": "NullVariable"} + ).json() + + assert response["valid"] is True + assert response["errors"] == [] + assert any("audio" in warning for warning in response["warnings"]) + + def test_an_inline_document_with_no_arguments_is_a_warning_too(self, server): + with server() as client: + response = client.post( + "/api/validate", json={"workflow": null_variable_workflow()} + ).json() + + assert response["valid"] is True + assert any("audio" in warning for warning in response["warnings"]) + + def test_an_explicit_empty_arguments_dict_is_still_a_hard_error(self, server): + """`{}` names a run with no values supplied, distinct from omitting + `arguments` entirely - the variable is still null for that run.""" + with server() as client: + response = client.post( + "/api/validate", + json={"workflow_path": "NullVariable", "arguments": {}}, + ).json() + + assert response["valid"] is False + assert response["errors"][0]["variable"] == "audio" + + def test_arguments_that_still_leave_it_null_are_a_hard_error(self, server): + with server() as client: + response = client.post( + "/api/validate", + json={ + "workflow_path": "NullVariable", + "arguments": {"audio": None}, + }, + ).json() + + assert response["valid"] is False + assert response["errors"][0]["variable"] == "audio" + + def test_arguments_that_supply_a_value_pass(self, server): + with server() as client: + response = client.post( + "/api/validate", + json={ + "workflow_path": "NullVariable", + "arguments": {"audio": "asset:iris.png"}, + }, + ).json() + + assert response["valid"] is True + assert response["errors"] == [] + + def test_save_workflow_accepts_the_document_with_a_warning(self, server): + with server() as client: + response = client.put( + "/api/workflows/NullVariableSaved", + json={"workflow": null_variable_workflow()}, + ) + + assert response.status_code == 200 + assert any("audio" in warning for warning in response.json()["warnings"]) + + def test_an_entry_key_no_step_reads_is_a_warning_not_an_error(server): workflow = { "id": "cut", From d5f8d5d64b9d7c6518fd0e425f0688ee2d5777c9 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 10:26:19 -0500 Subject: [PATCH 050/181] fix(engine): #368 - bound step_cache's retained media by size, not just entry count Every shot@ for_each member of dialogue-short's `shot` step is legitimately retained by the step cache (the gather step genuinely reads all of them), but nothing ever released that retention once the run that needed it finished - a handful of members held several GB of decoded video resident in this process indefinitely, past LRU eviction only kicking in at the 128-entry cap. StepCache now also tracks an approximate byte size for each retained entry (_result_bytes, walking arrays/tensors/PIL images/dataclasses the way deep_equal walks step_data) and evicts LRU past a DEFAULT_MAX_RETAINED_BYTES budget (4 GiB) the same way it already evicts past the entry cap, never evicting the sole remaining entry. Co-Authored-By: Claude Sonnet 5 --- dw/step_cache.py | 94 +++++++++++++++++++++++++++++++----- tests/test_step_cache.py | 64 +++++++++++++++++++++++- tests/test_worker_execute.py | 1 + 3 files changed, 147 insertions(+), 12 deletions(-) diff --git a/dw/step_cache.py b/dw/step_cache.py index b8aaec53..5382fb1e 100644 --- a/dw/step_cache.py +++ b/dw/step_cache.py @@ -36,12 +36,19 @@ a sub-workflow sharing a name with its parent) hit and republish the other workflow's file paths while writing none of its own. -The cache is per-process and bounded (DEFAULT_MAX_ENTRIES, LRU): it holds -realized media, so unbounded growth would work against the OOM avoidance -release_unreferenced_results exists for. The bound is sized so a maximal -for_each run (32 entries over two groups plus fixed steps, ~70 members) -never evicts its own earlier members before it ends - a smaller cap would -turn a long list-driven run into one that thrashes its own cache. +The cache is per-process and bounded two ways, LRU either way: by entry +count (DEFAULT_MAX_ENTRIES), sized so a maximal for_each run (32 entries +over two groups plus fixed steps, ~70 members) never evicts its own earlier +members before it ends - a smaller cap would turn a long list-driven run +into one that thrashes its own cache - and by the approximate byte size of +the retained media itself (DEFAULT_MAX_RETAINED_BYTES). A for_each group +whose members are each a full decoded video (dialogue-short's `shot`, gathered +by `episode`) is retained in full for every member - correctly, since the +final step genuinely reads all of them - but nothing ever released that once +the run finished, so a handful of members held several GB resident +indefinitely (#368). The byte cap does not know which entries a future run +would most want back; it just keeps the newest-used bytes under budget the +same way the entry cap keeps the newest-used count under one, oldest first. """ import copy @@ -187,6 +194,47 @@ def deep_equal(a, b): return False +def _approx_bytes(value, seen): + """Approximate resident size of a retained result item, in bytes. + + Walks the same shapes a pipeline output actually takes - a dict/dataclass + of arrays wrapping frames and audio, nested lists of per-frame images - + rather than every Python object, so an unrecognized type (a bare string, + a plain number) costs nothing rather than raising. `seen` is shared + across one Result's whole result_list so an object two artifacts both + reference (get_artifact_list's fitted-in-place audio, say) is not + double-counted. + """ + key = id(value) + if key in seen: + return 0 + seen.add(key) + if isinstance(value, torch.Tensor): + return value.element_size() * value.nelement() + if isinstance(value, np.ndarray): + return value.nbytes + if isinstance(value, Image.Image): + bands = len(value.getbands()) or 1 + return value.width * value.height * bands + if isinstance(value, (bytes, bytearray)): + return len(value) + if isinstance(value, dict): + return sum(_approx_bytes(v, seen) for v in value.values()) + if isinstance(value, (list, tuple)): + return sum(_approx_bytes(v, seen) for v in value) + if dataclasses.is_dataclass(value) and not isinstance(value, type): + return sum( + _approx_bytes(getattr(value, f.name), seen) + for f in dataclasses.fields(value) + ) + return 0 + + +def _result_bytes(result): + seen = set() + return sum(_approx_bytes(item, seen) for item in result.result_list) + + class StepCache: """Per-process cache of the last Result produced for each (workflow_id, step_name). @@ -200,20 +248,32 @@ class StepCache: """ DEFAULT_MAX_ENTRIES = 128 + # 4 GiB: enough for several retained shot@ videos at once, small next to + # the VRAM/RAM a generation step itself needs, and never the only thing + # standing between a run and OOM - release_unreferenced_results and the + # entry cap both still apply + DEFAULT_MAX_RETAINED_BYTES = 4 * 1024**3 - def __init__(self, max_entries=None): + def __init__(self, max_entries=None, max_retained_bytes=None): # (workflow_id, step_name) -> {"step_data", "step_seed", "result", - # "output_dir", "generation", "upstream_generations", "retained"}, - # ordered least- to most-recently-used + # "output_dir", "generation", "upstream_generations", "retained", + # "size"}, ordered least- to most-recently-used self._entries = OrderedDict() self.max_entries = ( self.DEFAULT_MAX_ENTRIES if max_entries is None else max_entries ) + self.max_retained_bytes = ( + self.DEFAULT_MAX_RETAINED_BYTES + if max_retained_bytes is None + else max_retained_bytes + ) + self._retained_bytes = 0 def clear(self): # The generation counter deliberately survives: it only has to be # monotonic, and restarting it could make a stale reference match self._entries.clear() + self._retained_bytes = 0 def get( self, workflow_id, step_data, step_seed, hits_this_run, output_dir, needs_result @@ -301,6 +361,10 @@ def put(self, workflow_id, step_data, step_seed, result, output_dir, retain_resu retain_result = retain_result and getattr(result, "retainable", True) name = step_data["name"] key = (workflow_id, name) + size = _result_bytes(result) if retain_result else 0 + previous = self._entries.get(key) + if previous is not None: + self._retained_bytes -= previous["size"] self._entries[key] = { "step_data": step_data, "step_seed": step_seed, @@ -309,10 +373,18 @@ def put(self, workflow_id, step_data, step_seed, result, output_dir, retain_resu "output_dir": output_dir, "generation": next(_generations), "upstream_generations": self._upstream_generations(workflow_id, step_data), + "size": size, } + self._retained_bytes += size self._entries.move_to_end(key) - while len(self._entries) > self.max_entries: - (evicted_workflow, evicted_step), _ = self._entries.popitem(last=False) + while len(self._entries) > 1 and ( + len(self._entries) > self.max_entries + or self._retained_bytes > self.max_retained_bytes + ): + (evicted_workflow, evicted_step), evicted = self._entries.popitem( + last=False + ) + self._retained_bytes -= evicted["size"] logger.debug( "Step cache full - evicting least recently used " f"'{evicted_workflow}/{evicted_step}'" diff --git a/tests/test_step_cache.py b/tests/test_step_cache.py index aa289aee..efc7ff8c 100644 --- a/tests/test_step_cache.py +++ b/tests/test_step_cache.py @@ -1,3 +1,4 @@ +import numpy as np import pytest from dw.for_each import MAX_FOR_EACH_ENTRIES @@ -6,15 +7,22 @@ deep_equal, reference_resolves_to, normalized_downstream, + _result_bytes, ) class FakeResult: - def __init__(self, label, saved_files=None): + def __init__(self, label, saved_files=None, result_list=None): self.label = label # Default to no files: a hit verifies every saved file still exists, # and most of these tests are about key matching, not disk state self.saved_files = [] if saved_files is None else list(saved_files) + self.result_list = [] if result_list is None else result_list + + +def _frames(count, height=64, width=64, channels=3): + # A stand-in for decoded video frames - real weight, not a mock of it + return [np.zeros((height, width, channels), dtype=np.uint8) for _ in range(count)] def test_deep_equal_matches_identical_nested_dicts(): @@ -262,6 +270,60 @@ def test_the_default_cap_fits_a_maximal_for_each_run(): assert StepCache.DEFAULT_MAX_ENTRIES >= 2 * MAX_FOR_EACH_ENTRIES + 8 +def test_result_bytes_sums_array_frames_and_ignores_scalars(): + result = FakeResult("video", result_list=[{"videos": _frames(10), "fps": 24}]) + assert _result_bytes(result) == 10 * 64 * 64 * 3 + + +def test_result_bytes_does_not_double_count_a_shared_object(): + # get_artifact_list's fitted-in-place audio can be referenced by more + # than one artifact in the same result_list - the byte budget should + # not charge for it twice + audio = np.zeros(1000, dtype=np.float32) + result = FakeResult("av", result_list=[{"audio": audio}, {"audio": audio}]) + assert _result_bytes(result) == audio.nbytes + + +def test_step_cache_evicts_retained_entries_over_the_byte_budget(): + """Every shot@ member of a for_each group is legitimately retained (the + gather step genuinely reads all of them), but nothing bounded how much + decoded media that retention pins resident once the run that needed it + is done (#368) - the byte budget is the bound, on top of the entry cap.""" + frame_bytes = 64 * 64 * 3 + cache = StepCache(max_entries=128, max_retained_bytes=4 * frame_bytes) + for name in ("shot@a", "shot@b", "shot@c"): + result = FakeResult(name, result_list=[{"videos": _frames(2)}]) + cache.put("w", {"name": name}, 1, result, "/out", True) + + # 3 entries x 2 frames each = 6 frames worth, over the 4-frame budget - + # the least recently used (shot@a) is evicted despite being well under + # the entry-count cap + assert cache.get("w", {"name": "shot@a"}, 1, set(), "/out", True) is None + assert cache.get("w", {"name": "shot@b"}, 1, set(), "/out", True) is not None + assert cache.get("w", {"name": "shot@c"}, 1, set(), "/out", True) is not None + + +def test_step_cache_has_a_default_byte_budget(): + cache = StepCache() + assert cache.max_retained_bytes == StepCache.DEFAULT_MAX_RETAINED_BYTES + + +def test_step_cache_never_evicts_the_only_retained_entry_over_budget(): + cache = StepCache(max_retained_bytes=1) + result = FakeResult("big", result_list=[{"videos": _frames(5)}]) + cache.put("w", {"name": "only"}, 1, result, "/out", True) + + assert cache.get("w", {"name": "only"}, 1, set(), "/out", True) is result + + +def test_step_cache_unretained_result_does_not_count_against_the_byte_budget(): + cache = StepCache(max_retained_bytes=1) + heavy = FakeResult("heavy", result_list=[{"videos": _frames(5)}]) + cache.put("w", {"name": "unretained"}, 1, heavy, "/out", False) + + assert cache._retained_bytes == 0 + + def test_step_cache_clear(): cache = StepCache() step_data = {"name": "gen", "pipeline": {"arguments": {"prompt": "a cat"}}} diff --git a/tests/test_worker_execute.py b/tests/test_worker_execute.py index c17b9a74..20857864 100644 --- a/tests/test_worker_execute.py +++ b/tests/test_worker_execute.py @@ -146,6 +146,7 @@ def test_shutdown_during_run_cancels_then_flags_shutdown(): class StubResult: saved_files = [] + result_list = [] def test_full_cleanup_clears_step_cache(): From 25e145c01b6e61b64239eb7f6e5afb3c96ddf7d9 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 10:33:49 -0500 Subject: [PATCH 051/181] fix(engine): #370 - correct out-of-range at/frame: message and error A fractional bare number (e.g. 6.5) past a clip's end is a seconds overshoot, not a frame index in disguise, so the out-of-range message no longer suggests frame:N for it - only for an integer bare value that itself fits the clip's frame range. And a non-integer frame:N value (following that old advice) now gets a consumer-facing message instead of a raw int() ValueError. Co-Authored-By: Claude Sonnet 5 --- dw/media_frames.py | 17 +++++++++++++---- tests/test_media_frames.py | 37 +++++++++++++++++++++++++++++++++++++ 2 files changed, 50 insertions(+), 4 deletions(-) diff --git a/dw/media_frames.py b/dw/media_frames.py index b348ba9e..fa839629 100644 --- a/dw/media_frames.py +++ b/dw/media_frames.py @@ -215,7 +215,11 @@ def fit(image, _index): def _moment_to_index(moment, shape): total = shape["frame_count"] if isinstance(moment, str) and moment.startswith("frame:"): - index = int(moment[len("frame:") :]) + raw = moment[len("frame:") :] + try: + index = int(raw) + except ValueError: + raise ValueError(f'"{moment}" - frame index must be a whole number') if index < 0: index += total if not 0 <= index < total: @@ -235,11 +239,16 @@ def _moment_to_index(moment, shape): index += total if not 0 <= index < total: duration = total / fps - raise ValueError( + message = ( f"{moment!r} s is past the end of a {duration:.2f} s " - f"({total}-frame) clip - bare numbers in 'at' are seconds; " - f'use "frame:{_format_moment(moment)}" for a frame index' + f"({total}-frame) clip - bare numbers in 'at' are seconds" ) + # A fractional value is a seconds overshoot, not a frame index in + # disguise - the frame: hint only makes sense for a whole number + # that would itself be a valid frame index. + if seconds.is_integer() and 0 <= int(seconds) < total: + message += f'; use "frame:{_format_moment(moment)}" for a frame index' + raise ValueError(message) return index diff --git a/tests/test_media_frames.py b/tests/test_media_frames.py index 9825f6eb..03551bdb 100644 --- a/tests/test_media_frames.py +++ b/tests/test_media_frames.py @@ -138,6 +138,43 @@ def test_a_moment_past_the_end_is_refused(tmp_path): frames_at(str(tmp_path / "ramp.mp4"), ["frame:6"]) +def test_a_fractional_moment_past_the_end_gives_no_frame_hint(tmp_path): + # #370: a fractional overshoot is a seconds problem, not a frame index + # in disguise - "frame:6.5" would only fail again. + write_ramp_mp4(tmp_path / "ramp.mp4", frames=141, fps=24) + + with pytest.raises(ValueError, match="past the end") as excinfo: + frames_at(str(tmp_path / "ramp.mp4"), [6.5]) + assert "frame:" not in str(excinfo.value) + + +def test_an_integer_moment_past_the_end_in_both_units_gives_no_frame_hint(tmp_path): + write_ramp_mp4(tmp_path / "ramp.mp4", frames=141, fps=24) + + with pytest.raises(ValueError, match="past the end") as excinfo: + frames_at(str(tmp_path / "ramp.mp4"), [500]) + assert "frame:" not in str(excinfo.value) + + +def test_an_integer_moment_past_the_end_in_seconds_but_valid_as_a_frame_suggests_it( + tmp_path, +): + write_ramp_mp4(tmp_path / "ramp.mp4", frames=141, fps=24) + + with pytest.raises(ValueError, match="past the end") as excinfo: + frames_at(str(tmp_path / "ramp.mp4"), [130]) + assert 'use "frame:130" for a frame index' in str(excinfo.value) + + +def test_a_non_integer_frame_reference_gets_a_consumer_message(tmp_path): + write_ramp_mp4(tmp_path / "ramp.mp4", frames=141, fps=24) + + with pytest.raises( + ValueError, match=r'"frame:6\.5" - frame index must be a whole number' + ): + frames_at(str(tmp_path / "ramp.mp4"), ["frame:6.5"]) + + def test_a_contact_sheet_tiles_evenly_spaced_frames(tmp_path): write_ramp_mp4(tmp_path / "ramp.mp4", frames=24, fps=6, width=64, height=32) From 4aaeef7f74237b3bd255880ae28956105eafb639 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 10:40:14 -0500 Subject: [PATCH 052/181] fix(mcp): #372 - get_gallery_metadata audio hint matches -3 clip fix The next hint's prescription for a clipped audio/video file still said peak_dbfs: -1 and "a few tenths" of encoder overshoot, both stale since #362 moved the templates to -3 after measuring 0.56-1.56 dB overshoot on Music 3 mp3s. Match the hint to the audio_clipped warning's own -3 prescription and the measured range. Co-Authored-By: Claude Sonnet 5 --- dw_mcp/catalog.py | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 06272061..49b3a23c 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -308,9 +308,10 @@ def get_gallery_metadata(client, name, envelope=False, workspace=None): "the level normalize_audio would be given, and the range has two " "ends: mean_dbfs below -40 on a track that should be full is a " "near-silent render, and peak_dbfs at or above 0 is a deliverable " - "at or over full scale - a decoded lossy file overshoots by a few " - "tenths legitimately, but a figure of +1 or more is a mix with no " - "headroom, and 'normalize_audio' (peak_dbfs: -1) before the saving " - "step is what fixes it." + "at or over full scale - a decoded lossy file overshoots by up to " + "a couple dB legitimately (0.59-1.56 dB measured on Music 3 " + "mp3s), but a figure of +1 or more is a mix with no headroom, and " + "'normalize_audio' (peak_dbfs: -3) before the saving step is what " + "fixes it." ) return body From 7f5b63c6b583efc90ce8a782539db05076289d22 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 12:01:34 -0500 Subject: [PATCH 053/181] fix(engine): #343 - flat downloads stall, bytes never shrink, decimal units, cancel aborts - Count hf_xet's network transfer bytes via XetDownloadProgressReporter, not only the file on disk, which hf_xet holds flat while buffering - Measure published blobs through their symlink into the shared store and keep downloaded_bytes monotonic, so bytes_per_second is never negative - A zero-growth download_progress still fires (phase_detail reads "no bytes for Ns", seconds_since_bytes_changed on the event) but does not count as progress, so phase_stall fires once bytes stop - Size labels divide by 1000 to match GB/MB - A cancel during loading aborts the xet session / stops an HTTP download at its next chunk and surfaces as WorkflowCancelled Co-Authored-By: Claude Opus 5.5 --- dw/download_watch.py | 265 ++++++++++++++++++++----- dw/events.py | 15 +- dw/pipeline_processors/pipeline.py | 4 + dw/server/jobs.py | 1 + tests/test_download_watch.py | 309 ++++++++++++++++++++--------- 5 files changed, 442 insertions(+), 152 deletions(-) diff --git a/dw/download_watch.py b/dw/download_watch.py index a426401a..3a15114c 100644 --- a/dw/download_watch.py +++ b/dw/download_watch.py @@ -2,21 +2,39 @@ `from_pretrained` downloads with no hook of its own, and a pull that runs long looks identical to a hang: the phase-stall watchdog (dw/events.py) has -nothing but silence to report. Rather than intercepting the download - which -would mean second-guessing what `from_pretrained` fetches, and getting it -wrong for a variant or an alternate weight file it would not have pulled - -this only watches the Hugging Face cache directory the download writes into -while a `loading` phase is in progress, the same directory `hub_cache.py` -scans for `list_downloads`. Byte growth there is real progress whoever -triggered it, and is reported as such. - -A `download_progress` event fires on a fixed cadence, not only when the -watched size has grown (#343 follow-up): an xet-backed file reconstructs -against its local CAS cache in bursts, and can hold an unchanged size on -disk for well past the phase-stall threshold while a transfer is genuinely -still running underneath. Ticking on a timer reports that honestly - a -quiet interval is `bytes_per_second=0`, not silence - which is what keeps -the watchdog from mistaking it for a hang. +nothing but silence to report. This does not intercept or pre-fetch +anything - `from_pretrained` fetches exactly what it would have - it only +listens to the byte counts the download already reports while a `loading` +phase is in progress. + +Two signals, and the larger is reported: + +- The hub's own xet progress report (`XetDownloadProgressReporter`), which + carries the bytes *received from the network* as well as the bytes written + to disk. The network count is the one that matters: hf_xet buffers and + writes a file in order, and on lem held a 2.8 GB file at 67 MB on disk + while 700 MB had arrived, then wrote the rest at the very end. No location + on disk tells that apart from a hang; the transfer count does. Hooked + whether or not the hub's progress bars are displayed. +- The repo's cache directory, for a plain HTTP download (no hf_xet), which + writes its `.incomplete` file as bytes arrive. A published blob is a + symlink into the hub's shared store (`/blobs/xx/`), so the + size is taken through the link - measured beside it, the total fell by the + size of every file that finished (#343, a negative rate). + +Both only ever count up, so `downloaded_bytes` never goes backwards. + +A `download_progress` event fires about every EMIT_INTERVAL_SECONDS while a +download is underway. One reporting growth is progress. One reporting none +is not: it still fires, so `phase_detail` says `no bytes for 45s` instead of +repeating the last healthy rate, but the stall watchdog counts it as silence +and says `phase_stall` once bytes have been flat for its threshold - a +download and a hang no longer read the same. + +A cancel during a download aborts it: the xet session is aborted (the hub's +own KeyboardInterrupt path, `abort_xet_session`) and a plain HTTP download is +stopped at its next chunk, and the load surfaces as `WorkflowCancelled` +rather than running to the end of a multi-gigabyte file first. """ import logging @@ -28,12 +46,13 @@ from huggingface_hub.file_download import repo_folder_name from huggingface_hub.utils import HFValidationError, validate_repo_id +from .events import WorkflowCancelled, get_context + logger = logging.getLogger("dw") -# How often the watcher re-measures the cache directory, and the minimum gap -# between emitted events - well under the phase-stall threshold (30s) so a -# real, ongoing download never trips it, but not so tight that a run emits a -# progress event on every tick of a fast-growing directory. +# How often the watcher re-measures, and the gap between emitted events - +# well under the phase-stall threshold (30s) so a real, ongoing download +# never trips it, but not so tight that a run emits an event per chunk. CHECK_INTERVAL_SECONDS = 1.0 EMIT_INTERVAL_SECONDS = 5.0 @@ -49,31 +68,128 @@ def is_watchable_repo_id(name): def _blob_dir_size(blob_dir): + """Bytes under a repo's blobs/ directory, counting a published blob + through its symlink into the shared store.""" total = 0 try: with os.scandir(blob_dir) as entries: for entry in entries: try: - if entry.is_file(follow_symlinks=False): - total += entry.stat(follow_symlinks=False).st_size + if entry.is_file(): + total += entry.stat().st_size except OSError: - # A blob renamed or removed mid-scan (a completed - # .incomplete file, for instance) is not a fault + # A blob renamed or published mid-scan is not a fault continue except FileNotFoundError: return 0 return total +# The watches currently open and the hub hooks they share. One job runs at a +# time in the worker, but a hook is process-wide, so it is installed on the +# first open watch and removed with the last. +_lock = threading.Lock() +_active = [] +_originals = {} + + +def _note_hub_bytes(transferred, written): + with _lock: + watches = list(_active) + for watch_ in watches: + watch_._add_hub_bytes(transferred, written) + + +def _any_cancelled(): + with _lock: + return any(w._context.cancelled for w in _active) + + +def _install_hooks(): + try: + from huggingface_hub.utils import _xet_progress_reporting as xet_reporting + + reporter = xet_reporting.XetDownloadProgressReporter + original = reporter.update_progress + last_seen = {} + + def update_progress(self, group_report, *args, **kwargs): + # Runs on hf_xet's callback thread, which prints and swallows + # anything raised here - so nothing may be raised, and the + # hub's own bar update is guarded along with the counting + try: + key = id(self) + previous = last_seen.get(key, (0, 0)) + transferred = group_report.total_transfer_bytes_completed + written = group_report.total_bytes_completed + last_seen[key] = ( + max(previous[0], transferred), + max(previous[1], written), + ) + _note_hub_bytes( + max(0, transferred - previous[0]), max(0, written - previous[1]) + ) + except Exception as e: + logger.debug(f"Download watch could not read xet progress: {e}") + try: + return original(self, group_report, *args, **kwargs) + except Exception as e: + logger.debug(f"xet progress update raised: {e}") + + _patch(reporter, "update_progress", update_progress) + except (ImportError, AttributeError) as e: + logger.debug(f"No xet progress reporter to hook: {e}") + + try: + from importlib import import_module + + hub_tqdm = import_module("huggingface_hub.utils.tqdm").tqdm + original_update = hub_tqdm.update + + def update(self, n=1): + # A plain HTTP download calls this once per chunk on the thread + # running from_pretrained - the one place that can stop it + if _any_cancelled(): + raise WorkflowCancelled("Workflow run was cancelled") + return original_update(self, n) + + _patch(hub_tqdm, "update", update) + except (ImportError, AttributeError) as e: + logger.debug(f"No hub progress bar to hook: {e}") + + +def _patch(owner, name, replacement): + # Remembers whether the class defined the method itself or inherited it, + # so removing the hook restores exactly what was there + _originals[(owner, name)] = vars(owner).get(name) + setattr(owner, name, replacement) + + +def _remove_hooks(): + for (owner, name), original in _originals.items(): + if original is None: + delattr(owner, name) + else: + setattr(owner, name, original) + _originals.clear() + + +def _abort_xet_downloads(): + try: + from huggingface_hub.utils._xet import abort_xet_session + + abort_xet_session() + except Exception as e: + logger.debug(f"Could not abort the xet session: {e}") + + class DownloadWatch: - """Context manager that watches one repo's cache directory for growth - for as long as it is open, emitting a throttled `download_progress` - event on the active run context while bytes are arriving. + """Context manager that reports one repo's download on the active run + context for as long as it is open. Not a download itself - `from_pretrained` runs unmodified inside the - `with` block. Growth with nothing watching it (a cache miss the caller - already had, or a repo id from_pretrained resolves differently than - expected) simply emits nothing, same as no watch at all. + `with` block. A load that downloads nothing (every file cached) emits + nothing, same as no watch at all. """ def __init__(self, repo_id, context, cache_dir=None, repo_type="model"): @@ -87,47 +203,87 @@ def __init__(self, repo_id, context, cache_dir=None, repo_type="model"): ) self._stop = threading.Event() self._thread = None + self._counts_lock = threading.Lock() + self._transferred = 0 + self._written = 0 + self._aborted = False + + def _add_hub_bytes(self, transferred, written): + with self._counts_lock: + self._transferred += transferred + self._written += written + + def _hub_bytes(self): + with self._counts_lock: + return max(self._transferred, self._written) def __enter__(self): + with _lock: + if not _active: + _install_hooks() + _active.append(self) self._thread = threading.Thread( target=self._run, daemon=True, name="dw-download-watch" ) self._thread.start() return self - def __exit__(self, *exc_info): + def __exit__(self, exc_type, exc, traceback): self._stop.set() if self._thread is not None: self._thread.join(timeout=CHECK_INTERVAL_SECONDS * 2) + with _lock: + if self in _active: + _active.remove(self) + if not _active: + _remove_hooks() + if ( + exc is not None + and not isinstance(exc, WorkflowCancelled) + and self._context.cancelled + ): + # The abort surfaces as whatever the download library raises + # (hf_xet: "RuntimeError: Operation cancelled") - it is the + # cancel, not a load failure, and must read as one + raise WorkflowCancelled("Workflow run was cancelled") from exc return False def _run(self): baseline = _blob_dir_size(self._blob_dir) - last_emitted_size = baseline + disk_high_water = 0 + downloaded = 0 + emitted = 0 last_emit_at = time.monotonic() + last_change_at = last_emit_at while not self._stop.wait(CHECK_INTERVAL_SECONDS): try: - size = _blob_dir_size(self._blob_dir) + if self._context.cancelled and not self._aborted: + self._aborted = True + _abort_xet_downloads() + disk_high_water = max( + disk_high_water, _blob_dir_size(self._blob_dir) - baseline + ) now = time.monotonic() + current = max(downloaded, disk_high_water, self._hub_bytes()) + if current > downloaded: + downloaded = current + last_change_at = now elapsed = now - last_emit_at - if elapsed < EMIT_INTERVAL_SECONDS: + if downloaded == 0 or elapsed < EMIT_INTERVAL_SECONDS: continue - # Emitted on a timer, not on growth: a large xet-backed file - # reconstructs in bursts (dedup against its CAS cache) and - # can sit with an unchanged size on disk for well over - # EMIT_INTERVAL_SECONDS while genuinely still transferring - # (#343) - waiting for growth before emitting reproduces the - # exact silence the phase-stall watchdog is meant to catch. - # A tick with no growth still reports honestly: 0 B/s, not a - # guessed or carried-over rate. - rate = (size - last_emitted_size) / elapsed if elapsed > 0 else None - last_emitted_size = size + growth = downloaded - emitted + emitted = downloaded last_emit_at = now + # Flat bytes still report, so phase_detail stops showing the + # last healthy rate - but do not count as progress, so the + # stall watchdog sees the silence a hang is self._context.emit( "download_progress", + counts_as_progress=growth > 0, repo_id=self._repo_id, - downloaded_bytes=max(0, size - baseline), - bytes_per_second=rate, + downloaded_bytes=downloaded, + bytes_per_second=growth / elapsed, + seconds_since_bytes_changed=round(now - last_change_at, 1), ) except Exception as e: # A watch that breaks must not take the download down with it @@ -136,26 +292,31 @@ def _run(self): def _format_bytes(n): + # Decimal units, as the labels say and as the hub itself reports sizes - + # 1,879,623,333 bytes is 1.9 GB (it is 1.75 GiB) for unit in ("B", "KB", "MB", "GB", "TB"): - if n < 1024 or unit == "TB": + if n < 1000 or unit == "TB": return f"{n:.1f} {unit}" if unit != "B" else f"{n:.0f} {unit}" - n /= 1024 + n /= 1000 -def format_progress(repo_id, downloaded_bytes, bytes_per_second): +def format_progress( + repo_id, downloaded_bytes, bytes_per_second, seconds_since_bytes_changed=None +): """The `phase_detail` text a job's progress reports while this repo is - downloading - `downloading : 12.3 GB, 38 MB/s`, per #343.""" - text = f"downloading {repo_id}: {_format_bytes(downloaded_bytes)}" + downloading - `downloading : 12.3 GB, 38.0 MB/s` per #343, or + `downloading : 12.3 GB, no bytes for 45s` once bytes stop.""" + text = f"downloading {repo_id}: {_format_bytes(downloaded_bytes or 0)}" if bytes_per_second: text += f", {_format_bytes(bytes_per_second)}/s" + elif seconds_since_bytes_changed: + text += f", no bytes for {seconds_since_bytes_changed:.0f}s" return text def watch(repo_id, cache_dir=None, repo_type="model"): """A DownloadWatch on the active run context, or a no-op context manager when repo_id is not shaped like a hub repo (a local path, for instance).""" - from .events import get_context - if not is_watchable_repo_id(repo_id): return _NULL_WATCH return DownloadWatch( diff --git a/dw/events.py b/dw/events.py index 8ad10de2..57fc1d90 100644 --- a/dw/events.py +++ b/dw/events.py @@ -73,12 +73,17 @@ def __init__(self, on_event=None): self._watchdog_thread = None self._watchdog_stop = threading.Event() - def emit(self, event_type, **data): + def emit(self, event_type, counts_as_progress=True, **data): + """Send one event to the sink. counts_as_progress=False is for a + status report that says nothing has moved - a download reporting + zero new bytes (#343): the caller should see it, but the stall + watchdog must still see the silence, so it bumps neither clock.""" now = time.monotonic() - self._last_event_at = now - if not (event_type == "warning" and data.get("kind") == "phase_stall"): - self._last_progress_at = now - self._last_progress_kind = event_type + if counts_as_progress: + self._last_event_at = now + if not (event_type == "warning" and data.get("kind") == "phase_stall"): + self._last_progress_at = now + self._last_progress_kind = event_type if self._on_event is None: return try: diff --git a/dw/pipeline_processors/pipeline.py b/dw/pipeline_processors/pipeline.py index c7295983..ff0680c2 100644 --- a/dw/pipeline_processors/pipeline.py +++ b/dw/pipeline_processors/pipeline.py @@ -1879,6 +1879,10 @@ def load_component( component, component_name, configuration, device, components_manager ) + except WorkflowCancelled: + # A cancel that aborted a download (dw/download_watch.py) - not a + # load failure, so no error log + raise except Exception as e: # 401/403 from the Hub means the account behind whatever token (or # lack of one) HfApi is using cannot read this repo - almost always diff --git a/dw/server/jobs.py b/dw/server/jobs.py index 2b923671..9d741919 100644 --- a/dw/server/jobs.py +++ b/dw/server/jobs.py @@ -591,6 +591,7 @@ def _note_progress(self, event): event.get("repo_id"), event.get("downloaded_bytes"), event.get("bytes_per_second"), + event.get("seconds_since_bytes_changed"), ) elif kind == "warning": # Both channels, on purpose: the event log keeps the moment it diff --git a/tests/test_download_watch.py b/tests/test_download_watch.py index c208d471..4b01ce1c 100644 --- a/tests/test_download_watch.py +++ b/tests/test_download_watch.py @@ -1,34 +1,44 @@ -"""A download from_pretrained triggers is watched, not intercepted (#343): -byte growth in the Hugging Face cache counts as progress, so a real download -does not trip the phase-stall watchdog. - -Progress ticks on a timer, not only on growth (#343 follow-up): an -xet-backed transfer reconstructs against its local CAS cache in bursts and -can leave the watched size unchanged on disk for well past the phase-stall -threshold while genuinely still running underneath - confirmed by polling a -real large download (stabilityai/sd-turbo's non-fp16 unet, ~3.5GB) directly, -which sat with an unchanged blob size for 20+ seconds more than once before -jumping hundreds of MB at a time. There is no on-disk signal, in the blob -directory or in HF_XET_CACHE, that fills that gap, so the trade this file -makes deliberately: a watched download cannot read as a stall for as long as -DownloadWatch's thread is alive, in exchange for never mistaking a -reconstruction pause for one. What that costs is real: an `from_pretrained` -call that hangs with no eventual timeout or exception, for the length of the -hang, does not raise a phase-stall warning either - orthogonal to what #343 -reported, and a smaller loss than 17.8 minutes of false ones.""" +"""A download from_pretrained triggers is watched, not intercepted (#343). +What the tester held the earlier fixes to, and what these pin: + +- bytes arriving count as progress, so a real download does not trip the + phase-stall watchdog - measured through the hub's own xet progress + reporter, which counts network bytes while hf_xet holds the file on disk + unchanged (on lem, 67 MB written while 700 MB had arrived); +- bytes that stop arriving *do* stall: a flat download still reports, so + phase_detail reads `no bytes for Ns`, but the watchdog sees the silence; +- `downloaded_bytes` never goes backwards when a finished blob is published + into the shared store and replaced by a symlink; +- the size labels are the decimal units they claim to be; +- a cancel during a download aborts it rather than waiting it out. + +The reporter driven here is huggingface_hub's real +`XetDownloadProgressReporter`, fed the group reports hf_xet would send - not +a stub of DownloadWatch's own counting. +""" + +import logging +import os import time +from types import SimpleNamespace from unittest.mock import patch +import pytest +from huggingface_hub.utils import _xet_progress_reporting as xet_reporting +from huggingface_hub.utils.tqdm import tqdm as hub_tqdm + from dw import download_watch from dw import events as events_module -from dw.events import RunContext +from dw.download_watch import format_progress +from dw.events import RunContext, WorkflowCancelled +REPO_ID = "some-org/some-model" -# Scaled down from production's 5s emit / 30s stall, keeping the margin -# between them wide enough that scheduler jitter cannot open a stall-sized -# gap between two progress events. Each test still runs longer than the -# threshold, so silence would trip the watchdog. + +# Scaled down from production's 1s check / 5s emit / 30s stall, keeping the +# margin between emit and stall wide enough that scheduler jitter cannot +# open a stall-sized gap between two growing progress events. def _fast_watchdog(): return patch.multiple( events_module, @@ -37,110 +47,219 @@ def _fast_watchdog(): ) +@pytest.fixture(autouse=True) def _fast_watch(monkeypatch): monkeypatch.setattr(download_watch, "CHECK_INTERVAL_SECONDS", 0.02) monkeypatch.setattr(download_watch, "EMIT_INTERVAL_SECONDS", 0.05) -def test_growing_download_emits_progress_and_suppresses_stall(tmp_path, monkeypatch): - _fast_watch(monkeypatch) - repo_id = "some-org/some-model" - blob_dir = tmp_path / "models--some-org--some-model" / "blobs" - blob_dir.mkdir(parents=True) - blob_file = blob_dir / "abc123.incomplete" +def _group_report(transferred, written, total): + return SimpleNamespace( + total_bytes_completed=written, + total_transfer_bytes_completed=transferred, + total_bytes=total, + total_bytes_completion_rate=None, + total_transfer_bytes_completion_rate=None, + ) - events = [] + +def _reporter(total): + return xet_reporting.XetDownloadProgressReporter( + reconstruction_desc="model.safetensors: reconstructing file", + transfer_desc="model.safetensors: downloading bytes", + total=total, + log_level=logging.INFO, + name="huggingface_hub.xet_get", + ) + + +def _run_loading(events, body, cache_dir): context = RunContext(on_event=events.append) with _fast_watchdog(): context.enter_run() try: context.note_phase("loading") - with download_watch.DownloadWatch( - repo_id, context, cache_dir=str(tmp_path) - ): - for _ in range(20): - with open(blob_file, "ab") as f: - f.write(b"x" * 4096) - time.sleep(0.04) + with download_watch.DownloadWatch(REPO_ID, context, cache_dir=cache_dir): + body(context) finally: context.exit_run() + return context + + +def _progress(events): + return [e for e in events if e["event"] == "download_progress"] - progress = [e for e in events if e["event"] == "download_progress"] - assert progress, "growing bytes must be reported as download_progress" - assert progress[0]["repo_id"] == repo_id - assert all(e["downloaded_bytes"] > 0 for e in progress) - stalls = [e for e in events if e.get("kind") == "phase_stall"] - assert not stalls, "a real, ongoing download must not read as a stall" +def _stalls(events): + return [e for e in events if e.get("kind") == "phase_stall"] -def test_unchanging_size_still_heartbeats_and_does_not_stall(tmp_path, monkeypatch): - # A size that never changes for the life of the watch is exactly what a - # healthy xet reconstruction pause looks like on disk (#343 follow-up) - - # there is no growth-based signal available to tell it apart from a - # genuine hang, so the watch must still speak up on its own cadence. - _fast_watch(monkeypatch) - repo_id = "some-org/some-model" +def test_xet_transfer_counts_while_the_file_on_disk_is_flat(tmp_path): + # hf_xet's real shape: network bytes climb steadily while the bytes + # written to disk sit still until the end - the disk alone reads as a hang + total = 3_000_000_000 + events = [] + + def body(_context): + reporter = _reporter(total) + for i in range(1, 25): + reporter.update_progress(_group_report(i * 20_000_000, 67_000_000, total)) + time.sleep(0.04) + reporter.update_progress(_group_report(480_000_000, total, total)) + + _run_loading(events, body, str(tmp_path)) + + progress = _progress(events) + assert progress, "transfer bytes must be reported as download_progress" + assert progress[0]["repo_id"] == REPO_ID + sizes = [e["downloaded_bytes"] for e in progress] + assert sizes == sorted(sizes) + # Climbs well past the 67 MB on disk, and in steps, not one jump at the end + assert max(sizes) >= 300_000_000 + assert len(set(sizes)) >= 3 + assert all(e["bytes_per_second"] >= 0 for e in progress) + assert not _stalls(events), "a download receiving bytes must not stall" + + +def test_bytes_that_stop_arriving_stall_and_say_so(tmp_path): + events = [] + + def body(_context): + reporter = _reporter(3_000_000_000) + for i in range(1, 6): + reporter.update_progress(_group_report(i * 10_000_000, 0, 3_000_000_000)) + time.sleep(0.04) + time.sleep(1.2) # three stall thresholds with nothing arriving + + _run_loading(events, body, str(tmp_path)) + + flat = [e for e in _progress(events) if e["bytes_per_second"] == 0] + assert flat, "a flat download still reports, so phase_detail can say so" + assert flat[-1]["seconds_since_bytes_changed"] >= 0.8 + assert flat[-1]["downloaded_bytes"] == 50_000_000 + assert "no bytes for" in format_progress( + REPO_ID, + flat[-1]["downloaded_bytes"], + flat[-1]["bytes_per_second"], + flat[-1]["seconds_since_bytes_changed"], + ) + + stalls = _stalls(events) + assert stalls, "zero-growth reports must not keep the watchdog quiet" + assert stalls[0]["phase"] == "loading" + assert "download_progress" in stalls[0]["message"] + + +def test_publishing_a_blob_to_the_shared_store_does_not_shrink_the_total(tmp_path): + # The HTTP path, on disk as huggingface_hub 1.32 lays it out: the file + # grows as blobs/.incomplete, is renamed to blobs/, then moved + # into /blobs/xx/ and replaced by a symlink to it blob_dir = tmp_path / "models--some-org--some-model" / "blobs" blob_dir.mkdir(parents=True) - blob_file = blob_dir / "abc123.incomplete" - blob_file.write_bytes(b"x" * 4096) # written once, never grows again - + shared = tmp_path / "blobs" / "13" + shared.mkdir(parents=True) events = [] - context = RunContext(on_event=events.append) - with _fast_watchdog(): - context.enter_run() - try: - context.note_phase("loading") - with download_watch.DownloadWatch( - repo_id, context, cache_dir=str(tmp_path) - ): - time.sleep(0.8) - finally: - context.exit_run() - progress = [e for e in events if e["event"] == "download_progress"] - assert progress, "an active watch must heartbeat even with no growth" - assert all(e["bytes_per_second"] == 0 for e in progress) + def body(_context): + incomplete = blob_dir / "etag1.a1b2c3d4.incomplete" + for _ in range(8): + with open(incomplete, "ab") as f: + f.write(b"x" * 1_000_000) + time.sleep(0.04) + done = blob_dir / "etag1" + os.rename(incomplete, done) + target = shared / "1392abc" + os.rename(done, target) + os.symlink(os.path.relpath(target, blob_dir), done) + time.sleep(0.2) + + _run_loading(events, body, str(tmp_path)) - stalls = [e for e in events if e.get("kind") == "phase_stall"] - assert not stalls, "a watched download must not read as a stall" + progress = _progress(events) + sizes = [e["downloaded_bytes"] for e in progress] + assert sizes and sizes == sorted(sizes), f"downloaded_bytes went backwards: {sizes}" + assert sizes[-1] == 8_000_000 + assert all(e["bytes_per_second"] >= 0 for e in progress) -def test_large_file_freeze_then_burst_does_not_stall(tmp_path, monkeypatch): - # Reproduces the shape measured against a real ~3.5GB xet-backed download - # (stabilityai/sd-turbo unet): the tracked size holds flat for stretches - # well past one emit interval, then jumps hundreds of MB at once when - # reconstruction catches up - not a mock of DownloadWatch's own size - # function, a real directory written on that timing. - _fast_watch(monkeypatch) - repo_id = "some-org/some-model" +def test_a_cached_load_reports_nothing(tmp_path): blob_dir = tmp_path / "models--some-org--some-model" / "blobs" blob_dir.mkdir(parents=True) - blob_file = blob_dir / "abc123.incomplete" - blob_file.write_bytes(b"x" * 1024) + (blob_dir / "etag1").write_bytes(b"x" * 4096) + events = [] + _run_loading(events, lambda _context: time.sleep(0.2), str(tmp_path)) + assert not _progress(events) + + +def test_cancel_aborts_an_xet_download_and_reads_as_cancelled(tmp_path, monkeypatch): + aborted = [] + monkeypatch.setattr( + download_watch, "_abort_xet_downloads", lambda: aborted.append(True) + ) + events = [] + + def body(context): + context.cancel() + deadline = time.monotonic() + 2 + while not aborted and time.monotonic() < deadline: + time.sleep(0.01) + # What hf_xet raises in the loading thread once its session aborts + raise RuntimeError("Operation cancelled: Task cancelled") + + with pytest.raises(WorkflowCancelled): + _run_loading(events, body, str(tmp_path)) + assert aborted, "a cancel during loading must abort the in-flight download" + +def test_cancel_stops_a_plain_http_download_at_its_next_chunk(tmp_path): + events = [] + + def body(context): + bar = hub_tqdm(total=10, unit="B", disable=True) + bar.update(1) + context.cancel() + bar.update(1) + + with pytest.raises(WorkflowCancelled): + _run_loading(events, body, str(tmp_path)) + + +def test_hooks_are_removed_when_the_last_watch_closes(tmp_path): + reporter_method = vars(xet_reporting.XetDownloadProgressReporter)["update_progress"] + had_update = "update" in vars(hub_tqdm) + _run_loading([], lambda _context: None, str(tmp_path)) + assert ( + vars(xet_reporting.XetDownloadProgressReporter)["update_progress"] + is reporter_method + ) + assert ("update" in vars(hub_tqdm)) == had_update + + +def test_sizes_are_labelled_in_the_units_they_are(): + # The tester's own reading: 1,879,623,333 bytes rendered as "1.8 GB" + assert download_watch._format_bytes(1_879_623_333) == "1.9 GB" + assert download_watch._format_bytes(38_000_000) == "38.0 MB" + assert ( + format_progress(REPO_ID, 12_300_000_000, 38_000_000) + == f"downloading {REPO_ID}: 12.3 GB, 38.0 MB/s" + ) + assert ( + format_progress(REPO_ID, 12_300_000_000, 0, 45.2) + == f"downloading {REPO_ID}: 12.3 GB, no bytes for 45s" + ) + + +def test_a_non_progress_event_does_not_reset_the_watchdog(): events = [] context = RunContext(on_event=events.append) with _fast_watchdog(): context.enter_run() try: context.note_phase("loading") - with download_watch.DownloadWatch( - repo_id, context, cache_dir=str(tmp_path) - ): - # Freeze well past one emit interval (0.05s) ... - time.sleep(0.3) - # ... then a burst, then freeze again. - with open(blob_file, "ab") as f: - f.write(b"x" * (16 * 1024 * 1024)) - time.sleep(0.3) + for _ in range(40): + context.emit("download_progress", counts_as_progress=False, x=1) + time.sleep(0.02) finally: context.exit_run() - - progress = [e for e in events if e["event"] == "download_progress"] - assert progress, "growth and quiet stretches must both be reported" - assert progress[-1]["downloaded_bytes"] >= 16 * 1024 * 1024 - - stalls = [e for e in events if e.get("kind") == "phase_stall"] - assert not stalls, "a freeze-then-burst download must not read as a stall" + assert _stalls(events) + assert all("counts_as_progress" not in e for e in events) From 5ccd1fc349e96098699f84509c57acb7254df11e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 12:17:32 -0500 Subject: [PATCH 054/181] fix(engine): #349 - grade accepts video input, enforces temperature/tint range grade's image-only parameter forced every string argument through the image loader, refusing a video passed as `image` and dropping audio when passed as `video` (both key names trigger automatic, name-based media loading in realize_args). Renamed the parameter to `media`, which escapes both conventions, and dispatch on file extension in _handle_grade: a video path loads via load_audio_video (keeping fps and audio, mirroring stabilize_video's own escape of the same convention), an image path via fetch_image, and an already-loaded object (from a previous step or an in-memory call) passes straight to _per_frame as before. temperature and tint are documented as a -1.0..1.0 scale but were not domain-checked, so an out-of-range value (e.g. temperature=5) silently produced an unmodelled, visibly wrong result instead of being refused. Added a CLOSED_UNIT domain to dw/task_domains.py, declared for both arguments, checked statically (validate_workflow) and at run time (check_arguments) like every other domain. domain_violation's refusal text is now looked up per domain instead of hardcoding audio-specific wording, so grade's message talks about a documented range rather than "a plausible-looking track of the wrong length or speed". Also updated the handler's summary to mention video alongside image. Co-Authored-By: Claude Sonnet 5 --- dw/task_domains.py | 37 ++++++++++++++++++++----- dw/tasks/grade.py | 14 +++++----- dw/tasks/task.py | 21 ++++++++++++--- tests/test_grade.py | 55 ++++++++++++++++++++++++++++++++------ tests/test_task_domains.py | 24 +++++++++++++++-- 5 files changed, 124 insertions(+), 27 deletions(-) diff --git a/dw/task_domains.py b/dw/task_domains.py index 3d79cb9e..ecf48d90 100644 --- a/dw/task_domains.py +++ b/dw/task_domains.py @@ -21,9 +21,12 @@ the command would refuse anyway - it is the earlier of two answers, not a second opinion. -Only numbers whose domain is not a judgement call are listed. A level in dBFS, -a gain, a colour: those are the command's business, and a command that wants -to refuse something subtler does it in its own body. +Only numbers whose domain is not a judgement call are listed. Whether a level +in dBFS or a gain is the *right* one for a mix is the command's business, and +a command that wants to refuse something subtler does it in its own body - +but a value's documented range is not a judgement call either: grade's +temperature and tint are only defined from -1.0 to 1.0, and a value outside +it is not a bolder version of the effect, just an unmodelled one (#349). """ import logging @@ -42,13 +45,31 @@ # (0 is full scale, positive is not a level any of these commands can reach), # shared here with normalize_audio's target_lufs NON_POSITIVE = "non_positive" +# A value documented as a scale from -1.0 to 1.0 - grade's temperature and +# tint, whose linear interpolation is only defined inside that range; outside +# it the same formula still runs and produces a value, just not the one the +# documented scale promised +CLOSED_UNIT = "closed_unit" _DOMAIN_TEXT = { POSITIVE: "above zero", NON_NEGATIVE: "zero or above", NON_POSITIVE: "at or below full scale (0)", + CLOSED_UNIT: "between -1.0 and 1.0", } +_DOMAIN_REASON = { + CLOSED_UNIT: ( + "A value outside that range is refused rather than extrapolated - " + "the documented scale is only defined inside it" + ), +} +_DEFAULT_REASON = ( + "A value outside that range is refused rather than interpreted - " + "a negative count or a zero rate would otherwise produce a " + "plausible-looking track of the wrong length or speed" +) + # command -> argument -> domain. Every entry here is pinned to a real command # and a real parameter of it by tests/test_task_domains.py, so a renamed # argument cannot leave a domain checking nothing @@ -125,6 +146,8 @@ "grade": { "contrast": NON_NEGATIVE, "saturation": NON_NEGATIVE, + "temperature": CLOSED_UNIT, + "tint": CLOSED_UNIT, }, } @@ -139,6 +162,8 @@ def in_domain(value, domain): return number > 0 if domain == NON_POSITIVE: return number <= 0 + if domain == CLOSED_UNIT: + return -1.0 <= number <= 1.0 return number >= 0 @@ -186,11 +211,9 @@ def domain_violation(command, name, value, domain): if in_domain(item, domain): continue label = f"{name}[{index}]" if index is not None else name + reason = _DOMAIN_REASON.get(domain, _DEFAULT_REASON) message = ( - f"{command} needs '{label}' {_DOMAIN_TEXT[domain]}, got {item!r}. " - f"A value outside that range is refused rather than interpreted - " - f"a negative count or a zero rate would otherwise produce a " - f"plausible-looking track of the wrong length or speed" + f"{command} needs '{label}' {_DOMAIN_TEXT[domain]}, got {item!r}. {reason}" ) return index, message return None diff --git a/dw/tasks/grade.py b/dw/tasks/grade.py index 136b8802..91570459 100644 --- a/dw/tasks/grade.py +++ b/dw/tasks/grade.py @@ -23,14 +23,14 @@ def grade_image( - image, + media, exposure=0.0, contrast=1.0, saturation=1.0, temperature=0.0, tint=0.0, ): - """Adjust exposure, contrast, saturation and white balance of an image. + """Adjust exposure, contrast, saturation and white balance of a single frame. Every parameter is optional; omitting one leaves that adjustment at its identity value, so calling with no arguments returns the input pixels @@ -38,7 +38,9 @@ def grade_image( then contrast, then temperature/tint, then saturation. Args: - image: PIL Image to grade. + media: PIL Image to grade. A video is dispatched to this one frame at + a time by the command handler (task.py's _per_frame), so this + function itself only ever sees a single frame. exposure: Stops to brighten (positive) or darken (negative) by, applied as a multiply of 2**exposure. 0.0 (default) is identity. contrast: Multiplier applied around the mid grey point (0.5). @@ -60,10 +62,10 @@ def grade_image( channel, if the input has one, passes through untouched. """ alpha = None - if image.mode in ("RGBA", "LA"): - alpha = image.getchannel("A") + if media.mode in ("RGBA", "LA"): + alpha = media.getchannel("A") - rgb = image.convert("RGB") + rgb = media.convert("RGB") array = np.asarray(rgb, dtype=np.float32) / 255.0 if exposure != 0.0: diff --git a/dw/tasks/task.py b/dw/tasks/task.py index b2d1b1e1..5c190fc7 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -428,12 +428,25 @@ def _handle_restore_faces(task, arguments, previous_pipelines): @register_command("grade", implementation="dw.tasks.grade.grade_image") def _handle_grade(task, arguments, previous_pipelines): - """Adjust exposure, contrast, saturation and white balance of an image or video""" - logger.debug("Grading image") - image = arguments.pop("image") + """Adjust exposure, contrast, saturation and white balance of an image or a video""" + logger.debug("Grading media") + media = arguments.pop("media") from .grade import grade_image - return _per_frame(image, lambda frame: grade_image(frame, **arguments)) + if isinstance(media, str): + import os + + from ..security import ALLOWED_VIDEO_EXTENSIONS + from .video_utils import load_audio_video + + if os.path.splitext(media)[1].lower() in ALLOWED_VIDEO_EXTENSIONS: + media = load_audio_video(media) + else: + from ..arguments import fetch_image + + media = fetch_image(media) + + return _per_frame(media, lambda frame: grade_image(frame, **arguments)) @register_command( diff --git a/tests/test_grade.py b/tests/test_grade.py index cedade35..33d51b3c 100644 --- a/tests/test_grade.py +++ b/tests/test_grade.py @@ -117,7 +117,7 @@ def video(self): def test_grade_dispatches_over_every_frame_and_keeps_audio_and_fps(self): task = Task({"command": "grade", "arguments": {}}, "cpu") - result = task.run({"image": self.video(), "exposure": 1.0}) + result = task.run({"media": self.video(), "exposure": 1.0}) assert isinstance(result, AudioVideo) assert len(result.frames) == 3 @@ -130,6 +130,36 @@ def test_grade_dispatches_over_every_frame_and_keeps_audio_and_fps(self): numpy.asarray(source_frame), numpy.asarray(graded_frame) ) + def test_a_video_file_path_is_loaded_with_its_audio(self, tmp_path): + # The regression this guards: `media` is deliberately not called + # "image" or "video" - either name would make the engine's own + # key-convention loading grab the string first (an image-only + # loader that refuses a .mp4 extension, or a frame-only loader that + # silently drops the audio), before the command ever saw it. + import av + + path = tmp_path / "clip.mp4" + container = av.open(str(path), mode="w") + stream = container.add_stream("libx264rgb", rate=24) + stream.width, stream.height = 4, 4 + stream.pix_fmt = "rgb24" + for i in range(3): + frame = av.VideoFrame.from_ndarray( + numpy.full((4, 4, 3), i * 40, dtype=numpy.uint8), format="rgb24" + ) + for packet in stream.encode(frame): + container.mux(packet) + for packet in stream.encode(): + container.mux(packet) + container.close() + + task = Task({"command": "grade", "arguments": {}}, "cpu") + result = task.run({"media": str(path), "exposure": 1.0}) + + assert isinstance(result, AudioVideo) + assert len(result.frames) == 3 + assert result.fps == 24 + class TestDomains: def _errors(self, arguments): @@ -146,15 +176,23 @@ def _errors(self, arguments): return task_argument_errors(definition) def test_a_negative_contrast_is_refused_at_its_path(self): - errors = self._errors({"image": "asset:a.png", "contrast": -0.5}) + errors = self._errors({"media": "asset:a.png", "contrast": -0.5}) assert [e["path"] for e in errors] == ["steps[0].task.arguments.contrast"] def test_a_negative_saturation_is_refused(self): - errors = self._errors({"image": "asset:a.png", "saturation": -1.0}) + errors = self._errors({"media": "asset:a.png", "saturation": -1.0}) assert [e["path"] for e in errors] == ["steps[0].task.arguments.saturation"] def test_zero_saturation_is_a_legitimate_request(self): - assert self._errors({"image": "asset:a.png", "saturation": 0.0}) == [] + assert self._errors({"media": "asset:a.png", "saturation": 0.0}) == [] + + def test_an_out_of_range_temperature_is_refused(self): + errors = self._errors({"media": "asset:a.png", "temperature": 5.0}) + assert [e["path"] for e in errors] == ["steps[0].task.arguments.temperature"] + + def test_an_out_of_range_tint_is_refused(self): + errors = self._errors({"media": "asset:a.png", "tint": -1.5}) + assert [e["path"] for e in errors] == ["steps[0].task.arguments.tint"] def test_validate_workflow_refuses_it_before_a_run(self): workflow = Workflow( @@ -165,7 +203,7 @@ def test_validate_workflow_refuses_it_before_a_run(self): "name": "grade", "task": { "command": "grade", - "arguments": {"image": "asset:a.png", "contrast": -1.0}, + "arguments": {"media": "asset:a.png", "contrast": -1.0}, }, "result": {"content_type": "image/png"}, } @@ -177,13 +215,14 @@ def test_validate_workflow_refuses_it_before_a_run(self): errors = workflow.validation_errors() assert [e["path"] for e in errors] == ["steps[0].task.arguments.contrast"] - def test_exposure_and_temperature_are_unconstrained(self): + def test_exposure_is_unconstrained_temperature_and_tint_are_closed_unit(self): from dw.introspection import describe_task + from dw.task_domains import CLOSED_UNIT parameters = {p["name"]: p for p in describe_task("grade")["parameters"]} assert "domain" not in parameters["exposure"] - assert "domain" not in parameters["temperature"] - assert "domain" not in parameters["tint"] + assert parameters["temperature"]["domain"] == CLOSED_UNIT + assert parameters["tint"]["domain"] == CLOSED_UNIT if __name__ == "__main__": diff --git a/tests/test_task_domains.py b/tests/test_task_domains.py index 15e00170..c071187e 100644 --- a/tests/test_task_domains.py +++ b/tests/test_task_domains.py @@ -15,6 +15,7 @@ from dw.task_domains import ( TASK_ARGUMENT_DOMAINS, + CLOSED_UNIT, NON_NEGATIVE, NON_POSITIVE, POSITIVE, @@ -60,9 +61,14 @@ def test_every_argument_is_a_parameter_of_its_command(self): parameters = {p["name"] for p in describe_task(command)["parameters"]} assert set(domains) <= parameters, command - def test_every_domain_is_one_of_the_three(self): + def test_every_domain_is_one_of_the_declared_kinds(self): for domains in TASK_ARGUMENT_DOMAINS.values(): - assert set(domains.values()) <= {POSITIVE, NON_NEGATIVE, NON_POSITIVE} + assert set(domains.values()) <= { + POSITIVE, + NON_NEGATIVE, + NON_POSITIVE, + CLOSED_UNIT, + } class TestAsNumber: @@ -108,6 +114,20 @@ def test_a_positive_target_lufs_is_refused(self): assert len(errors) == 1 assert errors[0]["path"] == "steps[0].task.arguments.target_lufs" + def test_an_out_of_range_temperature_is_refused(self): + errors = errors_for("grade", {"media": "asset:a.png", "temperature": 5.0}) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].task.arguments.temperature" + assert "-1.0 and 1.0" in errors[0]["message"] + + def test_an_out_of_range_tint_is_refused(self): + errors = errors_for("grade", {"media": "asset:a.png", "tint": -1.5}) + assert len(errors) == 1 + assert errors[0]["path"] == "steps[0].task.arguments.tint" + + def test_a_boundary_temperature_is_fine(self): + assert errors_for("grade", {"media": "asset:a.png", "temperature": 1.0}) == [] + def test_a_zero_offset_is_fine(self): assert ( errors_for( From 4a09cadb75cd7cdf71260eb41cb615443c27c432 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 12:25:03 -0500 Subject: [PATCH 055/181] fix(mcp): #354 - complete the transcription recipe with validate + acknowledged_cost The docstring and both skills told callers to run templates/transcribe-audio directly, but run_workflow refuses without an acknowledged_cost bound to a plan - that call was refused as documented (per the #354 bounce). Point at validate_workflow first, run_workflow with acknowledged_cost bound to its plan, then get_output_text and delete_output. The plan's basis is "unknown" (minutes: null), so tell callers to quote seconds, not minutes - it takes a few seconds in practice. Drops the "two calls" undercount. Co-Authored-By: Claude Sonnet 5 --- dw_mcp/server.py | 18 ++++++++++++------ plugins/dw/skills/minimax-h3/SKILL.md | 10 +++++----- plugins/dw/skills/series-episodes/SKILL.md | 12 ++++++++---- 3 files changed, 25 insertions(+), 15 deletions(-) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index aee628ae..ffb70253 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -534,12 +534,18 @@ def get_output_audio( says what was cut. To *see* a video, `get_output_frames`. To confirm the *words* an output speaks rather than hear it - a text-only client can't consume the `AudioContent` block this - returns - run `run_workflow(name="templates/transcribe-audio", - arguments={"input_audio": "output:"}, wait_seconds=55)` (it - takes an audio file or a video's muxed soundtrack directly) then - `get_output_text` on the result, and `delete_output` the scratch - run afterward. Two calls and a short wait, not a GPU-spending read - tool - keep the normal queue rather than adding one. + returns - `validate_workflow(name="templates/transcribe-audio", + arguments={"input_audio": "output:"})` first (free; it takes + an audio file or a video's muxed soundtrack directly), then + `run_workflow(..., acknowledged_cost={"fingerprint": ..., + "minutes": ..., "downloads": [...]})` bound to that plan with + `wait_seconds=55`, then `get_output_text` on the result, and + `delete_output(job_id=...)` the scratch run afterward. This + workflow's plan comes back `basis: "unknown"` with `minutes: null` + - nothing is curated or observed for it - so quote what it actually + takes rather than the plan: seconds, not minutes (a few seconds per + clip in practice). Four calls and a short wait, not a GPU-spending + read tool - keep the normal queue rather than adding one. `workspace` pins this call to another workspace.""" result = media.get_output_audio( diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 88fc6ac8..f0049bfd 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -180,11 +180,11 @@ itself - or every shot inherits the portrait's composition. everywhere), a portrait imposing its framing on every shot - `at` late in a chain for drift sharpening to noise, and `get_output_audio` for a voice-over without affect. `get_output_audio` returns sound, not text; to - confirm a line rendered rather than judge its delivery, - `run_workflow(name="templates/transcribe-audio", - arguments={"input_audio": "output:"}, wait_seconds=55)` then - `get_output_text` on the result, and `delete_output` the scratch run. - Then `get_gallery_metadata` for duration and whether audio is present, + confirm a line rendered, use its docstring's transcription route + (`validate_workflow` on `templates/transcribe-audio`, `run_workflow` with + `acknowledged_cost` bound, `get_output_text`, `delete_output`) - `basis: + "unknown"`, quote seconds not minutes. + Then `get_gallery_metadata` for duration/audio presence, and hand the user the gallery `url` (`list_gallery`). `get_output_image` works only on image steps - the Z-Image portraits and boards of `dialogue-short`, `storyboard`, diff --git a/plugins/dw/skills/series-episodes/SKILL.md b/plugins/dw/skills/series-episodes/SKILL.md index 2f5732c8..b94d5bff 100644 --- a/plugins/dw/skills/series-episodes/SKILL.md +++ b/plugins/dw/skills/series-episodes/SKILL.md @@ -130,10 +130,14 @@ mismatch between episodes shows up before a viewer notices it - its as a whole; `integrated_lufs` (BS.1770, whole-track) is the field that answers that, and is what to compare across episodes. To confirm a line actually rendered rather than judging it by ear, `get_output_audio` -returns sound, not text: `run_workflow(name="templates/transcribe-audio", -arguments={"input_audio": "output:"}, wait_seconds=55)` then -`get_output_text` on the result (it takes the episode's muxed soundtrack -directly), and delete the scratch run afterward. +returns sound, not text: `validate_workflow(name="templates/transcribe-audio", +arguments={"input_audio": "output:"})` first (free; it takes the +episode's muxed soundtrack directly), then `run_workflow(..., +acknowledged_cost=, +wait_seconds=55)`, then `get_output_text` on the result, and +`delete_output(job_id=...)` the scratch run afterward. This workflow's +plan is `basis: "unknown"` with `minutes: null` - quote seconds, not +minutes; it runs in a few seconds. ## Sources From 95391f28242277f9e938c07bc831416262006a9e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 12:36:51 -0500 Subject: [PATCH 056/181] fix(worker): #368 - release host caches at every job end, not just clear_memory The prior fix (d5f8d5d) targeted step_cache's retained Results, but those are only 0.4-1.6GB for dialogue-short's two shot videos - far under the 4GiB budget, so eviction never fired and RSS still grew ~3.6GB/shot with ~10.6GB resident after H3 release. The real gap: release_pipeline's gc.collect() + empty_device_cache() frees GPU memory but never touches two host-side caches - torch's pinned-host allocator (group_offload's staging buffers) and glibc's malloc arenas the SDNQ weights were read into. release_host_caches() (dw/host_memory.py, added for #98) already frees exactly these and is documented safe to call with models still resident, since it only touches blocks nothing is using. It was only ever called from _cleanup_all(), which also clears loaded_pipelines/shared_components/step_cache - explaining why clear_memory reclaimed the memory but a normal job end did not. _cleanup_between_runs() now calls release_host_caches() too, without touching the warm caches that exist so a rerun doesn't reload. Co-Authored-By: Claude Sonnet 5 --- dw/worker.py | 15 +++++++++++++++ tests/test_worker_execute.py | 22 ++++++++++++++++++++++ 2 files changed, 37 insertions(+) diff --git a/dw/worker.py b/dw/worker.py index ad382b7a..325552c2 100644 --- a/dw/worker.py +++ b/dw/worker.py @@ -540,6 +540,21 @@ def _cleanup_between_runs(self): except Exception as e: logger.warning(f"Could not clean GPU cache: {e}") + # release_host_caches() only touches blocks nothing is using - the + # pinned-host staging buffers of a step's group_offload and the + # glibc arenas a released pipeline's weights were read into - so it + # is safe here even though loaded_pipelines/shared_components are + # still warm for the next run. Without it those two caches are the + # gap between what a job's own cleanup releases and what an explicit + # clear_memory does (#368): a released pipeline's GPU memory drops at + # `release_pipeline`, but the host arenas it staged through, plus any + # SDNQ/group_offload residue from steps that never released at all + # because the job ended first, stay resident until something calls + # this. + released = release_host_caches() + if released: + logger.info(f"Inter-run cleanup returned {released:.0f} MB to the OS") + # Check for memory growth current_memory = self._get_gpu_memory_mb() if current_memory > 0: diff --git a/tests/test_worker_execute.py b/tests/test_worker_execute.py index 20857864..826da9ec 100644 --- a/tests/test_worker_execute.py +++ b/tests/test_worker_execute.py @@ -393,3 +393,25 @@ def test_a_workflow_switch_forgets_the_prior_keys(): command={"workflow_path": "other.json", "arguments": {}, "output_dir": "/tmp"}, ) assert worker.prior_step_keys == {} + + +def test_between_run_cleanup_releases_host_caches_without_clearing_pipelines(): + """#368: a job's own cleanup left ~10GB resident that only clear_memory + reclaimed - the pinned-host staging buffers of group_offload and the + glibc arenas a released pipeline's weights were read into. Neither is + touched by gc.collect()/empty_device_cache() alone, so the light, + every-job cleanup must also call release_host_caches() - and must keep + loaded_pipelines/shared_components warm while doing it, since those + exist for exactly this (inter-run) cleanup to leave alone. + """ + worker = _make_worker() + worker.loaded_pipelines["warm-key"] = object() + worker.shared_components["warm-component"] = object() + + with patch("dw.worker.release_host_caches", return_value=512.0) as released: + worker._cleanup_between_runs() + + released.assert_called_once() + # the whole point: still-warm state for the next run survives this call + assert "warm-key" in worker.loaded_pipelines + assert "warm-component" in worker.shared_components From afd695348aa4a04e6f3909550fb2c7459729cdbe Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 13:06:11 -0500 Subject: [PATCH 057/181] fix(workflow): #368 - return host caches at pipeline release, not only at job end release_pipeline, release_models and a superseded pipeline's release now call release_host_caches() after dropping the model, so the pinned staging buffers and heap arenas it held leave RSS before pipeline_released is announced, rather than staying through every later step until the worker's between-run cleanup. Co-Authored-By: Claude Opus 5.5 --- dw/workflow.py | 26 +++++++ tests/test_pipeline_caching.py | 134 +++++++++++++++++++++++++++++++++ 2 files changed, 160 insertions(+) diff --git a/dw/workflow.py b/dw/workflow.py index 38c8df2a..129396ab 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -86,6 +86,7 @@ from .tasks.model_cache import clear_model_cache from .tasks.task import Task from . import get_device, empty_device_cache, device_memory_stats +from .host_memory import release_host_caches from .security import ( validate_path, validate_workflow_path, @@ -252,6 +253,21 @@ def _allocated_mb(): return stats["allocated_mb"] if stats["available"] else None +def _release_host_caches(step_name): + """Hand the host memory a release freed back to the OS, not at job end. + + `release_host_caches` only touches blocks nothing is using, so anything + still loaded is undisturbed. A cleanup is never worth failing a run for. + """ + try: + released = release_host_caches() + except Exception as e: + logger.debug(f"Could not release host caches after {step_name}: {e}") + return + if released: + logger.info(f"Release after {step_name} returned {released:.0f} MB to the OS") + + def selected_field(step_data, selected): """The manifest/step_end 'selected' block for a step's Result.selected. @@ -1418,6 +1434,13 @@ def run( step_action = None gc.collect() empty_device_cache() + # The device cache is not the only one the release fills: + # the pinned-host staging buffers the pipeline offloaded + # through and the heap arenas its weights were read into + # stay in this process's RSS until they are handed back, + # which otherwise waits for the end of the job - ~10 GB + # held through every step after the release (#368) + _release_host_caches(step.name) # Say so on the event stream. The release is otherwise # invisible to a consumer: it sits inside the sub-second # window between a step's generation and its files @@ -1512,6 +1535,8 @@ def run( if step_data.get("release_models", False): logger.info(f"Releasing task models for step: {step.name}") clear_model_cache() + gc.collect() + _release_host_caches(step.name) # Cleanup between steps (but keep pipelines loaded). Returning # cached blocks to the device lets the next step's differently @@ -1738,6 +1763,7 @@ def create_step_action( previous_pipelines.pop(prior_key, None) gc.collect() empty_device_cache() + _release_host_caches(step_name) # Say so on the event stream, for the same reason the explicit # release does: without it a reload-on-top-of-a-resident-model # is indistinguishable from a cold load, and the difference is diff --git a/tests/test_pipeline_caching.py b/tests/test_pipeline_caching.py index dae119d1..62bb0814 100644 --- a/tests/test_pipeline_caching.py +++ b/tests/test_pipeline_caching.py @@ -661,3 +661,137 @@ def test_superseded_release_is_reported_on_the_event_stream(): assert len(released) == 1, f"expected one release event, got {events}" assert released[0]["step"] == "gen" assert released[0]["reason"] == "superseded" + + +def test_release_pipeline_returns_host_caches_before_announcing_it(): + """The release hands pinned/arena host memory back then, not at job end. + + `pipeline_released` dropped the device memory but left ~10 GB of pinned + staging buffers and heap arenas in RSS through every later step of the + job, until the worker's between-run cleanup finally returned them (#368). + The host caches are emptied after the pipeline has left the cache, and + before the event, so the memory reading that follows it shows the drop. + """ + from dw.events import RunContext + + events = [] + pipeline_cache = {} + workflow = Workflow(_release_workflow_def(), "/tmp/test_output", "test.json") + at_release = [] + + def fake_release(): + key = workflow._pipeline_keys_by_step.get("generate") + at_release.append( + { + "released_cached": key in pipeline_cache, + "announced": any(e["event"] == "pipeline_released" for e in events), + } + ) + return 0.0 + + def mock_pipeline_load(self, shared_components): + self.pipeline = MagicMock() + + with patch.object(Pipeline, "load", mock_pipeline_load): + with patch.object( + Step, "run", lambda self, *args, **kwargs: MagicMock(result_list=[]) + ): + with patch("dw.workflow.empty_device_cache"): + with patch("dw.workflow.release_host_caches", fake_release): + workflow.run( + {}, + previous_pipelines=pipeline_cache, + context=RunContext(on_event=events.append), + ) + + assert at_release == [{"released_cached": False, "announced": False}] + # the step that did not ask for a release leaves its pipeline warm + assert workflow._pipeline_keys_by_step["keep"] in pipeline_cache + + +def test_release_models_returns_host_caches(): + """release_models is the task-model sibling of release_pipeline (#368).""" + workflow = Workflow( + _release_models_workflow_def(True), "/tmp/test_output", "test.json" + ) + calls = [] + + def mock_pipeline_load(self, shared_components): + self.pipeline = MagicMock() + + cached_model(("text_generation", "some-model", "cuda"), MagicMock) + try: + with patch.object(Pipeline, "load", mock_pipeline_load): + with patch.object( + Step, "run", lambda self, *args, **kwargs: MagicMock(result_list=[]) + ): + with patch("dw.workflow.empty_device_cache"): + with patch( + "dw.workflow.release_host_caches", + lambda: calls.append(bool(_model_cache)) or 0.0, + ): + workflow.run({}, previous_pipelines={}) + finally: + clear_model_cache() + + assert calls == [False], "emptied once, after the task models were dropped" + + +def test_superseded_release_returns_host_caches(): + """A redefined step's old pipeline hands its host memory back before the + new one loads, not once the job ends (#368).""" + from dw.workflow import pipeline_cache_key + + old_def = { + "configuration": {"component_type": "{Mock}"}, + "from_pretrained_arguments": {"model_name": "old-model"}, + "arguments": {}, + } + new_step = { + "name": "gen", + "pipeline": { + "configuration": {"component_type": "{Mock}"}, + "from_pretrained_arguments": {"model_name": "new-model"}, + "arguments": {}, + }, + } + old_key = pipeline_cache_key(old_def) + cache = {old_key: MagicMock()} + workflow = Workflow({"id": "swap", "steps": []}, "/tmp/test_output", "t.json") + workflow._prior_step_keys = {"gen": old_key} + order = [] + + def mock_load(self, shared_components): + order.append("load") + self.pipeline = MagicMock() + + with patch.object(Pipeline, "load", mock_load): + with patch( + "dw.workflow.release_host_caches", + lambda: order.append(("release", old_key in cache)) or 0.0, + ): + workflow.create_step_action(new_step, {}, cache, 1, "cpu") + + assert order == [("release", False), "load"] + + +def test_release_host_caches_runs_for_real_on_the_release_path(): + """The unmocked call is harmless where there is no CUDA or glibc.""" + from dw.host_memory import release_host_caches + + workflow = Workflow(_release_workflow_def(), "/tmp/test_output", "test.json") + + def mock_pipeline_load(self, shared_components): + self.pipeline = MagicMock() + + with patch.object(Pipeline, "load", mock_pipeline_load): + with patch.object( + Step, "run", lambda self, *args, **kwargs: MagicMock(result_list=[]) + ): + with patch( + "dw.workflow.release_host_caches", + wraps=release_host_caches, + ) as real: + workflow.run({}, previous_pipelines={}) + + assert real.call_count == 1 From ce61b7b89c41691cb4cfe18ad43d335f968542a6 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 13:20:22 -0500 Subject: [PATCH 058/181] fix(plan): #382 - downloads_required checks completeness, not just presence scan_models lists a repo whose pull was interrupted the same as a fully cached one, so a plan answering downloads_required: [] let run_workflow silently pull several GB after the caller acknowledged zero downloads. repo_download_incomplete (dw/hub_cache.py) checks the local cache only, never the hub: any *.incomplete blob under the repo's blobs/ marks it incomplete outright; past that, a diffusers pipeline repo (one with a model_index.json in the snapshot refs/main points at) is checked component by component against the requested variant, and a plain checkpoint/LoRA repo has nothing further to check. downloads_required now calls it for every repo scan_models does list, and _collect_sources threads each repo's from_pretrained_arguments.variant through so the check knows which tagged files to look for. Co-Authored-By: Claude Sonnet 5 --- dw/hub_cache.py | 99 +++++++++++++++++++++++++++++++++++++ dw/plan.py | 37 +++++++++----- tests/test_hub_cache.py | 107 +++++++++++++++++++++++++++++++++++++++- tests/test_plan.py | 33 +++++++++++++ 4 files changed, 264 insertions(+), 12 deletions(-) diff --git a/dw/hub_cache.py b/dw/hub_cache.py index 655bf23d..348de781 100644 --- a/dw/hub_cache.py +++ b/dw/hub_cache.py @@ -8,7 +8,9 @@ revisions it found inside the cache directory. """ +import json import logging +import os import shutil import threading import time @@ -88,6 +90,103 @@ def scan_models(cache_dir=None): } +_WEIGHT_SUFFIXES = (".safetensors", ".bin", ".msgpack", ".onnx", ".pt") + + +def _repo_folder_name(repo_id, repo_type="model"): + return f"{repo_type}s--{repo_id.replace('/', '--')}" + + +def _snapshot_dir(repo_dir): + """The snapshot directory `refs/main` currently points at, or None - + a repo dir with no readable ref or no matching snapshot folder is + itself a sign the pull never finished.""" + try: + with open(os.path.join(repo_dir, "refs", "main")) as f: + commit = f.read().strip() + except OSError: + return None + snapshot = os.path.join(repo_dir, "snapshots", commit) + return snapshot if os.path.isdir(snapshot) else None + + +def _component_incomplete(folder, variant): + """Whether a diffusers pipeline component's folder is missing files a + load would need. A component with no weight files at all (a scheduler + or tokenizer config) is complete once its folder exists; one with + weight files is checked against the requested `variant` only when that + variant is actually in play, since an unvarianted component (most + schedulers, safety checkers) never carries a tagged file.""" + if not os.path.isdir(folder): + return True + try: + entries = os.listdir(folder) + except OSError: + return True + if not entries: + return True + if not variant: + return False + weight_files = [e for e in entries if e.endswith(_WEIGHT_SUFFIXES)] + if not weight_files: + return False + return not any(f".{variant}." in name for name in weight_files) + + +def repo_download_incomplete(repo_id, cache_dir=None, variant=None): + """Whether repo_id, though listed by `scan_models`, is left over from an + interrupted pull rather than fully fetched (#382): a cancelled + `snapshot_download` leaves the revision folder in place - the repo still + shows up in `scan_cache_dir` - but with an in-progress blob or with whole + components never fetched. + + Two checks, in order of how cheap they are to rule out: any + `*.incomplete` blob (huggingface_hub's own marker for a file mid-transfer) + anywhere in the repo's `blobs/` makes the repo incomplete outright. Past + that, a diffusers pipeline repo (one with a `model_index.json` in its + current snapshot) is checked component by component, for the requested + `variant`; a plain checkpoint/LoRA repo (no `model_index.json`) has + nothing further to check beyond its blobs. This is not a full hub + file-list diff - it does not fetch the repo's file list from the hub, so + a file that was never attempted at all on an otherwise-untouched repo + (nothing downloaded, no blobs, no snapshot) is caught by the "no snapshot" + case, not enumerated. + """ + resolved = _resolved_cache_dir(cache_dir) + repo_dir = os.path.join(resolved, _repo_folder_name(repo_id)) + if not os.path.isdir(repo_dir): + return True + + blobs_dir = os.path.join(repo_dir, "blobs") + if os.path.isdir(blobs_dir): + try: + if any(name.endswith(".incomplete") for name in os.listdir(blobs_dir)): + return True + except OSError: + return True + + snapshot = _snapshot_dir(repo_dir) + if snapshot is None: + return True + + index_path = os.path.join(snapshot, "model_index.json") + if not os.path.isfile(index_path): + # Not a diffusers pipeline repo - blob completeness is all there is + return False + try: + with open(index_path) as f: + index = json.load(f) + except (OSError, ValueError): + return False + + for key, value in index.items(): + if key.startswith("_") or not isinstance(value, list): + continue + if _component_incomplete(os.path.join(snapshot, key), variant): + return True + return False + + def delete_model(repo_id, cache_dir=None): """Delete every cached revision of repo_id. Returns the bytes freed. diff --git a/dw/plan.py b/dw/plan.py index b2096e54..c5e3117a 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -23,7 +23,7 @@ from huggingface_hub.utils import GatedRepoError, HFValidationError, validate_repo_id from .elision import elide_definition -from .hub_cache import scan_models +from .hub_cache import repo_download_incomplete, scan_models from .realize import ( BUILTIN_PREFIX, VARIABLE_PREFIX, @@ -741,25 +741,30 @@ def _repriced(minutes, list_entries, measured_entries): # An adapter names its repo directly, not through from_pretrained_arguments LORAS_KEY = "loras" SINGLE_FILE_KEY = "from_single_file" +VARIANT_KEY = "variant" def downloads_required(expanded, base_dir, workflow_dir, cache_dir, lookup_sizes): """The hub repos and checkpoint URLs the run would fetch before its first step: every `model_name` in the expanded definition (and in each - composed child) that `scan_models` does not find, plus every - `from_single_file` that is a URL. Sizes come from the hub when asked - and are None whenever it does not answer - an offline box is a state, - not an error, so nothing here raises or logs above debug. + composed child) that `scan_models` does not find, plus every one it does + find but whose cached copy is left over from an interrupted pull + (#382) - a component the load needs is missing, or a blob is still + `.incomplete` - plus every `from_single_file` that is a URL. Sizes come + from the hub when asked and are None whenever it does not answer - an + offline box is a state, not an error, so nothing here raises or logs + above debug. """ names = [] + variants = {} urls = [] - _collect_sources(expanded, names, urls) + _collect_sources(expanded, names, variants, urls) for path, _arguments in _sub_workflow_paths(expanded): raw = read_sub_workflow(path, base_dir, workflow_dir) if raw is None: continue try: - _collect_sources(json.loads(raw), names, urls) + _collect_sources(json.loads(raw), names, variants, urls) except ValueError: continue present = {repo.get("repo_id") for repo in scan_models(cache_dir).get("repos", [])} @@ -768,7 +773,11 @@ def downloads_required(expanded, base_dir, workflow_dir, cache_dir, lookup_sizes # A name not shaped like a hub id is a local checkout - decided by # shape, never by touching the disk: the name came from the request # body, and a free pre-flight must not be a directory-existence oracle - if name in present or not _is_repo_id(name): + if not _is_repo_id(name): + continue + if name in present and not repo_download_incomplete( + name, cache_dir, variant=variants.get(name) + ): continue entry = {"repo": name, "gb": None, "gated": None, "access_blocked": None} if lookup_sizes: @@ -787,8 +796,11 @@ def downloads_required(expanded, base_dir, workflow_dir, cache_dir, lookup_sizes return required -def _collect_sources(tree, names, urls): +def _collect_sources(tree, names, variants, urls): """Every from_pretrained source in a tree, first-seen order, deduplicated. + `variants` collects each name's requested `variant` (first-seen), used to + tell a repo that is merely missing its fp16 files from one that is fully + cached without them. A `loras` entry counts too. It carries its repo under `model_name` directly rather than inside a `from_pretrained_arguments` block, so the @@ -807,14 +819,17 @@ def _collect_sources(tree, names, urls): name = source.get(MODEL_NAME_KEY) if isinstance(name, str) and name not in names: names.append(name) + variant = source.get(VARIANT_KEY) + if isinstance(variant, str): + variants[name] = variant single = source.get(SINGLE_FILE_KEY) if isinstance(single, str) and _is_url(single) and single not in urls: urls.append(single) for value in tree.values(): - _collect_sources(value, names, urls) + _collect_sources(value, names, variants, urls) elif isinstance(tree, list): for value in tree: - _collect_sources(value, names, urls) + _collect_sources(value, names, variants, urls) def gate_warnings(downloads_required): diff --git a/tests/test_hub_cache.py b/tests/test_hub_cache.py index 7de8f25c..b6d1b1f2 100644 --- a/tests/test_hub_cache.py +++ b/tests/test_hub_cache.py @@ -1,9 +1,11 @@ """The hub cache manager: scanning what from_pretrained left on disk and deleting it through huggingface_hub's own strategy.""" +import json + import pytest -from dw.hub_cache import scan_models, delete_model +from dw.hub_cache import scan_models, delete_model, repo_download_incomplete def make_repo(cache_dir, name="tiny", commit="aaaa1111", size=64): @@ -49,6 +51,109 @@ def test_delete_of_an_unknown_repo_is_refused_and_deletes_nothing(self, tmp_path assert len(scan_models(tmp_path)["repos"]) == 1 +def make_pipeline_repo(cache_dir, name="pipe", commit="aaaa1111", components=None, variant=None): + """A repo shaped like a diffusers pipeline: model_index.json at the + snapshot root plus one folder per component. `components` maps a + component name to its file list (weight-suffixed names get the + `.{variant}.` tag when `variant` is given); a component omitted from + `components` gets no folder at all, as an interrupted pull would leave it. + """ + repo = cache_dir / f"models--acme--{name}" + snapshot = repo / "snapshots" / commit + snapshot.mkdir(parents=True) + (repo / "blobs").mkdir() + (repo / "refs").mkdir() + (repo / "refs" / "main").write_text(commit) + index = {"_class_name": "SomePipeline"} + for component in components or {}: + index[component] = ["diffusers", "SomeComponent"] + (snapshot / "model_index.json").write_text(json.dumps(index)) + for component, files in (components or {}).items(): + folder = snapshot / component + folder.mkdir() + for file_name in files: + tagged = file_name + if variant and file_name.endswith((".safetensors", ".bin")): + base, _, ext = file_name.rpartition(".") + tagged = f"{base}.{variant}.{ext}" + (folder / tagged).write_bytes(b"x") + return repo + + +class TestRepoDownloadIncomplete: + def test_a_repo_with_no_cache_trace_at_all_is_incomplete(self, tmp_path): + assert repo_download_incomplete("acme/never-pulled", cache_dir=tmp_path) is True + + def test_an_incomplete_blob_marks_the_repo_incomplete(self, tmp_path): + repo = make_repo(tmp_path) + (repo / "blobs").mkdir() + (repo / "blobs" / "deadbeef.incomplete").write_bytes(b"x") + assert repo_download_incomplete("acme/tiny", cache_dir=tmp_path) is True + + def test_a_plain_checkpoint_with_no_model_index_is_complete_once_present(self, tmp_path): + # make_repo has no model_index.json - a bare checkpoint/LoRA repo + make_repo(tmp_path) + assert repo_download_incomplete("acme/tiny", cache_dir=tmp_path) is False + + def test_a_repo_whose_ref_points_at_no_snapshot_is_incomplete(self, tmp_path): + repo = make_repo(tmp_path) + (repo / "refs" / "main").write_text("some-other-commit-never-fetched") + assert repo_download_incomplete("acme/tiny", cache_dir=tmp_path) is True + + def test_a_pipeline_with_every_component_present_is_complete(self, tmp_path): + make_pipeline_repo( + tmp_path, + components={ + "unet": ["diffusion_pytorch_model.safetensors"], + "scheduler": ["scheduler_config.json"], + }, + ) + assert repo_download_incomplete("acme/pipe", cache_dir=tmp_path) is False + + def test_a_pipeline_missing_a_component_folder_entirely_is_incomplete(self, tmp_path): + # model_index.json lists "vae" but the folder was never fetched + repo = make_pipeline_repo( + tmp_path, + components={"unet": ["diffusion_pytorch_model.safetensors"]}, + ) + snapshot = repo / "snapshots" / "aaaa1111" + index = json.loads((snapshot / "model_index.json").read_text()) + index["vae"] = ["diffusers", "AutoencoderKL"] + (snapshot / "model_index.json").write_text(json.dumps(index)) + assert repo_download_incomplete("acme/pipe", cache_dir=tmp_path) is True + + def test_a_component_with_only_config_files_is_complete_regardless_of_variant(self, tmp_path): + make_pipeline_repo( + tmp_path, + components={"scheduler": ["scheduler_config.json"]}, + variant="fp16", + ) + assert ( + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is False + ) + + def test_a_component_missing_the_requested_variant_is_incomplete(self, tmp_path): + # weights were fetched in the default dtype, not the fp16 variant the + # workflow's from_pretrained_arguments asks for + make_pipeline_repo( + tmp_path, + components={"unet": ["diffusion_pytorch_model.safetensors"]}, + ) + assert ( + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is True + ) + + def test_a_component_with_the_requested_variant_present_is_complete(self, tmp_path): + make_pipeline_repo( + tmp_path, + components={"unet": ["diffusion_pytorch_model.safetensors"]}, + variant="fp16", + ) + assert ( + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is False + ) + + class FakeSibling: def __init__(self, size): self.size = size diff --git a/tests/test_plan.py b/tests/test_plan.py index ebe3f7c3..95a7642b 100644 --- a/tests/test_plan.py +++ b/tests/test_plan.py @@ -875,8 +875,40 @@ def test_a_cached_repo_is_not(self, plan, monkeypatch): "scan_models", lambda cache_dir=None: {"repos": [{"repo_id": "org/still-model"}]}, ) + monkeypatch.setattr( + dw.plan, "repo_download_incomplete", lambda *a, **k: False + ) assert plan()["downloads_required"] == [] + def test_a_present_but_incomplete_repo_is_still_required(self, plan, monkeypatch): + """#382: a repo scan_models lists (an interrupted pull left the + revision folder in place) but whose snapshot is missing files the + load needs stays in downloads_required rather than reading as + cached.""" + import dw.plan + + monkeypatch.setattr( + dw.plan, + "scan_models", + lambda cache_dir=None: {"repos": [{"repo_id": "org/still-model"}]}, + ) + seen = [] + + def fake_incomplete(name, cache_dir, variant=None): + seen.append((name, variant)) + return True + + monkeypatch.setattr(dw.plan, "repo_download_incomplete", fake_incomplete) + assert plan()["downloads_required"] == [ + { + "repo": "org/still-model", + "gb": None, + "gated": None, + "access_blocked": None, + } + ] + assert seen == [("org/still-model", None)] + def test_cache_dir_reaches_scan_models(self, plan, monkeypatch): import dw.plan @@ -1297,6 +1329,7 @@ def repos(self, present, monkeypatch): "scan_models", lambda cache_dir=None: {"repos": [{"repo_id": name} for name in present]}, ) + monkeypatch.setattr(dw.plan, "repo_download_incomplete", lambda *a, **k: False) return [ entry["repo"] for entry in dw.plan.downloads_required( From 803026304d6ffb983139e17be556e3856695700d Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 13:21:32 -0500 Subject: [PATCH 059/181] fix(plan): #382 - apply ruff format to new tests Co-Authored-By: Claude Sonnet 5 --- tests/test_hub_cache.py | 25 ++++++++++++++++++------- tests/test_plan.py | 4 +--- 2 files changed, 19 insertions(+), 10 deletions(-) diff --git a/tests/test_hub_cache.py b/tests/test_hub_cache.py index b6d1b1f2..b7bc5643 100644 --- a/tests/test_hub_cache.py +++ b/tests/test_hub_cache.py @@ -51,7 +51,9 @@ def test_delete_of_an_unknown_repo_is_refused_and_deletes_nothing(self, tmp_path assert len(scan_models(tmp_path)["repos"]) == 1 -def make_pipeline_repo(cache_dir, name="pipe", commit="aaaa1111", components=None, variant=None): +def make_pipeline_repo( + cache_dir, name="pipe", commit="aaaa1111", components=None, variant=None +): """A repo shaped like a diffusers pipeline: model_index.json at the snapshot root plus one folder per component. `components` maps a component name to its file list (weight-suffixed names get the @@ -90,7 +92,9 @@ def test_an_incomplete_blob_marks_the_repo_incomplete(self, tmp_path): (repo / "blobs" / "deadbeef.incomplete").write_bytes(b"x") assert repo_download_incomplete("acme/tiny", cache_dir=tmp_path) is True - def test_a_plain_checkpoint_with_no_model_index_is_complete_once_present(self, tmp_path): + def test_a_plain_checkpoint_with_no_model_index_is_complete_once_present( + self, tmp_path + ): # make_repo has no model_index.json - a bare checkpoint/LoRA repo make_repo(tmp_path) assert repo_download_incomplete("acme/tiny", cache_dir=tmp_path) is False @@ -110,7 +114,9 @@ def test_a_pipeline_with_every_component_present_is_complete(self, tmp_path): ) assert repo_download_incomplete("acme/pipe", cache_dir=tmp_path) is False - def test_a_pipeline_missing_a_component_folder_entirely_is_incomplete(self, tmp_path): + def test_a_pipeline_missing_a_component_folder_entirely_is_incomplete( + self, tmp_path + ): # model_index.json lists "vae" but the folder was never fetched repo = make_pipeline_repo( tmp_path, @@ -122,14 +128,17 @@ def test_a_pipeline_missing_a_component_folder_entirely_is_incomplete(self, tmp_ (snapshot / "model_index.json").write_text(json.dumps(index)) assert repo_download_incomplete("acme/pipe", cache_dir=tmp_path) is True - def test_a_component_with_only_config_files_is_complete_regardless_of_variant(self, tmp_path): + def test_a_component_with_only_config_files_is_complete_regardless_of_variant( + self, tmp_path + ): make_pipeline_repo( tmp_path, components={"scheduler": ["scheduler_config.json"]}, variant="fp16", ) assert ( - repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is False + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") + is False ) def test_a_component_missing_the_requested_variant_is_incomplete(self, tmp_path): @@ -140,7 +149,8 @@ def test_a_component_missing_the_requested_variant_is_incomplete(self, tmp_path) components={"unet": ["diffusion_pytorch_model.safetensors"]}, ) assert ( - repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is True + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") + is True ) def test_a_component_with_the_requested_variant_present_is_complete(self, tmp_path): @@ -150,7 +160,8 @@ def test_a_component_with_the_requested_variant_present_is_complete(self, tmp_pa variant="fp16", ) assert ( - repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") is False + repo_download_incomplete("acme/pipe", cache_dir=tmp_path, variant="fp16") + is False ) diff --git a/tests/test_plan.py b/tests/test_plan.py index 95a7642b..b6a3b55b 100644 --- a/tests/test_plan.py +++ b/tests/test_plan.py @@ -875,9 +875,7 @@ def test_a_cached_repo_is_not(self, plan, monkeypatch): "scan_models", lambda cache_dir=None: {"repos": [{"repo_id": "org/still-model"}]}, ) - monkeypatch.setattr( - dw.plan, "repo_download_incomplete", lambda *a, **k: False - ) + monkeypatch.setattr(dw.plan, "repo_download_incomplete", lambda *a, **k: False) assert plan()["downloads_required"] == [] def test_a_present_but_incomplete_repo_is_still_required(self, plan, monkeypatch): From 11c436925edae32dd655d6d5eeedc57e39a06e27 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 13:31:18 -0500 Subject: [PATCH 060/181] fix(tasks): #383 - grade's get_task docs describe the command, not the per-frame function The summary and media parameter text get_task reported for grade came from grade_image's docstring, which describes the single PIL-frame function _handle_grade dispatches to - not the command a caller invokes, which also accepts a video path or asset:/output: reference. register_command gains optional summary/parameter_descriptions overrides (the same fix #366 already applied to get_first_frame/ get_last_frame via _VIDEO_PROCESSOR_INFO) and grade uses them. Also makes the non_negative/positive/non_positive domain refusal wording in task_domains.py domain-neutral - it previously talked about a "track of the wrong length or speed" even when refusing grade's contrast/saturation, which are multipliers, not audio. Co-Authored-By: Claude Sonnet 5 --- dw/introspection.py | 6 ++++++ dw/task_domains.py | 10 +++++++--- dw/tasks/task.py | 36 ++++++++++++++++++++++++++++++++++-- 3 files changed, 47 insertions(+), 5 deletions(-) diff --git a/dw/introspection.py b/dw/introspection.py index e32b3634..8a0bdd87 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -479,6 +479,12 @@ def describe_task(command): if domain is not None: parameter["domain"] = domain + parameter_descriptions = info.get("parameter_descriptions") or {} + for parameter in parameters: + description = parameter_descriptions.get(parameter["name"]) + if description: + parameter["description"] = description + summary = info.get("summary") if not summary: summary = _first_paragraph(inspect.getdoc(implementation)) diff --git a/dw/task_domains.py b/dw/task_domains.py index ecf48d90..94c1b130 100644 --- a/dw/task_domains.py +++ b/dw/task_domains.py @@ -64,10 +64,14 @@ "the documented scale is only defined inside it" ), } +# Shared by POSITIVE, NON_NEGATIVE and NON_POSITIVE, which span both counts/ +# rates (audio, frame) and multipliers (grade's contrast, saturation) - kept +# neutral rather than naming either, since a wording specific to one reads as +# nonsense on the other (#383) _DEFAULT_REASON = ( - "A value outside that range is refused rather than interpreted - " - "a negative count or a zero rate would otherwise produce a " - "plausible-looking track of the wrong length or speed" + "A value outside that range is refused rather than interpreted - it " + "would otherwise produce a plausible-looking result outside the " + "documented range" ) # command -> argument -> domain. Every entry here is pinned to a real command diff --git a/dw/tasks/task.py b/dw/tasks/task.py index 5c190fc7..e566107c 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -38,6 +38,8 @@ def register_command( provided=(), consumes_device=False, returns="artifact", + summary=None, + parameter_descriptions=None, ): """ Decorator to register a command handler function. @@ -64,6 +66,16 @@ def register_command( `validation_errors` (dw/scalar_result_validation.py, #212) rather than reaching `save_artifact` at run time, where a float has nothing left identifying which command produced it + summary: Overrides the command's `get_task` summary, which otherwise + reads the implementation function's docstring. For a command + whose handler dispatches its implementation per video frame + (`_per_frame`), that docstring describes the single-frame + function rather than the command a caller invokes - same reason + `_VIDEO_PROCESSOR_INFO` overrides `get_first_frame`/ + `get_last_frame` (#366, #383) + parameter_descriptions: Overrides one or more of the implementation's + per-parameter `get_task` descriptions by name, for the same + single-frame-vs-command reason as `summary` Returns: Decorator function @@ -80,12 +92,17 @@ def handler(task, arguments, previous_pipelines, _func=func): return _func(task, arguments, previous_pipelines) _COMMAND_REGISTRY[command_name] = handler - _COMMAND_INFO[command_name] = { + info = { "kind": "command", "implementation": implementation, "provided": tuple(provided), "returns": returns, } + if summary: + info["summary"] = summary + if parameter_descriptions: + info["parameter_descriptions"] = dict(parameter_descriptions) + _COMMAND_INFO[command_name] = info logger.debug(f"Registered command handler: {command_name}") return func @@ -426,7 +443,22 @@ def _handle_restore_faces(task, arguments, previous_pipelines): ) -@register_command("grade", implementation="dw.tasks.grade.grade_image") +@register_command( + "grade", + implementation="dw.tasks.grade.grade_image", + summary=( + "Adjust exposure, contrast, saturation and white balance of an " + "image or a video." + ), + parameter_descriptions={ + "media": ( + "Image or video to grade. An image is a PIL Image; a video is a " + "file path or an asset:/output: reference, read with its audio " + "and graded frame by frame, keeping its frame rate and audio " + "unchanged." + ), + }, +) def _handle_grade(task, arguments, previous_pipelines): """Adjust exposure, contrast, saturation and white balance of an image or a video""" logger.debug("Grading media") From 416c6ba5692c5f8114f11ebf0c4cdc50bb7a576a Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 18:31:34 -0500 Subject: [PATCH 061/181] fix(mcp,plugin): correct stale and contradictory agent-facing text A prompt audit of the MCP tool descriptions and plugin skills. - export_job told the agent to fetch the zip URL, but never mentioned auth_required/open_url; a token-gated zip can't be fetched by the agent - script-to-video put H3 shots at 4-6 s; the floor is 124 frames, 5.17 s - run_workflow/rerun_job said pass acknowledged_cost=true, then said to bind it to the plan; they now point at the bound form - get_gallery_metadata still described an agent that cannot listen - all-caps cost warnings lowered to sentence case (reasons and the go-ahead instruction kept; the server gate is enforced in code) - "always run this" / "as before" wording; script-to-video's unattended bullet contradicted step 5's pre-authorization allowance Co-Authored-By: Claude Opus 5.5 (1M context) --- dw_mcp/server.py | 42 +++++++++++----------- plugins/dw/skills/script-to-video/SKILL.md | 8 ++--- tests/test_mcp_server.py | 4 +-- 3 files changed, 27 insertions(+), 27 deletions(-) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 447a6f4c..0278c98f 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -408,7 +408,7 @@ def get_gallery_metadata( not a summary, so a result can be reproduced or a failed run's definition edited and re-run. For audio and video the `media` block carries duration, sample rate, channels, fps, size and level - - the checks an agent that cannot listen makes on a deliverable. + - the numbers to check a deliverable by, beside listening to it. `envelope=true` adds that level second by second (`media.envelope.rms_dbfs` / `peak_dbfs`), which says *where* in a track something is: whether a shot still sounds at its last frame, @@ -685,7 +685,7 @@ def download_output( get_output_text for inline content, or the `url` that `list_gallery` reports for each entry, which already carries the workspace selector - do not build an /outputs URL by hand). This is - also NOT how a generated file becomes an input for a later + also not how a generated file becomes an input for a later workflow: use `keep_output`, which links it inside the workspace under an "asset:" name, rather than writing into the server's asset directory behind the API's back. Unlike the inline tools, this works @@ -872,7 +872,7 @@ def validate_workflow( arguments: dict | None = None, ) -> dict: """Check a workflow against the schema and against real pipeline - signatures. Free and instant - always run this before run_workflow. + signatures. Free and instant - run it before run_workflow. Give exactly one of `workflow` or `name` (a stored workflow as `list_workflows` reports it); `run_workflow`'s `inline_workflow` and `workflow_path` spellings are accepted here too. `workflow` may @@ -1020,7 +1020,7 @@ def enhance_prompt( acknowledged_cost: bool = False, ) -> dict: """Expand a short idea into a full prompt with a language model. - THIS COSTS TIME ON THE ENGINE: it queues a real job, and the engine + This costs time on the engine: it queues a real job, and the engine runs one at a time, so a generation waiting behind it is delayed. Tell the user what will be enhanced and get their go-ahead, then pass acknowledged_cost=true. Returns as soon as the job is queued; @@ -1052,10 +1052,10 @@ def run_workflow( workspace: str | None = None, wait_seconds: int = 0, ) -> dict: - """Queue a workflow for generation. THIS COSTS GPU TIME: a run + """Queue a workflow for generation. This costs GPU time: a run occupies the machine for minutes and the engine runs one job at a time. Tell the user what will run and get their go-ahead, then pass - acknowledged_cost=true. Returns as soon as the job is queued; + acknowledged_cost as below. Returns as soon as the job is queued; follow it with `wait_for_job`, then `get_job` for the manifest - or fold that first wait in with `wait_seconds` above 0, which waits on the job exactly as `wait_for_job(job_id, @@ -1063,7 +1063,7 @@ def run_workflow( its fields to the result (`still_running`, `waited_seconds`, `timeout_*`, the slim `job`). If the cap covers the job's runtime one call is enough; on `still_running: true` call - `wait_for_job` as before. Give exactly one of `workflow_path` - a + `wait_for_job` again. Give exactly one of `workflow_path` - a catalog name from `list_workflows`, with or without .json, or a path on the server - or `inline_workflow`, a full definition nothing stored covers; `validate_workflow` calls these `name` and @@ -1194,10 +1194,10 @@ def cancel_job(job_id: str) -> dict: def rerun_job( job_id: str, acknowledged_cost: bool | dict = False, new_seed: bool = False ) -> dict: - """Queue a fresh job from a previous job's stored specification. THIS - COSTS GPU TIME: a rerun is a run - it occupies the machine for + """Queue a fresh job from a previous job's stored specification. This + costs GPU time: a rerun is a run - it occupies the machine for minutes and the engine runs one job at a time. Tell the user what - will run and get their go-ahead, then pass acknowledged_cost=true. + will run and get their go-ahead, then pass acknowledged_cost. Pass new_seed=true for a different image: a workflow that pins its seed reruns to the same pixels, and the step cache serves that whole @@ -1231,11 +1231,13 @@ def export_job(job_id: str, overwrite: bool = False) -> dict: what was copied. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. - THE DIRECTORY IS ON THE MACHINE RUNNING THE SERVER, not on yours. To - give the user the files, fetch the zip URL and unpack it into - exports/ under the session's working directory - it is the user's - deliverable, not a temp file; the archive already unpacks into one - folder named after the job id, so do not create that folder first. + The directory is on the machine running the server, not yours. + `open_url` is the zip: with `auth_required` false, fetch it and + unpack it into exports/ under the session's working directory - the + user's deliverable, not a temp file; it unpacks into a folder named + after the job id, so do not create that folder first. With it true + the zip needs a token you cannot attach, so hand `open_url` to the + person. Refuses a job that is still running; refuses an existing export unless overwrite=true.""" return exports.export_job(client, job_id, overwrite=overwrite) @@ -1250,8 +1252,8 @@ def export_job(job_id: str, overwrite: bool = False) -> dict: # -------------------------------------------------------------- models def download_model(repo_id: str, acknowledged_cost: bool = False) -> dict: - """Fetch a model repo into the Hugging Face cache. THIS COSTS DISK - AND BANDWIDTH: a model repo is commonly tens of gigabytes. Check + """Fetch a model repo into the Hugging Face cache. This costs disk + and bandwidth: a model repo is commonly tens of gigabytes. Check list_models first - it may already be cached. Tell the user what you are about to fetch and get their go-ahead, then pass acknowledged_cost=true. Returns as soon as the download starts; poll @@ -1270,8 +1272,8 @@ def cancel_download(download_id: str) -> dict: return models.cancel_download(client, download_id) def delete_model(repo: str, acknowledged_cost: bool = False) -> dict: - """Delete every cached revision of one model repo. THIS IS NOT - RECOVERABLE: getting the model back means downloading it again. Tell + """Delete every cached revision of one model repo. This is not + recoverable: getting the model back means downloading it again. Tell the user which repo and how much it frees, get their go-ahead, then pass acknowledged_cost=true. Refused while a job or download is active.""" @@ -1282,7 +1284,7 @@ def get_diffusers_state() -> dict: return models.get_diffusers_state(client) def update_diffusers(acknowledged_cost: bool = False) -> dict: - """Upgrade diffusers to GitHub HEAD. THIS CAN BREAK THE INSTALL: it + """Upgrade diffusers to GitHub HEAD. This can break the install: it installs an untagged development build that workflows running today may not survive, and this tool cannot undo it. Report the current version, explain why the update is worth it, get the user's diff --git a/plugins/dw/skills/script-to-video/SKILL.md b/plugins/dw/skills/script-to-video/SKILL.md index 0a28e1f5..0ae59fa8 100644 --- a/plugins/dw/skills/script-to-video/SKILL.md +++ b/plugins/dw/skills/script-to-video/SKILL.md @@ -22,9 +22,9 @@ does not parse screenplay formats. ## 2. Decompose into a shot list, not a shot count The unit that matters is *shots*, not lines or scenes: one scene may be -several shots, and `dw:minimax-h3`'s sweet spot is 4-6 second shots (the -`17*n+5` frame grid). For each shot, name it, note which character(s) -appear, what happens, and roughly how long it runs. +several shots, and a `dw:minimax-h3` shot runs 5.17 to 14.4 seconds (the +`17*n+5` frame grid, 124 to 345 frames). For each shot, name it, note +which character(s) appear, what happens, and roughly how long it runs. ## 3. Cast recurring characters once, before any shot generates @@ -92,8 +92,6 @@ already owns this step - do not re-derive it here. catalog templates; a script whose shape no template covers is a "propose a new template" conversation, not something to improvise at run time. -- **Full unattended autonomy.** Step 5's cost acknowledgment is a real gate - every time, not a one-time setup step. ## Sources diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 3a367d8e..e11191f0 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -258,7 +258,7 @@ async def test_run_workflow_advertises_its_cost(): tools = await tools_of(server_over(ok({}))) description = tools["run_workflow"].description - assert "COSTS GPU TIME" in description + assert "costs gpu time" in description.lower() assert "acknowledged_cost" in description @@ -915,7 +915,7 @@ async def test_rerun_job_advertises_its_cost(): tools = await tools_of(server_over(ok({}))) description = tools["rerun_job"].description - assert "COSTS GPU TIME" in description + assert "costs gpu time" in description.lower() assert "acknowledged_cost" in description From 4331715d99da2038081a724d01c19f08d7672baa Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 18:56:32 -0500 Subject: [PATCH 062/181] docs(proposals): #375 - decline workspace-folders, move to declined/ with the reason Co-Authored-By: Claude Opus 5.5 (1M context) --- .../{ => declined}/workspace-folders.md | 24 +++++++++++++++---- docs/proposals/todo.md | 14 +++++++---- 2 files changed, 30 insertions(+), 8 deletions(-) rename docs/proposals/{ => declined}/workspace-folders.md (90%) diff --git a/docs/proposals/workspace-folders.md b/docs/proposals/declined/workspace-folders.md similarity index 90% rename from docs/proposals/workspace-folders.md rename to docs/proposals/declined/workspace-folders.md index 36f6723f..cfb9dc6b 100644 --- a/docs/proposals/workspace-folders.md +++ b/docs/proposals/declined/workspace-folders.md @@ -1,9 +1,25 @@ # Design: hierarchical workspace names -Status: **proposal, not implemented**. Written for issue #167. Researched -2026-09-16 (feasibility comment on the issue, model claude-sonnet-5 via -provider anthropic); this document is the triaged follow-up (model -claude-sonnet-5, provider anthropic). +Status: **declined 2026-09-23** (issue #375, closed not planned). Written for +issue #167 and researched 2026-09-16 by claude-sonnet-5 via anthropic. + +Why it was declined: plan v2 on #375 weighed value against build cost, +inertia and reversibility. `lem` held 20 workspaces, 14 of them unrelated +projects, and the `QA-EP*` clutter behind the ask was down to one. Against +that, the change loosens `WORKSPACE_NAME_PATTERN`, a deliberate security +boundary, and it cannot be backed out without a breaking change once +grouped names reach job records. Don: "don't build". + +The design below still has two errors, corrected in #375's plan. Read that +plan before reviving this: +- `DELETE /api/workspaces/{name}` needs `{name:path}`. +- Reserved names must be checked per segment: whole-string, `outputs/x` + passes. + +Revisit if a workflow starts making per-episode workspaces again, or the +list reaches about 40 with real prefix families. If only the look is +wanted, a display-only prefix grouping in the UI sidebar gets most of it +with no engine change. ## The ask, as filed diff --git a/docs/proposals/todo.md b/docs/proposals/todo.md index 6f0a1eb0..43bc94ce 100644 --- a/docs/proposals/todo.md +++ b/docs/proposals/todo.md @@ -10,7 +10,7 @@ original Tier 1 items and both fully-finished proposals (`score-and-select`, `script-to-video-agent-skill`) were removed. Since 2026-09-23 every open item below is also a GitHub issue labeled -`feature` (#374–#380, and #244 for `resume.md`), parked with Don. Its +`feature` (#374 and #376–#380, and #244 for `resume.md`; #375 was declined), parked with Don. Its `priority:N` label mirrors the tier here. Work starts from the issue. ## Tier 1 — do these first (small, scoped, clear payoff) @@ -24,9 +24,6 @@ remaining deferred fix is recorded in 2. **orphaned-run-directories.md** — real, recurring disk-usage annoyance (leftover manifests invisible to gallery/asset listings); the proposal already recommends the simple option (A). Moderate but bounded work. -3. **workspace-folders.md** — one regex relax + a depth-2 listing walk + UI - grouping. Low complexity, meaningful convenience for the growing - series-episodes workflow. 4. **mcp-context-cost-partial.md** — cheap and safe, but the remaining payoff is small (~1.1k tokens of connect cost), and recommendation 6 in `mcp-context-cost-complete.md` suggests the harness-level fix (deferred @@ -57,6 +54,15 @@ remaining deferred fix is recorded in 2026-09-20; the maintenance/observability page itself (orphan listing, disk usage, job pruning) is the remaining, much bigger ask. +## Declined + +Kept in `declined/` with the reason at the top, so a revival starts from +the analysis rather than repeating it. + +- **declined/workspace-folders.md** — grouped workspace names (`QA/EP1`). + Declined 2026-09-23 on #375: thin demonstrated value against a loosened + security boundary and a change that can't be taken back. + ## Backlog ideas with no doc on file - **A larger Qwen-Image catalog entry** — the 20B, Apache-2.0 Qwen-Image From c0e3e4ef5be6020659f359e8da301da9259c3377 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 19:04:35 -0500 Subject: [PATCH 063/181] docs(readme): ask driving sessions to file feedback with the field-report label Co-Authored-By: Claude Opus 5.5 (1M context) --- README.md | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/README.md b/README.md index ffcb9491..89ef6c6c 100644 --- a/README.md +++ b/README.md @@ -122,6 +122,19 @@ and outputs — so two agents, or an agent and you in the browser, share the GPU without saving over each other. An agent calls `use_workspace` once and the rest of the session lands there. +**Feedback from a session.** At the end of a working session, ask the agent +what got in its way: bugs, gaps, misleading skill text, tools it reached for +and couldn't find. Have it file each one as an issue on +`dkackman/diffusers-workflow` with the `field-report` label, for example: + +> File each bug or gap you hit as an issue on dkackman/diffusers-workflow +> with the label `field-report`. + +The label marks a report as coming from real use, not from the automated test +loop. The agent loop (see [Agent Loop](docs/AGENT_LOOP.md)) picks the report up +like any other issue. The label is also what feature planning reads as +evidence of demand. + The complete tool reference, client configuration for other MCP hosts, and the troubleshooting table: [MCP Server](docs/MCP.md). Workspaces in depth: [Workspaces](docs/WORKSPACES.md). From 525e4701115a538609f0a26208fa3d7a5b89b6ba Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 20:44:08 -0500 Subject: [PATCH 064/181] fix(mcp): #384 - get_gallery_metadata points at the actual recovery route metadata is null for every audio/video file by design (only PNG/JPEG/WebP embed it), but next left that null unexplained instead of saying where the recipe actually lives. next now names get_job_workflow(job_id) when the job is known, and says a kept asset carries no provenance at all when it isn't - matching the triage scope on #384. Docstrings updated in dw_mcp/server.py, the REST route in dw/server/app.py, and docs/MCP.md; two new catalog tests exercise both branches through the real handler. Co-Authored-By: Claude Sonnet 5 --- docs/MCP.md | 2 +- dw/server/app.py | 6 +++++- dw_mcp/catalog.py | 40 +++++++++++++++++++++++++++++++------ dw_mcp/server.py | 19 +++++++++--------- tests/test_mcp_catalog.py | 42 +++++++++++++++++++++++++++++++++++++++ 5 files changed, 92 insertions(+), 17 deletions(-) diff --git a/docs/MCP.md b/docs/MCP.md index 9346c21b..d0067bf3 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -227,7 +227,7 @@ when no single workflow covers it. | `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in | | `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds | | `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None, folder=None, version=None)` | `limit`, `subfolder`, `only_orphans`, `workspace`, `folder`, `version` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. `folder` with `version` lists that one run's files, and `output:/v5/` names one in a workflow; every other tool takes `name`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | -| `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). Only an image (PNG/JPEG/WebP) embeds `metadata` this way — it is always null for audio and video, and `next` then points at `get_job_workflow(job_id)` when `job` is known, or says a kept asset carries no provenance at all when it isn't. `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | ### Media diff --git a/dw/server/app.py b/dw/server/app.py index 3a3f1591..256eefd0 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -2982,7 +2982,11 @@ def gallery_metadata( it is the full definition the editor can reopen), plus the job that produced the file when history remembers one, plus - for audio and video - what the file itself holds: duration, format and level, - which is how an agent that cannot listen checks a track. + which is how an agent that cannot listen checks a track. Only an + image embeds 'metadata' this way - it is always null for audio and + video, since neither format has a slot this writer uses; recover + the recipe from 'job' (GET /api/jobs/{id}/workflow) when one is + known, or from nothing when it isn't (a kept asset has no job). `envelope=true` adds the soundtrack's level second by second, which is what says *where* in a track something is - whether a shot is diff --git a/dw_mcp/catalog.py b/dw_mcp/catalog.py index 49b3a23c..8405628b 100644 --- a/dw_mcp/catalog.py +++ b/dw_mcp/catalog.py @@ -270,10 +270,20 @@ def list_gallery( def get_gallery_metadata(client, name, envelope=False, workspace=None): - """Metadata embedded in a saved file: the full workflow that made it, - plus the job that produced it when history remembers one, plus for - audio and video what the file holds - duration, sample rate, channels, - fps, size, peak and mean level in dBFS. + """Metadata embedded in a saved file: the full workflow that made it - + the exact workflow, arguments and seed, so a result can be reproduced + or a failed run's definition edited and re-run - plus the job that + produced it when history remembers one, plus for audio and video what + the file holds - duration, sample rate, channels, fps, size, peak and + mean level in dBFS. + + Only an image (PNG/JPEG/WebP) carries embedded metadata; `metadata` is + always null for audio and video, since neither format has a slot this + writer uses. `job` is the fallback recipe when one is known - `next` + then names `get_job_workflow(job_id)`, which reads the run's realized + workflow instead. A kept asset (`source: "asset"`) has no job at all, + so nothing on the server remembers which run made it; `next` says so + rather than pretending a lookup exists. `name` is a gallery name - the `name` field `list_gallery` reports, not its `label` (a display-only basename that is not a valid reference) - @@ -290,8 +300,24 @@ def get_gallery_metadata(client, name, envelope=False, workspace=None): workspace=workspace, ) media = body.get("media") + job = body.get("job") + hints = [] + if body.get("metadata") is None: + if job: + hints.append( + "metadata is null because only an image (PNG/JPEG/WebP) " + "carries it embedded - this file's job is known, and " + f'get_job_workflow(job_id="{job["id"]}") returns the exact ' + "workflow, arguments and seed that produced it." + ) + elif body.get("source") == "asset": + hints.append( + "metadata is null and this is a kept asset, which carries " + "no provenance - nothing on the server remembers which job, " + "if any, produced the file it was kept from." + ) if media and body.get("source") == "asset": - body["next"] = ( + hints.append( "These are the numbers a workflow's arguments have to match " "before the run, not after: frame_count and fps decide a cut's " "'total_frames', sample_rate decides what its audio is mixed " @@ -301,7 +327,7 @@ def get_gallery_metadata(client, name, envelope=False, workspace=None): "make a longer bed with the 'loop_audio' task instead." ) elif media and media.get("kind") in ("audio", "video"): - body["next"] = ( + hints.append( "Check duration_seconds against what was asked for: a Music 3 " "track that lands within 0.2 s of its audio_duration ceiling was " "cut off, one well short of it finished naturally. peak_dbfs is " @@ -314,4 +340,6 @@ def get_gallery_metadata(client, name, envelope=False, workspace=None): "'normalize_audio' (peak_dbfs: -3) before the saving step is what " "fixes it." ) + if hints: + body["next"] = " ".join(hints) return body diff --git a/dw_mcp/server.py b/dw_mcp/server.py index ffb70253..de538f2d 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -411,21 +411,22 @@ def get_gallery_metadata( """Get the metadata embedded in a generated file: the exact workflow, arguments and seed that produced it - the definition, not a summary, so a result can be reproduced or a failed run's - definition edited and re-run. For audio and video the `media` - block carries duration, sample rate, channels, fps, size and level - - the checks an agent that cannot listen makes on a deliverable. + definition edited and re-run. Only an image embeds it; for audio + and video `metadata` is null and `next` names + `get_job_workflow(job_id)` when known, else a kept asset has no + provenance. `media` itself carries duration, sample rate, + channels, fps, size and level - the checks an agent that cannot + listen makes on a deliverable. `envelope=true` adds that level second by second (`media.envelope.rms_dbfs` / `peak_dbfs`), which says *where* in a - track something is: whether a shot still sounds at its last frame, - how deep the hole at a seam goes, where a score goes quiet. Leave - it off unless the question is about a position - a long track is - a long list. + track something is: a shot's last frame, a seam's hole, where a + score goes quiet. Leave it off unless it's about position - a long + track is a long list. `media.peak_dbfs` is what the job's `audio_no_headroom` (-0.5 dBFS, pre-encode) and `audio_clipped` (0.0 dBFS, post-encode) warnings read - see `normalize_audio` under "Video Processing" in the tasks - guide. A mux emits only the second, so a peak between the two is - clean. + guide. A mux emits only the second. `name` may be an "asset:" reference instead of a gallery name, and then it describes that input asset - how many frames a shot is, diff --git a/tests/test_mcp_catalog.py b/tests/test_mcp_catalog.py index 7ca94199..a09459cb 100644 --- a/tests/test_mcp_catalog.py +++ b/tests/test_mcp_catalog.py @@ -274,3 +274,45 @@ def test_gallery_metadata_reads_an_asset_and_says_the_numbers_are_inputs(): assert result["media"]["duration_seconds"] == 3.3 assert "loop_audio" in result["next"] assert "audio_duration" not in result["next"] + + +def test_gallery_metadata_points_at_get_job_workflow_when_embedded_is_null(): + """#384: a video output's 'metadata' is always null (only an image + embeds it), and the job that made it is known - the hint has to name + the actual recovery route rather than leave the caller with null.""" + body = { + "name": "templates/minimax/dialogue-short/20260923-185419-650e12db/final/x.mp4", + "source": "output", + "metadata": None, + "job": {"id": "c68bc29607ec", "status": "succeeded"}, + "media": {"kind": "video", "duration_seconds": 6.0}, + } + client, _ = scripted( + {("GET", "/api/gallery/" + body["name"] + "/metadata"): (200, body)} + ) + + result = catalog.get_gallery_metadata(client, body["name"]) + + assert "get_job_workflow" in result["next"] + assert "c68bc29607ec" in result["next"] + + +def test_gallery_metadata_says_a_kept_asset_has_no_provenance(): + """#384: a kept asset has 'job: null' by construction - nothing traces + it back to the run that made it, and the hint has to say that rather + than staying silent about the null 'metadata'.""" + body = { + "name": "asset:qa-cast/ep37-shot2-alibi.mp4", + "source": "asset", + "metadata": None, + "job": None, + "media": {"kind": "video", "duration_seconds": 4.2}, + } + client, _ = scripted( + {("GET", "/api/gallery/asset:qa-cast/ep37-shot2-alibi.mp4/metadata"): (200, body)} + ) + + result = catalog.get_gallery_metadata(client, "asset:qa-cast/ep37-shot2-alibi.mp4") + + assert "no provenance" in result["next"] + assert "get_job_workflow" not in result["next"] From b2b8b544b5e9d49137df820326c32686394d9e57 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 20:45:12 -0500 Subject: [PATCH 065/181] style(mcp): #384 - ruff format test_mcp_catalog.py Co-Authored-By: Claude Sonnet 5 --- tests/test_mcp_catalog.py | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/tests/test_mcp_catalog.py b/tests/test_mcp_catalog.py index a09459cb..01c03d64 100644 --- a/tests/test_mcp_catalog.py +++ b/tests/test_mcp_catalog.py @@ -309,7 +309,12 @@ def test_gallery_metadata_says_a_kept_asset_has_no_provenance(): "media": {"kind": "video", "duration_seconds": 4.2}, } client, _ = scripted( - {("GET", "/api/gallery/asset:qa-cast/ep37-shot2-alibi.mp4/metadata"): (200, body)} + { + ("GET", "/api/gallery/asset:qa-cast/ep37-shot2-alibi.mp4/metadata"): ( + 200, + body, + ) + } ) result = catalog.get_gallery_metadata(client, "asset:qa-cast/ep37-shot2-alibi.mp4") From 897dc162a4034a71759d7f2d61d8e0c3ee9b93e9 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 21:04:58 -0500 Subject: [PATCH 066/181] fix(mcp): #389 - scope /api/server directories to the requested workspace download_output over a mounted dw.serve --mcp endpoint confines its write via dw_mcp/media.py's _remote_root, which read directories.workspace off GET /api/server. That route ignored ?workspace= entirely and always answered with the server's own default directories, so a session pinned to a named workspace (or a download_output(workspace=...) override) had its write silently confined to and landed in the default workspace's output tree instead. /api/server now takes the same selected_workspace dependency the other scoped routes use, and _remote_root/download_output forward an explicit per-call workspace override into the /api/server lookup rather than relying only on the client's session-pin fallback. upload_asset's file_path confinement needed no change: it has no per-call workspace argument, so the shared server-side fix already covers it. Also guards ConfiguredWorkspace._prompts against a None prompt_dir, matching the existing _assets guard - needed once /api/server's directories are built from a Workspace instance whose prompts library may not be configured at all. Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 31 ++++++++++++++++----------- dw/workspace.py | 6 +++++- dw_mcp/media.py | 12 ++++++++--- tests/test_mcp_media.py | 36 ++++++++++++++++++++++++++++++++ tests/test_server_info.py | 44 +++++++++++++++++++++++++++++++++++++++ 5 files changed, 113 insertions(+), 16 deletions(-) diff --git a/dw/server/app.py b/dw/server/app.py index 256eefd0..13d925d1 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -4155,7 +4155,7 @@ def health(): } @app.get("/api/server") - def server_info(): + def server_info(ws: Workspace = Depends(selected_workspace)): """How this server is reachable, for the UI's Server page: what it is bound to, whether a token is needed, whether MCP is mounted, and the addresses another machine could name it by. @@ -4164,6 +4164,12 @@ def server_info(): and `mcp.path` - and the token itself is never reported in any form, only whether one is required. An interface enumeration failure is not a server failure: `addresses` comes back empty. + + `directories` is scoped to the `?workspace=` a caller names (or the + session's own pin, via `_scoped`) - a mounted `download_output` + confines a write to *that* workspace's output tree, so reporting + the server's own default here regardless of the selector sent a + caller pinned elsewhere writing into `default` without any error (#389). """ import socket @@ -4196,17 +4202,18 @@ def server_info(): # answers, e.g. whether bitsandbytes is even installed (#222) "runtime": runtime_info(), "directories": { - # The workspace the three below default to folders of; an - # individually overridden folder still reports its own path - "workspace": app.state.workspace, - "workflows": os.path.abspath(app.state.workflow_dir), - "assets": app.state.asset_dir, - "outputs": os.path.abspath(manager.output_dir), - "prompts": ( - os.path.abspath(app.state.prompt_dir) - if app.state.prompt_dir - else None - ), + # ws's properties are already absolute (Workspace and + # ConfiguredWorkspace both resolve at construction). This + # "workspace" is the root path a mounted download_output + # confines a write to (dw_mcp/media.py's _remote_root) - + # None for a default workspace configured from individual + # directory overrides with no --workspace root, same as + # before this route was workspace-aware + "workspace": ws.root, + "workflows": ws.workflows, + "assets": ws.assets, + "outputs": ws.outputs, + "prompts": ws.prompts, }, } diff --git a/dw/workspace.py b/dw/workspace.py index c7972582..dff323fb 100644 --- a/dw/workspace.py +++ b/dw/workspace.py @@ -671,7 +671,11 @@ def __init__(self, workflows, assets, outputs, prompts, root=None): os.path.abspath(os.path.expanduser(str(assets))) if assets else None ) self._outputs = os.path.abspath(os.path.expanduser(str(outputs))) - self._prompts = os.path.abspath(os.path.expanduser(str(prompts))) + # prompts is optional too - a server configured with no prompt + # library at all (app.state.prompt_dir can be None) + self._prompts = ( + os.path.abspath(os.path.expanduser(str(prompts))) if prompts else None + ) @property def is_default(self): diff --git a/dw_mcp/media.py b/dw_mcp/media.py index ca28c75a..fc6b1373 100644 --- a/dw_mcp/media.py +++ b/dw_mcp/media.py @@ -442,18 +442,24 @@ def delete_output(client, name=None, workspace=None, job_id=None): return {**deleted, "job_id": job_id, "run_dir": run_dir} -def _remote_root(client): +def _remote_root(client, workspace=None): """The workspace a remote write is confined to, or None when local. Only the mounted MCP surface is remote: there the tool runs inside dw.serve, so the path a caller names is a path on the operator's box rather than on its own machine. A stdio `dw-mcp` returns None and keeps writing wherever the user can. + + `workspace` is an explicit per-call override (download_output's own + `workspace` argument); when omitted, `client.get_json`'s `_scoped` + already falls back to the session's own pin (#389). """ if not getattr(client, "mounted", False): return None - directories = (client.get_json("/api/server").get("directories")) or {} + directories = ( + client.get_json("/api/server", workspace=workspace).get("directories") + ) or {} root = directories.get("workspace") if not root: raise DwApiError( @@ -524,7 +530,7 @@ def download_output(client, name, destination=None, overwrite=False, workspace=N omitted `destination` keeps defaulting to the current working directory, because there "local disk" is genuinely their own. """ - root = _remote_root(client) + root = _remote_root(client, workspace=workspace) if destination is None: if root: raise DwApiError( diff --git a/tests/test_mcp_media.py b/tests/test_mcp_media.py index 5839bf0a..c5a53705 100644 --- a/tests/test_mcp_media.py +++ b/tests/test_mcp_media.py @@ -897,6 +897,42 @@ def test_a_mounted_server_refuses_an_overwrite_outside_the_workspace(tmp_path): assert victim.read_text() == "mine" +def test_a_mounted_server_confines_a_per_call_workspace_override(tmp_path): + """#389: download_output's own `workspace` argument overrides the + session's pin for the download itself (stream_to_file already forwarded + it), but _remote_root asked /api/server with no workspace at all, so the + confinement root stayed the session's - a relative destination under a + workspace= override landed in the wrong tree with no error.""" + default_ws = tmp_path / "default" + default_ws.mkdir() + other_ws = tmp_path / "other" + other_ws.mkdir() + + def handler(request): + if request.url.path == "/api/server": + requested = httpx.QueryParams(request.url.query.decode()) + root = other_ws if requested.get("workspace") == "other" else default_ws + return httpx.Response( + 200, + json={"directories": {"workspace": str(root)}}, + headers={"content-type": "application/json"}, + ) + return httpx.Response( + 200, content=png_bytes(4, 4), headers={"content-type": "image/png"} + ) + + client = DwClient(transport=httpx.MockTransport(handler)) + client.mounted = True + + result = download_output( + client, "run/probe.jpg", destination="kept/probe.jpg", workspace="other" + ) + + assert result["saved_to"] == str(other_ws / "kept" / "probe.jpg") + assert (other_ws / "kept" / "probe.jpg").read_bytes() == png_bytes(4, 4) + assert not (default_ws / "kept" / "probe.jpg").exists() + + def test_a_mounted_server_writes_a_relative_destination_into_its_workspace(tmp_path): """And the default keeps working: a relative destination is joined onto the workspace rather than onto whatever the server's cwd happens to be.""" diff --git a/tests/test_server_info.py b/tests/test_server_info.py index 7aaf3e93..c878e946 100644 --- a/tests/test_server_info.py +++ b/tests/test_server_info.py @@ -9,6 +9,7 @@ from dw.server.app import create_app from dw.server.jobs import JobManager from dw.server import netinfo +from dw.workspace import Workspace, create_workspace from tests.test_server import ScriptedWorkerManager, success_script @@ -99,6 +100,49 @@ def test_prompt_dir_may_be_absent(tmp_path): assert c.get("/api/server").json()["directories"]["prompts"] is None +def test_directories_are_scoped_to_the_requested_workspace(tmp_path): + """#389: a mounted download_output confines a write against + directories.workspace, so this route has to answer per the ?workspace= + a caller (or the client's session pin) actually names, not the server's + own default - a caller pinned to a named workspace was writing into the + default workspace's tree with no error.""" + root = Workspace(tmp_path / "studio", "flag").ensure() + named = create_workspace(root, "session-a") + manager = JobManager( + root.outputs, + worker_manager=ScriptedWorkerManager(success_script), + history_path=str(tmp_path / "jobs.sqlite"), + workflow_dir=root.workflows, + ) + app = create_app( + workflow_dir=root.workflows, + output_dir=root.outputs, + prompt_dir=root.prompts, + asset_dir=root.assets, + job_manager=manager, + workspace=root.root, + ) + with TestClient(app, base_url="http://localhost") as c: + default_directories = c.get("/api/server").json()["directories"] + scoped_directories = c.get("/api/server?workspace=session-a").json()[ + "directories" + ] + + assert default_directories["workspace"] == root.root + assert default_directories["workflows"] == root.workflows + assert default_directories["assets"] == root.assets + assert default_directories["outputs"] == root.outputs + assert default_directories["prompts"] == root.prompts + + assert scoped_directories["workspace"] == named.root + assert scoped_directories["workflows"] == named.workflows + assert scoped_directories["assets"] == named.assets + assert scoped_directories["outputs"] == named.outputs + assert scoped_directories["prompts"] == named.prompts + + assert scoped_directories["workspace"] != default_directories["workspace"] + + def test_auth_required_and_token_never_disclosed(tmp_path): token = "s3cr3t-token-value" with client(tmp_path, token=token) as c: From 97ef14b204391e5321b9aa2bae4b680b730b88ed Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 21:27:04 -0500 Subject: [PATCH 067/181] feat(engine): #385 - record shot boundaries on joined videos and in the manifest concat_videos, dissolve_videos and run_chain measure where each shot landed, in frames and in samples; the other AudioVideo constructors carry, rescale, re-measure or drop them. The manifest entry and step_end carry shots named by the step's previous_result:shot@ references; get_gallery_metadata reports media.shots and get_output_frames(seams=true) falls back to them. Co-Authored-By: Claude Opus 5.5 --- dw/pipeline_processors/chain.py | 46 +++++++- dw/result.py | 32 ++++-- dw/runs.py | 31 ++++++ dw/server/app.py | 35 +++++- dw/shots.py | 188 ++++++++++++++++++++++++++++++++ dw/tasks/audio_utils.py | 10 +- dw/tasks/concat_videos.py | 27 ++++- dw/tasks/dissolve_videos.py | 66 ++++++++++- dw/tasks/interpolate_frames.py | 10 +- dw/tasks/pair_audio.py | 13 ++- dw/tasks/stabilize.py | 10 +- dw/tasks/task.py | 10 +- dw/tasks/video_utils.py | 4 +- dw/workflow.py | 22 ++++ dw_mcp/media.py | 5 +- dw_mcp/server.py | 8 +- tests/test_mcp_server.py | 10 +- 17 files changed, 488 insertions(+), 39 deletions(-) create mode 100644 dw/shots.py diff --git a/dw/pipeline_processors/chain.py b/dw/pipeline_processors/chain.py index 6e8811a2..da346458 100644 --- a/dw/pipeline_processors/chain.py +++ b/dw/pipeline_processors/chain.py @@ -41,6 +41,7 @@ get_artifact_list, output_file_path, ) +from ..shots import shot_record, without_samples from ..tasks.audio_utils import ( as_channels_samples, equal_power_crossfade_join, @@ -315,6 +316,11 @@ def run_chain(pipeline, chain_definition, arguments): audio = None # joined generated audio, (channels, samples) float32 audio_rate = None carry = None + # One shot per segment, measured as the picture and track grow (#378); + # the frames are counted rather than read off `frames`, which a spilled + # chain never fills + shots = [] + frame_count = 0 for segment in config.plan: segment_arguments = dict(arguments) @@ -349,6 +355,16 @@ def run_chain(pipeline, chain_definition, arguments): segment_audio, segment_rate = _generated_audio(artifact) kept_frames = segment_frames[segment.head_trim :] + start_sample = audio.shape[1] if audio is not None else 0 + shots.append( + shot_record( + f"segment {segment.index + 1}", + frame_count, + len(kept_frames), + start_sample, + ) + ) + frame_count += len(kept_frames) if spill is not None: spill.write( kept_frames, @@ -363,6 +379,10 @@ def run_chain(pipeline, chain_definition, arguments): audio, audio_rate, segment_audio, segment_rate, segment, config ) + shots[-1]["num_samples"] = ( + audio.shape[1] if audio is not None else 0 + ) - start_sample + # The segment's raw output is finished with - the frames live on # (in RAM or on disk) and the carry frame is extracted. Free it # before the next segment needs the accelerator. @@ -387,10 +407,32 @@ def run_chain(pipeline, chain_definition, arguments): if spill is None: frames = frames[: config.total_frames] return AudioVideo( - frames, config.source_audio, config.source_rate, fps=config.fps + frames, + config.source_audio, + config.source_rate, + fps=config.fps, + shots=_trimmed_shots(shots, config.total_frames), ) - return AudioVideo(frames, audio, audio_rate, fps=config.fps) + if audio is None: + shots = without_samples(shots) + return AudioVideo(frames, audio, audio_rate, fps=config.fps, shots=shots) + + +def _trimmed_shots(shots, total_frames): + """A match_audio chain's shots, cut where its overshooting picture is. + + The soundtrack is the caller's own, laid under whole rather than built + segment by segment, so no shot has a stretch of it to measure: the sample + side is cleared, not derived. + """ + trimmed = [] + for shot in without_samples(shots): + if shot["start_frame"] >= total_frames: + break + end = min(shot["start_frame"] + shot["num_frames"], total_frames) + trimmed.append({**shot, "num_frames": end - shot["start_frame"]}) + return trimmed class ChainConfig: diff --git a/dw/result.py b/dw/result.py index d6d9d2aa..1b875ef8 100644 --- a/dw/result.py +++ b/dw/result.py @@ -461,7 +461,7 @@ class AudioVideo: the result mux them into one file instead of dropping the audio on the floor. """ - def __init__(self, frames, audio, sample_rate, fps=None): + def __init__(self, frames, audio, sample_rate, fps=None, shots=None): """ Args: frames: The video, as PIL images or an array of frames @@ -474,11 +474,16 @@ def __init__(self, frames, audio, sample_rate, fps=None): joins 24 fps shots writing them at 8 is three times slow with its audio still the right length (#84). A declared `result.fps` still wins over this + shots: Where each input landed, for a video a step joined from + several - a list of shot records (dw/shots.py), or None for a + video that is one shot. Carried into the step's manifest + entry when the video is saved (#378) """ self.frames = frames self.audio = audio self.sample_rate = sample_rate self.fps = fps + self.shots = shots class AudioTrack: @@ -537,6 +542,10 @@ def __init__(self, result_definition, consumed_by_normalizer=False): self.result_list = [] self.metadata = None self.saved_files = [] + # The shots (dw/shots.py) of each saved file whose artifact carried + # any, keyed by the path in saved_files - plain data, so a step cache + # hit's stripped copy still reports them (#378) + self.saved_shots = {} # Set when a select step's Selected wrapper flows through # add_result - the winning position/score, replayable in the # manifest and step_end alongside the unwrapped value (#119) @@ -715,6 +724,7 @@ def save(self, output_dir, default_base_name): if not self.result_definition.get("save", True) or content_type is None: logger.debug("Skipping save - disabled or no content type specified") self.saved_files = [] + self.saved_shots = {} return self.saved_files # Determine base filename with validation. A file_base_name *replaces* @@ -743,6 +753,7 @@ def save(self, output_dir, default_base_name): # Save each result, collecting the paths written as the step's manifest saved_files = [] + saved_shots = {} for i, result in enumerate(self.result_list): if content_type.endswith("json"): # Handle JSON content type @@ -756,16 +767,19 @@ def save(self, output_dir, default_base_name): else: # Handle other content types for j, artifact in enumerate(self._artifacts_for(result)): - saved_files.extend( - self.save_artifact( - validated_output_dir, - artifact, - f"{file_base_name}-{i}.{j}", - content_type, - extension, - ) + paths = self.save_artifact( + validated_output_dir, + artifact, + f"{file_base_name}-{i}.{j}", + content_type, + extension, ) + shots = getattr(artifact, "shots", None) + if shots: + saved_shots.update((path, shots) for path in paths) + saved_files.extend(paths) self.saved_files = saved_files + self.saved_shots = saved_shots return saved_files def save_artifact( diff --git a/dw/runs.py b/dw/runs.py index 919a3437..6d252e05 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -649,3 +649,34 @@ def manifest_relative_files(files, run_dir): path if relative.startswith(os.pardir) else relative.replace(os.sep, "/") ) return recorded + + +def recorded_shots(output_root, relative_path): + """The shot boundaries the run's manifest records for one of its files. + + None when the file is not in a run directory (the flat layout), its run + has no readable manifest, or no step recorded shots for it - a file that + was not joined from shots, or one written before shots were recorded. + """ + from .shots import shots_for_file + + folder, run_id, _subfolder = split_run_path(relative_path) + if not run_id: + return None + run_dir = os.path.join(output_root, folder, run_id) + manifest = _read_manifest(run_dir) + if manifest is None: + return None + prefix = f"{folder}/{run_id}/" if folder else f"{run_id}/" + own = relative_path[len(prefix) :] if relative_path.startswith(prefix) else None + if own is None: + return None + for entry in manifest.get("steps") or []: + if not isinstance(entry, dict) or entry.get("reused"): + continue + files = entry.get("files") or [] + if own in files: + shots = shots_for_file(entry.get("shots"), own, files) + if shots: + return shots + return None diff --git a/dw/server/app.py b/dw/server/app.py index 13d925d1..00ab1821 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -103,6 +103,7 @@ record_run_versions, resolve_output_reference, run_versions, + recorded_shots, split_run_path, ) from ..workspace import ( @@ -3031,6 +3032,10 @@ def gallery_metadata( if MEDIA_KINDS.get(extension) in ("audio", "video") else None ) + if media is not None and source == "output": + # Where each shot of a joined video sits, as the run that wrote + # it recorded (dw/shots.py) - null for a file not joined from shots + media["shots"] = recorded_shots(ws.outputs, name) return { "name": name, "source": source, @@ -3158,8 +3163,10 @@ def gallery_frames( numbers) for the last frame before and first frame after each boundary, side by side. `boundaries` is the comma list of frame indexes each shot after the first starts at, and `names` the - shots' names; both are required with `seams` until a joined file - carries its own (stage 2 of docs/proposals/output-assessment.md). + shots' names. Without `boundaries`, an output's seams are the shots + its run's manifest recorded for it (a `concat_videos`, + `dissolve_videos` or chained step), named as recorded unless `names` + is given; a file with none recorded still needs `boundaries`. Tiles are downscaled to `max_dimension` on their longest side. `crop` is `x,y,width,height` in the video's own source pixels (`video_shape`'s `width`/`height`) - resolved once and cut from @@ -3232,15 +3239,31 @@ def gallery_frames( ) ] else: - if not boundaries: + recorded = ( + None + if boundaries or is_asset_reference(name) + else recorded_shots(ws.outputs, name) + ) + if recorded: + # The file's own seams, from its run's manifest + starts = [shot["start_frame"] for shot in recorded[1:]] + shot_names = ( + [n.strip() for n in names.split(",")] + if names + else [shot["name"] for shot in recorded] + ) + elif not boundaries: raise HTTPException( status_code=400, detail="`seams` needs `boundaries`: the frame index each " "shot after the first starts at, comma-separated - this " - "file carries none of its own", + "file's run recorded no shots for it", + ) + else: + starts = [int(b) for b in boundaries.split(",") if b.strip()] + shot_names = ( + [n.strip() for n in names.split(",")] if names else None ) - starts = [int(b) for b in boundaries.split(",") if b.strip()] - shot_names = [n.strip() for n in names.split(",")] if names else None wanted = ( None if seams.lower() == "true" diff --git a/dw/shots.py b/dw/shots.py new file mode 100644 index 00000000..174bc992 --- /dev/null +++ b/dw/shots.py @@ -0,0 +1,188 @@ +"""Shot boundaries a joined video carries: where each input landed in it. + +A step that joins shots - `concat_videos`, `dissolve_videos`, a chained +pipeline - knows exactly where every seam fell, in frames and in samples, and +used to throw that away: a consumer checking a cut had to re-derive the seams +from arguments, and a shot whose track ran 267 samples long drifted the rest +of the cut with nothing saying where (#378). The join now records one entry +per shot on the `AudioVideo` it returns (`AudioVideo.shots`), the step's +manifest entry carries them, and `get_gallery_metadata` reads them back. + +A shot is a dict: + +- `name` - which input it was: `shot@` when the step named a `for_each` + member, else the path it was given, else `video N` / `segment N` +- `start_frame`, `num_frames` - its place on the joined picture. The shots + partition the frames: the counts add up to the file's frame count +- `start_sample`, `num_samples` - its place on the joined track, *measured* + from the waveform the join built rather than derived from the frame + numbers, so an overrun shows up as a count that disagrees with the frames'. + None when the video has no track, or when the track is one the join did + not build shot by shot (a chain's `match_audio`) +- `overlap_frames` - a dissolve's head: the frames at its start that are + blended with the shot before it + +Every other `AudioVideo` constructor either carries the list (same frames), +rescales it (`interpolate_frames`), re-measures the sample side for a new +track (`pair_audio`), or builds a video with no shots at all. +`tests/test_shots.py` fails on a constructor site nobody decided for. +""" + +import copy + +from .arguments import PREVIOUS_RESULT_PREFIX + +# The step names a for_each member as `@`; only a member of the +# group conventionally called `shot` names a shot +SHOT_REFERENCE_PREFIX = f"{PREVIOUS_RESULT_PREFIX}shot@" + + +def shot_record( + name, start_frame, num_frames, start_sample=None, num_samples=None, **extra +): + """One shot's entry, in the key order the manifest shows.""" + record = { + "name": name, + "start_frame": int(start_frame), + "num_frames": int(num_frames), + "start_sample": None if start_sample is None else int(start_sample), + "num_samples": None if num_samples is None else int(num_samples), + } + record.update(extra) + return record + + +def carried_shots(source): + """The shots of a video whose frames a step kept one for one, copied.""" + shots = getattr(source, "shots", None) + return copy.deepcopy(shots) if shots else None + + +def without_samples(shots): + """The shots with their sample side cleared - a track that is gone.""" + return [{**shot, "start_sample": None, "num_samples": None} for shot in shots] + + +def rescaled_shots(shots, multiplier): + """The shots of a video whose frames were multiplied by interpolation. + + N frames become (N - 1) * multiplier + 1: every frame but the last gains + multiplier - 1 frames after it. A shot starting at frame s now starts at + s * multiplier, and the last shot keeps the one frame nothing follows. + Interpolation drops the track, so the sample side goes with it. + """ + if not shots: + return None + total = sum(shot["num_frames"] for shot in shots) + rescaled = [] + for shot in shots: + start = shot["start_frame"] * multiplier + end = shot["start_frame"] + shot["num_frames"] + new_end = (end - 1) * multiplier + 1 if end == total else end * multiplier + entry = { + **shot, + "start_frame": start, + "num_frames": new_end - start, + "start_sample": None, + "num_samples": None, + } + if shot.get("overlap_frames"): + entry["overlap_frames"] = shot["overlap_frames"] * multiplier + rescaled.append(entry) + return rescaled + + +def remeasured_shots(shots, fps, sample_rate, total_samples): + """The shots of a video laid over a new track by pair_audio. + + The frame side is unchanged. The new track was not built shot by shot, so + each shot's samples are the stretch of the track its frames play over: a + shot starts at start_frame / fps seconds, and the last one runs to the end + of the track as written - with `fit: "video"` that is the fitted length. + Without a frame rate there is no way to place a frame on the track, so the + sample side is cleared rather than guessed. + """ + if not shots: + return None + if not fps or not sample_rate or total_samples is None: + return without_samples(shots) + starts = [ + min(int(round(shot["start_frame"] / fps * sample_rate)), total_samples) + for shot in shots + ] + ends = starts[1:] + [total_samples] + return [ + {**shot, "start_sample": start, "num_samples": max(end - start, 0)} + for shot, start, end in zip(shots, starts, ends) + ] + + +def shot_reference_names(references): + """`shot@` for each entry of a step's list naming a shot, else None. + + A step's `videos` argument is written as a list of `previous_result:` + references; by the time the task runs `gather:` has expanded into exactly + such a list, so the entry at position i names the video the join put at + position i. Only a reference to a `shot@` member names a shot - anything + else keeps the name the join gave it. + """ + if not isinstance(references, list): + return None + names = [] + for reference in references: + if isinstance(reference, str) and reference.startswith(SHOT_REFERENCE_PREFIX): + # `previous_result:shot@x.field` names the member, not the field + member = reference[len(PREVIOUS_RESULT_PREFIX) :] + names.append(member.split(".", 1)[0]) + else: + names.append(None) + return names + + +def named_shots(shots, names): + """The shots with each positional one renamed where the step named it.""" + if not shots or not names or len(names) != len(shots): + return shots + return [ + {**shot, "name": name} if name else shot for shot, name in zip(shots, names) + ] + + +def step_shots(saved_shots, saved_files, references=None): + """The `shots` a step's manifest entry and step_end carry, or None. + + `saved_shots` maps each file the step wrote to the shots its video + carried (`Result.saved_shots`). A step that wrote one file lists that + file's shots; one that wrote several marks each shot with the `file` it + belongs to, so no shot is ever read against the wrong file. `references` + is the step's `videos` argument as written, which names the shots. + """ + if not saved_shots: + return None + names = shot_reference_names(references) + files = [path for path in saved_files or [] if path in saved_shots] + if len(files) == 1 and len(saved_files) == 1: + return named_shots(copy.deepcopy(saved_shots[files[0]]), names) + return [ + {**shot, "file": path} + for path in files + for shot in named_shots(copy.deepcopy(saved_shots[path]), names) + ] + + +def shots_for_file(shots, path, step_files): + """The shots of one file out of a manifest entry's `shots`, or None. + + `path` and `step_files` are as the manifest records them, so a + run-relative path matches a run-relative `file`. + """ + if not shots: + return None + if any("file" in shot for shot in shots): + own = [ + {key: value for key, value in shot.items() if key != "file"} + for shot in shots + if shot.get("file") == path + ] + return own or None + return shots if list(step_files or []) == [path] else None diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index 5584bc0d..06a1edb3 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -296,17 +296,23 @@ def bleed_join( return numpy.concatenate([previous, following], axis=1) -def crossfade_concat(waveforms, sample_rate, crossfade_ms): +def crossfade_concat(waveforms, sample_rate, crossfade_ms, starts=None): """Concatenate waveforms, overlapping each seam by an equal-power crossfade. The classic crossfade: each seam overlaps the two waveforms by the fade window, so the result is shorter than the plain sum by one window per seam. + + `starts`, when given a list, is filled with the sample each waveform + begins at in the result - where its crossfade opens - measured as the + result grows rather than worked out from the lengths (#378). """ waveforms = [as_channels_samples(waveform) for waveform in waveforms] if not waveforms: raise ValueError("No waveforms to concatenate") result = waveforms[0] + if starts is not None: + starts.append(0) for following in waveforms[1:]: result, following = _matched_channels(result, following) window = min( @@ -314,6 +320,8 @@ def crossfade_concat(waveforms, sample_rate, crossfade_ms): result.shape[1], following.shape[1], ) + if starts is not None: + starts.append(result.shape[1] - window) if window == 0: result = _declick_join(result, following, sample_rate) continue diff --git a/dw/tasks/concat_videos.py b/dw/tasks/concat_videos.py index 3ecb12fd..aec91d8c 100644 --- a/dw/tasks/concat_videos.py +++ b/dw/tasks/concat_videos.py @@ -12,6 +12,7 @@ from ..events import emit_warning from ..result import AudioVideo +from ..shots import shot_record from .audio_utils import ( as_channels_samples, bleed_join, @@ -184,10 +185,21 @@ def concat_videos( frames = [] audio = None audio_native_rate = None + # Where each video landed, measured on the joined picture and track as + # they grow - never derived from the frame count, so a track that runs + # long shows up here as the samples it actually took (#378) + shots = [] for index, (video, clip) in enumerate(zip(videos, clips)): head_trim = trim_frames if index > 0 else 0 + start_frame = len(frames) + start_sample = audio.shape[1] if audio is not None else 0 frames.extend(clip[head_trim:]) + shots.append( + shot_record( + names[index], start_frame, len(frames) - start_frame, start_sample + ) + ) if waveforms[index] is None: continue @@ -227,6 +239,14 @@ def concat_videos( ) audio_native_rate = video.sample_rate + for shot, following in zip(shots, shots[1:] + [None]): + # A seam's crossfade leaves the samples before it where they were, so + # a shot's track is everything up to where the next one's began + end = following["start_sample"] if following else _length(audio) + shot["num_samples"] = None if audio is None else end - shot["start_sample"] + if audio is None: + shot["start_sample"] = None + logger.debug(f"Concatenated {len(videos)} videos into {len(frames)} frames") # The rate the caller declared, else the rate the first input carries - # either beats the result's 8 fps default (#84) @@ -234,4 +254,9 @@ def concat_videos( (v.fps for v in videos if getattr(v, "fps", None)), None, ) - return AudioVideo(frames, audio, sample_rate, fps=written_fps) + return AudioVideo(frames, audio, sample_rate, fps=written_fps, shots=shots) + + +def _length(audio): + """How many samples a joined track holds, 0 for none.""" + return 0 if audio is None else audio.shape[1] diff --git a/dw/tasks/dissolve_videos.py b/dw/tasks/dissolve_videos.py index c8061a16..e387e2e2 100644 --- a/dw/tasks/dissolve_videos.py +++ b/dw/tasks/dissolve_videos.py @@ -19,6 +19,7 @@ from ..events import emit_warning from ..result import AudioVideo +from ..shots import shot_record from .audio_utils import ( as_channels_samples, crossfade_concat, @@ -98,7 +99,10 @@ def dissolve_videos( ) joined = clips[0] + # Where each clip's first frame landed - the start of its dissolve + frame_starts = [0] for clip in clips[1:]: + frame_starts.append(len(joined) - dissolve_frames) joined = _dissolve_join(joined, clip, dissolve_frames) if fade_in_frames or fade_out_frames: @@ -116,8 +120,23 @@ def dissolve_videos( ) frames = [Image.fromarray(frame) for frame in joined.round().astype(numpy.uint8)] + sample_starts = [] audio, sample_rate = _dissolve_audio( - loaded, dissolve_frames, fps, match_levels, match_levels_dbfs, sample_rate + loaded, + dissolve_frames, + fps, + match_levels, + match_levels_dbfs, + sample_rate, + sample_starts, + ) + shots = _dissolve_shots( + video_names(videos), + frame_starts, + len(frames), + sample_starts, + audio, + dissolve_frames, ) logger.info( f"Dissolved {len(clips)} videos into {len(frames)} frames " @@ -127,7 +146,39 @@ def dissolve_videos( (v.fps for v in loaded if getattr(v, "fps", None)), None, ) - return AudioVideo(frames, audio, sample_rate, fps=written_fps) + return AudioVideo(frames, audio, sample_rate, fps=written_fps, shots=shots) + + +def _dissolve_shots( + names, frame_starts, total_frames, sample_starts, audio, dissolve_frames +): + """One shot per video, partitioning the dissolved picture and track. + + A dissolve belongs to the shot coming in: each shot runs from where its + dissolve opens to where the next one's does, so the counts add up to the + file's and `overlap_frames` says how much of its head is blended. + """ + frame_ends = frame_starts[1:] + [total_frames] + if audio is None: + sample_starts = [None] * len(frame_starts) + sample_ends = sample_starts + else: + sample_ends = sample_starts[1:] + [audio.shape[1]] + shots = [] + for index, name in enumerate(names): + start_sample = sample_starts[index] + shots.append( + shot_record( + name, + frame_starts[index], + frame_ends[index] - frame_starts[index], + start_sample, + None if start_sample is None else sample_ends[index] - start_sample, + ) + ) + if index and dissolve_frames: + shots[-1]["overlap_frames"] = dissolve_frames + return shots def _dissolve_join(previous, following, overlap): @@ -159,8 +210,12 @@ def _dissolve_audio( match_levels=None, match_levels_dbfs=None, sample_rate=None, + starts=None, ): - """Crossfade every video's track over the seams' own span.""" + """Crossfade every video's track over the seams' own span. + + `starts` is filled with where each track begins in the joined one. + """ tracks = [v for v in videos if isinstance(v, AudioVideo) and v.audio is not None] if len(tracks) != len(videos): if tracks: @@ -212,4 +267,7 @@ def _dissolve_audio( ) else: warn_on_level_spread(waveforms, "dissolve_videos") - return crossfade_concat(waveforms, sample_rate, crossfade_ms), sample_rate + return ( + crossfade_concat(waveforms, sample_rate, crossfade_ms, starts), + sample_rate, + ) diff --git a/dw/tasks/interpolate_frames.py b/dw/tasks/interpolate_frames.py index c0cfbe5f..b383cfcf 100644 --- a/dw/tasks/interpolate_frames.py +++ b/dw/tasks/interpolate_frames.py @@ -11,6 +11,7 @@ import torch from ..result import AudioVideo +from ..shots import rescaled_shots from .tensor_image import pil_to_float_tensor as _pil_to_tensor, float_tensor_to_pil from .video_utils import frames_as_pil_list @@ -50,6 +51,7 @@ def interpolate_frames(video, device="cpu", **kwargs): # through by identity. The soundtrack does not survive - the frame count # changes, so pair_audio is how it comes back source_fps = getattr(video, "fps", None) + source_shots = getattr(video, "shots", None) video = frames_as_pil_list(video) if len(video) < 2: raise ValueError(f"Need at least 2 frames to interpolate, got {len(video)}") @@ -74,8 +76,14 @@ def interpolate_frames(video, device="cpu", **kwargs): # was given, so playing them back at the source rate would run the clip # `multiplier` times long. The rate that keeps the source's duration is # the source's times the multiplier (#84) + # The shots stretch with the frames between them; the track is gone, so + # their sample side goes with it return AudioVideo( - frames, None, None, fps=source_fps * multiplier if source_fps else None + frames, + None, + None, + fps=source_fps * multiplier if source_fps else None, + shots=rescaled_shots(source_shots, multiplier), ) diff --git a/dw/tasks/pair_audio.py b/dw/tasks/pair_audio.py index 320b2833..07e4a701 100644 --- a/dw/tasks/pair_audio.py +++ b/dw/tasks/pair_audio.py @@ -11,6 +11,7 @@ from ..events import emit_warning from ..result import AudioVideo +from ..shots import remeasured_shots from .audio_utils import as_channels_samples logger = logging.getLogger("dw") @@ -186,4 +187,14 @@ def pair_audio(video, audio, sample_rate=None, fps=None, fit=None): waveform = _fit_to_video( as_channels_samples(waveform), rate, frames, frame_rate, fit ) - return AudioVideo(frames, waveform, rate, fps=getattr(video, "fps", None)) + # The picture's shots survive; their samples are re-measured on the new + # track, which was laid under whole rather than built shot by shot + return AudioVideo( + frames, + waveform, + rate, + fps=getattr(video, "fps", None), + shots=remeasured_shots( + getattr(video, "shots", None), frame_rate, rate, waveform.shape[1] + ), + ) diff --git a/dw/tasks/stabilize.py b/dw/tasks/stabilize.py index e2fcc2e4..64204a78 100644 --- a/dw/tasks/stabilize.py +++ b/dw/tasks/stabilize.py @@ -18,6 +18,7 @@ from PIL import Image from ..result import AudioVideo +from ..shots import carried_shots from .video_utils import frames_as_pil_list, load_audio_video logger = logging.getLogger("dw") @@ -117,5 +118,12 @@ def stabilize_video(clip, smooth=0): ) if isinstance(clip, AudioVideo): - return AudioVideo(held, clip.audio, clip.sample_rate, fps=clip.fps) + # Same frames, one for one, so every shot boundary still holds + return AudioVideo( + held, + clip.audio, + clip.sample_rate, + fps=clip.fps, + shots=carried_shots(clip), + ) return held diff --git a/dw/tasks/task.py b/dw/tasks/task.py index e566107c..c633ba4e 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -380,6 +380,7 @@ def _per_frame(image, process): carried through untouched. A single image is processed as itself. """ from ..result import AudioVideo + from ..shots import carried_shots from .video_utils import frames_as_pil_list, is_video if not is_video(image): @@ -387,7 +388,14 @@ def _per_frame(image, process): frames = [process(frame) for frame in frames_as_pil_list(image)] audio = getattr(image, "audio", None) sample_rate = getattr(image, "sample_rate", None) - return AudioVideo(frames, audio, sample_rate, fps=getattr(image, "fps", None)) + # One frame out per frame in, so the shot boundaries carry through too + return AudioVideo( + frames, + audio, + sample_rate, + fps=getattr(image, "fps", None), + shots=carried_shots(image), + ) @register_command( diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index 361f3ffa..aeffb992 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -527,7 +527,9 @@ def _decode_audio_video(handle): f"{audio.shape[1] if audio is not None else 0} audio samples" ) # The file's own rate travels with it: a step that joins videos read - # from disk knows what to write them back at without being told (#84) + # from disk knows what to write them back at without being told (#84). + # A file carries no shots - the manifest that recorded them is the run's, + # not the file's return AudioVideo( frames, audio, sample_rate if audio is not None else None, fps=frame_rate ) diff --git a/dw/workflow.py b/dw/workflow.py index 129396ab..1b72e092 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -42,6 +42,7 @@ resolve_constraint_references, ) from .result_fps import fps_errors +from .shots import step_shots from .subfolders import step_subfolder, subfolder_errors from .reference_names import reference_name_errors from .video_extensions import video_extension_errors @@ -268,6 +269,15 @@ def _release_host_caches(step_name): logger.info(f"Release after {step_name} returned {released:.0f} MB to the OS") +def _relative_shots(entry, run_dir): + """A manifest entry's `shots`, each `file` made relative as `files` is.""" + shots = entry.get("shots") + if not shots or not any("file" in shot for shot in shots): + return {} + files = manifest_relative_files([shot["file"] for shot in shots], run_dir) + return {"shots": [{**shot, "file": f} for shot, f in zip(shots, files)]} + + def selected_field(step_data, selected): """The manifest/step_end 'selected' block for a step's Result.selected. @@ -1493,6 +1503,15 @@ def run( manifest_entry["reused"] = True if selected is not None: manifest_entry["selected"] = selected + # Where each joined shot sits in the file, named by the + # step's own `videos` references (dw/shots.py) + shots = step_shots( + getattr(result, "saved_shots", None), + saved_files, + step_data.get("task", {}).get("arguments", {}).get("videos"), + ) + if shots: + manifest_entry["shots"] = shots # No entry at all for a step the parent saves for: the # parent's own entry names the same files, under the step # name the caller wrote (#92) @@ -1506,6 +1525,8 @@ def run( step_end_data["reused"] = True if selected is not None: step_end_data["selected"] = selected + if shots: + step_end_data["shots"] = shots run_context.emit( "step_end", workflow=workflow_id, @@ -1651,6 +1672,7 @@ def _write_run_manifest( "files": manifest_relative_files( entry.get("files"), self._run_dir ), + **_relative_shots(entry, self._run_dir), } for entry in self.manifest ], diff --git a/dw_mcp/media.py b/dw_mcp/media.py index fc6b1373..baccda8e 100644 --- a/dw_mcp/media.py +++ b/dw_mcp/media.py @@ -231,8 +231,9 @@ def get_output_frames( the last frame before and the first frame after each boundary, side by side). `boundaries` is the list of frame indexes each shot after the first starts at - the running sum of the shots' `frame_count` from - `get_gallery_metadata` on their own files - `names` the shots' names - - both needed with `seams` until a joined file carries its own. + `get_gallery_metadata` on their own files - `names` the shots' names. + Without `boundaries`, an output joined from shots uses the boundaries + its run recorded (`get_gallery_metadata`'s `media.shots`). `crop` is `[x, y, width, height]` in the video's own source pixels - the same convention `get_output_image` uses - resolved once against diff --git a/dw_mcp/server.py b/dw_mcp/server.py index de538f2d..c7bd2994 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -416,7 +416,8 @@ def get_gallery_metadata( `get_job_workflow(job_id)` when known, else a kept asset has no provenance. `media` itself carries duration, sample rate, channels, fps, size and level - the checks an agent that cannot - listen makes on a deliverable. + listen makes on a deliverable; `media.shots` places a joined + video's shots. `envelope=true` adds that level second by second (`media.envelope.rms_dbfs` / `peak_dbfs`), which says *where* in a track something is: a shot's last frame, a seam's hole, where a @@ -579,9 +580,8 @@ def get_output_frames( """See a generated video as frames - no video content type exists over MCP. One selector: `count` (contact sheet), `at` (seconds or "frame:N"), or `seams` (true, or seam numbers from 1) for each - join's frame pair. `seams` needs `boundaries` - each later shot's - first frame, running sum of `get_gallery_metadata`'s `frame_count`; - `names` names the shots. Over budget, tiles shrink together. + join's frame pair, at a joined output's `media.shots`; else + `boundaries` (each later shot's first frame) and `names`. Over budget, tiles shrink together. `hear=N` adds N seconds of soundtrack around each `at`. `crop` is `[x, y, width, height]` in the video's own source pixels, cut from every frame before any downscale, like diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index 2ad02bad..871d15b2 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -470,14 +470,14 @@ async def test_the_media_tool_descriptions_say_what_they_hand_back(): @pytest.mark.asyncio async def test_the_frames_tool_says_where_boundaries_come_from(): - """Until a joined file carries its own shots (stage 2), the agent has to - derive seam boundaries; the tool has to say from what, or `seams` is a - parameter nobody can fill in.""" + """A joined output's seams come from the shots its run recorded (#385); + the tool has to say so, and say what to pass for a file with none, or + `seams` is a parameter nobody can fill in.""" server = server_over(ok({})) tools = await tools_of(server) text = tools["get_output_frames"].description - assert "get_gallery_metadata" in text - assert "frame_count" in text + assert "media.shots" in text + assert "boundaries" in text # Every tool, the arguments a client would send, and the one API call it is From 75566731720bdeb90501a58e1cefb5b0a55c76a7 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 21:28:06 -0500 Subject: [PATCH 068/181] docs: #385 - document recorded shot boundaries Co-Authored-By: Claude Opus 5.5 --- CLAUDE.md | 15 +++++++++++++++ docs/MCP.md | 4 ++-- docs/WORKFLOW_GUIDE.md | 35 +++++++++++++++++++++++++++++++++++ 3 files changed, 52 insertions(+), 2 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 06e89aa8..dbefee67 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -593,6 +593,21 @@ same reason - default setup cannot load a pack. name into the JSON. An entry violation is reported at `arguments.shots[0].num_frames`, and the rule is reported beside the field in the catalog's `lists` block as well as in `constraints` +- **A joined video records its shots, measured** — `concat_videos`, + `dissolve_videos` and both `run_chain` returns set `AudioVideo.shots` + (`dw/shots.py`): one `{name, start_frame, num_frames, start_sample, + num_samples}` per input. The frames are partitioned, and the samples are + read off the waveform the join built, never derived from the frames, so a + shot's overrun stays visible (#385). Every other `AudioVideo` constructor + carries, rescales (`interpolate_frames`), re-measures (`pair_audio`) or + drops them, and `tests/test_shots.py` enumerates the constructor sites with + `ast`, so a new one fails until someone decides for it. `Result.save` + keeps them as plain data in `saved_shots` (path -> shots), which survives + the step cache's stripped copy. The manifest entry and `step_end` carry + `shots`, renamed `shot@` from the step's `videos` references (as + `selected_field` does). `recorded_shots` (`dw/runs.py`) reads them back for + `get_gallery_metadata`'s `media.shots` and for `get_output_frames(seams=true)` + without `boundaries`. The mp4 itself carries nothing yet - **Step cache**: a process-wide singleton (`dw/step_cache.py`) consulted by every `Workflow.run`, including server jobs; entries are keyed by `(workflow id, step name)` and validated against the output *root*, never the per-run directory - a run directory is new every execution and would defeat the cache; disabled entirely when the workflow sets no `seed`; a hit reports the earlier run's files with `reused: true` and writes nothing new; `memory clear` drops it. This is why "Run again" on a seeded workflow finishes instantly and generates nothing - the job page says so when every step was reused, and `POST /api/jobs/{id}/rerun` with `{"new_seed": true}` (MCP `rerun_job(new_seed=True)`) draws a fresh seed into the workflow's seed variable, which is the way to get a different image diff --git a/docs/MCP.md b/docs/MCP.md index d0067bf3..92e60957 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -227,7 +227,7 @@ when no single workflow covers it. | `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in | | `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds | | `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None, folder=None, version=None)` | `limit`, `subfolder`, `only_orphans`, `workspace`, `folder`, `version` | List generated output files, newest first. A name is `//`, where `` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. `folder` with `version` lists that one run's files, and `output:/v5/` names one in a workflow; every other tool takes `name`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | -| `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). Only an image (PNG/JPEG/WebP) embeds `metadata` this way — it is always null for audio and video, and `next` then points at `get_job_workflow(job_id)` when `job` is known, or says a kept asset carries no provenance at all when it isn't. `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | +| `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). Only an image (PNG/JPEG/WebP) embeds `metadata` this way — it is always null for audio and video, and `next` then points at `get_job_workflow(job_id)` when `job` is known, or says a kept asset carries no provenance at all when it isn't. `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `media.shots` is set on an output joined from shots: one `{name, start_frame, num_frames, start_sample, num_samples}` per shot, as the join measured them (see the run manifest in docs/WORKFLOW_GUIDE.md), else null. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | ### Media @@ -235,7 +235,7 @@ when no single workflow covers it. | --- | --- | --- | | `get_output_image(name, max_dimension=768, workspace=None, crop=None)` | `name`, `max_dimension`, `workspace`, `crop` | Look at a generated image, downscaled to `max_dimension` on its longest side. Returns the image plus a text part reporting `original_size`, `returned_size` and `bytes`, so a downscale is never silent. `crop` is `[x, y, width, height]` in the original's pixels, cut before the downscale and reported back clamped - the way to see a region of a 2K still at 100%, where the whole would be shrunk past what a small element or a decode-tiling seam can be judged at. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `get_output_audio(name, start=None, duration=None, workspace=None)` | `name`, `start`, `duration`, `workspace` | Listen to a generated soundtrack as base64 - an audio output, or the track muxed into a video (#193) - in its own encoding when served whole, WAV when extracted or excerpted. No downscale exists for audio, so a whole clip over the 4MB budget is refused rather than cut (#204); ask for the part instead with `start` and `duration` in seconds, and the text part names what was cut (`excerpt: 2.0s from 10.0s of 240.0s`) so a slice is never mistaken for the whole. `get_gallery_metadata`'s envelope says where in a track to look. `workspace` names the workspace for this one call without switching the session to it | -| `get_output_frames(name, at=None, seams=None, count=None, boundaries=None, names=None, max_dimension=512, hear=None, workspace=None)` | `name`, `at`, `seams`, `count`, `boundaries`, `names`, `max_dimension`, `hear`, `workspace` | See a generated video as frames, since there is no video content type over MCP (#193). One selector per call: `count` for an evenly spaced contact sheet, `at` for moments (seconds or `"frame:N"`), `seams` (true, or seam numbers from 1) for the last frame before and first frame after each join side by side (each pair carries `difference`, the mean pixel change across the join, 0-255 - rank seams by it and look at the worst) - `boundaries` is each later shot's first frame - the running sum of the shots' `frame_count` from `get_gallery_metadata` on their own `intermediate/` files - and `names` names them. Tiles are fitted to `max_dimension` and, when the set would exceed the 4MB budget, shrunk together rather than dropped; the text part lists each tile and says so. `hear=N` also returns N seconds of soundtrack centred on each `at` moment, after its image - the way to check a hit point or lip-sync without reconciling two clocks; a mute clip keeps its frames and says `no soundtrack` | +| `get_output_frames(name, at=None, seams=None, count=None, boundaries=None, names=None, max_dimension=512, hear=None, workspace=None)` | `name`, `at`, `seams`, `count`, `boundaries`, `names`, `max_dimension`, `hear`, `workspace` | See a generated video as frames, since there is no video content type over MCP (#193). One selector per call: `count` for an evenly spaced contact sheet, `at` for moments (seconds or `"frame:N"`), `seams` (true, or seam numbers from 1) for the last frame before and first frame after each join side by side (each pair carries `difference`, the mean pixel change across the join, 0-255 - rank seams by it and look at the worst). On an output joined from shots (`concat_videos`, `dissolve_videos`, a chained step) `seams` alone is enough: the seams and their names are the `media.shots` its run recorded. For any other file pass `boundaries` - each later shot's first frame, the running sum of the shots' `frame_count` from `get_gallery_metadata` on their own `intermediate/` files - and `names` to name them; either one given overrides the recorded value. Tiles are fitted to `max_dimension` and, when the set would exceed the 4MB budget, shrunk together rather than dropped; the text part lists each tile and says so. `hear=N` also returns N seconds of soundtrack centred on each `at` moment, after its image - the way to check a hit point or lip-sync without reconciling two clocks; a mute clip keeps its frames and says `no soundtrack` | | `get_output_text(name, max_characters=20000, workspace=None)` | `name`, `max_characters`, `workspace` | Read a text output — a prompt enhancement, or any step whose result is `text/plain` or JSON. Reports the file's real length and whether it was truncated. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `download_output(name, destination=None, overwrite=False, workspace=None)` | `name`, `destination`, `overwrite`, `workspace` | Save one output file to local disk, of any content type. `destination` may be a full path or a directory; `~` expands and missing parent directories are created. `overwrite=True` is required to replace a file already at the resolved path. Over the stdio `dw-mcp`, omitting `destination` saves under the output's own name in the current working directory. Over a `dw.serve --mcp` endpoint the file lands on the server confined to that workspace (a relative path is joined onto it), and `destination` is required there - an omitted one is refused rather than dropped loose in the workspace root, where nothing can find or delete it later (#353); use the `url` `list_gallery` reports, `get_output_image`/`get_output_audio`/`get_output_frames` for inline content, or `keep_output` to make it a named asset instead. Returns nothing to the conversation but where the file landed — unlike the other media tools, the point is a file on disk, not a payload in context. Writes on the machine running the MCP server - over `dw.serve --mcp` that is the GPU box. A write that fails there (a path that exists only on the client, for instance) comes back as an error naming the server-side write and the client-side alternatives, not as an anonymous tool failure. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it | | `delete_output(name=None, workspace=None, job_id=None)` | exactly one of `name` / `job_id`, `workspace` | Permanently remove one generated file from the output directory, or - with `name` a `/` run directory, or with `job_id` - a whole run. By `job_id` the run directory is read from the job record (`run_dir`) and the reply adds `job_id` and the resolved `run_dir` to the usual `name` / `deleted` / `run_swept`; a job that never wrote a run directory, or an unknown one, is an error. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it; a `job_id` delete with no `workspace` goes to the workspace the job ran in | diff --git a/docs/WORKFLOW_GUIDE.md b/docs/WORKFLOW_GUIDE.md index 4a82ef6d..29fb78f6 100644 --- a/docs/WORKFLOW_GUIDE.md +++ b/docs/WORKFLOW_GUIDE.md @@ -1469,6 +1469,41 @@ each already names something pinned by the asset library or by the manifest's SHA-256 recorded in the manifest. The manifest also lists which stored prompts were inlined, since inlining loses the name. +A step that joins shots (`concat_videos`, `dissolve_videos`, or a pipeline +step with a `chain`) also records where each one landed, as `shots` on its +manifest entry (and on its `step_end` event): + +```json +{ + "step": "cut", + "files": ["final/film.mp4"], + "subfolder": "final", + "shots": [ + {"name": "shot@a", "start_frame": 0, "num_frames": 121, "start_sample": 0, "num_samples": 242267}, + {"name": "shot@b", "start_frame": 121, "num_frames": 97, "start_sample": 242267, "num_samples": 194000} + ] +} +``` + +The shots partition the file's frames: the `num_frames` add up to the frame +count. The sample fields are *measured* off the track the join built, not +worked out from the frames. That means a shot whose track ran long shows it +here: the first shot above is 267 samples longer than 121 frames at 24 fps. +They are null when the video has no track, and for a chain that uses +`match_audio`. A shot is named `shot@` when the step's `videos` entry was +a `previous_result:shot@` reference, else by its input's position +(`video N`, a chain's `segment N`). A dissolve's shots after the first carry +`overlap_frames`, the head they share with the shot before. A step that wrote +several joined files marks each shot with its `file`. + +The steps that keep the frames pass `shots` on. `stabilize` and the per-frame +tasks keep them as they are. `interpolate_frames` rescales them to the new +frame count and clears the samples. `pair_audio` measures the samples again +against the new track. Everything else drops them: an audio task's track, +say, or a video read back from a file. `get_gallery_metadata` reports the +recorded shots as `media.shots`, and `get_output_frames(seams=true)` uses them +when you pass no `boundaries`. + The file is a valid workflow, and running it again is `python -m dw.run workflow.json` or handing its contents to `run_workflow` as `inline_workflow` — but either way the `asset:` and `output:` names in it resolve against the From 2473820f97b7fdeecf19a3c21bc018162244cae1 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 21:59:39 -0500 Subject: [PATCH 069/181] test(engine): #385 - pin shot boundaries: constructors, joins, overrun, pair_audio, slice_audio, round-trip Co-Authored-By: Claude Sonnet 5 --- tests/test_shots.py | 591 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 591 insertions(+) create mode 100644 tests/test_shots.py diff --git a/tests/test_shots.py b/tests/test_shots.py new file mode 100644 index 00000000..886779eb --- /dev/null +++ b/tests/test_shots.py @@ -0,0 +1,591 @@ +""" +Unit tests for shot boundaries (#385): where each input landed on a video a +step joined from several, recorded on the `AudioVideo` a join returns +(`dw/shots.py`), carried into the step's manifest entry and read back by +`dw.runs.recorded_shots`. + +Covers: every `AudioVideo(` constructor site in `dw/` is accounted for by an +explicit decision (populates / carries / rescales / remeasures / none); +`concat_videos`, `dissolve_videos` and `run_chain` populate shots that +partition their output; a track measured longer than the frames imply is +recorded as measured, not derived; `pair_audio` re-measures the sample side +against a new track; `slice_audio` drops shots entirely (its output is audio, +which has no picture to partition); `rescaled_shots` is exercised directly for +`interpolate_frames`; and a join's shots survive a save/manifest/ +`recorded_shots` round trip. +""" + +import ast +import json +import os +from unittest.mock import patch + +import numpy +from PIL import Image + +from dw.pipeline_processors.chain import run_chain +from dw.result import AudioVideo, Result +from dw.runs import MANIFEST_FILE_NAME, recorded_shots +from dw.shots import ( + carried_shots, + named_shots, + remeasured_shots, + rescaled_shots, + shot_reference_names, + shot_record, + shots_for_file, + step_shots, +) +from dw.tasks.audio_utils import slice_audio +from dw.tasks.concat_videos import concat_videos +from dw.tasks.dissolve_videos import dissolve_videos +from dw.tasks.pair_audio import pair_audio + +REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +DW_ROOT = os.path.join(REPO_ROOT, "dw") + + +def frames(count, color=(0, 0, 0)): + return [Image.new("RGB", (4, 4), color) for _ in range(count)] + + +def audio_video(num_frames, level, fps=4, sample_rate=100): + samples = int(num_frames / fps * sample_rate) + audio = numpy.full((2, samples), float(level), dtype=numpy.float32) + return AudioVideo(frames(num_frames), audio, sample_rate, fps=fps) + + +# --------------------------------------------------------------------------- +# 1. Every AudioVideo(...) construction site in dw/ is accounted for +# --------------------------------------------------------------------------- + +# (relative path, enclosing function name) -> (decision, expected call count) +# +# A decision is one of: +# "populates" - builds a fresh shots list for a video it joined +# "carries" - copies an unchanged input's shots across (carried_shots) +# "rescales" - stretches an input's shots to a new frame count (rescaled_shots) +# "remeasures" - keeps the frame side, re-measures the sample side (remeasured_shots) +# "none" - the video is not joined from named inputs; no shots kwarg at all +EXPECTED_SITES = { + ("dw/tasks/concat_videos.py", "concat_videos"): ("populates", 1), + ("dw/tasks/dissolve_videos.py", "dissolve_videos"): ("populates", 1), + ("dw/pipeline_processors/chain.py", "run_chain"): ("populates", 2), + ("dw/tasks/task.py", "_per_frame"): ("carries", 1), + ("dw/tasks/stabilize.py", "stabilize_video"): ("carries", 1), + ("dw/tasks/interpolate_frames.py", "interpolate_frames"): ("rescales", 1), + ("dw/tasks/pair_audio.py", "pair_audio"): ("remeasures", 1), + ("dw/tasks/video_utils.py", "_decode_audio_video"): ("none", 1), + ("dw/result.py", "pair_audio_with_frames"): ("none", 1), +} + +_SHOTS_HELPER_BY_DECISION = { + "populates": None, # builds its own list - no single shared helper + "carries": "carried_shots", + "rescales": "rescaled_shots", + "remeasures": "remeasured_shots", +} + + +def _iter_python_files(root): + for dirpath, _dirnames, filenames in os.walk(root): + for name in filenames: + if name.endswith(".py"): + yield os.path.join(dirpath, name) + + +def _enclosing_function(tree, call_node): + """The innermost function/method def that contains `call_node`, or None.""" + best = None + for node in ast.walk(tree): + if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)): + if node.lineno <= call_node.lineno and getattr( + node, "end_lineno", node.lineno + ) >= getattr(call_node, "end_lineno", call_node.lineno): + if best is None or node.lineno > best.lineno: + best = node + return best.name if best else None + + +def _is_audio_video_call(call_node): + func = call_node.func + if isinstance(func, ast.Name): + return func.id == "AudioVideo" + if isinstance(func, ast.Attribute): + return func.attr == "AudioVideo" + return False + + +def _find_audio_video_calls(): + """Every `AudioVideo(...)` call in dw/, as (relative path, function, node).""" + found = [] + for path in _iter_python_files(DW_ROOT): + with open(path, encoding="utf-8") as handle: + source = handle.read() + try: + tree = ast.parse(source, filename=path) + except SyntaxError: + continue + relative = os.path.relpath(path, REPO_ROOT).replace(os.sep, "/") + for node in ast.walk(tree): + if isinstance(node, ast.Call) and _is_audio_video_call(node): + function_name = _enclosing_function(tree, node) + found.append((relative, function_name, node)) + return found + + +def test_every_audio_video_constructor_site_is_decided(): + """Every `AudioVideo(` call site in dw/ is one of the sites this module + decided for (dw/shots.py's own module docstring: "tests/test_shots.py + fails on a constructor site nobody decided for.").""" + calls = _find_audio_video_calls() + + counts = {} + for relative, function_name, _node in calls: + counts[(relative, function_name)] = counts.get((relative, function_name), 0) + 1 + + unexpected = sorted(set(counts) - set(EXPECTED_SITES)) + assert not unexpected, ( + "AudioVideo(...) constructed at a site with no recorded decision - " + "decide for it in tests/test_shots.py and dw/shots.py: " + f"{unexpected}" + ) + + missing = sorted(set(EXPECTED_SITES) - set(counts)) + assert not missing, ( + "A previously-decided AudioVideo(...) site has disappeared - update " + f"EXPECTED_SITES in tests/test_shots.py: {missing}" + ) + + for site, (_decision, expected_count) in EXPECTED_SITES.items(): + assert counts[site] == expected_count, ( + f"{site} now constructs AudioVideo {counts[site]} time(s), " + f"expected {expected_count}" + ) + + +def test_every_non_none_site_passes_shots_explicitly(): + """A site that decided its video is (or is not) joined from named inputs + says so with an explicit `shots=` keyword - "none" is the one decision + that is allowed to omit it entirely.""" + calls = _find_audio_video_calls() + for relative, function_name, node in calls: + decision, _count = EXPECTED_SITES[(relative, function_name)] + keyword_names = {kw.arg for kw in node.keywords if kw.arg is not None} + if decision == "none": + continue + assert "shots" in keyword_names, ( + f"{relative}:{node.lineno} ({function_name}, decision={decision!r}) " + "does not pass shots= explicitly" + ) + + +# --------------------------------------------------------------------------- +# 2 & 3. concat_videos populates shots that partition its output, and an +# overrun is measured rather than derived +# --------------------------------------------------------------------------- + + +class TestConcatVideosShots: + def test_shots_partition_frames_and_samples(self): + videos = [audio_video(4, 1), audio_video(6, 2), audio_video(5, 3)] + + result = concat_videos(videos, fps=4) + + assert [shot["name"] for shot in result.shots] == [ + "video 1", + "video 2", + "video 3", + ] + starts = [shot["start_frame"] for shot in result.shots] + counts = [shot["num_frames"] for shot in result.shots] + assert starts == [0, 4, 10] + assert sum(counts) == len(result.frames) + + sample_starts = [shot["start_sample"] for shot in result.shots] + sample_counts = [shot["num_samples"] for shot in result.shots] + assert sample_starts[0] == 0 + assert sum(sample_counts) == result.audio.shape[1] + + def test_shot_reference_names_name_shot_at_members(self): + videos = [audio_video(4, 1), audio_video(4, 2)] + + result = concat_videos(videos, fps=4) + named = named_shots( + result.shots, + shot_reference_names( + ["previous_result:shot@a.frames", "previous_result:shot@b.frames"] + ), + ) + + # The prefix strips only "previous_result:" - the member keeps its + # "shot@" marker, which is how a for_each member is named elsewhere + assert [shot["name"] for shot in named] == ["shot@a", "shot@b"] + + def test_no_audio_input_leaves_sample_fields_none(self): + result = concat_videos([frames(4), frames(3)]) + + assert result.shots is not None + for shot in result.shots: + assert shot["start_sample"] is None + assert shot["num_samples"] is None + + def test_an_overrun_track_is_measured_not_derived(self): + """One input's audio runs 267 samples longer than its frames alone + would imply - concat_videos records the measured length of the + joined track, not a count derived from frame/fps arithmetic.""" + fps, sample_rate = 4, 100 + first = audio_video(4, 1, fps=fps, sample_rate=sample_rate) + second = audio_video(4, 2, fps=fps, sample_rate=sample_rate) + overrun = 267 + second.audio = numpy.concatenate( + [ + second.audio, + numpy.full((2, overrun), 2.0, dtype=numpy.float32), + ], + axis=1, + ) + + result = concat_videos([first, second], fps=fps) + + frame_derived_samples = int(4 / fps * sample_rate) + second_shot = result.shots[1] + assert second_shot["num_samples"] == frame_derived_samples + overrun + assert second_shot["num_samples"] == result.audio.shape[1] - int( + 4 / fps * sample_rate + ) + + +# --------------------------------------------------------------------------- +# 4. dissolve_videos +# --------------------------------------------------------------------------- + + +class TestDissolveVideosShots: + def test_shots_partition_frames_with_overlap_recorded(self): + result = dissolve_videos([frames(10), frames(10), frames(10)], 3) + + starts = [shot["start_frame"] for shot in result.shots] + counts = [shot["num_frames"] for shot in result.shots] + assert sum(counts) == len(result.frames) + assert starts[0] == 0 + assert starts[1] == starts[0] + counts[0] + assert starts[2] == starts[1] + counts[1] + + assert "overlap_frames" not in result.shots[0] + assert result.shots[1]["overlap_frames"] == 3 + assert result.shots[2]["overlap_frames"] == 3 + + def test_sample_fields_partition_the_track_when_audio_is_present(self): + def clip(level): + samples = 250 + audio = numpy.full((2, samples), float(level), dtype=numpy.float32) + return AudioVideo(frames(10), audio, 100, fps=4) + + result = dissolve_videos([clip(1), clip(2), clip(3)], 3, fps=4) + + sample_starts = [shot["start_sample"] for shot in result.shots] + sample_counts = [shot["num_samples"] for shot in result.shots] + assert sample_starts[0] == 0 + assert sum(sample_counts) == result.audio.shape[1] + + +# --------------------------------------------------------------------------- +# 5. run_chain +# --------------------------------------------------------------------------- + + +class _FakePipeline: + def __init__(self, output_factory): + self.output_factory = output_factory + self.calls = [] + + def _run_once(self, arguments): + self.calls.append(arguments) + return self.output_factory(arguments, len(self.calls) - 1) + + +def _video_output(arguments, index, num_frames=4): + color = (50 * index % 256, 100, 150) + made = frames(num_frames, color) + if "image" in arguments: + made[0] = arguments["image"] + from types import SimpleNamespace + + return SimpleNamespace(frames=[made]) + + +class TestRunChainShots: + def test_segments_are_named_and_partition_the_output(self): + pipeline = _FakePipeline(_video_output) + + result = run_chain(pipeline, {"segments": 3, "trim_frames": 1}, {}) + + assert [shot["name"] for shot in result.shots] == [ + "segment 1", + "segment 2", + "segment 3", + ] + counts = [shot["num_frames"] for shot in result.shots] + starts = [shot["start_frame"] for shot in result.shots] + assert counts == [4, 3, 3] + assert starts == [0, 4, 7] + assert sum(counts) == len(result.frames) + + def test_video_only_chain_has_no_sample_fields(self): + pipeline = _FakePipeline(_video_output) + + result = run_chain(pipeline, {"segments": 2}, {}) + + for shot in result.shots: + assert shot["start_sample"] is None + assert shot["num_samples"] is None + + +# --------------------------------------------------------------------------- +# 6. pair_audio recomputes the sample fields +# --------------------------------------------------------------------------- + + +class TestPairAudioShots: + def test_recomputes_sample_fields_against_the_new_track(self): + fps, rate = 4, 100 + shots = [ + shot_record("a", 0, 4, start_sample=999, num_samples=999), + shot_record("b", 4, 4, start_sample=999, num_samples=999), + ] + video = AudioVideo(frames(8), None, None, fps=fps, shots=shots) + new_track = numpy.zeros((2, 250), dtype=numpy.float32) + + paired = pair_audio(video, new_track, sample_rate=rate) + + # Frame side is untouched + assert [shot["start_frame"] for shot in paired.shots] == [0, 4] + assert [shot["num_frames"] for shot in paired.shots] == [4, 4] + + assert paired.shots[0]["start_sample"] == round(0 / fps * rate) + assert paired.shots[1]["start_sample"] == round(4 / fps * rate) + # The last shot ends at the waveform's own length + last = paired.shots[-1] + assert last["start_sample"] + last["num_samples"] == new_track.shape[1] + assert sum(shot["num_samples"] for shot in paired.shots) == new_track.shape[1] + + def test_no_frame_rate_clears_the_sample_side(self): + shots = [shot_record("a", 0, 4, start_sample=1, num_samples=2)] + video = AudioVideo(frames(4), None, None, fps=None, shots=shots) + new_track = numpy.zeros((2, 100), dtype=numpy.float32) + + paired = pair_audio(video, new_track, sample_rate=100) + + assert paired.shots[0]["start_sample"] is None + assert paired.shots[0]["num_samples"] is None + + def test_remeasured_shots_directly(self): + """dw.shots.remeasured_shots in isolation, the function pair_audio calls.""" + shots = [shot_record("a", 0, 5), shot_record("b", 5, 5)] + + remeasured = remeasured_shots(shots, fps=5, sample_rate=10, total_samples=20) + + assert remeasured[0]["start_sample"] == 0 + assert remeasured[1]["start_sample"] == 10 + assert remeasured[1]["num_samples"] == 10 + + assert remeasured_shots(shots, fps=None, sample_rate=10, total_samples=20) == [ + {**shot, "start_sample": None, "num_samples": None} for shot in shots + ] + assert remeasured_shots(None, fps=5, sample_rate=10, total_samples=20) is None + + +# --------------------------------------------------------------------------- +# 7. slice_audio drops shots +# --------------------------------------------------------------------------- + + +class TestSliceAudioDropsShots: + def test_slicing_a_joined_video_returns_a_track_with_no_shots(self): + shots = [shot_record("a", 0, 4, 0, 100), shot_record("b", 4, 4, 100, 100)] + video = AudioVideo( + frames(8), + numpy.zeros((2, 200), dtype=numpy.float32), + 100, + fps=4, + shots=shots, + ) + + sliced = slice_audio(video, start_frame=0, num_frames=4, fps=4) + + assert getattr(sliced, "shots", None) is None + + +# --------------------------------------------------------------------------- +# 8. interpolate_frames / rescaled_shots +# --------------------------------------------------------------------------- + + +class TestRescaledShots: + """interpolate_frames calls rescaled_shots(source_shots, multiplier) + directly (dw/tasks/interpolate_frames.py); driving it through the real + task needs a loaded RIFE model, so the function is exercised here.""" + + def test_frame_counts_and_partition_after_doubling(self): + shots = [shot_record("a", 0, 5), shot_record("b", 5, 5)] + + rescaled = rescaled_shots(shots, multiplier=2) + + # (10 - 1) * 2 + 1 = 19 total frames + assert rescaled[0]["start_frame"] == 0 + assert rescaled[1]["start_frame"] == 10 + last_end = rescaled[1]["start_frame"] + rescaled[1]["num_frames"] + assert last_end == 19 + + def test_sample_fields_are_cleared(self): + shots = [shot_record("a", 0, 5, start_sample=0, num_samples=50)] + + rescaled = rescaled_shots(shots, multiplier=2) + + assert rescaled[0]["start_sample"] is None + assert rescaled[0]["num_samples"] is None + + def test_overlap_frames_scales_with_the_multiplier(self): + shots = [ + shot_record("a", 0, 5), + {**shot_record("b", 5, 5), "overlap_frames": 3}, + ] + + rescaled = rescaled_shots(shots, multiplier=2) + + assert rescaled[1]["overlap_frames"] == 6 + + def test_empty_input_returns_none(self): + assert rescaled_shots(None, multiplier=2) is None + assert rescaled_shots([], multiplier=2) is None + + +class TestCarriedShots: + def test_deep_copies_so_the_source_is_unaffected(self): + source = AudioVideo( + frames(4), None, None, shots=[shot_record("a", 0, 4, 0, 10)] + ) + + copied = carried_shots(source) + copied[0]["num_frames"] = 999 + + assert source.shots[0]["num_frames"] == 4 + + def test_no_shots_returns_none(self): + source = AudioVideo(frames(4), None, None) + + assert carried_shots(source) is None + + +# --------------------------------------------------------------------------- +# 9. Round trip: join -> Result.save -> manifest -> recorded_shots +# --------------------------------------------------------------------------- + + +class TestRoundTrip: + """Exercises the same functions Workflow.run does (Result.save, + dw.shots.step_shots, and dw.runs.recorded_shots) with a real join's + output, without spinning up a full model-backed Workflow.run - the + workflow's own step loop (dw/workflow.py) glues these together with no + logic of its own beyond what is called here.""" + + def test_shots_survive_save_and_recorded_shots(self, tmp_path): + videos = [audio_video(4, 1), audio_video(6, 2)] + joined = concat_videos(videos, fps=4) + + result = Result({"content_type": "video/mp4", "save": True}) + result.add_result(joined) + + run_dir = tmp_path / "concat-demo" / "20260923-120000-abcdef01" + run_dir.mkdir(parents=True) + # The actual mux is exercised by dw's own result-saving tests + # (tests/test_concat_videos.py's TestPreviousResultChainAudioFit + # patches the same three names); what this test pins is what + # Result.save records in saved_shots and how the manifest step reads + # it back, not the codec. + with ( + patch("dw.result.encode_video"), + patch("dw.result.export_to_video"), + patch("dw.result.is_av_available", return_value=True), + ): + saved_files = result.save(str(run_dir), "concat-demo-join.0") + + assert saved_files + assert result.saved_shots + + relative_files = [os.path.relpath(path, run_dir) for path in saved_files] + manifest_shots = step_shots(result.saved_shots, saved_files) + assert manifest_shots is not None + assert [shot["name"] for shot in manifest_shots] == ["video 1", "video 2"] + assert sum(shot["num_frames"] for shot in manifest_shots) == len(joined.frames) + + manifest = { + "steps": [ + { + "step": "join", + "files": relative_files, + "shots": manifest_shots, + } + ] + } + with open(run_dir / MANIFEST_FILE_NAME, "w") as handle: + json.dump(manifest, handle) + + relative_path = f"concat-demo/20260923-120000-abcdef01/{relative_files[0]}" + read_back = recorded_shots(str(tmp_path), relative_path) + + assert read_back == manifest_shots + + def test_recorded_shots_is_none_outside_a_run_directory(self, tmp_path): + # The flat layout - no run id segment - has no manifest to read shots from + assert recorded_shots(str(tmp_path), "workflow/still.png") is None + + def test_recorded_shots_is_none_when_the_step_was_reused(self, tmp_path): + run_dir = tmp_path / "wf" / "20260923-120000-abcdef01" + run_dir.mkdir(parents=True) + manifest = { + "steps": [ + { + "step": "join", + "files": ["out.mp4"], + "shots": [shot_record("a", 0, 4, 0, 10)], + "reused": True, + } + ] + } + with open(run_dir / MANIFEST_FILE_NAME, "w") as handle: + json.dump(manifest, handle) + + assert ( + recorded_shots(str(tmp_path), "wf/20260923-120000-abcdef01/out.mp4") is None + ) + + +# --------------------------------------------------------------------------- +# 10. shots_for_file with multi-file entries +# --------------------------------------------------------------------------- + + +class TestShotsForFile: + def test_multi_file_entry_returns_only_that_files_shots_without_file_key(self): + shots = [ + {**shot_record("a", 0, 4, 0, 10), "file": "one.mp4"}, + {**shot_record("b", 0, 5, 0, 12), "file": "two.mp4"}, + ] + + own = shots_for_file(shots, "two.mp4", ["one.mp4", "two.mp4"]) + + assert len(own) == 1 + assert own[0]["name"] == "b" + assert "file" not in own[0] + + def test_single_file_entry_matches_by_the_step_files_list(self): + shots = [shot_record("a", 0, 4, 0, 10)] + + assert shots_for_file(shots, "solo.mp4", ["solo.mp4"]) == shots + assert shots_for_file(shots, "other.mp4", ["solo.mp4"]) is None + + def test_no_shots_returns_none(self): + assert shots_for_file(None, "solo.mp4", ["solo.mp4"]) is None + assert shots_for_file([], "solo.mp4", ["solo.mp4"]) is None From 2c802c281a1c263c77a497d4ece7826e3630f4a1 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 22:05:37 -0500 Subject: [PATCH 070/181] test(engine): #385 - shots through Workflow.run's manifest, and the seams route's fallback Co-Authored-By: Claude Opus 5.5 --- tests/test_server.py | 58 ++++++++++++++++++++++++++++ tests/test_shots.py | 90 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 148 insertions(+) diff --git a/tests/test_server.py b/tests/test_server.py index a8608cdc..5d27691e 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -1563,6 +1563,64 @@ def test_gallery_frames_returns_the_moments_asked_for(server, tmp_path): assert _png_of(body["tiles"][0]).size == (32, 16) +def test_gallery_frames_seams_read_a_joined_outputs_recorded_shots(server, tmp_path): + """#385: `seams` without `boundaries` was a 400 - the file carried no + seams of its own. An output whose run recorded shots for it now answers + from the manifest, named as recorded; `media.shots` in the metadata is + the same list. A file whose run recorded none still needs `boundaries`.""" + import json + + from tests.test_media_frames import write_ramp_mp4 + + with server(success_script) as client: + run = tmp_path / "outputs" / "cut" / "20260924-000000-0123abcd" + run.mkdir(parents=True) + write_ramp_mp4(run / "cut.mp4", frames=24, fps=6, width=64, height=32) + write_ramp_mp4(run / "other.mp4", frames=24, fps=6, width=64, height=32) + shots = [ + { + "name": "shot@wide", + "start_frame": 0, + "num_frames": 10, + "start_sample": None, + "num_samples": None, + }, + { + "name": "shot@close", + "start_frame": 10, + "num_frames": 14, + "start_sample": None, + "num_samples": None, + }, + ] + (run / "manifest.json").write_text( + json.dumps( + { + "steps": [ + {"step": "cut", "files": ["cut.mp4"], "shots": shots}, + {"step": "other", "files": ["other.mp4"]}, + ] + } + ) + ) + name = "cut/20260924-000000-0123abcd/cut.mp4" + + response = client.get(f"/api/gallery/{name}/frames", params={"seams": "true"}) + assert response.status_code == 200, response.text + (tile,) = response.json()["tiles"] + assert tile["label"] == "seam 1: shot@wide | shot@close" + + metadata = client.get(f"/api/gallery/{name}/metadata").json() + assert metadata["media"]["shots"] == shots + + refused = client.get( + "/api/gallery/cut/20260924-000000-0123abcd/other.mp4/frames", + params={"seams": "true"}, + ) + assert refused.status_code == 400 + assert "boundaries" in refused.json()["detail"] + + def test_gallery_frames_crop_names_the_same_source_region_at_any_max_dimension( server, tmp_path ): diff --git a/tests/test_shots.py b/tests/test_shots.py index 886779eb..7d6afe4c 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -589,3 +589,93 @@ def test_single_file_entry_matches_by_the_step_files_list(self): def test_no_shots_returns_none(self): assert shots_for_file(None, "solo.mp4", ["solo.mp4"]) is None assert shots_for_file([], "solo.mp4", ["solo.mp4"]) is None + + +# 11. Workflow.run: the manifest entry and step_end carry the shots, named by +# the `gather:shot` the step wrote + + +class _ShotResult: + """Stands in for dw.result.Result: writes one file and, for the join, + reports the shots its video carried the way Result.save does.""" + + def __init__(self, shots=None): + self.result_list = [] + self.saved_files = [] + self.saved_shots = {} + self.selected = None + self._shots = shots + + def save(self, output_dir, base_name): + os.makedirs(output_dir, exist_ok=True) + path = os.path.join(output_dir, f"{base_name}.mp4") + with open(path, "wb") as handle: + handle.write(b"") + self.saved_files = [path] + if self._shots: + self.saved_shots = {path: self._shots} + return self.saved_files + + +def test_workflow_run_names_the_joined_shots_by_their_members(tmp_path): + """The real Workflow.run path: `gather:shot` expands to + `previous_result:shot@` references, and the join's positional + shots are renamed from them in the manifest, manifest.json and + step_end - and read back by recorded_shots.""" + from dw.events import RunContext + from dw.step import Step + from dw.step_cache import step_cache + from dw.workflow import Workflow + + joined = [ + shot_record("video 1", 0, 24, 0, 16000), + shot_record("video 2", 24, 30, 16000, 20267), + ] + definition = { + "id": "shots_round_trip", + "steps": [ + { + "name": "shot", + "for_each": [{"name": "wide"}, {"name": "close"}], + "task": {"command": "stabilize_video", "arguments": {"clip": "x"}}, + "result": {"content_type": "video/mp4"}, + }, + { + "name": "cut", + "task": { + "command": "concat_videos", + "arguments": {"videos": "gather:shot"}, + }, + "result": {"content_type": "video/mp4"}, + }, + ], + } + + def fake_step_run(self, previous_results, previous_pipelines, step_action): + return _ShotResult(joined if self.name == "cut" else None) + + step_cache.clear() + events = [] + workflow = Workflow(definition, str(tmp_path), str(tmp_path / "shots.json")) + with patch.object(Step, "run", fake_step_run): + workflow.run({}, context=RunContext(on_event=events.append)) + + (entry,) = [e for e in workflow.manifest if e["step"] == "cut"] + names = [shot["name"] for shot in entry["shots"]] + assert names == ["shot@wide", "shot@close"] + assert [shot["num_samples"] for shot in entry["shots"]] == [16000, 20267] + + (step_end,) = [ + e for e in events if e["event"] == "step_end" and e.get("step") == "cut" + ] + assert step_end["shots"] == entry["shots"] + + with open(os.path.join(workflow._run_dir, "manifest.json")) as handle: + manifest = json.load(handle) + (written,) = [e for e in manifest["steps"] if e["step"] == "cut"] + assert written["shots"] == entry["shots"] + + relative = os.path.relpath( + os.path.join(workflow._run_dir, written["files"][0]), str(tmp_path) + ).replace(os.sep, "/") + assert recorded_shots(str(tmp_path), relative) == entry["shots"] From 8c49c69f7e36c9c29340e91cd968e60b3d4e5028 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 22:20:45 -0500 Subject: [PATCH 071/181] fix(engine): #390 - name a joined shot by its file, not its resolved absolute path concat_videos/dissolve_videos named a string-input shot by the exact value they received, but by the time the task runs an asset:/output: reference has already been resolved to its absolute server path (realize_args), so it leaked server layout into the manifest, get_gallery_metadata's media.shots and the seams route's labels. video_names() now trims a string input to its file name (os.path.basename), which is shared by concat_videos and dissolve_videos and also fixes the sample-rate-mismatch warning's per-video labels. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/concat_videos.py | 9 +++++++-- tests/test_shots.py | 34 ++++++++++++++++++++++++++++++++++ 2 files changed, 41 insertions(+), 2 deletions(-) diff --git a/dw/tasks/concat_videos.py b/dw/tasks/concat_videos.py index aec91d8c..07b7e7b3 100644 --- a/dw/tasks/concat_videos.py +++ b/dw/tasks/concat_videos.py @@ -9,6 +9,7 @@ """ import logging +import os from ..events import emit_warning from ..result import AudioVideo @@ -33,10 +34,14 @@ def video_names(videos): A caller passes a path, or a previous step's result; only the path says anything by itself, so the rest are named by position - which is what a six-entry `shots` list needs to be actionable ("24000 then 32000" does - not say which entry to fix). + not say which entry to fix). By the time this runs, an `asset:`/`output:` + reference has already been resolved to its absolute path on this server + (#390) - naming a shot by that path leaked server layout onto a consumer + surface, so a path is trimmed to its file name, the one part that means + anything off this box. """ return [ - original if isinstance(original, str) else f"video {index + 1}" + os.path.basename(original) if isinstance(original, str) else f"video {index + 1}" for index, original in enumerate(videos) ] diff --git a/tests/test_shots.py b/tests/test_shots.py index 7d6afe4c..62705a62 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -230,6 +230,24 @@ def test_no_audio_input_leaves_sample_fields_none(self): assert shot["start_sample"] is None assert shot["num_samples"] is None + def test_asset_literal_shots_are_named_by_file_not_absolute_path(self): + """An `asset:`/`output:` reference is resolved to its absolute + server path before the task ever runs (dw/workflow.py's + realize_args), so naming a shot by the string the join received + leaked that path onto every consumer of the shots - the manifest, + get_gallery_metadata, and the seams route's labels (#390). Only the + file name means anything off this box.""" + resolved = "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot1-incident.mp4" + videos = [resolved, audio_video(4, 2)] + + with patch( + "dw.tasks.concat_videos.load_audio_video", return_value=audio_video(4, 1) + ): + result = concat_videos(videos, fps=4) + + assert result.shots[0]["name"] == "ep3-shot1-incident.mp4" + assert "/" not in result.shots[0]["name"] + def test_an_overrun_track_is_measured_not_derived(self): """One input's audio runs 267 samples longer than its frames alone would imply - concat_videos records the measured length of the @@ -289,6 +307,22 @@ def clip(level): assert sample_starts[0] == 0 assert sum(sample_counts) == result.audio.shape[1] + def test_asset_literal_shots_are_named_by_file_not_absolute_path(self): + """Same leak as concat_videos (#390): a resolved asset: reference + arrives here as an absolute server path, and only its file name + belongs on a consumer-facing shot name.""" + resolved = "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot2-reply.mp4" + videos = [resolved, frames(10)] + + with patch( + "dw.tasks.dissolve_videos.load_audio_video", + return_value=frames(10), + ): + result = dissolve_videos(videos, 3) + + assert result.shots[0]["name"] == "ep3-shot2-reply.mp4" + assert "/" not in result.shots[0]["name"] + # --------------------------------------------------------------------------- # 5. run_chain From 9c7899341afe5677070741471848023e5b7bc807 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 22:21:45 -0500 Subject: [PATCH 072/181] style(engine): #390 - ruff format Co-Authored-By: Claude Sonnet 5 --- dw/tasks/concat_videos.py | 4 +++- tests/test_shots.py | 8 ++++++-- 2 files changed, 9 insertions(+), 3 deletions(-) diff --git a/dw/tasks/concat_videos.py b/dw/tasks/concat_videos.py index 07b7e7b3..a433fc2c 100644 --- a/dw/tasks/concat_videos.py +++ b/dw/tasks/concat_videos.py @@ -41,7 +41,9 @@ def video_names(videos): anything off this box. """ return [ - os.path.basename(original) if isinstance(original, str) else f"video {index + 1}" + os.path.basename(original) + if isinstance(original, str) + else f"video {index + 1}" for index, original in enumerate(videos) ] diff --git a/tests/test_shots.py b/tests/test_shots.py index 62705a62..3c447330 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -237,7 +237,9 @@ def test_asset_literal_shots_are_named_by_file_not_absolute_path(self): leaked that path onto every consumer of the shots - the manifest, get_gallery_metadata, and the seams route's labels (#390). Only the file name means anything off this box.""" - resolved = "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot1-incident.mp4" + resolved = ( + "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot1-incident.mp4" + ) videos = [resolved, audio_video(4, 2)] with patch( @@ -311,7 +313,9 @@ def test_asset_literal_shots_are_named_by_file_not_absolute_path(self): """Same leak as concat_videos (#390): a resolved asset: reference arrives here as an absolute server path, and only its file name belongs on a consumer-facing shot name.""" - resolved = "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot2-reply.mp4" + resolved = ( + "/home/don/diffusers-workspace/common/assets/qa-cast/ep3-shot2-reply.mp4" + ) videos = [resolved, frames(10)] with patch( From 1c5ff80f7b6890bc471f4f0475d17d56b1b8ac21 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 073/181] test(security): token, Origin/Host and ungated static routes Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_auth.py | 457 ++++++++++++++++++++++++++++++++++++ 1 file changed, 457 insertions(+) create mode 100644 tests/test_security_auth.py diff --git a/tests/test_security_auth.py b/tests/test_security_auth.py new file mode 100644 index 00000000..ebbbd4b1 --- /dev/null +++ b/tests/test_security_auth.py @@ -0,0 +1,457 @@ +"""The HTTP server's own boundary: token, Origin/Host, the ungated routes. + +tests/test_server.py covers the happy shapes - a missing or wrong token is a +401, a foreign Origin a 403, a foreign Host a 400, the query-token allowance +is per route - and tests/test_serve_main.py the `--mcp` refusal on 0.0.0.0. +This file is the adversarial side of the same boundary: path spellings that +might slip past a prefix check, credentials in the wrong place, Origin and +Host values built to fool a hostname comparison, and the routes that are +ungated *on purpose* (`/outputs`, `/inputs`, `/exports`), pinned to serving +their own roots and nothing else. + +Every probe target is under `tmp_path`; "outside" means a sibling directory +the server was never told about. +""" + +import json + +import pytest +from fastapi.testclient import TestClient + +from dw.server.app import create_app +from dw.server.jobs import JobManager + +from .test_server import ScriptedWorkerManager, success_script, valid_workflow + +TOKEN = "s3cr3t-token" +SECRET = "outside-the-roots-probe" + + +@pytest.fixture +def layout(tmp_path): + """A workspace-shaped tree plus a sibling directory holding a secret.""" + paths = { + "workflows": tmp_path / "ws" / "workflows", + "outputs": tmp_path / "ws" / "outputs", + "assets": tmp_path / "ws" / "assets", + "prompts": tmp_path / "ws" / "prompts", + "outside": tmp_path / "outside", + } + for path in paths.values(): + path.mkdir(parents=True) + (paths["workflows"] / "Basic.json").write_text(json.dumps(valid_workflow("b"))) + (paths["outputs"] / "run.png").write_bytes(b"generated") + (paths["assets"] / "iris.png").write_bytes(b"asset") + (paths["outside"] / "secret.txt").write_text(SECRET) + (paths["outside"] / "secret.png").write_text(SECRET) + paths["root"] = tmp_path / "ws" + return paths + + +@pytest.fixture +def make_client(layout, tmp_path): + def make(token=None, host="127.0.0.1", base_url="http://localhost"): + manager = JobManager( + str(layout["outputs"]), + worker_manager=ScriptedWorkerManager(success_script), + history_path=str(tmp_path / "jobs.sqlite"), + ) + app = create_app( + workflow_dir=str(layout["workflows"]), + output_dir=str(layout["outputs"]), + job_manager=manager, + prompt_dir=str(layout["prompts"]), + asset_dir=str(layout["assets"]), + workspace=str(layout["root"]), + token=token, + host=host, + ) + return TestClient(app, base_url=base_url) + + return make + + +def _jobs(client, headers=None): + listing = client.get("/api/jobs", headers=headers or {}).json() + return listing.get("jobs", listing) if isinstance(listing, dict) else listing + + +def _leaks(response): + return response.status_code == 200 and SECRET in response.text + + +# ---------------------------------------------------------------- the token + + +class TestTheTokenGate: + @pytest.mark.parametrize( + "path", + [ + "/api/health", + "/api/health/", + "/api/workflows", + "/api/jobs", + "/api/server", + "/api/gallery", + "/api/prompts", + "/api/assets", + "/%61pi/health", + "/api/%68ealth", + "/api/./health", + "/mcp", + "/mcp/", + ], + ) + def test_every_api_spelling_needs_the_token(self, make_client, path): + with make_client(token=TOKEN) as client: + response = client.get(path) + assert response.status_code == 401, (path, response.status_code) + + @pytest.mark.parametrize("path", ["//api/health", "/API/health", "/api"]) + def test_a_spelling_that_misses_the_prefix_reaches_no_api_route( + self, make_client, path + ): + """If the gate's prefix check misses a spelling, the router must miss + it too - otherwise an ungated spelling of a gated route exists.""" + with make_client(token=TOKEN) as client: + response = client.get(path) + assert response.status_code == 404, (path, response.status_code) + + @pytest.mark.parametrize( + "authorization", + [ + "", + "Bearer", + "Bearer ", + f"Basic {TOKEN}", + f"Token {TOKEN}", + f"{TOKEN}", + f"Bearer {TOKEN}x", + f"Bearer x{TOKEN}", + f"Bearer {TOKEN[:-1]}", + f"Bearer {TOKEN.upper()}", + f"Bearer {TOKEN}\x00", + ], + ) + def test_anything_but_the_exact_bearer_token_is_a_401( + self, make_client, authorization + ): + with make_client(token=TOKEN) as client: + try: + response = client.get( + "/api/health", headers={"Authorization": authorization} + ) + except Exception: + # httpx refuses to send some header values at all - nothing + # reached the server, which is the outcome the test wants + return + assert response.status_code == 401 + + def test_the_scheme_is_case_insensitive_but_the_token_is_not(self, make_client): + with make_client(token=TOKEN) as client: + assert ( + client.get( + "/api/health", headers={"Authorization": f"bearer {TOKEN}"} + ).status_code + == 200 + ) + + @pytest.mark.parametrize( + "method, path", + [ + ("GET", "/api/health"), + ("GET", "/api/workflows"), + ("GET", "/api/jobs"), + ("POST", "/api/validate"), + ("DELETE", "/api/prompts/x"), + ], + ) + def test_the_query_token_is_refused_where_it_is_not_allowed( + self, make_client, method, path + ): + """?token= exists for /EventSource GETs on marked routes only; a + route that is not marked must not accept it.""" + with make_client(token=TOKEN) as client: + response = client.request(method, f"{path}?token={TOKEN}", json={}) + assert response.status_code == 401 + + def test_the_query_token_is_get_only_even_on_a_marked_route( + self, make_client, layout + ): + with make_client(token=TOKEN) as client: + assert ( + client.get(f"/api/gallery/run.png/download?token={TOKEN}").status_code + == 200 + ) + assert ( + client.delete(f"/api/gallery/run.png?token={TOKEN}").status_code == 401 + ) + assert (layout["outputs"] / "run.png").exists() + + def test_a_valid_token_does_not_excuse_a_foreign_origin(self, make_client): + """The Origin check is not an auth check the token can satisfy: a + page that somehow holds the token still may not drive the API.""" + with make_client(token=TOKEN) as client: + response = client.post( + "/api/jobs", + json={"workflow": valid_workflow()}, + headers={ + "Authorization": f"Bearer {TOKEN}", + "Origin": "https://evil.example", + }, + ) + assert response.status_code == 403 + + def test_a_refused_request_queues_nothing(self, make_client): + with make_client(token=TOKEN) as client: + client.post("/api/jobs", json={"workflow": valid_workflow()}) + assert _jobs(client, {"Authorization": f"Bearer {TOKEN}"}) == [] + + +# ---------------------------------------------------------- Origin and Host + + +class TestOriginAndHostSpoofing: + @pytest.mark.parametrize( + "origin", + [ + "http://localhost@evil.example", + "http://localhost:8765@evil.example", + "http://localhost.evil.example", + "http://127.0.0.1.evil.example", + "http://evil.example#localhost", + "http://evil.example/localhost", + "http://evil.example?localhost", + "http://evil-localhost", + "null", + "file://", + ], + ) + def test_a_lookalike_origin_is_refused(self, make_client, origin): + with make_client() as client: + response = client.post( + "/api/jobs", + json={"workflow": valid_workflow()}, + headers={"Origin": origin}, + ) + assert response.status_code == 403, origin + + @pytest.mark.parametrize( + "origin", ["http://[::1].evil.example", "http://[localhost]", "http://["] + ) + def test_an_unparseable_origin_is_not_processed(self, make_client, origin): + """urlparse raises on these; whatever the middleware answers, the + request must not reach a route.""" + app = make_client().app + with TestClient( + app, base_url="http://localhost", raise_server_exceptions=False + ) as client: + response = client.post( + "/api/jobs", + json={"workflow": valid_workflow()}, + headers={"Origin": origin}, + ) + assert response.status_code >= 400 + assert _jobs(client) == [] + + @pytest.mark.xfail( + strict=True, + reason="reject_foreign_origins calls urlparse outside a try, so a " + "bracketed non-IPv6 Origin is a 500 from an unhandled ValueError " + "rather than the 403 every other refused Origin gets", + ) + def test_an_unparseable_origin_is_a_403_not_a_500(self, make_client): + app = make_client().app + with TestClient( + app, base_url="http://localhost", raise_server_exceptions=False + ) as client: + response = client.post( + "/api/jobs", + json={"workflow": valid_workflow()}, + headers={"Origin": "http://[::1].evil.example"}, + ) + assert response.status_code == 403 + + @pytest.mark.parametrize( + "host", + [ + "localhost.evil.example", + "127.0.0.1.evil.example", + "evil.example", + "evil.example:8765", + "127.0.0.2", + "0.0.0.0", + ], + ) + def test_a_lookalike_host_is_refused_on_a_loopback_bind(self, make_client, host): + with make_client() as client: + response = client.get("/api/health", headers={"Host": host}) + assert response.status_code == 400, host + + @pytest.mark.parametrize("path", ["/outputs/run.png", "/inputs/iris.png"]) + def test_the_host_check_covers_the_ungated_routes_too(self, make_client, path): + """DNS rebinding reads through whatever route answers: the static + routes need the Host check as much as /api does.""" + with make_client() as client: + assert client.get(path).status_code == 200 + assert ( + client.get(path, headers={"Host": "rebind.evil.example"}).status_code + == 400 + ) + + +# ------------------------------------------------ the ungated static routes + + +class TestUngatedRoutesServeOnlyTheirRoots: + """/outputs, /inputs and /exports take no token by design - an or + a download link cannot attach one. What they serve is therefore exactly + what anyone who can reach the port can read, so pin it: the workspace's + outputs, its asset search path, and finished exports, never a byte from + anywhere else.""" + + def test_what_they_serve_without_a_token(self, make_client): + with make_client(token=TOKEN) as client: + assert client.get("/outputs/run.png").content == b"generated" + assert client.get("/inputs/iris.png").content == b"asset" + assert client.get("/api/gallery").status_code == 401 + + @pytest.mark.parametrize( + "path", + [ + "/outputs/../outside/secret.txt", + "/outputs/..%2foutside/secret.txt", + "/outputs/%2e%2e/outside/secret.txt", + "/outputs/%2e%2e%2foutside%2fsecret.txt", + "/outputs/..%5coutside%5csecret.txt", + "/outputs/....//outside/secret.txt", + "/outputs//etc/hostname", + "/outputs/%2fetc%2fhostname", + "/outputs/~/secret.txt", + "/inputs/../outside/secret.png", + "/inputs/..%2foutside/secret.png", + "/inputs/%2e%2e%2f%2e%2e%2foutside%2fsecret.png", + "/inputs/..%5coutside%5csecret.png", + "/exports/..%2f..%2foutside.zip", + "/exports/%2e%2e.zip", + ], + ) + def test_traversal_spellings_serve_nothing_outside(self, make_client, path): + with make_client(token=TOKEN) as client: + response = client.get(path) + assert not _leaks(response), path + assert response.status_code in (400, 404, 405), (path, response.status_code) + + def test_an_absolute_path_in_the_name_is_not_joined_as_absolute( + self, make_client, layout + ): + """os.path.join(root, '/abs') is '/abs' - a name that starts with a + separator must still resolve under the root.""" + secret = layout["outside"] / "secret.txt" + with make_client(token=TOKEN) as client: + for route in ("/outputs", "/inputs"): + response = client.get(f"{route}/{secret}") + assert not _leaks(response), route + response = client.get(f"{route}/%2F{str(secret).lstrip('/')}") + assert not _leaks(response), route + + @pytest.mark.parametrize("workspace", ["..", "../outside", "%2e%2e", "/tmp", "."]) + def test_a_hostile_workspace_parameter_is_refused(self, make_client, workspace): + with make_client(token=TOKEN) as client: + response = client.get(f"/outputs/secret.txt?workspace={workspace}") + assert not _leaks(response) + assert response.status_code in (400, 404, 422) + + def test_the_gallery_download_route_is_contained_as_well(self, make_client): + with make_client(token=TOKEN) as client: + for name in ("..%2Foutside%2Fsecret.txt", "%2e%2e/outside/secret.txt"): + response = client.get(f"/api/gallery/{name}/download?token={TOKEN}") + assert not _leaks(response), name + + +# ------------------------------------------------------------- dw.serve CLI + + +@pytest.fixture +def serve(monkeypatch, tmp_path): + """dw.serve.main with the app factory, uvicorn and startup replaced.""" + import uvicorn + + import dw + import dw.serve as serve_module + from dw.server import app as app_module + + calls = {} + + def fake_create_app(**kwargs): + calls["create_app"] = kwargs + return object() + + monkeypatch.setattr(app_module, "create_app", fake_create_app) + monkeypatch.setattr(uvicorn, "run", lambda app, **kwargs: calls.update(ran=True)) + monkeypatch.setattr(dw, "startup", lambda *a, **k: calls.update(started=True)) + monkeypatch.delenv("DW_API_TOKEN", raising=False) + monkeypatch.setenv("DW_PROMPT_DIR", str(tmp_path / "prompts")) + monkeypatch.setenv("DW_ASSET_DIR", str(tmp_path / "assets")) + monkeypatch.setenv("DW_WORKSPACE", str(tmp_path / "workspace")) + monkeypatch.setenv("DW_WORKSPACE_SOURCE", "flag") + (tmp_path / "workflows").mkdir() + + def run(*argv): + calls.clear() + monkeypatch.setattr( + "sys.argv", + ["dw-serve", "--workflow-dir", str(tmp_path / "workflows"), *argv], + ) + serve_module.main() + return calls + + return run + + +class TestServeRefusesAnOpenMcpEndpoint: + @pytest.mark.parametrize( + "host", + [ + "0.0.0.0", + "::", + "10.0.0.2", + "192.168.1.5", + "gpu-box.local", + # loopback in fact, but not a name the check knows - the + # conservative direction, pinned so it stays conservative + "127.0.0.2", + "LOCALHOST", + ], + ) + def test_no_token_means_no_start(self, serve, host): + with pytest.raises(SystemExit) as exit_info: + serve("--mcp", "--host", host) + assert exit_info.value.code == 2 + + @pytest.mark.parametrize("token_args", [["--token", ""], []]) + def test_an_empty_token_is_no_token(self, serve, monkeypatch, token_args): + monkeypatch.setenv("DW_API_TOKEN", "") + with pytest.raises(SystemExit): + serve("--mcp", "--host", "0.0.0.0", *token_args) + + def test_nothing_starts_before_the_refusal(self, serve): + calls = {} + try: + calls = serve("--mcp", "--host", "0.0.0.0") + except SystemExit: + pass + assert "started" not in calls and "ran" not in calls + + def test_the_environment_token_satisfies_it_and_reaches_the_app( + self, serve, monkeypatch + ): + monkeypatch.setenv("DW_API_TOKEN", "from-env") + calls = serve("--mcp", "--host", "0.0.0.0") + assert calls["create_app"]["token"] == "from-env" + assert calls["create_app"]["mcp"] is True + + def test_the_flag_token_wins_over_the_environment(self, serve, monkeypatch): + monkeypatch.setenv("DW_API_TOKEN", "from-env") + calls = serve("--mcp", "--host", "0.0.0.0", "--token", "from-flag") + assert calls["create_app"]["token"] == "from-flag" From 851c051f716336e8c786229c705df58a232a64a8 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 074/181] test(security): trust gate refuses before any import, and trust lets it through Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_trust_gate.py | 560 ++++++++++++++++++++++++++++++ 1 file changed, 560 insertions(+) create mode 100644 tests/test_security_trust_gate.py diff --git a/tests/test_security_trust_gate.py b/tests/test_security_trust_gate.py new file mode 100644 index 00000000..3b15f5f3 --- /dev/null +++ b/tests/test_security_trust_gate.py @@ -0,0 +1,560 @@ +"""The code-execution gate, from both sides, proven by what did *not* happen. + +tests/test_workflow_trust.py shows each gate raises. What it cannot show is +*when*: a refusal that arrives as UntrustedWorkflowError after the module was +imported, its source file opened, or the Hub asked for code has already let +the code run - "refused too late" is the same as not refused. So every +untrusted case here names a probe module that exists on sys.path and would +write a marker file the moment its top-level code ran, installs an import +hook that records any attempt to find it, and fails the test on either. + +The trusted half is the other boundary: `--trust-workflows` and +`DW_TRUST_WORKFLOWS` must let each surface through to a real import, or the +documented escape hatch is broken and operators reach for worse ones. + +tests/conftest.py trusts every test by default; each test here states the +posture it runs under. +""" + +import copy +import importlib.abc +import socket +import sys +import textwrap +from unittest.mock import MagicMock + +import pytest + +from dw.security import ( + TRUST_WORKFLOWS_ENV_VAR, + UntrustedWorkflowError, + set_trust_workflows, + workflows_are_trusted, +) + +PROBE = "dw_untrusted_probe_module" + + +class _ImportRecorder(importlib.abc.MetaPathFinder): + """First on sys.meta_path: sees every import that reaches the finders, + records the ones naming the probe, and lets the real finders answer.""" + + def __init__(self, watched): + self.watched = watched + self.attempts = [] + + def find_spec(self, fullname, path=None, target=None): + if fullname.split(".")[0] == self.watched: + self.attempts.append(fullname) + return None + + +@pytest.fixture +def probe(tmp_path, monkeypatch): + """A module that would run if anything imported it, and the evidence. + + Yields an object with `.attempts` (import lookups for the probe) and + `.executed()` (whether its top-level code ran). The module is on + sys.path for real, so a gate that failed open would import it rather + than fail on ModuleNotFoundError and look like a refusal. + """ + package = tmp_path / "probe_path" + package.mkdir() + marker = tmp_path / "EXECUTED" + (package / f"{PROBE}.py").write_text( + textwrap.dedent( + f""" + import pathlib + pathlib.Path({str(marker)!r}).write_text("ran") + + class Thing: + def __init__(self, *args, **kwargs): + self.args = args + self.kwargs = kwargs + + VALUE = 7 + """ + ) + ) + monkeypatch.syspath_prepend(str(package)) + for name in list(sys.modules): + if name == PROBE or name.startswith(PROBE + "."): + monkeypatch.delitem(sys.modules, name) + + recorder = _ImportRecorder(PROBE) + monkeypatch.setattr(sys, "meta_path", [recorder, *sys.meta_path]) + recorder.executed = marker.exists + yield recorder + sys.modules.pop(PROBE, None) + + +@pytest.fixture +def untrusted(monkeypatch): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + + +@pytest.fixture +def no_network(monkeypatch): + """Any attempt to resolve or dial is a failure, not a slow test.""" + attempts = [] + + def refuse(*args, **kwargs): + attempts.append(args) + raise AssertionError("network access attempted") + + monkeypatch.setattr(socket.socket, "connect", refuse) + monkeypatch.setattr(socket, "create_connection", refuse) + monkeypatch.setattr(socket, "getaddrinfo", refuse) + return attempts + + +def assert_nothing_happened(probe, no_network=()): + assert probe.attempts == [], f"import attempted before refusal: {probe.attempts}" + assert not probe.executed(), "the probe module's code ran" + assert list(no_network) == [], "the network was touched before refusal" + + +# ------------------------------------------------------------------ surfaces + + +def _realize(arguments): + """realize_args on a copy - it converts in place, and a parametrized + dict realized once would reach the next case already a class.""" + from dw.arguments import realize_args + + arguments = copy.deepcopy(arguments) + realize_args(arguments) + return arguments + + +TYPE_REFERENCES = [ + pytest.param({"scheduler_type": f"{PROBE}.Thing"}, id="_type"), + pytest.param({"component_type": f"{PROBE}.Thing"}, id="component_type"), + pytest.param({"torch_dtype": f"{PROBE}.Thing"}, id="_dtype"), + pytest.param({"dtype": f"{PROBE}.Thing"}, id="dtype"), + pytest.param( + { + "quantization_config": { + "configuration": {"config_type": f"{PROBE}.Thing"}, + "arguments": {}, + } + }, + id="config_type", + ), + pytest.param( + {"outer": {"inner": [{"weights_dtype": f"{PROBE}.Thing"}]}}, + id="nested", + ), + pytest.param( + {"scheduler_type": f"{PROBE}.sub.Thing"}, + id="submodule", + ), +] + + +def _pipeline(configuration=None, from_pretrained=None, **blocks): + from dw.pipeline_processors.pipeline import Pipeline + + definition = { + "configuration": configuration or {}, + "from_pretrained_arguments": from_pretrained or {"model_name": "a/b"}, + "arguments": {}, + **blocks, + } + return Pipeline(definition, 0, "cpu") + + +class TestUntrustedRefusesBeforeImport: + @pytest.mark.parametrize("arguments", TYPE_REFERENCES) + def test_a_dotted_type_reference(self, untrusted, probe, no_network, arguments): + with pytest.raises(UntrustedWorkflowError, match="trust-workflows"): + _realize(arguments) + assert_nothing_happened(probe, no_network) + + @pytest.mark.parametrize( + "modules", + [ + [PROBE], + [f"{PROBE}.sub"], + # a trusted entry first must not import before the untrusted + # one is looked at: every entry is checked, then any imported + ["json", PROBE], + ], + ) + def test_pre_load_modules(self, untrusted, probe, no_network, modules): + pipeline = _pipeline(configuration={"pre_load_modules": modules}) + imported_before = set(sys.modules) + with pytest.raises(UntrustedWorkflowError, match="pre_load_modules"): + pipeline.load(shared_components={}) + assert_nothing_happened(probe, no_network) + assert PROBE not in set(sys.modules) - imported_before + + def test_pre_load_modules_are_refused_by_the_preflight_too( + self, untrusted, probe, no_network + ): + pipeline = _pipeline(configuration={"pre_load_modules": [PROBE]}) + with pytest.raises(UntrustedWorkflowError): + pipeline.check_trusted() + assert_nothing_happened(probe, no_network) + + @pytest.mark.parametrize( + "reference", + [f"constant:{PROBE}.VALUE", f"constant:{PROBE}.sub.VALUE"], + ) + def test_a_constant_reference(self, untrusted, probe, no_network, reference): + from dw.arguments import fetch_constant + + with pytest.raises(UntrustedWorkflowError): + fetch_constant(reference) + assert_nothing_happened(probe, no_network) + + def test_a_constant_inside_arguments(self, untrusted, probe, no_network): + with pytest.raises(UntrustedWorkflowError): + _realize({"sigmas": f"constant:{PROBE}.VALUE"}) + assert_nothing_happened(probe, no_network) + + @pytest.mark.parametrize( + "extra", + [ + {"trust_remote_code": True}, + {"trust_remote_code": 1}, + {"trust_remote_code": "yes"}, + {"custom_pipeline": "someone/remote-pipeline"}, + {"custom_pipeline": "PROBE_PATH"}, + ], + ) + def test_remote_code_arguments(self, untrusted, probe, no_network, tmp_path, extra): + """Neither goes through our importlib, so the proof is that + from_pretrained - the thing that would fetch and import - is never + called, and no socket is opened.""" + from dw.pipeline_processors.pipeline import load_component + + if extra.get("custom_pipeline") == "PROBE_PATH": + # a local custom pipeline is a .py diffusers would import + extra = {"custom_pipeline": str(tmp_path / "probe_path")} + component_type = MagicMock() + component_type.__name__ = "ProbePipeline" + with pytest.raises(UntrustedWorkflowError): + load_component( + "pipeline", + {"component_type": component_type}, + {"model_name": "a/b", **extra}, + "cpu", + ) + component_type.from_pretrained.assert_not_called() + assert_nothing_happened(probe, no_network) + + @pytest.mark.parametrize("key", ["trust_remote_code", "custom_pipeline"]) + def test_remote_code_in_any_nested_block_is_seen_by_the_preflight( + self, untrusted, no_network, key + ): + """A component inside a list (controlnets, loras, text encoders) is + still a from_pretrained call the gate must see.""" + pipeline = _pipeline( + controlnets=[ + { + "configuration": {}, + "from_pretrained_arguments": {"model_name": "c/d", key: "x/y"}, + } + ], + ) + with pytest.raises(UntrustedWorkflowError, match=key): + pipeline.check_trusted() + assert list(no_network) == [] + + +class TestValidationRefusesBeforeImport: + """POST /api/validate and the pre-queue check call validation_errors - + a refusal there must not itself have imported the module to decide.""" + + def _workflow(self, pipeline, variables=None): + from dw.workflow import Workflow + + definition = { + "id": "trust_probe", + "steps": [ + { + "name": "gen", + "pipeline": pipeline, + "result": {"content_type": "image/png"}, + } + ], + } + if variables: + definition["variables"] = variables + return definition, Workflow + + # The pipeline itself is an escaped '{Fake}' rather than a real diffusers + # class: resolving a real one imports bitsandbytes, whose CPU backend + # asks the Hub for a kernel at import time - a network call these tests + # must not make, and nothing to do with the gate under test + @pytest.mark.parametrize( + "pipeline, path", + [ + pytest.param( + { + "configuration": {"component_type": f"{PROBE}.Thing"}, + "from_pretrained_arguments": {"model_name": "a/b"}, + "arguments": {}, + }, + "steps[0].pipeline.configuration.component_type", + id="component_type", + ), + pytest.param( + { + "configuration": {"component_type": "{Fake}"}, + "from_pretrained_arguments": {"model_name": "a/b"}, + "scheduler": { + "configuration": {"scheduler_type": f"{PROBE}.Thing"} + }, + "arguments": {}, + }, + "steps[0].pipeline.scheduler.configuration.scheduler_type", + id="scheduler_type", + ), + ], + ) + def test_a_dotted_type_is_a_validation_error( + self, untrusted, probe, no_network, tmp_path, pipeline, path + ): + definition, Workflow = self._workflow(pipeline) + errors = Workflow(definition, str(tmp_path), "").validation_errors() + refusals = [e for e in errors if "trust-workflows" in e["message"]] + assert [e["path"] for e in refusals] == [path] + assert_nothing_happened(probe, no_network) + + def test_a_constant_default_is_a_validation_error( + self, untrusted, probe, no_network, tmp_path + ): + definition, Workflow = self._workflow( + { + "configuration": {"component_type": "{Fake}"}, + "from_pretrained_arguments": {"model_name": "a/b"}, + "arguments": {"prompt": "variable:p"}, + }, + variables={"p": f"constant:{PROBE}.VALUE"}, + ) + errors = Workflow(definition, str(tmp_path), "").validation_errors() + assert [e["path"] for e in errors] == ["variables.p"] + assert_nothing_happened(probe, no_network) + + +class TestTheAllowlistIsNotAnEscapeHatch: + """In-ecosystem names are allowed untrusted by top-level package. That is + only safe if nothing reachable under those packages hands a workflow the + code execution the gate exists to deny.""" + + @pytest.mark.xfail( + strict=True, + reason="config_type is called with workflow kwargs and 'torch' is " + "allowlisted, so torch.hub.load(repo_or_dir=..., trust_repo=True) " + "runs a GitHub repo's hubconf.py untrusted", + ) + def test_config_type_cannot_name_a_code_loader_in_an_allowed_package( + self, untrusted, monkeypatch, no_network + ): + import torch.hub + + from dw.pipeline_processors.config_objects import create_quantization_config + + loader = MagicMock(name="torch.hub.load") + monkeypatch.setattr(torch.hub, "load", loader) + definition = { + "configuration": {"config_type": "torch.hub.load"}, + "arguments": { + "repo_or_dir": "attacker/repo", + "model": "anything", + "trust_repo": True, + }, + } + try: + realized = _realize({"quantization_config": definition}) + create_quantization_config(realized["quantization_config"]) + except UntrustedWorkflowError: + pass + loader.assert_not_called() + + @pytest.mark.xfail( + strict=True, + reason="load_constant_from_name walks attributes past the allowlisted " + "top-level module, so constant:torch.os.environ reads the server's " + "environment (DW_API_TOKEN, HF_TOKEN) untrusted", + ) + def test_a_constant_cannot_walk_out_of_an_allowed_package( + self, untrusted, monkeypatch + ): + from dw.arguments import fetch_constant + + monkeypatch.setenv("DW_API_TOKEN", "server-secret-probe") + leaked = None + try: + leaked = fetch_constant("constant:torch.os.environ") + except (UntrustedWorkflowError, ValueError): + pass + assert leaked is None or "server-secret-probe" not in str(leaked) + + @pytest.mark.parametrize( + "name", + [ + "os.system", + "builtins.eval", + "subprocess.Popen", + "importlib.import_module", + ".os", + " torch.os", + "torchx.Thing", + "diffusersx.Thing", + "dwx.Thing", + "Torch.nn.Linear", + ], + ) + def test_near_misses_of_an_allowed_name_are_refused(self, untrusted, name): + from dw.type_helpers import load_type_from_full_name + + with pytest.raises(UntrustedWorkflowError): + load_type_from_full_name(name) + + +# ------------------------------------------------------------------ trusted + + +@pytest.fixture(params=["env", "flag"]) +def trusted(request, monkeypatch): + """Both ways a process ends up trusting workflows: the environment + variable a spawned worker inherits, and the call the CLI flag makes.""" + if request.param == "env": + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "1") + else: + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + set_trust_workflows(True) + assert workflows_are_trusted() + return request.param + + +class TestTrustedLetsEachSurfaceThrough: + @pytest.mark.parametrize( + "arguments, find", + [ + ({"scheduler_type": f"{PROBE}.Thing"}, lambda a: a["scheduler_type"]), + ({"torch_dtype": f"{PROBE}.Thing"}, lambda a: a["torch_dtype"]), + ({"dtype": f"{PROBE}.Thing"}, lambda a: a["dtype"]), + ], + ) + def test_a_dotted_type_imports(self, trusted, probe, arguments, find): + realized = _realize(arguments) + assert find(realized).__name__ == "Thing" + assert probe.executed() + + def test_a_config_type_imports_and_builds(self, trusted, probe): + from dw.pipeline_processors.config_objects import create_quantization_config + + definition = { + "configuration": {"config_type": f"{PROBE}.Thing"}, + "arguments": {"bits": 4}, + } + realized = _realize({"quantization_config": definition}) + built = create_quantization_config(realized["quantization_config"]) + assert built.kwargs == {"bits": 4} + assert probe.executed() + + def test_pre_load_modules_import(self, trusted, probe, monkeypatch): + from dw.pipeline_processors.pipeline import Pipeline + + pipeline = _pipeline(configuration={"pre_load_modules": [PROBE]}) + monkeypatch.setattr( + Pipeline, + "populate_from_pretrained_arguments", + MagicMock(side_effect=RuntimeError("stop - past the gate")), + ) + with pytest.raises(RuntimeError, match="stop"): + pipeline.load(shared_components={}) + assert PROBE in probe.attempts + assert probe.executed() + + def test_a_constant_reads(self, trusted, probe): + from dw.arguments import fetch_constant + + assert fetch_constant(f"constant:{PROBE}.VALUE") == 7 + assert probe.executed() + + @pytest.mark.parametrize( + "extra", + [{"trust_remote_code": True}, {"custom_pipeline": "someone/pipeline"}], + ) + def test_remote_code_arguments_reach_from_pretrained(self, trusted, extra): + from dw.pipeline_processors.pipeline import load_component + + component_type = MagicMock() + component_type.__name__ = "ProbePipeline" + load_component( + "pipeline", + {"component_type": component_type}, + {"model_name": "a/b", **extra}, + "cpu", + ) + component_type.from_pretrained.assert_called_once() + _, kwargs = component_type.from_pretrained.call_args + for key, value in extra.items(): + assert kwargs[key] == value + + +class TestHowAProcessBecomesTrusted: + """Only the flag trusts: an unset or unrecognized variable is untrusted, + and a server started without --trust-workflows does not inherit trust + from whatever launched it.""" + + @pytest.mark.parametrize("value", ["", "0", "true", "yes", "TRUE", " 1", "1 "]) + def test_only_exactly_1_trusts(self, monkeypatch, value): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, value) + assert workflows_are_trusted() is False + + @pytest.fixture + def serve(self, monkeypatch, tmp_path): + import uvicorn + + import dw + import dw.serve as serve_module + from dw.server import app as app_module + + monkeypatch.setattr(app_module, "create_app", lambda **kwargs: object()) + monkeypatch.setattr(uvicorn, "run", lambda app, **kwargs: None) + monkeypatch.setattr(dw, "startup", lambda *args, **kwargs: None) + monkeypatch.delenv("DW_API_TOKEN", raising=False) + monkeypatch.setenv("DW_PROMPT_DIR", str(tmp_path / "prompts")) + monkeypatch.setenv("DW_ASSET_DIR", str(tmp_path / "assets")) + monkeypatch.setenv("DW_WORKSPACE", str(tmp_path / "workspace")) + monkeypatch.setenv("DW_WORKSPACE_SOURCE", "flag") + (tmp_path / "workflows").mkdir() + + def run(*argv): + monkeypatch.setattr( + "sys.argv", + ["dw-serve", "--workflow-dir", str(tmp_path / "workflows"), *argv], + ) + serve_module.main() + return workflows_are_trusted() + + return run + + def test_serve_without_the_flag_overrides_an_inherited_1(self, serve, monkeypatch): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "1") + assert serve() is False + + def test_serve_with_the_flag_trusts(self, serve, monkeypatch): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + assert serve("--trust-workflows") is True + + def test_validate_cli_without_the_flag_is_untrusted(self, monkeypatch, tmp_path): + """dw.validate is how an operator vets a file before running it - + it must vet it under the posture the run will have.""" + import dw.validate as validate_module + + workflow = tmp_path / "w.json" + workflow.write_text('{"id": "w", "steps": []}') + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "1") + monkeypatch.setattr("sys.argv", ["dw-validate", str(workflow)]) + try: + validate_module.main() + except SystemExit: + pass + assert workflows_are_trusted() is False From 5b651869665588fa7bb93cca82a1ce91f7664d46 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 075/181] test(security): planted symlinks and download_output destinations Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_symlinks.py | 638 ++++++++++++++++++++++++++++++++ 1 file changed, 638 insertions(+) create mode 100644 tests/test_security_symlinks.py diff --git a/tests/test_security_symlinks.py b/tests/test_security_symlinks.py new file mode 100644 index 00000000..8a51c974 --- /dev/null +++ b/tests/test_security_symlinks.py @@ -0,0 +1,638 @@ +"""A symlink planted inside a workspace must not carry anything out of it. + +Every root the server works in - outputs, assets, the shared asset library, +workflows, prompts, exports - is a directory a local user (or a mounted +volume, or an unpacked archive) can put a symlink into. `validate_path` +resolves symlinks before its containment check, so a path that is *checked* +is safe; the question this file asks is which code paths reach the disk +without being checked. For each root: read, write, delete and enumeration +through a link pointing at a sibling directory the server was never told +about, and the same for `download_output`'s destination over a mounted MCP +endpoint (gap 7 of the live suite's "Not covered here" list). + +Everything is under `tmp_path`: "outside" is a sibling of the workspace, and +the "secret" is a marker string, so a failing boundary leaks a marker and +overwrites a scratch file. +""" + +import io +import json +import os +import zipfile + +import httpx +import pytest +from fastapi.testclient import TestClient +from PIL import Image + +from dw.security import TRUST_WORKFLOWS_ENV_VAR +from dw.server.app import create_app +from dw.server.jobs import JobManager + +from .test_server import ScriptedWorkerManager, success_script, valid_workflow + +SECRET = "outside-the-roots-probe" + +pytestmark = pytest.mark.skipif( + not hasattr(os, "symlink") or os.name == "nt", + reason="symlinks need POSIX semantics", +) + + +def _png(path, color="red"): + Image.new("RGB", (4, 4), color).save(path) + return path + + +@pytest.fixture +def tree(tmp_path): + """/ws is the workspace root (default workspace), /outside the + directory nothing may reach. Each outside file carries SECRET.""" + root = tmp_path / "ws" + paths = { + "root": root, + "workflows": root / "workflows", + "outputs": root / "outputs", + "assets": root / "assets", + "prompts": root / "prompts", + "common": root / "common" / "assets", + "outside": tmp_path / "outside", + } + for path in paths.values(): + path.mkdir(parents=True, exist_ok=True) + outside = paths["outside"] + _png(outside / "secret.png") + (outside / "secret.txt").write_text(SECRET) + (outside / "secret.json").write_text( + json.dumps( + { + "id": "secret", + "description": SECRET, + "variables": {SECRET.replace("-", "_"): 1}, + "steps": [], + } + ) + ) + (outside / "prompt.json").write_text(json.dumps({"text": SECRET})) + (outside / "victim.txt").write_text("untouched") + (outside / "dir").mkdir() + _png(outside / "dir" / "inner.png") + (outside / "dir" / "inner.txt").write_text(SECRET) + (paths["workflows"] / "Basic.json").write_text(json.dumps(valid_workflow("b"))) + return paths + + +def link(at, target): + at.parent.mkdir(parents=True, exist_ok=True) + os.symlink(target, at) + return at + + +@pytest.fixture +def client(tree, tmp_path, monkeypatch): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + manager = JobManager( + str(tree["outputs"]), + worker_manager=ScriptedWorkerManager(success_script), + history_path=str(tmp_path / "jobs.sqlite"), + ) + app = create_app( + workflow_dir=str(tree["workflows"]), + output_dir=str(tree["outputs"]), + job_manager=manager, + prompt_dir=str(tree["prompts"]), + asset_dir=str(tree["assets"]), + workspace=str(tree["root"]), + ) + with TestClient(app, base_url="http://localhost") as test_client: + yield test_client + + +def _leaks(response): + return response.status_code < 300 and SECRET.encode() in response.content + + +def _outside_untouched(tree): + outside = tree["outside"] + assert (outside / "victim.txt").read_text() == "untouched" + assert (outside / "secret.txt").read_text() == SECRET + assert (outside / "dir" / "inner.txt").read_text() == SECRET + assert (outside / "secret.png").exists() + assert sorted(p.name for p in outside.iterdir()) == [ + "dir", + "prompt.json", + "secret.json", + "secret.png", + "secret.txt", + "victim.txt", + ], "something was written into the outside directory" + + +# ------------------------------------------------------------------ outputs + + +class TestOutputs: + @pytest.fixture(autouse=True) + def plant(self, tree): + link(tree["outputs"] / "leak.png", tree["outside"] / "secret.png") + link(tree["outputs"] / "leak.txt", tree["outside"] / "secret.txt") + link(tree["outputs"] / "linked_run", tree["outside"] / "dir") + + @pytest.mark.parametrize( + "path", + [ + "/outputs/leak.txt", + "/outputs/leak.png", + "/outputs/linked_run/inner.txt", + "/api/gallery/leak.txt/download", + "/api/gallery/leak.png/download", + "/api/gallery/linked_run/inner.txt/download", + "/api/gallery/leak.png/metadata", + "/api/gallery/leak.png/thumbnail", + "/api/gallery/linked_run/inner.png/thumbnail", + ], + ) + def test_reads_do_not_follow_the_link(self, client, path): + response = client.get(path) + assert not _leaks(response), path + assert response.status_code >= 400, (path, response.status_code) + + def test_the_archive_route_does_not_follow_the_link(self, client): + response = client.post("/api/gallery/archive", json={"names": ["leak.txt"]}) + assert response.status_code >= 400 + assert not _leaks(response) + + @pytest.mark.xfail( + strict=True, + reason="_iter_gallery_files walks outputs with os.walk and os.stat, " + "so GET /api/gallery lists a symlink pointing outside, with the " + "target's size and mtime", + ) + def test_the_gallery_listing_does_not_enumerate_the_link(self, client): + names = [entry["name"] for entry in client.get("/api/gallery").json()["files"]] + assert "leak.png" not in names + + def test_delete_removes_nothing_outside(self, client, tree): + for name in ("leak.txt", "linked_run/inner.txt", "linked_run"): + client.delete(f"/api/gallery/{name}") + _outside_untouched(tree) + + def test_keep_output_does_not_copy_the_target_into_assets(self, client, tree): + response = client.post( + "/api/assets/keep", json={"name": "leak.png", "asset_name": "kept.png"} + ) + assert response.status_code >= 400 + assert not (tree["assets"] / "kept.png").exists() + + def test_an_output_reference_does_not_follow_a_linked_run_directory(self, tree): + from dw.runs import resolve_output_reference + from dw.security import SecurityError + + with pytest.raises((SecurityError, ValueError)): + resolve_output_reference( + "output:linked_run/inner.png", root=str(tree["outputs"]) + ) + + +# ------------------------------------------------------------------- assets + + +class TestAssets: + @pytest.fixture(autouse=True) + def plant(self, tree): + link(tree["assets"] / "leak.png", tree["outside"] / "secret.png") + link(tree["assets"] / "cast", tree["outside"] / "dir") + link(tree["common"] / "shared-leak.png", tree["outside"] / "secret.png") + link(tree["common"] / "shared-dir", tree["outside"] / "dir") + + @pytest.mark.parametrize( + "path", + [ + "/inputs/leak.png", + "/inputs/cast/inner.txt", + "/inputs/cast/inner.png", + "/inputs/shared-leak.png", + "/inputs/shared-dir/inner.txt", + ], + ) + def test_reads_do_not_follow_the_link(self, client, path): + response = client.get(path) + assert response.status_code >= 400, (path, response.status_code) + + def test_the_asset_archive_does_not_follow_the_link(self, client): + for names in (["leak.png"], ["cast/inner.txt"], ["shared-dir/inner.txt"]): + response = client.post("/api/assets/archive", json={"names": names}) + assert response.status_code >= 400, names + assert not _leaks(response) + + @pytest.mark.parametrize( + "reference", + [ + "asset:leak.png", + "asset:cast/inner.png", + "asset:shared-leak.png", + "asset:shared-dir/inner.png", + ], + ) + def test_an_asset_reference_does_not_follow_the_link( + self, tree, monkeypatch, reference + ): + from dw.assets import resolve_asset_reference + from dw.security import SecurityError + + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + monkeypatch.setenv("DW_ASSET_DIR", str(tree["assets"])) + monkeypatch.setenv("DW_ASSET_PATH", str(tree["common"])) + with pytest.raises((SecurityError, ValueError)): + resolve_asset_reference(reference) + + @pytest.mark.xfail( + strict=True, + reason="the asset listing walks the library with os.walk/os.stat and " + "lists a symlink that resolves outside it", + ) + def test_the_asset_listing_does_not_enumerate_the_link(self, client): + listing = client.get("/api/assets").json() + assert "leak.png" not in json.dumps(listing) + + def test_an_upload_does_not_write_through_a_linked_name(self, client, tree): + """Uploads land in /uploads/, so that is where a link that + could redirect one would be planted.""" + link(tree["assets"] / "uploads" / "victim.png", tree["outside"] / "victim.txt") + response = client.post( + "/api/uploads?filename=victim.png&asset_name=victim.png", + content=b"\x89PNG overwritten", + ) + assert response.status_code >= 400 + _outside_untouched(tree) + + def test_an_upload_does_not_write_into_a_linked_folder(self, client, tree): + link(tree["assets"] / "uploads" / "cast", tree["outside"] / "dir") + response = client.post( + "/api/uploads?filename=new.png&asset_name=cast/new.png", + content=b"\x89PNG new", + ) + assert response.status_code >= 400 + assert not (tree["outside"] / "dir" / "new.png").exists() + _outside_untouched(tree) + + def test_a_shared_upload_does_not_write_through_a_link_in_common( + self, client, tree + ): + link(tree["common"] / "uploads" / "victim.png", tree["outside"] / "victim.txt") + response = client.post( + "/api/uploads?filename=victim.png&asset_name=victim.png&shared=true", + content=b"\x89PNG overwritten", + ) + assert response.status_code >= 400 + _outside_untouched(tree) + + def test_keep_output_does_not_write_through_a_linked_destination( + self, client, tree + ): + _png(tree["outputs"] / "real.png", "blue") + link(tree["assets"] / "dest.png", tree["outside"] / "victim.txt") + for overwrite in (False, True): + client.post( + "/api/assets/keep", + json={ + "name": "real.png", + "asset_name": "dest.png", + "overwrite": overwrite, + }, + ) + client.post( + "/api/assets/keep", + json={ + "name": "real.png", + "asset_name": "cast/new.png", + "overwrite": overwrite, + }, + ) + _outside_untouched(tree) + + def test_delete_removes_nothing_outside(self, client, tree): + for name in ("leak.png", "cast/inner.png", "cast", "shared-leak.png"): + client.delete(f"/api/assets/{name}") + _outside_untouched(tree) + + +# ------------------------------------------------------ gather_images globs + + +class TestGlobs: + """tests/test_locations.py drops a *file* match that escapes; a linked + *directory* in the pattern's fixed part is refused before expansion.""" + + @pytest.fixture(autouse=True) + def untrusted(self, monkeypatch, tree): + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + monkeypatch.setenv("DW_ASSET_DIR", str(tree["assets"])) + link(tree["assets"] / "cast", tree["outside"] / "dir") + link(tree["assets"] / "leak.png", tree["outside"] / "secret.png") + _png(tree["assets"] / "own.png") + + def test_a_linked_directory_in_the_pattern_is_refused(self, tree): + from dw.security import PathTraversalError + from dw.tasks.gather import gather_images + + with pytest.raises(PathTraversalError): + gather_images(glob=str(tree["assets"] / "cast" / "*.png")) + + def test_a_wildcard_does_not_descend_through_a_linked_directory(self, tree): + """The only match is under the linked directory; dropped, it leaves + nothing, which gather_images reports as no images.""" + from dw.tasks.gather import gather_images + + with pytest.raises(ValueError, match="No images found"): + gather_images(glob=str(tree["assets"] / "*" / "*.png")) + + def test_a_linked_file_match_is_dropped(self, tree): + from dw.tasks.gather import gather_images + + images = gather_images(glob=str(tree["assets"] / "*.png")) + assert len(images) == 1 + + def test_gather_videos_drops_a_linked_match(self, tree): + from dw.tasks.gather import gather_videos + + (tree["outside"] / "clip.mp4").write_bytes(b"not really a video") + link(tree["assets"] / "clip.mp4", tree["outside"] / "clip.mp4") + # a decode error here would mean the linked file was opened + with pytest.raises(ValueError, match="No videos found"): + gather_videos(glob=str(tree["assets"] / "*.mp4")) + + +# ------------------------------------------------------- workflows, prompts + + +class TestWorkflows: + @pytest.fixture(autouse=True) + def plant(self, tree): + link(tree["workflows"] / "leak.json", tree["outside"] / "secret.json") + link(tree["workflows"] / "linked", tree["outside"]) + + @pytest.mark.parametrize( + "path", + [ + "/api/workflows/leak", + "/api/workflows/leak/download", + "/api/workflows/leak/variables", + "/api/workflows/linked/secret", + ], + ) + def test_reads_do_not_follow_the_link(self, client, path): + response = client.get(path) + assert not _leaks(response), path + + @pytest.mark.xfail( + strict=True, + reason="workflow_details opens every *.json os.walk finds without a " + "containment check, so GET /api/workflows reads a linked file's " + "description and variable names", + ) + def test_the_listing_does_not_read_through_the_link(self, client): + response = client.get("/api/workflows") + assert SECRET not in response.text + assert SECRET.replace("-", "_") not in response.text + + def test_a_save_does_not_write_through_the_link(self, client, tree): + for name in ("leak", "linked/victim"): + client.put(f"/api/workflows/{name}", json=valid_workflow("overwrite")) + assert json.loads((tree["outside"] / "secret.json").read_text())["id"] == ( + "secret" + ) + _outside_untouched(tree) + + def test_a_delete_removes_nothing_outside(self, client, tree): + for name in ("leak", "linked/secret"): + client.delete(f"/api/workflows/{name}") + _outside_untouched(tree) + + +class TestPrompts: + @pytest.fixture(autouse=True) + def plant(self, tree): + link(tree["prompts"] / "leak.json", tree["outside"] / "prompt.json") + link(tree["prompts"] / "linked", tree["outside"]) + + @pytest.mark.parametrize( + "path", ["/api/prompts/leak", "/api/prompts/leak/download"] + ) + def test_reads_do_not_follow_the_link(self, client, path): + assert not _leaks(client.get(path)), path + + @pytest.mark.xfail( + strict=True, + reason="list_prompts hands every *.json workflow_names finds to " + "prompt_details, which opens it without a containment check - the " + "listing carries a linked file's text", + ) + def test_the_listing_does_not_read_through_the_link(self, client): + assert SECRET not in client.get("/api/prompts").text + + def test_a_prompt_reference_does_not_follow_the_link(self, tree, monkeypatch): + from dw.prompts import fetch_prompt + from dw.security import SecurityError + + for reference in ("prompt:leak", "prompt:linked/prompt"): + with pytest.raises((SecurityError, ValueError)): + fetch_prompt(reference, prompt_dir=str(tree["prompts"])) + + def test_a_save_does_not_write_through_the_link(self, client, tree): + for name in ("leak", "linked/victim"): + client.put(f"/api/prompts/{name}", json={"prompt": {"text": "overwrite"}}) + assert json.loads((tree["outside"] / "prompt.json").read_text()) == { + "text": SECRET + } + _outside_untouched(tree) + + def test_a_delete_removes_nothing_outside(self, client, tree): + for name in ("leak", "linked/prompt"): + client.delete(f"/api/prompts/{name}") + assert (tree["outside"] / "prompt.json").exists() + _outside_untouched(tree) + + +# -------------------------------------------------------- exports, workspaces + + +class TestExportsAndWorkspaces: + @pytest.mark.xfail( + strict=True, + reason="GET /exports/.zip (no token) zips the export directory " + "with os.walk + ZipFile.write, which follows a planted file symlink " + "and archives the target's bytes", + ) + def test_the_export_zip_does_not_follow_a_planted_link(self, client, tree): + export = tree["root"] / "exports" / "job-1" + export.mkdir(parents=True) + (export / "README.md").write_text("an export") + link(export / "leak.txt", tree["outside"] / "secret.txt") + + response = client.get("/exports/job-1.zip") + assert response.status_code == 200 + archive = zipfile.ZipFile(io.BytesIO(response.content)) + for name in archive.namelist(): + assert SECRET.encode() not in archive.read(name), name + + def test_the_export_zip_does_not_follow_a_linked_export_directory( + self, client, tree + ): + link(tree["root"] / "exports" / "job-2", tree["outside"] / "dir") + response = client.get("/exports/job-2.zip") + assert response.status_code == 404 + assert SECRET.encode() not in response.content + + def test_deleting_a_workspace_does_not_follow_a_link_inside_it(self, client, tree): + assert ( + client.post("/api/workspaces", json={"name": "doomed"}).status_code == 201 + ) + doomed = tree["root"] / "doomed" + link(doomed / "outputs" / "out", tree["outside"]) + link(doomed / "assets" / "cast", tree["outside"] / "dir") + client.delete("/api/workspaces/doomed?acknowledged=true") + _outside_untouched(tree) + + def test_the_workspace_size_count_does_not_follow_a_link(self, client, tree): + assert client.post("/api/workspaces", json={"name": "sized"}).status_code == 201 + link(tree["root"] / "sized" / "outputs" / "big", tree["outside"]) + response = client.delete("/api/workspaces/sized") + contents = response.json()["detail"]["contents"] + assert all(entry.get("files", 0) == 0 for entry in contents.values()), contents + + +# ------------------------------------------------------ download_output (7) + + +def _mounted(workspace_root, body=b"downloaded-bytes"): + """A DwClient shaped like the one dw.serve builds for its own /mcp.""" + from dw_mcp.client import DwClient + + def handler(request): + if request.url.path == "/api/server": + return httpx.Response( + 200, json={"directories": {"workspace": str(workspace_root)}} + ) + return httpx.Response(200, content=body, headers={"content-type": "image/png"}) + + client = DwClient(transport=httpx.MockTransport(handler)) + client.mounted = True + return client + + +class TestDownloadOutputDestination: + """Over a mounted endpoint the file lands on the server, so the + destination is confined to the workspace (#113). A legal relative + destination writes inside it and nowhere else.""" + + @pytest.fixture + def workspace(self, tree): + return tree["root"] + + def _download(self, workspace, destination, overwrite=False): + from dw_mcp.media import download_output + + return download_output( + _mounted(workspace), + "run/probe.png", + destination=destination, + overwrite=overwrite, + ) + + @pytest.mark.parametrize( + "destination", + ["kept/probe.png", "probe.png", "kept/", "deep/er/still/probe.png"], + ) + def test_a_relative_destination_lands_inside(self, workspace, tree, destination): + before = {p for p in tree["outside"].rglob("*")} + result = self._download(workspace, destination) + saved = os.path.realpath(result["saved_to"]) + assert saved.startswith(os.path.realpath(workspace) + os.sep) + assert open(saved, "rb").read() == b"downloaded-bytes" + assert {p for p in tree["outside"].rglob("*")} == before + + @pytest.mark.parametrize( + "destination", + [ + "../outside/escaped.png", + "kept/../../outside/escaped.png", + "OUTSIDE_ABSOLUTE", + "~/escaped.png", + ], + ) + def test_dot_dot_absolute_and_home_are_refused( + self, workspace, tree, monkeypatch, destination + ): + from dw_mcp.client import DwApiError + + monkeypatch.setenv("HOME", str(tree["outside"])) + if destination == "OUTSIDE_ABSOLUTE": + destination = str(tree["outside"] / "escaped.png") + with pytest.raises(DwApiError): + self._download(workspace, destination) + assert not (tree["outside"] / "escaped.png").exists() + _outside_untouched(tree) + + @pytest.mark.parametrize("overwrite", [False, True]) + def test_a_linked_parent_is_refused(self, workspace, tree, overwrite): + from dw_mcp.client import DwApiError + + link(workspace / "kept", tree["outside"]) + with pytest.raises(DwApiError): + self._download(workspace, "kept/escaped.png", overwrite=overwrite) + with pytest.raises(DwApiError): + self._download(workspace, "kept/victim.txt", overwrite=overwrite) + _outside_untouched(tree) + + def test_a_linked_grandparent_is_refused_even_when_the_parent_is_missing( + self, workspace, tree + ): + """The parent does not exist yet, so realpath of the destination + alone would stop at the missing segment - the check has to resolve + the nearest existing ancestor, which is the link.""" + from dw_mcp.client import DwApiError + + link(workspace / "kept", tree["outside"]) + with pytest.raises(DwApiError): + self._download(workspace, "kept/new/deeper/escaped.png") + assert not (tree["outside"] / "new").exists() + _outside_untouched(tree) + + def test_overwrite_cannot_clobber_a_linked_file(self, workspace, tree): + from dw_mcp.client import DwApiError + + link(workspace / "victim.png", tree["outside"] / "victim.txt") + with pytest.raises(DwApiError): + self._download(workspace, "victim.png", overwrite=True) + _outside_untouched(tree) + + def test_a_dangling_link_is_replaced_not_followed(self, workspace, tree): + """A link to a file that does not exist yet: writing through it + would create the file outside. The write must replace the link (or + refuse), never create its target.""" + target = tree["outside"] / "created-by-download.png" + link(workspace / "dangling.png", target) + try: + self._download(workspace, "dangling.png", overwrite=True) + except Exception: + pass + assert not target.exists() + _outside_untouched(tree) + + def test_a_workspace_root_that_is_itself_a_link_still_confines( + self, tree, tmp_path + ): + """The operator's workspace may be a symlink (a data volume); the + root is resolved, and a destination is confined to what it resolves + to - not to the link's own parent.""" + from dw_mcp.client import DwApiError + + linked_root = link(tmp_path / "ws-link", tree["root"]) + result = self._download(linked_root, "kept/probe.png") + assert os.path.realpath(result["saved_to"]).startswith( + os.path.realpath(tree["root"]) + os.sep + ) + with pytest.raises(DwApiError): + self._download(linked_root, str(tmp_path / "outside" / "escaped.png")) + _outside_untouched(tree) From 066d20659e088d8d6c7917b56166f2a2736c0847 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 076/181] test(security): decoder bombs through get_output_image and gallery routes Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_decoder_bombs.py | 195 +++++++++++++++++++++++++++ 1 file changed, 195 insertions(+) create mode 100644 tests/test_security_decoder_bombs.py diff --git a/tests/test_security_decoder_bombs.py b/tests/test_security_decoder_bombs.py new file mode 100644 index 00000000..d851ae9b --- /dev/null +++ b/tests/test_security_decoder_bombs.py @@ -0,0 +1,195 @@ +"""A crafted PNG that is small on disk and enormous once decoded. + +`get_output_image` (MCP) and the gallery thumbnail route both decode a whole +image a caller names. A PNG header can claim any size; the pixels are a +zlib stream that compresses a flat colour about a thousand to one. Pillow's +own guard (`Image.MAX_IMAGE_PIXELS`, ~89.5M) raises only above *twice* that +and only warns in between - so without a clamp of dw's own, an image just +under 179M pixels is decoded in full: half a gigabyte of RGB per request. + +The probe never lets a decode happen, whatever the boundary does: the +headers here are real, but `ImageFile.load` is replaced for the test with +one that records the attempt and raises before allocating. A test fails when +that record is non-empty - "refused" means refused *before* the decode. +""" + +import struct +import zlib + +import httpx +import pytest +from fastapi.testclient import TestClient +from PIL import Image, ImageFile + +from dw.server.app import create_app +from dw.server.jobs import JobManager +from dw_mcp.client import DwApiError, DwClient +from dw_mcp.media import get_output_image + +from .test_server import ScriptedWorkerManager, success_script + +# Well above anything the tool legitimately returns (a 4K frame is 8.3M) and +# well below Pillow's default, where only dw's own clamp can stop it +UNDER_PILLOWS_ERROR = (12_000, 12_000) # 144M pixels: Pillow only warns +OVER_PILLOWS_ERROR = (20_000, 20_000) # 400M pixels: Pillow refuses on open +DECODE_LIMIT = 50_000_000 + + +def _chunk(kind, data): + body = kind + data + return struct.pack(">I", len(data)) + body + struct.pack(">I", zlib.crc32(body)) + + +def bomb_png(width, height): + """A syntactically valid PNG header claiming width x height grayscale + pixels, with a token IDAT. Tiny on disk; never decoded by this file.""" + header = struct.pack(">IIBBBBB", width, height, 8, 0, 0, 0, 0) + return ( + b"\x89PNG\r\n\x1a\n" + + _chunk(b"IHDR", header) + + _chunk(b"IDAT", zlib.compress(b"\x00" * 64)) + + _chunk(b"IEND", b"") + ) + + +@pytest.fixture +def decodes(monkeypatch): + """Every attempt to decode an image over DECODE_LIMIT pixels, recorded + and stopped before the allocation.""" + attempts = [] + original = ImageFile.ImageFile.load + + def guarded(self): + width, height = self.size + if width * height > DECODE_LIMIT: + attempts.append(self.size) + raise MemoryError("probe: a decompression bomb was decoded") + return original(self) + + monkeypatch.setattr(ImageFile.ImageFile, "load", guarded) + return attempts + + +def test_the_probe_png_parses_to_the_size_it_claims(): + """The header is honest enough for Pillow to believe it - otherwise the + tests below would pass on a parse error rather than a clamp.""" + import io + import warnings + + with warnings.catch_warnings(): + warnings.simplefilter("ignore", Image.DecompressionBombWarning) + with Image.open(io.BytesIO(bomb_png(*UNDER_PILLOWS_ERROR))) as image: + assert image.size == UNDER_PILLOWS_ERROR + + +def test_pillows_guard_has_not_been_disabled(): + """Everything above 2 x MAX_IMAGE_PIXELS rests on Pillow's default; a + module that set it to None would open every size below.""" + assert Image.MAX_IMAGE_PIXELS is not None + assert Image.MAX_IMAGE_PIXELS <= 100_000_000 + + +def _serving(body): + def handler(request): + return httpx.Response(200, content=body, headers={"content-type": "image/png"}) + + return DwClient(transport=httpx.MockTransport(handler)) + + +class TestGetOutputImage: + def test_a_bomb_over_pillows_limit_is_refused_before_decode(self, decodes): + with pytest.raises((DwApiError, Image.DecompressionBombError)): + get_output_image(_serving(bomb_png(*OVER_PILLOWS_ERROR)), "bomb.png") + assert decodes == [] + + @pytest.mark.xfail( + strict=True, + reason="no pixel clamp - gap: get_output_image calls image.load() on " + "whatever Pillow opens, and Pillow only warns below 2x MAX_IMAGE_PIXELS", + ) + def test_a_bomb_under_pillows_limit_is_refused_before_decode(self, decodes): + try: + get_output_image(_serving(bomb_png(*UNDER_PILLOWS_ERROR)), "bomb.png") + except DwApiError: + pass + assert decodes == [] + + @pytest.mark.xfail( + strict=True, + reason="no pixel clamp - gap: a crop is cut from the fully decoded " + "image, so a small crop still decodes the whole bomb", + ) + def test_a_crop_does_not_decode_the_whole_bomb(self, decodes): + try: + get_output_image( + _serving(bomb_png(*UNDER_PILLOWS_ERROR)), "bomb.png", crop=[0, 0, 8, 8] + ) + except DwApiError: + pass + assert decodes == [] + + +@pytest.fixture +def gallery(tmp_path): + outputs = tmp_path / "outputs" + outputs.mkdir() + (tmp_path / "workflows").mkdir() + manager = JobManager( + str(outputs), + worker_manager=ScriptedWorkerManager(success_script), + history_path=str(tmp_path / "jobs.sqlite"), + ) + app = create_app( + workflow_dir=str(tmp_path / "workflows"), + output_dir=str(outputs), + job_manager=manager, + prompt_dir=str(tmp_path / "prompts"), + ) + client = TestClient(app, base_url="http://localhost", raise_server_exceptions=False) + return client, outputs + + +class TestGalleryThumbnail: + def test_a_bomb_over_pillows_limit_is_refused_before_decode(self, gallery, decodes): + client, outputs = gallery + (outputs / "bomb.png").write_bytes(bomb_png(*OVER_PILLOWS_ERROR)) + with client: + response = client.get("/api/gallery/bomb.png/thumbnail") + assert response.status_code >= 400 + assert decodes == [] + + @pytest.mark.xfail( + strict=True, + reason="no pixel clamp - gap: gallery_thumbnail's draft() is a no-op " + "for PNG, so thumbnail() decodes the full image first", + ) + def test_a_bomb_under_pillows_limit_is_refused_before_decode( + self, gallery, decodes + ): + client, outputs = gallery + (outputs / "bomb.png").write_bytes(bomb_png(*UNDER_PILLOWS_ERROR)) + with client: + client.get("/api/gallery/bomb.png/thumbnail") + assert decodes == [] + + @pytest.mark.xfail( + strict=True, + reason="no pixel clamp - gap: read_embedded_metadata (dw/result.py) " + "reads PngImageFile.text, which makes Pillow load() the whole image " + "to reach text chunks after IDAT", + ) + def test_the_gallery_metadata_route_does_not_decode_it(self, gallery, decodes): + """Width and height are in the header; nothing about a request for + metadata needs every pixel of a bomb.""" + client, outputs = gallery + (outputs / "bomb.png").write_bytes(bomb_png(*UNDER_PILLOWS_ERROR)) + with client: + client.get("/api/gallery/bomb.png/metadata") + assert decodes == [] + + def test_the_gallery_listing_does_not_decode_it(self, gallery, decodes): + client, outputs = gallery + (outputs / "bomb.png").write_bytes(bomb_png(*UNDER_PILLOWS_ERROR)) + with client: + assert client.get("/api/gallery").status_code == 200 + assert decodes == [] From 7ee72d253418f8bb344d2b54692d61cbd1009ce7 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 077/181] test(security): SSRF address spellings, redirects and HF token scope Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_ssrf.py | 599 ++++++++++++++++++++++++++++++++++++ 1 file changed, 599 insertions(+) create mode 100644 tests/test_security_ssrf.py diff --git a/tests/test_security_ssrf.py b/tests/test_security_ssrf.py new file mode 100644 index 00000000..18765f01 --- /dev/null +++ b/tests/test_security_ssrf.py @@ -0,0 +1,599 @@ +"""SSRF beyond loopback, and where the HuggingFace token may travel. + +tests/test_locations.py pins the host policy for the spellings a workflow +author would write by accident - 127.0.0.1, localhost, 169.254.169.254, one +RFC1918 address each. This file covers the spellings an attacker writes on +purpose: the rest of the private and link-local space, IPv6 and IPv4-mapped +IPv6, the decimal/octal/hex spellings glibc's resolver accepts for an IPv4 +address, DNS names that answer with any of those, and redirects. The second +half pins where `remote_text_encoder` sends this machine's HuggingFace token +(#112) against lookalike hosts, URL-parser tricks and a cross-host redirect. + +Nothing here touches the network. Name resolution goes through +`fake_resolver`, which answers a numeric spelling the way glibc's +`getaddrinfo` does (via `inet_aton`, which is pure computation) and a name +only from its own table; HTTP goes through `_Transport`, which stands in for +requests' `HTTPAdapter.send` - the one place every requests call in the +engine, and diffusers' own `load_image`, reaches the wire. +""" + +import io +import socket +from unittest.mock import patch +from urllib.parse import urlparse + +import pytest +import requests +from PIL import Image + +from dw.locations import token_host_allowed, validate_media_url +from dw.security import InvalidInputError, TRUST_WORKFLOWS_ENV_VAR + + +@pytest.fixture +def untrusted(monkeypatch): + """The posture a deployed server runs on (conftest defaults to trusted).""" + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + + +def fake_resolver(names=None): + """A getaddrinfo that never leaves the process. + + A numeric host resolves the way glibc resolves it - `inet_aton` accepts + '2130706433', '0x7f000001', '0177.0.0.1' and '127.1', exactly the + spellings a string-matching SSRF filter misses. Any other name answers + only from `names`, and an unknown one fails the way DNS does. + """ + names = names or {} + + def getaddrinfo(host, port=None, *args, **kwargs): + addresses = names.get(host) + if addresses is None: + try: + addresses = [socket.inet_ntoa(socket.inet_aton(host))] + except OSError: + raise socket.gaierror(socket.EAI_NONAME, "Name or service not known") + answers = [] + for address in addresses: + family = socket.AF_INET6 if ":" in address else socket.AF_INET + sockaddr = (address, port or 0, 0, 0) if ":" in address else (address, 0) + answers.append((family, socket.SOCK_STREAM, 6, "", sockaddr)) + return answers + + return getaddrinfo + + +@pytest.fixture +def no_real_sockets(monkeypatch): + """Belt and braces: a test that forgot a mock fails instead of dialing.""" + + def refuse(*args, **kwargs): + raise AssertionError("a test in this file tried to open a real connection") + + monkeypatch.setattr(socket.socket, "connect", refuse) + monkeypatch.setattr(socket, "create_connection", refuse) + + +def _refused(url, names=None): + with patch("dw.locations.socket.getaddrinfo", fake_resolver(names)): + with pytest.raises(InvalidInputError) as refusal: + validate_media_url(url, "an image argument") + return str(refusal.value) + + +class TestInternalAddressSpellings: + """Every spelling of an address inside the deployment is refused.""" + + @pytest.mark.parametrize( + "url", + [ + # link-local, including the cloud metadata address on other ports + "http://169.254.169.254/latest/meta-data/iam/", + "http://169.254.169.254:8080/", + "http://169.254.0.1/", + # the whole of RFC1918, both ends of each block + "http://10.0.0.1/", + "http://10.255.255.254/", + "http://172.16.0.1/", + "http://172.31.255.254/", + "http://192.168.0.1/", + "http://192.168.255.254/", + # the rest of loopback, and 'this host' + "http://127.0.0.2/", + "http://127.255.255.254/", + "http://0.0.0.0/", + ], + ) + def test_ipv4_literals(self, untrusted, no_real_sockets, url): + assert "inside this deployment" in _refused(url) + + @pytest.mark.parametrize( + "url", + [ + "http://[::1]/", + "http://[0:0:0:0:0:0:0:1]/", + "http://[::]/", + # unique local, including AWS's IPv6 metadata endpoint + "http://[fd00::1]/", + "http://[fd00:ec2::254]/latest/meta-data/", + "http://[fc00::1]/", + # link-local, with and without a zone id + "http://[fe80::1]/", + "http://[fe80::1%25eth0]/", + ], + ) + def test_ipv6_literals(self, untrusted, no_real_sockets, url): + assert "inside this deployment" in _refused(url) + + @pytest.mark.parametrize( + "url", + [ + "http://[::ffff:127.0.0.1]/", + "http://[::ffff:7f00:1]/", + "http://[::ffff:169.254.169.254]/", + "http://[::ffff:a9fe:a9fe]/", + "http://[::ffff:10.0.0.1]/", + "http://[::ffff:192.168.1.1]/", + ], + ) + def test_ipv4_mapped_ipv6(self, untrusted, no_real_sockets, url): + assert "inside this deployment" in _refused(url) + + @pytest.mark.parametrize( + "url", + [ + # 127.0.0.1 as one decimal, as hex, as octal, and shortened + "http://2130706433/", + "http://0x7f000001/", + "http://0x7f.0x0.0x0.0x1/", + "http://0177.0.0.1/", + "http://017700000001/", + "http://127.1/", + "http://0/", + # 169.254.169.254 the same ways + "http://2852039166/", + "http://0xa9fea9fe/", + "http://0251.0376.0251.0376/", + # 10.0.0.1 and 192.168.0.1 + "http://167772161/", + "http://0xc0a80001/", + ], + ) + def test_numeric_spellings_the_resolver_accepts( + self, untrusted, no_real_sockets, url + ): + """Refused because the check runs on what the name resolves to, not + on the string - these never look like an IP address to a regex.""" + assert "inside this deployment" in _refused(url) + + @pytest.mark.parametrize( + "address", + [ + "127.0.0.1", + "169.254.169.254", + "10.1.2.3", + "172.20.0.5", + "192.168.1.10", + "::1", + "fd12:3456::1", + "fe80::1", + "::ffff:127.0.0.1", + ], + ) + def test_a_name_that_resolves_inside(self, untrusted, no_real_sockets, address): + assert "inside this deployment" in _refused( + "http://innocent.example.com/x.png", + {"innocent.example.com": [address]}, + ) + + def test_one_internal_answer_among_public_ones_is_enough( + self, untrusted, no_real_sockets + ): + """The fetch may connect to any of the answers, so every one counts.""" + assert "inside this deployment" in _refused( + "http://round-robin.example.com/x.png", + {"round-robin.example.com": ["93.184.216.34", "10.0.0.7"]}, + ) + + def test_userinfo_does_not_hide_the_real_host(self, untrusted, no_real_sockets): + """'user@host' - the part after the '@' is where the request goes.""" + assert "inside this deployment" in _refused( + "http://example.com@169.254.169.254/latest/meta-data/" + ) + + def test_the_metadata_hostname_is_refused_by_what_it_answers( + self, untrusted, no_real_sockets + ): + assert "inside this deployment" in _refused( + "http://metadata.google.internal/computeMetadata/v1/", + {"metadata.google.internal": ["169.254.169.254"]}, + ) + + @pytest.mark.xfail( + strict=True, + reason="100.64.0.0/10 (CGNAT, Alibaba's 100.100.100.200 metadata) is " + "not is_private in ipaddress, and _is_internal never asks is_global", + ) + def test_shared_address_space_metadata_is_refused(self, untrusted, no_real_sockets): + assert "inside this deployment" in _refused( + "http://100.100.100.200/latest/meta-data/" + ) + + @pytest.mark.parametrize( + "url, names", + [ + ("http://93.184.216.34/x.png", None), + ("https://example.com/x.png", {"example.com": ["93.184.216.34"]}), + ("https://example.com/x.png", {"example.com": ["2606:2800:21f::1"]}), + ], + ) + def test_a_public_address_is_still_allowed( + self, untrusted, no_real_sockets, url, names + ): + """The policy is not 'refuse everything': a public host passes.""" + with patch("dw.locations.socket.getaddrinfo", fake_resolver(names)): + assert validate_media_url(url, "an image argument") == url + + +# ---------------------------------------------------------------- transport + + +def _png_bytes(): + buffer = io.BytesIO() + Image.new("RGB", (2, 2), "red").save(buffer, format="PNG") + return buffer.getvalue() + + +class _Transport: + """Stands in for requests' HTTPAdapter.send: answers from a table of + {url: (status, headers, body)} and records every request that would + have gone on the wire, headers included.""" + + def __init__(self, routes): + self.routes = routes + self.sent = [] + + def send(self, adapter, request, **kwargs): + self.sent.append(request) + status, headers, body = self.routes.get(request.url, (404, {}, b"not found")) + response = requests.Response() + response.status_code = status + response.headers.update(headers) + response.url = request.url + response.request = request + response.raw = io.BytesIO(body) + response.reason = "scripted" + return response + + def install(self, monkeypatch): + transport = self + + def send(adapter, request, **kwargs): + return transport.send(adapter, request, **kwargs) + + monkeypatch.setattr(requests.adapters.HTTPAdapter, "send", send) + return self + + def hosts(self): + return [urlparse(request.url).hostname for request in self.sent] + + +PUBLIC = {"cdn.example.com": ["93.184.216.34"], "evil.example": ["45.33.32.156"]} + + +class TestRedirects: + """A URL is checked once, before the fetch - but the fetch follows + redirects, and the redirect target was never in the document to check.""" + + @pytest.mark.xfail( + strict=True, + reason="fetch_image validates the first URL only; requests follows " + "a 302 from a public host to 169.254.169.254 unchecked", + ) + def test_fetch_image_does_not_follow_a_redirect_inside( + self, untrusted, no_real_sockets, monkeypatch, tmp_path + ): + from dw.arguments import fetch_image + + transport = _Transport( + { + "https://cdn.example.com/x.png": ( + 302, + {"Location": "http://169.254.169.254/latest/meta-data/"}, + b"", + ), + "http://169.254.169.254/latest/meta-data/": ( + 200, + {"Content-Type": "image/png"}, + _png_bytes(), + ), + } + ).install(monkeypatch) + + with patch("dw.locations.socket.getaddrinfo", fake_resolver(PUBLIC)): + try: + fetch_image("https://cdn.example.com/x.png", str(tmp_path)) + except Exception: + pass + + assert "169.254.169.254" not in transport.hosts() + + @pytest.mark.xfail( + strict=True, + reason="audio fetch validates the first URL only; requests follows " + "a redirect to loopback unchecked", + ) + def test_audio_fetch_does_not_follow_a_redirect_inside( + self, untrusted, no_real_sockets, monkeypatch, tmp_path + ): + from dw.tasks.audio_utils import load_audio + + transport = _Transport( + { + "https://cdn.example.com/a.wav": ( + 307, + {"Location": "http://127.0.0.1:8765/api/server"}, + b"", + ), + } + ).install(monkeypatch) + + with patch("dw.locations.socket.getaddrinfo", fake_resolver(PUBLIC)): + try: + load_audio("https://cdn.example.com/a.wav", str(tmp_path)) + except Exception: + pass + + assert "127.0.0.1" not in transport.hosts() + + @pytest.mark.xfail( + strict=True, + reason="validate_media_url checks urlparse's host (example.com, after " + "the backslash) while requests dials urllib3's (169.254.169.254)", + ) + def test_a_backslash_does_not_split_the_checked_host_from_the_dialed_one( + self, untrusted, no_real_sockets, monkeypatch, tmp_path + ): + """Parser differential: the host policy and the HTTP client must + agree on which host a URL names, or the policy checks one and the + request goes to the other.""" + from dw.arguments import fetch_image + + transport = _Transport({}).install(monkeypatch) + with patch( + "dw.locations.socket.getaddrinfo", + fake_resolver({"example.com": ["93.184.216.34"]}), + ): + try: + fetch_image( + "http://169.254.169.254\\@example.com/latest/meta-data/", + str(tmp_path), + ) + except Exception: + pass + assert "169.254.169.254" not in transport.hosts() + + def test_a_refused_first_url_sends_nothing( + self, untrusted, no_real_sockets, monkeypatch, tmp_path + ): + """The refusal comes before the request, not after the response.""" + from dw.arguments import fetch_image + + transport = _Transport({}).install(monkeypatch) + with patch("dw.locations.socket.getaddrinfo", fake_resolver()): + with pytest.raises(InvalidInputError): + fetch_image("http://2852039166/latest/meta-data/", str(tmp_path)) + assert transport.sent == [] + + +# ------------------------------------------------------------- the HF token + +HF_TOKEN = "hf_test_token_never_real" + +HF_HOSTS = { + "api-inference.huggingface.co": ["18.0.0.1"], + "huggingface.co": ["18.0.0.2"], + "abc.us-east-1.aws.endpoints.huggingface.cloud": ["18.0.0.3"], + "someone-encoder.hf.space": ["18.0.0.4"], + "huggingface.co.evil.example": ["45.33.32.156"], + "evilhuggingface.co": ["45.33.32.156"], + "huggingface.co-evil.example": ["45.33.32.156"], + "evil.example": ["45.33.32.156"], + "hf.space.evil.example": ["45.33.32.156"], + "notreallyhf.space": ["45.33.32.156"], + "huggingface.cloud.evil.example": ["45.33.32.156"], +} + + +class _Embeds: + def to(self, device): + return self + + +def _encode(url, monkeypatch, routes=None): + """Run remote_text_encoder against a scripted transport and return what + went on the wire. The token is a fixed fake; torch.load never sees a + real body.""" + from dw.pipeline_processors import remote + + transport = _Transport( + routes + or { + url: (200, {"Content-Type": "application/octet-stream"}, b"embeds"), + } + ).install(monkeypatch) + monkeypatch.setattr(remote, "get_token", lambda: HF_TOKEN) + monkeypatch.setattr(remote.torch, "load", lambda *a, **k: _Embeds()) + with patch("dw.locations.socket.getaddrinfo", fake_resolver(HF_HOSTS)): + try: + remote.remote_text_encoder(["a prompt"], url, "cpu") + except RuntimeError: + # a scripted non-200 at the end of a redirect chain + pass + return transport + + +def _carried_token(request): + return HF_TOKEN in (request.headers.get("Authorization") or "") + + +class TestHuggingFaceTokenScope: + @pytest.mark.parametrize( + "host", + [ + "huggingface.co", + "api-inference.huggingface.co", + "abc.us-east-1.aws.endpoints.huggingface.cloud", + "huggingface.cloud", + "hf.space", + "someone-encoder.hf.space", + "HuggingFace.CO", + ], + ) + def test_the_token_hosts(self, host): + assert token_host_allowed(host) + + @pytest.mark.parametrize( + "host", + [ + "huggingface.co.evil.example", + "evilhuggingface.co", + "huggingface.co-evil.example", + "huggingface-co.evil.example", + "hf.space.evil.example", + "notreallyhf.space", + "huggingface.cloud.evil.example", + "xhuggingface.cloud", + "huggingface.co.", + "huggingface", + "co", + "", + None, + ], + ) + def test_lookalikes_do_not_get_it(self, host): + assert not token_host_allowed(host) + + @pytest.mark.parametrize( + "url", + [ + "https://api-inference.huggingface.co/models/x", + "https://abc.us-east-1.aws.endpoints.huggingface.cloud/", + "https://someone-encoder.hf.space/encode", + ], + ) + def test_a_huggingface_endpoint_gets_the_token( + self, untrusted, no_real_sockets, monkeypatch, url + ): + transport = _encode(url, monkeypatch) + assert len(transport.sent) == 1 + assert _carried_token(transport.sent[0]) + + @pytest.mark.parametrize( + "url", + [ + "https://huggingface.co.evil.example/encode", + "https://evilhuggingface.co/encode", + "https://huggingface.co-evil.example/encode", + "https://evil.example/huggingface.co/encode", + "https://evil.example/?host=huggingface.co", + # userinfo: everything before '@' is credentials, not the host + "https://huggingface.co@evil.example/encode", + "https://huggingface.co:443@evil.example/encode", + "https://api-inference.huggingface.co%2f@evil.example/encode", + # a fragment or a backslash that a browser would read differently + "https://evil.example#@huggingface.co/encode", + "https://evil.example/\\@huggingface.co/encode", + ], + ) + def test_a_lookalike_url_never_carries_it( + self, untrusted, no_real_sockets, monkeypatch, url + ): + transport = _encode(url, monkeypatch) + for request in transport.sent: + assert not _carried_token(request), request.url + + @pytest.mark.parametrize( + "url", + [ + "https://huggingface.co@evil.example/encode", + pytest.param( + "https://evil.example\\@huggingface.co/encode", + marks=pytest.mark.xfail( + strict=True, + reason="remote.py decides on urlparse's host (huggingface.co, " + "after the backslash) but requests dials urllib3's " + "(evil.example) - the token leaves with the request", + ), + ), + "https://evil.example%5c@huggingface.co/encode", + "https://evil.example%40huggingface.co/encode", + ], + ) + def test_the_token_goes_only_where_the_connection_goes( + self, untrusted, no_real_sockets, monkeypatch, url + ): + """Parser differential: the decision reads urllib.parse's idea of the + host, the connection uses requests'/urllib3's. Whatever either one + makes of a URL, a request that carries the token must be addressed + to a HuggingFace host by the parser that actually dials.""" + from urllib3.util import parse_url + + transport = _encode(url, monkeypatch) + for request in transport.sent: + if _carried_token(request): + assert token_host_allowed(parse_url(request.url).host), request.url + + def test_a_redirect_to_another_host_drops_the_token( + self, untrusted, no_real_sockets, monkeypatch + ): + """A HuggingFace endpoint answering 307 elsewhere must not hand the + credential on with the replayed POST.""" + start = "https://api-inference.huggingface.co/models/x" + transport = _encode( + start, + monkeypatch, + routes={ + start: (307, {"Location": "https://evil.example/collect"}, b""), + "https://evil.example/collect": ( + 200, + {"Content-Type": "application/octet-stream"}, + b"embeds", + ), + }, + ) + assert transport.hosts() == ["api-inference.huggingface.co", "evil.example"] + assert _carried_token(transport.sent[0]) + assert not _carried_token(transport.sent[1]) + + def test_a_redirect_to_http_on_the_same_host_drops_the_token( + self, untrusted, no_real_sockets, monkeypatch + ): + """A downgrade to cleartext is a different origin too.""" + start = "https://api-inference.huggingface.co/models/x" + downgrade = "http://api-inference.huggingface.co/models/x" + transport = _encode( + start, + monkeypatch, + routes={ + start: (307, {"Location": downgrade}, b""), + downgrade: ( + 200, + {"Content-Type": "application/octet-stream"}, + b"embeds", + ), + }, + ) + cleartext = [r for r in transport.sent if r.url.startswith("http://")] + assert cleartext, "the scripted redirect was not followed" + assert not any(_carried_token(r) for r in cleartext) + + def test_trust_lifts_the_scope(self, no_real_sockets, monkeypatch): + """--trust-workflows lifts the scope: the documented escape hatch, + pinned so it cannot widen silently into the untrusted default.""" + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "1") + transport = _encode("https://evil.example/encode", monkeypatch) + assert _carried_token(transport.sent[0]) + monkeypatch.setenv(TRUST_WORKFLOWS_ENV_VAR, "0") + transport = _encode("https://evil.example/encode", monkeypatch) + assert not _carried_token(transport.sent[0]) From b5aa656f3145f64ea63e04b7cab28dc4c1fd7d36 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 03:27:14 +0000 Subject: [PATCH 078/181] test(security): documented input caps refused before any work Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01NSpbAKgGb282ixhsEqGQ52 --- tests/test_security_input_caps.py | 253 ++++++++++++++++++++++++++++++ 1 file changed, 253 insertions(+) create mode 100644 tests/test_security_input_caps.py diff --git a/tests/test_security_input_caps.py b/tests/test_security_input_caps.py new file mode 100644 index 00000000..f3f08d9f --- /dev/null +++ b/tests/test_security_input_caps.py @@ -0,0 +1,253 @@ +"""The documented size limits, refused at validation - before any work. + +`MAX_VARIABLE_VALUE_LENGTH` (20,000 characters), the variable-name cap, the +32-entry `for_each` ceiling, the 50MB workflow-file ceiling and the 200MB +upload ceiling are each a promise that an oversized input costs the server +nothing but the refusal. The live suite can see the refusal; it cannot see +whether a worker was handed the job first, a file was read whole, or a body +was buffered. These tests watch for exactly that: a ScriptedWorkerManager +that records every command it is sent, a sparse file that would be 50MB to +read, and an ASGI receive() that records whether the body was pulled. +""" + +import asyncio +import json + +import pytest +from fastapi.testclient import TestClient + +from dw.security import ( + MAX_JSON_SIZE, + MAX_VARIABLE_NAME_LENGTH, + MAX_VARIABLE_VALUE_LENGTH, + InvalidInputError, +) +from dw.server.app import create_app +from dw.server.jobs import JobManager +from dw.variables import argument_errors + +from .test_server import ScriptedWorkerManager, success_script + +OVERSIZED = "x" * 25_000 + + +def _workflow(variables=None, for_each=None): + step = { + "name": "t", + "task": {"command": "compose_text", "arguments": {"parts": ["variable:p"]}}, + "result": {"content_type": "text/plain"}, + } + if for_each is not None: + step["for_each"] = for_each + step["task"]["arguments"]["parts"] = ["item:prompt"] + return { + "id": "caps", + "variables": {"p": "short", "shots": [], **(variables or {})}, + "steps": [step], + } + + +@pytest.fixture +def server(tmp_path): + (tmp_path / "workflows").mkdir() + worker = ScriptedWorkerManager(success_script) + manager = JobManager( + str(tmp_path / "outputs"), + worker_manager=worker, + history_path=str(tmp_path / "jobs.sqlite"), + ) + app = create_app( + workflow_dir=str(tmp_path / "workflows"), + output_dir=str(tmp_path / "outputs"), + job_manager=manager, + prompt_dir=str(tmp_path / "prompts"), + asset_dir=str(tmp_path / "assets"), + ) + with TestClient(app, base_url="http://localhost") as client: + yield client, worker + + +def _executed(worker): + return [c for c in worker.commands if c.get("type") == "execute"] + + +def test_the_documented_limit_is_the_one_under_test(): + assert MAX_VARIABLE_VALUE_LENGTH == 20_000 + assert len(OVERSIZED) > MAX_VARIABLE_VALUE_LENGTH + + +class TestVariableValueLength: + def test_the_limit_itself_is_accepted(self): + assert ( + argument_errors(_workflow(), {"p": "x" * MAX_VARIABLE_VALUE_LENGTH}) == [] + ) + + def test_one_over_is_refused(self): + errors = argument_errors( + _workflow(), {"p": "x" * (MAX_VARIABLE_VALUE_LENGTH + 1)} + ) + assert [e["path"] for e in errors] == ["arguments.p"] + + @pytest.mark.parametrize( + "arguments, path", + [ + ({"p": OVERSIZED}, "arguments.p"), + ({"shots": [{"name": "a", "prompt": OVERSIZED}]}, "arguments.shots"), + ({"shots": [[OVERSIZED]]}, "arguments.shots"), + ( + {"shots": [{"name": "a", "nested": {"deep": OVERSIZED}}]}, + "arguments.shots", + ), + ], + ) + def test_25000_characters_is_refused_wherever_it_sits(self, arguments, path): + errors = argument_errors(_workflow(), arguments) + assert [e["path"] for e in errors] == [path] + assert "too long" in errors[0]["message"] + + def test_the_validate_route_refuses_it(self, server): + client, worker = server + response = client.post( + "/api/validate", + json={"workflow": _workflow(), "arguments": {"p": OVERSIZED}}, + ) + body = response.json() + assert body["valid"] is False + assert any(e["path"] == "arguments.p" for e in body["errors"]) + assert _executed(worker) == [] + + def test_the_jobs_route_refuses_it_before_the_worker_sees_it(self, server): + client, worker = server + response = client.post( + "/api/jobs", json={"workflow": _workflow(), "arguments": {"p": OVERSIZED}} + ) + assert response.status_code == 400 + assert _executed(worker) == [] + assert client.get("/api/jobs").json() in ([], {"jobs": []}) or not ( + client.get("/api/jobs").json().get("jobs") + ) + + def test_the_cli_refuses_it_before_loading_anything(self, monkeypatch, tmp_path): + import dw.run as run_module + + loaded = [] + monkeypatch.setattr( + run_module, + "workflow_from_file", + lambda *a, **k: loaded.append(a), + raising=False, + ) + workflow = tmp_path / "w.json" + workflow.write_text(json.dumps(_workflow())) + monkeypatch.setattr("sys.argv", ["dw-run", str(workflow), f"p={OVERSIZED}"]) + with pytest.raises(SystemExit) as exit_info: + run_module.main() + assert exit_info.value.code != 0 + assert loaded == [] + + @pytest.mark.xfail( + strict=True, + reason="the cap is applied to caller arguments only - a 25,000-character " + "default written into an inline workflow's own variables validates clean", + ) + def test_a_default_in_the_definition_is_held_to_the_same_cap(self, server): + client, worker = server + response = client.post( + "/api/validate", json={"workflow": _workflow({"p": OVERSIZED})} + ) + assert response.json()["valid"] is False + + +class TestOtherDocumentedLimits: + def test_an_over_long_variable_name_is_refused(self): + name = "v" * (MAX_VARIABLE_NAME_LENGTH + 1) + errors = argument_errors(_workflow({name: "d"}), {name: "value"}) + assert [e["path"] for e in errors] == [f"arguments.{name}"] + + def test_33_for_each_entries_are_a_validation_error(self, server): + client, worker = server + entries = [{"name": f"e{i}", "prompt": "p"} for i in range(33)] + response = client.post( + "/api/validate", + json={"workflow": _workflow(for_each=entries)}, + ) + body = response.json() + assert body["valid"] is False + assert any(e["path"] == "steps[0].for_each" for e in body["errors"]) + assert _executed(worker) == [] + + def test_33_entries_through_arguments_are_refused_before_the_queue(self, server): + client, worker = server + entries = [{"name": f"e{i}", "prompt": "p"} for i in range(33)] + response = client.post( + "/api/jobs", + json={ + "workflow": _workflow(for_each="variable:shots"), + "arguments": {"shots": entries}, + }, + ) + assert response.status_code == 400 + assert _executed(worker) == [] + + def test_an_oversized_workflow_file_is_refused_without_reading_it( + self, tmp_path, monkeypatch + ): + """A sparse file costs no disk; the check is on its size, so a + refusal that opened it first would show up as an open() call.""" + from dw.workflow import workflow_from_file + + big = tmp_path / "big.json" + with open(big, "wb") as file: + file.truncate(MAX_JSON_SIZE + 1) + + opened = [] + real_open = open + + def recording_open(path, *args, **kwargs): + if str(path) == str(big) or str(path).endswith("big.json"): + opened.append(path) + return real_open(path, *args, **kwargs) + + monkeypatch.setattr("builtins.open", recording_open) + with pytest.raises(InvalidInputError, match="too large"): + workflow_from_file(str(big), str(tmp_path / "out")) + assert opened == [] + + def test_an_upload_over_the_cap_is_refused_from_its_declared_length(self, server): + """The 200MB ceiling is checked on Content-Length before the body is + read - driven over raw ASGI so the declared length can exceed what + is actually sent, and so the test sees whether receive() was called.""" + client, worker = server + app = client.app + pulled = [] + sent = [] + + async def receive(): + pulled.append(True) + return {"type": "http.request", "body": b"x", "more_body": False} + + async def send(message): + sent.append(message) + + scope = { + "type": "http", + "asgi": {"version": "3.0"}, + "http_version": "1.1", + "method": "POST", + "scheme": "http", + "path": "/api/uploads", + "raw_path": b"/api/uploads", + "query_string": b"filename=huge.png", + "root_path": "", + "headers": [ + (b"host", b"localhost"), + (b"content-length", str(300 * 1024 * 1024).encode()), + (b"content-type", b"application/octet-stream"), + ], + "client": ("127.0.0.1", 1), + "server": ("localhost", 80), + } + asyncio.run(app(scope, receive, send)) + start = next(m for m in sent if m["type"] == "http.response.start") + assert start["status"] == 413 + assert pulled == [] From b4a08b33527e6dad36f2e922c0a727579f356ac1 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 22:35:18 -0500 Subject: [PATCH 079/181] feat(engine): #387 - assessment probes, rules table, json returns kind Co-Authored-By: Claude Opus 5.5 --- dw/assessment_rules.py | 152 ++++++++ dw/introspection.py | 20 +- dw/runs.py | 44 +++ dw/scalar_result_validation.py | 31 +- dw/tasks/assess.py | 660 +++++++++++++++++++++++++++++++++ dw/tasks/task.py | 36 ++ ui/src/lib/api.ts | 1 + 7 files changed, 936 insertions(+), 8 deletions(-) create mode 100644 dw/assessment_rules.py create mode 100644 dw/tasks/assess.py diff --git a/dw/assessment_rules.py b/dw/assessment_rules.py new file mode 100644 index 00000000..0911f7d4 --- /dev/null +++ b/dw/assessment_rules.py @@ -0,0 +1,152 @@ +"""The rules the assessment probes read their measurements against (#387). + +A probe (`dw/tasks/assess.py`) measures a finished file and reports every +number it took; a rule here names one of those numbers and the threshold +past which it is worth a look, and a crossing becomes a `finding` in the +probe's answer. Findings are places to look, not verdicts: nothing in the +engine acts on one, no run fails for one, and a finding a person has looked +at and accepted is simply left alone. That is the authority rule, and it is +why the thresholds sit in one table - a number an agent is told to trust +has to be one somebody can find, argue with and change in one place. + +Each entry names the probe that reports the field, the field (a key of the +probe's per-shot or per-seam record, or of its answer itself - +`tests/test_assessment_rules.py` pins every one to a real probe field so a +rename cannot leave a rule reading nothing), how the value is compared, +the threshold, the severity and what a crossing says. `magnitude` compares +the value's absolute size, for a signed measurement whose direction is not +the problem. `shot_level_spread` shares its threshold with the run-time +`LEVEL_SPREAD_WARN_DB`, so the join-time warning and the after-the-fact +probe cannot disagree about the same cut. + +Two rules carry a guard the table alone cannot express, applied by the +probe and named here in `unless` so it is written down beside the number: +`seam_hole` holds only while both sides of the seam are voiced, and +`seam_frame_jump` does not fire at a seam whose incoming shot is marked +`hard_cut: true` - a cut meant as a cut. +""" + +from .tasks.audio_utils import LEVEL_SPREAD_WARN_DB + +SEVERITIES = ("info", "warn") + +# How loud both sides of a seam have to be for a quiet join to be a hole +# rather than a pause the shots themselves hold +HOLE_VOICED_DBFS = -30.0 + +COMPARATORS = { + ">": lambda value, threshold: value > threshold, + ">=": lambda value, threshold: value >= threshold, + "<": lambda value, threshold: value < threshold, +} + +RULES = ( + { + "name": "shot_level_spread", + "probe": "analyze_shots", + "field": "rms_range_db", + "comparator": ">=", + "threshold": LEVEL_SPREAD_WARN_DB, + "severity": "warn", + "says": "the shots sit this far apart in level - a jump a listener hears at the cut", + }, + { + "name": "seam_level_step", + "probe": "analyze_seams", + "field": "level_step_db", + "comparator": ">", + "threshold": 3.0, + "severity": "warn", + "says": "the level steps this much across the seam", + }, + { + "name": "seam_click", + "probe": "analyze_seams", + "field": "click_db", + "comparator": ">", + "threshold": 12.0, + "severity": "warn", + "says": "the join peaks this far above the audio either side of it - an audible click", + }, + { + "name": "seam_hole", + "probe": "analyze_seams", + "field": "floor_dbfs", + "comparator": "<", + "threshold": -50.0, + "severity": "warn", + "says": "the track drops out at the join while both sides are voiced", + "unless": f"either side's rms is at or below {HOLE_VOICED_DBFS} dBFS", + }, + { + "name": "seam_frame_jump", + "probe": "analyze_seams", + "field": "jump_ratio", + "comparator": ">", + "threshold": 8.0, + "severity": "info", + "says": "the picture changes this many times more across the seam than inside either shot", + "unless": "the incoming shot is marked hard_cut: true", + }, + { + "name": "sync_drift", + "probe": "analyze_sync_drift", + "field": "end_offset_ms", + "comparator": ">", + "threshold": 40.0, + "severity": "warn", + "magnitude": True, + "says": "by this shot's end the audio sits this far off the picture", + }, + { + "name": "sync_length", + "probe": "analyze_sync_drift", + "field": "length_delta_ms", + "comparator": ">", + "threshold": 40.0, + "severity": "warn", + "magnitude": True, + "says": "the soundtrack and the picture differ in length by this much", + }, +) + +RULES_BY_NAME = {rule["name"]: rule for rule in RULES} + + +def rules_for(probe): + """The rules one probe's measurements are read against, in table order.""" + return [rule for rule in RULES if rule["probe"] == probe] + + +def crosses(rule, value): + """Whether a measured value crosses a rule's threshold. None never does - + a measurement that could not be taken (a silent window, no soundtrack) + is not a finding.""" + if value is None: + return False + if rule.get("magnitude"): + value = abs(value) + return COMPARATORS[rule["comparator"]](value, rule["threshold"]) + + +def finding(rule, value, at): + """A crossing as the probe reports it.""" + return { + "rule": rule["name"], + "severity": rule["severity"], + "at": at, + "value": value, + "threshold": rule["threshold"], + "says": rule["says"], + } + + +__all__ = [ + "HOLE_VOICED_DBFS", + "RULES", + "RULES_BY_NAME", + "SEVERITIES", + "crosses", + "finding", + "rules_for", +] diff --git a/dw/introspection.py b/dw/introspection.py index 8a0bdd87..cf4eb50a 100644 --- a/dw/introspection.py +++ b/dw/introspection.py @@ -370,14 +370,30 @@ def unknown_call_arguments(name, argument_names): def list_tasks(): - """Every task command a workflow's task step can name.""" - from .tasks.task import _COMMAND_REGISTRY, _VIDEO_PROCESSOR_COMMANDS + """Every task command a workflow's task step can name. + + `assessment` names the probes among the commands (#387) - the ones that + answer a JSON document of measurements about a finished file rather than + make one - so a caller looking for a way to check a cut finds them + without reading every command's schema. They stay in `commands` too, + since a step still names one as its `command`. + """ + from .tasks.task import ( + _COMMAND_INFO, + _COMMAND_REGISTRY, + _VIDEO_PROCESSOR_COMMANDS, + ) from .tasks.image_utils import available_processors return { "commands": sorted(_COMMAND_REGISTRY.keys()), "image_processors": sorted(available_processors()), "video_processors": list(_VIDEO_PROCESSOR_COMMANDS), + "assessment": sorted( + name + for name, info in _COMMAND_INFO.items() + if info.get("returns") == "json" + ), } diff --git a/dw/runs.py b/dw/runs.py index 6d252e05..0ab2d4ee 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -680,3 +680,47 @@ def recorded_shots(output_root, relative_path): if shots: return shots return None + + +# How far up from a file its run's manifest can sit: the run directory, a +# subfolder (`final/`) and the subfolder's own nesting, which +# SUBFOLDER_PATTERN caps well below this +MANIFEST_SEARCH_DEPTH = 8 + + +def shots_beside(path): + """The shots a run's manifest records for a file named by its absolute path. + + `recorded_shots` needs the output root and a gallery name; a task that + was handed a resolved `output:` path has neither, only the file. The run + directory is the nearest parent holding a manifest, and the file's name + inside it is what the manifest records. None when no manifest is found + within MANIFEST_SEARCH_DEPTH parents, or it records no shots for the + file. + """ + from .shots import shots_for_file + + path = os.path.abspath(path) + run_dir = os.path.dirname(path) + for _ in range(MANIFEST_SEARCH_DEPTH): + if os.path.isfile(os.path.join(run_dir, MANIFEST_FILE_NAME)): + break + parent = os.path.dirname(run_dir) + if parent == run_dir: + return None + run_dir = parent + else: + return None + manifest = _read_manifest(run_dir) + if manifest is None: + return None + own = os.path.relpath(path, run_dir).replace(os.sep, "/") + for entry in manifest.get("steps") or []: + if not isinstance(entry, dict) or entry.get("reused"): + continue + files = entry.get("files") or [] + if own in files: + shots = shots_for_file(entry.get("shots"), own, files) + if shots: + return shots + return None diff --git a/dw/scalar_result_validation.py b/dw/scalar_result_validation.py index 8628c91d..0faf482d 100644 --- a/dw/scalar_result_validation.py +++ b/dw/scalar_result_validation.py @@ -10,16 +10,24 @@ own declared `returns` kind (`dw/tasks/task.py`'s `register_command`) rather than a name match on `judge`, so a future scalar-returning task is covered by declaring itself rather than by a second special case here. + +A "json" command (the assessment probes, #387) answers a dict of +measurements, which `Result.save` writes whole only under +`application/json`; any other content type explodes the dict key by key +into files, or dies on a number. So a `result` on one is allowed and must +say `application/json`. """ from .for_each import MEMBER_SEPARATOR, render_path from .tasks.task import task_command_info RESULT_KEY = "result" +JSON_CONTENT_TYPE = "application/json" def scalar_result_errors(workflow_definition, source_indices=None): - """Every `result` block on a scalar-returning command, as [{path, message}]. + """Every `result` block a command's `returns` kind cannot save, as + [{path, message}]. The definition handed here has already been substituted and expanded, so a step's `task.command` is literal. `source_indices`, when given, is @@ -49,7 +57,21 @@ def scalar_result_errors(workflow_definition, source_indices=None): info = task_command_info(command) except ValueError: continue - if info.get("returns") != "scalar": + returns = info.get("returns") + if returns == "json": + if result.get("content_type") == JSON_CONTENT_TYPE: + continue + message = ( + f"{command} answers a JSON document - 'result' must set " + f"content_type '{JSON_CONTENT_TYPE}', not " + f"{result.get('content_type')!r}" + ) + elif returns == "scalar": + message = ( + f"{command} returns a number, not an artifact - " + f"'result' cannot be saved" + ) + else: continue source = ( @@ -66,10 +88,7 @@ def scalar_result_errors(workflow_definition, source_indices=None): errors.append( { "path": render_path(("steps", source, RESULT_KEY)), - "message": ( - f"{command} returns a number, not an artifact - " - f"'result' cannot be saved{where}" - ), + "message": f"{message}{where}", } ) return errors diff --git a/dw/tasks/assess.py b/dw/tasks/assess.py new file mode 100644 index 00000000..e867fab0 --- /dev/null +++ b/dw/tasks/assess.py @@ -0,0 +1,660 @@ +"""Assessment probes: measure a finished cut and say where to look (#387). + +Three read-only task commands - `analyze_shots`, `analyze_seams` and +`analyze_sync_drift` - each take a video (a path, or the AudioVideo an +earlier step returned) and answer a JSON-safe dict: every measurement it +took, the `findings` its rules raised (`dw/assessment_rules.py`), the names +of the rules it read the measurements against and where the shot +boundaries came from. Nothing acts on a finding; they are places to look. + +The reader streams. A cut is minutes of full-resolution picture, and every +one of these measurements needs only a 64x36 grey thumbnail of each frame +and the soundtrack, so `read_media` decodes the file once and keeps exactly +that - never a frame list, and never `load_audio`'s 0.25 s fit of the track +to the picture, which would move the very sample count `analyze_sync_drift` +is measuring. The soundtrack is trimmed to the audio stream's own duration, +which the container already reports net of the encoder's priming, so a +lossy codec's padding is not read as drift. + +Shot boundaries come, in order, from the `shots` argument, the AudioVideo's +own `shots`, the run manifest beside a file (`dw.runs.shots_beside`), and +otherwise the whole file is one shot. `shots_source` says which. A shot +record may carry `hard_cut: true` - a cut meant as a cut - and the frame +jump rule does not fire at the seam that shot opens. +""" + +import logging +import math + +import numpy + +from ..assessment_rules import HOLE_VOICED_DBFS, crosses, finding, rules_for + +logger = logging.getLogger("dw") + +THUMB_WIDTH = 64 +THUMB_HEIGHT = 36 + +# Audio windows around a seam, in seconds +LEVEL_WINDOW = 0.25 +FLOOR_WINDOW = 0.02 +CLICK_WINDOW = 0.002 +CLICK_NEIGHBOUR_WINDOW = 0.01 + +# The inter-frame difference a shot is expected to show, below which its +# own motion is not a baseline: a static shot's 90th percentile is ~0, and +# dividing a seam's change by it would call any change at all a jump. +# Grey levels on the 0-255 scale. +TYPICAL_DELTA_FLOOR = 2.0 +TYPICAL_DELTA_PERCENTILE = 90 + +# A ratio against a zero neighbour has no size; report it capped +CLICK_CAP_DB = 120.0 + +_SILENCE = 1e-10 + + +class Media: + """What the probes read from a video: thumbnails, soundtrack, timing. + + thumbs: (frames, THUMB_HEIGHT, THUMB_WIDTH) uint8 grey, or None + audio: (channels, samples) float32, or None + video_seconds / audio_seconds: each stream's own duration, or None + """ + + def __init__( + self, + thumbs, + audio, + sample_rate, + fps, + video_seconds=None, + audio_seconds=None, + shots=None, + ): + self.thumbs = thumbs + self.audio = audio + self.sample_rate = sample_rate + self.fps = fps + self.video_seconds = video_seconds + self.audio_seconds = audio_seconds + self.shots = shots + + @property + def frame_count(self): + return 0 if self.thumbs is None else int(self.thumbs.shape[0]) + + +def _stream_seconds(stream): + if stream is None or stream.duration is None or stream.time_base is None: + return None + return float(stream.duration * stream.time_base) + + +def _as_float_samples(samples): + if samples.dtype.kind == "u": + iinfo = numpy.iinfo(samples.dtype) + half = (iinfo.max + 1) / 2 + return (samples.astype(numpy.float32) - half) / half + if samples.dtype.kind == "i": + return samples.astype(numpy.float32) / numpy.iinfo(samples.dtype).max + return samples.astype(numpy.float32) + + +def read_media(path): + """Stream a video file once into a Media. + + Each decoded frame is reduced to a grey thumbnail as it arrives and then + dropped, so memory holds thumbnails, not pictures. The soundtrack is + kept whole (the probes cut windows out of it anywhere) and trimmed to + the audio stream's reported duration. + """ + import av + + from ..media_info import _as_frame_samples + + thumbs = [] + chunks = [] + with av.open(path) as container: + video = container.streams.video[0] if container.streams.video else None + audio = container.streams.audio[0] if container.streams.audio else None + if video is None and audio is None: + raise ValueError(f"{path} has neither a video nor an audio stream") + fps = float(video.average_rate) if video and video.average_rate else None + video_seconds = _stream_seconds(video) + audio_seconds = _stream_seconds(audio) + channels = int(audio.channels) if audio is not None else 0 + streams = [s for s in (video, audio) if s is not None] + for frame in container.decode(*streams): + if isinstance(frame, av.VideoFrame): + thumbs.append( + frame.reformat( + width=THUMB_WIDTH, height=THUMB_HEIGHT, format="gray" + ).to_ndarray()[:THUMB_HEIGHT, :THUMB_WIDTH] + ) + elif isinstance(frame, av.AudioFrame): + samples = _as_float_samples(frame.to_ndarray()) + chunks.append(_as_frame_samples(samples, channels)) + sample_rate = int(audio.rate) if audio is not None else None + + waveform = None + if chunks: + waveform = numpy.ascontiguousarray(numpy.concatenate(chunks, axis=0).T) + if audio_seconds is not None: + waveform = waveform[:, : int(round(audio_seconds * sample_rate))] + if video_seconds is None and fps and thumbs: + video_seconds = len(thumbs) / fps + return Media( + numpy.stack(thumbs) if thumbs else None, + waveform, + sample_rate, + fps, + video_seconds, + audio_seconds, + ) + + +def _thumb(frame): + """One in-memory frame (PIL image, or an HxWxC array or tensor, uint8 or + float in [0, 1]) as a grey thumbnail.""" + from PIL import Image + + if not isinstance(frame, Image.Image): + array = frame + if hasattr(array, "detach"): + array = array.detach().cpu().numpy() + array = numpy.asarray(array) + if array.dtype.kind == "f": + array = numpy.clip(array * 255.0, 0, 255) + array = array.astype(numpy.uint8) + if array.ndim == 3 and array.shape[-1] == 1: + array = array[..., 0] + frame = Image.fromarray(array) + return numpy.asarray( + frame.convert("L").resize((THUMB_WIDTH, THUMB_HEIGHT), Image.BILINEAR), + dtype=numpy.uint8, + ) + + +def _in_memory_thumbs(frames): + """Thumbnails of an AudioVideo's frames, one frame at a time: a list, an + (N, H, W, C) array, or a SegmentedFrames replaying (N, H, W, 3) chunks.""" + from ..pipeline_processors.chain import SegmentedFrames + + thumbs = [] + if isinstance(frames, SegmentedFrames): + for chunk in frames: + for frame in chunk: + thumbs.append(_thumb(frame)) + else: + for frame in frames: + thumbs.append(_thumb(frame)) + return numpy.stack(thumbs) if thumbs else None + + +def media_from(video): + """A Media from a path or an in-memory AudioVideo.""" + if isinstance(video, str): + from ..locations import validate_media_path + + return read_media(validate_media_path(video, None, "a video to assess")) + if hasattr(video, "frames") or hasattr(video, "audio"): + from .audio_utils import as_channels_samples + + audio = getattr(video, "audio", None) + waveform = None if audio is None else as_channels_samples(audio) + thumbs = ( + _in_memory_thumbs(video.frames) + if getattr(video, "frames", None) is not None + else None + ) + sample_rate = getattr(video, "sample_rate", None) + fps = getattr(video, "fps", None) + media = Media( + thumbs, + waveform, + sample_rate, + fps, + video_seconds=(thumbs.shape[0] / fps) + if thumbs is not None and fps + else None, + audio_seconds=( + waveform.shape[1] / sample_rate + if waveform is not None and sample_rate + else None + ), + shots=getattr(video, "shots", None), + ) + return media + raise ValueError( + "a probe takes a video file's path or the video an earlier step " + f"returned, not {type(video).__name__}" + ) + + +def resolve_shots(video, media, shots=None): + """The shot records to measure against, and where they came from.""" + if shots: + return [dict(shot) for shot in shots], "argument" + if media.shots: + return [dict(shot) for shot in media.shots], "artifact" + if isinstance(video, str): + from ..locations import validate_media_path + from ..runs import shots_beside + + recorded = shots_beside(validate_media_path(video, None, "a video to assess")) + if recorded: + return [dict(shot) for shot in recorded], "manifest" + return None, "none" + + +def _whole_file_shot(media): + samples = media.audio.shape[1] if media.audio is not None else None + return { + "name": "whole", + "start_frame": 0, + "num_frames": media.frame_count, + "start_sample": 0 if samples is not None else None, + "num_samples": samples, + } + + +def _sample_span(shot, media): + """A shot's (start, count, source) on the soundtrack: recorded, else + derived from its frames.""" + start = shot.get("start_sample") + count = shot.get("num_samples") + if start is not None and count is not None: + return int(start), int(count), "recorded" + if media.fps and media.sample_rate: + scale = media.sample_rate / media.fps + return ( + int(round(shot.get("start_frame", 0) * scale)), + int(round(shot.get("num_frames", 0) * scale)), + "derived", + ) + return None, None, None + + +def _db(value): + return None if value is None or value <= _SILENCE else 20.0 * math.log10(value) + + +def _rms(window): + if window is None or window.size == 0: + return None + return math.sqrt(float(numpy.mean(numpy.square(window, dtype=numpy.float64)))) + + +def _peak(window): + if window is None or window.size == 0: + return None + return float(numpy.max(numpy.abs(window))) + + +def _clip(media, start, end): + total = media.audio.shape[1] + start = max(0, min(total, int(start))) + end = max(start, min(total, int(end))) + return media.audio[:, start:end] + + +def _round(value, places=2): + return None if value is None else round(float(value), places) + + +def _findings(probe, record, at, skip=()): + found = [] + for rule in rules_for(probe): + if rule["name"] in skip or rule["field"] not in record: + continue + if crosses(rule, record[rule["field"]]): + found.append(finding(rule, record[rule["field"]], at)) + return found + + +def _answer(probe, measurements, findings, shots_source): + return { + **measurements, + "findings": findings, + "rules_applied": [rule["name"] for rule in rules_for(probe)], + "shots_source": shots_source, + } + + +def analyze_shots(video, shots=None): + """Task command: each shot's level and spectral balance, and how far + apart the shots sit. + + Args: + video: A video file's path, or the video an earlier step returned + shots: Shot records to measure by, overriding any the video carries + + Returns: + {shots: [{name, peak_dbfs, rms_dbfs, crest_db, low_dbfs, mid_dbfs, + high_dbfs, samples}], rms_range_db, findings, rules_applied, + shots_source} + """ + from .audio_utils import _spectral_balance + + media = media_from(video) + records, source = resolve_shots(video, media, shots) + if media.audio is None: + return _answer( + "analyze_shots", + {"shots": [], "rms_range_db": None, "has_audio": False}, + [], + source, + ) + records = records or [_whole_file_shot(media)] + + measured = [] + for shot in records: + start, count, samples_source = _sample_span(shot, media) + window = _clip(media, start, start + count) if start is not None else None + peak = _db(_peak(window)) + rms = _db(_rms(window)) + balance = ( + _spectral_balance(window, media.sample_rate) + if window is not None and window.size + else {"low_dbfs": None, "mid_dbfs": None, "high_dbfs": None} + ) + measured.append( + { + "name": shot.get("name"), + "peak_dbfs": _round(peak), + "rms_dbfs": _round(rms), + "crest_db": _round(None if peak is None or rms is None else peak - rms), + **{key: _round(value) for key, value in balance.items()}, + "samples": samples_source, + } + ) + + voiced = [shot for shot in measured if shot["rms_dbfs"] is not None] + rms_range = None + at = None + if len(voiced) > 1: + loudest = max(voiced, key=lambda shot: shot["rms_dbfs"]) + quietest = min(voiced, key=lambda shot: shot["rms_dbfs"]) + rms_range = _round(loudest["rms_dbfs"] - quietest["rms_dbfs"]) + at = {"between": [loudest["name"], quietest["name"]]} + elif voiced: + rms_range = 0.0 + answer = {"shots": measured, "rms_range_db": rms_range, "has_audio": True} + return _answer( + "analyze_shots", + answer, + _findings("analyze_shots", answer, at) if at else [], + source, + ) + + +def _band_shares(window, sample_rate): + from .audio_utils import _spectral_balance + + if window is None or window.size == 0: + return None + bands = _spectral_balance(window, sample_rate) + energies = { + key: (10.0 ** (value / 10.0) if value is not None else 0.0) + for key, value in bands.items() + } + total = sum(energies.values()) + if total <= 0.0: + return None + return {key: value / total for key, value in energies.items()} + + +def _seam_audio(media, before_end, after_start): + """Audio measurements at a seam: the level windows end at `before_end` + and open at `after_start` (the same sample at a cut, either side of the + fade at a dissolve), and the join is what lies between them - or the + FLOOR_WINDOW centred on the cut.""" + rate = media.sample_rate + level = int(round(LEVEL_WINDOW * rate)) + before = _clip(media, before_end - level, before_end) + after = _clip(media, after_start, after_start + level) + before_rms = _db(_rms(before)) + after_rms = _db(_rms(after)) + + centre = (before_end + after_start) // 2 + half_floor = max(1, int(round(FLOOR_WINDOW * rate / 2))) + join = ( + _clip(media, before_end, after_start) + if after_start - before_end > 2 * half_floor + else _clip(media, centre - half_floor, centre + half_floor) + ) + floor = _db(_rms(join)) + + half_click = max(1, int(round(CLICK_WINDOW * rate / 2))) + neighbour = int(round(CLICK_NEIGHBOUR_WINDOW * rate)) + click_peak = _peak(_clip(media, centre - half_click, centre + half_click)) + neighbour_peak = max( + _peak(_clip(media, centre - half_click - neighbour, centre - half_click)) + or 0.0, + _peak(_clip(media, centre + half_click, centre + half_click + neighbour)) + or 0.0, + ) + click = None + if click_peak is not None and click_peak > _SILENCE: + click = ( + CLICK_CAP_DB + if neighbour_peak <= _SILENCE + else min(CLICK_CAP_DB, 20.0 * math.log10(click_peak / neighbour_peak)) + ) + + shares_before = _band_shares(before, rate) + shares_after = _band_shares(after, rate) + spectral_shift = ( + sum(abs(shares_after[key] - shares_before[key]) for key in shares_before) / 2.0 + if shares_before and shares_after + else None + ) + return { + "before_rms_dbfs": _round(before_rms), + "after_rms_dbfs": _round(after_rms), + "level_step_db": _round( + None + if before_rms is None or after_rms is None + else abs(after_rms - before_rms) + ), + "floor_dbfs": _round(floor), + "click_db": _round(click), + "spectral_shift": _round(spectral_shift, 3), + } + + +def _typical_delta(thumbs, start, end): + """The TYPICAL_DELTA_PERCENTILE of frame-to-frame change inside + thumbs[start:end], or None for fewer than two frames.""" + span = thumbs[max(0, start) : max(0, end)] + if span.shape[0] < 2: + return None + deltas = numpy.abs(numpy.diff(span.astype(numpy.int16), axis=0)).mean(axis=(1, 2)) + return float(numpy.percentile(deltas, TYPICAL_DELTA_PERCENTILE)) + + +def _seam_video(media, previous, shot, fade): + """Picture measurements at the seam `shot` opens: the largest single-frame + change across it (at a dissolve, across the whole fade - each step of a + fade is small, which is what a dissolve is), against the larger of the + two shots' own typical change, floored.""" + thumbs = media.thumbs + start = int(shot.get("start_frame", 0)) + first = max(0, start - 1) + last = min(thumbs.shape[0] - 1, start + max(0, fade - 1) if fade else start) + if last <= first: + return {"frame_delta": None, "typical_delta": None, "jump_ratio": None} + across = numpy.abs( + numpy.diff(thumbs[first : last + 1].astype(numpy.int16), axis=0) + ).mean(axis=(1, 2)) + frame_delta = float(across.max()) + typical = [ + _typical_delta( + thumbs, + int(previous.get("start_frame", 0)) + + int(previous.get("overlap_frames") or 0), + start, + ), + _typical_delta(thumbs, start + fade, start + int(shot.get("num_frames", 0))), + ] + typical = max( + [TYPICAL_DELTA_FLOOR] + [value for value in typical if value is not None] + ) + return { + "frame_delta": _round(frame_delta), + "typical_delta": _round(typical), + "jump_ratio": _round(frame_delta / typical), + } + + +def analyze_seams(video, shots=None): + """Task command: measure every seam between shots, audio and picture. + + Args: + video: A video file's path, or the video an earlier step returned + shots: Shot records to measure by, overriding any the video carries. + A shot marked `hard_cut: true` opens a seam meant as a cut + + Returns: + {seams: [{seam, between, seconds, kind, level_step_db, floor_dbfs, + click_db, spectral_shift, before_rms_dbfs, after_rms_dbfs, + frame_delta, typical_delta, jump_ratio}], findings, rules_applied, + shots_source} + """ + media = media_from(video) + records, source = resolve_shots(video, media, shots) + if not records or len(records) < 2: + return _answer("analyze_seams", {"seams": []}, [], source) + + seams = [] + findings = [] + for index in range(1, len(records)): + previous, shot = records[index - 1], records[index] + fade = int(shot.get("overlap_frames") or 0) + start_frame = int(shot.get("start_frame", 0)) + seam_frame = start_frame + fade / 2.0 + seconds = seam_frame / media.fps if media.fps else None + record = { + "seam": index, + "between": [previous.get("name"), shot.get("name")], + "seconds": _round(seconds, 3), + "kind": "dissolve" if fade else "cut", + "hard_cut": bool(shot.get("hard_cut")), + } + skip = set() + if media.audio is not None and media.sample_rate: + start, _count, _source = _sample_span(shot, media) + if start is not None: + fade_samples = ( + int(round(fade / media.fps * media.sample_rate)) + if fade and media.fps + else 0 + ) + record.update(_seam_audio(media, start, start + fade_samples)) + if ( + record["before_rms_dbfs"] is None + or record["after_rms_dbfs"] is None + or record["before_rms_dbfs"] <= HOLE_VOICED_DBFS + or record["after_rms_dbfs"] <= HOLE_VOICED_DBFS + ): + skip.add("seam_hole") + else: + skip.add("seam_hole") + if media.thumbs is not None: + record.update(_seam_video(media, previous, shot, fade)) + if record["hard_cut"]: + skip.add("seam_frame_jump") + seams.append(record) + findings.extend( + _findings( + "analyze_seams", + record, + { + "seam": index, + "between": record["between"], + "seconds": record["seconds"], + }, + skip, + ) + ) + return _answer("analyze_seams", {"seams": seams}, findings, source) + + +def analyze_sync_drift(video, shots=None): + """Task command: how far the soundtrack sits from the picture, shot by + shot and over the whole file. + + Args: + video: A video file's path, or the video an earlier step returned + shots: Shot records to measure by, overriding any the video carries + + Returns: + {shots: [{name, start_offset_ms, end_offset_ms}], max_offset_ms, + video_seconds, audio_seconds, length_delta_ms, findings, + rules_applied, shots_source} + """ + media = media_from(video) + records, source = resolve_shots(video, media, shots) + measured = [] + findings = [] + rate = media.sample_rate + fps = media.fps + for shot in records or []: + start, count = shot.get("start_sample"), shot.get("num_samples") + if start is None or count is None or not rate or not fps: + continue + start_frame = int(shot.get("start_frame", 0)) + end_frame = start_frame + int(shot.get("num_frames", 0)) + record = { + "name": shot.get("name"), + "start_offset_ms": _round((int(start) / rate - start_frame / fps) * 1000.0), + "end_offset_ms": _round( + ((int(start) + int(count)) / rate - end_frame / fps) * 1000.0 + ), + } + measured.append(record) + findings.extend( + _findings( + "analyze_sync_drift", + record, + { + "shot": record["name"], + "seconds": _round(end_frame / fps, 3), + }, + ) + ) + + length_delta = None + if media.audio_seconds is not None and media.video_seconds is not None: + length_delta = _round((media.audio_seconds - media.video_seconds) * 1000.0) + answer = { + "shots": measured, + "max_offset_ms": ( + max((record["end_offset_ms"] for record in measured), key=abs) + if measured + else None + ), + "video_seconds": _round(media.video_seconds, 4), + "audio_seconds": _round(media.audio_seconds, 4), + "length_delta_ms": length_delta, + } + findings.extend( + _findings( + "analyze_sync_drift", + {"length_delta_ms": length_delta}, + {"file": True}, + ) + ) + return _answer("analyze_sync_drift", answer, findings, source) + + +__all__ = [ + "Media", + "analyze_seams", + "analyze_shots", + "analyze_sync_drift", + "media_from", + "read_media", + "resolve_shots", +] diff --git a/dw/tasks/task.py b/dw/tasks/task.py index c633ba4e..5f240f52 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -66,6 +66,10 @@ def register_command( `validation_errors` (dw/scalar_result_validation.py, #212) rather than reaching `save_artifact` at run time, where a float has nothing left identifying which command produced it + - or "json" for a command answering a JSON-safe dict of + measurements (the assessment probes, `dw/tasks/assess.py`): its + `result` may only be `application/json`, since any other content + type would explode the dict key by key into files summary: Overrides the command's `get_task` summary, which otherwise reads the implementation function's docstring. For a command whose handler dispatches its implementation per video frame @@ -335,6 +339,38 @@ def _handle_analyze_audio(task, arguments, previous_pipelines): return analyze_audio(**arguments) +@register_command( + "analyze_shots", implementation="dw.tasks.assess.analyze_shots", returns="json" +) +def _handle_analyze_shots(task, arguments, previous_pipelines): + """Measure each shot of a cut's soundtrack and how far apart they sit""" + from .assess import analyze_shots + + return analyze_shots(**arguments) + + +@register_command( + "analyze_seams", implementation="dw.tasks.assess.analyze_seams", returns="json" +) +def _handle_analyze_seams(task, arguments, previous_pipelines): + """Measure every seam of a cut - level step, hole, click, frame jump""" + from .assess import analyze_seams + + return analyze_seams(**arguments) + + +@register_command( + "analyze_sync_drift", + implementation="dw.tasks.assess.analyze_sync_drift", + returns="json", +) +def _handle_analyze_sync_drift(task, arguments, previous_pipelines): + """Measure how far a cut's soundtrack sits from its picture""" + from .assess import analyze_sync_drift + + return analyze_sync_drift(**arguments) + + @register_command("compose_text", implementation="dw.tasks.compose_text.compose_text") def _handle_compose_text(task, arguments, previous_pipelines): """Join parts written once into one block of text""" diff --git a/ui/src/lib/api.ts b/ui/src/lib/api.ts index f3261548..7a567808 100644 --- a/ui/src/lib/api.ts +++ b/ui/src/lib/api.ts @@ -344,6 +344,7 @@ export const api = { commands: string[] image_processors: string[] video_processors: string[] + assessment: string[] }>('/api/tasks'), describeTask: (command: string) => request(`/api/tasks/${encodeURIComponent(command)}`), From f34d7e6fd139f540f1fe2adb90e6251903b7eb36 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 22:51:35 -0500 Subject: [PATCH 080/181] feat(engine): #387 - probe tests, rules table pinned to probe fields, docs Co-Authored-By: Claude Opus 5.5 --- CLAUDE.md | 15 + docs/TASKS.md | 112 +++++ docs/WORKFLOW_GUIDE.md | 9 + tests/test_assess.py | 560 +++++++++++++++++++++++++ tests/test_assessment_rules.py | 190 +++++++++ tests/test_scalar_result_validation.py | 162 +++++++ 6 files changed, 1048 insertions(+) create mode 100644 tests/test_assess.py create mode 100644 tests/test_assessment_rules.py diff --git a/CLAUDE.md b/CLAUDE.md index dbefee67..01644e69 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -611,6 +611,21 @@ same reason - default setup cannot load a pack. - **Step cache**: a process-wide singleton (`dw/step_cache.py`) consulted by every `Workflow.run`, including server jobs; entries are keyed by `(workflow id, step name)` and validated against the output *root*, never the per-run directory - a run directory is new every execution and would defeat the cache; disabled entirely when the workflow sets no `seed`; a hit reports the earlier run's files with `reused: true` and writes nothing new; `memory clear` drops it. This is why "Run again" on a seeded workflow finishes instantly and generates nothing - the job page says so when every step was reused, and `POST /api/jobs/{id}/rerun` with `{"new_seed": true}` (MCP `rerun_job(new_seed=True)`) draws a fresh seed into the workflow's seed variable, which is the way to get a different image +- **Assessment probes measure a finished file and say where to look, and + decide nothing** (`dw/tasks/assess.py`, #387) - `analyze_shots`, + `analyze_seams` and `analyze_sync_drift` each read a video streaming + (thumbnails only, never the full frame list, so a long cut is cheap) and + answer a JSON dict of measurements plus `findings`, the ones that crossed a + threshold in `dw/assessment_rules.py`'s table; nothing in the engine acts + on a finding. These are the `returns: "json"` task kind, listed separately + in `list_tasks`' `assessment` (probes stay in `commands` too), and a step + on one must save `"result": {"content_type": "application/json"}` - + anything else fails validation. Shot boundaries resolve in order: the step's `shots` argument, + the video's own carried shots, the run manifest beside the file, else the + whole file as one shot (`shots_source` says which). A shot's `hard_cut: + true` field suppresses the `seam_frame_jump` rule at the seam it opens - a + cut meant as a cut. `tests/test_assessment_rules.py` pins the rules table + to real probe fields, so a rename cannot leave a rule reading nothing. ## JSON Workflow Structure diff --git a/docs/TASKS.md b/docs/TASKS.md index bc9b388c..2ce3b2e4 100644 --- a/docs/TASKS.md +++ b/docs/TASKS.md @@ -857,6 +857,118 @@ scale as `rms_dbfs` (their powers sum to it), so the loudest band sits near `threshold_dbfs`. A silent track, or a band with no content at the track's sample rate, reads as `null` rather than `-inf`. +## Assessment Probes + +Three read-only commands measure a finished cut and say where to look - +`analyze_shots`, `analyze_seams`, `analyze_sync_drift`. Each takes a video +(a path, or the video an earlier step returned) and answers one JSON +document: every measurement it took, plus `findings` (the measurements that +crossed a rule in the table below), `rules_applied` (the rule names the probe +checked) and `shots_source` (where the shot list came from). A probe reads +the file streaming - a 64x36 grey thumbnail per frame and the soundtrack, +never a full frame list - so it runs on a cut of any length. + +Findings are places to look, not verdicts: nothing in the engine acts on +one, no run fails for one, and a finding someone has looked at and accepted +is simply left alone. + +A probe's `result` must save as JSON: + +```json +{ + "task": { + "command": "analyze_seams", + "arguments": { + "video": "output:/latest/final/cut.mp4" + } + }, + "result": { "content_type": "application/json" } +} +``` + +Any other `content_type` (or none) fails validation - a JSON document can +only be saved whole under `application/json`; every other content type +would explode it key by key or die trying to write a number. + +Shot boundaries come, in order: the step's own `shots` argument, the shots +carried by the video an earlier step returned, the run manifest beside the +file, and otherwise the whole file is treated as one shot. `shots_source` +reports which - `argument`, `artifact`, `manifest`, or `none`. + +### analyze_shots + +Each shot's level and spectral balance, and how far apart the shots sit: + +| Field | Meaning | +| ----- | ------- | +| `shots[].name` | The shot's name | +| `shots[].peak_dbfs` | Peak level within the shot | +| `shots[].rms_dbfs` | RMS level within the shot | +| `shots[].crest_db` | `peak_dbfs` minus `rms_dbfs` | +| `shots[].low_dbfs` / `mid_dbfs` / `high_dbfs` | Spectral balance (20-250 Hz / 250-4000 Hz / 4000-20000 Hz), on the same scale as `rms_dbfs` | +| `shots[].samples` | Whether the shot's sample span was `recorded` (carried by the shot record) or `derived` (scaled from its frames) | +| `rms_range_db` | The spread between the loudest and quietest voiced shot | +| `has_audio` | Whether the file carries a soundtrack at all | + +### analyze_seams + +Every seam between shots, audio and picture: + +| Field | Meaning | +| ----- | ------- | +| `seams[].seam` | The seam's index (1-based) | +| `seams[].between` | `[previous shot name, next shot name]` | +| `seams[].seconds` | Where the seam sits in the file | +| `seams[].kind` | `cut` or `dissolve` (a dissolve has `overlap_frames`) | +| `seams[].hard_cut` | Whether the incoming shot is marked `hard_cut: true` | +| `seams[].before_rms_dbfs` / `after_rms_dbfs` | RMS level either side of the seam | +| `seams[].level_step_db` | The absolute level jump across the seam | +| `seams[].floor_dbfs` | RMS level of the join itself (the fade, or a short window centred on a cut) | +| `seams[].click_db` | How far a spike at the join peaks above its immediate neighbours | +| `seams[].spectral_shift` | How much the low/mid/high balance shifts across the seam (0-1) | +| `seams[].frame_delta` | The largest single-frame picture change across the seam | +| `seams[].typical_delta` | The larger shot's own typical frame-to-frame change, floored | +| `seams[].jump_ratio` | `frame_delta` divided by `typical_delta` | + +### analyze_sync_drift + +How far the soundtrack sits from the picture, shot by shot and over the +whole file: + +| Field | Meaning | +| ----- | ------- | +| `shots[].name` | The shot's name | +| `shots[].start_offset_ms` | How far the audio sits from the picture at the shot's start | +| `shots[].end_offset_ms` | How far the audio sits from the picture at the shot's end | +| `max_offset_ms` | The largest `end_offset_ms` across all shots, by magnitude | +| `video_seconds` / `audio_seconds` | Each stream's own duration | +| `length_delta_ms` | `audio_seconds` minus `video_seconds` | + +### Rules + +Each rule names the probe and field it reads, how the value is compared to +its threshold, and the severity of a crossing: + +| Rule | Probe | Field | Threshold | Severity | +| ---- | ----- | ----- | --------- | -------- | +| `shot_level_spread` | `analyze_shots` | `rms_range_db` | >= 6.0 dB | warn | +| `seam_level_step` | `analyze_seams` | `level_step_db` | > 3.0 dB | warn | +| `seam_click` | `analyze_seams` | `click_db` | > 12.0 dB | warn | +| `seam_hole` | `analyze_seams` | `floor_dbfs` | < -50.0 dBFS | warn | +| `seam_frame_jump` | `analyze_seams` | `jump_ratio` | > 8.0 | info | +| `sync_drift` | `analyze_sync_drift` | `end_offset_ms` | > 40.0 ms (magnitude) | warn | +| `sync_length` | `analyze_sync_drift` | `length_delta_ms` | > 40.0 ms (magnitude) | warn | + +Two rules carry a guard beyond the threshold: `seam_hole` only fires while +both sides of the seam are voiced above -30 dBFS (a quiet join between two +quiet shots is not a hole, it's a pause the shots themselves hold), and +`seam_frame_jump` is skipped at a seam whose incoming shot is marked +`hard_cut: true` - a cut meant as a cut. + +`list_tasks` names the probes in their own `assessment` list, alongside +`commands`, so a caller looking for a way to check a cut can find them +without reading every command's schema. + ## Data Gathering ### gather_images diff --git a/docs/WORKFLOW_GUIDE.md b/docs/WORKFLOW_GUIDE.md index 29fb78f6..84fa3589 100644 --- a/docs/WORKFLOW_GUIDE.md +++ b/docs/WORKFLOW_GUIDE.md @@ -798,6 +798,15 @@ that works. Supported content types: `image/jpeg`, `image/png`, `image/webp`, `image/gif`, `video/mp4`, `audio/wav`, `audio/flac`, `audio/mpeg` (mp3), `audio/ogg`, `audio/opus`, `audio/aiff`, `application/json`, `text/plain` (plus the common aliases `audio/x-wav`, `audio/mp3`, `audio/vorbis`). +A task command's implementation declares what it hands back - most answer an +`artifact` (a file `result` saves in one of the media content types above), +some (`judge`) answer a bare `scalar` that cannot be saved at all, and some +(the assessment probes in [TASKS.md](TASKS.md)) answer a `json` document - +every measurement taken, in one dict. A step on a `json` command must set +`content_type` to `application/json`, which saves it as one document; a step on a `scalar` command +may not carry a `result` at all. Both are checked in validation, by the +command's own declared kind rather than a name match. + `subfolder` places the step's files in a subfolder of the run directory - see *Saying which output is the deliverable* above. `file_base_name` is the base name the step's files are written under, replacing the name derived from the workflow and step; it may not contain a path separator. For video, `"fps"` is the rate the file is written at. It is rarely needed: diff --git a/tests/test_assess.py b/tests/test_assess.py new file mode 100644 index 00000000..ad81cdad --- /dev/null +++ b/tests/test_assess.py @@ -0,0 +1,560 @@ +"""Tests for the assessment probes (dw/tasks/assess.py, #387). + +Every probe call here is the real function - no mocking - against synthetic +media built in-process (PIL frames + numpy sine tones) or real mp4 files +written with PyAV, following the fixture style of tests/test_shots.py and +tests/test_media_info.py. +""" + +import json + +import numpy +import pytest +from PIL import Image + +from dw.result import AudioVideo +from dw.runs import MANIFEST_FILE_NAME, shots_beside +from dw.shots import shot_record +from dw.tasks.assess import ( + THUMB_HEIGHT, + THUMB_WIDTH, + analyze_seams, + analyze_shots, + analyze_sync_drift, + read_media, + resolve_shots, +) +from dw.tasks.dissolve_videos import dissolve_videos + + +# --------------------------------------------------------------------------- +# Synthetic media helpers +# --------------------------------------------------------------------------- + + +def make_frames(num_frames, base_grey, noise=2.0, size=(64, 64), seed=0): + """`num_frames` RGB frames around one grey level, with per-pixel noise so + consecutive frames have a nonzero typical delta.""" + rng = numpy.random.default_rng(seed) + width, height = size + frames = [] + for _ in range(num_frames): + arr = base_grey + rng.normal(0, noise, size=(height, width, 3)) + arr = numpy.clip(arr, 0, 255).astype(numpy.uint8) + frames.append(Image.fromarray(arr, mode="RGB")) + return frames + + +def make_tone(num_samples, sample_rate=48000, freq=440.0, amplitude=0.2, channels=2): + t = numpy.arange(num_samples) / sample_rate + tone = (amplitude * numpy.sin(2 * numpy.pi * freq * t)).astype(numpy.float32) + return numpy.tile(tone, (channels, 1)) + + +def write_mp4( + path, frames=48, fps=24, width=64, height=64, sample_rate=48000, seconds=None +): + """A real mp4 with a moving picture and a sine soundtrack, for the tests + that exercise read_media on an actual encoded file. Mirrors + tests/test_media_info.py's write_mp4 helper, with a picture that moves + frame to frame instead of staying flat.""" + import av + + container = av.open(str(path), "w") + video = container.add_stream("libx264", rate=fps) + video.width, video.height, video.pix_fmt = width, height, "yuv420p" + audio = container.add_stream("aac", rate=sample_rate) + audio.layout = "stereo" + + rng = numpy.random.default_rng(1) + for _ in range(frames): + arr = rng.integers(0, 255, size=(height, width, 3), dtype=numpy.uint8) + frame = av.VideoFrame.from_ndarray(arr, format="rgb24") + for packet in video.encode(frame): + container.mux(packet) + + total_samples = ( + seconds and int(seconds * sample_rate) or int(sample_rate * frames / fps) + ) + tone = make_tone(total_samples, sample_rate=sample_rate, amplitude=0.2) + for start in range(0, total_samples, 1024): + chunk = av.AudioFrame.from_ndarray( + numpy.ascontiguousarray(tone[:, start : start + 1024]), + format="fltp", + layout="stereo", + ) + chunk.sample_rate = sample_rate + chunk.pts = start + for packet in audio.encode(chunk): + container.mux(packet) + for packet in audio.encode(): + container.mux(packet) + for packet in video.encode(): + container.mux(packet) + container.close() + + +def cut_shots(names, frames_per_shot, samples_per_shot): + """Shot records for a hard-cut concatenation: contiguous, no overlap.""" + shots = [] + for index, name in enumerate(names): + shots.append( + shot_record( + name, + index * frames_per_shot, + frames_per_shot, + index * samples_per_shot, + samples_per_shot, + ) + ) + return shots + + +# --------------------------------------------------------------------------- +# 1. A 6 dB step flags seam_level_step at one seam only, and shot_level_spread +# --------------------------------------------------------------------------- + + +class TestLevelStep: + def test_step_flags_one_seam_and_shot_spread(self): + fps = 24 + sample_rate = 48000 + frames_per_shot = 48 + samples_per_shot = sample_rate * frames_per_shot // fps + + names = ["s0", "s1", "s2"] + amplitudes = [0.2, 0.2, 0.5] # s2 is ~8 dB louder + + frames = [] + tracks = [] + for index, amp in enumerate(amplitudes): + frames.extend( + make_frames(frames_per_shot, base_grey=120, noise=3.0, seed=index) + ) + tracks.append(make_tone(samples_per_shot, sample_rate, amplitude=amp)) + audio = numpy.concatenate(tracks, axis=1) + + shots = cut_shots(names, frames_per_shot, samples_per_shot) + video = AudioVideo(frames, audio, sample_rate, fps=fps, shots=shots) + + seams = analyze_seams(video) + assert seams["shots_source"] == "artifact" + step_findings = [f for f in seams["findings"] if f["rule"] == "seam_level_step"] + assert len(step_findings) == 1 + assert step_findings[0]["at"]["seam"] == 2 + + seam1_rules = {f["rule"] for f in seams["findings"] if f["at"]["seam"] == 1} + assert "seam_level_step" not in seam1_rules + + shots_answer = analyze_shots(video) + assert shots_answer["has_audio"] is True + spread_findings = [ + f for f in shots_answer["findings"] if f["rule"] == "shot_level_spread" + ] + assert len(spread_findings) == 1 + assert shots_answer["rms_range_db"] >= 6.0 + + # JSON-serializable, per requirement 10 + json.dumps(seams) + json.dumps(shots_answer) + + +# --------------------------------------------------------------------------- +# 2. A dissolve does not flag a level step, click, hole or frame jump +# --------------------------------------------------------------------------- + + +class TestDissolveDoesNotFlag: + def test_dissolve_seams_are_clean(self): + fps = 24 + sample_rate = 48000 + frames_per_clip = 48 + samples_per_clip = sample_rate * frames_per_clip // fps + + clips = [] + for index, base in enumerate((100, 130, 160)): + clip_frames = make_frames( + frames_per_clip, base_grey=base, noise=3.0, seed=10 + index + ) + audio = make_tone(samples_per_clip, sample_rate, amplitude=0.3) + clips.append(AudioVideo(clip_frames, audio, sample_rate, fps=fps)) + + joined = dissolve_videos(clips, dissolve_frames=12, fps=fps) + assert joined.shots is not None + + seams = analyze_seams(joined) + assert seams["shots_source"] == "artifact" + assert len(seams["seams"]) == 2 + for seam in seams["seams"]: + assert seam["kind"] == "dissolve" + + offending = [ + f + for f in seams["findings"] + if f["rule"] in ("seam_level_step", "seam_click", "seam_hole") + ] + assert offending == [] + assert seams["findings"] == [], seams["findings"] + + json.dumps(seams) + + +# --------------------------------------------------------------------------- +# 3. hard_cut suppresses seam_frame_jump; the same media without it fires +# --------------------------------------------------------------------------- + + +class TestHardCut: + def _build(self, hard_cut): + fps = 24 + frames_per_shot = 24 + frames = make_frames(frames_per_shot, base_grey=0.3 * 255, noise=2.0, seed=1) + frames += make_frames(frames_per_shot, base_grey=0.8 * 255, noise=2.0, seed=2) + video = AudioVideo(frames, None, None, fps=fps) + extra = {"hard_cut": True} if hard_cut else {} + shots = [ + shot_record("s0", 0, frames_per_shot), + shot_record("s1", frames_per_shot, frames_per_shot, **extra), + ] + return video, shots + + def test_hard_cut_suppresses_the_finding(self): + video, shots = self._build(hard_cut=True) + answer = analyze_seams(video, shots=shots) + assert answer["shots_source"] == "argument" + assert answer["seams"][0]["hard_cut"] is True + jump_findings = [ + f for f in answer["findings"] if f["rule"] == "seam_frame_jump" + ] + assert jump_findings == [] + + def test_without_hard_cut_the_jump_fires(self): + video, shots = self._build(hard_cut=False) + answer = analyze_seams(video, shots=shots) + assert answer["seams"][0]["hard_cut"] is False + jump_findings = [ + f for f in answer["findings"] if f["rule"] == "seam_frame_jump" + ] + assert len(jump_findings) == 1 + assert jump_findings[0]["severity"] == "info" + assert jump_findings[0]["at"]["seam"] == 1 + + +# --------------------------------------------------------------------------- +# 4. A static shot into a modest cut does not flag seam_frame_jump +# --------------------------------------------------------------------------- + + +class TestStaticShotThenModestCut: + def test_no_jump_for_a_modest_change_after_a_static_shot(self): + fps = 24 + frames_per_shot = 24 + # shot 0: nearly static (tiny noise) + frames = make_frames(frames_per_shot, base_grey=128, noise=0.3, seed=5) + # shot 1: moderate, ordinary frame-to-frame motion + frames += make_frames(frames_per_shot, base_grey=140, noise=5.0, seed=6) + video = AudioVideo(frames, None, None, fps=fps) + shots = [ + shot_record("s0", 0, frames_per_shot), + shot_record("s1", frames_per_shot, frames_per_shot), + ] + + answer = analyze_seams(video, shots=shots) + seam = answer["seams"][0] + assert seam["typical_delta"] is not None + assert seam["jump_ratio"] is not None + assert seam["jump_ratio"] <= 8.0 + jump_findings = [ + f for f in answer["findings"] if f["rule"] == "seam_frame_jump" + ] + assert jump_findings == [] + + +# --------------------------------------------------------------------------- +# 5. Drift accumulates and crosses sync_drift once it passes 40 ms +# --------------------------------------------------------------------------- + + +class TestSyncDrift: + def test_drift_accumulates_and_flags(self): + fps = 24 + sample_rate = 48000 + frames_per_shot = 48 # 2s + nominal_samples = sample_rate * frames_per_shot // fps # 96000 + overrun = 267 + num_shots = 9 + + names = [f"s{i}" for i in range(num_shots)] + start_sample = 0 + shots = [] + expected_offsets = [] + for i, name in enumerate(names): + num_samples = nominal_samples + overrun + shots.append( + shot_record( + name, + i * frames_per_shot, + frames_per_shot, + start_sample, + num_samples, + ) + ) + end_frame = (i + 1) * frames_per_shot + offset_ms = ( + (start_sample + num_samples) / sample_rate - end_frame / fps + ) * 1000.0 + expected_offsets.append(offset_ms) + start_sample += num_samples + + total_frames = num_shots * frames_per_shot + total_samples = start_sample + frames = make_frames(total_frames, base_grey=100, noise=2.0, seed=7) + audio = make_tone(total_samples, sample_rate, amplitude=0.2) + video = AudioVideo(frames, audio, sample_rate, fps=fps) + + answer = analyze_sync_drift(video, shots=shots) + assert answer["shots_source"] == "argument" + measured_offsets = [shot["end_offset_ms"] for shot in answer["shots"]] + + # growing, per shot, and matching the arithmetic above + assert measured_offsets == sorted(measured_offsets) + for measured, expected in zip(measured_offsets, expected_offsets): + assert measured == pytest.approx(expected, abs=0.01) + + assert answer["max_offset_ms"] == pytest.approx(expected_offsets[-1], abs=0.01) + assert answer["max_offset_ms"] > 40.0 + + drift_findings = [f for f in answer["findings"] if f["rule"] == "sync_drift"] + assert drift_findings + assert all(f["severity"] == "warn" for f in drift_findings) + + json.dumps(answer) + + def test_a_clean_cut_reports_no_findings(self): + fps = 24 + sample_rate = 48000 + frames_per_shot = 48 + samples_per_shot = sample_rate * frames_per_shot // fps + names = ["a", "b", "c"] + + shots = cut_shots(names, frames_per_shot, samples_per_shot) + total_frames = len(names) * frames_per_shot + total_samples = len(names) * samples_per_shot + frames = make_frames(total_frames, base_grey=100, noise=2.0, seed=8) + audio = make_tone(total_samples, sample_rate, amplitude=0.2) + video = AudioVideo(frames, audio, sample_rate, fps=fps) + + answer = analyze_sync_drift(video, shots=shots) + assert answer["findings"] == [] + for shot in answer["shots"]: + assert shot["end_offset_ms"] == pytest.approx(0.0, abs=0.01) + + +# --------------------------------------------------------------------------- +# 6. Memory bound: read_media never holds the full decoded frame list +# --------------------------------------------------------------------------- + + +class TestMemoryBound: + def test_read_media_stays_far_below_the_full_frame_list( + self, tmp_path, monkeypatch + ): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + import tracemalloc + + path = tmp_path / "big.mp4" + frames, width, height = 240, 256, 256 + write_mp4(path, frames=frames, fps=24, width=width, height=height) + + tracemalloc.start() + media = read_media(str(path)) + _current, peak = tracemalloc.get_traced_memory() + tracemalloc.stop() + + full_frame_list_bytes = frames * height * width * 3 + assert peak < full_frame_list_bytes / 4 + + assert media.thumbs is not None + assert media.thumbs.shape[1] == THUMB_HEIGHT + assert media.thumbs.shape[2] == THUMB_WIDTH + assert media.thumbs.shape[0] == frames + + +# --------------------------------------------------------------------------- +# 7. read_media + probes on a real encoded file; AAC priming doesn't drift +# --------------------------------------------------------------------------- + + +class TestRealFile: + def test_probes_on_an_encoded_file(self, tmp_path, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + path = tmp_path / "cut.mp4" + write_mp4(path, frames=48, fps=24, width=64, height=64, sample_rate=48000) + + media = read_media(str(path)) + assert media.thumbs is not None + assert media.audio is not None + + shots_answer = analyze_shots(str(path)) + assert shots_answer["shots_source"] == "none" + + seams_answer = analyze_seams(str(path)) + assert seams_answer["shots_source"] == "none" + assert seams_answer["seams"] == [] + + drift_answer = analyze_sync_drift(str(path)) + assert drift_answer["shots_source"] == "none" + assert drift_answer["length_delta_ms"] is not None + assert abs(drift_answer["length_delta_ms"]) < 40.0 + + json.dumps(shots_answer) + json.dumps(seams_answer) + json.dumps(drift_answer) + + +# --------------------------------------------------------------------------- +# 8. shots_beside reads a manifest, and a probe reports shots_source manifest +# --------------------------------------------------------------------------- + + +class TestShotsBeside: + def test_manifest_shots_are_read_back(self, tmp_path, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + run_dir = tmp_path / "demo-workflow" / "20260923-120000-abcdef01" + run_dir.mkdir(parents=True) + media_path = run_dir / "final" / "cut.mp4" + media_path.parent.mkdir(parents=True) + write_mp4(media_path, frames=48, fps=24, width=64, height=64, sample_rate=48000) + + manifest_shots = [ + shot_record("a", 0, 24, 0, 48000), + shot_record("b", 24, 24, 48000, 48000), + ] + own = "final/cut.mp4" + manifest = { + "steps": [{"step": "join", "files": [own], "shots": manifest_shots}] + } + with open(run_dir / MANIFEST_FILE_NAME, "w") as handle: + json.dump(manifest, handle) + + read_back = shots_beside(str(media_path)) + assert read_back == manifest_shots + + answer = analyze_seams(str(media_path)) + assert answer["shots_source"] == "manifest" + assert len(answer["seams"]) == 1 + + +# --------------------------------------------------------------------------- +# 9. resolve_shots order: argument > artifact > manifest > none +# --------------------------------------------------------------------------- + + +class TestResolveShotsOrder: + def _media(self, video): + from dw.tasks.assess import media_from + + return media_from(video) + + def test_argument_wins_over_artifact(self): + frames = make_frames(8, base_grey=100, noise=1.0, seed=20) + artifact_shots = [shot_record("artifact", 0, 8)] + video = AudioVideo(frames, None, None, fps=4, shots=artifact_shots) + media = self._media(video) + argument_shots = [shot_record("argument", 0, 8)] + + records, source = resolve_shots(video, media, shots=argument_shots) + assert source == "argument" + assert records[0]["name"] == "argument" + + def test_artifact_wins_over_manifest(self, tmp_path, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + run_dir = tmp_path / "wf" / "20260923-120000-abcdef02" + run_dir.mkdir(parents=True) + media_path = run_dir / "cut.mp4" + write_mp4(media_path, frames=8, fps=4, width=32, height=32, sample_rate=8000) + manifest = { + "steps": [ + { + "step": "join", + "files": ["cut.mp4"], + "shots": [shot_record("manifest", 0, 8)], + } + ] + } + with open(run_dir / MANIFEST_FILE_NAME, "w") as handle: + json.dump(manifest, handle) + + from dw.tasks.assess import media_from + + media = media_from(str(media_path)) + media.shots = [shot_record("artifact", 0, 8)] + + records, source = resolve_shots(str(media_path), media, shots=None) + assert source == "artifact" + assert records[0]["name"] == "artifact" + + def test_manifest_wins_over_none(self, tmp_path, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + run_dir = tmp_path / "wf" / "20260923-120000-abcdef03" + run_dir.mkdir(parents=True) + media_path = run_dir / "cut.mp4" + write_mp4(media_path, frames=8, fps=4, width=32, height=32, sample_rate=8000) + manifest = { + "steps": [ + { + "step": "join", + "files": ["cut.mp4"], + "shots": [shot_record("manifest", 0, 8)], + } + ] + } + with open(run_dir / MANIFEST_FILE_NAME, "w") as handle: + json.dump(manifest, handle) + + from dw.tasks.assess import media_from + + media = media_from(str(media_path)) + records, source = resolve_shots(str(media_path), media, shots=None) + assert source == "manifest" + assert records[0]["name"] == "manifest" + + def test_none_when_nothing_resolves(self, tmp_path, monkeypatch): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + media_path = tmp_path / "solo.mp4" + write_mp4(media_path, frames=8, fps=4, width=32, height=32, sample_rate=8000) + + from dw.tasks.assess import media_from + + media = media_from(str(media_path)) + records, source = resolve_shots(str(media_path), media, shots=None) + assert records is None + assert source == "none" + + +# --------------------------------------------------------------------------- +# 10. Answers are JSON-serializable and carry the common fields +# --------------------------------------------------------------------------- + + +class TestAnswerShape: + def test_answers_carry_the_common_fields(self): + fps = 24 + sample_rate = 48000 + frames_per_shot = 24 + samples_per_shot = sample_rate * frames_per_shot // fps + shots = cut_shots(["a", "b"], frames_per_shot, samples_per_shot) + frames = make_frames(frames_per_shot * 2, base_grey=100, noise=2.0, seed=30) + audio = make_tone(samples_per_shot * 2, sample_rate, amplitude=0.2) + video = AudioVideo(frames, audio, sample_rate, fps=fps, shots=shots) + + for answer in ( + analyze_shots(video), + analyze_seams(video), + analyze_sync_drift(video), + ): + json.dumps(answer) + assert "findings" in answer + assert "rules_applied" in answer + assert "shots_source" in answer + assert answer["shots_source"] == "artifact" diff --git a/tests/test_assessment_rules.py b/tests/test_assessment_rules.py new file mode 100644 index 00000000..a80f3fae --- /dev/null +++ b/tests/test_assessment_rules.py @@ -0,0 +1,190 @@ +"""Tests for #387: the assessment rules table and the probes it is read +against. + +`dw/assessment_rules.py` names, for each rule, the probe that reports the +field, the field itself (a key of that probe's per-shot/per-seam record or +of its answer), how the value is compared and the threshold. A rule reading +a field a probe does not actually emit, or naming a probe that is not a +real `returns="json"` task command, would silently find nothing - this pins +every rule to the real probes in `dw/tasks/assess.py`. +""" + +import numpy +import pytest +from PIL import Image + +from dw.assessment_rules import ( + COMPARATORS, + RULES, + SEVERITIES, + crosses, + finding, + rules_for, +) +from dw.introspection import list_tasks +from dw.result import AudioVideo +from dw.shots import shot_record +from dw.tasks.assess import analyze_seams, analyze_shots, analyze_sync_drift +from dw.tasks.audio_utils import LEVEL_SPREAD_WARN_DB +from dw.tasks.task import task_command_info + +FPS = 24 +SAMPLE_RATE = 8000 +FRAMES_PER_SHOT = 24 +SAMPLES_PER_SHOT = SAMPLE_RATE # 1 second, matching FRAMES_PER_SHOT at FPS + +# Distinct brightness per shot so a seam shows a real frame jump; constant +# within a shot so its own typical delta stays at the floor +_SHOT_LEVELS = (40, 210, 40) +# Distinct tone amplitude per shot so the shots differ in level +_SHOT_AMPLITUDES = (0.02, 0.5, 0.02) + + +def _frame(level): + array = numpy.full((16, 16), level, dtype=numpy.uint8) + return Image.fromarray(array, mode="L") + + +def _tone(amplitude, samples, sample_rate=SAMPLE_RATE, frequency=440.0): + t = numpy.arange(samples, dtype=numpy.float32) / sample_rate + return (amplitude * numpy.sin(2 * numpy.pi * frequency * t)).astype(numpy.float32) + + +def _synthetic_video(): + """A 3-shot AudioVideo small enough to run through the real probes.""" + frames = [] + audio_chunks = [] + shots = [] + for index, (level, amplitude) in enumerate(zip(_SHOT_LEVELS, _SHOT_AMPLITUDES)): + frames.extend(_frame(level) for _ in range(FRAMES_PER_SHOT)) + audio_chunks.append(_tone(amplitude, SAMPLES_PER_SHOT)) + shots.append( + shot_record( + name=f"shot{index}", + start_frame=index * FRAMES_PER_SHOT, + num_frames=FRAMES_PER_SHOT, + start_sample=index * SAMPLES_PER_SHOT, + num_samples=SAMPLES_PER_SHOT, + ) + ) + audio = numpy.concatenate(audio_chunks) + return AudioVideo(frames, audio, SAMPLE_RATE, fps=FPS, shots=shots) + + +PROBES = { + "analyze_shots": analyze_shots, + "analyze_seams": analyze_seams, + "analyze_sync_drift": analyze_sync_drift, +} + + +def _fields(answer): + """Every key a probe's answer or its per-shot/per-seam records carry.""" + fields = set(answer.keys()) + for records_key in ("shots", "seams"): + for record in answer.get(records_key) or []: + fields.update(record.keys()) + return fields + + +@pytest.fixture(scope="module") +def probe_answers(): + video = _synthetic_video() + return {name: probe(video) for name, probe in PROBES.items()} + + +class TestRuleProbesAreRealCommands: + def test_every_rule_names_a_registered_json_command(self): + assessment = list_tasks()["assessment"] + for rule in RULES: + assert rule["probe"] in assessment + info = task_command_info(rule["probe"]) + assert info["returns"] == "json" + + def test_assessment_is_exactly_the_three_probes(self): + assert list_tasks()["assessment"] == sorted( + ["analyze_shots", "analyze_seams", "analyze_sync_drift"] + ) + + def test_probes_stay_in_commands_too(self): + commands = list_tasks()["commands"] + for name in list_tasks()["assessment"]: + assert name in commands + + +class TestRuleFieldsAreReal: + def test_every_rule_field_is_a_real_key_the_probe_emits(self, probe_answers): + for rule in RULES: + answer = probe_answers[rule["probe"]] + assert rule["field"] in _fields(answer), ( + f"{rule['name']}: {rule['field']!r} is not a key {rule['probe']} emits" + ) + + def test_every_rule_name_appears_in_its_probe_answer(self, probe_answers): + for rule in RULES: + answer = probe_answers[rule["probe"]] + assert rule["name"] in answer["rules_applied"] + + def test_rules_applied_matches_rules_for(self, probe_answers): + for probe_name, answer in probe_answers.items(): + assert answer["rules_applied"] == [r["name"] for r in rules_for(probe_name)] + + +class TestShotLevelSpreadThreshold: + def test_shares_the_run_time_warning_threshold(self): + rule = next(r for r in RULES if r["name"] == "shot_level_spread") + assert rule["threshold"] is LEVEL_SPREAD_WARN_DB + + +class TestTableShape: + def test_severities_are_declared(self): + for rule in RULES: + assert rule["severity"] in SEVERITIES + + def test_comparators_are_declared(self): + for rule in RULES: + assert rule["comparator"] in COMPARATORS + + def test_rule_names_are_unique(self): + names = [rule["name"] for rule in RULES] + assert len(names) == len(set(names)) + + +class TestCrosses: + def test_none_never_crosses(self): + for rule in RULES: + assert crosses(rule, None) is False + + def test_magnitude_uses_abs(self): + rule = {"comparator": ">", "threshold": 10.0, "magnitude": True} + assert crosses(rule, -20.0) is True + assert crosses(rule, -5.0) is False + + def test_greater_than(self): + rule = {"comparator": ">", "threshold": 10.0} + assert crosses(rule, 10.0) is False + assert crosses(rule, 10.1) is True + + def test_greater_than_or_equal(self): + rule = {"comparator": ">=", "threshold": 10.0} + assert crosses(rule, 10.0) is True + assert crosses(rule, 9.9) is False + + def test_less_than(self): + rule = {"comparator": "<", "threshold": 10.0} + assert crosses(rule, 9.9) is True + assert crosses(rule, 10.0) is False + + +class TestFinding: + def test_shape(self): + rule = RULES[0] + result = finding(rule, 12.5, {"seam": 1}) + assert result == { + "rule": rule["name"], + "severity": rule["severity"], + "at": {"seam": 1}, + "value": 12.5, + "threshold": rule["threshold"], + "says": rule["says"], + } diff --git a/tests/test_scalar_result_validation.py b/tests/test_scalar_result_validation.py index 8d87325f..d6766c78 100644 --- a/tests/test_scalar_result_validation.py +++ b/tests/test_scalar_result_validation.py @@ -5,11 +5,17 @@ clean and then died at run time inside `save_artifact` with `write() argument must be str, not float` after the fan-out ahead of it had already generated (#212). These are the free pre-flight versions. + +The same check covers the `returns="json"` assessment probes (#387): a +"result" on one of those may only be `application/json`, since any other +content type would explode the answer dict key by key into files or die on +a number. """ import unittest from dw.scalar_result_validation import scalar_result_errors +from dw.workflow import Workflow def _step(name, command, result=None, arguments=None): @@ -99,5 +105,161 @@ def test_non_dict_step_is_skipped(self): self.assertEqual(errors, []) +class TestJsonReturningCommands(unittest.TestCase): + """A `returns="json"` command (the assessment probes, #387) may only be + saved as `application/json` - any other content type would explode the + answer dict key by key into files, or die on a number.""" + + def test_application_json_is_fine(self): + definition = { + "steps": [ + _step( + "seams", + "analyze_seams", + result={"content_type": "application/json"}, + ) + ] + } + + errors = scalar_result_errors(definition, source_indices=[0]) + + self.assertEqual(errors, []) + + def test_text_plain_is_an_error_naming_application_json(self): + definition = { + "steps": [ + _step( + "seams", + "analyze_seams", + result={"content_type": "text/plain"}, + ) + ] + } + + errors = scalar_result_errors(definition, source_indices=[0]) + + self.assertEqual(len(errors), 1) + self.assertEqual(errors[0]["path"], "steps[0].result") + self.assertIn("application/json", errors[0]["message"]) + self.assertIn("analyze_seams", errors[0]["message"]) + + def test_image_png_is_an_error_naming_application_json(self): + definition = { + "steps": [ + _step( + "seams", + "analyze_seams", + result={"content_type": "image/png"}, + ) + ] + } + + errors = scalar_result_errors(definition, source_indices=[0]) + + self.assertEqual(len(errors), 1) + self.assertEqual(errors[0]["path"], "steps[0].result") + self.assertIn("application/json", errors[0]["message"]) + + def test_scalar_kind_behaviour_is_unchanged(self): + # judge still refuses a result block outright, regardless of + # content_type - the json check is additional, not a replacement + definition = { + "steps": [ + _step( + "score", + "judge", + result={"content_type": "application/json"}, + ) + ] + } + + errors = scalar_result_errors(definition, source_indices=[0]) + + self.assertEqual(len(errors), 1) + self.assertIn("not an artifact", errors[0]["message"]) + + def test_artifact_kind_behaviour_is_unchanged(self): + definition = { + "steps": [ + _step( + "shot", + "gather_images", + result={"content_type": "application/json"}, + ) + ] + } + + errors = scalar_result_errors(definition, source_indices=[0]) + + self.assertEqual(errors, []) + + +class TestThroughValidationErrors(unittest.TestCase): + """The same check, exercised through the public `validation_errors` on a + full workflow definition rather than the pre-substituted form directly.""" + + def _step(self, content_type): + return { + "name": "seams", + "task": { + "command": "analyze_seams", + "arguments": {"video": "asset:cut.mp4"}, + }, + "result": {"content_type": content_type}, + } + + def _workflow(self, content_type): + return Workflow( + {"id": "assess", "steps": [self._step(content_type)]}, + "outputs", + None, + ) + + def test_application_json_validates_clean(self): + self.assertEqual(self._workflow("application/json").validation_errors(), []) + + def test_text_plain_is_refused(self): + errors = self._workflow("text/plain").validation_errors() + + messages = [e["message"] for e in errors if e["path"] == "steps[0].result"] + self.assertTrue(messages) + self.assertTrue(any("application/json" in m for m in messages)) + + def test_image_png_is_refused(self): + errors = self._workflow("image/png").validation_errors() + + messages = [e["message"] for e in errors if e["path"] == "steps[0].result"] + self.assertTrue(messages) + self.assertTrue(any("application/json" in m for m in messages)) + + def test_scalar_command_through_validation_errors_is_unchanged(self): + workflow = Workflow( + { + "id": "assess", + "steps": [ + { + "name": "score", + "task": { + "command": "judge", + "arguments": { + "image": "asset:cut.png", + "prompt": "a cat", + }, + }, + "result": {"content_type": "text/plain"}, + } + ], + }, + "outputs", + None, + ) + + errors = workflow.validation_errors() + + messages = [e["message"] for e in errors if e["path"] == "steps[0].result"] + self.assertTrue(messages) + self.assertTrue(any("not an artifact" in m for m in messages)) + + if __name__ == "__main__": unittest.main() From d9d1143a3fc0dd2261b7ebb64d82e918800d1c01 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 23:09:43 -0500 Subject: [PATCH 081/181] fix(engine): #392 - normalize_audio/grade/loop_audio emit structured logs of what they applied A save:false step's effect was unobservable over MCP: normalize_audio, grade and loop_audio only used logger.debug, which reaches the server log but not get_job_events. All three now emit_log a summary - normalize_audio reports measured peak/LUFS, the gain applied, and which constraint (target_lufs, peak_ceiling, or peak_dbfs when no target_lufs was given) set it; loop_audio reports the lap count and resulting length; grade reports the non-identity parameters actually applied. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/audio_utils.py | 32 ++++++++++++--- dw/tasks/task.py | 20 ++++++++++ tests/test_audio_utils.py | 84 +++++++++++++++++++++++++++++++++++++++ tests/test_grade.py | 36 +++++++++++++++++ 4 files changed, 166 insertions(+), 6 deletions(-) diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index 06a1edb3..8ef69cb7 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -1094,15 +1094,21 @@ def loop_audio( window = min(int(crossfade_ms / 1000.0 * sample_rate), waveform.shape[1] // 2) bed = waveform + laps = 1 # Each lap after the first overlaps the one before it by the crossfade, so # a lap adds (source - window) samples rather than a whole source while bed.shape[1] < length: bed = crossfade_concat( [bed, waveform], sample_rate, window / sample_rate * 1000.0 ) - logger.debug( - f"loop_audio: {waveform.shape[1]} samples at {sample_rate}Hz looped to " - f"{length} ({bed.shape[1]} before trimming)" + laps += 1 + emit_log( + f"loop_audio: {waveform.shape[1]} samples at {sample_rate}Hz looped " + f"{laps}x to {length} samples ({length / sample_rate:.2f} s)", + command="loop_audio", + laps=laps, + output_samples=length, + output_seconds=round(length / sample_rate, 2), ) return _as_track(bed[:, :length], sample_rate, "loop_audio") @@ -1346,12 +1352,15 @@ def normalize_audio(audio, peak_dbfs=-1.0, target_lufs=None, sample_rate=None): logger.warning("normalize_audio: the track is silent - left unchanged") return _as_track(waveform, sample_rate, "normalize_audio") + peak_db = 20 * numpy.log10(peak) + measured_lufs = None if target_lufs is None: - gain_db = peak_dbfs - 20 * numpy.log10(peak) + constraint = "peak_dbfs" + gain_db = peak_dbfs - peak_db else: - peak_db = 20 * numpy.log10(peak) ceiling_gain_db = peak_dbfs - peak_db current_lufs = integrated_lufs(waveform.T, sample_rate) + measured_lufs = current_lufs if current_lufs is None: emit_warning( f"normalize_audio: target_lufs={target_lufs} was given, but the " @@ -1363,9 +1372,11 @@ def normalize_audio(audio, peak_dbfs=-1.0, target_lufs=None, sample_rate=None): target_lufs=target_lufs, ) gain_db = ceiling_gain_db + constraint = "peak_ceiling" else: target_gain_db = target_lufs - current_lufs gain_db = min(target_gain_db, ceiling_gain_db) + constraint = "peak_ceiling" if gain_db < target_gain_db else "target_lufs" if gain_db < target_gain_db: emit_warning( f"normalize_audio: target_lufs={target_lufs} would need " @@ -1379,7 +1390,16 @@ def normalize_audio(audio, peak_dbfs=-1.0, target_lufs=None, sample_rate=None): shortfall_lu=target_gain_db - gain_db, ) gain = 10 ** (gain_db / 20) - logger.debug(f"normalize_audio: peak {peak:.3f}, gain {gain_db:+.1f} dB") + emit_log( + f"normalize_audio: measured {peak_db:.1f} dBFS peak" + + ("" if measured_lufs is None else f", {measured_lufs:.1f} LUFS") + + f" -> gain {gain_db:+.1f} dB, set by {constraint}", + command="normalize_audio", + measured_peak_dbfs=round(peak_db, 1), + measured_lufs=round(measured_lufs, 1) if measured_lufs is not None else None, + gain_db=round(gain_db, 1), + constraint=constraint, + ) return _as_track( (waveform * gain).astype(numpy.float32), sample_rate, "normalize_audio" ) diff --git a/dw/tasks/task.py b/dw/tasks/task.py index 5f240f52..674e5ec0 100644 --- a/dw/tasks/task.py +++ b/dw/tasks/task.py @@ -2,6 +2,7 @@ from typing import Callable, Dict from .. import resolve_device +from ..events import emit_log from .qr_code import get_qrcode_image from .image_utils import process_image from .video_utils import process_video @@ -522,6 +523,25 @@ def _handle_grade(task, arguments, previous_pipelines): media = fetch_image(media) + defaults = { + "exposure": 0.0, + "contrast": 1.0, + "saturation": 1.0, + "temperature": 0.0, + "tint": 0.0, + } + applied = { + name: arguments.get(name, default) + for name, default in defaults.items() + if arguments.get(name, default) != default + } + emit_log( + f"grade: applied {applied}" + if applied + else "grade: no adjustment (all identity)", + command="grade", + **applied, + ) return _per_frame(media, lambda frame: grade_image(frame, **arguments)) diff --git a/tests/test_audio_utils.py b/tests/test_audio_utils.py index 4837dde6..7da2e452 100644 --- a/tests/test_audio_utils.py +++ b/tests/test_audio_utils.py @@ -702,6 +702,70 @@ def test_default_behavior_is_unchanged_without_target_lufs(self): assert numpy.array_equal(with_default, explicit_none) + def test_logs_peak_only_constraint_without_target_lufs(self): + # #392: without target_lufs the peak ceiling is the only constraint, + # and a caller reading job events should see that named explicitly + from dw.events import RunContext, activate_context, deactivate_context + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone() + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + normalize_audio(track, peak_dbfs=-3.0, sample_rate=rate) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert logs[0]["constraint"] == "peak_dbfs" + assert logs[0]["measured_lufs"] is None + assert logs[0]["gain_db"] is not None + + def test_logs_measured_lufs_and_target_lufs_constraint(self): + # #392: the caller needs to know the gain was set by the LUFS target, + # not just that a gain was applied + from dw.events import RunContext, activate_context, deactivate_context + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone(density=0.1) + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + normalize_audio(track, peak_dbfs=0.0, target_lufs=-16.0, sample_rate=rate) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert logs[0]["constraint"] == "target_lufs" + # measured_lufs is the *input's* loudness before the gain was applied, + # not the target - the scaled track's loudness is what test_target_lufs_* + # already checks lands on target + assert logs[0]["measured_lufs"] is not None + assert logs[0]["gain_db"] == pytest.approx(-16.0 - logs[0]["measured_lufs"]) + + def test_logs_peak_ceiling_constraint_when_it_caps_the_target(self): + # #392: the ceiling-capped case (already warned via target_lufs_capped) + # should also name peak_ceiling as the constraint in the summary log + from dw.events import RunContext, activate_context, deactivate_context + from dw.tasks.audio_utils import normalize_audio + + track, rate = self._tone(amplitude=0.5, density=1.0) + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + normalize_audio(track, peak_dbfs=-1.0, target_lufs=-1.0, sample_rate=rate) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert logs[0]["constraint"] == "peak_ceiling" + class TestAudioTasksTakeAnAudioVideo: """Every audio task accepts the video an earlier step generated with its @@ -860,6 +924,26 @@ def test_a_length_in_frames_matches_a_cut_exactly(self): assert samples(bed).shape == (200, 1) + def test_logs_the_loop_count_and_output_length(self): + # #392: with save:false a caller can only see what loop_audio did + # through job events, so the lap count and resulting length must + # reach the log rather than only logger.debug + from dw.events import RunContext, activate_context, deactivate_context + from dw.tasks.audio_utils import loop_audio + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + loop_audio(self.tone(), duration_seconds=3.5, sample_rate=100) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert logs[0]["laps"] > 1 + assert logs[0]["output_samples"] == 350 + assert logs[0]["output_seconds"] == pytest.approx(3.5) + def test_a_source_longer_than_the_bed_is_trimmed(self): from dw.tasks.audio_utils import loop_audio diff --git a/tests/test_grade.py b/tests/test_grade.py index 33d51b3c..7785a43a 100644 --- a/tests/test_grade.py +++ b/tests/test_grade.py @@ -161,6 +161,42 @@ def test_a_video_file_path_is_loaded_with_its_audio(self, tmp_path): assert result.fps == 24 +class TestAppliedParametersAreLogged: + def test_logs_the_non_identity_parameters_applied(self): + # #392: with save:false a caller can only see what grade did through + # job events, so the parameters actually applied must reach the log + from dw.events import RunContext, activate_context, deactivate_context + + task = Task({"command": "grade", "arguments": {}}, "cpu") + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + task.run({"media": _swatch(), "exposure": 1.0, "saturation": 0.5}) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert logs[0]["exposure"] == 1.0 + assert logs[0]["saturation"] == 0.5 + assert "contrast" not in logs[0] + + def test_logs_no_adjustment_when_every_parameter_is_identity(self): + from dw.events import RunContext, activate_context, deactivate_context + + task = Task({"command": "grade", "arguments": {}}, "cpu") + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + task.run({"media": _swatch()}) + finally: + deactivate_context(token) + + logs = [e for e in events if e.get("event") == "log"] + assert len(logs) == 1 + assert "no adjustment" in logs[0]["message"] + + class TestDomains: def _errors(self, arguments): definition = { From 9589caedda0fad59a6503b160b3836b742221c43 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 23:17:23 -0500 Subject: [PATCH 082/181] feat(engine): #387 - seam_level_step compares shot levels, not the seam's edge windows A take's own tail and head can sit 20 dB apart, which flagged every seam of a cut made of one clip and added that natural difference to a real duck. analyze_shots now echoes each shot's start_frame/num_frames. Co-Authored-By: Claude Opus 5.5 --- docs/TASKS.md | 6 ++++-- dw/assessment_rules.py | 2 +- dw/tasks/assess.py | 47 +++++++++++++++++++++++++++++++++--------- tests/test_assess.py | 40 +++++++++++++++++++++++++++++++++++ 4 files changed, 82 insertions(+), 13 deletions(-) diff --git a/docs/TASKS.md b/docs/TASKS.md index 2ce3b2e4..d62435fb 100644 --- a/docs/TASKS.md +++ b/docs/TASKS.md @@ -902,6 +902,7 @@ Each shot's level and spectral balance, and how far apart the shots sit: | Field | Meaning | | ----- | ------- | | `shots[].name` | The shot's name | +| `shots[].start_frame` / `num_frames` | The shot's frame range, as the shot record gave it | | `shots[].peak_dbfs` | Peak level within the shot | | `shots[].rms_dbfs` | RMS level within the shot | | `shots[].crest_db` | `peak_dbfs` minus `rms_dbfs` | @@ -921,8 +922,9 @@ Every seam between shots, audio and picture: | `seams[].seconds` | Where the seam sits in the file | | `seams[].kind` | `cut` or `dissolve` (a dissolve has `overlap_frames`) | | `seams[].hard_cut` | Whether the incoming shot is marked `hard_cut: true` | -| `seams[].before_rms_dbfs` / `after_rms_dbfs` | RMS level either side of the seam | -| `seams[].level_step_db` | The absolute level jump across the seam | +| `seams[].before_shot_rms_dbfs` / `after_shot_rms_dbfs` | RMS level of the whole shot either side of the seam | +| `seams[].level_step_db` | The absolute difference between those two shot levels. Shot against shot, not the audio at the seam's edges: a take's own tail and head can sit 20 dB apart, which is not a step the cut made | +| `seams[].before_rms_dbfs` / `after_rms_dbfs` | RMS level of the 0.25 s either side of the seam - what `seam_hole`'s both-sides-voiced guard reads | | `seams[].floor_dbfs` | RMS level of the join itself (the fade, or a short window centred on a cut) | | `seams[].click_db` | How far a spike at the join peaks above its immediate neighbours | | `seams[].spectral_shift` | How much the low/mid/high balance shifts across the seam (0-1) | diff --git a/dw/assessment_rules.py b/dw/assessment_rules.py index 0911f7d4..121f7f84 100644 --- a/dw/assessment_rules.py +++ b/dw/assessment_rules.py @@ -57,7 +57,7 @@ "comparator": ">", "threshold": 3.0, "severity": "warn", - "says": "the level steps this much across the seam", + "says": "the shots either side of the seam sit this far apart in level", }, { "name": "seam_click", diff --git a/dw/tasks/assess.py b/dw/tasks/assess.py index e867fab0..769cc2fd 100644 --- a/dw/tasks/assess.py +++ b/dw/tasks/assess.py @@ -35,7 +35,12 @@ THUMB_WIDTH = 64 THUMB_HEIGHT = 36 -# Audio windows around a seam, in seconds +# Audio windows around a seam, in seconds. The edge windows say whether the +# join itself is voiced (the hole guard) and what the balance does across it; +# they are not the level step, which is shot against shot (`_shot_rms`): a +# shot's own last and first quarter-second differ by whatever the take does +# there - 20 dB on a line that trails off and opens on a breath - and read +# as a step at every seam of a cut made of one clip (#387's bounce). LEVEL_WINDOW = 0.25 FLOOR_WINDOW = 0.02 CLICK_WINDOW = 0.002 @@ -331,7 +336,7 @@ def analyze_shots(video, shots=None): shots: Shot records to measure by, overriding any the video carries Returns: - {shots: [{name, peak_dbfs, rms_dbfs, crest_db, low_dbfs, mid_dbfs, + {shots: [{name, start_frame, num_frames, peak_dbfs, rms_dbfs, crest_db, low_dbfs, mid_dbfs, high_dbfs, samples}], rms_range_db, findings, rules_applied, shots_source} """ @@ -362,6 +367,8 @@ def analyze_shots(video, shots=None): measured.append( { "name": shot.get("name"), + "start_frame": shot.get("start_frame"), + "num_frames": shot.get("num_frames"), "peak_dbfs": _round(peak), "rms_dbfs": _round(rms), "crest_db": _round(None if peak is None or rms is None else peak - rms), @@ -405,11 +412,20 @@ def _band_shares(window, sample_rate): return {key: value / total for key, value in energies.items()} -def _seam_audio(media, before_end, after_start): - """Audio measurements at a seam: the level windows end at `before_end` +def _shot_rms(media, shot): + """A shot's RMS level over its whole sample span, in dBFS, or None.""" + start, count, _source = _sample_span(shot, media) + if start is None: + return None + return _db(_rms(_clip(media, start, start + count))) + + +def _seam_audio(media, before_end, after_start, previous_rms, next_rms): + """Audio measurements at a seam: the edge windows end at `before_end` and open at `after_start` (the same sample at a cut, either side of the fade at a dissolve), and the join is what lies between them - or the - FLOOR_WINDOW centred on the cut.""" + FLOOR_WINDOW centred on the cut. The level step is between the two + shots' own levels, `previous_rms` and `next_rms`.""" rate = media.sample_rate level = int(round(LEVEL_WINDOW * rate)) before = _clip(media, before_end - level, before_end) @@ -453,10 +469,12 @@ def _seam_audio(media, before_end, after_start): return { "before_rms_dbfs": _round(before_rms), "after_rms_dbfs": _round(after_rms), + "before_shot_rms_dbfs": _round(previous_rms), + "after_shot_rms_dbfs": _round(next_rms), "level_step_db": _round( None - if before_rms is None or after_rms is None - else abs(after_rms - before_rms) + if previous_rms is None or next_rms is None + else abs(next_rms - previous_rms) ), "floor_dbfs": _round(floor), "click_db": _round(click), @@ -517,8 +535,9 @@ def analyze_seams(video, shots=None): A shot marked `hard_cut: true` opens a seam meant as a cut Returns: - {seams: [{seam, between, seconds, kind, level_step_db, floor_dbfs, - click_db, spectral_shift, before_rms_dbfs, after_rms_dbfs, + {seams: [{seam, between, seconds, kind, level_step_db, + before_shot_rms_dbfs, after_shot_rms_dbfs, floor_dbfs, click_db, + spectral_shift, before_rms_dbfs, after_rms_dbfs, frame_delta, typical_delta, jump_ratio}], findings, rules_applied, shots_source} """ @@ -551,7 +570,15 @@ def analyze_seams(video, shots=None): if fade and media.fps else 0 ) - record.update(_seam_audio(media, start, start + fade_samples)) + record.update( + _seam_audio( + media, + start, + start + fade_samples, + _shot_rms(media, previous), + _shot_rms(media, shot), + ) + ) if ( record["before_rms_dbfs"] is None or record["after_rms_dbfs"] is None diff --git a/tests/test_assess.py b/tests/test_assess.py index ad81cdad..b4c53930 100644 --- a/tests/test_assess.py +++ b/tests/test_assess.py @@ -158,6 +158,46 @@ def test_step_flags_one_seam_and_shot_spread(self): json.dumps(seams) json.dumps(shots_answer) + @staticmethod + def _one_clip_three_times(duck_db=0.0): + """One take cut after itself three times, as C-F099 builds it: the + take trails off (-46 dBFS tail) and opens near silence (-66 dBFS + head) around a voiced body, so its own edges sit 20 dB apart. The + third shot is optionally ducked by `duck_db`.""" + fps, sample_rate, frames_per_shot = 24, 48000, 48 + samples_per_shot = sample_rate * frames_per_shot // fps + edge = sample_rate // 4 + take = make_tone(samples_per_shot, sample_rate, amplitude=0.2) + take[:, :edge] *= 10 ** (-66 / 20) / (0.2 / numpy.sqrt(2)) + take[:, -edge:] *= 10 ** (-46 / 20) / (0.2 / numpy.sqrt(2)) + ducked = take * 10 ** (-duck_db / 20) + audio = numpy.concatenate([take, take, ducked], axis=1) + frames = [] + for index in range(3): + frames.extend(make_frames(frames_per_shot, 120, noise=3.0, seed=index)) + shots = cut_shots(["a", "b", "c"], frames_per_shot, samples_per_shot) + return AudioVideo(frames, audio, sample_rate, fps=fps, shots=shots) + + def test_a_take_with_quiet_edges_cut_after_itself_is_clean(self): + seams = analyze_seams(self._one_clip_three_times()) + assert all(seam["before_rms_dbfs"] < -40 for seam in seams["seams"]) + assert all(seam["level_step_db"] == 0.0 for seam in seams["seams"]) + assert not [f for f in seams["findings"] if f["rule"] == "seam_level_step"] + + def test_a_duck_reports_its_own_size_at_its_seam_only(self): + seams = analyze_seams(self._one_clip_three_times(duck_db=12.0)) + steps = [f for f in seams["findings"] if f["rule"] == "seam_level_step"] + assert [f["at"]["seam"] for f in steps] == [2] + assert steps[0]["value"] == pytest.approx(12.0, abs=0.1) + + def test_shots_echo_their_frame_range(self): + answer = analyze_shots(self._one_clip_three_times()) + assert [(s["start_frame"], s["num_frames"]) for s in answer["shots"]] == [ + (0, 48), + (48, 48), + (96, 48), + ] + # --------------------------------------------------------------------------- # 2. A dissolve does not flag a level step, click, hole or frame jump From eda82ba5dfa65789a4d84d9a7e642e0da90716e1 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 23:29:53 -0500 Subject: [PATCH 083/181] feat(engine): #387 - probes read an asset:/output: video from the file, not a decoded frame list A probe's 'video' went through the eager fetch_video path, which decodes to a FrameList with no soundtrack, and every probe then refused it. The probes join get_frame's by-reference realization (VideoFileReference), so a stored file is streamed by read_media and its manifest shots are found beside it. Co-Authored-By: Claude Opus 5.5 --- docs/TASKS.md | 4 +- dw/arguments.py | 23 +++++++--- dw/tasks/assess.py | 32 +++++++++---- dw/tasks/video_utils.py | 5 ++- tests/test_assess.py | 99 +++++++++++++++++++++++++++++++++++++++++ 5 files changed, 146 insertions(+), 17 deletions(-) diff --git a/docs/TASKS.md b/docs/TASKS.md index d62435fb..5dd65180 100644 --- a/docs/TASKS.md +++ b/docs/TASKS.md @@ -861,7 +861,9 @@ sample rate, reads as `null` rather than `-inf`. Three read-only commands measure a finished cut and say where to look - `analyze_shots`, `analyze_seams`, `analyze_sync_drift`. Each takes a video -(a path, or the video an earlier step returned) and answers one JSON +(a stored file - `asset:`, `output:` or a path, read straight from disk +rather than decoded first - or the video an earlier step returned; not a +URL, whose download is a bare frame list with no soundtrack) and answers one JSON document: every measurement it took, plus `findings` (the measurements that crossed a rule in the table below), `rules_applied` (the rule names the probe checked) and `shots_source` (where the shot list came from). A probe reads diff --git a/dw/arguments.py b/dw/arguments.py index 70867641..21a07818 100644 --- a/dw/arguments.py +++ b/dw/arguments.py @@ -1126,17 +1126,30 @@ def fetch_video(video_spec, base_dir=None): raise -# get_frame and its two fixed-index siblings - see _realize_lazy_frame_arguments -_LAZY_FRAME_COMMANDS = frozenset({"get_frame", "get_first_frame", "get_last_frame"}) +# get_frame and its two fixed-index siblings, and the assessment probes +# (dw/tasks/assess.py), which stream the file themselves - decoding it to a +# frame list first dropped the soundtrack they measure and failed every probe +# on an asset:/output: video (#387) - see _realize_lazy_frame_arguments +_LAZY_FRAME_COMMANDS = frozenset( + { + "get_frame", + "get_first_frame", + "get_last_frame", + "analyze_shots", + "analyze_seams", + "analyze_sync_drift", + } +) def _realize_lazy_frame_arguments(arguments, base_dir): - """Realize a get_frame/get_first_frame/get_last_frame step's arguments, - reading a file-based 'video' by reference rather than decoding it (#367). + """Realize a get_frame/get_first_frame/get_last_frame or probe step's + arguments, reading a file-based 'video' by reference rather than decoding + it (#367, #387). Everything but 'video' is realized the ordinary way. A 'video' naming a real file or an asset/output path becomes a VideoFileReference the task - reads one frame out of by seeking; a 'previous_result:'/'variable:' + reads one frame out of by seeking, or a probe streams; a 'previous_result:'/'variable:' reference is still deferred, and a URL still goes through the ordinary eager fetch_video, since a seek needs a local, seekable file. """ diff --git a/dw/tasks/assess.py b/dw/tasks/assess.py index 769cc2fd..d656de9d 100644 --- a/dw/tasks/assess.py +++ b/dw/tasks/assess.py @@ -197,12 +197,25 @@ def _in_memory_thumbs(frames): return numpy.stack(thumbs) if thumbs else None -def media_from(video): - """A Media from a path or an in-memory AudioVideo.""" +def _file_path(video): + """The validated file a probe's 'video' names, or None for an in-memory + clip: a path, or the VideoFileReference an asset:/output: reference or a + literal path is realized to (dw/arguments.py, #387).""" + from ..locations import validate_media_path + from .video_utils import VideoFileReference + + if isinstance(video, VideoFileReference): + video = video.path if isinstance(video, str): - from ..locations import validate_media_path + return validate_media_path(video, None, "a video to assess") + return None + - return read_media(validate_media_path(video, None, "a video to assess")) +def media_from(video): + """A Media from a path, a VideoFileReference or an in-memory AudioVideo.""" + path = _file_path(video) + if path is not None: + return read_media(path) if hasattr(video, "frames") or hasattr(video, "audio"): from .audio_utils import as_channels_samples @@ -232,8 +245,9 @@ def media_from(video): ) return media raise ValueError( - "a probe takes a video file's path or the video an earlier step " - f"returned, not {type(video).__name__}" + "a probe takes a stored video (asset:, output: or a path) or the " + f"video an earlier step returned, not {type(video).__name__} - a URL " + "downloads as bare frames with no soundtrack to measure" ) @@ -243,11 +257,11 @@ def resolve_shots(video, media, shots=None): return [dict(shot) for shot in shots], "argument" if media.shots: return [dict(shot) for shot in media.shots], "artifact" - if isinstance(video, str): - from ..locations import validate_media_path + path = _file_path(video) + if path is not None: from ..runs import shots_beside - recorded = shots_beside(validate_media_path(video, None, "a video to assess")) + recorded = shots_beside(path) if recorded: return [dict(shot) for shot in recorded], "manifest" return None, "none" diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index aeffb992..bdf4a4d1 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -42,8 +42,9 @@ class VideoFileReference: """A 'video' argument realized to a file on disk rather than an in-memory clip - built by dw/arguments.py's _realize_lazy_frame_arguments so get_frame can seek to the one frame it needs instead of decoding the - whole file (#367). Not a public shape; nothing else constructs or - consumes one.""" + whole file (#367), and so an assessment probe streams the file, soundtrack + and all (#387). Not a public shape; nothing else constructs or consumes + one.""" __slots__ = ("path",) diff --git a/tests/test_assess.py b/tests/test_assess.py index b4c53930..7813c521 100644 --- a/tests/test_assess.py +++ b/tests/test_assess.py @@ -598,3 +598,102 @@ def test_answers_carry_the_common_fields(self): assert "rules_applied" in answer assert "shots_source" in answer assert answer["shots_source"] == "artifact" + + +# --------------------------------------------------------------------------- +# 10. A probe step on a stored video: asset: and output: stream the file (#387) +# --------------------------------------------------------------------------- + + +class TestStoredMediaInAWorkflow: + """A probe's 'video' naming a stored file used to be decoded to a frame + list first - dropping the soundtrack - and every probe then refused it + ("not FrameList"). It is now read by reference, like get_frame's.""" + + def probe_workflow(self, video, command="analyze_shots"): + return { + "id": "probe-stored", + "steps": [ + { + "name": "probe", + "task": {"command": command, "arguments": {"video": video}}, + "result": {"content_type": "application/json"}, + } + ], + } + + def run(self, definition, output_dir): + from dw.workflow import Workflow + + results = Workflow(definition, str(output_dir), "").run({}) + answers = list(results.values()) if isinstance(results, dict) else results + return answers + + def saved_answer(self, output_dir): + saved = [ + p + for p in output_dir.rglob("*.json") + if p.name not in ("manifest.json", "workflow.json") + ] + assert len(saved) == 1, saved + return json.loads(saved[0].read_text()) + + def test_an_asset_video_is_probed_from_the_file(self, tmp_path, monkeypatch): + import dw.arguments as arguments_module + + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + assets = tmp_path / "assets" / "cast" + assets.mkdir(parents=True) + write_mp4(assets / "clip.mp4", frames=48, fps=24) + monkeypatch.setenv("DW_ASSET_DIR", str(tmp_path / "assets")) + + def _boom(*args, **kwargs): + raise AssertionError("a probe's video must not be decoded eagerly") + + monkeypatch.setattr(arguments_module, "load_video", _boom) + + outputs = tmp_path / "outputs" + self.run(self.probe_workflow("asset:cast/clip.mp4"), outputs) + + answer = self.saved_answer(outputs) + assert answer["shots_source"] == "none" + assert len(answer["shots"]) == 1 + # The soundtrack was read, not dropped with a frame-list decode + assert answer["shots"][0]["rms_dbfs"] is not None + + def test_an_output_video_is_probed_with_its_manifest_shots( + self, tmp_path, monkeypatch + ): + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + outputs = tmp_path / "outputs" + run_dir = outputs / "cutter" / "20260923-120000-abcdef01" + (run_dir / "final").mkdir(parents=True) + write_mp4(run_dir / "final" / "cut.mp4", frames=48, fps=24) + (run_dir / "manifest.json").write_text( + json.dumps( + { + "steps": [ + { + "step": "cut", + "files": ["final/cut.mp4"], + "shots": [ + shot_record("shot@a", 0, 24, 0, 48000), + shot_record("shot@b", 24, 24, 48000, 48000), + ], + } + ], + } + ) + ) + + self.run( + self.probe_workflow( + "output:cutter/20260923-120000-abcdef01/final/cut.mp4", + command="analyze_seams", + ), + outputs, + ) + + answer = self.saved_answer(outputs / "probe-stored") + assert answer["shots_source"] == "manifest" + assert len(answer["seams"]) == 1 From fdaa348846a5b5f184482d0d21a0cc45fa9e0f49 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 23:50:42 -0500 Subject: [PATCH 084/181] fix(engine): #393 - keep_output carries a run's shot records into the asset keep_output_as_asset linked or copied only a file's bytes, so a kept multi-shot cut lost the boundaries concat_videos/dissolve_videos/run_chain had recorded for it: analyze_seams on the asset reported shots_source: "none" and every seam rule as applied with findings: [] - a false clean, because nothing had actually been checked. record_kept_shots (dw/runs.py) writes a manifest.json sidecar beside a kept asset in the same shape a run's own manifest uses, reusing the convention shots_beside already reads - so no change was needed in the assessment probes themselves. keep_output_as_asset looks up the source run's shots via recorded_shots and calls it after the link/copy succeeds; get_gallery_metadata now reports media.shots for an asset the same way it already does for an output, via shots_beside. Co-Authored-By: Claude Sonnet 5 --- dw/runs.py | 35 ++++++++++++++++++++ dw/server/app.py | 16 +++++++++ tests/test_server_workspaces.py | 57 +++++++++++++++++++++++++++++++++ 3 files changed, 108 insertions(+) diff --git a/dw/runs.py b/dw/runs.py index 0ab2d4ee..686f966d 100644 --- a/dw/runs.py +++ b/dw/runs.py @@ -682,6 +682,41 @@ def recorded_shots(output_root, relative_path): return None +def record_kept_shots(directory, file_name, shots): + """Write or update the manifest sidecar beside a kept asset so + `shots_beside` can read the shot boundaries the source run recorded for + it (#393). + + Keeping a file copies its bytes but not the run directory it lived in, + so a join's shot records - `pair_audio`'s picture is unchanged, but + nothing carried them past `keep_output` - were unreachable from the + asset and every probe saw `shots_source: "none"`. One manifest per + directory, keyed by file name, in the same shape a run's own + `manifest.json` uses, so the existing manifest-reading path (used by + both outputs and assets) picks it up with no change of its own. A + re-keep replaces the entry for that name rather than leaving a stale + one from a differently-shot source; `shots` of None or [] removes it. + """ + manifest_path = os.path.join(directory, MANIFEST_FILE_NAME) + manifest = _read_manifest(directory) or {} + steps = [ + entry + for entry in manifest.get("steps") or [] + if not (isinstance(entry, dict) and entry.get("files") == [file_name]) + ] + if shots: + steps.append({"step": "keep_output", "files": [file_name], "shots": shots}) + if not steps: + try: + os.remove(manifest_path) + except OSError: + pass + return + manifest["steps"] = steps + with open(manifest_path, "w") as handle: + json.dump(manifest, handle) + + # How far up from a file its run's manifest can sit: the run directory, a # subfolder (`final/`) and the subfolder's own nesting, which # SUBFOLDER_PATTERN caps well below this diff --git a/dw/server/app.py b/dw/server/app.py index 00ab1821..867dfdec 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -100,10 +100,12 @@ REALIZED_FILE_NAME, is_output_reference, is_run_id, + record_kept_shots, record_run_versions, resolve_output_reference, run_versions, recorded_shots, + shots_beside, split_run_path, ) from ..workspace import ( @@ -3036,6 +3038,11 @@ def gallery_metadata( # Where each shot of a joined video sits, as the run that wrote # it recorded (dw/shots.py) - null for a file not joined from shots media["shots"] = recorded_shots(ws.outputs, name) + elif media is not None and source == "asset": + # keep_output carries the source run's shots into a sidecar + # manifest beside the asset (#393); a file kept before that fix, + # or never joined from shots, has none + media["shots"] = shots_beside(path) return { "name": name, "source": source, @@ -3912,6 +3919,15 @@ def keep_output_as_asset( shutil.copy2(source, destination) linked = False + # The source run's shot boundaries - carrying bytes without them left + # a kept multi-shot cut looking like one shot to every probe, with no + # sign anything was missing (#393) + record_kept_shots( + os.path.dirname(destination), + os.path.basename(destination), + recorded_shots(ws.outputs, kept_name), + ) + logger.info(f"Kept output {request.name} as asset:{asset_name}") return { "reference": f"asset:{asset_name}", diff --git a/tests/test_server_workspaces.py b/tests/test_server_workspaces.py index 420e43a0..8be35e99 100644 --- a/tests/test_server_workspaces.py +++ b/tests/test_server_workspaces.py @@ -2,6 +2,7 @@ root, each with its own workflows, assets and outputs, all sharing the one prompt library.""" +import json import os import pytest @@ -516,6 +517,62 @@ def test_a_destination_cannot_leave_the_library( ) assert response.status_code == 400 + def test_a_kept_outputs_shots_survive_and_report_through_the_gallery( + self, server, workspace_root + ): + """#393: keeping a cut copied only its bytes, so a joined video's shot + boundaries were unreachable from the asset it became - the gallery + metadata for a kept asset carried no `shots` at all, where the same + file's metadata as an output did. keep_output now carries the run's + recorded shots into a manifest sidecar beside the asset, which is the + same convention `shots_beside` (and so every assessment probe) already + reads.""" + from .test_media_info import write_mp4 + + run_dir = os.path.join(workspace_root.outputs, "Gyre/20260905-101500-aaaaaaaa") + os.makedirs(run_dir, exist_ok=True) + write_mp4(os.path.join(run_dir, "cut.mp4"), frames=18, fps=6) + shots = [ + { + "name": "a", + "start_frame": 0, + "num_frames": 10, + "start_sample": 0, + "num_samples": 100, + }, + { + "name": "b", + "start_frame": 10, + "num_frames": 8, + "start_sample": 100, + "num_samples": 80, + }, + ] + manifest = { + "steps": [{"step": "concat", "files": ["cut.mp4"], "shots": shots}] + } + with open(os.path.join(run_dir, "manifest.json"), "w") as handle: + json.dump(manifest, handle) + + with server() as client: + kept = client.post( + "/api/assets/keep", + json={ + "name": "Gyre/20260905-101500-aaaaaaaa/cut.mp4", + "asset_name": "qa-cast/cut.mp4", + }, + ) + assert kept.status_code == 201 + + metadata = client.get( + "/api/gallery/asset:qa-cast/cut.mp4/metadata" + ).json() + assert metadata["source"] == "asset" + assert metadata["media"]["shots"] == shots + + sidecar = os.path.join(workspace_root.assets, "qa-cast", "manifest.json") + assert os.path.isfile(sidecar) + def test_keeping_stays_inside_the_workspace(self, server, workspace_root): """The source is read from the named workspace's outputs and the copy lands in its assets - neither reaches the default workspace.""" From 4740f2ec39515b0dcd0e2e54607c0a29cb84d02c Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Wed, 23 Sep 2026 23:51:40 -0500 Subject: [PATCH 085/181] style(tests): #393 - ruff format the keep_output shots test Co-Authored-By: Claude Sonnet 5 --- tests/test_server_workspaces.py | 8 ++------ 1 file changed, 2 insertions(+), 6 deletions(-) diff --git a/tests/test_server_workspaces.py b/tests/test_server_workspaces.py index 8be35e99..872c7a2c 100644 --- a/tests/test_server_workspaces.py +++ b/tests/test_server_workspaces.py @@ -548,9 +548,7 @@ def test_a_kept_outputs_shots_survive_and_report_through_the_gallery( "num_samples": 80, }, ] - manifest = { - "steps": [{"step": "concat", "files": ["cut.mp4"], "shots": shots}] - } + manifest = {"steps": [{"step": "concat", "files": ["cut.mp4"], "shots": shots}]} with open(os.path.join(run_dir, "manifest.json"), "w") as handle: json.dump(manifest, handle) @@ -564,9 +562,7 @@ def test_a_kept_outputs_shots_survive_and_report_through_the_gallery( ) assert kept.status_code == 201 - metadata = client.get( - "/api/gallery/asset:qa-cast/cut.mp4/metadata" - ).json() + metadata = client.get("/api/gallery/asset:qa-cast/cut.mp4/metadata").json() assert metadata["source"] == "asset" assert metadata["media"]["shots"] == shots From 8c86cbafd6884f07d0b00d30a5554748d4c4b36c Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 00:04:17 -0500 Subject: [PATCH 086/181] fix(engine): #394 - analyze_seams/analyze_shots report rules_skipped, not rules_applied, when shots_source is none A file with no shot records at all (an asset kept before #393, an upload, a cut joined outside dw) read as a clean pass: rules_applied listed every seam or shot-spread rule with findings: [], even though nothing was measured. _answer() now takes the probe's shot-dependent rule names and, when shots_source is "none", moves them to rules_skipped with a "no shot boundaries" reason and emits a job warning naming shots= as the fix. analyze_sync_drift is untouched - its rules are file-level, not shot-boundary-dependent, and the issue didn't name it. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/assess.py | 45 +++++++++++++++++++++++++++++++++++------- tests/test_assess.py | 47 ++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 85 insertions(+), 7 deletions(-) diff --git a/dw/tasks/assess.py b/dw/tasks/assess.py index d656de9d..9e8f552c 100644 --- a/dw/tasks/assess.py +++ b/dw/tasks/assess.py @@ -29,6 +29,7 @@ import numpy from ..assessment_rules import HOLE_VOICED_DBFS, crosses, finding, rules_for +from ..events import emit_warning logger = logging.getLogger("dw") @@ -332,11 +333,37 @@ def _findings(probe, record, at, skip=()): return found -def _answer(probe, measurements, findings, shots_source): +def _answer(probe, measurements, findings, shots_source, shot_dependent=()): + """A probe's answer, with `rules_applied` cut down to the rules that + actually ran. `shot_dependent` names (or `True` for all of the probe's + rules) the ones that only mean anything measured shot against shot; with + no shot boundaries at all (`shots_source == "none"`) those are reported + as `rules_skipped` instead of `rules_applied`, and a run warning says why + - without this a shotless file (an asset kept before #393, an upload, a + cut joined outside dw) read as a clean pass with nothing measured (#394). + """ + names = [rule["name"] for rule in rules_for(probe)] + dependent = set(names) if shot_dependent is True else set(shot_dependent) + applied, skipped = names, [] + if shots_source == "none" and dependent: + applied = [name for name in names if name not in dependent] + skipped = [ + {"rule": name, "reason": "no shot boundaries"} + for name in names + if name in dependent + ] + emit_warning( + f"{probe} found no shot boundaries for this file, so " + f"{', '.join(sorted(dependent))} could not be measured - pass " + "shots= to supply them", + kind="no_shot_boundaries", + probe=probe, + ) return { **measurements, "findings": findings, - "rules_applied": [rule["name"] for rule in rules_for(probe)], + "rules_applied": applied, + "rules_skipped": skipped, "shots_source": shots_source, } @@ -352,7 +379,7 @@ def analyze_shots(video, shots=None): Returns: {shots: [{name, start_frame, num_frames, peak_dbfs, rms_dbfs, crest_db, low_dbfs, mid_dbfs, high_dbfs, samples}], rms_range_db, findings, rules_applied, - shots_source} + rules_skipped, shots_source} """ from .audio_utils import _spectral_balance @@ -364,6 +391,7 @@ def analyze_shots(video, shots=None): {"shots": [], "rms_range_db": None, "has_audio": False}, [], source, + shot_dependent={"shot_level_spread"}, ) records = records or [_whole_file_shot(media)] @@ -407,6 +435,7 @@ def analyze_shots(video, shots=None): answer, _findings("analyze_shots", answer, at) if at else [], source, + shot_dependent={"shot_level_spread"}, ) @@ -553,12 +582,12 @@ def analyze_seams(video, shots=None): before_shot_rms_dbfs, after_shot_rms_dbfs, floor_dbfs, click_db, spectral_shift, before_rms_dbfs, after_rms_dbfs, frame_delta, typical_delta, jump_ratio}], findings, rules_applied, - shots_source} + rules_skipped, shots_source} """ media = media_from(video) records, source = resolve_shots(video, media, shots) if not records or len(records) < 2: - return _answer("analyze_seams", {"seams": []}, [], source) + return _answer("analyze_seams", {"seams": []}, [], source, shot_dependent=True) seams = [] findings = [] @@ -619,7 +648,9 @@ def analyze_seams(video, shots=None): skip, ) ) - return _answer("analyze_seams", {"seams": seams}, findings, source) + return _answer( + "analyze_seams", {"seams": seams}, findings, source, shot_dependent=True + ) def analyze_sync_drift(video, shots=None): @@ -633,7 +664,7 @@ def analyze_sync_drift(video, shots=None): Returns: {shots: [{name, start_offset_ms, end_offset_ms}], max_offset_ms, video_seconds, audio_seconds, length_delta_ms, findings, - rules_applied, shots_source} + rules_applied, rules_skipped, shots_source} """ media = media_from(video) records, source = resolve_shots(video, media, shots) diff --git a/tests/test_assess.py b/tests/test_assess.py index 7813c521..5bb1fe99 100644 --- a/tests/test_assess.py +++ b/tests/test_assess.py @@ -451,6 +451,53 @@ def test_probes_on_an_encoded_file(self, tmp_path, monkeypatch): json.dumps(seams_answer) json.dumps(drift_answer) + def test_shotless_file_reports_skipped_rules_not_a_clean_pass( + self, tmp_path, monkeypatch + ): + """#394: shots_source "none" used to report every rule as applied + with findings: [] - a false clean, since no seam or shot spread was + actually measured. It must say the rules were skipped instead, and + warn that shots= would supply the missing boundaries.""" + monkeypatch.setenv("DW_TRUST_WORKFLOWS", "1") + import dw.tasks.assess as assess_module + + warnings = [] + monkeypatch.setattr( + assess_module, + "emit_warning", + lambda message, **data: warnings.append((message, data)), + ) + + path = tmp_path / "two-shots.mp4" + write_mp4(path, frames=48, fps=24, width=64, height=64, sample_rate=48000) + + seams_answer = analyze_seams(str(path)) + assert seams_answer["shots_source"] == "none" + assert seams_answer["rules_applied"] == [] + assert seams_answer["rules_skipped"] == [ + {"rule": name, "reason": "no shot boundaries"} + for name in ( + "seam_level_step", + "seam_click", + "seam_hole", + "seam_frame_jump", + ) + ] + assert any("no shot boundaries" in w[0].lower() for w in warnings) + assert any(w[1].get("kind") == "no_shot_boundaries" for w in warnings) + warnings.clear() + + shots_answer = analyze_shots(str(path)) + assert shots_answer["shots_source"] == "none" + assert "shot_level_spread" not in shots_answer["rules_applied"] + assert shots_answer["rules_skipped"] == [ + {"rule": "shot_level_spread", "reason": "no shot boundaries"} + ] + assert any("no shot boundaries" in w[0].lower() for w in warnings) + + json.dumps(seams_answer) + json.dumps(shots_answer) + # --------------------------------------------------------------------------- # 8. shots_beside reads a manifest, and a probe reports shots_source manifest From 612d91fbb531ed93eb0ee139604df5672b14578e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 00:17:29 -0500 Subject: [PATCH 087/181] fix(engine): #395 - gain_audio with no region gains the whole track validate_workflow let a region-less gain_audio step through clean and the run then failed 1s in with "needs either 'start_seconds'/'duration_seconds' or 'start_frame'/'num_frames'/'fps'". With every region argument omitted, gain_audio now applies the gain to the whole track - the same "no region means everything" reading mix_audio's gains already use, and the obvious meaning of "duck this clip by 8 dB". Docstring, get_task's parameter docs (via the docstring) and docs/TASKS.md say so; no new validate rule needed since there is no longer a missing-region case to refuse. Co-Authored-By: Claude Sonnet 5 --- docs/TASKS.md | 15 ++++++++------- dw/tasks/audio_utils.py | 21 +++++++++++---------- tests/test_audio_utils.py | 12 +++++++++++- 3 files changed, 30 insertions(+), 18 deletions(-) diff --git a/docs/TASKS.md b/docs/TASKS.md index 5dd65180..dabeda3c 100644 --- a/docs/TASKS.md +++ b/docs/TASKS.md @@ -523,15 +523,16 @@ material: | -------- | -------- | ----------- | | `audio` | Yes | Path or URL of an audio file (or of a video file, whose soundtrack is taken), a waveform from a previous step, or an earlier step's video generated with a soundtrack (which brings its sample rate along) | | `gain_db` | Yes | Gain to apply within the region, in decibels - negative ducks it, positive boosts it | -| `start_seconds` / `duration_seconds` | One pair | The region in seconds; either may be omitted | -| `start_frame` / `num_frames` / `fps` | One pair | The region in video frames; `fps` is required, start and count may be omitted | +| `start_seconds` / `duration_seconds` | No | The region in seconds; either may be omitted | +| `start_frame` / `num_frames` / `fps` | No | The region in video frames; `fps` is required if either is given, start and count may be omitted | | `sample_rate` | With a waveform | Sample rate of a directly passed waveform (files carry their own) | -One pair is required — there is no separate "whole track" mode — but the -whole track is still one step: give just `start_seconds: 0` and leave -`duration_seconds` unset (or `start_frame: 0` + `fps` and leave `num_frames` -unset), which runs to the end of the track without needing to already know -how long that is. +No region argument is required: with every one of them omitted, the gain +applies to the whole track (#395) - the same "no region means everything" +reading `mix_audio`'s gains use. To gain everything from some point on +instead, give just `start_seconds: 0` and leave `duration_seconds` unset (or +`start_frame: 0` + `fps` and leave `num_frames` unset), which runs to the +end of the track without needing to already know how long that is. ### crossfade_audio diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index 8ef69cb7..c42d2c2f 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -548,11 +548,13 @@ def gain_audio( passed through unchanged, so ducking a scene under another is one step rather than the slice/gain/mix/rejoin/pair_audio chain that was previously the only way to apply a gain to part of a track rather than - all of it (#187). At least one of the two pairs is required - there is - no separate "whole track" mode - but the whole track is still one step: - give just start_seconds=0 (or start_frame=0 + fps) and leave - duration_seconds/num_frames unset, which runs to the end of the track - without the caller needing to already know how long that is. + all of it (#187). With no region given at all, the gain applies to the + whole track - the same "no region means everything" reading mix_audio's + gains use, and the obvious meaning of "duck this clip by 8 dB" (#395). + To gain everything from some point on, give just start_seconds=0 (or + start_frame=0 + fps) and leave duration_seconds/num_frames unset, which + runs to the end of the track without the caller needing to already know + how long that is. Unlike slice_audio, a region reaching past the end of the track is clipped to it rather than zero-padded: there is no silence there to @@ -569,7 +571,8 @@ def gain_audio( soundtrack, or a waveform (which needs sample_rate alongside it) gain_db: Gain to apply within the region, in decibels - negative ducks it, positive boosts it - start_seconds: Start of the region, in seconds + start_seconds: Start of the region, in seconds. Omitted along with + every other region argument, the gain applies to the whole track duration_seconds: Length of the region, in seconds start_frame: Start of the region, in video frames num_frames: Length of the region, in video frames @@ -622,10 +625,8 @@ def gain_audio( else frames_to_samples(num_frames, fps, sample_rate) ) else: - raise ValueError( - "gain_audio needs either 'start_seconds'/'duration_seconds' or " - "'start_frame'/'num_frames'/'fps' to address the region to gain" - ) + start = 0 + length = total region_start = max(0, min(start, total)) region_end = max(region_start, min(start + max(length, 0), total)) diff --git a/tests/test_audio_utils.py b/tests/test_audio_utils.py index 7da2e452..70ee39f2 100644 --- a/tests/test_audio_utils.py +++ b/tests/test_audio_utils.py @@ -590,8 +590,18 @@ def test_the_applied_gain_is_logged(self): assert logs[0]["duration_seconds"] == pytest.approx(0.5) assert logs[0]["sample_rate"] == 100 + def test_no_region_gains_the_whole_track(self): + # #395: validate_workflow let a region-less gain_audio step through + # clean and the run then failed - the fix is to gain everything, + # matching mix_audio's "no region means everything" reading + from dw.tasks.audio_utils import gain_audio + + track = numpy.ones((1, 100), dtype=numpy.float32) + + gained = samples(gain_audio(track, gain_db=-6.0, sample_rate=100)) + + assert numpy.allclose(gained, 10 ** (-6.0 / 20)) -class TestNormalizeAudio: def test_the_peak_lands_on_the_target(self): from dw.tasks.audio_utils import normalize_audio From 6b0e943f1d1b6a904b5d941fe0d9990ac1623095 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 00:27:10 -0500 Subject: [PATCH 088/181] fix(engine): #396 - name a joined shot after its previous_result step shot_reference_names only recognized a shot@ for_each member; a plain previous_result: reference (e.g. to a pair_audio or chain step) fell through to concat_videos'/dissolve_videos' "video N" fallback in the manifest and in every probe finding built from it, with nothing to trace the shot back to the step that produced it. Now any previous_result: reference names the shot after the step it points at. --- dw/shots.py | 16 +++++++++++++--- tests/test_shots.py | 19 +++++++++++++++++++ 2 files changed, 32 insertions(+), 3 deletions(-) diff --git a/dw/shots.py b/dw/shots.py index 174bc992..f4ee68b4 100644 --- a/dw/shots.py +++ b/dw/shots.py @@ -118,13 +118,17 @@ def remeasured_shots(shots, fps, sample_rate, total_samples): def shot_reference_names(references): - """`shot@` for each entry of a step's list naming a shot, else None. + """A name per entry of a step's list naming a shot, else None. A step's `videos` argument is written as a list of `previous_result:` references; by the time the task runs `gather:` has expanded into exactly such a list, so the entry at position i names the video the join put at - position i. Only a reference to a `shot@` member names a shot - anything - else keeps the name the join gave it. + position i. A reference to a `shot@` member keeps that name (the + for_each entry, not the field read off it); any other `previous_result:` + reference is named after the step it points at, so a shot generated by an + ordinary step (a `pair_audio`, a chain) is traceable in the manifest and + in a probe finding the same way (#396). Anything else (an `asset:` path, + a literal video) keeps the name the join gave it. """ if not isinstance(references, list): return None @@ -134,6 +138,12 @@ def shot_reference_names(references): # `previous_result:shot@x.field` names the member, not the field member = reference[len(PREVIOUS_RESULT_PREFIX) :] names.append(member.split(".", 1)[0]) + elif isinstance(reference, str) and reference.startswith( + PREVIOUS_RESULT_PREFIX + ): + # `previous_result:step.field` names the step, not the field + step = reference[len(PREVIOUS_RESULT_PREFIX) :] + names.append(step.split(".", 1)[0]) else: names.append(None) return names diff --git a/tests/test_shots.py b/tests/test_shots.py index 3c447330..1ee28744 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -222,6 +222,25 @@ def test_shot_reference_names_name_shot_at_members(self): # "shot@" marker, which is how a for_each member is named elsewhere assert [shot["name"] for shot in named] == ["shot@a", "shot@b"] + def test_shot_reference_names_name_an_ordinary_step(self): + """A `previous_result:` reference that does not name a `shot@` + for_each member still names a shot - after the step it points at, + so a shot from an ordinary step (a `pair_audio`, a chain) is + traceable rather than falling back to "video N" (#396).""" + videos = [audio_video(4, 1), audio_video(4, 2)] + + result = concat_videos(videos, fps=4) + named = named_shots( + result.shots, + shot_reference_names( + ["asset:ep31-shot1-return.mp4", "previous_result:shot2d"] + ), + ) + + # position 0 keeps whatever name the join itself gave it - only the + # previous_result reference at position 1 is renamed + assert [shot["name"] for shot in named] == ["video 1", "shot2d"] + def test_no_audio_input_leaves_sample_fields_none(self): result = concat_videos([frames(4), frames(3)]) From c30b1d2e875c79bc0d62055064807b252789a35e Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 00:47:36 -0500 Subject: [PATCH 089/181] fix(engine): #396 - reach the live artifact when naming a joined shot The prior fix (6b0e943) renamed shots only in the deep copy step_shots returns for the manifest, but Result.save stores saved_shots[path] as the artifact's own .shots list, not a copy - and that same object is what a later previous_result: step reads. A probe (analyze_seams, analyze_sync_drift) run against previous_result:cut therefore still saw the join's "video N" fallback even though the manifest was correct. step_shots now renames the shot dicts in place before deep-copying for the manifest, so the rename reaches both the manifest and any later previous_result: consumer, including probe findings. Co-Authored-By: Claude Sonnet 5 --- dw/shots.py | 23 +++++++++++++++++++++-- tests/test_shots.py | 39 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 60 insertions(+), 2 deletions(-) diff --git a/dw/shots.py b/dw/shots.py index f4ee68b4..8bc776c4 100644 --- a/dw/shots.py +++ b/dw/shots.py @@ -158,6 +158,23 @@ def named_shots(shots, names): ] +def _rename_in_place(shots, names): + """Write the step-named `name`s back onto the shot dicts themselves. + + `Result.save` stores `saved_shots[path]` as the artifact's own `.shots` + list, not a copy (`self._artifacts_for` / `getattr(artifact, "shots")`), + and that same artifact is what a later `previous_result:` step reads + (`results[step.name]` holds it directly). Renaming only the manifest's + deep copy left that artifact carrying the join's `video N` fallback, so + a probe reading `previous_result:cut` still saw the unnamed shot even + after the manifest was fixed (#396 follow-up). Mutating here reaches + both. + """ + for shot, named in zip(shots, named_shots(shots, names)): + if named is not shot: + shot["name"] = named["name"] + + def step_shots(saved_shots, saved_files, references=None): """The `shots` a step's manifest entry and step_end carry, or None. @@ -171,12 +188,14 @@ def step_shots(saved_shots, saved_files, references=None): return None names = shot_reference_names(references) files = [path for path in saved_files or [] if path in saved_shots] + for path in files: + _rename_in_place(saved_shots[path], names) if len(files) == 1 and len(saved_files) == 1: - return named_shots(copy.deepcopy(saved_shots[files[0]]), names) + return copy.deepcopy(saved_shots[files[0]]) return [ {**shot, "file": path} for path in files - for shot in named_shots(copy.deepcopy(saved_shots[path]), names) + for shot in copy.deepcopy(saved_shots[path]) ] diff --git a/tests/test_shots.py b/tests/test_shots.py index 1ee28744..47cc3eaf 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -594,6 +594,45 @@ def test_shots_survive_save_and_recorded_shots(self, tmp_path): assert read_back == manifest_shots + def test_step_shots_renames_reach_the_artifact_a_later_step_would_read( + self, tmp_path + ): + """A later step's `previous_result:cut` reads the same artifact + object `Result.save` extracted (`Result._artifacts_for`'s cache, by + identity) - not a fresh copy - so a probe reading `previous_result: + cut` after this step must see the step-named shot too, not just the + manifest built alongside it. Renaming only a deep copy for the + manifest left the artifact itself carrying the join's `video 2` + fallback (#396 follow-up).""" + videos = [audio_video(4, 1), audio_video(4, 2)] + joined = concat_videos(videos, fps=4) + + result = Result({"content_type": "video/mp4", "save": True}) + result.add_result(joined) + + run_dir = tmp_path / "cut-demo" / "20260924-000000-abcdef01" + run_dir.mkdir(parents=True) + with ( + patch("dw.result.encode_video"), + patch("dw.result.export_to_video"), + patch("dw.result.is_av_available", return_value=True), + ): + saved_files = result.save(str(run_dir), "cut-demo-join.0") + + step_shots( + result.saved_shots, + saved_files, + references=["asset:ep31-shot1-return.mp4", "previous_result:shot2d"], + ) + + # What a later step reads via previous_result:cut (get_artifacts, the + # same-identity artifact) - not the manifest's deep copy + artifact = result.get_artifacts()[0] + assert [shot["name"] for shot in artifact.shots] == [ + "video 1", + "shot2d", + ] + def test_recorded_shots_is_none_outside_a_run_directory(self, tmp_path): # The flat layout - no run id segment - has no manifest to read shots from assert recorded_shots(str(tmp_path), "workflow/still.png") is None From df796cf252a16362438f413f3f402f76db52d6a8 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 01:27:34 -0500 Subject: [PATCH 090/181] fix(mcp): #397 - suggest catalog names for an unresolved workflow get_workflow/validate_workflow/run_workflow answered a bare "Unknown workflow: x" when x was a short name or typo for a real catalog entry (e.g. "dialogue-short" for "templates/minimax/dialogue-short"), forcing a list_workflows round trip to find it. suggest_workflow_names() resolves a unique catalog path suffix, falling back to a close spelling match, and both resolve_readable_workflow (get_workflow, validate_workflow(name=), delete/download routes) and resolve_workflow_reference (run_workflow/validate_workflow's workflow_path) now append "- did you mean ...?" when it finds one. Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 30 ++++++++++++++++++++++++------ dw/workflow_sources.py | 25 +++++++++++++++++++++++++ tests/test_server.py | 24 ++++++++++++++++++++++++ tests/test_workflow_sources.py | 32 ++++++++++++++++++++++++++++++++ 4 files changed, 105 insertions(+), 6 deletions(-) diff --git a/dw/server/app.py b/dw/server/app.py index 867dfdec..3a73e7c6 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -134,6 +134,7 @@ resolve_in_source, resolve_sub_workflow, source_for_path, + suggest_workflow_names, workflow_names, workflow_sources, writable_source, @@ -449,6 +450,21 @@ def _write_bytes(path, data): f.write(data) +def _unknown_workflow_detail(sources, name): + """'Unknown workflow: x', with a '- did you mean ...?' pointer when the + catalog holds something `name` could be short for or a typo of (#397) - + otherwise a caller has to spend a list_workflows call and guess the + right shape/traits to find the entry it already knows by its short + name.""" + detail = f"Unknown workflow: {name}" + suggestions = suggest_workflow_names(sources, name) + if len(suggestions) == 1: + detail += f" - did you mean {suggestions[0]}?" + elif suggestions: + detail += f" - did you mean one of: {', '.join(suggestions)}?" + return detail + + def resolve_readable_workflow(sources, name): """The path a name has anywhere on the search path, and its source. @@ -458,7 +474,7 @@ def resolve_readable_workflow(sources, name): """ path, source = find_workflow(sources, name) if path is None: - raise HTTPException(status_code=404, detail=f"Unknown workflow: {name}") + raise HTTPException(status_code=404, detail=_unknown_workflow_detail(sources, name)) return path, source @@ -521,11 +537,13 @@ def resolve_workflow_reference(workflow_path, sources): confined = None if confined is not None and os.path.isfile(confined): return confined, source - raise HTTPException( - status_code=400, - detail=f"workflow_path must name a workflow the server can reach: " - f"{workflow_path}", - ) + detail = f"workflow_path must name a workflow the server can reach: {workflow_path}" + suggestions = suggest_workflow_names(sources, workflow_path) + if len(suggestions) == 1: + detail += f" - did you mean {suggestions[0]}?" + elif suggestions: + detail += f" - did you mean one of: {', '.join(suggestions)}?" + raise HTTPException(status_code=400, detail=detail) # What each prompt says about itself, for listing cards - cached by mtime diff --git a/dw/workflow_sources.py b/dw/workflow_sources.py index 62309d37..69b9c860 100644 --- a/dw/workflow_sources.py +++ b/dw/workflow_sources.py @@ -174,6 +174,31 @@ def find_workflow(sources, name): return None, None +def suggest_workflow_names(sources, name, limit=3): + """Catalog names an unresolved `name` might have meant, for an error + message rather than a second round trip. + + The catalog is organised in directories (`templates/minimax/dialogue-short`) + and a caller - a skill, an earlier turn - often has only the trailing + name (`dialogue-short`). Preferred answer: every catalog entry `name` is + a unique path suffix of, since that is unambiguous; failing that, a + close spelling match (`difflib`), for a typo rather than a shortened + path. Empty when neither finds anything worth naming. + """ + import difflib + + stripped = name[: -len(".json")] if name.endswith(".json") else name + catalog_names = list(listing(sources).keys()) + suffix_matches = [ + candidate + for candidate in catalog_names + if candidate == stripped or candidate.endswith(f"/{stripped}") + ] + if suffix_matches: + return suffix_matches[:limit] + return difflib.get_close_matches(stripped, catalog_names, n=limit, cutoff=0.6) + + def listing(sources): """Every name the search path offers, each with the source it comes from - a name in an earlier source shadowing the same name later.""" diff --git a/tests/test_server.py b/tests/test_server.py index 5d27691e..aae85318 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -536,6 +536,16 @@ def test_submit_rejects_a_traversal_shaped_name(server): assert client.app.state.job_manager.worker_manager.commands == [] +def test_submit_names_a_suggestion_for_a_short_workflow_path(server): + """#397: a workflow_path that's a catalog entry's trailing segment (the + name a skill or earlier turn is likely to say) points at the entry + rather than a bare 400.""" + with server(success_script) as client: + response = client.post("/api/jobs", json={"workflow_path": "asic"}) + assert response.status_code == 400 + assert "did you mean Basic?" in response.json()["detail"] + + def test_validate_accepts_a_stored_workflow_name(server, tmp_path): with server(success_script) as client: result = client.post("/api/validate", json={"workflow_path": "Basic"}).json() @@ -682,6 +692,20 @@ def test_workflow_browsing_and_confinement(server): assert client.get("/api/workflows/../secret").status_code == 404 assert client.get("/api/workflows/nope").status_code == 404 + # #397: a name that is a catalog entry's trailing segment points at + # the entry it's short for, rather than a bare 404 + client.put("/api/workflows/sub/Basic", json={"workflow": valid_workflow()}) + missed = client.get("/api/workflows/nope2") + assert missed.status_code == 404 + assert "did you mean" not in missed.json()["detail"] + + short = client.get("/api/workflows/Basic") + assert short.status_code == 200 # the top-level name still shadows the nested one + + suggested = client.get("/api/workflows/sub%2FBasik") + assert suggested.status_code == 404 + assert "sub/Basic" in suggested.json()["detail"] + def test_configures_resolves_against_the_listing(server): with server(success_script) as client: diff --git a/tests/test_workflow_sources.py b/tests/test_workflow_sources.py index 4eb478be..a1ac4568 100644 --- a/tests/test_workflow_sources.py +++ b/tests/test_workflow_sources.py @@ -17,6 +17,7 @@ resolve_in_source, resolve_sub_workflow, source_for_path, + suggest_workflow_names, workflow_names, workflow_sources, writable_source, @@ -111,6 +112,37 @@ def test_a_path_knows_which_source_it_belongs_to(self, roots, tmp_path): assert source_for_path(sources, str(tmp_path / "elsewhere.json")) is None +class TestSuggestions: + """#397: a caller who knows a catalog entry by its short name gets a + pointer to the real one rather than a bare 404.""" + + def test_a_unique_path_suffix_is_suggested(self, roots): + workspace, examples = roots + sources = workflow_sources(str(workspace), [str(examples)]) + assert suggest_workflow_names(sources, "Gyre") == ["ltx2/Gyre"] + + def test_a_typo_falls_back_to_a_close_spelling_match(self, roots): + workspace, examples = roots + sources = workflow_sources(str(workspace), [str(examples)]) + assert suggest_workflow_names(sources, "Shered") == ["Shared"] + + def test_nothing_close_suggests_nothing(self, roots): + workspace, examples = roots + sources = workflow_sources(str(workspace), [str(examples)]) + assert suggest_workflow_names(sources, "zzz-completely-unrelated") == [] + + def test_the_real_catalog_suggests_the_full_template_path(self): + # #397's own repro: a skill or an earlier turn names a template by + # its short id, not its catalog path + repo_root = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) + sources = workflow_sources( + os.path.join(repo_root, "workflows"), include_builtin=True + ) + assert suggest_workflow_names(sources, "dialogue-short") == [ + "templates/minimax/dialogue-short" + ] + + class TestSubWorkflowResolution: """A composed step's relative path is confined to the root it is handed back with, so a name that climbs out of the catalog is never resolved - From 57750da8dd171aafc52b9262658c98790bec1fb6 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 01:28:53 -0500 Subject: [PATCH 091/181] style(mcp): #397 - satisfy ruff format on the unknown-workflow suggestion Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 4 +++- tests/test_server.py | 4 +++- 2 files changed, 6 insertions(+), 2 deletions(-) diff --git a/dw/server/app.py b/dw/server/app.py index 3a73e7c6..01ca8044 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -474,7 +474,9 @@ def resolve_readable_workflow(sources, name): """ path, source = find_workflow(sources, name) if path is None: - raise HTTPException(status_code=404, detail=_unknown_workflow_detail(sources, name)) + raise HTTPException( + status_code=404, detail=_unknown_workflow_detail(sources, name) + ) return path, source diff --git a/tests/test_server.py b/tests/test_server.py index aae85318..33779315 100644 --- a/tests/test_server.py +++ b/tests/test_server.py @@ -700,7 +700,9 @@ def test_workflow_browsing_and_confinement(server): assert "did you mean" not in missed.json()["detail"] short = client.get("/api/workflows/Basic") - assert short.status_code == 200 # the top-level name still shadows the nested one + assert ( + short.status_code == 200 + ) # the top-level name still shadows the nested one suggested = client.get("/api/workflows/sub%2FBasik") assert suggested.status_code == 404 From 83df0e62f29d51f6afb9234bb694bb094d27d867 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 01:41:32 -0500 Subject: [PATCH 092/181] fix(engine): #397 - measure a close-match typo against the catalog entry's own name difflib.get_close_matches compared a bare query like "dialog-short" against the full catalog path "templates/minimax/dialogue-short" (ratio 0.55, below the 0.6 cutoff), missing a real one-letter typo the issue's own follow-up flagged. Compare basenames on both sides instead (query's own trailing segment against each catalog entry's trailing segment), mapping back to the full path - "dialog-short" now scores 0.92 against "dialogue-short" and surfaces the suggestion. Co-Authored-By: Claude Sonnet 5 --- dw/workflow_sources.py | 21 ++++++++++++++++++++- tests/test_workflow_sources.py | 12 ++++++++++++ 2 files changed, 32 insertions(+), 1 deletion(-) diff --git a/dw/workflow_sources.py b/dw/workflow_sources.py index 69b9c860..8f2167bb 100644 --- a/dw/workflow_sources.py +++ b/dw/workflow_sources.py @@ -196,7 +196,26 @@ def suggest_workflow_names(sources, name, limit=3): ] if suffix_matches: return suffix_matches[:limit] - return difflib.get_close_matches(stripped, catalog_names, n=limit, cutoff=0.6) + + # A typo is measured against the catalog entry's own name, not against + # its directory prefix - "dialog-short" scores 0.92 against + # "dialogue-short" and 0.55 against "templates/minimax/dialogue-short", + # so comparing full paths lets a real typo miss the cutoff (#397). The + # query's own prefix is stripped the same way, so "sub/Basik" is + # measured as "Basik" against "Basic" rather than against "sub/Basic". + query_basename = stripped.rsplit("/", 1)[-1] + by_basename = {} + for candidate in catalog_names: + by_basename.setdefault(candidate.rsplit("/", 1)[-1], []).append(candidate) + close_bases = difflib.get_close_matches( + query_basename, list(by_basename.keys()), n=limit, cutoff=0.6 + ) + matches = [] + for base in close_bases: + for candidate in by_basename[base]: + if candidate not in matches: + matches.append(candidate) + return matches[:limit] def listing(sources): diff --git a/tests/test_workflow_sources.py b/tests/test_workflow_sources.py index a1ac4568..168acbb0 100644 --- a/tests/test_workflow_sources.py +++ b/tests/test_workflow_sources.py @@ -142,6 +142,18 @@ def test_the_real_catalog_suggests_the_full_template_path(self): "templates/minimax/dialogue-short" ] + def test_a_typo_on_a_short_name_still_finds_the_full_catalog_path(self): + # The tester's own follow-up: "dialog-short" scores 0.92 against the + # entry's own name but 0.55 against the full path, so comparing + # full paths missed a real typo entirely + repo_root = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) + sources = workflow_sources( + os.path.join(repo_root, "workflows"), include_builtin=True + ) + assert suggest_workflow_names(sources, "dialog-short") == [ + "templates/minimax/dialogue-short" + ] + class TestSubWorkflowResolution: """A composed step's relative path is confined to the root it is handed From 25fd20c93c6636bae0a5ab83dea99dc1b0129644 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 01:57:04 -0500 Subject: [PATCH 093/181] fix(engine): #398 - pair_audio carries a loaded video's recorded shots A video loaded by path (an asset:/output: reference) went through diffusers' own load_video, which answers a bare frame list with no shots - even though pair_audio already knew how to remeasure them (#385) and shots_beside(path) already finds them, in a run manifest or a kept-asset sidecar (#393). Nothing wired the two together, so rescoring a kept episode lost media.shots and left analyze_seams blind. FrameList (added for the analogous fps loss, #104) now also carries shots, and _with_frame_rate looks them up beside the fps for any local path. Co-Authored-By: Claude Sonnet 5 --- dw/arguments.py | 25 ++++++++--- dw/tasks/video_utils.py | 12 ++++- tests/test_video_utils.py | 94 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 122 insertions(+), 9 deletions(-) diff --git a/dw/arguments.py b/dw/arguments.py index 21a07818..5fb56111 100644 --- a/dw/arguments.py +++ b/dw/arguments.py @@ -1029,19 +1029,30 @@ def fetch_image(img_spec, base_dir=None): def _with_frame_rate(frames, location): - """The loaded frames carrying the rate their file declares. - - `load_video` reads frames and drops the rate, so a step that paired a - 24 fps file with a soundtrack wrote it back at 8 - three times long, - silently (#104). The rate is read from the container without decoding - anything, and a file that will not say stays a plain list. + """The loaded frames carrying the rate their file declares, and the + shot boundaries its run (or kept-asset sidecar) recorded for it. + + `load_video` reads frames and drops both: a step that paired a 24 fps + file with a soundtrack wrote it back at 8 - three times long, silently + (#104) - and a video loaded from an `asset:`/`output:` path had no + `shots` to hand `pair_audio`, even when the server had them on file + for that exact video (#398). The rate is read from the container + without decoding anything; the shots come from `shots_beside`, which + only looks at a real local path, so a URL carries none. A file that + says neither stays a plain list. """ + from .runs import shots_beside from .tasks.video_utils import FrameList, file_fps if not isinstance(frames, list): return frames fps = file_fps(location) - return FrameList(frames, fps) if fps else frames + shots = ( + shots_beside(location) + if not (location.startswith("http://") or location.startswith("https://")) + else None + ) + return FrameList(frames, fps, shots) if (fps or shots) else frames def fetch_video(video_spec, base_dir=None): diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index bdf4a4d1..fd55326b 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -411,7 +411,8 @@ def _to_pil(frame): class FrameList(list): - """The frames of a video file, carrying the rate the file plays at. + """The frames of a video file, carrying the rate the file plays at and + the shot boundaries its run recorded, if any. `load_video` answers a plain list of images, which is what every pipeline argument and every task wants - and which says nothing about @@ -422,11 +423,18 @@ class FrameList(list): every consumer working unchanged while `getattr(video, "fps", None)` - the question AudioVideo, concat_videos and interpolate_frames already ask - gets a real answer. + + `shots` is the same idea for the boundaries `dw.runs.shots_beside` + finds beside the file: a video loaded from an `asset:`/`output:` path + carried no way to answer `getattr(video, "shots", None)`, so + `pair_audio` had nothing to remeasure even though the file's own + manifest (or its kept-asset sidecar) already held them (#398). """ - def __init__(self, frames, fps=None): + def __init__(self, frames, fps=None, shots=None): super().__init__(frames) self.fps = fps + self.shots = shots def file_fps(path): diff --git a/tests/test_video_utils.py b/tests/test_video_utils.py index 456049fe..13b06236 100644 --- a/tests/test_video_utils.py +++ b/tests/test_video_utils.py @@ -361,6 +361,100 @@ def test_the_rate_reaches_the_paired_video(self, tmp_path): assert paired.fps == 24 + def test_a_loaded_video_argument_carries_the_run_s_shots(self, tmp_path): + """A video loaded by path (an asset:/output: reference, already + resolved to a local file by the time fetch_video sees it) carries + the shots its run's manifest recorded, the way it already carries + the file's fps - #398.""" + import json + + from dw.arguments import fetch_video + from dw.runs import MANIFEST_FILE_NAME + from dw.tasks.video_utils import FrameList + + run_dir = tmp_path / "ep42" / "20260101-000000-abcdef01" + run_dir.mkdir(parents=True) + path = self.write_video(run_dir / "ep42-film.mp4", fps=24, num_frames=24) + shots = [ + { + "name": "shot@accuse", + "start_frame": 0, + "num_frames": 12, + "start_sample": None, + "num_samples": None, + }, + { + "name": "shot@deflect", + "start_frame": 12, + "num_frames": 12, + "start_sample": None, + "num_samples": None, + }, + ] + manifest = { + "steps": [ + {"step": "concat_videos", "files": ["ep42-film.mp4"], "shots": shots} + ] + } + (run_dir / MANIFEST_FILE_NAME).write_text(json.dumps(manifest)) + + frames = fetch_video(path) + + assert isinstance(frames, FrameList) + assert [shot["name"] for shot in frames.shots] == [ + "shot@accuse", + "shot@deflect", + ] + + def test_pair_audio_remeasures_the_shots_a_loaded_video_carries(self, tmp_path): + """The other half of #398: pair_audio's own remeasuring, fed a + video loaded from a path rather than built by an earlier step in + the same workflow.""" + import json + + from dw.arguments import fetch_video + from dw.runs import MANIFEST_FILE_NAME + from dw.tasks.pair_audio import pair_audio + + run_dir = tmp_path / "ep42" / "20260101-000000-abcdef01" + run_dir.mkdir(parents=True) + path = self.write_video(run_dir / "ep42-film.mp4", fps=24, num_frames=24) + shots = [ + { + "name": "shot@accuse", + "start_frame": 0, + "num_frames": 12, + "start_sample": None, + "num_samples": None, + }, + { + "name": "shot@deflect", + "start_frame": 12, + "num_frames": 12, + "start_sample": None, + "num_samples": None, + }, + ] + manifest = { + "steps": [ + {"step": "concat_videos", "files": ["ep42-film.mp4"], "shots": shots} + ] + } + (run_dir / MANIFEST_FILE_NAME).write_text(json.dumps(manifest)) + + paired = pair_audio( + video=fetch_video(path), + audio=numpy.zeros((2, 16000), dtype=numpy.float32), + sample_rate=16000, + ) + + assert [shot["name"] for shot in paired.shots] == [ + "shot@accuse", + "shot@deflect", + ] + assert paired.shots[0]["start_sample"] == 0 + assert paired.shots[1]["start_sample"] == round(12 / 24 * 16000) + def test_audio_is_fitted_to_the_frames_own_duration(self, tmp_path): """The codec pads the last block; joined shot after shot that padding would walk the sound off the picture.""" From 43502d94844649037d7a70c5c0b5a54c19475a70 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 02:29:16 -0500 Subject: [PATCH 094/181] fix(engine): #399 - nest an input's own shots when concat_videos/dissolve_videos join it An input that is itself an earlier join's output carried its own media.shots, which both joins collapsed to one record per file - so analyze_seams on a join-of-joins silently skipped the inner seams. concat_videos and dissolve_videos now flatten an input's inner shots into the output, offset onto where the whole input landed (dw/shots.py's new nested_shots), instead of falling back to shot_record for it. Frames are always exact; concat_videos additionally clips/drops an inner shot straddling its head trim and clears its sample fields there (trimmed_shots), since the crossfade drawn from trimmed material makes their position in the joined track unmeasurable. measured_num_samples is shared by both joins so a partially-nested shots list (some entries with a known start_sample, some without) is filled in the same way. load_audio_video now attaches a locally-loaded file's own recorded shots (shots_beside) so a join-of-joins read back from disk sees them too. An input without shots stays a single record, as before. Co-Authored-By: Claude Sonnet 5 --- dw/shots.py | 84 ++++++++++++++++++++++++++++++++ dw/tasks/concat_videos.py | 32 ++++++++----- dw/tasks/dissolve_videos.py | 70 ++++++++++++++++++--------- dw/tasks/video_utils.py | 20 +++++--- tests/test_shots.py | 95 +++++++++++++++++++++++++++++++++++++ 5 files changed, 260 insertions(+), 41 deletions(-) diff --git a/dw/shots.py b/dw/shots.py index 8bc776c4..59a63501 100644 --- a/dw/shots.py +++ b/dw/shots.py @@ -117,6 +117,90 @@ def remeasured_shots(shots, fps, sample_rate, total_samples): ] +def trimmed_shots(shots, head_trim): + """The shots of a video after dropping `head_trim` frames off its start. + + concat_videos trims the head of every video after the first before + joining it. A shot entirely inside the trim never reaches the joined + picture and is dropped; one straddling the cut survives, clipped to what + is left and re-based to start at 0, so a later frame offset places it + correctly. The crossfade drawn from the trimmed material makes the + surviving samples' position in the joined track unmeasurable, so the + sample side is cleared regardless of rate. + """ + if not head_trim: + return shots + clipped = [] + for shot in shots: + end = shot["start_frame"] + shot["num_frames"] + if end <= head_trim: + continue + start = max(shot["start_frame"], head_trim) + clipped.append( + { + **shot, + "start_frame": start - head_trim, + "num_frames": end - start, + "start_sample": None, + "num_samples": None, + } + ) + return clipped + + +def nested_shots(shots, frame_offset, sample_offset, native_rate, target_rate): + """An input's own shots, offset onto where the whole input landed in a join. + + Frames are exact: a join only ever adds frames before an input, never + inside it, so `start_frame + frame_offset` is where each inner shot now + sits. Samples are only ever offset when the join measured where the + input's own track landed (`sample_offset`) and both rates are known - + resampling a partial waveform inside the crossfaded region is not a + measurement, so trimmed_shots already clears those before this runs. + Otherwise the sample side is cleared, same as without_samples. + """ + rescale = ( + target_rate / native_rate + if sample_offset is not None and native_rate and target_rate + else None + ) + offset = [] + for shot in shots: + entry = {**shot, "start_frame": shot["start_frame"] + frame_offset} + start_sample = shot.get("start_sample") + if rescale is not None and start_sample is not None: + entry["start_sample"] = sample_offset + round(start_sample * rescale) + else: + entry["start_sample"] = None + entry["num_samples"] = None + offset.append(entry) + return offset + + +def measured_num_samples(shots, total_samples): + """Fill each shot's `num_samples` from where the next measured one starts. + + A shot's track runs up to wherever the next shot with a known + `start_sample` begins, or to the end of the joined track for the last + one - shared by concat_videos and dissolve_videos so nesting an input's + shots (which can leave some entries with no `start_sample`) is handled + the same way in both. + """ + for index, shot in enumerate(shots): + if total_samples is None or shot["start_sample"] is None: + shot["num_samples"] = None + if total_samples is None: + shot["start_sample"] = None + continue + end = total_samples + for following in shots[index + 1 :]: + if following["start_sample"] is not None: + end = following["start_sample"] + break + shot["num_samples"] = end - shot["start_sample"] + return shots + + def shot_reference_names(references): """A name per entry of a step's list naming a shot, else None. diff --git a/dw/tasks/concat_videos.py b/dw/tasks/concat_videos.py index a433fc2c..17b22c9b 100644 --- a/dw/tasks/concat_videos.py +++ b/dw/tasks/concat_videos.py @@ -13,7 +13,7 @@ from ..events import emit_warning from ..result import AudioVideo -from ..shots import shot_record +from ..shots import measured_num_samples, nested_shots, shot_record, trimmed_shots from .audio_utils import ( as_channels_samples, bleed_join, @@ -202,11 +202,23 @@ def concat_videos( start_frame = len(frames) start_sample = audio.shape[1] if audio is not None else 0 frames.extend(clip[head_trim:]) - shots.append( - shot_record( - names[index], start_frame, len(frames) - start_frame, start_sample + inner = getattr(video, "shots", None) + if inner: + shots.extend( + nested_shots( + trimmed_shots(inner, head_trim), + start_frame, + start_sample if waveforms[index] is not None else None, + getattr(video, "sample_rate", None), + sample_rate, + ) + ) + else: + shots.append( + shot_record( + names[index], start_frame, len(frames) - start_frame, start_sample + ) ) - ) if waveforms[index] is None: continue @@ -246,13 +258,9 @@ def concat_videos( ) audio_native_rate = video.sample_rate - for shot, following in zip(shots, shots[1:] + [None]): - # A seam's crossfade leaves the samples before it where they were, so - # a shot's track is everything up to where the next one's began - end = following["start_sample"] if following else _length(audio) - shot["num_samples"] = None if audio is None else end - shot["start_sample"] - if audio is None: - shot["start_sample"] = None + # A seam's crossfade leaves the samples before it where they were, so a + # shot's track is everything up to where the next measured one began + measured_num_samples(shots, _length(audio) if audio is not None else None) logger.debug(f"Concatenated {len(videos)} videos into {len(frames)} frames") # The rate the caller declared, else the rate the first input carries - diff --git a/dw/tasks/dissolve_videos.py b/dw/tasks/dissolve_videos.py index e387e2e2..afa74ead 100644 --- a/dw/tasks/dissolve_videos.py +++ b/dw/tasks/dissolve_videos.py @@ -19,7 +19,7 @@ from ..events import emit_warning from ..result import AudioVideo -from ..shots import shot_record +from ..shots import measured_num_samples, nested_shots, shot_record from .audio_utils import ( as_channels_samples, crossfade_concat, @@ -131,11 +131,13 @@ def dissolve_videos( sample_starts, ) shots = _dissolve_shots( + loaded, video_names(videos), frame_starts, len(frames), sample_starts, audio, + sample_rate, dissolve_frames, ) logger.info( @@ -150,34 +152,58 @@ def dissolve_videos( def _dissolve_shots( - names, frame_starts, total_frames, sample_starts, audio, dissolve_frames + videos, + names, + frame_starts, + total_frames, + sample_starts, + audio, + sample_rate, + dissolve_frames, ): - """One shot per video, partitioning the dissolved picture and track. + """One shot per video - or, for one that already carries its own, one per + inner shot - partitioning the dissolved picture and track. - A dissolve belongs to the shot coming in: each shot runs from where its - dissolve opens to where the next one's does, so the counts add up to the - file's and `overlap_frames` says how much of its head is blended. + A dissolve belongs to the shot coming in: each video's own frames map + onto the joined picture by the exact offset frame_starts[index] gives (a + dissolve blends in place rather than dropping frames, unlike + concat_videos' head trim), so an input that is itself an earlier join's + output keeps its inner seams rather than collapsing to one record + (#399). `overlap_frames` marks how much of the shot's own head - the + first inner one, when it nests - is blended with what came before. """ - frame_ends = frame_starts[1:] + [total_frames] - if audio is None: - sample_starts = [None] * len(frame_starts) - sample_ends = sample_starts - else: - sample_ends = sample_starts[1:] + [audio.shape[1]] + total_samples = audio.shape[1] if audio is not None else None shots = [] - for index, name in enumerate(names): - start_sample = sample_starts[index] - shots.append( - shot_record( + for index, (video, name) in enumerate(zip(videos, names)): + inner = getattr(video, "shots", None) + sample_offset = sample_starts[index] if audio is not None else None + if inner: + nested = nested_shots( + inner, + frame_starts[index], + sample_offset, + getattr(video, "sample_rate", None), + sample_rate, + ) + if index and dissolve_frames and nested: + nested[0]["overlap_frames"] = dissolve_frames + shots.extend(nested) + else: + frame_end = ( + frame_starts[index + 1] + if index + 1 < len(frame_starts) + else total_frames + ) + shot = shot_record( name, frame_starts[index], - frame_ends[index] - frame_starts[index], - start_sample, - None if start_sample is None else sample_ends[index] - start_sample, + frame_end - frame_starts[index], + sample_offset, ) - ) - if index and dissolve_frames: - shots[-1]["overlap_frames"] = dissolve_frames + if index and dissolve_frames: + shot["overlap_frames"] = dissolve_frames + shots.append(shot) + measured_num_samples(shots, total_samples) return shots diff --git a/dw/tasks/video_utils.py b/dw/tasks/video_utils.py index fd55326b..38c59f66 100644 --- a/dw/tasks/video_utils.py +++ b/dw/tasks/video_utils.py @@ -470,7 +470,10 @@ def load_audio_video(location, base_dir=None): Returns: An AudioVideo holding the frames as PIL images and, when the file carries an audio stream, its waveform as a (channels, samples) float32 - array with the stream's sample rate + array with the stream's sample rate. A local file also carries the + shots its own run manifest recorded for it (`shots_beside`), so a + join of a file that is itself an earlier join's output can see the + seams inside it (#399); a URL carries none. """ from ..security import ALLOWED_VIDEO_EXTENSIONS, validate_file_extension from ..locations import validate_media_path, validate_media_url @@ -488,13 +491,16 @@ def load_audio_video(location, base_dir=None): response = requests.get(validated_url, timeout=300) response.raise_for_status() handle = io.BytesIO(response.content) - else: - validated_path = validate_media_path(location, base_dir, "a video argument") - validate_file_extension(validated_path, ALLOWED_VIDEO_EXTENSIONS) - logger.debug(f"Reading video from {validated_path}") - handle = validated_path + return _decode_audio_video(handle) - return _decode_audio_video(handle) + validated_path = validate_media_path(location, base_dir, "a video argument") + validate_file_extension(validated_path, ALLOWED_VIDEO_EXTENSIONS) + logger.debug(f"Reading video from {validated_path}") + video = _decode_audio_video(validated_path) + from ..runs import shots_beside + + video.shots = shots_beside(validated_path) + return video def _decode_audio_video(handle): diff --git a/tests/test_shots.py b/tests/test_shots.py index 47cc3eaf..0097546d 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -294,6 +294,64 @@ def test_an_overrun_track_is_measured_not_derived(self): 4 / fps * sample_rate ) + def test_an_inner_input_nests_its_own_shots(self): + """#399: a video that is itself an earlier join's output carries its + own `.shots` - concat_videos flattens those into the joined output, + offset onto where the whole input landed, rather than collapsing + them to one record for the whole file.""" + fps, sample_rate = 4, 100 + inner_shots = [ + shot_record("shot@accuse", 0, 4, start_sample=0, num_samples=100), + shot_record("shot@deflect", 4, 4, start_sample=100, num_samples=100), + ] + nested = AudioVideo( + frames(8), + numpy.full((2, 200), 1.0, dtype=numpy.float32), + sample_rate, + fps=fps, + shots=inner_shots, + ) + trailing = audio_video(4, 2, fps=fps, sample_rate=sample_rate) + + result = concat_videos([nested, trailing], fps=fps) + + names = [shot["name"] for shot in result.shots] + assert names == ["shot@accuse", "shot@deflect", "video 2"] + starts = [shot["start_frame"] for shot in result.shots] + assert starts == [0, 4, 8] + sample_starts = [shot["start_sample"] for shot in result.shots] + assert sample_starts[:2] == [0, 100] + + def test_a_trimmed_inner_input_clips_and_clears_its_shots(self): + """The same nested input, but trimmed as the second video - the + trim can cut into or through an inner shot; one entirely inside it + is dropped, one straddling it survives clipped and re-based, and + either way its sample fields are cleared because the crossfade + draws from the trimmed material (dw/shots.py's trimmed_shots).""" + fps, sample_rate = 4, 100 + inner_shots = [ + shot_record("shot@a", 0, 2, start_sample=0, num_samples=50), + shot_record("shot@b", 2, 6, start_sample=50, num_samples=150), + ] + nested = AudioVideo( + frames(8), + numpy.full((2, 200), 1.0, dtype=numpy.float32), + sample_rate, + fps=fps, + shots=inner_shots, + ) + leading = audio_video(4, 1, fps=fps, sample_rate=sample_rate) + + result = concat_videos([leading, nested], trim_frames=3, fps=fps) + + names = [shot["name"] for shot in result.shots] + assert names == ["video 1", "shot@b"] + trimmed = result.shots[1] + assert trimmed["start_frame"] == 4 + assert trimmed["num_frames"] == 5 + assert trimmed["start_sample"] is None + assert trimmed["num_samples"] is None + # --------------------------------------------------------------------------- # 4. dissolve_videos @@ -346,6 +404,43 @@ def test_asset_literal_shots_are_named_by_file_not_absolute_path(self): assert result.shots[0]["name"] == "ep3-shot2-reply.mp4" assert "/" not in result.shots[0]["name"] + def test_an_inner_input_nests_its_own_shots(self): + """#399: same fix as concat_videos - an input already carrying + `.shots` from an earlier join keeps its inner seams instead of + collapsing to a single record for the whole file. A dissolve maps a + clip's own frame j to joined frame frame_starts[index] + j exactly + (no head trim), so the offset is a straight add with no clipping.""" + fps, sample_rate = 4, 100 + inner_shots = [ + shot_record("shot@accuse", 0, 4, start_sample=0, num_samples=100), + shot_record("shot@deflect", 4, 4, start_sample=100, num_samples=100), + ] + nested = AudioVideo( + frames(8), + numpy.full((2, 200), 1.0, dtype=numpy.float32), + sample_rate, + fps=fps, + shots=inner_shots, + ) + leading = AudioVideo( + frames(10), + numpy.full((2, 250), 2.0, dtype=numpy.float32), + sample_rate, + fps=fps, + ) + + result = dissolve_videos([leading, nested], 3, fps=fps) + + names = [shot["name"] for shot in result.shots] + assert names == ["video 1", "shot@accuse", "shot@deflect"] + frame_offset = 10 - 3 # frame_starts[1] + starts = [shot["start_frame"] for shot in result.shots] + assert starts[1:] == [frame_offset, frame_offset + 4] + # overlap_frames marks the dissolve's head - the first nested + # sub-shot only, since the seam is between videos, not inside one + assert result.shots[1]["overlap_frames"] == 3 + assert "overlap_frames" not in result.shots[2] + # --------------------------------------------------------------------------- # 5. run_chain From 0fd6043149da3288044a94b109e98a719899fafb Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 02:56:09 -0500 Subject: [PATCH 095/181] fix(engine): #400 - refuse a dissolve_videos overlap wider than a resolvable input's frame count at validate time dissolve_videos only discovers a too-short input after decoding every video in the step, so a chain that generates each shot before joining them could spend many GPU minutes reaching a step that was always going to fail. dissolve_frame_errors (dw/dissolve_frame_errors.py) mirrors that run-time check in validation_errors for the cases the frame count is already knowable - a literal path, or an asset:/output: reference, paired with a literal dissolve_frames - and leaves a previous_result:/variable:/item:/gather: reference or non-literal dissolve_frames to the existing run-time check, since the length isn't known yet. concat_videos' trim_frames and crossfade_audio's crossfade window were each checked for the same gap the issue asked about; neither raises on a too-short input (they truncate/clamp silently instead), so neither has an equivalent bug shape to move earlier. Co-Authored-By: Claude Sonnet 5 --- dw/dissolve_frame_errors.py | 150 ++++++++++++++++++++++++++++ dw/workflow.py | 8 ++ tests/test_dissolve_frame_errors.py | 146 +++++++++++++++++++++++++++ 3 files changed, 304 insertions(+) create mode 100644 dw/dissolve_frame_errors.py create mode 100644 tests/test_dissolve_frame_errors.py diff --git a/dw/dissolve_frame_errors.py b/dw/dissolve_frame_errors.py new file mode 100644 index 00000000..a9969725 --- /dev/null +++ b/dw/dissolve_frame_errors.py @@ -0,0 +1,150 @@ +"""A `dissolve_videos` overlap too wide for one of its own inputs, refused +before the run when the frame counts are already knowable. + +`dissolve_videos` (`dw/tasks/dissolve_videos.py`) raises once it has decoded +every input: a video with fewer frames than its share of `dissolve_frames` +overlaps (`seams * dissolve_frames`) fails with "video N has M frames, too +few for its S dissolve(s) of F frames". That is correct, but late - a chain +that generates each shot before joining them can spend many GPU minutes +reaching a step that was always going to fail, for an arithmetic mistake +visible from the workflow document alone (#400). + +Moved here, into `validation_errors`, for exactly the cases where a video's +frame count is knowable without running anything: a literal file path, or an +`asset:`/`output:` reference, with a literal `dissolve_frames`. +`resolve_path_references` is what turns either into a real path before the +run reads it; `probe_media` decodes that file the same way `dw/server/app.py` +already does for gallery metadata. A `previous_result:` (or any reference +`expand_for_each` left unresolved), a `variable:`/`item:`/`gather:` reference, +or a non-literal `dissolve_frames`, names no frame count yet and is left to +the existing run-time check - silence there is correct, not a gap, since the +length is not known until the step that produces it runs. + +`concat_videos`'s `trim_frames` and `crossfade_audio`'s crossfade window were +each considered for the same treatment - the issue that motivated this module +asked whether they "probably have the same gap". They do not: neither raises +when an input is too short. `concat_videos` silently truncates +(`frames.extend(clip[head_trim:])`), and `crossfade_audio` silently clamps +its window to the shortest side (`crossfade_concat`) - a different, and +already silent, shape of problem with no run-time error to move earlier. +""" + +import os + +from .arguments import resolve_path_references +from .assets import is_asset_reference +from .for_each import MEMBER_SEPARATOR, render_path +from .media_info import probe_media +from .runs import is_output_reference + +# Left to the run-time check: not yet resolved to a real file at the point +# validation walks the expanded definition. +_UNRESOLVED_PREFIXES = ("previous_result:", "variable:", "item:", "gather:") + + +def _resolve_video_path(value, base_dir): + """The local file `value` names, or None when it is not yet resolvable, + is not a local file, or does not exist - any of which defers the check + to the run, exactly as `dissolve_videos` itself would then load it.""" + if not isinstance(value, str): + return None + if value.startswith(_UNRESOLVED_PREFIXES): + return None + if value.startswith(("http://", "https://")): + return None + if is_asset_reference(value) or is_output_reference(value): + try: + value = resolve_path_references(value, base_dir) + except Exception: + # Existence/traversal problems belong to reference_name_errors + # and reference resolution at run time, not to this check + return None + if not isinstance(value, str): + return None + return value if os.path.isfile(value) else None + + +def _frame_count(path): + """The frame count `dissolve_videos` would see for this file, or None + when it cannot be probed or carries no video stream.""" + info = probe_media(path) + if info is None or info.get("kind") != "video": + return None + return info.get("frame_count") + + +def dissolve_frame_errors(workflow_definition, source_indices=None, base_dir=None): + """Every `dissolve_videos` step whose overlap already exceeds a + statically-resolvable input's real frame count, as [{path, message}]. + + Walks the substituted, expanded definition, the same convention + `video_extension_errors` and `task_argument_errors` follow: + `source_indices` maps an expanded step back to the one the author wrote, + and a path inside a `for_each` member names the member. + """ + steps = workflow_definition.get("steps") + if not isinstance(steps, list): + return [] + + errors = [] + for index, step in enumerate(steps): + if not isinstance(step, dict): + continue + task = step.get("task") + if not isinstance(task, dict) or task.get("command") != "dissolve_videos": + continue + arguments = task.get("arguments") + if not isinstance(arguments, dict): + continue + videos = arguments.get("videos") + if not isinstance(videos, list) or len(videos) < 2: + continue + dissolve_frames = arguments.get("dissolve_frames", 12) + if not isinstance(dissolve_frames, (int, float)) or isinstance( + dissolve_frames, bool + ): + continue + if dissolve_frames <= 0: + continue + + problems = [] + for video_index, video in enumerate(videos): + path = _resolve_video_path(video, base_dir) + if path is None: + continue + frame_count = _frame_count(path) + if frame_count is None: + continue + seams = (video_index > 0) + (video_index < len(videos) - 1) + needed = seams * dissolve_frames + if frame_count < needed: + problems.append( + f"video {video_index} has {frame_count} frames, too few " + f"for its {seams} dissolve(s) of {dissolve_frames} frames" + ) + if not problems: + continue + + source = ( + source_indices[index] + if source_indices is not None and index < len(source_indices) + else index + ) + name = step.get("name") + where = ( + f" in member '{name}'" + if isinstance(name, str) and MEMBER_SEPARATOR in name + else "" + ) + errors.append( + { + "path": render_path( + ("steps", source, "task", "arguments", "dissolve_frames") + ), + "message": f"dissolve_videos: {'; '.join(problems)}{where}", + } + ) + return errors + + +__all__ = ["dissolve_frame_errors"] diff --git a/dw/workflow.py b/dw/workflow.py index 1b72e092..513f7bde 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -33,6 +33,7 @@ from .adapter_compatibility import adapter_errors, warn_adapters from .elision import elide_definition, warn_elided from .introspection import task_signature_errors, component_type_errors +from .dissolve_frame_errors import dissolve_frame_errors from .task_domains import task_argument_errors from .select_validation import select_errors from .variable_constraints import ( @@ -728,6 +729,13 @@ def validation_errors(self, arguments=None, composing=None): # sample rate a silent fallback to 44100 (dw/task_domains.py, # #139, #140) + task_argument_errors(expanded, source_indices) + # A dissolve_videos overlap wider than a statically-resolvable + # input's real frame count decoded clean past the queue and + # failed only after every upstream step had already generated - + # refused here for a literal dissolve_frames against an asset:/ + # output:/literal-path video, the cases the frame count is + # already knowable (dw/dissolve_frame_errors.py, #400) + + dissolve_frame_errors(expanded, source_indices, base_dir) # A select step whose rule is misspelled, or whose # threshold/index does not match its rule, validated clean and # died on select's own run-time ValueError after the fan-out diff --git a/tests/test_dissolve_frame_errors.py b/tests/test_dissolve_frame_errors.py new file mode 100644 index 00000000..d57cdf17 --- /dev/null +++ b/tests/test_dissolve_frame_errors.py @@ -0,0 +1,146 @@ +"""A dissolve_videos overlap wider than a statically-resolvable input's real +frame count, refused before the run - #400. + +`dissolve_videos` itself only discovers a too-short input after decoding +every video in the step; these tests exercise the real decode path +(`probe_media` against a genuine mp4) rather than a mock of it, so a fixture +short of its declared dissolve_frames is the same file the run itself would +have failed on. +""" + +import os +import tempfile + +import numpy + +from dw.dissolve_frame_errors import dissolve_frame_errors +from dw.runs import activate_output_root, deactivate_output_root +from dw.workflow import workflow_from_definition + + +def write_mp4(path, frames=12, fps=6, width=32, height=16): + import av + + container = av.open(str(path), "w") + video = container.add_stream("libx264", rate=fps) + video.width, video.height, video.pix_fmt = width, height, "yuv420p" + for _ in range(frames): + frame = av.VideoFrame.from_ndarray( + numpy.zeros((height, width, 3), numpy.uint8), format="rgb24" + ) + for packet in video.encode(frame): + container.mux(packet) + for packet in video.encode(): + container.mux(packet) + container.close() + + +def workflow_dir_with_asset(monkeypatch, *names_and_frames): + """A base_dir whose assets/ subfolder holds the given fixtures, pinned + via DW_ASSET_DIR so it is found ahead of this checkout's own assets/ + (discover_library's './assets' in the working directory would otherwise + shadow it).""" + base_dir = tempfile.mkdtemp() + asset_dir = os.path.join(base_dir, "assets") + os.makedirs(asset_dir) + for name, frames in names_and_frames: + write_mp4(os.path.join(asset_dir, name), frames=frames) + monkeypatch.setenv("DW_ASSET_DIR", asset_dir) + return base_dir + + +def dissolve_workflow(videos, dissolve_frames=12): + return { + "id": "dissolving", + "steps": [ + { + "name": "join", + "task": { + "command": "dissolve_videos", + "arguments": {"videos": videos, "dissolve_frames": dissolve_frames}, + }, + "result": {"content_type": "video/mp4", "fps": 6}, + } + ], + } + + +class TestTheCheck: + def test_an_asset_too_short_for_its_dissolve_is_refused(self, monkeypatch): + base_dir = workflow_dir_with_asset(monkeypatch, ("a.mp4", 124), ("b.mp4", 124)) + definition = dissolve_workflow( + ["asset:a.mp4", "asset:b.mp4"], dissolve_frames=130 + ) + + problems = dissolve_frame_errors(definition, base_dir=base_dir) + + assert len(problems) == 1 + assert problems[0]["path"] == "steps[0].task.arguments.dissolve_frames" + assert "124 frames" in problems[0]["message"] + assert "130 frames" in problems[0]["message"] + + def test_enough_frames_validates_clean(self, monkeypatch): + base_dir = workflow_dir_with_asset(monkeypatch, ("a.mp4", 124), ("b.mp4", 124)) + definition = dissolve_workflow( + ["asset:a.mp4", "asset:b.mp4"], dissolve_frames=12 + ) + + assert dissolve_frame_errors(definition, base_dir=base_dir) == [] + + def test_a_previous_result_video_is_left_to_the_run(self): + definition = dissolve_workflow( + ["previous_result:make_a", "previous_result:make_b"], + dissolve_frames=130, + ) + + assert dissolve_frame_errors(definition) == [] + + def test_an_output_reference_too_short_for_its_dissolve_is_refused(self): + output_root = tempfile.mkdtemp() + run_dir = os.path.join(output_root, "clip", "20260101-000000-abc") + os.makedirs(run_dir) + write_mp4(os.path.join(run_dir, "a.mp4"), frames=124) + write_mp4(os.path.join(run_dir, "b.mp4"), frames=124) + definition = dissolve_workflow( + [ + "output:clip/20260101-000000-abc/a.mp4", + "output:clip/20260101-000000-abc/b.mp4", + ], + dissolve_frames=130, + ) + + token = activate_output_root(output_root) + try: + problems = dissolve_frame_errors(definition) + finally: + deactivate_output_root(token) + + assert len(problems) == 1 + assert "124 frames" in problems[0]["message"] + + def test_nothing_is_reported_for_a_definition_with_no_dissolve_step(self): + assert dissolve_frame_errors({"steps": [{"name": "a", "task": {}}]}) == [] + + def test_a_single_video_needs_no_dissolve(self, monkeypatch): + base_dir = workflow_dir_with_asset(monkeypatch, ("a.mp4", 5)) + definition = dissolve_workflow(["asset:a.mp4"], dissolve_frames=130) + + assert dissolve_frame_errors(definition, base_dir=base_dir) == [] + + +class TestTheValidationPass: + def test_wired_into_validation_errors(self, monkeypatch): + base_dir = workflow_dir_with_asset(monkeypatch, ("a.mp4", 124), ("b.mp4", 124)) + definition = dissolve_workflow( + ["asset:a.mp4", "asset:b.mp4"], dissolve_frames=130 + ) + workflow = workflow_from_definition( + definition, os.path.join(base_dir, "workflow.json") + ) + + problems = workflow.validation_errors() + + assert any( + problem["path"] == "steps[0].task.arguments.dissolve_frames" + for problem in problems + ) From 0a73ce3dd03aecf54c64bb3bdb2ee534e73a1c96 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 03:42:53 -0500 Subject: [PATCH 096/181] fix(engine): #401 - round dissolve_videos' crossfade window like every other frame->sample conversion crossfade_concat computed its overlap window with plain int() truncation while frames_to_samples and remeasured_shots (pair_audio's own conversion) round - so a dissolve's recorded shot start_sample could land one sample below what pair_audio would recompute for the same frame boundary after re-pairing the video with a new track. Rounding the window brings dissolve_videos onto the same conversion rule as the rest of the shot-record contract. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/audio_utils.py | 2 +- tests/test_shots.py | 50 ++++++++++++++++++++++++++++++++++++++++- 2 files changed, 50 insertions(+), 2 deletions(-) diff --git a/dw/tasks/audio_utils.py b/dw/tasks/audio_utils.py index c42d2c2f..e712ac93 100644 --- a/dw/tasks/audio_utils.py +++ b/dw/tasks/audio_utils.py @@ -316,7 +316,7 @@ def crossfade_concat(waveforms, sample_rate, crossfade_ms, starts=None): for following in waveforms[1:]: result, following = _matched_channels(result, following) window = min( - int(crossfade_ms / 1000.0 * sample_rate), + int(round(crossfade_ms / 1000.0 * sample_rate)), result.shape[1], following.shape[1], ) diff --git a/tests/test_shots.py b/tests/test_shots.py index 0097546d..e590f1ef 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -36,7 +36,7 @@ shots_for_file, step_shots, ) -from dw.tasks.audio_utils import slice_audio +from dw.tasks.audio_utils import frames_to_samples, slice_audio from dw.tasks.concat_videos import concat_videos from dw.tasks.dissolve_videos import dissolve_videos from dw.tasks.pair_audio import pair_audio @@ -441,6 +441,54 @@ def test_an_inner_input_nests_its_own_shots(self): assert result.shots[1]["overlap_frames"] == 3 assert "overlap_frames" not in result.shots[2] + def test_seam_sample_start_matches_the_frame_to_sample_conversion(self): + """#401: dissolve_videos' crossfade window used to floor its + ms->sample conversion while every other tool that places a frame on + a track (frames_to_samples, remeasured_shots) rounds - so a shot's + start_sample recorded here could land one sample below what + pair_audio would recompute for the same frame boundary. The seam's + recorded start_sample must agree with frames_to_samples for the same + frame offset, fps and rate.""" + fps, sample_rate, dissolve_frames = 24, 32000, 8 + + def clip(num_frames, level): + samples = frames_to_samples(num_frames, fps, sample_rate) + audio = numpy.full((2, samples), float(level), dtype=numpy.float32) + return AudioVideo(frames(num_frames), audio, sample_rate, fps=fps) + + result = dissolve_videos( + [clip(20, 1), clip(20, 2)], dissolve_frames, fps=fps + ) + + frame_starts = [shot["start_frame"] for shot in result.shots] + assert frame_starts[1] == 12 # 20 - dissolve_frames + + expected_sample_start = frames_to_samples(frame_starts[1], fps, sample_rate) + assert result.shots[1]["start_sample"] == expected_sample_start + + def test_shots_survive_a_pair_audio_round_trip_unchanged(self): + """#401: a shot passed through pair_audio unchanged in frames must + come out unchanged in samples - the invariant the two tools' sample + math is required to agree on.""" + fps, sample_rate, dissolve_frames = 24, 32000, 8 + + def clip(num_frames, level): + samples = frames_to_samples(num_frames, fps, sample_rate) + audio = numpy.full((2, samples), float(level), dtype=numpy.float32) + return AudioVideo(frames(num_frames), audio, sample_rate, fps=fps) + + joined = dissolve_videos([clip(20, 1), clip(20, 2)], dissolve_frames, fps=fps) + + new_track = numpy.zeros((2, joined.audio.shape[1]), dtype=numpy.float32) + paired = pair_audio(joined, new_track, sample_rate=sample_rate, fps=fps) + + assert [shot["start_frame"] for shot in paired.shots] == [ + shot["start_frame"] for shot in joined.shots + ] + assert [shot["start_sample"] for shot in paired.shots] == [ + shot["start_sample"] for shot in joined.shots + ] + # --------------------------------------------------------------------------- # 5. run_chain From 3f1c84df475baddba2710089396302e52909ebf9 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 03:44:05 -0500 Subject: [PATCH 097/181] style(tests): #401 - ruff format test_shots.py Co-Authored-By: Claude Sonnet 5 --- tests/test_shots.py | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/tests/test_shots.py b/tests/test_shots.py index e590f1ef..02d254d8 100644 --- a/tests/test_shots.py +++ b/tests/test_shots.py @@ -456,9 +456,7 @@ def clip(num_frames, level): audio = numpy.full((2, samples), float(level), dtype=numpy.float32) return AudioVideo(frames(num_frames), audio, sample_rate, fps=fps) - result = dissolve_videos( - [clip(20, 1), clip(20, 2)], dissolve_frames, fps=fps - ) + result = dissolve_videos([clip(20, 1), clip(20, 2)], dissolve_frames, fps=fps) frame_starts = [shot["start_frame"] for shot in result.shots] assert frame_starts[1] == 12 # 20 - dissolve_frames From a22ddf2844d18704774ca92cf9aa71c4db034ba2 Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 04:07:50 -0500 Subject: [PATCH 098/181] fix(engine): #401 - derive dissolve_videos' shot start_sample instead of measuring it Summing the individually-rounded per-clip lengths a real dissolve crossfade measures does not equal rounding the cumulative frame count in one step, so a shot whose frames never changed could disagree with pair_audio's remeasured_shots by a sample on re-pairing (#401, deeper than the crossfade-window rounding already fixed for this issue). _dissolve_shots now derives each shot's start_sample with frames_to_samples - the same rule remeasured_shots uses - rather than reading it off where the crossfade actually landed; the crossfade itself still blends the real, measured audio bytes, only the recorded seam position changed. crossfade_concat's now-unused starts parameter is dropped from the call. concat_videos has a structurally similar drift but deliberately measures rather than derives (#378, to catch a genuine audio overrun), so it is left alone here and noted separately on the issue. Co-Authored-By: Claude Sonnet 5 --- dw/tasks/dissolve_videos.py | 39 +++++++++++++++++++++++-------------- 1 file changed, 24 insertions(+), 15 deletions(-) diff --git a/dw/tasks/dissolve_videos.py b/dw/tasks/dissolve_videos.py index afa74ead..10b6b9e7 100644 --- a/dw/tasks/dissolve_videos.py +++ b/dw/tasks/dissolve_videos.py @@ -23,6 +23,7 @@ from .audio_utils import ( as_channels_samples, crossfade_concat, + frames_to_samples, match_levels as match_track_levels, resample_waveform, warn_on_level_spread, @@ -120,7 +121,10 @@ def dissolve_videos( ) frames = [Image.fromarray(frame) for frame in joined.round().astype(numpy.uint8)] - sample_starts = [] + written_fps = fps or next( + (v.fps for v in loaded if getattr(v, "fps", None)), + None, + ) audio, sample_rate = _dissolve_audio( loaded, dissolve_frames, @@ -128,14 +132,13 @@ def dissolve_videos( match_levels, match_levels_dbfs, sample_rate, - sample_starts, ) shots = _dissolve_shots( loaded, video_names(videos), frame_starts, len(frames), - sample_starts, + written_fps, audio, sample_rate, dissolve_frames, @@ -144,10 +147,6 @@ def dissolve_videos( f"Dissolved {len(clips)} videos into {len(frames)} frames " f"({dissolve_frames}-frame seams)" ) - written_fps = fps or next( - (v.fps for v in loaded if getattr(v, "fps", None)), - None, - ) return AudioVideo(frames, audio, sample_rate, fps=written_fps, shots=shots) @@ -156,7 +155,7 @@ def _dissolve_shots( names, frame_starts, total_frames, - sample_starts, + fps, audio, sample_rate, dissolve_frames, @@ -171,12 +170,26 @@ def _dissolve_shots( output keeps its inner seams rather than collapsing to one record (#399). `overlap_frames` marks how much of the shot's own head - the first inner one, when it nests - is blended with what came before. + + Each shot's `start_sample` is *derived* from its frame offset + (frames_to_samples), the same rule pair_audio's remeasured_shots uses, + rather than read off where the crossfade actually landed: summing the + individually-rounded per-clip lengths a real dissolve measures does not + equal rounding the cumulative frame count in one step, so the two tools + disagreed by a sample on a shot whose frames never changed (#401). The + crossfade itself still blends the real, measured audio - only the + recorded seam position is derived, so it matches whatever a later + pair_audio recomputes for the same boundary. """ total_samples = audio.shape[1] if audio is not None else None shots = [] for index, (video, name) in enumerate(zip(videos, names)): inner = getattr(video, "shots", None) - sample_offset = sample_starts[index] if audio is not None else None + sample_offset = ( + min(frames_to_samples(frame_starts[index], fps, sample_rate), total_samples) + if audio is not None and fps + else None + ) if inner: nested = nested_shots( inner, @@ -236,12 +249,8 @@ def _dissolve_audio( match_levels=None, match_levels_dbfs=None, sample_rate=None, - starts=None, ): - """Crossfade every video's track over the seams' own span. - - `starts` is filled with where each track begins in the joined one. - """ + """Crossfade every video's track over the seams' own span.""" tracks = [v for v in videos if isinstance(v, AudioVideo) and v.audio is not None] if len(tracks) != len(videos): if tracks: @@ -294,6 +303,6 @@ def _dissolve_audio( else: warn_on_level_spread(waveforms, "dissolve_videos") return ( - crossfade_concat(waveforms, sample_rate, crossfade_ms, starts), + crossfade_concat(waveforms, sample_rate, crossfade_ms), sample_rate, ) From 49f3749f81bb9bcb027acb587cec737d3382afdf Mon Sep 17 00:00:00 2001 From: Don Kackman Date: Thu, 24 Sep 2026 04:34:13 -0500 Subject: [PATCH 099/181] fix(server): #402 - warn at validate time when slice_audio's source is known short validate_workflow was silent when a slice_audio step's requested region reaches past a source it can already probe (an asset:/output: reference), even though the run itself warns via _warn_on_slice_past_end once it generates. dw/slice_preflight.py walks the expanded definition, resolves asset:/output: audio references to a real file with resolve_path_references, and probes its duration with probe_media - the same resolution and decode the run would do, just ahead of the queue. Wired into Workflow.slice_past_end_warnings() (mirroring adapter_warnings) and into POST /api/validate's warnings. Deliberately narrower than the run-time check: previous_result:, literal paths, remote URLs and anything probe_media can't read stay silent here, same as the #400 precedent (dissolve_frame_errors) this mirrors. Co-Authored-By: Claude Sonnet 5 --- dw/server/app.py | 7 +- dw/slice_preflight.py | 176 ++++++++++++++++++++++++++++++++++ dw/workflow.py | 22 +++++ tests/test_slice_preflight.py | 148 ++++++++++++++++++++++++++++ 4 files changed, 352 insertions(+), 1 deletion(-) create mode 100644 dw/slice_preflight.py create mode 100644 tests/test_slice_preflight.py diff --git a/dw/server/app.py b/dw/server/app.py index 01ca8044..81e814e0 100644 --- a/dw/server/app.py +++ b/dw/server/app.py @@ -1833,7 +1833,12 @@ def validate_workflow( + candidate.null_variable_argument_warnings(caller_arguments) # An argument a sub-workflow step passes to a workflow that # declares no variable for it - dropped in silence at run time - + candidate.sub_workflow_warnings(), + + candidate.sub_workflow_warnings() + # A slice_audio source whose real duration is already knowable + # (an asset:/output: reference validate can already probe) and + # whose requested slice reaches past it - zero-padded rather than + # refused, but previously said only by the run itself (#402) + + candidate.slice_past_end_warnings(request.arguments), } if request.arguments: # Naming what was checked is the difference between 'the stored diff --git a/dw/slice_preflight.py b/dw/slice_preflight.py new file mode 100644 index 00000000..0182499b --- /dev/null +++ b/dw/slice_preflight.py @@ -0,0 +1,176 @@ +"""A `slice_audio` step whose source duration validate can already learn, +sliced past where that source ends, warned about before the run (#402). + +`slice_audio` zero-pads a slice that reaches past its source and only says so +at run time (`_warn_on_slice_past_end`, `dw/tasks/audio_utils.py`) - a +correct message, but late when the slice sits downstream of a long render +(#402's repro: `assemble-and-score` scoring a 15.5 s cut with a 4.96 s +`score` asset, `validate_workflow` answering clean). Mirrors +`dissolve_frame_errors.py` (#400): walk the expanded definition, +`resolve_path_references` an `asset:`/`output:` audio into a real path, and +`probe_media` it - the same resolution and decode the run itself would do, +just ahead of the queue. + +Deliberately narrower than the run-time check, same as #400's: a +`previous_result:` audio (nothing written yet), a literal path, a remote URL, +or a source `probe_media` cannot read, all answer "unknown" rather than +guessing - silence here is correct, not a gap, since the run-time warning +still fires once the file exists. `variable:` needs no hop of its own: by the +time `validation_errors`/`adapter_warnings` hand this module the *expanded* +definition, `replace_variables` has already substituted every `variable:` +reference (or the run cannot start at all), so what is left unresolved is +only a reference that genuinely cannot resolve yet. The threshold +(`SLICE_PAD_WARN_MS`) and the requested-length arithmetic mirror +`slice_audio`'s own two argument shapes, so the two agree on the same +padding for the same arguments. +""" + +import os + +from .arguments import resolve_path_references +from .assets import is_asset_reference +from .for_each import MEMBER_SEPARATOR, render_path +from .media_info import probe_media +from .runs import is_output_reference +from .tasks.audio_utils import SLICE_PAD_WARN_MS + +# Left to the run-time check: not yet resolved to a real file at the point +# validation walks the expanded definition. +_UNRESOLVED_PREFIXES = ("previous_result:", "variable:", "item:", "gather:") + + +def _resolve_audio_path(value, base_dir): + """The local file `value` names, or None when it is not yet resolvable, + is not a local file, or does not exist - any of which defers the check + to the run, exactly as `slice_audio` itself would then load it.""" + if not isinstance(value, str): + return None + if value.startswith(_UNRESOLVED_PREFIXES): + return None + if value.startswith(("http://", "https://")): + return None + if is_asset_reference(value) or is_output_reference(value): + try: + value = resolve_path_references(value, base_dir) + except Exception: + # Existence/traversal problems belong to reference_name_errors + # and reference resolution at run time, not to this check + return None + if not isinstance(value, str): + return None + return value if os.path.isfile(value) else None + + +def _source_seconds(path): + """The duration `slice_audio` would see for this file, or None when it + cannot be probed or carries no audio.""" + info = probe_media(path) + if info is None or info.get("kind") != "audio": + return None + return info.get("duration_seconds") + + +def _as_number(value): + if isinstance(value, bool) or not isinstance(value, (int, float, str)): + return None + try: + return float(value) + except (TypeError, ValueError): + return None + + +def _requested_region(task_args): + """The (start_seconds, length_seconds) `slice_audio` would compute for + these arguments, or None when the shape given cannot be resolved to a + length without running anything - mirrors `slice_audio`'s own branch + order in `dw/tasks/audio_utils.py`.""" + start_seconds = _as_number(task_args.get("start_seconds")) + duration_seconds = _as_number(task_args.get("duration_seconds")) + start_frame = _as_number(task_args.get("start_frame")) + num_frames = _as_number(task_args.get("num_frames")) + fps = _as_number(task_args.get("fps")) + + if ( + task_args.get("start_seconds") is not None + or task_args.get("duration_seconds") is not None + ): + if duration_seconds is None: + # Runs to the source's own end - cannot overrun it + return None + return start_seconds or 0.0, duration_seconds + if ( + task_args.get("start_frame") is not None + or task_args.get("num_frames") is not None + ): + if num_frames is None or not fps: + return None + return (start_frame or 0.0) / fps, num_frames / fps + return None + + +def slice_past_end_warnings(workflow_definition, source_indices=None, base_dir=None): + """Every `slice_audio` step whose source's real duration is already + knowable and whose requested slice reaches past it, as messages. + + Walks the substituted, expanded definition, the same convention + `dissolve_frame_errors` follows: `source_indices` maps an expanded step + back to the one the author wrote, and a path inside a `for_each` member + names the member. + """ + steps = workflow_definition.get("steps") + if not isinstance(steps, list): + return [] + + warnings = [] + for index, step in enumerate(steps): + if not isinstance(step, dict): + continue + task = step.get("task") + if not isinstance(task, dict) or task.get("command") != "slice_audio": + continue + task_args = task.get("arguments") + if not isinstance(task_args, dict): + continue + + path = _resolve_audio_path(task_args.get("audio"), base_dir) + if path is None: + continue + source_seconds = _source_seconds(path) + if not source_seconds: + continue + + region = _requested_region(task_args) + if region is None: + continue + requested_start, requested_length = region + + available = max(0.0, min(source_seconds - requested_start, requested_length)) + padded_seconds = requested_length - available + if padded_seconds * 1000.0 < SLICE_PAD_WARN_MS: + continue + + source = ( + source_indices[index] + if source_indices is not None and index < len(source_indices) + else index + ) + name = step.get("name") + where = ( + f" in member '{name}'" + if isinstance(name, str) and MEMBER_SEPARATOR in name + else "" + ) + path_str = render_path(("steps", source, "task", "arguments", "audio")) + warnings.append( + f"{path_str}: slice_audio will run {padded_seconds:.2f} s past " + f"the end of a {source_seconds:.2f} s source ({task_args.get('audio')}), " + f"so that much of the {requested_start + requested_length:.2f} s " + f"requested will be digital silence{where}. If you meant to fill " + f"a cut of this length, make a bed with the 'loop_audio' task " + f"('target_frames' + 'fps' matches one exactly) and slice that; " + f"if you meant the tail pad, nothing is wrong." + ) + return warnings + + +__all__ = ["slice_past_end_warnings"] diff --git a/dw/workflow.py b/dw/workflow.py index 513f7bde..21a911e1 100644 --- a/dw/workflow.py +++ b/dw/workflow.py @@ -805,6 +805,28 @@ def adapter_warnings(self, arguments=None): supplied=set(arguments or {}), ) + def slice_past_end_warnings(self, arguments=None): + """Every `slice_audio` step whose source's real duration is already + knowable and whose requested slice reaches past it - valid, padded + with silence rather than refused, but worth saying before the run + rather than only after it (#402). + + Best effort: a definition the schema or the expander refuses has its + own errors to report and none of them are this one. + """ + from .slice_preflight import slice_past_end_warnings + + try: + source_indices = [] + expanded = self.expanded_definition(arguments, source_indices) + except Exception: + logger.debug("No slice_past_end warnings available", exc_info=True) + return [] + base_dir = ( + os.path.dirname(os.path.abspath(self.file_spec)) if self.file_spec else None + ) + return slice_past_end_warnings(expanded, source_indices, base_dir) + def null_variable_argument_warnings(self, arguments=None): """Every required task argument fed by `variable:name` where name's value is null - downgraded out of `validation_errors` when diff --git a/tests/test_slice_preflight.py b/tests/test_slice_preflight.py new file mode 100644 index 00000000..db338658 --- /dev/null +++ b/tests/test_slice_preflight.py @@ -0,0 +1,148 @@ +"""A slice_audio source shorter than the requested slice, warned about at +validate time rather than only at run time - #402. + +Exercises the real decode path (`probe_media` against a genuine wav file) +rather than a mock of it, so a fixture shorter than its declared slice is +the same file the run itself would have padded with silence. +""" + +import os +import tempfile +import wave + +import numpy + +from dw.runs import activate_output_root, deactivate_output_root +from dw.slice_preflight import slice_past_end_warnings +from dw.workflow import workflow_from_definition + + +def write_wav(path, seconds=2.0, sample_rate=8000): + t = numpy.arange(int(seconds * sample_rate)) / sample_rate + samples = (numpy.sin(2 * numpy.pi * 220 * t) * 0.5 * 32767).astype(" Date: Thu, 24 Sep 2026 04:58:53 -0500 Subject: [PATCH 100/181] feat(mcp): #403 - trim model narrative from MCP descriptions, move the transcription walkthrough to the guide get_output_audio's walkthrough moves to WORKFLOW_GUIDE's "The loop" step 6; the H3 frame-grid example, the prompt family list and get_job_events' examples drop from their descriptions, every rule kept. diagnose.py's internal docstring points at the model skill for figures. Surface 13_883 -> 13_670.5; SURFACE_BUDGET stays 13_890, room reserved for #388. Co-Authored-By: Claude Opus 5.5 --- docs/WORKFLOW_GUIDE.md | 17 +++++++++++++++++ dw_mcp/diagnose.py | 13 ++++++------- dw_mcp/server.py | 33 ++++++++++----------------------- tests/test_mcp_server.py | 38 ++++++++++++++++++++++++++++++++++++++ 4 files changed, 71 insertions(+), 30 deletions(-) diff --git a/docs/WORKFLOW_GUIDE.md b/docs/WORKFLOW_GUIDE.md index 84fa3589..a7373ce7 100644 --- a/docs/WORKFLOW_GUIDE.md +++ b/docs/WORKFLOW_GUIDE.md @@ -642,6 +642,23 @@ the entry an item needs. like it does. 6. `get_output_image` to look at what was actually made, and say whether it answers the request. Nothing before this step establishes that it does. + `get_output_frames` looks at a video and `get_output_audio` listens to a + soundtrack. + + To confirm the words a clip speaks - a text-only client can't consume the + `AudioContent` block `get_output_audio` returns - transcribe it instead. + `validate_workflow(name="templates/transcribe-audio", + arguments={"input_audio": "output:"})` first (free; it takes an + audio file or a video's muxed soundtrack directly), then + `run_workflow(..., acknowledged_cost={"fingerprint": ..., "minutes": ..., + "downloads": [...]})` bound to that plan with `wait_seconds=55`, then + `get_output_text` on the result, and `delete_output(job_id=...)` the + scratch run afterward. This workflow's plan comes back + `basis: "unknown"` with `minutes: null` - nothing is curated or observed + for it - so quote what it actually takes rather than the plan: seconds, + not minutes (a few seconds per clip in practice). Four calls and a short + wait, not a GPU-spending read tool - keep the normal queue rather than + adding one. 7. Getting the files to the user's machine. `download_output` and `export_job` write on the machine running `dw.serve`, which over a remote `--mcp` endpoint is the GPU box. The last mile of every deliverable is the `url` diff --git a/dw_mcp/diagnose.py b/dw_mcp/diagnose.py index 00ab0420..fb11b619 100644 --- a/dw_mcp/diagnose.py +++ b/dw_mcp/diagnose.py @@ -258,8 +258,8 @@ def wait_for_job(client, job_id, timeout_seconds=20): named in `phase_detail`, `seconds_in_phase`, `seconds_since_event`, and `denoise_step`/`denoise_total_steps`, which are null until the denoise loop starts. `denoise_total_steps` is the schedule that actually runs, - which is not always the `num_inference_steps` asked for - MiniMax H3 - runs N-1 evaluations for N (#110). Two calls with the same phase and a growing + which is not always the `num_inference_steps` asked for (#110). Two + calls with the same phase and a growing `seconds_in_phase` but a moving `denoise_step` is a slow run; one where `denoise_step` is a number that does not move while `seconds_since_event` climbs is a stuck one. @@ -267,17 +267,16 @@ def wait_for_job(client, job_id, timeout_seconds=20): `denoise_step: null` under `generating` is neither: it is the lead-in the pipeline runs before the loop - encoding the prompt and every reference - which emits nothing and is well over a minute on a large - video model. Its length follows what it has to encode: ~90 s on - MiniMax H3 for a prompt with an image or audio reference, ~10 min once - a *video* reference is among them (measured 629 s for one 5 s 960x544 - clip on an RTX 3090). Silence there is expected, and `get_job_events` + video model. Its length follows what it has to encode, and a *video* + reference makes it much longer; the measured figures are the model + skill's (`minimax-h3` for H3). Silence there is expected, and `get_job_events` says which block it is inside while it lasts - one `log` line per top-level block of a modular pipeline. `seconds_since_event` only says something once `denoise_step` is a number, or in any other phase. Even then it is coarse: where a transformer block cache is configured the denoise steps are uneven - several cheap ones, then a full one - - so on H3 a 140 s gap between steps is a healthy run. Read liveness as + so a long gap between steps can be a healthy run. Read liveness as `denoise_step` having moved between polls minutes apart rather than as silence under a fixed threshold.""" requested = max(0.0, float(timeout_seconds)) diff --git a/dw_mcp/server.py b/dw_mcp/server.py index 4d0770fa..ee613f2b 100644 --- a/dw_mcp/server.py +++ b/dw_mcp/server.py @@ -157,7 +157,7 @@ def list_workflows( entry (`{variable, minutes, entries}`), so a run over a different-length list can be priced from it. `constraints`, present for a workflow that bounds a variable, is the rule each bounded one - has to satisfy, terse (`17*n+5, 124-345, rounds up`) - pass an + has to satisfy, terse - pass an `arguments` value outside it and `validate_workflow` refuses it for free, instead of the run failing after the weights are loaded. Templates only by default; `configures=