Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
fe7ce60
docs: list-driven steps proposal revised after review, with the stage…
dkackman Sep 11, 2026
3b7e856
feat(for_each): expansion pass - member naming, ceiling, item: substi…
dkackman Sep 11, 2026
2a2ea70
test(for_each): gather:, same-key siblings and the directed group errors
dkackman Sep 11, 2026
60358b3
test(for_each): music-video and dialogue-short expand back to their h…
dkackman Sep 11, 2026
d4ca7a1
feat(schema): for_each on a step
dkackman Sep 11, 2026
83e5bec
feat(for_each): expand in Workflow.run and validate the expanded defi…
dkackman Sep 11, 2026
2cc659b
feat(server): POST /api/validate expands for_each over the caller's list
dkackman Sep 11, 2026
171c0b7
docs(for_each): the item:/gather:/for_each conventions for agents, CL…
dkackman Sep 11, 2026
0c4ebb3
test(for_each): a list-driven job end to end; proposal marks stage 1 …
dkackman Sep 11, 2026
5b30f92
fix(for_each): copy leaves only inside a member, blame the right vari…
dkackman Sep 11, 2026
69ce983
docs(for_each): the new validation contract, paired by key, stage-2 n…
dkackman Sep 11, 2026
faa27e1
Merge list-driven-steps: for_each stage 1 - expansion pass, validatio…
dkackman Sep 12, 2026
2ca4eb8
plan: for_each stage 2 - the cut templates on a shots list
dkackman Sep 12, 2026
29e8b6b
fix(mcp): T020 reference check at submission, T021 progress inside a …
dkackman Sep 12, 2026
10b9d27
Merge mcp-feedback-round-5: T020 submission reference check, T021 den…
dkackman Sep 12, 2026
e68bbf5
feat(variables): an entry of a list-valued variable may reference ano…
dkackman Sep 12, 2026
3ba8605
fix(catalog): a concat over gather:<step> derives sequence
dkackman Sep 12, 2026
36a7383
fix(variables): review round 1 - black formatting, cycle as validatio…
dkackman Sep 12, 2026
5d8bfcd
feat(templates): music-video's slices and shots are two for_each grou…
dkackman Sep 12, 2026
3259339
feat(templates): dialogue-short's five shots are one for_each group o…
dkackman Sep 12, 2026
c8023f4
docs(for_each): the cut templates take one shots list; entries may na…
dkackman Sep 12, 2026
e33642a
docs(for_each): the listing's cost is the default list's total, divid…
dkackman Sep 12, 2026
5173a48
fix(mcp): T022 - denoise counter present-but-null through the generat…
dkackman Sep 12, 2026
0c4ecb2
Merge mcp-feedback-round-6: T022 denoise counter through the generati…
dkackman Sep 12, 2026
61cfe8b
test(for_each): pin the shot settings and the run-time realization or…
dkackman Sep 12, 2026
93f8cc2
fix: pre-flight references inside list arguments; cycle error type; e…
dkackman Sep 12, 2026
9cc4e7c
Merge branch 'list-driven-templates' into develop
dkackman Sep 12, 2026
5adcc92
plan: for_each stage 3 - catalog lists and per_entry cost, entry-key …
dkackman Sep 12, 2026
9f84bb0
ui(flow): edges for gather: and from_previous_result, so a list-drive…
dkackman Sep 12, 2026
1c329ac
fix: add 'denoise' to cSpell words in settings
dkackman Sep 12, 2026
1fb25ca
feat(server): stamp every job event with seconds since the job started
dkackman Sep 12, 2026
68231ca
feat(catalog): derive what an entry of a list-driven variable carries
dkackman Sep 12, 2026
e5ea96b
feat(validate): an entry key no for_each step reads is a warning at t…
dkackman Sep 12, 2026
67d87f8
feat(catalog): per_entry cost on a measured entry; variables preview …
dkackman Sep 12, 2026
82f97cc
fix(for_each): an empty list is an error; validation realizes constan…
dkackman Sep 12, 2026
f6cc9ef
fix(for_each): validation reports a failed constant: default instead …
dkackman Sep 12, 2026
70d1c2c
docs(for_each): the listing's lists block and per_entry cost, quoted …
dkackman Sep 12, 2026
a743885
fix(cli,mcp): dw.validate gates untrusted constant defaults; MCP docs…
dkackman Sep 12, 2026
e95aeb8
docs: reconcile SERVER/MCP/WORKFLOW_GUIDE with the lists/per_entry/wa…
dkackman Sep 12, 2026
e0ed4e9
fix: match dw.validate's --trust-workflows help text to dw.run's, ver…
dkackman Sep 12, 2026
d4e0e6e
Merge branch 'list-driven-stage-3' into develop
dkackman Sep 12, 2026
3e7e735
docs
dkackman Sep 12, 2026
5bb6e8e
minor
dkackman Sep 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/skills/model-family-onboarding/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ next family gets it in an afternoon rather than a rediscovery.
independent sources repeat it *and* nothing primary contradicts it.
- **Model knowledge is data, never engine code.** Template JSON, stored
prompts, a builtin's system prompt, README prose, a skill. No per-model
Python (proposal `docs/proposals/agent-catalog-legibility.md`, "Principle").
Python (proposal `docs/proposals/agent-catalog-legibility-complete.md`, "Principle").
- **Do not transcribe a prompt format the vendor publishes.** Point at it. If
the vendor ships an agent skill (MiniMax does) or a system-prompt constant
inside diffusers (Lightricks does), the dw skill says "use that" and a test
Expand Down
1 change: 1 addition & 0 deletions .vscode/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
"aiohttp",
"controlnet",
"cuda",
"denoise",
"dtype",
"exif",
"fullgraph",
Expand Down
47 changes: 44 additions & 3 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,7 +120,7 @@ read an inferred workspace back as one the user named - `get_prompt_dir` yields
to its older discovery (`./prompts`, then the walk up from the workflow file)
for an inferred workspace but not for an explicit one. `--workflow-dir`,
`--output-dir` and `--prompt-dir` each still override one folder. See
docs/WORKSPACES.md, and docs/proposals/server-workspaces.md for the later stages
docs/WORKSPACES.md, and docs/proposals/server-workspaces-complete.md for the later stages
(workflow search path, run directories, `asset:`/`output:` references).

### Type System
Expand Down Expand Up @@ -154,6 +154,27 @@ docs/WORKSPACES.md, and docs/proposals/server-workspaces.md for the later stages
rooted at the library rather than the workflow file. The library is `DW_PROMPT_DIR` /
`--prompt-dir`, else `./prompts` if it exists, else found by walking up from the
workflow file's directory
- A step carrying `for_each` (a list, or `variable:` naming one) is expanded by
`expand_for_each` (`dw/for_each.py`) into one ordinary step per entry, named
`<step>@<entry name or index>`, immediately after `replace_variables` in
`Workflow.run` and, with the caller's arguments folded, in `validation_errors`.
Inside a member `item:` / `item:field` is the entry (any type, spliced whole);
a later step reads the group with `gather:<step>` (a list; splices inside a
list); two groups over the same list pair by key (`slice` inside `shot@x` is
`slice@x`). `previous_result:` naming a group is a directed error. `@` is
reserved in step names; entry names are validated and unique; 32 entries max;
`release_pipeline`/`release_models` survive on the last member only. The
realized workflow keeps `for_each`; the manifest names the members. An entry
of a list-valued variable may reference another variable
(`"from_file": "variable:character_a_voice"`); `resolve_variable_values`
(`dw/variables.py`) replaces those once, before `realize_args`, refusing a
cycle, and `undeclared_variable_references` walks inside list/dict variable
values too. The catalog derives `lists` (`list_fields`, `dw/for_each.py`):
the fields an entry takes are the `item:` references the steps make, `name`
first; an entry key no step reads is a validation warning
(`entry_field_warnings`). A `cost` entry may carry `per_entry`
(`{variable, minutes, entries}`), measured, never derived. An empty
`for_each` list is an error; `expanded_definition` realizes constants first.
- Every run directory holds `workflow.json` beside its manifest: the *realized*
workflow, with the run's arguments folded into the variable defaults, the seed
it used, stored prompt text inlined and `output:.../latest/...` pinned to the
Expand Down Expand Up @@ -219,8 +240,28 @@ same reason - default setup cannot load a pack.
literal `previous_result:` or `from_previous_result` naming no *earlier* step, with
the JSON path it sits at. References otherwise resolve lazily per step, so a step
renamed in one place and not another failed only when the run reached it, after
every step before it had generated. A reference spelled by a `variable:` is left
alone - what it names is not knowable before substitution
every step before it had generated. The definition is substituted before the check,
so a reference spelled by a *declared* variable is checked by its value; one spelled
by an undeclared variable is itself a validation error (below)
- **`for_each` expands before the reference check** — `validation_errors` substitutes
(the caller's `arguments` when they are all good, else the defaults) and expands
first, so `gather:` and `item:` errors carry the path of the template step
(`steps[0].for_each[1].name`). Expansion records each expanded step's *source* index
(`expand_for_each(definition, source_indices)`), so a reference error always carries a
path in the file the author wrote, and one inside a member names the member in its
message. An undeclared `variable:` is a validation error at the path it sits at, not a
warning and not a complaint about the `for_each` list that did substitute: once a
`variables` block exists, `replace_variables` refuses an undeclared reference, so it is
a run that cannot start
- **The two MiniMax cut templates take one `shots` list** — since the stage-2
rewrite (2026-09-11) `templates/minimax/dialogue-short` and `music-video`
have no `shot_N_*` variables; a scripted caller passes `shots` (entries
`{name, prompt, references, num_frames}` and `{name, prompt, start_frame}`).
The members are `shot@<name>` in the manifest and the gallery. This is the
breaking change the next release note should name. The CLI and REPL only
take `name=value` strings, and a string handed to a list variable is
comma-split - so `shots` can only be supplied over the API/MCP (a JSON
body); `python -m dw.run` runs the templates' default list
- **Cartesian product explosion** — multiple `previous_result` references multiply: 4 images × 3 masks = 12 iterations
- **Component sharing requires exact key matching** between `shared_components` and `reused_components`
- **Built-in workflows** need explicit argument mapping: `"prompt": "variable:prompt"`
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@ once; without it the run fails partway through with a 401/403 from the Hub.

## Drive it from an agent

Then just ask. The agent has 50 tools covering the whole surface — the
Then just ask. The agent has 55 tools covering the whole surface — the
workflow catalog, the real diffusers pipeline signatures, the job queue, the
gallery, the model cache:

Expand Down
31 changes: 24 additions & 7 deletions docs/MCP.md
Original file line number Diff line number Diff line change
Expand Up @@ -187,7 +187,7 @@ Nothing in this sequence costs GPU time.

## Tool reference

52 tools in six groups. Names and arguments below are transcribed from
55 tools in six groups. Names and arguments below are transcribed from
`dw_mcp/server.py` — nothing here is renamed or reshaped for the docs.

### Catalog (read-only)
Expand All @@ -212,8 +212,8 @@ when no single workflow covers it.
| --- | --- | --- |
| `list_guides()` | — | List the documentation the engine serves: each guide's name, what it covers, and its section headings. The index is the routing table - match a request's shape against a heading rather than guessing |
| `get_guide(name, section=None)` | `name`, `section` | Get one guide whole, or one section of it. Prefer a section: a guide runs to thousands of lines. Section names match loosely, so a heading copied approximately still resolves |
| `list_workflows(shape=None, traits=None, configures=None, include_models=False)` | `shape`, `traits`, `configures`, `include_models` | List stored workflows. Always the server's compact view: each entry carries `summary`, `shape`, `traits`, `cost`, `kinds`, `variable_names`, and `configures` only when set - `get_workflow` has the full description and definition. `shape` keeps one of `image`, `image-set`, `image-edit`, `shot`, `sequence`, `audio`, `text`, `utility`; `traits` is comma-separated and every one listed must match (`has-audio`, `chained`, `image-conditioned`, `identity-referenced`, `needs-input-media`, `composes-workflows`); an unknown value in either is a 400 listing the vocabulary. Templates only by default - `configures=<template>` lists the checkpoint configs tuned for one, `include_models=true` lists them all. The first call to make for a request an existing workflow might cover |
| `get_workflow(name, variables_only=False)` | `name` | Get one stored workflow's full JSON definition. `variables_only=true` answers with just its variables and their defaults (long strings cut to 200 characters, the cut ones named in `truncated`) — the cheap way to confirm what a variable defaults to |
| `list_workflows(shape=None, traits=None, configures=None, include_models=False)` | `shape`, `traits`, `configures`, `include_models` | List stored workflows. Always the server's compact view: each entry carries `summary`, `shape`, `traits`, `cost`, `kinds`, `variable_names`, `lists`, and `configures` only when set - `get_workflow` has the full description and definition. `lists`, present for a list-driven workflow, names the fields an entry of each list takes, the steps over it and the default's length; `cost` may carry `per_entry`, the measured cost of one entry so a run over a different-length list can be priced from it. `shape` keeps one of `image`, `image-set`, `image-edit`, `shot`, `sequence`, `audio`, `text`, `utility`; `traits` is comma-separated and every one listed must match (`has-audio`, `chained`, `image-conditioned`, `identity-referenced`, `needs-input-media`, `composes-workflows`); an unknown value in either is a 400 listing the vocabulary. Templates only by default - `configures=<template>` lists the checkpoint configs tuned for one, `include_models=true` lists them all. The first call to make for a request an existing workflow might cover |
| `get_workflow(name, variables_only=False)` | `name` | Get one stored workflow's full JSON definition. `variables_only=true` answers with just its variables and their defaults (long strings cut to 200 characters, the cut ones named in `truncated`, including strings inside a list default, named like `shots[0].prompt`) — the cheap way to confirm what a variable defaults to |
| `get_schema()` | — | Get the JSON schema every workflow definition must satisfy |
| `list_pipelines()` | — | List every diffusers pipeline class this installation provides |
| `get_pipeline_signature(name)` | `name` | Get a pipeline's real call arguments |
Expand Down Expand Up @@ -249,7 +249,7 @@ The session starts in `default` and stays there unless it is told otherwise.

| Tool | Arguments | Purpose |
| --- | --- | --- |
| `validate_workflow(workflow=None, name=None, workspace=None, arguments=None)` | exactly one of `workflow` (inline definition) or `name` (a stored workflow, as `list_workflows` reports it), optional `workspace`, optional `arguments` | Check a workflow against the schema and against real pipeline signatures. Free and instant. Validating by name uses the workflow file's own directory as the base directory, so it sees what a run would. Returns every schema violation in `errors`, each with the JSON path it sits at, so a draft is fixed in one pass, and a `previous_result:` that names no earlier step is one of them. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's. Pass the same `arguments` you will pass to `run_workflow` and they are checked too - an undeclared or renamed variable name, a value that will not coerce to the declared type, and an `asset:`, `prompt:` or `output:` reference that names nothing this workspace can reach, each reported at `arguments.<name>`. `checked_arguments` lists what was covered, so a `valid: true` about the stored defaults cannot be mistaken for one about your values. `run_workflow` makes the same check and refuses a bad argument rather than queuing a job that fails on its first step |
| `validate_workflow(workflow=None, name=None, workspace=None, arguments=None)` | exactly one of `workflow` (inline definition) or `name` (a stored workflow, as `list_workflows` reports it), optional `workspace`, optional `arguments` | Check a workflow against the schema and against real pipeline signatures. Free and instant. Validating by name uses the workflow file's own directory as the base directory, so it sees what a run would. Returns every schema violation in `errors`, each with the JSON path it sits at, so a draft is fixed in one pass, and a `previous_result:` that names no earlier step is one of them. `warnings` covers what still runs but is probably wrong - a signature mismatch, and, for a list-driven variable, an entry key no step reads, at the entry's path. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's. Pass the same `arguments` you will pass to `run_workflow` and they are checked too - an undeclared or renamed variable name, a value that will not coerce to the declared type, and an `asset:`, `prompt:` or `output:` reference that names nothing this workspace can reach, each reported at `arguments.<name>`. `checked_arguments` lists what was covered, so a `valid: true` about the stored defaults cannot be mistaken for one about your values. `run_workflow` makes the same check and refuses a bad argument rather than queuing a job that fails on its first step |
| `list_workspaces()` | — | The server's workspaces and which one this session is using. Each has its own workflows, assets and outputs; the prompt library is shared by all of them |
| `use_workspace(name)` | `name` | Work in that workspace for the rest of the session - every later call reads and writes there. This is how to keep your work out of another agent's namespace rather than sharing the default one. Checked against the server, so a typo fails here rather than scoping every later call to nothing |
| `create_workspace(name, use=False)` | `name`, `use` | Create a workspace. Pass use=true to switch this session to it as well; otherwise the session stays where it was and the result says so |
Expand Down Expand Up @@ -283,11 +283,11 @@ references written in the same session.
| Tool | Arguments | Purpose |
| --- | --- | --- |
| `run_workflow(workflow_path=None, inline_workflow=None, arguments=None, acknowledged_cost=False, workspace=None)` | exactly one of `workflow_path` (a catalog name from `list_workflows`, with or without `.json`, or a path to a workflow file on the server) or `inline_workflow`, optional `arguments`, `acknowledged_cost`, `workspace` | Queue a workflow for generation. Returns as soon as the job is queued. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's |
| `get_job(job_id)` | `job_id` | Get a job's status, warnings, output manifest, error and traceback |
| `get_job(job_id)` | `job_id` | Get a job's status, warnings, output manifest, error and traceback. A running job also carries `progress` (below) |
| `get_job_workflow(job_id)` | `job_id` | The workflow the job actually ran. `realized: true` means every mutable input is pinned (arguments, seed, prompts, `output:latest`); `false` means the job predates run tracking and this is the definition as submitted. Pass it to `save_workflow` to keep it under a name |
| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `<workspace>/exports/<job id>/` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination - fetch the zip URL and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file); the archive already unpacks into one folder named after the job id |
| `get_job_events(job_id, after=-1, limit=200)` | `job_id`, `after`, `limit` | Get a page of a job's progress events |
| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, and the manifest once finished - without the arguments; `get_job` has those |
| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) |
| `cancel_job(job_id)` | `job_id` | Ask a queued or running job to stop |
| `rerun_job(job_id, acknowledged_cost=False, new_seed=False)` | `job_id`, `acknowledged_cost`, `new_seed` | Queue a fresh job from a previous job's stored specification. Costs GPU time, so it passes the same gate as `run_workflow`. `new_seed=true` draws a fresh seed into the workflow's seed variable — without it a seeded workflow's rerun repeats its arguments exactly and the step cache serves the whole run from the earlier one's files (`reused: true`), generating nothing. `get_job_workflow`'s `seed_variable` says whether there is one |
| `move_job(job_id, direction)` | `job_id`, `direction` (`up`\|`down`\|`front`\|`back`) | Reorder a queued job |
Expand Down Expand Up @@ -353,11 +353,28 @@ The intended loop:
again if it comes back `still_running: true` — or
`get_job_events(job_id)` repeatedly, passing back the previous call's
`last_seq` as `after`, for incremental progress instead of just a
terminal/not-terminal status
terminal/not-terminal status. Each event carries `at`, seconds since the
job started, so where a step's time went is a subtraction between two
events - `step_start` to `generating` is the lead-in a reused pipeline
still pays, `generating` to the first `pipeline_step` the encoding
4. `get_job(job_id)` for the finished manifest (or the error and traceback,
if it failed)
5. `get_output_image(name)` to look at a result image

While a job runs, `get_job` and `wait_for_job` carry a `progress` block -
the step being run, the phase (`loading`, `generating`, `decoding`,
`saving`) with the model in `phase_detail`, `seconds_in_phase`,
`seconds_since_event`, and `denoise_step`/`denoise_total_steps`, which are
null until the denoise loop starts. A single-step generation is minutes of
one phase, so two polls otherwise come back identical: read `denoise_step`
moving (slow but healthy) against a `denoise_step` that is a number and
stays put while `seconds_since_event` climbs (nothing is happening). A null
`denoise_step` under `generating` is neither - it is the lead-in the
pipeline runs before the loop, encoding the prompt and any reference image
or audio, ~90 s on MiniMax H3 with nothing emitted, so silence there is
expected. `cancel_job` stops at the next denoise or
step boundary, which `denoise_step` is also the measure of.

## Security

The MCP server adds no authentication of its own — it inherits the REST
Expand Down
Loading