Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 42 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,7 +153,9 @@ references) are documented above in *Workflow sources* and *Type System*.
`"output:ltx2/Gyre/latest/still.png"`. The name is `<workflow identity>/<run id>/<file>`
under the output root, and `latest` in the run-id position picks the newest run that
holds the file (run ids sort by their UTC timestamp; a failed or fully-cached run holds
only a manifest and is skipped). Resolved in `realize_args` beside `asset:` (`dw/runs.py`),
only a manifest and is skipped), and `v<N>` there picks the run whose version is N
(below) - exactly that run, with no fallback to an older one. Either is a selector only
where run directories are, and the realized workflow pins both to the run id. Resolved in `realize_args` beside `asset:` (`dw/runs.py`),
against the output root `Workflow.run` activates, and confined to it
- A generated file becomes a stable input with `POST /api/assets/keep` (gallery "Keep as
asset", MCP `keep_output`): it is hard-linked, else copied, from the workspace's outputs
Expand Down Expand Up @@ -368,6 +370,45 @@ same reason - default setup cannot load a pack.
`JobManager.realized` finds the file. `exports` is a reserved workspace name:
`POST /api/jobs/{id}/export` gathers one finished job into
`<workspace>/exports/<job id>/` and `GET /exports/<job id>.zip` streams it.
- **A run has a number, and it is not derived from the listing** - a file's name
is per *step*, so four runs of one workflow write four files called
`AcornWarsCutAndScore-film.7-0.0.mp4` and the gallery drew four identical
captions: the run id told them apart but is not something anyone says out
loud, so an agent had no way to name one of them to a person. Every run now
takes an ordinal, `assign_run_version` (`dw/runs.py`) at the moment
`Workflow.run` opens the run directory, recorded as `version` in
`manifest.json` and read back by `run_versions`. Assigned once and never
recomputed, which is the point: deleting a middle run leaves a gap rather
than sliding every later number down, so "version 5" still means the same
run tomorrow. Assignment is `max(recorded) + 1` over *every* sibling
manifest, not one past the newest - run ids are chronological only to the
second, and within one second the spec digest decides the sort, which is
exactly what three quick reruns hit. The number is on disk from the moment
the run opens - a `status: "running"` manifest is written before the first
step and rewritten in full at the end - so a hard kill does not lose it and
a second process opening a run of the same workflow sees it. A run with no
recorded number (made before the field, or killed before even that first
manifest) is ranked: the unrecorded runs older than every recorded one take
the numbers beneath the lowest, later ones continue from the highest before
them. A ranked number would move when an older sibling is deleted, so
`record_run_versions` writes it into the manifest on the two write paths -
a run opening and a run directory being deleted; the listing never writes,
and a run with no manifest at all is left ranked. A gap in the numbers is
not only a deletion: a failed run or a fully cached rerun takes a number
and may have no media for the gallery to show under it. `GET
/api/gallery` and the metadata route carry `version` and `run_id`
(`run_versions` read once per identity per listing, not per file), and
`?folder=&version=` lists one run's files. The number is also a name:
`output:<identity>/v4/<file>`. The `run_start` event carries it, the job
records it (`run_version`, a `jobs.sqlite` column) and the export README and
zip download name (`<identity>-v4-<job id>.zip`) carry it too. MCP
`list_gallery` teaches the vocabulary and takes `folder`/`version`, and the
web UI reads the field only - a `v4` chip on the gallery card, the jobs list
and the job page, the run id in the gallery's detail pane. Nothing on disk is
renamed, so `output:` references, the step cache and `keep_output` are
untouched. Two limits taken deliberately: deleting the *newest* run frees
its number for reuse (the high-water mark lived in the manifest that went
with it), and the flat layout has no runs, so `version` is null there.
- **Result subfolders**: a step's `result.subfolder` (`dw/subfolders.py`) puts its files
in a subfolder of the run directory - `<run>/final/x.mp4` - by convention `final` or
`intermediate`; the engine treats no name specially and there is no default.
Expand Down
4 changes: 2 additions & 2 deletions docs/MCP.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,7 +226,7 @@ when no single workflow covers it.
| `get_health()` | — | Check that the server is alive, and which machine answered: `version`, `device`, whether a model process is currently resident (`worker_alive`), the job running now and the queue depth. `worker_alive: false` is the normal idle state on a server that has not run a job since startup or the last memory clear - not a fault - the on-demand worker starts with the next job (#206) |
| `get_server_info()` | — | What this installation can do and where it keeps things: `device` (the accelerator a run will use), `version`, the `workspace` this session is working in and the workflow/asset/output/prompt `directories` of *that* workspace, the bind address and port, whether a token is required, and whether MCP is mounted. Check the device before authoring - a CUDA-only choice (bitsandbytes, `torch.compile`, flash attention) is not available on an `mps` or `cpu` server. `runtime` (#222) reports the Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (when `nvidia-smi` is reachable), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (`null` for one not installed) - for diagnosing an environment mismatch between boxes without shelling in |
| `list_jobs(limit=20, status=None, workspace=None)` | optional `limit` (newest N), `status` (one state or a comma-separated set of `queued`, `running`, `succeeded`, `failed`, `cancelled`), `workspace` | List queued, running and recent jobs, **newest first**. Bounded by default: the unbounded listing was over a client's tool-result limit on a server with a few months of history, which made it a tool that could not be called at all. `total` says how many matched and `truncated`/`next` say so when the answer was cut - raise `limit` or narrow with `status`. Without `workspace`, a named workspace lists its own jobs and the default one lists every job the server holds |
| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None)` | `limit`, `subfolder`, `only_orphans`, `workspace` | List generated output files, newest first. A name is `<workflow>/<run id>/<file>`, where `<file>` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/<name>` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it |
| `list_gallery(limit=50, subfolder=None, only_orphans=False, workspace=None, folder=None, version=None)` | `limit`, `subfolder`, `only_orphans`, `workspace`, `folder`, `version` | List generated output files, newest first. A name is `<workflow>/<run id>/<file>`, where `<file>` may sit in the subfolder the step chose (`final/episode.mp4`); each entry carries `folder` (the workflow) and `subfolder` (by convention `final` or `intermediate`, `''` when the step chose none, any path the workflow wrote otherwise), and `subfolder="final"` lists only deliverables. Each entry carries `run_id` and `version` - that run's ordinal among the workflow's runs, which is how one of several runs that wrote the same basename is named to a person: the web UI labels the same file `v5`. The number is assigned when the run opens and never renumbered, so deleting a run leaves a gap rather than sliding the rest down (a failed run, or a rerun that reused every step, leaves one too - it took a number and may have nothing to list), and it is `null` under the flat output layout, which has no runs. `folder` with `version` lists that one run's files, and `output:<folder>/v5/<file>` names one in a workflow; every other tool takes `name`. Each entry also carries a ready-made `url`, already scoped to the workspace that made it - a hand-built `/outputs/<name>` URL 404s for anything but the default workspace. `only_orphans=True` inverts the call: instead of files, it returns run directories with no media anywhere under them (`runs`, each `{name, mtime}`) - a run whose output was deleted before `delete_output` could remove it by name, or one that failed before writing anything; `subfolder` does not apply in this mode, and `name` is exactly what `delete_output` accepts (#170). `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it |
| `get_gallery_metadata(name, envelope=False, workspace=None)` | `name`, `workspace` | Get the metadata embedded in a generated file — or, when `name` is an `asset:` reference, what an *input* asset holds (`source` says which; `job` is null for an asset). Reading an input's duration, frame count, fps and sample rate before a run is how a caller learns the `total_frames`, `fps` and `sample_rate` a workflow expects it to supply: the exact workflow and arguments that produced it, and, for audio/video, a `media` block (duration, rate, channels, fps, size, peak/mean dBFS). `envelope=true` adds `media.envelope` — `rms_dbfs` and `peak_dbfs` one entry per second — which is what locates something in a track rather than measuring the whole of it. `workspace` names the workspace for this one call without switching the session to it - the same pin `run_workflow` takes, so a job run into another workspace stays reachable from the session that queued it |

### Media
Expand Down Expand Up @@ -306,7 +306,7 @@ references written in the same session.
| `get_job_workflow(job_id)` | `job_id` | The REST equivalent is `GET /api/jobs/{id}/workflow` (see [SERVER.md](SERVER.md#jobs-api)). The workflow the job actually ran. `realized: true` means every mutable input is pinned (arguments, seed, prompts, `output:latest`); `false` means the job predates run tracking and this is the definition as submitted. Pass it to `save_workflow` to keep it under a name |
| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `<workspace>/exports/<job id>/` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination - fetch the zip URL and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file); the archive already unpacks into one folder named after the job id |
| `get_job_events(job_id, after=-1, limit=200)` | `job_id`, `after`, `limit` | Get a page of a job's progress events |
| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) |
| `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most 55 seconds** — a larger `timeout_seconds` is clamped, not honoured, because no MCP client holds a tool call open for a generation's real runtime, so budget one call per ~55s of the job. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, `run_id` and `run_version` (the run's `v5`, as the gallery labels it), and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) |
| `cancel_job(job_id)` | `job_id` | Ask a queued or running job to stop |
| `clear_memory()` | — | Drop every loaded pipeline and the step cache, freeing VRAM/RAM immediately instead of waiting for the next job to evict one model for another. Also drops the step cache, so a seeded workflow that would otherwise reuse cached results regenerates on its next run. Refused with a 409 while a job is running or queued - the queue is FIFO, so wait for it to finish and retry rather than expecting this call to block until it does (#221) |
| `rerun_job(job_id, acknowledged_cost=False, new_seed=False)` | `job_id`, `acknowledged_cost`, `new_seed` | Queue a fresh job from a previous job's stored specification. Costs GPU time, so it passes the same gate as `run_workflow`. `new_seed=true` draws a fresh seed into the workflow's seed variable — without it a seeded workflow's rerun repeats its arguments exactly and the step cache serves the whole run from the earlier one's files (`reused: true`), generating nothing. `get_job_workflow`'s `seed_variable` says whether there is one - `acknowledged_cost` is `true` or the bound `{fingerprint, minutes, downloads}` from the validate plan; a 409 means the plan changed and the message carries the new estimate |
Expand Down
3 changes: 2 additions & 1 deletion docs/SERVER.md
Original file line number Diff line number Diff line change
Expand Up @@ -441,7 +441,8 @@ The editor's forms come from these; they are just as usable from scripts:
gallery entry carries `folder` (the workflow identity, the run id dropped)
and `subfolder` (what followed the run id - the `final`/`intermediate` a
step's `result.subfolder` chose, `''` when it chose none); `?folder=` and
`?subfolder=` filter independently, and the reply's `folders` and
`?subfolder=` filter independently (`?version=` too - with `?folder=`,
the one run the gallery labels `v4`), and the reply's `folders` and
`subfolders` list every distinct value over the whole tree, `''` always a
member of each so root-level files stay selectable
- `GET /api/gallery/{name:path}/download` — download an output file
Expand Down
13 changes: 9 additions & 4 deletions docs/WORKFLOW_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -289,7 +289,8 @@ for existence.
to the same file.
- `output:` — `output:<workflow identity>/<run id>/<file>` is a file an earlier
run wrote, under the output root and confined to it. `latest` in the run-id position
picks the newest run that holds that file. A run id is not stable against
picks the newest run that holds that file; `v<N>` picks the run the gallery labels
`v<N>` (`list_gallery`'s `version`), and only that run. A run id is not stable against
pruning: to depend on a generated file, promote it with `keep_output` and
reference the `asset:` name instead.
- `prompt:` — `prompt:name` or `prompt:folder/name` is a stored prompt's
Expand Down Expand Up @@ -1459,7 +1460,7 @@ Beside that manifest the run also writes `workflow.json` — the *realized*
workflow, meaning the one that actually ran. Every mutable input is pinned into
it: the caller's `arguments` folded into the `variables` defaults, the seed the
run used, each `prompt:` reference replaced by the stored text, and each
`output:<identity>/latest/<file>` rewritten to the run id it resolved to.
`output:<identity>/latest/<file>` (or `/v<N>/`) rewritten to the run id it resolved to.
`asset:`, `constant:`, `previous_result:` and `builtin:` are kept as written —
each already names something pinned by the asset library or by the manifest's
`dw_version` — and a sub-workflow named by local path is kept with its file's
Expand Down Expand Up @@ -1657,8 +1658,12 @@ second-stage workflow name the first stage's product without being edited after
run - and keeps working when the newest run failed part way, or reused every step from
the cache and so wrote nothing of its own but a manifest. Runs sort by their id, which
starts with a UTC timestamp, so "newest" needs no file timestamps and survives a
directory being copied. `latest` only selects a run where run directories are; a
workflow or file that happens to be called `latest` is still named as itself.
directory being copied. `v<N>` in the same position names the run whose version is N -
the `v4` the gallery labels its files with - so the number a person was told is a name
a workflow can take. Unlike `latest` it picks exactly one run: `v4` not holding the file
is an error, not a reason to try `v3`. `latest` and `v<N>` only select a run where run
directories are; a workflow or file that happens to be called either is still named as
itself.

Like `asset:`, a reference resolves to a path and then whatever loads paths loads it, so
it works under `image`, `video`, a `from_file`, or a list of them. The audio tasks take
Expand Down
18 changes: 12 additions & 6 deletions dw/realize.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@
is_output_reference,
output_root as default_output_root,
resolve_output_reference,
version_selector,
)
from .security import SecurityError, validate_workflow_path
from .workflow_sources import resolve_sub_workflow, SubWorkflowNotFound
Expand Down Expand Up @@ -170,14 +171,19 @@ def _inline_prompt(reference, annotations, prompt_dir, base_dir):


def _pin_output(reference, output_root):
"""'output:<identity>/latest/<file>' rewritten to the run it resolved to.

An explicit run id is already pinned, so it is returned untouched without
touching the disk - realizing must not fail on a reference the run has
not reached yet.
"""'output:<identity>/latest/<file>' - or '/v4/' - rewritten to the run it
resolved to.

A version is stable, but deleting the newest run frees its number for
reuse, so the realized copy names the run id either way. An explicit run
id is already pinned, so it is returned untouched without touching the
disk - realizing must not fail on a reference the run has not reached
yet.
"""
name = reference.removeprefix(OUTPUT_PREFIX).strip()
if LATEST not in name.split("/"):
if not any(
part == LATEST or version_selector(part) is not None for part in name.split("/")
):
return reference
root = output_root or default_output_root()
try:
Expand Down
Loading
Loading