diff --git a/.claude/skills/model-family-onboarding/SKILL.md b/.claude/skills/model-family-onboarding/SKILL.md index 75bcb574..43906c4b 100644 --- a/.claude/skills/model-family-onboarding/SKILL.md +++ b/.claude/skills/model-family-onboarding/SKILL.md @@ -111,7 +111,11 @@ triggering; call `get_server_info` and a shape-filtered `list_workflows` before trusting any name; the shape decision as choices; the hard numeric rules; the vendor pointer for prompts and nothing else; validate, quote cost, run, look, and the family's failure modes; sources with dates. Near the size -of the README it derives from, under the 12 KB cap. +of the README it derives from, `SKILL.md` under the 12 KB cap. What only some +requests need (a long vendor spec, a multi-shot recipe, checkpoint swaps) goes +in `references/.md` beside it, named by path at the point an agent +needs it; the tests read references as part of the skill and fail on a +reference no `SKILL.md` links. Add the family's numbers to `tests/test_plugin_skills.py`, each checked against the diffusers module that enforces it, and any quoted vendor text to diff --git a/docs/MCP.md b/docs/MCP.md index 3e8c4eeb..0437a64c 100644 --- a/docs/MCP.md +++ b/docs/MCP.md @@ -308,7 +308,7 @@ references written in the same session. | `run_workflow(workflow_path=None, inline_workflow=None, arguments=None, acknowledged_cost=False, workspace=None, wait_seconds=0)` | exactly one of `workflow_path` (a catalog name from `list_workflows`, with or without `.json`, or a path to a workflow file on the server) or `inline_workflow`, optional `arguments`, `acknowledged_cost`, `workspace`, `wait_seconds` | Queue a workflow for generation. Returns as soon as the job is queued - unless `wait_seconds` is above 0, in which case the call then waits on the queued job exactly as `wait_for_job(job_id, timeout_seconds=wait_seconds)` would (same per-call cap, clamped not honoured) and the result carries the queued-job fields plus that wait's (`status`, `still_running`, `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds`, `timeout_capped`, the slim `job`); when the cap covers the job's runtime one call is the run and the wait, and a `still_running: true` result is followed with `wait_for_job` as before. `workspace` names the workspace for this one call without switching the session to it - use it to pin a job whose `output:` or `asset:` references live in a workspace other than the session's - `acknowledged_cost` is `true` or the bound `{fingerprint, minutes, downloads}` from the validate plan; a 409 means the plan changed and the message carries the new estimate, and nothing is waited on | | `get_job(job_id)` | `job_id` | Get a job's status, warnings, output manifest, error and traceback; each manifest entry's `subfolder` is the in-run subfolder the step declared - by convention `final` for the deliverable, `intermediate` for scratch, `''` for none. A running job also carries `progress` (below) | | `get_job_workflow(job_id)` | `job_id` | The REST equivalent is `GET /api/jobs/{id}/workflow` (see [SERVER.md](SERVER.md#jobs-api)). The workflow the job actually ran. `realized: true` means every mutable input is pinned (arguments, seed, prompts, `output:latest`); `false` means the job predates run tracking and this is the definition as submitted. Pass it to `save_workflow` to keep it under a name | -| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `/exports//` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination. `auth_required` says whether opening the zip needs this server's bearer token, which this agent cannot attach to someone else's fetch (#353): when false, fetch `open_url` yourself and unpack it into `exports/` under the session's working directory (a deliverable, not a temp file) - the archive already unpacks into one folder named after the job id, so do not create that folder first; when true, hand `open_url` to the person instead of fetching it | +| `export_job(job_id, overwrite=False)` | `job_id`, `overwrite` | Gather one finished job into `/exports//` on the server: the realized workflow, the run's manifest, the job row, a README, and copies of the assets, earlier-run inputs and outputs. Returns the directory, a zip URL, the file list with sizes and the total. The three JSON files are in the zip, not repeated here - get_job_workflow and get_job serve them individually. **The directory is on the machine running the server**, like `download_output`'s destination. The zip needs no token (`/exports/*.zip` is ungated like `/outputs`, #592), so `auth_required` is false and the agent fetches `open_url` itself (prefixing a relative one with the server address) and unpacks it into `exports/` under the session's working directory (a deliverable, not a temp file) - the archive already unpacks into one folder named after the job id, so do not create that folder first; only an agent without HTTP gives the user `open_url`. `true` is kept as a forward guard for a gated zip, when `open_url` is handed to the person instead | | `get_job_events(job_id, after=-1, limit=200)` | `job_id`, `after`, `limit` | Get a page of a job's progress events | | `wait_for_job(job_id, timeout_seconds=20)` | `job_id`, `timeout_seconds` | Block until a job reaches a terminal status, or `timeout_seconds` elapses. **One call blocks for at most the server's cap** — 55 seconds unless the deployment sets `DW_MCP_MAX_WAIT_SECONDS` higher, and the tool's description states the live value; a larger `timeout_seconds` is clamped, not honoured. Ask for the job's `plan.estimate` plus a margin: under the cap that is one call for the whole job, over it one call per cap. A deployment raises the cap only as far as its clients (and anything between them and the server) hold one silent HTTP request open - Claude Code's limit is `MCP_TOOL_TIMEOUT`. Every reply carries `waited_seconds`, `timeout_requested_seconds`, `timeout_applied_seconds` and `timeout_capped`, so a capped return is distinguishable from an elapsed one. Use instead of hand-polling `get_job`/`get_job_events` in a loop; if it returns `still_running: true`, call it again. Returns a slim job - status, warnings, error, `run_id` and `run_version` (the run's `v5`, as the gallery labels it), and the manifest once finished - without the arguments; `get_job` has those. A running job also carries `progress` (below) | | `cancel_job(job_id)` | `job_id` | Ask a queued or running job to stop | diff --git a/docs/RELEASING.md b/docs/RELEASING.md index 29a25a3a..56fb1296 100644 --- a/docs/RELEASING.md +++ b/docs/RELEASING.md @@ -7,6 +7,63 @@ notes from commits at tag time (see below). This section is a scratch pad for items a branch's author wants the next release note to name; clear it when a release ships. +### 0.9.0 + + + +A smaller range than 0.8.0: the LoRA catalog grows to cover three more +bases, the cost planner is corrected for list-driven and `for_each` steps, +and `export_job` reports its zip's real auth gating. No upgrade-order +changes. + +**LoRA catalog** (docs/LORAS.md) + +- 11 LTX-2.5 entries added: IC-LoRAs for alpha generation, clean plate, + colorization, day-to-night, layout-to-render, detail refinement, + restoration, SDR-to-HDR and water simulation, plus cinemagraph and + slow-motion control. +- MiniMax-H3 goes to 13 trial and 3 rejected entries beyond the 0.8.0 set: + styles, speech, orbits, action and motion, and a diffusers-native 4-step + turbo. FastVideo FastH3 (`.diff` keys), RAVEN (unrecognised prefix, loads + nothing silently) and TaoMate (lowercase `lora_a`/`lora_b`) are rejected. +- New bases: Qwen-Image-2.1 (10 trial, including three few-step distills + that need their scheduler and sigma overrides, plus edit LoRAs; Fun-Acc + rejected) and Z-Image-Turbo (13 trial). The Z-Image entries match + `Tongyi-MAI/Z-Image-Turbo` only; whether they apply to the SDNQ + checkpoint is untested. +- Every new entry pins the repo's current sha, and its header was read + against the installed diffusers converter. + +**Cost planning** + +- `other_device` figures are re-priced for `for_each` counts and shifted + list drivers (#589, #590). +- A shifted per-entry field in a list driver resets to unknown, and a + summed child figure takes the children's basis (#593). A numeric string + in a list-driver entry compares as its number (#593). +- Measured MPS cost entries for `templates/ltx2/text-to-video` and Music 3 + (#590). + +**Fixes** + +- `assemble-and-score` threads `sample_rate` into its edit join, and mixed-rate + shots resampled to a pinned `sample_rate` no longer draw a warning advising + you to pass it (#594). +- `export_job` reports the zip's real gating: `/exports/*.zip` is ungated + like `/outputs`, so `auth_required` is false whether or not the server has + a token. The MCP `next` text and the skills now say to fetch `open_url` and + unpack into `exports/` to bring a project home (#595, #592). + +**Plugin** + +- The `ltx-2.5`, `minimax-h3` and `minimax-music3` skills moved + request-specific detail (the LTX caption spec, H3 checkpoint and LoRA + combinations and cuts recipes, Music 3 loudness) into `references/` + beside each `SKILL.md`, which are read when the skill points there. The + 12 KiB cap applies to `SKILL.md` alone; a new test fails on an unlinked + reference or a dead link. +- The H3 and LTX skills point at `list_loras`. + ### 0.8.0 @@ -38,6 +95,9 @@ upscaler and two security fixes. - GHSA-crqf-hw9p-r739: `create_workspace`'s MCP result is `name`, `default`, `current` and `next` only; `list_workspaces(detail=true)` is the opt-in for folder paths. `POST /api/workspaces` is unchanged. +- `GET /api/loras/recommend` and `recommend_loras`: `hub_error` names the + exception type, or the HTTP status, never the exception's text, which + could name the server's HF cache directory. The log keeps the full error. - UI lockfile bumps for open Dependabot alerts (devalue, dompurify, brace-expansion, undici). diff --git a/docs/REMOTE.md b/docs/REMOTE.md index 2e26a1b0..7200039c 100644 --- a/docs/REMOTE.md +++ b/docs/REMOTE.md @@ -48,6 +48,11 @@ responses add an `absolute_url` / `absolute_zip_url` built from it; nothing guesses this from request headers, so an unconfigured server omits the field rather than composing a wrong origin. +The export zip, like `/outputs` and `/inputs` files, needs no token (#592), so +an agent with HTTP fetches `export_job`'s `open_url` itself, prefixing a +relative one with the address it reaches the server at; `DW_PUBLIC_URL` is for +the URLs handed to a person. + ## Browser Open `http://:8765`. Click the key icon next to the theme toggle, diff --git a/docs/proposals/complete/job-export-bulk-download-complete.md b/docs/proposals/complete/job-export-bulk-download-complete.md new file mode 100644 index 00000000..ba0fdc2e --- /dev/null +++ b/docs/proposals/complete/job-export-bulk-download-complete.md @@ -0,0 +1,128 @@ +# Bulk download of a job's outputs: export_job reports the zip's real gating (#592) + +Written by model `claude-opus-5-5` via provider `anthropic` (close-out, +2026-10-05). Plan v2 was approved by Don on 2026-10-05 (Q1 and Q2 kept +their defaults). Stage 1, #595, shipped the same day at `develop` @ +`69a3ac8a` and was verified on its first pass. + +## The idea + +A field report from an agent job on mini-ai: bringing 41 shots and 7 music +cues back to the Mac meant reading each job's file list and curling every +file. The issue asked for a "fetch everything from this job" call, and +whether it should extend `export_job` or be a new tool, and how large +binary transfer should work over MCP. + +## Verdict + +**Build smaller: one stage, no new tool.** The bulk download already +existed. `POST /api/jobs/{id}/export` builds `exports//` (outputs, +inputs, assets, workflow, manifest, job record, README) and returns +`zip_url` (`/exports/.zip`), and `/exports` is ungated by design, beside +`/outputs` (`dw/server/routes/files.py`, pinned in +`tests/test_security_auth.py`, stated in `docs/SERVER.md`). + +What broke it was one field. The export response set +`auth_required = bool(state.api_token)`: whether the *server* has a token, +not whether the *zip URL* needs one. `dw_mcp/exports.py` then told the agent +"do not fetch it; hand open_url to the person". On any token-bearing server +(mini-ai, lem) an agent was told not to fetch a URL that needs no token, +which left the per-file curl loop the report describes. #353 had introduced +the flag on the premise that the zip sat behind the token; the code said it +didn't. + +**Binary transfer over MCP itself: don't build.** A tool result lands in +the agent's context, so base64 video would swamp it (#203 capped inline +upload at 4 MB for the same reason), and MCP has no channel to the client's +disk (`download_output` over a mounted endpoint writes on the server). +Over MCP, big files travel by URL; the tool's job is to hand over one URL +that works. + +## Decisions (Don, 2026-10-05) + +- **Q1. Keep `/exports/*.zip` ungated?** Yes. It serves nothing `/outputs` + doesn't already serve ungated. Don required the test pinning "token + configured → `auth_required` false → unauthenticated GET returns 200", + so the field can't drift from the route again. +- **Q2. Keep `auth_required` (always false today)?** Yes: no shape change, + and it guards a future gating of `/exports`. +- **Q3. Surface removal:** none found. + +## What was built (stage 1, #595) + +- **Server** (`dw/server/routes/jobs.py`): the export response sets + `auth_required = False`, with a comment tying it to `files.py`'s ungated + list. Response shape unchanged, so the OpenAPI dump didn't move. +- **MCP `next`** (`dw_mcp/exports.py`): when `auth_required` is false, + `next` says to fetch `open_url` with whatever HTTP you have (it needs no + token; prefix a relative URL with the server address), and unpack into + `exports/` under the working directory without pre-creating the job-id + folder. Only an agent without HTTP is told to give the user `open_url`. + The `true` branch is unchanged, as the forward guard. +- **MCP description** (`dw_mcp/tools_jobs.py`): "do NOT fetch it" is gone. + It shrank from 1429 to 1313 characters, so `SURFACE_BUDGET` was left + alone. +- **Skills:** `script-to-video` and `series-episodes` end their delivery + sections with "take the project home": `export_job` once per job, then + fetch each zip into `exports/`. `minimax-music3/references/keeping-a-run.md` + lost its stale "can't attach the token" wording. `minimax-h3` already + keyed on the field; `minimax-music3/SKILL.md` and `ltx-2.5` had no export + wording. +- **Docs:** the `export_job` row in `docs/MCP.md`; a paragraph in + `docs/REMOTE.md` saying the zip needs no token and an agent with HTTP + fetches it itself. +- **Tests:** + - `tests/test_server_exports.py::…test_with_a_token_the_zip_is_not_auth_required_and_opens_without_one`: + with a token, `auth_required` is false and a token-less GET of + `zip_url` returns 200 with `/manifest.json` inside; a token-less + export POST is still 401 (Don's Q1). + - `tests/test_mcp_exports.py`: both `next` branches. + - `tests/test_mcp_server.py::test_export_job_description_says_to_fetch_the_ungated_zip` + replaces `…_is_auth_aware`. + +### Deviations from the plan + +- **No plugin version bump.** `plugins/dw/README.md` ties the plugin's + version to the engine's, bumped only by the release script. The tester's + plugin tree follows `origin/develop`, so the skill text was live without + one. +- The no-HTTP fallback reads "give the user open_url to open" rather than + the old "hand open_url to the person". + +## Verified + +On lem (token set) at `69a3ac8a`, cases C-F180–C-F182: `export_job` of a +finished job returns `auth_required: false` and the fetch `next`; unknown +id, running *and queued* jobs, and a repeat without `overwrite` are all +refused as before; `overwrite=true` replaces the export identically; the +served description and skills no longer steer away from fetching. The +token-less GET of the zip was not checked over MCP (the tester has no HTTP +client); it rests on the pinned unit test above. + +Left behind: no MCP tool reaches `exports/`, so each run of C-F180/C-F181 +leaves an `exports//` on the test server. + +## Bounces per stage + +- **Stage 1 (#595): 0.** The architecture review passed and the tester + verified on the first pass. + +No `usage:` figures were recorded on the stage, so cost is left out (the +plan estimated about $3–4). + +## Deferred + +- **A multi-job bundle** (`export_jobs([...])` or a series zip): a new tool, + a new naming scheme under `exports/`, a UI flow and new symlink cases, + about 4 stages and pressure on the surface budget. **Bring it back** when + a field report shows a project spread over enough jobs that one zip per + job is the friction; `script-to-video`'s one-job-per-shot film is the + likely trigger. +- **An outputs-only zip** (no `assets/` and `inputs/` copies, so no doubled + server disk): `POST /api/gallery/archive` does name-based zips but is + gated, has no job filter and no MCP tool. **Bring it back** on a report + of server disk pressure from `exports/`. +- **Removing an export over MCP:** no tool reaches `exports/` today, so + exports accumulate on the server. Not in scope; worth a fix if the + outputs-only zip above comes back, or if the test server's `exports/` + grows. diff --git a/docs/proposals/todo.md b/docs/proposals/todo.md index 5002f9ae..678427ad 100644 --- a/docs/proposals/todo.md +++ b/docs/proposals/todo.md @@ -87,6 +87,13 @@ remaining deferred fix is recorded in soft in the faces. Record, including what was deferred (the refine pass, persisted latents, task weights in `downloads_required`): `complete/h3-latent-upscale-complete.md`. +- **Bulk download of a job's outputs** (#592, stage #595), shipped + 2026-10-05 with no new tool: `export_job` now reports the ungated + `/exports` zip as `auth_required: false` on a token server and tells the + agent to fetch it, and the multi-job skills name it as how a project goes + home. Record, including what was deferred (a multi-job bundle, an + outputs-only zip, removing an export over MCP): + `complete/job-export-bulk-download-complete.md`. ## Declined diff --git a/dw/plan.py b/dw/plan.py index fd74f857..7b69eebd 100644 --- a/dw/plan.py +++ b/dw/plan.py @@ -17,6 +17,7 @@ import hashlib import json import logging +import math import os from huggingface_hub import model_info @@ -340,6 +341,36 @@ def _driver_comparable(value): return json.dumps(value, sort_keys=True, default=str) +def _numeric_fields(entries): + """{field: {values}} over the numeric fields of a list's dict entries. A + numeric string ("243") counts as its number, as `_driver_comparable` does + for a scalar driver.""" + fields = {} + if not isinstance(entries, list): + return fields + for entry in entries: + if not isinstance(entry, dict): + continue + for key, value in entry.items(): + number = _driver_comparable(value) + if isinstance(number, float) and math.isfinite(number): + fields.setdefault(key, set()).add(number) + return fields + + +def _list_entry_field_shifted(default_entries, effective_entries): + """Whether a list driver's entries carry a numeric field (a per-shot + `num_frames`, say) with a value none of the default entries had. The + curated figure was measured over the default entries' values, so a value + outside them has no matching bucket - the #267 rule applied inside a + list driver (#593).""" + measured = _numeric_fields(default_entries) + for field, values in _numeric_fields(effective_entries).items(): + if field in measured and not values <= measured[field]: + return True + return False + + def _scalar_driver_shifted(definition, expanded, list_entries): """Whether a declared, non-list `cost_driver` was overridden away from the default value the curated `cost` was measured against (#267). @@ -352,10 +383,12 @@ def _scalar_driver_shifted(definition, expanded, list_entries): defaults = definition.get("variables") or {} effective = expanded.get("variables") or {} for name in _declared_drivers(definition): - if name in list_entries: - continue default_value = defaults.get(name) if isinstance(default_value, list): + if _list_entry_field_shifted(default_value, effective.get(name)): + return True + continue + if name in list_entries: continue if _driver_comparable(effective.get(name)) != _driver_comparable(default_value): return True @@ -365,7 +398,7 @@ def _scalar_driver_shifted(definition, expanded, list_entries): def _own_price(definition, expanded, list_entries, device, measured_entries): """The workflow's own price, reset to unknown when a scalar driver moved.""" own = _price(definition.get("cost"), device, list_entries, measured_entries or {}) - if own["basis"] == CATALOG and _scalar_driver_shifted( + if own["basis"] in (CATALOG, OTHER_DEVICE) and _scalar_driver_shifted( definition, expanded, list_entries ): # A scalar cost_driver (H3's num_frames, say) moved away from the @@ -434,7 +467,7 @@ def _child_catalog_price(child_definition, child_cost, step_arguments, device): child_list_entries = _list_entries(child_definition, child_expanded) child = _price(child_cost, device, child_list_entries, child_measured_entries) if ( - child["basis"] == CATALOG + child["basis"] in (CATALOG, OTHER_DEVICE) and child_definition is not None and _scalar_driver_shifted(child_definition, child_expanded, child_list_entries) ): @@ -458,6 +491,7 @@ def __init__(self, minutes, partial, unpriced): self.all_observed = True self.runs = [] self.measured_on = set() + self.bases = [] def add(self, path, child, observed): self.had_child = True @@ -469,7 +503,9 @@ def add(self, path, child, observed): if child["minutes"] is None: self.partial = True self.unpriced.append(path) - elif self.minutes is not None: + return + self.bases.append((child["basis"], child.get("measured_on"))) + if self.minutes is not None: self.minutes += child["minutes"] else: self.minutes = child["minutes"] @@ -509,6 +545,20 @@ def _rolled_up_estimate(own, totals, device, cached_steps, total_steps): top_measured_on = ( next(iter(totals.measured_on)) if len(totals.measured_on) == 1 else None ) + if top_basis == UNKNOWN and minutes is not None and totals.bases: + # A number must not carry the basis a caller reads as "no number": + # the parent has no cost of its own, but its priced children do, so + # the sum takes their basis (#593) + kinds = {basis for basis, _ in totals.bases} + if OTHER_DEVICE in kinds: + top_basis = OTHER_DEVICE + elif kinds == {CATALOG}: + top_basis = CATALOG + else: + top_basis = DERIVED + devices = {on for _, on in totals.bases} + if len(devices) == 1: + top_measured_on = next(iter(devices)) rounded = round(minutes, 1) if minutes is not None else None result = { "minutes": rounded, @@ -696,18 +746,22 @@ def _price(cost, device, list_entries, measured_entries): basis = OTHER_DEVICE minutes = float(chosen.get("minutes", 0)) per = chosen.get("per_entry") - if ( - basis == CATALOG - and isinstance(per, dict) - and per.get("variable") in list_entries - ): + # Another device's figure is re-priced for the list length too (#589): it + # stays `other_device` (still not a measurement here), but a 32-entry + # for_each must not be quoted at the one-clip figure + if isinstance(per, dict) and per.get("variable") in list_entries: count = list_entries[per["variable"]] each = float(per.get("minutes", 0)) measured_with = int(per.get("entries", 0)) minutes = max(0.0, (minutes - each * measured_with) + each * count) - basis = PER_ENTRY - elif basis == CATALOG: - minutes, basis = _repriced(minutes, list_entries, measured_entries) + if basis == CATALOG: + basis = PER_ENTRY + else: + minutes, repriced_basis = _repriced(minutes, list_entries, measured_entries) + if repriced_basis == UNKNOWN: + return {"minutes": None, "basis": UNKNOWN, "measured_on": None} + if basis == CATALOG: + basis = repriced_basis return {"minutes": minutes, "basis": basis, "measured_on": chosen.get("name")} diff --git a/dw/server/routes/jobs.py b/dw/server/routes/jobs.py index 3eff0476..2ca9f43a 100644 --- a/dw/server/routes/jobs.py +++ b/dw/server/routes/jobs.py @@ -455,11 +455,13 @@ def export_job_route( absolute_zip_url = absolute_served_url(zip_path, ws) if absolute_zip_url is not None: body["absolute_zip_url"] = absolute_zip_url - # Same rule get_server_info's field states (#353): whether the zip - # URL above needs a bearer token an MCP-only agent has no way to - # attach itself, which is what tells the caller whether to fetch it - # or hand it to the person. - body["auth_required"] = bool(state.api_token) + # Whether the zip URL above needs a bearer token (#353, #592) - a + # property of the route, not of whether this server has a token. The + # auth middleware covers /api/ and /mcp only, and /exports/*.zip sits + # with /outputs and /inputs on dw/server/routes/files.py's ungated + # list, so it is false. export_job's `next` keeps its hand-it-over + # branch for the day /exports is gated. + body["auth_required"] = False for key, name in ( ("workflow", "workflow.json"), ("manifest", "manifest.json"), diff --git a/dw/tasks/joins.py b/dw/tasks/joins.py index 3fca16c1..cf9660f0 100644 --- a/dw/tasks/joins.py +++ b/dw/tasks/joins.py @@ -94,6 +94,7 @@ def reconcile_sample_rates(command, videos, names, waveforms, sample_rate=None): for video, waveform in zip(videos, waveforms) if waveform is not None ] + pinned = bool(sample_rate) sample_rate = sample_rate or (max(set(rates)) if rates else None) if rates and len(set(rates)) == 1 and rates[0] != sample_rate: # The inputs agree and the caller pinned another rate: converting @@ -104,6 +105,15 @@ def reconcile_sample_rates(command, videos, names, waveforms, sample_rate=None): command=command, sample_rate=sample_rate, ) + elif rates and pinned and any(rate != sample_rate for rate in rates): + # Mixed inputs converted to the rate the caller (or the template) + # pinned: expected, and 'pass sample_rate' advice would be wrong (#594) + emit_log( + f"{command}: resampling tracks at {sorted(set(rates))} Hz to the " + f"requested {sample_rate} Hz", + command=command, + sample_rate=sample_rate, + ) elif rates and any(rate != sample_rate for rate in rates): # emit_warning rather than logger.warning: resampling every track is # an audio decision made on the caller's behalf, and a caller reading diff --git a/dw_mcp/exports.py b/dw_mcp/exports.py index 07896a91..f7b14273 100644 --- a/dw_mcp/exports.py +++ b/dw_mcp/exports.py @@ -3,9 +3,9 @@ The one thing this module has to keep saying: the directory it makes is on the machine running dw.serve, which over a `dw.serve --mcp` endpoint is the GPU box and not where the agent is. The zip URL is the way to it from -anywhere else - but when the server requires a bearer token (#353), that URL -is for the person to open, not for this agent to fetch on their behalf; see -`export_job`'s `auth_required` / `open_url`. +anywhere else, and it needs no token: /exports/*.zip is ungated like +/outputs (#592), so `auth_required` is false and the agent fetches it itself. +The hand-it-to-the-person branch (#353) stays for the day it is gated. """ from dw_mcp.client import api_path @@ -22,10 +22,10 @@ def export_job(client, job_id, overwrite=False): here. The directory is on the server machine, not this one. `auth_required` says whether the zip needs this server's bearer token - to open - a token this agent has no way to attach to a browser or hand - to someone else's tooling. When it is true, `open_url` is for the - *person* to open, not for this agent to fetch: hand it to them (see - `next`). When it is false, `open_url` may be fetched directly. It is + to open. The server reports false - the zip route is ungated, token or + not (#592) - and `next` says to fetch `open_url` and unpack it. True + is a forward guard: a gated zip is for the *person* to open, since + this agent cannot attach the token to their browser. `open_url` is `absolute_zip_url` when the server has one configured (`DW_PUBLIC_URL` / the `public_url` setting), else the relative `zip_url`.""" body = client.post_json( @@ -58,11 +58,14 @@ def export_job(client, job_id, overwrite=False): open_url = absolute_zip_url or zip_url next_text = ( "The directory is on the server. To give the user the files, " - "fetch open_url and unpack it into exports/ under the session's " - "working directory - it is the user's deliverable, not a " - "temporary file, so not a scratch or temp directory. The archive " - "already unpacks into one folder named after the job id; do not " - "create that folder first or the id is doubled in the path. " + "fetch open_url with whatever HTTP you have - it needs no token; " + "prefix a relative one with the address you reach this server at " + "- and unpack it into exports/ under the session's working " + "directory - it is the user's deliverable, not a temporary file, " + "so not a scratch or temp directory. The archive already unpacks " + "into one folder named after the job id; do not create that " + "folder first or the id is doubled in the path. Only if you " + "cannot make HTTP requests, give the user open_url to open. " "workflow.json, manifest.json and job.json are inside it - they " "are not repeated here; get_job_workflow and get_job serve them " "individually." diff --git a/dw_mcp/tools_jobs.py b/dw_mcp/tools_jobs.py index 72610cb6..3aae9cbe 100644 --- a/dw_mcp/tools_jobs.py +++ b/dw_mcp/tools_jobs.py @@ -216,16 +216,15 @@ def export_job(self, job_id: str, overwrite: bool = False) -> dict: repeated here - get_job_workflow and get_job serve them individually. THE DIRECTORY IS ON THE MACHINE RUNNING THE SERVER, not on yours. - `auth_required` says whether opening the zip needs this server's - bearer token, a token you cannot attach to someone else's browser - or tooling. When it is false, fetch open_url yourself and unpack - it into exports/ under the session's working directory - it is - the user's deliverable, not a temp file; the archive already - unpacks into one folder named after the job id, so do not create that folder first. - When it is true, do NOT fetch it: hand open_url to the person and let them open it - (`next` says whether it is already absolute or needs the server's - address told to them). Individual results stay reachable inline - via get_output_image/get_output_audio/get_output_frames either - way. Refuses a job that is still running; refuses an existing + open_url is the zip and needs no token: fetch it with any HTTP + you have (prefix a relative one with the server's address) and + unpack it into exports/ under the session's working directory - + the user's deliverable, not a temp file; it already unpacks into + one folder named after the job id, so do not create that folder first. + Only without HTTP, give the user open_url to open. If `auth_required` + is ever true, the zip is gated: hand it over (see `next`). + Individual results stay reachable inline via + get_output_image/get_output_audio/get_output_frames + without the zip. Refuses a job that is still running; refuses an existing export unless overwrite=true.""" return exports.export_job(self.client, job_id, overwrite=overwrite) diff --git a/loras/ltx-2.5/cinemagraph.json b/loras/ltx-2.5/cinemagraph.json new file mode 100644 index 00000000..829d204e --- /dev/null +++ b/loras/ltx-2.5/cinemagraph.json @@ -0,0 +1,26 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-LoRA-Cinemagraph", + "weight_name": "ltx-2.5-22b-lora-cinemagraph-0.9.safetensors", + "revision": "8e5e88efcf59ff0efa6f9dcf7164c7a7d72c4022", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "Plain image-to-video LoRA that animates a still into a looping cinemagraph - one element moves (water, light), the rest stays frozen. Prompt for a locked-off camera, what stays frozen, what moves and a seamless loop. A 0.9 preview weight trained at 512x704, 25 frames on static-camera footage with no moving people; the card recommends STG scale 1.0 on block 29", + "use_when": "Animate a still image into a looping cinemagraph", + "trigger": "CINEMAGRAPH_MOTION", + "scale": { + "default": 1.0, + "range": [ + 1.0, + 1.2 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "cinemagraph", + "loop", + "image-to-video" + ] +} diff --git a/loras/ltx-2.5/ic-alpha-gen.json b/loras/ltx-2.5/ic-alpha-gen.json new file mode 100644 index 00000000..13b39df5 --- /dev/null +++ b/loras/ltx-2.5/ic-alpha-gen.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Alpha-Gen", + "weight_name": "ltx-2.5-22b-ic-lora-alpha-gen-0.9.safetensors", + "revision": "6184df14b1b560bf9b6447d3ee126bd7e4a88513", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that generates an alpha matte from an RGB clip - no green screen, mask or prompt; handles hair, smoke, fire and glass. The RGB clip is the reference video; the prompt must be empty. Output is a grayscale matte video (white foreground) aligned frame-for-frame, to be combined with the source RGB for compositing. A 0.9 preview weight; up to 145 frames (8n+1) at up to 1920x1088, stage-1 only at native size", + "use_when": "Pull a matte to cut a subject out of a clip for compositing", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "matte", + "alpha", + "compositing" + ] +} diff --git a/loras/ltx-2.5/ic-clean-plate.json b/loras/ltx-2.5/ic-clean-plate.json new file mode 100644 index 00000000..dc253c45 --- /dev/null +++ b/loras/ltx-2.5/ic-clean-plate.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Clean-Plate", + "weight_name": "ltx-2.5-22b-ic-lora-clean-plate-1.0.safetensors", + "revision": "5403b2b77e9994e0ae2e549e9933d06c9e8df26f", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that removes people and vehicles from a clip and keeps the static background - a clean plate. The original clip is the reference video; no mask. Describe the empty scene rather than giving removal commands, and name what to remove in the negative prompt too. Trained at 1024x576 / 576x1024, 49 frames @ 25 fps. The reference needs an audio track (add a silent one); works for partial or edge occluders, not a subject filling the frame", + "use_when": "Remove passers-by or vehicles to get an empty background plate", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "remove", + "clean-plate", + "edit" + ] +} diff --git a/loras/ltx-2.5/ic-colorization.json b/loras/ltx-2.5/ic-colorization.json new file mode 100644 index 00000000..87f29a1f --- /dev/null +++ b/loras/ltx-2.5/ic-colorization.json @@ -0,0 +1,26 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Colorization", + "weight_name": "ltx-2.5-22b-ic-lora-colorization-0.9.safetensors", + "revision": "a2b3d526bfe3b251ba61cdfbf93deb51eb607198", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that restores natural colour to grayscale or desaturated footage, keeping subject, composition and motion. The monochrome clip is the reference video at 1x output size. Prompt in the trained dual-panel form: 'Reference shows ... Edited shows the same scene with natural colors restored. COLORIZE ... only color information differs.' A 0.9 preview weight trained at 960x544, 121 frames @ 24 fps; far above that size weakens it; lower toward 0.8 if it oversaturates", + "use_when": "Colorize a black-and-white or washed-out clip", + "trigger": "COLORIZE", + "scale": { + "default": 1.0, + "range": [ + 0.8, + 1.0 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "restore", + "color", + "edit" + ] +} diff --git a/loras/ltx-2.5/ic-day-to-night.json b/loras/ltx-2.5/ic-day-to-night.json new file mode 100644 index 00000000..f3c5b65f --- /dev/null +++ b/loras/ltx-2.5/ic-day-to-night.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Day-To-Night", + "weight_name": "ltx-2.5-22b-ic-lora-day-to-night-0.9.safetensors", + "revision": "52aac758c153f9a74a6f6fe08cf92d2167f3c184", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that relights a daytime clip as night, keeping the composition, camera move and subject motion. The day clip is the reference video at 1x output size; the prompt steers the night look (moonlight, shadows, colour temperature). A 0.9 preview weight trained at 768x448 / 448x768, 97 frames @ 24 fps; re-encode the reference to 24 fps first", + "use_when": "Turn a daytime shot into the same shot at night", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "relight", + "night", + "edit" + ] +} diff --git a/loras/ltx-2.5/ic-layout-to-render.json b/loras/ltx-2.5/ic-layout-to-render.json new file mode 100644 index 00000000..cba9ad36 --- /dev/null +++ b/loras/ltx-2.5/ic-layout-to-render.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Layout-To-Render", + "weight_name": "ltx-2.5-22b-ic-lora-layout-to-render-1.0.safetensors", + "revision": "349625f9bad639dc41f1c70d2c63f75a02a09c59", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that turns a 3D viewport playblast - clay render or blocky shapes from Blender or Unreal - into a finished shot, keeping the camera move and object placement. The playblast is the reference video and a styled first-frame image (art direction, not the raw clay frame) is a keyframe condition. One or two sentences describing the finished shot; never mention clay, 3D or the tool. 960x544 stage 1, 1920x1088 stage 2, 24 fps; width and height divisible by 64", + "use_when": "Render a 3D blockout or previs animation as a finished shot", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "previs", + "layout", + "render" + ] +} diff --git a/loras/ltx-2.5/ic-refine-details.json b/loras/ltx-2.5/ic-refine-details.json new file mode 100644 index 00000000..f01021e1 --- /dev/null +++ b/loras/ltx-2.5/ic-refine-details.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details", + "weight_name": "ltx-2.5-22b-ic-lora-refine-details-1.0.safetensors", + "revision": "4912c478b35dd96b6cdcc5f7c7ff9d72c2c57929", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that rebuilds fine detail - texture, edges, grain - in a soft, compressed or upscaled clip. The source, lanczos-upscaled to the output canvas, is the reference video. Prompts stay generic and about rendering (sharpness, texture, light), not subjects. Trained on 97-frame 1024x576 / 576x1024 tiles @ 24 fps; the card runs it tiled (50% overlap) at every output size, which dw does not do, so only a clip near the tile size is run as trained", + "use_when": "Add native-looking detail to a soft or upscaled clip", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "refine", + "detail", + "upscale" + ] +} diff --git a/loras/ltx-2.5/ic-restore.json b/loras/ltx-2.5/ic-restore.json new file mode 100644 index 00000000..d0186afa --- /dev/null +++ b/loras/ltx-2.5/ic-restore.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Restore", + "weight_name": "ltx-2.5-22b-ic-lora-restore-1.0.safetensors", + "revision": "ed04134f5512bca3f7f36118d36737da09569c35", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that restores archive footage - compression damage, tape and sepia casts, flicker, dirt, scratches - and colorizes it as a clean modern capture. The archive clip is the reference video; describe period, place, lighting and materials, and use the negative prompt against modern objects. Trained at 960x544, 49 or 97 frames @ 24 fps on the distilled transformer; deinterlace first. Scale is a switch: below 0.7 it does nothing. The card runs it tiled at 960x544, which dw does not do. Broader than restore-deblur / restore-decompression, which each invert one defect", + "use_when": "Restore and colorize old archive or tape footage", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "restore", + "archive", + "color" + ] +} diff --git a/loras/ltx-2.5/ic-sdr-to-hdr.json b/loras/ltx-2.5/ic-sdr-to-hdr.json new file mode 100644 index 00000000..72f0471c --- /dev/null +++ b/loras/ltx-2.5/ic-sdr-to-hdr.json @@ -0,0 +1,25 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-SDR-To-HDR", + "weight_name": "ltx-2.5-22b-ic-lora-sdr-to-hdr-1.0.safetensors", + "revision": "c7578ebddb54f4210e534e4d8cad698a5a520692", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "Converts 8-bit SDR video to scene-linear HDR; needs LTX2HDRPipeline with precomputed connector embeddings and an HDR export that dw does not wire", + "use_when": "Do not use under dw", + "scale": { + "default": 1.0 + }, + "status": "rejected", + "evidence": [ + { + "note": "Needs diffusers' LTX2HDRPipeline fed the repo's second file (ltx-2.5-22b-ic-lora-sdr-to-hdr-scene-emb.safetensors) as connector_video_embeds / connector_audio_embeds in place of a prompt, and writes linear HDR (ACEScg EXR, HLG BT.2020 HEVC); dw has no step that loads those tensors or exports HDR" + } + ], + "license": "other", + "tags": [ + "rejected", + "hdr" + ] +} diff --git a/loras/ltx-2.5/ic-water-simulation.json b/loras/ltx-2.5/ic-water-simulation.json new file mode 100644 index 00000000..a53b4a47 --- /dev/null +++ b/loras/ltx-2.5/ic-water-simulation.json @@ -0,0 +1,26 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-IC-LoRA-Water-Simulation", + "weight_name": "ltx-2.5-22b-ic-lora-water-simulation-0.9.safetensors", + "revision": "f698d6801a673c636a2c5165ffa07737d7964691", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "IC-LoRA that adds water - rain, rivers, floods, splashes - to a dry clip while keeping identity, clothing, pose and framing. The dry clip is the reference video at 1x output size. Prompt in the trained dual-panel form: 'Reference shows ... Edited shows the same scene with water added. ADD WATER ... Subject identity, clothing, framing, and background geometry are identical to the reference.' A 0.9 preview weight trained at 960x544, 97-121 frames @ 24 fps; scale above 1.0 for more water", + "use_when": "Add rain, flooding or splashing water to an existing clip", + "trigger": "ADD WATER", + "scale": { + "default": 1.0, + "range": [ + 0.9, + 1.45 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "water", + "vfx", + "edit" + ] +} diff --git a/loras/ltx-2.5/slow-motion-control.json b/loras/ltx-2.5/slow-motion-control.json new file mode 100644 index 00000000..c99e19cf --- /dev/null +++ b/loras/ltx-2.5/slow-motion-control.json @@ -0,0 +1,21 @@ +{ + "model_name": "Lightricks/LTX-2.5-22b-LoRA-Slow-Motion-Control", + "weight_name": "ltx-2.5-22b-lora-slow-motion-control-1.0.safetensors", + "revision": "c7a014b985023874de6000fc791b959dbd46a0a8", + "base_models": [ + "Lightricks/LTX-2.5-Diffusers" + ], + "workflow": null, + "description": "Plain LoRA for high-speed-camera slow motion without smearing. Speed is set by frame rate alone: the pipeline's frame_rate is the motion fps (24 / speed, speed 1.0 down to 0.025) while the file is saved at 24 fps. Unverified under dw: frame_rate also sets the audio length, so the soundtrack will not match. Regular captions, identical across speeds, no speed wording. Trained at 1920x1088, 121 frames; width and height divisible by 64", + "use_when": "Generate a clip in true slow motion", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "slow-motion", + "speed", + "motion" + ] +} diff --git a/loras/minimax-h3/16bit-pixel.json b/loras/minimax-h3/16bit-pixel.json new file mode 100644 index 00000000..70e57335 --- /dev/null +++ b/loras/minimax-h3/16bit-pixel.json @@ -0,0 +1,24 @@ +{ + "model_name": "KennethFal/16bit-pixel-lora-minimax-h3", + "weight_name": "16bit-pixel.safetensors", + "revision": "ae7d92788e8ff760e75639d5a413f5f460b29d37", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "SNES-era 16-bit pixel-art video. The trigger is not published (fal applies it server-side), so the prompt has to find it by experiment; scale not stated. fal-trainer format, attention only - the proven realism-people format", + "use_when": "Render a clip as 16-bit pixel art", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "style", + "pixel-art", + "retro" + ] +} diff --git a/loras/minimax-h3/360-orbit-loop.json b/loras/minimax-h3/360-orbit-loop.json new file mode 100644 index 00000000..8a866727 --- /dev/null +++ b/loras/minimax-h3/360-orbit-loop.json @@ -0,0 +1,23 @@ +{ + "model_name": "pablodawson/MiniMax-H3-360-Orbit-LoRA", + "weight_name": "minimax_h3_flf2v_lora_v1.safetensors", + "revision": "5ddbc2dbbe95edbbdaf5017c3e934b1d01791697", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": "fl2va", + "description": "Frozen-time 360-degree camera orbit; with the same image as first and last frame the clip loops seamlessly. Use the card's instance prompt verbatim ('One frozen instant. Only the camera moves. In a continuous 360 orbit. ...'). Trained at 768x768, 73 frames, 28 steps on the adaln-pruned FL2VA base; no adaln keys, so it loads on the full one, quality there unproven", + "use_when": "A looping frozen-moment orbit around a still image", + "trigger": "One frozen instant. Only the camera moves. In a continuous 360 orbit.", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "camera", + "orbit", + "360", + "loop" + ] +} diff --git a/loras/minimax-h3/anime-motion.json b/loras/minimax-h3/anime-motion.json new file mode 100644 index 00000000..242759db --- /dev/null +++ b/loras/minimax-h3/anime-motion.json @@ -0,0 +1,22 @@ +{ + "model_name": "prithivMLmods/MiniMax-H3-I2V-Anime-Motion-LoRA", + "weight_name": "MiniMax-H3-I2V-Anime-Motion-LoRA-1400.safetensors", + "revision": "ca20427ab7488b80fc3a14363561431e27a8f974", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": "fl2va", + "description": "Subtle looping anime motion for animating an anime still from its first frame. DiffSynth format covering blocks 25-49 only, r16; experimental, scale not stated", + "use_when": "Animate an anime still with subtle motion", + "trigger": "Anime-Motion:", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "anime", + "motion", + "image-to-video" + ] +} diff --git a/loras/minimax-h3/better-human-motion.json b/loras/minimax-h3/better-human-motion.json new file mode 100644 index 00000000..ba279c25 --- /dev/null +++ b/loras/minimax-h3/better-human-motion.json @@ -0,0 +1,28 @@ +{ + "model_name": "vpakarinen/better-human-motion-h3-lora", + "weight_name": "better_motion_h3_lora_v1_500.safetensors", + "revision": "659eba44ffc9f1c7563ebd2a848a2447432e89c5", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "More realistic human motion and temporal consistency. No trigger. ai-toolkit format, attention and MLP, r32; 500 steps, thin card", + "use_when": "People moving unnaturally or inconsistently", + "scale": { + "default": 0.6, + "range": [ + 0.4, + 0.8 + ] + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "motion", + "people", + "quality" + ] +} diff --git a/loras/minimax-h3/facial-realism-closeup.json b/loras/minimax-h3/facial-realism-closeup.json new file mode 100644 index 00000000..4fdd3f5f --- /dev/null +++ b/loras/minimax-h3/facial-realism-closeup.json @@ -0,0 +1,25 @@ +{ + "model_name": "prithivMLmods/MiniMax-H3-Facial-Realism-CloseUp", + "weight_name": "minimax-h3-facial-realism-closeup-cp2000.safetensors", + "revision": "c89656bd0cb71873133dd2c7a925f306a21c633c", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "Close-up facial micro-expression realism. DiffSynth format covering blocks 25-49 only, r16 - a format not yet run here. Overlaps realism-people; the card's samples ran a 4-6 step turbo, scale not stated", + "use_when": "A close-up of a face that needs realistic micro-expressions", + "trigger": "Facial Realism:", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "realism", + "face", + "close-up" + ] +} diff --git a/loras/minimax-h3/fasth3-4step.json b/loras/minimax-h3/fasth3-4step.json new file mode 100644 index 00000000..9da086d2 --- /dev/null +++ b/loras/minimax-h3/fasth3-4step.json @@ -0,0 +1,25 @@ +{ + "model_name": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA", + "weight_name": "dense-datafree/adapter_model.safetensors", + "revision": "f509e629374cac104e7f62daecce6d1488a3041d", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": null, + "description": "FastH3 4-step DMD2 distill; carries full-weight delta keys diffusers' converter refuses", + "use_when": "Do not use under dw", + "scale": { + "default": 1.0 + }, + "status": "rejected", + "evidence": [ + { + "note": "diffusers' MiniMax-H3 converter rejects its 85 .diff/.diff_b full-weight keys (state_dict should be empty) - the same failure as fasth3-preview; the repo's vsa-* files add to_gate_compress.set_weight keys and need FastVideo's VSA kernel" + } + ], + "license": "other", + "tags": [ + "rejected", + "speed" + ] +} diff --git a/loras/minimax-h3/instantx-turbo-4step.json b/loras/minimax-h3/instantx-turbo-4step.json new file mode 100644 index 00000000..da6ae6b7 --- /dev/null +++ b/loras/minimax-h3/instantx-turbo-4step.json @@ -0,0 +1,24 @@ +{ + "model_name": "InstantX/MiniMax-H3-Turbo-Lora-Diffusers", + "weight_name": "minimax_h3_turbo_4step_ckpt500_diffusers.safetensors", + "revision": "d290a021e3c927b9396f3d20fc8de4a567c603dd", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "larryvrh's 4-step turbo in native diffusers format (transformer.-prefixed, adaln and norm_out included). num_inference_steps 5 (4 evaluations). The card calls it an under-trained prototype; an alternative to the proven lightx2v turbo-4step-keyframe, not a replacement. Needs the unpruned transformer", + "use_when": "A 4-step keyframe render, as an alternative to turbo-4step-keyframe", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "speed", + "turbo", + "distill" + ] +} diff --git a/loras/minimax-h3/motion-repair.json b/loras/minimax-h3/motion-repair.json new file mode 100644 index 00000000..e9700ef7 --- /dev/null +++ b/loras/minimax-h3/motion-repair.json @@ -0,0 +1,29 @@ +{ + "model_name": "JOKER141/MiniMax-H3-General-Motion-Continuity-Repair", + "weight_name": "Motion_Repair_V2.safetensors", + "revision": "8a5126beb17b7de5642ee056ff5a7b60ae0915c7", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "Repairs broken fast motion - flips, spins, slow-motion drift, missing transitions. About 0.9 alone, 0.5-0.7 stacked with an action LoRA such as wushu-action. The card states no partition or license; tagged for the keyframe transformer since it is ai-toolkit's default and the wushu card pairs it there", + "use_when": "Fast action whose motion breaks or skips", + "trigger": "bunny_crisp_motion", + "scale": { + "default": 0.9, + "range": [ + 0.5, + 0.9 + ] + }, + "status": "trial", + "license": "unknown", + "tags": [ + "motion", + "repair", + "action" + ] +} diff --git a/loras/minimax-h3/natural-face-speech.json b/loras/minimax-h3/natural-face-speech.json new file mode 100644 index 00000000..64169606 --- /dev/null +++ b/loras/minimax-h3/natural-face-speech.json @@ -0,0 +1,29 @@ +{ + "model_name": "vpakarinen/natural-face-speech-h3-lora", + "weight_name": "natural_face_speech_h3_lora_v2_500.safetensors", + "revision": "3349b1f21fa5ca597d364aaa038c29dcec765b85", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "Facial muscle dynamics and clear lip-synced English speech. No trigger; the card's prompts follow '... speaking into a mic ... saying: '...''. ai-toolkit format, attention and MLP, r32, v2 at 500 steps", + "use_when": "A character speaking a line on camera", + "scale": { + "default": 0.7, + "range": [ + 0.6, + 0.8 + ] + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "speech", + "dialogue", + "face", + "lip-sync" + ] +} diff --git a/loras/minimax-h3/orb360.json b/loras/minimax-h3/orb360.json new file mode 100644 index 00000000..95529f7c --- /dev/null +++ b/loras/minimax-h3/orb360.json @@ -0,0 +1,23 @@ +{ + "model_name": "MATLOWAI/MiniMax-H3-ORB360-CardSpin", + "weight_name": "minimax_h3_orb360_step1500.safetensors", + "revision": "de8aa4db6f9a0d949855f3dc62948ffc691aa6ce", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": "ref2va", + "description": "Clockwise 360-degree camera orbit around a frozen subject from 1-4 reference photos. Use the structured Ref2VA prompts in the repo's prompts/*.txt verbatim - their timestamps assume 124 frames; about 20 steps. Trained on the adaln-pruned base but carries no adaln keys, so it loads on the full one; quality there is unproven. The repo's cardspin_v2 file (trigger ORB360_CARDSPIN) spins the photo like a card instead", + "use_when": "Orbit the camera all the way round a subject from reference photos", + "trigger": "ORB360_CW", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "camera", + "orbit", + "360", + "reference" + ] +} diff --git a/loras/minimax-h3/raven-streaming.json b/loras/minimax-h3/raven-streaming.json new file mode 100644 index 00000000..5e267958 --- /dev/null +++ b/loras/minimax-h3/raven-streaming.json @@ -0,0 +1,25 @@ +{ + "model_name": "mvp-lab/MiniMax-H3-RAVEN-Streaming-LoRA", + "weight_name": "minimax_h3_raven_streaming_lora_4nfe_preview.safetensors", + "revision": "90f30b11639e73ce0e3f2ab6ed70d1a133a66caf", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": null, + "description": "Causal streaming generation; its key prefix matches no format the converter knows, so it loads nothing", + "use_when": "Do not use under dw", + "scale": { + "default": 1.0 + }, + "status": "rejected", + "evidence": [ + { + "note": "Keys are prefixed base_model.model.dit., which matches no format diffusers' MiniMax-H3 converter knows: the load warns and applies nothing. It also needs RAVEN's causal chunked sampler" + } + ], + "license": "other", + "tags": [ + "rejected", + "streaming" + ] +} diff --git a/loras/minimax-h3/studio-1939.json b/loras/minimax-h3/studio-1939.json new file mode 100644 index 00000000..1dcf0f4f --- /dev/null +++ b/loras/minimax-h3/studio-1939.json @@ -0,0 +1,29 @@ +{ + "model_name": "lovis93/studio-1939-old-animation-lora-minimax-h3", + "weight_name": "studio1939-strong.safetensors", + "revision": "19214d4c3989de6caca673d534d0d4b16b73b0f7", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "1930s hand-painted cel animation look (the full-cel r64 file; the repo's studio1939-light is a softer painterly r16). fal-trainer format, attention only - the same format as the proven realism-people entry. Turn prompt expansion off so the trigger stays first", + "use_when": "Render a clip as 1930s hand-painted cel animation", + "trigger": "gulliv3r, ", + "scale": { + "default": 1.0, + "range": [ + 0.4, + 1.0 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "style", + "animation", + "vintage" + ] +} diff --git a/loras/minimax-h3/taomate.json b/loras/minimax-h3/taomate.json new file mode 100644 index 00000000..1e725849 --- /dev/null +++ b/loras/minimax-h3/taomate.json @@ -0,0 +1,26 @@ +{ + "model_name": "TaoLiveAIGC/TaoMate-H3", + "weight_name": "adapter_model.safetensors", + "revision": "7d7a51f3e63972138882ec2f0e4d1d47728110cf", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": null, + "description": "3-step streaming text-to-video for the TaoMate runtime; lowercase lora_a/lora_b keys the converter refuses", + "use_when": "Do not use under dw", + "scale": { + "default": 1.0 + }, + "status": "rejected", + "evidence": [ + { + "note": "Lowercase .lora_a/.lora_b keys are left over by diffusers' MiniMax-H3 converter (state_dict should be empty); it also needs the TaoMate streaming runtime" + } + ], + "license": "other", + "tags": [ + "rejected", + "streaming", + "speed" + ] +} diff --git a/loras/minimax-h3/vh5tape-vhs.json b/loras/minimax-h3/vh5tape-vhs.json new file mode 100644 index 00000000..f533c127 --- /dev/null +++ b/loras/minimax-h3/vh5tape-vhs.json @@ -0,0 +1,30 @@ +{ + "model_name": "KennethFal/vh5tape-vhs-lora-minimax-h3", + "weight_name": "vh5tape.safetensors", + "revision": "d25a4ab75aba899aa508220321525c8465ac11af", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "Worn 1980s VHS tape look, picture and soundtrack both. The trigger goes first, then a damage phrase from the card (light / medium / heavy). fal-trainer format, attention only - the proven realism-people format. vh5tape-comfyui.safetensors in the same repo is the same weights with alpha scalars", + "use_when": "Make a clip look and sound like a worn VHS tape", + "trigger": "vh5tape", + "scale": { + "default": 1.0, + "range": [ + 0.7, + 1.5 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "style", + "vhs", + "vintage", + "audio" + ] +} diff --git a/loras/minimax-h3/wushu-action-reference.json b/loras/minimax-h3/wushu-action-reference.json new file mode 100644 index 00000000..f321d3bb --- /dev/null +++ b/loras/minimax-h3/wushu-action-reference.json @@ -0,0 +1,27 @@ +{ + "model_name": "Jojocodex/wushu-action-v7-minimax-h3-fl2va-ref2va-lora", + "weight_name": "wushu_h3_r2v_v7_2000.safetensors", + "revision": "1aab496db1c70df44b3edb2cf4e90d8ae3c5b26d", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": "ref2va", + "description": "The ref2va file of the wushu fight-choreography LoRA - the same moves with the subject held from references. The trigger must start the prompt. diffusion_model. format, r32", + "use_when": "A martial-arts fight scene with a referenced subject", + "trigger": "wushu_action", + "scale": { + "default": 1.0, + "range": [ + 0.6, + 1.0 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "action", + "fight", + "martial-arts", + "reference" + ] +} diff --git a/loras/minimax-h3/wushu-action.json b/loras/minimax-h3/wushu-action.json new file mode 100644 index 00000000..056faee2 --- /dev/null +++ b/loras/minimax-h3/wushu-action.json @@ -0,0 +1,30 @@ +{ + "model_name": "Jojocodex/wushu-action-v7-minimax-h3-fl2va-ref2va-lora", + "weight_name": "wushu_action_v7_fl2va_aitoolkit_adaln_full-int8convrot_bf16te_2000step.safetensors", + "revision": "1aab496db1c70df44b3edb2cf4e90d8ae3c5b26d", + "base_models": [ + "MiniMaxAI/MiniMax-H3" + ], + "workflow": [ + "t2va", + "fl2va" + ], + "description": "Wushu / martial-arts fight choreography; the card's Chinese move vocabulary steers the moves. The trigger must start the prompt. Lower toward 0.6-0.8 if bodies or weapons break; soft at low step counts. Stacks with motion-repair. musubi format, r32", + "use_when": "A martial-arts fight scene", + "trigger": "wushu_action", + "scale": { + "default": 1.0, + "range": [ + 0.6, + 1.0 + ] + }, + "status": "trial", + "license": "other", + "tags": [ + "action", + "fight", + "martial-arts", + "motion" + ] +} diff --git a/loras/qwen-image-2.1/anime-consistency.json b/loras/qwen-image-2.1/anime-consistency.json new file mode 100644 index 00000000..e1e8e88c --- /dev/null +++ b/loras/qwen-image-2.1/anime-consistency.json @@ -0,0 +1,25 @@ +{ + "model_name": "WarmBloodAban/Qwen-Image-2.1-LoRAs", + "weight_name": "Qwen2.1_Anime_consistency.safetensors", + "revision": "c4ab5473bfdf585fc19cfd1e280f79d2b0c79947", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: keeps an anime character consistent across edits. The author marks it an experimental parameter test; base step count", + "use_when": "Edit an anime character without it drifting off-model", + "scale": { + "default": 0.7, + "range": [ + 0.6, + 0.8 + ] + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "edit", + "anime", + "consistency" + ] +} diff --git a/loras/qwen-image-2.1/clay-sculpture.json b/loras/qwen-image-2.1/clay-sculpture.json new file mode 100644 index 00000000..f7c06d2c --- /dev/null +++ b/loras/qwen-image-2.1/clay-sculpture.json @@ -0,0 +1,22 @@ +{ + "model_name": "prithivMLmods/Qwen-Image-2.1-Clay-Sculpture-Style", + "weight_name": "Qwen-Image-2.1-Clay-Sculpture-Style-3000.safetensors", + "revision": "0168d4806d9c5f0fc3067ed24c56df8fa2a8dca2", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: restyles an image as a clay sculpture. The latest of the repo's 1000/2000/3000 checkpoints; the card names none", + "use_when": "Restyle an image as a clay sculpture", + "trigger": "Transform the image into a clay sculpture style", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "style", + "clay" + ] +} diff --git a/loras/qwen-image-2.1/doodle-in.json b/loras/qwen-image-2.1/doodle-in.json new file mode 100644 index 00000000..25d91259 --- /dev/null +++ b/loras/qwen-image-2.1/doodle-in.json @@ -0,0 +1,22 @@ +{ + "model_name": "ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA", + "weight_name": "doodle_in_lora_qwen21_gate_up_split.safetensors", + "revision": "cd3fe4d2d568d7224e18d06a732c3b49e55defad", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: turns a magenta (255,0,255) scribble 3-7 px wide on a photo into the described object. Prompt ' Turn the magenta scribble into {caption}.'; 40 steps, cfg 1. Best stacked with viggle-turbo (both 1.0, 6 steps) - alone it only matches the base. Use this gate_up-split file: the repo's other files keep img_mlp.gate_up fused, which diffusers drops", + "use_when": "Add an object to a photo where a magenta scribble is drawn", + "trigger": "", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "insert", + "scribble" + ] +} diff --git a/loras/qwen-image-2.1/fewstep-8.json b/loras/qwen-image-2.1/fewstep-8.json new file mode 100644 index 00000000..fd57d46f --- /dev/null +++ b/loras/qwen-image-2.1/fewstep-8.json @@ -0,0 +1,21 @@ +{ + "model_name": "ThakiCloud/Qwen-Image-2.1-FewStep-v0.1", + "weight_name": "qif_qwen_image_2.1_8step_v0.1.safetensors", + "revision": "82327fb112551eba6ccd5f54e6d0779105452c66", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "8-step DMD2 distill, text-to-image only (editing unverified). Scheduler use_dynamic_shifting false, shift 1.0, shift_terminal null; sigmas [1, 14/15, 6/7, 10/13, 2/3, 6/11, 0.4, 2/9]; cfg 1; trained at 1024x1024 only. The repo's qif_..._5step file is the 5-step variant, sigmas [1, 0.94, 6/7, 2/3, 0.4]", + "use_when": "Fast Qwen-Image-2.1 text-to-image in 8 steps", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "speed", + "turbo", + "distill" + ] +} diff --git a/loras/qwen-image-2.1/fun-acc-4step.json b/loras/qwen-image-2.1/fun-acc-4step.json new file mode 100644 index 00000000..1e0efaef --- /dev/null +++ b/loras/qwen-image-2.1/fun-acc-4step.json @@ -0,0 +1,25 @@ +{ + "model_name": "alibaba-pai/Qwen-Image-2.1-Fun-Acc-LoRAs", + "weight_name": "models/Qwen-Image-2.1-Fun-Acc-4Step.safetensors", + "revision": "f7545234760e1847cd8e89e52bd951cb0b7e327f", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "4-step PDD distill in a custom format that needs alibaba-pai's own pipeline", + "use_when": "Do not use under dw", + "scale": { + "default": 1.0 + }, + "status": "rejected", + "evidence": [ + { + "note": "Custom qwenimage21_extracted_prefused_v1 format: .lora_down/.lora_up keys without .weight, plus full norm_q/norm_k/txt_in.text_norm and proj_out weights - diffusers raises Invalid LoRA checkpoint. Needs the vendor's qwenimage21_pdd.py pipeline" + } + ], + "license": "other", + "tags": [ + "rejected", + "speed" + ] +} diff --git a/loras/qwen-image-2.1/natural-exposure.json b/loras/qwen-image-2.1/natural-exposure.json new file mode 100644 index 00000000..a1a5d2e8 --- /dev/null +++ b/loras/qwen-image-2.1/natural-exposure.json @@ -0,0 +1,23 @@ +{ + "model_name": "prithivMLmods/Qwen-Image-2.1-Natural-Exposure-LoRA", + "weight_name": "Qwen-Image-2.1-Natural-Exposure-LoRA-4000.safetensors", + "revision": "382d066d079854a86c9513f3c5a4026ada21ddbe", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: corrects an over- or under-exposed photo to balanced, neutral exposure. Scale and steps not stated (try 1.0, 40 steps); a preview", + "use_when": "Fix the exposure of a photo", + "trigger": "Transform the image with balanced neutral exposure", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "exposure", + "photo", + "fix" + ] +} diff --git a/loras/qwen-image-2.1/object-mover.json b/loras/qwen-image-2.1/object-mover.json new file mode 100644 index 00000000..195cc3f7 --- /dev/null +++ b/loras/qwen-image-2.1/object-mover.json @@ -0,0 +1,22 @@ +{ + "model_name": "prithivMLmods/Qwen-Image-2.1-Object-Mover-Bbox-turbo", + "weight_name": "Qwen-Image-2.1-Object-Mover-Bbox-turbo-4000.safetensors", + "revision": "183eb9034853c03e6ec273d81c324edb3d033f34", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: moves the object in one red box to the spot marked by a second red box. Trained on only 45 pairs", + "use_when": "Move an object from one red box to another", + "trigger": "Move the object highlighted in the red box to the location indicated by the other red box in the scene.", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "move", + "bbox" + ] +} diff --git a/loras/qwen-image-2.1/object-remover.json b/loras/qwen-image-2.1/object-remover.json new file mode 100644 index 00000000..83ec1f57 --- /dev/null +++ b/loras/qwen-image-2.1/object-remover.json @@ -0,0 +1,22 @@ +{ + "model_name": "prithivMLmods/Qwen-Image-2.1-Object-Remover-Bbox-turbo", + "weight_name": "Qwen-Image-2.1-Object-Remover-Bbox-turbo-4000.safetensors", + "revision": "90f708f53feb896a2873378629ffe414e0e584c0", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: removes the object inside a red box drawn on the input image (red 239,68,68). The card's examples run 40 steps; 'turbo' means it also holds stacked with a turbo LoRA", + "use_when": "Remove an object marked with a red box", + "trigger": "Remove the red highlighted object from the scene", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "remove", + "bbox" + ] +} diff --git a/loras/qwen-image-2.1/turbo8.json b/loras/qwen-image-2.1/turbo8.json new file mode 100644 index 00000000..d55cf30e --- /dev/null +++ b/loras/qwen-image-2.1/turbo8.json @@ -0,0 +1,23 @@ +{ + "model_name": "chriswritescode/Turbo8-LoRA-Qwen-Image-2.1", + "weight_name": "turbo8_lora_step2500.safetensors", + "revision": "611a3f86d0cc53d21a1a09d07ea7dc82df3d68a1", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "8-step DMD2 distill covering text-to-image, editing (up to 10 references) and RGBA. 8 steps, cfg 1, scheduler shift_terminal null. For RGBA: 'This is an RGBA image with transparency. {description}. The image has alpha channel and the background is transparent.'", + "use_when": "Fast Qwen-Image-2.1 generation, edits or RGBA output in 8 steps", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "speed", + "turbo", + "distill", + "edit", + "rgba" + ] +} diff --git a/loras/qwen-image-2.1/viggle-turbo.json b/loras/qwen-image-2.1/viggle-turbo.json new file mode 100644 index 00000000..7284a1f2 --- /dev/null +++ b/loras/qwen-image-2.1/viggle-turbo.json @@ -0,0 +1,22 @@ +{ + "model_name": "Viggle/Qwen-Image-2.1-viggle-turbo", + "weight_name": "Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors", + "revision": "009a44a895ef85f7e643c80fdca9543795248867", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "6-step distill for text-to-image and editing with 1-3 references. Needs the repo's scheduler/ config (shift_terminal null - the base's 0.02 wrecks the last step), sigmas [1.0, 0.9375, 0.875, 0.75, 0.5, 0.25], true_cfg_scale 1.0 and no negative prompt; keep it unfused. Stacks with doodle-in", + "use_when": "Fast Qwen-Image-2.1 generation or edits in 6 steps", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "speed", + "turbo", + "distill", + "edit" + ] +} diff --git a/loras/qwen-image-2.1/voxel.json b/loras/qwen-image-2.1/voxel.json new file mode 100644 index 00000000..893099a1 --- /dev/null +++ b/loras/qwen-image-2.1/voxel.json @@ -0,0 +1,22 @@ +{ + "model_name": "prithivMLmods/Qwen-Image-2.1-Voxel-Style", + "weight_name": "Qwen-Image-2.1-Voxel-Style-3000.safetensors", + "revision": "b4ccfe9333f04c2f978f2d1bfdfdafe287cbfad1", + "base_models": [ + "Qwen/Qwen-Image-2.1" + ], + "workflow": null, + "description": "Edit: restyles an image as voxel art. The latest of the repo's 1000/2000/3000 checkpoints; the card names none", + "use_when": "Restyle an image as voxel art", + "trigger": "Transform the image into a voxel style", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "other", + "tags": [ + "edit", + "style", + "voxel" + ] +} diff --git a/loras/z-image/awportrait.json b/loras/z-image/awportrait.json new file mode 100644 index 00000000..3c117075 --- /dev/null +++ b/loras/z-image/awportrait.json @@ -0,0 +1,21 @@ +{ + "model_name": "Shakker-Labs/AWPortrait-Z", + "weight_name": "AWPortrait-Z.safetensors", + "revision": "9f264c4784cf3a9be4ba4769b9bb516d84905365", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Portrait quality: removes Turbo's skin grain, tones down the excess HDR, widens face diversity. No trigger, scale not stated; r128", + "use_when": "Portraits with cleaner skin and more natural tone", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "portrait", + "quality", + "realism" + ] +} diff --git a/loras/z-image/childrens-drawings.json b/loras/z-image/childrens-drawings.json new file mode 100644 index 00000000..509788f3 --- /dev/null +++ b/loras/z-image/childrens-drawings.json @@ -0,0 +1,21 @@ +{ + "model_name": "ostris/z_image_turbo_childrens_drawings", + "weight_name": "z_image_turbo_childrens_drawings.safetensors", + "revision": "7fcd66a99149c58741990fca28562a4a581af7a9", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Makes any prompt look like a child's drawing. No trigger, scale not stated; ostris's tutorial LoRA", + "use_when": "A picture that looks drawn by a child", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "drawing", + "naive" + ] +} diff --git a/loras/z-image/classic-painting.json b/loras/z-image/classic-painting.json new file mode 100644 index 00000000..9d25dfed --- /dev/null +++ b/loras/z-image/classic-painting.json @@ -0,0 +1,22 @@ +{ + "model_name": "renderartist/Classic-Painting-Z-Image-Turbo-LoRA", + "weight_name": "Classic_Painting_Z_Image_Turbo_v1_renderartist_1750.safetensors", + "revision": "e873cc474ce517d49d09ed1f64358ff43660b276", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Old Master oil painting, trained on public-domain Art Institute of Chicago works with no artist names. The trigger is taken from the card's examples", + "use_when": "An Old Master oil painting", + "trigger": "class1cpa1nt", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "painting", + "classical" + ] +} diff --git a/loras/z-image/coloring-book.json b/loras/z-image/coloring-book.json new file mode 100644 index 00000000..808dd91b --- /dev/null +++ b/loras/z-image/coloring-book.json @@ -0,0 +1,22 @@ +{ + "model_name": "renderartist/Coloring-Book-Z-Image-Turbo-LoRA", + "weight_name": "Coloring_Book_Z_Image_Turbo_v1_renderartist_2000.safetensors", + "revision": "3d65420a16668ec00847915f2f69c50854eb5d33", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Black-and-white coloring-book line art. The trigger is optional - 'black and white cartoon, simple, cute' also works", + "use_when": "A coloring-book page", + "trigger": "c0l0ringb00k", + "scale": { + "default": 0.7 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "line-art", + "coloring-book" + ] +} diff --git a/loras/z-image/dejpeg.json b/loras/z-image/dejpeg.json new file mode 100644 index 00000000..cafc2052 --- /dev/null +++ b/loras/z-image/dejpeg.json @@ -0,0 +1,21 @@ +{ + "model_name": "wcde/Z-Image-Turbo-DeJPEG-Lora", + "weight_name": "dejpeg_v3.safetensors", + "revision": "a963b554171a35895660af26fe9e931994fb9587", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Removes Turbo's JPEG-like block and ringing artifacts. v3 (r128) is the newest of the repo's files. No trigger, scale not stated. The repo declares no license", + "use_when": "Clean JPEG-like artifacts out of Z-Image-Turbo output", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "unknown", + "tags": [ + "quality", + "fix", + "artifacts" + ] +} diff --git a/loras/z-image/detailed-eyes.json b/loras/z-image/detailed-eyes.json new file mode 100644 index 00000000..2335a079 --- /dev/null +++ b/loras/z-image/detailed-eyes.json @@ -0,0 +1,26 @@ +{ + "model_name": "bdsqlsz/qinglong_DetailedEyes_Z-Image", + "weight_name": "qinglong_detailedeye_z-image.safetensors", + "revision": "cf2ef7daebdb5c2c042ea09b59b429aeba1e3bb9", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Eye-detail slider (trained 'girl' to 'extremely detailed gradient eyes'): scale works both ways, negative reduces it. No trigger, scale not stated; one user reports a reddish tint. kohya format, includes the refiners", + "use_when": "Sharper, more detailed eyes", + "scale": { + "default": 1.0, + "range": [ + -1.0, + 1.0 + ] + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "eyes", + "detail", + "slider", + "quality" + ] +} diff --git a/loras/z-image/historic-color.json b/loras/z-image/historic-color.json new file mode 100644 index 00000000..4333f018 --- /dev/null +++ b/loras/z-image/historic-color.json @@ -0,0 +1,23 @@ +{ + "model_name": "AlekseyCalvin/HistoricColor_Z-image-Turbo-LoRA", + "weight_name": "ZImage1HST.safetensors", + "revision": "c4f2cb7bfd1df640522cc4d5c5ed9fe061de6d61", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "1900s-1910s early colour photography in the Prokudin-Gorsky manner. The final file of the repo's checkpoints; scale not stated", + "use_when": "An early-1900s colour photograph", + "trigger": "HST photo", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "photo", + "vintage", + "historic" + ] +} diff --git a/loras/z-image/lenovo-ultrareal.json b/loras/z-image/lenovo-ultrareal.json new file mode 100644 index 00000000..18064c15 --- /dev/null +++ b/loras/z-image/lenovo-ultrareal.json @@ -0,0 +1,21 @@ +{ + "model_name": "Danrisi/Lenovo_UltraReal_Z_Image", + "weight_name": "lenovo_z.safetensors", + "revision": "9708e1e10d666d60ca9b97bfdbf1d697fa4acb16", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Amateur phone-photo realism, from the author's UltraReal series. The card is empty: trigger and scale unknown. Overlaps realism", + "use_when": "A casual amateur phone-camera photo", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "realism", + "photo", + "amateur" + ] +} diff --git a/loras/z-image/pixel-art.json b/loras/z-image/pixel-art.json new file mode 100644 index 00000000..b5cb2028 --- /dev/null +++ b/loras/z-image/pixel-art.json @@ -0,0 +1,21 @@ +{ + "model_name": "tarn59/pixel_art_style_lora_z_image_turbo", + "weight_name": "pixel_art_style_z_image_turbo.safetensors", + "revision": "0a5092d1619664d94a5a36784f92db84b3ae62bd", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Pixel art style. Scale not stated", + "use_when": "Render as pixel art", + "trigger": "Pixel art style.", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "pixel-art" + ] +} diff --git a/loras/z-image/realism.json b/loras/z-image/realism.json new file mode 100644 index 00000000..da3daac2 --- /dev/null +++ b/loras/z-image/realism.json @@ -0,0 +1,21 @@ +{ + "model_name": "suayptalha/Z-Image-Turbo-Realism-LoRA", + "weight_name": "pytorch_lora_weights.safetensors", + "revision": "8dc179cb56844d4ddc6e4527c30e98347eeae971", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Photo realism, trained on the cc12m-4mp-realistic subset. Prefix the prompt with the trigger. Scale not stated; r16", + "use_when": "A more photographic look", + "trigger": "Realism, ", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "realism", + "photo" + ] +} diff --git a/loras/z-image/reversal-film-gravure.json b/loras/z-image/reversal-film-gravure.json new file mode 100644 index 00000000..b675aa26 --- /dev/null +++ b/loras/z-image/reversal-film-gravure.json @@ -0,0 +1,23 @@ +{ + "model_name": "AIImageStudio/ReversalFilmGravure_z_Image_turbo", + "weight_name": "z_image_turbo_ReversalFilmGravure_v1.0.safetensors", + "revision": "f38c6abc8b9399ad43aaea75681629c1dfd23e4d", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Slide-film portrait look: more contrast, less noise. The card samples with euler / sgm_uniform", + "use_when": "A slide-film photographic portrait", + "trigger": "Reversal Film Gravure, analog film photography", + "scale": { + "default": 0.5 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "film", + "photo", + "portrait" + ] +} diff --git a/loras/z-image/sunbleached-flash.json b/loras/z-image/sunbleached-flash.json new file mode 100644 index 00000000..e015cd74 --- /dev/null +++ b/loras/z-image/sunbleached-flash.json @@ -0,0 +1,22 @@ +{ + "model_name": "Quorlen/z_image_turbo_Sunbleached_Protograph_Style_Lora", + "weight_name": "zimageturbo_Sunbleach_Photograph_Style_Lora_TAV2_000002500_(recommended).safetensors", + "revision": "aa74c8731efe332ebfb3397f71d4329ac25a6249", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Sun-bleached flash photograph: teal shadows, peach skin, halation. The trigger is optional; end the prompt with 'flash pop, warm peach skin, cyan grass, sunbleach style, red halation, light vignette'. The author's recommended checkpoint of twelve", + "use_when": "A sun-bleached flash photo look", + "trigger": "Act1vate!", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "mit", + "tags": [ + "style", + "photo", + "film" + ] +} diff --git a/loras/z-image/technically-color.json b/loras/z-image/technically-color.json new file mode 100644 index 00000000..8240eb35 --- /dev/null +++ b/loras/z-image/technically-color.json @@ -0,0 +1,23 @@ +{ + "model_name": "renderartist/Technically-Color-Z-Image-Turbo", + "weight_name": "Technically_Color_Z_Image_Turbo_v1_renderartist_2000.safetensors", + "revision": "193838d6d10d552514d51105fa1c07e4a7bbd7e6", + "base_models": [ + "Tongyi-MAI/Z-Image-Turbo" + ], + "workflow": null, + "description": "Classic three-strip Technicolor film look. The author's examples use a two-pass workflow", + "use_when": "A vintage Technicolor film look", + "trigger": "t3chnic4lly", + "scale": { + "default": 1.0 + }, + "status": "trial", + "license": "apache-2.0", + "tags": [ + "style", + "film", + "color", + "vintage" + ] +} diff --git a/plugins/dw/.claude-plugin/plugin.json b/plugins/dw/.claude-plugin/plugin.json index a6806093..d2c24cf3 100644 --- a/plugins/dw/.claude-plugin/plugin.json +++ b/plugins/dw/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "dw", "description": "Compose MiniMax H3 video, MiniMax Music 3 and LTX-2.5 workflows over a dw MCP server, and cut a multi-episode series from them: which template fits which shape, the hard rules, cost, and how to judge the output. Prompt format comes from the vendors' own guides.", - "version": "0.8.0", + "version": "0.9.0-alpha.1", "author": { "name": "Don Kackman" }, diff --git a/plugins/dw/README.md b/plugins/dw/README.md index eeea5d50..0b2dbfbf 100644 --- a/plugins/dw/README.md +++ b/plugins/dw/README.md @@ -36,7 +36,10 @@ Name the skill: the H3 repo also ships eight style packs alongside it. Each skill quotes catalog names and numeric rules that `tests/test_plugin_skills.py` -holds to the catalog and to the diffusers module that enforces them. The +holds to the catalog and to the diffusers module that enforces them. A +skill's `SKILL.md` stays under 12 KB because it loads whenever the skill +triggers; detail only some requests need is in its `references/`, read on +demand, and held to the same tests. The plugin's version is the engine's; the release script bumps both. Adding a family: copy a skill, follow its outline, add the family's rules to diff --git a/plugins/dw/skills/ltx-2.5/SKILL.md b/plugins/dw/skills/ltx-2.5/SKILL.md index 1b8cf390..baf0e52a 100644 --- a/plugins/dw/skills/ltx-2.5/SKILL.md +++ b/plugins/dw/skills/ltx-2.5/SKILL.md @@ -12,10 +12,10 @@ schedule that is not a knob. Every template here fits a 24 GB card. 1. `get_server_info`: the device and workspace. On `mps` they run slower than CUDA `cost` says (only text-to-video has run - there); on `cpu` say so and stop. + there); on `cpu` say so, stop. 2. `list_workflows(shape="shot")`: every template in the family carries - that shape. Take current names, `summary`, `traits` and `cost` from - the listing, trusting it over names quoted below. + that shape. Take names, `summary`, `traits` and `cost` from the + listing over names quoted below. 3. `get_workflow` on the one chosen, for its variables and defaults. 4. Before anything near the card's ceiling - a full-size refine, 481 frames, a 2x upscale - `get_memory` on an idle server; a leftover `live: true` figure @@ -33,7 +33,7 @@ schedule that is not a knob. Every template here fits a 24 GB card. - **Sharper at full size**: `templates/ltx2/two-stage` - eight sigmas at 768x448, a 2x latent upsample, then renoise and three stage-two sigmas at 1536x896 carrying audio latents through. The upsample alone is soft; - the refine pass supplies the detail. On the user's clip: `refine-clip`. + the refine adds the detail. On the user's clip: `refine-clip`. - `templates/ltx2/diffusion-decode` compares decoders. Do not offer it: without a `shi-labs/natten` build its fallback OOMs on 24GB at any size. - **A generative 2x render**: `templates/ltx2/generative-upscale` draws its own @@ -62,7 +62,8 @@ schedule that is not a knob. Every template here fits a 24 GB card. added a chunked one 2026-09-29); one 481-frame pass reaches 20 seconds first. If none fits, compose from `list_tasks` before authoring a new workflow; -read the `workflows` guide's authoring section first. +read the `workflows` guide's authoring section first. Other +LoRAs: `list_loras` first. ## Hard rules @@ -102,41 +103,8 @@ angle and lighting, reuse the same identifiers for recurring subjects, and say what the audio does. 2 to 4 shots; stay single-take for image-to-video, lip-synced dialogue or an unbroken camera move. Skip the enhancer. -The spec, from -`diffusers.pipelines.ltx2.utils.LTX2_5_T2V_DEFAULT_SYSTEM_PROMPT` (the -image-to-video variant, `LTX2_5_I2V_DEFAULT_SYSTEM_PROMPT`, adds the -image-grounding rule): - -## The trained caption spec - -``` -You are given a user's short text-to-video request. Write a single, highly detailed audio-visual caption describing the video that best fulfills that request, in the EXACT style of the training captions used for this video model. The generated video is scored against the user's ORIGINAL request, so preserve every element the user stated; expand faithfully into the full caption style without contradicting or dropping anything they asked for. - -Match this captioning style precisely: - -1. Begin immediately with the action or visual detail. Do NOT use "The scene opens…", "We see…", "There is…". - -2. Objective, observable description only. Do not infer emotions or intentions — describe what is visible and audible (e.g. not "he looks sad" but "his eyebrows angle downward and his lips are pressed together"). - -3. Full visual detail: environment (materials, textures, lighting, colors), character appearance (clothing, posture, facial details), and the spatial positioning of all elements. When a human appears, identify them specifically (gendered terms when clearly implied; differentiate multiple people consistently) and describe visible physical attributes — apparent gender presentation, skin tone, estimated age group, hair color/length/style, build, clothing and accessories. Do not infer ethnicity, nationality, religion, or culture. - -4. Precise motion and cinematic description. For every shot you MUST include, woven naturally into the prose (never as tags or labels): - - Shot type (exactly one: extreme wide shot / wide shot / medium shot / medium close-up / close-up / extreme close-up) - - Camera motion (always stated; if none, explicitly say the camera remains static). Camera movement is expected and good — match the user if they specified it, otherwise choose the treatment that best presents the requested scene. - - Camera viewpoint relative to subject (front-facing / back-facing / side view / over-the-shoulder / top-down / low-angle / high-angle). - Express these as flowing prose: "a medium shot frames…, captured from a front-facing angle as the camera slowly pans…". Never as "medium shot, static camera —". - -5. Complete soundscape, integrated naturally: any dialogue (quote it exactly, in the original language), tone of voice, background music (type, mood, volume changes), and environmental sounds (footsteps, wind, traffic, animals). If the request implies sound, describe it plausibly. - -6. Strict chronological, real-time flow using transitions like "Initially…", "A moment later…", "Simultaneously…". Keep every stated action in motion. - -7. One single continuous paragraph. No bullet points, no section headers, no labels like "Audio:" or "Visual:". Exhaustive and lossless — include background elements, subtle movements, lighting, secondary sounds — detailed enough to reconstruct the scene. Aim for a rich, complete paragraph (roughly 150–220 words). - -If the user wrote in another language, produce the English caption of the same content. Output ONLY the caption text — no JSON, no preamble. - -AESTHETIC QUALITY (in addition to the above, without breaking the objective caption style): render the described scene with strong visual production value — cinematic, film-grade color and contrast, beautiful natural lighting, crisp fine detail and texture, pleasing composition and depth. Weave these quality descriptors naturally into the same observable prose (e.g. "warm cinematic lighting", "richly saturated film-grade color", "crisp high-resolution detail") — describe how the exact requested scene LOOKS at its most visually striking, never adding new objects or actions. Keep everything else (framing triple, soundscape, chronological single paragraph, faithfulness) exactly as specified. - -``` +Before writing a caption, read `references/caption-spec.md`: the training +caption spec, verbatim from the pipeline, which the caption must follow. ## Run and judge diff --git a/plugins/dw/skills/ltx-2.5/references/caption-spec.md b/plugins/dw/skills/ltx-2.5/references/caption-spec.md new file mode 100644 index 00000000..e80866f0 --- /dev/null +++ b/plugins/dw/skills/ltx-2.5/references/caption-spec.md @@ -0,0 +1,39 @@ +# LTX-2.5: the trained caption spec + +Part of the `ltx-2.5` skill; its *Prompts* section says when to read this. + +## The trained caption spec + +The spec the model's captions were trained in, from +`diffusers.pipelines.ltx2.utils.LTX2_5_T2V_DEFAULT_SYSTEM_PROMPT` (the +image-to-video variant, `LTX2_5_I2V_DEFAULT_SYSTEM_PROMPT`, adds the +image-grounding rule): + +``` +You are given a user's short text-to-video request. Write a single, highly detailed audio-visual caption describing the video that best fulfills that request, in the EXACT style of the training captions used for this video model. The generated video is scored against the user's ORIGINAL request, so preserve every element the user stated; expand faithfully into the full caption style without contradicting or dropping anything they asked for. + +Match this captioning style precisely: + +1. Begin immediately with the action or visual detail. Do NOT use "The scene opens…", "We see…", "There is…". + +2. Objective, observable description only. Do not infer emotions or intentions — describe what is visible and audible (e.g. not "he looks sad" but "his eyebrows angle downward and his lips are pressed together"). + +3. Full visual detail: environment (materials, textures, lighting, colors), character appearance (clothing, posture, facial details), and the spatial positioning of all elements. When a human appears, identify them specifically (gendered terms when clearly implied; differentiate multiple people consistently) and describe visible physical attributes — apparent gender presentation, skin tone, estimated age group, hair color/length/style, build, clothing and accessories. Do not infer ethnicity, nationality, religion, or culture. + +4. Precise motion and cinematic description. For every shot you MUST include, woven naturally into the prose (never as tags or labels): + - Shot type (exactly one: extreme wide shot / wide shot / medium shot / medium close-up / close-up / extreme close-up) + - Camera motion (always stated; if none, explicitly say the camera remains static). Camera movement is expected and good — match the user if they specified it, otherwise choose the treatment that best presents the requested scene. + - Camera viewpoint relative to subject (front-facing / back-facing / side view / over-the-shoulder / top-down / low-angle / high-angle). + Express these as flowing prose: "a medium shot frames…, captured from a front-facing angle as the camera slowly pans…". Never as "medium shot, static camera —". + +5. Complete soundscape, integrated naturally: any dialogue (quote it exactly, in the original language), tone of voice, background music (type, mood, volume changes), and environmental sounds (footsteps, wind, traffic, animals). If the request implies sound, describe it plausibly. + +6. Strict chronological, real-time flow using transitions like "Initially…", "A moment later…", "Simultaneously…". Keep every stated action in motion. + +7. One single continuous paragraph. No bullet points, no section headers, no labels like "Audio:" or "Visual:". Exhaustive and lossless — include background elements, subtle movements, lighting, secondary sounds — detailed enough to reconstruct the scene. Aim for a rich, complete paragraph (roughly 150–220 words). + +If the user wrote in another language, produce the English caption of the same content. Output ONLY the caption text — no JSON, no preamble. + +AESTHETIC QUALITY (in addition to the above, without breaking the objective caption style): render the described scene with strong visual production value — cinematic, film-grade color and contrast, beautiful natural lighting, crisp fine detail and texture, pleasing composition and depth. Weave these quality descriptors naturally into the same observable prose (e.g. "warm cinematic lighting", "richly saturated film-grade color", "crisp high-resolution detail") — describe how the exact requested scene LOOKS at its most visually striking, never adding new objects or actions. Keep everything else (framing triple, soundscape, chronological single paragraph, faithfulness) exactly as specified. + +``` diff --git a/plugins/dw/skills/minimax-h3/SKILL.md b/plugins/dw/skills/minimax-h3/SKILL.md index 6a3fe663..eb44d81d 100644 --- a/plugins/dw/skills/minimax-h3/SKILL.md +++ b/plugins/dw/skills/minimax-h3/SKILL.md @@ -14,9 +14,9 @@ arguments; prompt format is MiniMax's, from its text not here. 1. `get_server_info`: the device (H3 is CUDA; quantized configs don't run on mps) and the session's workspace. 2. `list_workflows(shape="shot")`, `list_workflows(shape="sequence")` and - `list_workflows(shape="audio")`: the family's templates by current name, + `list_workflows(shape="audio")`: the family's templates by name, `summary`, `traits`, `cost`. Trust the listing over names below. -3. `get_workflow` on the one chosen, for its variables and defaults. +3. `get_workflow` on the one chosen: its variables and defaults. ## Which shape is the request @@ -54,23 +54,10 @@ arguments; prompt format is MiniMax's, from its text not here. `shots` entry, `concat_videos` splices) and `templates/minimax/music-video` (a song, one slice and one lip-synced shot per entry - `from_file` reuses an existing cast portrait and skips drawing). - `shots` is one list: a dialogue entry is `name`, `prompt`, `references` - and `num_frames`; a music-video entry is `name`, `prompt`, `start_frame`. - The listing's `lists` - block says what an entry carries; its `cost` carries `per_entry` when one - shot was measured - scale from that, else quote the default list's total. - A cut erases drift and each shot makes its audio: score with - `templates/minimax/music` and `templates/assemble-and-score`'s `pair_audio` - mix under the world sound. Keep one voice across shots with a repeated - audio reference, not a repeated description. Duck a score burying - voice-over rather than raising the world track - a `gain_audio` step per - shot, negative `gain_db`. -- **Dialogue into a song**, not a concat: sing each shot to a `slice_audio` - slice, the first at `cue_seconds` (song time on the first sung frame; it - enters that long before the cut), each next where the last ended; then - `join_into_song` (dialogue, `song_shots`, unbroken `song`, same - `cue_seconds`), `normalize_audio` to -3 dBFS, `pair_audio`. Recipe: - `workflows` guide, "A spoken scene breaking into a song". + Before composing one, read `references/cuts.md`: the `shots` list, + cost per entry, scoring and one voice across cuts. +- **Dialogue into a song**, not a concat: `slice_audio`, `join_into_song`, + `pair_audio`, in that reference's last section. - **Unrelated shots, no cut**: `templates/minimax/shots-batch` - one H3 step per `shots` entry, no shared cast, no concat. `keep_output` each clip, then `templates/assemble-and-score` cuts and scores. @@ -89,22 +76,9 @@ read the `workflows` guide's authoring section. - Canvas: 768-pixel short edge, at most 768x1344, in multiples of 32, aspect 1:4 to 4:1. Output audio is 32 kHz stereo. - A checkpoint comes with a canvas, a sigma shift and an alpha and they move - together - change one, change all. `video_shift`, `audio_shift` and - `lora_alpha` are variables everywhere, so a swap is arguments, not a file. - Three combinations are tested, plus two additions below, nothing else: 544p FL2VA turbo, 960x544, - shift 12/3, alpha unset - default; 768p FL2VA turbo, 1344x768, shift - **6**/3, alpha unset - `video-with-audio-768p`; 768p Ref2VA turbo, shift - 12/3, alpha unset - every `ref2va` template. The two 768p LoRAs differ in - shift; do not generalise. - Two 768p additions (#585, one prompt): a **4-step draft** on - `video-with-audio-768p` - `lora_weight_name=minimax_h3_fl2v_turbo_4step_v1.2_768p_bf16.safetensors`, - `num_inference_steps=5`, shift 6/3, alpha unset; ~25% faster, one cold run; - render at 8. **Realism People** as a second `loras` entry on its t2va - step (`fal/MiniMax-H3-Realism-People-LoRA`, - `h3-realism-people-t2v-i2v-r2v.safetensors`, scale 0.7, own `adapter_name`; - a saved copy): validate warns the name fits neither partition, so never on - a `ref2va` template. Acc PDD, HyperFlow and drozbay FastH3 need their own - loaders, not `loras`: don't try them. + together - change one, change all. Before swapping one or adding a LoRA, + read `references/checkpoints-and-loras.md`: the tested combinations, + the LoRAs that need their own loaders, and the catalog. Never put an FL2VA LoRA on a reference template: `ref2va` holds `transformer_ref` alone, so anything handed there degrades output. `validate_workflow` refuses it and warns on a `weight_name` naming neither. diff --git a/plugins/dw/skills/minimax-h3/references/checkpoints-and-loras.md b/plugins/dw/skills/minimax-h3/references/checkpoints-and-loras.md new file mode 100644 index 00000000..128a50b6 --- /dev/null +++ b/plugins/dw/skills/minimax-h3/references/checkpoints-and-loras.md @@ -0,0 +1,27 @@ +# MiniMax H3: checkpoints and LoRAs + +Part of the `minimax-h3` skill; its *Hard rules* say when to read this. + +A checkpoint comes with a canvas, a sigma shift and an alpha and they move +together - change one, change all. `video_shift`, `audio_shift` and +`lora_alpha` are variables everywhere, so a swap is arguments, not a file. + +Three combinations are tested, plus two additions below, nothing else: +544p FL2VA turbo, 960x544, shift 12/3, alpha unset - default; 768p FL2VA +turbo, 1344x768, shift **6**/3, alpha unset - `video-with-audio-768p`; 768p +Ref2VA turbo, shift 12/3, alpha unset - every `ref2va` template. The two 768p +LoRAs differ in shift; do not generalise. + +Two 768p additions (#585, one prompt): a **4-step draft** on +`video-with-audio-768p` - `lora_weight_name=minimax_h3_fl2v_turbo_4step_v1.2_768p_bf16.safetensors`, +`num_inference_steps=5`, shift 6/3, alpha unset; ~25% faster, one cold run; +render at 8. **Realism People** as a second `loras` entry on its t2va +step (`fal/MiniMax-H3-Realism-People-LoRA`, +`h3-realism-people-t2v-i2v-r2v.safetensors`, scale 0.7, own `adapter_name`; +a saved copy): validate warns the name fits neither partition, so never on +a `ref2va` template. + +Acc PDD, HyperFlow and drozbay FastH3 need their own loaders, not `loras`: +don't try them. Any other LoRA (a style, a motion fix): `list_loras` first - +each entry names its partition, trigger and scale; `trial` is unproven, +`rejected` says why it fails. diff --git a/plugins/dw/skills/minimax-h3/references/cuts.md b/plugins/dw/skills/minimax-h3/references/cuts.md new file mode 100644 index 00000000..d6d379e0 --- /dev/null +++ b/plugins/dw/skills/minimax-h3/references/cuts.md @@ -0,0 +1,27 @@ +# MiniMax H3: a piece with cuts + +Part of the `minimax-h3` skill; *Which shape is the request* says when to +read this. + +## The cuts templates + +`shots` is one list: a dialogue entry is `name`, `prompt`, `references` +and `num_frames`; a music-video entry is `name`, `prompt`, `start_frame`. +The listing's `lists` +block says what an entry carries; its `cost` carries `per_entry` when one +shot was measured - scale from that, else quote the default list's total. +A cut erases drift and each shot makes its audio: score with +`templates/minimax/music` and `templates/assemble-and-score`'s `pair_audio` +mix under the world sound. Keep one voice across shots with a repeated +audio reference, not a repeated description. Duck a score burying +voice-over rather than raising the world track - a `gain_audio` step per +shot, negative `gain_db`. + +## Dialogue into a song + +Not a concat: sing each shot to a `slice_audio` +slice, the first at `cue_seconds` (song time on the first sung frame; it +enters that long before the cut), each next where the last ended; then +`join_into_song` (dialogue, `song_shots`, unbroken `song`, same +`cue_seconds`), `normalize_audio` to -3 dBFS, `pair_audio`. Recipe: +`workflows` guide, "A spoken scene breaking into a song". diff --git a/plugins/dw/skills/minimax-music3/SKILL.md b/plugins/dw/skills/minimax-music3/SKILL.md index 381dfb89..b5b0b57a 100644 --- a/plugins/dw/skills/minimax-music3/SKILL.md +++ b/plugins/dw/skills/minimax-music3/SKILL.md @@ -119,7 +119,8 @@ from the templates' examples: `npx skills add MiniMax-AI/MiniMax-Music3 --skill music-caption-rewriter` installs it; this is the family where the local skill pays off, since a fetch of `SKILL.md` alone doesn't reach the templates. -2. Else read its `SKILL.md` and `references/genre-router.md` at that path. The +2. Else read `skills/music-caption-rewriter/SKILL.md` and + `skills/music-caption-rewriter/references/genre-router.md` in that repo. The contract is three headings in order - Global Metadata, Vocal Details, Arrangement - in 250-450 English words, with no title, no reasoning, and no lyric line copied into the caption. The pipeline strips markdown @@ -163,29 +164,17 @@ Control" section. 4. Judge it yourself. `get_gallery_metadata` for duration and sample rate: `media.duration_seconds` within 0.2 s of `audio_duration` means the ceiling cut the track (raise it and rerun); well short of it means the - song finished on its own. Its `peak_dbfs` is a single sample and doesn't - say how loud the song reads end to end - `integrated_lufs` (BS.1770, - whole-track) is the field for that, and what `normalize_audio`'s - `target_lufs` targets when a mix must match another track by ear, not - by peak. A master louder than its peak allows (-16 streaming): add - `limit: true`; `limiter_heavy` means lower the target. Then listen with - `get_output_audio` (a long - track in `start`/`duration` excerpts) for the family's failure modes: a + song finished on its own. For loudness (`peak_dbfs`, `integrated_lufs`, + matching a mix) read `references/loudness.md`. Then listen with + `get_output_audio` (a long track in `start`/`duration` excerpts) for the family's failure modes: a song gone instrumental (name the vocals in the caption), an ending cut mid-note (raise the ceiling, then trim), a structure ignoring the tags (fewer sections, plainer directions). Hand the user the gallery `url` (`list_gallery`, or the manifest's file name). 5. To use the track in a later workflow, `keep_output` makes it an `asset:`; to trim it in the same run, chain `templates/audio-trim-fade` on the output. -6. After an inline run worth keeping, `get_job_workflow` and `save_workflow` it, - so the next run is by name rather than pasting JSON; `export_job` bundles - the run — workflow, manifest, job row and media — for git, on the server. - If `auth_required` is false, fetch `open_url` and unpack it into - `exports/` under the session's working directory, never a temp directory - (the archive already unpacks into a job-id folder, don't make one first). - If true, this agent can't attach the token - hand `open_url` to - the person, and keep working via - `get_output_image`/`get_output_audio`/`get_output_frames`. +6. A run worth keeping: `references/keeping-a-run.md` saves it by name + and exports it. ## Sources diff --git a/plugins/dw/skills/minimax-music3/references/keeping-a-run.md b/plugins/dw/skills/minimax-music3/references/keeping-a-run.md new file mode 100644 index 00000000..cf8956a9 --- /dev/null +++ b/plugins/dw/skills/minimax-music3/references/keeping-a-run.md @@ -0,0 +1,15 @@ +# MiniMax Music 3: keeping a run + +Part of the `minimax-music3` skill; step 6 of *Run and judge* says when to +read this. + +After an inline run worth keeping, `get_job_workflow` and `save_workflow` it, +so the next run is by name rather than pasting JSON; `export_job` bundles +the run — workflow, manifest, job row and media — for git, on the server. +`auth_required` is false (the zip needs no token): fetch `open_url` +(prefix a relative one with the server address) and unpack it into +`exports/` under the session's working directory, never a temp directory +(the archive already unpacks into a job-id folder, don't make one first). +Only without HTTP, or if `auth_required` is true (a gated zip), hand +`open_url` to the person, and keep working via +`get_output_image`/`get_output_audio`/`get_output_frames`. diff --git a/plugins/dw/skills/minimax-music3/references/loudness.md b/plugins/dw/skills/minimax-music3/references/loudness.md new file mode 100644 index 00000000..87296a3e --- /dev/null +++ b/plugins/dw/skills/minimax-music3/references/loudness.md @@ -0,0 +1,11 @@ +# MiniMax Music 3: loudness + +Part of the `minimax-music3` skill; step 4 of *Run and judge* says when to +read this. + +`get_gallery_metadata`'s `peak_dbfs` is a single sample and doesn't say +how loud the song reads end to end - `integrated_lufs` (BS.1770, +whole-track) is the field for that, and what `normalize_audio`'s +`target_lufs` targets when a mix must match another track by ear, not +by peak. A master louder than its peak allows (-16 streaming): add +`limit: true`; `limiter_heavy` means lower the target. diff --git a/plugins/dw/skills/script-to-video/SKILL.md b/plugins/dw/skills/script-to-video/SKILL.md index a72fdc52..f0534eb0 100644 --- a/plugins/dw/skills/script-to-video/SKILL.md +++ b/plugins/dw/skills/script-to-video/SKILL.md @@ -83,6 +83,10 @@ Once every shot exists, hand off to `templates/assemble-and-score` (or the a series) for the recut/bed/match_levels/normalize/pair pass. That skill already owns this step - do not re-derive it here. +To take the project home, call `export_job` once per job of the project, then +fetch each zip's `open_url` and unpack it into `exports/` under the working +directory (per its `next`). + ## Not in scope - **Automatic model speed optimization.** If a model is slow, that is a diff --git a/plugins/dw/skills/series-episodes/SKILL.md b/plugins/dw/skills/series-episodes/SKILL.md index 72f6bf9c..071a0fee 100644 --- a/plugins/dw/skills/series-episodes/SKILL.md +++ b/plugins/dw/skills/series-episodes/SKILL.md @@ -171,6 +171,10 @@ wait_seconds=60)`, then `get_output_text` on the result, and plan is `basis: "unknown"` with `minutes: null` - quote seconds, not minutes; it runs in a few seconds. +To take the project home, call `export_job` once per job of the project (each +episode, and the cast run), then fetch each zip's `open_url` and unpack it into +`exports/` under the working directory (per its `next`). + ## Sources `workflows/templates/assemble-and-score.json`, `dw/tasks/audio_utils.py` diff --git a/pyproject.toml b/pyproject.toml index d351356c..299519dd 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -8,7 +8,7 @@ name = "diffusers-workflow" # runtime (static in TOML because importing dw at build time would drag # torch into the build environment). Release tags must match it - see # docs/RELEASING.md -version = "0.8.0" +version = "0.9.0-alpha.1" description = "Declarative workflow engine and web UI for Hugging Face Diffusers" readme = "README.md" requires-python = ">=3.10" diff --git a/tests/test_assemble_and_score_target_lufs.py b/tests/test_assemble_and_score_target_lufs.py index e48623fa..70cdd457 100644 --- a/tests/test_assemble_and_score_target_lufs.py +++ b/tests/test_assemble_and_score_target_lufs.py @@ -124,3 +124,15 @@ def test_the_description_places_the_limit_ceiling_on_the_mix_not_the_film(): description = load_definition()["description"] assert "holds on the mix 'balanced' writes, not on the film" in description assert "about 1 dB above it" in description + + +def test_edit_join_gets_template_sample_rate(): + """#594: the edit join resampled mixed-rate shots to its own 48 kHz default + and warned the caller to pass sample_rate, which the template never + threaded through.""" + from dw.variables import replace_variables, resolve_variable_values + + definition = load_definition() + merged = resolve_variable_values({**definition["variables"], "sample_rate": 32000}) + edit = steps_by_name(replace_variables(definition, merged))["edit"] + assert edit["task"]["arguments"]["sample_rate"] == 32000 diff --git a/tests/test_concat_videos.py b/tests/test_concat_videos.py index 5f878feb..297ff3ab 100644 --- a/tests/test_concat_videos.py +++ b/tests/test_concat_videos.py @@ -643,6 +643,23 @@ def test_the_resample_warning_is_emitted_as_an_event(self): assert "resampling them all to 200 Hz" in warnings[0]["message"] assert "resample_audio" in warnings[0]["message"] + def test_mixed_rates_resampled_to_a_pinned_rate_draw_no_warning(self): + """#594: a pinned sample_rate is the target; no advice to pass it.""" + from dw.events import RunContext, activate_context, deactivate_context + + events = [] + token = activate_context(RunContext(on_event=events.append)) + try: + result = concat_videos( + [audio_video(8, 0.5), audio_video(8, 0.5, sample_rate=200)], + sample_rate=150, + ) + finally: + deactivate_context(token) + + assert result.sample_rate == 150 + assert [e for e in events if e.get("event") == "warning"] == [] + def test_agreeing_rates_resampled_to_a_pinned_rate_draw_no_warning(self): """#453: inputs that agree, converted to the rate the caller pinned, are not 'videos at different sample rates'.""" diff --git a/tests/test_mcp_exports.py b/tests/test_mcp_exports.py index c899eb2e..55c7da99 100644 --- a/tests/test_mcp_exports.py +++ b/tests/test_mcp_exports.py @@ -106,9 +106,10 @@ def test_the_next_hint_sends_the_zip_to_the_working_directory(): def test_auth_required_tells_the_agent_to_hand_the_zip_to_a_person(): - """#353: when the server gates the zip with a bearer token, this agent - has no way to attach one to a fetch made on the person's behalf - the - hint has to say hand it over, not fetch it.""" + """#353, kept as a forward guard by #592: the server reports false + today, but if the zip is ever gated this agent has no way to attach the + token to a fetch made on the person's behalf - the hint has to say hand + it over, not fetch it.""" client, _ = exporting(body={**SUMMARY, "auth_required": True}) result = exports.export_job(client, "job-1") @@ -135,7 +136,9 @@ def test_an_absolute_open_url_is_preferred_and_said_to_be_absolute(): assert "already absolute" in result["next"] -def test_no_auth_required_still_fetches_the_zip_itself(): +def test_no_auth_required_fetches_the_zip_itself(): + """#592: the zip is ungated, so the server reports false and the hint + says fetch and unpack - not hand it over, which is the gated branch.""" client, _ = exporting(body={**SUMMARY, "auth_required": False}) result = exports.export_job(client, "job-1") @@ -143,6 +146,10 @@ def test_no_auth_required_still_fetches_the_zip_itself(): assert result["auth_required"] is False assert result["open_url"] == "/exports/job-1.zip" assert "fetch open_url" in result["next"] + assert "needs no token" in result["next"] + assert "exports/" in result["next"] + assert "hand open_url to the person" not in result["next"] + assert "do not fetch it" not in result["next"] def test_absolute_zip_url_is_omitted_rather_than_null_when_unconfigured(): diff --git a/tests/test_mcp_server.py b/tests/test_mcp_server.py index edde55be..fde28c6e 100644 --- a/tests/test_mcp_server.py +++ b/tests/test_mcp_server.py @@ -1056,17 +1056,18 @@ async def test_export_job_sends_the_zip_to_the_working_directory(): @pytest.mark.asyncio -async def test_export_job_description_is_auth_aware(): - """#353: the served tool description, not just the runtime `next` hint, - has to tell the agent not to fetch an auth-gated zip on the person's - behalf - the description is what the agent plans from before it ever - calls the tool and sees `next`.""" +async def test_export_job_description_says_to_fetch_the_ungated_zip(): + """#592: the zip route is ungated, so the served description - what the + agent plans from before it sees `next` - tells it to fetch open_url, + no longer the #353 "do NOT fetch it", and keeps auth_required's hand-over + as the forward guard for a gated zip.""" tools = await tools_of(server_over(ok({}))) description = tools["export_job"].description + assert "needs no token" in description + assert "fetch it with any HTTP" in description + assert "do NOT fetch it" not in description assert "auth_required" in description - assert "do NOT fetch it" in description - assert "hand open_url to the person" in description @pytest.mark.asyncio diff --git a/tests/test_plan.py b/tests/test_plan.py index 991f7875..7a018e71 100644 --- a/tests/test_plan.py +++ b/tests/test_plan.py @@ -365,14 +365,42 @@ def test_two_changed_lists_withhold_the_figure(self, plan): answer = plan(spec, arguments=arguments)["estimate"] assert (answer["minutes"], answer["basis"]) == (None, "unknown") - def test_another_devices_figure_is_not_extrapolated(self, plan): - """`other_device` already says the figure is not this machine's - - re-pricing it would dress a guess as arithmetic.""" + def test_another_devices_figure_is_re_priced_for_the_list(self, plan): + """#589: a 2-shot figure from another backend quoted verbatim for a + 10-shot run was the 1/100 gate. It stays `other_device` - not a + measurement here - but the entry count multiplies.""" spec = definition() spec["cost"] = [cost("mps", 40, name="M2")] shots = [{"name": f"s{n}", "prompt": "x"} for n in range(10)] answer = plan(spec, arguments={"shots": shots})["estimate"] - assert (answer["minutes"], answer["basis"]) == (40.0, "other_device") + assert (answer["minutes"], answer["basis"]) == (200.0, "other_device") + + def test_another_devices_per_entry_rate_prices_the_list(self, plan): + spec = definition() + spec["cost"] = [ + cost("mps", 10, {"variable": "shots", "minutes": 3, "entries": 2}) + ] + shots = [{"name": f"s{n}", "prompt": "x"} for n in range(5)] + answer = plan(spec, arguments={"shots": shots})["estimate"] + assert (answer["minutes"], answer["basis"]) == (19.0, "other_device") + + def test_another_devices_figure_with_two_changed_lists_is_unknown(self, plan): + spec = definition() + spec["cost"] = [cost("mps", 40, name="M2")] + spec["variables"]["angles"] = [{"name": "wide"}] + spec["steps"].append( + { + "name": "angle", + "for_each": "variable:angles", + "task": {"command": "x", "arguments": {"name": "item:name"}}, + } + ) + arguments = { + "shots": [{"name": f"s{n}", "prompt": "x"} for n in range(4)], + "angles": [{"name": "wide"}, {"name": "tight"}], + } + answer = plan(spec, arguments=arguments)["estimate"] + assert (answer["minutes"], answer["basis"]) == (None, "unknown") def test_minutes_is_rounded_to_one_decimal(self, plan): spec = definition() @@ -389,6 +417,15 @@ def test_a_shifted_scalar_cost_driver_falls_back_to_unknown(self, plan): answer = plan(spec, arguments={"frames": 9})["estimate"] assert (answer["minutes"], answer["basis"]) == (None, "unknown") + def test_a_shifted_scalar_driver_is_unknown_on_another_device_too(self, plan): + """#589: the other_device branch skipped the #267 check, so a 241-frame + run was quoted the 121-frame clip's figure.""" + spec = definition() + spec["cost"] = [cost("mps", 40, name="M2")] + spec["cost_drivers"] = ["frames"] + answer = plan(spec, arguments={"frames": 9})["estimate"] + assert (answer["minutes"], answer["basis"]) == (None, "unknown") + def test_an_unshifted_scalar_cost_driver_still_quotes_the_catalog_figure( self, plan ): @@ -407,6 +444,35 @@ def test_a_list_driver_still_reprices_instead_of_going_unknown(self, plan): answer = plan(spec, arguments={"shots": shots})["estimate"] assert (answer["minutes"], answer["basis"]) == (50.0, "derived") + def test_a_shifted_per_entry_field_in_a_list_driver_is_unknown(self, plan): + """#593: a per-shot num_frames outside the values the default entries + were measured with has no matching bucket, same as a scalar (#267).""" + spec = definition() + spec["cost"] = [cost("mps", 40, name="M2")] + spec["cost_drivers"] = ["shots"] + spec["variables"]["shots"] = [ + {"name": "a", "prompt": "p", "num_frames": 124}, + {"name": "b", "prompt": "p", "num_frames": 141}, + ] + same = [ + {"name": "x", "prompt": "p", "num_frames": 124}, + {"name": "y", "prompt": "p", "num_frames": 124}, + ] + long = [ + {"name": "x", "prompt": "p", "num_frames": 243}, + {"name": "y", "prompt": "p", "num_frames": 243}, + ] + kept = plan(spec, arguments={"shots": same})["estimate"] + assert (kept["minutes"], kept["basis"]) == (40.0, "other_device") + moved = plan(spec, arguments={"shots": long})["estimate"] + assert (moved["minutes"], moved["basis"]) == (None, "unknown") + as_text = [dict(entry, num_frames=str(entry["num_frames"])) for entry in long] + moved = plan(spec, arguments={"shots": as_text})["estimate"] + assert (moved["minutes"], moved["basis"]) == (None, "unknown") + same_text = [dict(entry, num_frames="124") for entry in same] + kept = plan(spec, arguments={"shots": same_text})["estimate"] + assert (kept["minutes"], kept["basis"]) == (40.0, "other_device") + def test_an_undeclared_variable_shift_is_not_a_driver_shift(self, plan): """Only a declared cost_driver triggers the fallback - any other variable overridden away from its default is none of this rule's @@ -580,6 +646,17 @@ def test_a_childs_catalog_cost_is_added(self, plan, tmp_path): assert (answer["minutes"], answer["partial"]) == (7.0, False) assert answer["unpriced"] == [] + def test_a_number_never_carries_the_unknown_basis(self, plan, tmp_path): + """#593: an uncosted parent that only composes priced children sums + to a figure, which takes the children's basis rather than `unknown`.""" + child = {"id": "child", "cost": [cost("cuda", 5)], "steps": []} + (tmp_path / "child.json").write_text(json.dumps(child)) + parent = composing("child.json") + del parent["cost"] + parent["steps"] = parent["steps"][1:] + answer = plan(parent)["estimate"] + assert (answer["minutes"], answer["basis"]) == (5.0, "catalog") + def test_a_child_without_a_cost_makes_the_estimate_partial(self, plan, tmp_path): (tmp_path / "child.json").write_text(json.dumps({"id": "child", "steps": []})) answer = plan(composing("child.json"))["estimate"] diff --git a/tests/test_plugin_skills.py b/tests/test_plugin_skills.py index d3a52a7e..11fb7c76 100644 --- a/tests/test_plugin_skills.py +++ b/tests/test_plugin_skills.py @@ -22,12 +22,28 @@ SKILL_SIZE_LIMIT = 12 * 1024 CATALOG_NAME = re.compile(r"`((?:templates|models)/[A-Za-z0-9_./-]+)`") +REFERENCE_LINK = re.compile(r"`(references/[A-Za-z0-9_.-]+\.md)`") -def skill_text(path): +def skill_body(path): + """SKILL.md alone - what loads whenever the skill triggers.""" return open(path, encoding="utf-8").read() +def skill_references(path): + """The files beside a skill that it sends an agent to read on demand.""" + folder = os.path.join(os.path.dirname(path), "references") + return sorted(glob.glob(os.path.join(folder, "*.md"))) + + +def skill_text(path): + """A skill and its references, so a rule moved out of SKILL.md is still + held to the library that enforces it.""" + parts = [skill_body(path)] + parts += [open(ref, encoding="utf-8").read() for ref in skill_references(path)] + return "\n".join(parts) + + def frontmatter(text): """The YAML block between the leading '---' lines, as a dict of the top-level 'key: value' pairs (the two the plugin format needs).""" @@ -85,7 +101,9 @@ def test_there_are_skills(): "path", SKILLS, ids=lambda p: os.path.basename(os.path.dirname(p)) ) def test_a_skill_has_a_triggering_description_under_the_size_cap(path): - text = skill_text(path) + """The cap is on SKILL.md, which loads every time the skill triggers; + detail only some requests need goes in `references/`, read on demand.""" + text = skill_body(path) fields = frontmatter(text) assert fields["name"] == os.path.basename(os.path.dirname(path)) @@ -95,6 +113,21 @@ def test_a_skill_has_a_triggering_description_under_the_size_cap(path): ) +@pytest.mark.parametrize( + "path", SKILLS, ids=lambda p: os.path.basename(os.path.dirname(p)) +) +def test_a_skill_links_each_reference_and_each_link_resolves(path): + """A reference no SKILL.md names is never read, and a link to a file that + is not there sends the agent nowhere.""" + linked = set(REFERENCE_LINK.findall(skill_body(path))) + shipped = {"references/" + os.path.basename(ref) for ref in skill_references(path)} + + assert linked == shipped, ( + f"{path}: linked but missing {sorted(linked - shipped)}, " + f"shipped but never linked {sorted(shipped - linked)}" + ) + + @pytest.mark.parametrize( "path", SKILLS, ids=lambda p: os.path.basename(os.path.dirname(p)) ) diff --git a/tests/test_server_exports.py b/tests/test_server_exports.py index 0706df1c..d2117142 100644 --- a/tests/test_server_exports.py +++ b/tests/test_server_exports.py @@ -354,12 +354,13 @@ def test_a_second_export_without_overwrite_is_409(self, server): forced = client.post(f"/api/jobs/{job_id}/export?overwrite=true") assert forced.status_code == 201 - def test_auth_required_reflects_whether_a_token_is_configured( + def test_with_a_token_the_zip_is_not_auth_required_and_opens_without_one( self, workspace_root, tmp_path ): - # #353: an MCP-only agent has no way to attach a bearer token to a - # fetch on the person's behalf, so export_job's `next` hint branches - # on this field rather than assuming the zip is open to fetch. + # #592 (Don, Q1): the field and the route are pinned together, so + # they cannot drift apart again the way #353's + # `bool(state.api_token)` did - with a token configured the export + # says auth_required false, and the zip it names opens without one. manager = JobManager( workspace_root.outputs, worker_manager=ScriptedWorkerManager(exporting_script), @@ -395,8 +396,16 @@ def test_auth_required_reflects_whether_a_token_is_configured( body = client.post( f"/api/jobs/{submitted['id']}/export", headers=headers ).json() - - assert body["auth_required"] is True + unauthenticated_export = client.post( + f"/api/jobs/{submitted['id']}/export?overwrite=true" + ) + zip_response = client.get(body["zip_url"]) + + assert unauthenticated_export.status_code == 401 + assert body["auth_required"] is False + assert zip_response.status_code == 200 + names = zipfile.ZipFile(io.BytesIO(zip_response.content)).namelist() + assert f"{submitted['id']}/manifest.json" in names def test_an_absolute_zip_url_is_added_when_a_public_url_is_configured( self, server, monkeypatch diff --git a/workflows/templates/assemble-and-score.json b/workflows/templates/assemble-and-score.json index fb232f1a..efede295 100644 --- a/workflows/templates/assemble-and-score.json +++ b/workflows/templates/assemble-and-score.json @@ -42,6 +42,7 @@ "arguments": { "videos": "variable:shots", "fps": "variable:fps", + "sample_rate": "variable:sample_rate", "audio_bleed_ms": "variable:audio_bleed_ms", "audio_bleed_gain_db": "variable:audio_bleed_gain_db", "seam_fade_ms": "variable:seam_fade_ms", diff --git a/workflows/templates/ltx2/text-to-video.json b/workflows/templates/ltx2/text-to-video.json index f800559d..c3e721d9 100644 --- a/workflows/templates/ltx2/text-to-video.json +++ b/workflows/templates/ltx2/text-to-video.json @@ -2,7 +2,8 @@ "id": "LTX2", "description": "The baseline for the LTX-2.5 family: text to video with a soundtrack generated alongside it, on a single 24GB card. The distilled transformer runs a fixed eight-step schedule with guidance off, quantized per component so it stays resident while the Gemma text encoder and the connectors are group offloaded around it. First load adds about 10 minutes.", "cost": [ - {"device": "cuda", "name": "RTX 3090", "vram_gb": 24, "minutes": 1.8} + {"device": "cuda", "name": "RTX 3090", "vram_gb": 24, "minutes": 1.8}, + {"device": "mps", "name": "M5 Pro 64GB", "vram_gb": 64, "minutes": 6} ], "cost_drivers": ["num_frames", "width", "height"], "vram_estimate": { diff --git a/workflows/templates/minimax/music.json b/workflows/templates/minimax/music.json index 6c9d9315..415e9d33 100644 --- a/workflows/templates/minimax/music.json +++ b/workflows/templates/minimax/music.json @@ -7,6 +7,12 @@ "name": "RTX 3090", "vram_gb": 24, "minutes": 3.6 + }, + { + "device": "mps", + "name": "M5 Pro 64GB", + "vram_gb": 64, + "minutes": 3.3 } ], "cost_drivers": [