diff --git a/.claude/commands/auto.md b/.claude/commands/auto.md index 87aa8f5a..dc373e7d 100644 --- a/.claude/commands/auto.md +++ b/.claude/commands/auto.md @@ -1,6 +1,6 @@ --- description: One-verb pipeline — drop a video, confirm strategy, render clips -allowed-tools: Read, Bash, mcp__podcli__transcribe_podcast, mcp__podcli__transcribe_start, mcp__podcli__transcribe_status, mcp__podcli__get_ui_state, mcp__podcli__set_video, mcp__podcli__suggest_clips, mcp__podcli__batch_create_clips, mcp__podcli__knowledge_base, mcp__podcli__clip_history +allowed-tools: Read, Bash, mcp__podcli__transcribe_podcast, mcp__podcli__transcribe_start, mcp__podcli__job_status, mcp__podcli__get_ui_state, mcp__podcli__set_video, mcp__podcli__suggest_clips, mcp__podcli__batch_create_clips, mcp__podcli__knowledge_base, mcp__podcli__clip_history, mcp__podcli__record_decisions argument-hint: [video-path-or-episode-slug] [optional: count e.g. "5 clips"] triggers: - auto @@ -21,7 +21,7 @@ This command orchestrates the existing MCP tools on top of the compact packed tr 1. **Read, don't watch.** Reason about clips from the packed markdown view — not raw segments, not frame dumps. 2. **Strategy first, render after.** Propose the cut list and WAIT for user confirmation before calling `batch_create_clips`. -3. **Knowledge base is context, not template.** If `` exists, read it for brand voice and format preferences. If not, infer from the content itself. +3. **Knowledge base is context, not template.** If `.podcli/knowledge/` exists, read it for brand voice and format preferences. If not, infer from the content itself. 4. **Never silently render.** Every clip that ships must appear in the proposal the user approved. 5. **Every clip carries its own context.** A stranger who never heard the episode has to follow it from the first second. If the moment is an answer, the question comes with it. @@ -46,14 +46,16 @@ This command orchestrates the existing MCP tools on top of the compact packed tr - Call `transcribe_start(file_path)` → returns `{job_id, cached, estimate}` immediately. - If `cached: true`, skip to step 3. - Otherwise emit a short status to the user: _"Transcription started — estimated {estimate}. I'll check progress every 30s."_ - - Loop: call `transcribe_status(job_id, wait_seconds: 30)`. Between calls, emit ONE terse line to the user like `"Progress: 47% — pyannote diarization"`. Keep it to one line per poll — no repeat prose. Exit the loop when `done: true`. + - Loop: call `job_status(job_id, wait_seconds: 30)`. Between calls, emit ONE terse line to the user like `"Progress: 47%, pyannote diarization"`. Keep it to one line per poll, no repeat prose. Exit the loop when `done: true`. - If `status: "error"`, stop and report the error. 3. Read the packed transcript: `get_ui_state(include_transcript: true)`. This returns a compact phrase-grouped view with speakers, silence gaps, and energy peaks. - **If the header says speakers: 0**, stop and tell the user before going further. Without speaker labels you cannot tell a question from an answer, so the whole question-with-the-answer rule below is inert and the picks will be worse. Offer to re-transcribe with `transcribe_start(file_path, enable_diarization: true)`. Only continue without it if the user says to. -4. If `` exists, read `01-brand-identity.md`, `02-voice-and-tone.md`, and `04-shorts-creation-guide.md` for show context. Skip silently if missing — `/auto` works on any content. +4. If `.podcli/knowledge/` exists, read `01-brand-identity.md`, `02-voice-and-tone.md`, and `04-shorts-creation-guide.md` for show context. Skip silently if missing: `/auto` works on any content. 5. Call `clip_history` to see what's already been shipped for this episode. Avoid duplicates in the proposal. -**Fallback**: if `transcribe_start` returns an error about the Web UI not running, tell the user and offer either (a) run `npm run ui` in another terminal then retry, or (b) fall back to the synchronous `transcribe_podcast` (no live progress, works silently). +**Fallback**: if `transcribe_start` returns an error about the Web UI not running, tell the user and offer either (a) start the Web UI in another terminal then retry (`podcli studio` for a launcher install, `npm run ui` in a source checkout), or (b) fall back to the synchronous `transcribe_podcast` (no live progress, works silently). + +6. **Ask once, reuse the answer.** `get_ui_state` lists this episode's unanswered decisions under `OPEN QUESTIONS` (clip count, duration range, captions, language, thumbnails, delivery target). Ask whichever are relevant to this run, batched, not one dialog box per field, then call `record_decisions(video_path, ...)` with the answers. On every later run against this same video, those fields are already answered and won't appear in `OPEN QUESTIONS` again. ### Phase 2 — Topic Map (silent) @@ -86,6 +88,7 @@ Work inside one topic at a time. Set boundaries by meaning, not by the clock. - The question has to be inside the clip. `context_line` is a note for the editor, not a fix: nothing burns it into the video yet, so a clip that relies on it still ships with no setup. - If the question rambles past roughly 8 seconds, use `segments` to keep the asked part and cut the rambling, or drop the moment. - Never open on a word pointing back before the cut: "that", "it", "they", "yeah", "so", "exactly", "right", "which is why". Widen the start until the reference is inside the clip. +- If the sharpest line sits mid-clip, you may pass it as `hook` (`{start, end, mode}`, 1-15 seconds) so it plays first. It must be a line actually spoken inside the clip, never invented text. `repeat` replays it in place; `move` lifts it out. **end_second** diff --git a/.claude/commands/bootstrap-knowledge.md b/.claude/commands/bootstrap-knowledge.md index caad6f39..a1a2695b 100644 --- a/.claude/commands/bootstrap-knowledge.md +++ b/.claude/commands/bootstrap-knowledge.md @@ -18,7 +18,7 @@ triggers: ## Before starting -1. If `` has no files, run `podcli knowledge init` first so all 14 templates exist. +1. If `.podcli/knowledge/` has no files, run `podcli knowledge init` first so all 14 templates exist. 2. Ask for whichever of these the user has not provided: - Channel or podcast URL (YouTube channel, Spotify show, RSS feed) - Or a few sentences about the show if nothing is published yet diff --git a/.claude/commands/generate-titles.md b/.claude/commands/generate-titles.md index 39224bc5..e91f2d3d 100644 --- a/.claude/commands/generate-titles.md +++ b/.claude/commands/generate-titles.md @@ -1,6 +1,6 @@ --- description: Generate 8 verified title options for a clip, moment, or episode -allowed-tools: Read +allowed-tools: Read, mcp__podcli__knowledge_base argument-hint: [clip-transcript-or-moment-brief] triggers: - titles for diff --git a/.claude/commands/plan-episode.md b/.claude/commands/plan-episode.md index b90f3e12..77d030c7 100644 --- a/.claude/commands/plan-episode.md +++ b/.claude/commands/plan-episode.md @@ -1,6 +1,6 @@ --- description: Design questions, story arc, and moment map BEFORE recording an episode -allowed-tools: Read, Write +allowed-tools: Read, Write, mcp__podcli__knowledge_base argument-hint: [guest-name-and-company] triggers: - plan episode diff --git a/.claude/commands/plan-thumbnails.md b/.claude/commands/plan-thumbnails.md index a665161c..6008e938 100644 --- a/.claude/commands/plan-thumbnails.md +++ b/.claude/commands/plan-thumbnails.md @@ -66,6 +66,7 @@ What is the single most compelling image or concept? - Guest photo requirements - Background suggestion - Special visual elements +- **Layout:** `single` (one face) or `pair` (two people). Propose `pair` for interview clips where the exchange is the point: a question and its answer, a disagreement, a reaction. podcli takes both faces from the clip itself, guest on the left and host on the right. Set it with `manage_thumbnail_config` (`set_layout`), the "Two people" toggle on the clip page, or `podcli thumbnails --layout pair`. Add `--swap` to flip sides, or `--left-image` and `--right-image` to name the people. It falls back to `single` when podcli cannot tell two people apart, so keep a single-face brief ready. ### Step 4: Quality Check - [ ] Readable at phone screen size @@ -88,6 +89,7 @@ What is the single most compelling image or concept? **Shorts (9:16):** - Text: "[LINE 1] / [LINE 2 — accent]" +- Layout: [single / pair: guest left, host right] - Visual: [action shot / dramatic imagery / B-roll] - Text position: Lower third, centered diff --git a/.claude/commands/process-transcript.md b/.claude/commands/process-transcript.md index bf609fb9..7426bf50 100644 --- a/.claude/commands/process-transcript.md +++ b/.claude/commands/process-transcript.md @@ -1,6 +1,6 @@ --- description: Extract, score, and classify the best moments from a raw podcast transcript -allowed-tools: Read, Write +allowed-tools: Read, Write, mcp__podcli__knowledge_base argument-hint: [transcript-file-or-paste] triggers: - transcript @@ -96,6 +96,8 @@ Before scoring, fix each flagged moment's edges and state its payoff. **Then run the standalone check.** Name what the viewer must already know. If it is anything other than nothing, the start moves back until the clip covers it. If it cannot, drop the moment. +**Optionally mark an opening hook.** If the sharpest line sits mid-clip, it can play first as a `hook` (1-15 seconds, `repeat` replays it in place, `move` lifts it out). It must be a line actually spoken inside the clip, quoted verbatim with its timestamps. Never invent hook text. + ### Phase 4: Score Each Moment For every flagged moment, score on four dimensions (1-5 each): @@ -182,6 +184,7 @@ Format: comma-separated, under 500 characters. **Payoff:** [What the viewer walks away with. One sentence, second person.] **Needs:** [nothing | what the viewer must already know] **Setup line:** [The question this answers, in one line, or omit when the clip carries its own setup] +**Opening hook:** [Optional. "Verbatim line" XX:XX-XX:XX, repeat or move. Omit when the clip opens strong on its own] **Why it works:** [One sentence explaining the appeal] diff --git a/.claude/commands/produce-shorts.md b/.claude/commands/produce-shorts.md index cb1f1aa5..682f9293 100644 --- a/.claude/commands/produce-shorts.md +++ b/.claude/commands/produce-shorts.md @@ -1,6 +1,6 @@ --- description: Full pipeline from transcript to publish-ready content package -allowed-tools: Read, Write, Edit, Task +allowed-tools: Read, Write, Edit, Task, mcp__podcli__knowledge_base, mcp__podcli__get_ui_state, mcp__podcli__record_decisions argument-hint: [transcript-file-or-episode-number] triggers: - process episode @@ -29,6 +29,8 @@ Read the full knowledge base with the `knowledge_base` MCP tool: - `07-thumbnail-guide.md` — visual specs - `13-learnings.md` — past retro patterns (what worked, what didn't) +If a video is already set (`get_ui_state`), check its `OPEN QUESTIONS` block for unanswered episode decisions (clip count, duration range, captions, language, thumbnails, delivery target). Ask whichever are relevant to this run, batched into one NEEDS_INPUT prompt rather than one per field, then call `record_decisions(video_path, ...)` so the next run against this video doesn't ask again. + --- ## Inputs @@ -52,6 +54,8 @@ Each phase calls the corresponding skill's logic. Each phase reports its own Com Extract guest info, flag 15-20 moments, anchor each one (boundaries by meaning, question pulled in or carried as a setup line, payoff written before any title), score them, select top moments, classify by content type, check for duplicates. +A moment may open with a `hook`: a 1-15 second line spoken inside the clip, played first. Quote it from the transcript. Never invent it. + ### Phase 2: Title Development *Runs `/generate-titles` logic per moment* diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index aa821459..dabe6200 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -87,6 +87,18 @@ jobs: - run: node scripts/gen-docs-manifest.mjs - run: node scripts/check-docs-drift.mjs + release-hygiene: + name: Release hygiene check + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 + with: + node-version: "20" + - run: node scripts/release-hygiene.mjs + python: name: Python tests runs-on: ${{ matrix.os }} @@ -102,6 +114,17 @@ jobs: with: python-version: "3.12" cache: "pip" + - name: Install ffmpeg + # Without this, every test that shells out to ffmpeg/ffprobe skips. + # About 22 tests were silently not running on any OS. + shell: bash + run: | + case "${{ runner.os }}" in + Linux) sudo apt-get update -qq && sudo apt-get install -y -qq ffmpeg ;; + macOS) brew install --quiet ffmpeg ;; + Windows) choco install ffmpeg -y --no-progress ;; + esac + ffmpeg -version | head -1 - name: Install minimal test deps # Skip whisper/torch — tests don't exercise them and they balloon CI time. run: | diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 00000000..c7b0d651 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,17 @@ +# AGENTS.md + +This file exists for coding agents (OpenAI Codex, opencode, Aider, Cursor Agent, +and others) that read `AGENTS.md` by convention. + +- `CLAUDE.md` is the primary instruction document: project layout, the MCP tool + table, the knowledge base, and the quality gate all live there and are not + repeated here. +- `AGENTS.podstack.md` covers cross-tool PodStack usage: how each host runs the + content-production commands (`/plan-episode`, `/process-transcript`, + `/generate-titles`, and the rest), and where each host installs them. +- `.claude/commands/*.md` are the PodStack command sources. Claude Code reads + them directly from that path; `podcli auto` (and the other PodStack + commands) also installs them as Codex skills under `~/.codex/skills/` when + the `codex` CLI is present. + +Start with `CLAUDE.md`, then `AGENTS.podstack.md` for command-by-command detail. diff --git a/AGENTS.podstack.md b/AGENTS.podstack.md index a1bae009..c21a55ec 100644 --- a/AGENTS.podstack.md +++ b/AGENTS.podstack.md @@ -8,7 +8,7 @@ PodStack turns your AI tool into a podcast content team: Episode Architect, Cont ## How to use -Each skill below is a self-contained instruction file in `commands/` (or `.claude/commands/`, `.codex/prompts/`, `.cursor/rules/`, `.opencode/commands/`, depending on which host installed it). +Each skill below is a self-contained instruction file in `commands/` (or `.claude/commands/`, `~/.codex/skills//SKILL.md`, `.cursor/rules/`, `.opencode/commands/`, depending on which host installed it). **To run a skill:** ask your agent to "run the [skill-name] skill" or invoke its slash command (`/[skill-name]`) where supported. The agent opens the corresponding file and follows it step by step. @@ -30,7 +30,7 @@ Each skill below is a self-contained instruction file in `commands/` (or `.claud - **role:** Episode Architect - **description:** Design questions, story arc, and moment map BEFORE recording -- **allowed-tools:** Read, Write +- **allowed-tools:** Read, Write, mcp__podcli__knowledge_base - **triggers:** plan episode, upcoming recording, guest prep, prepare for interview - **outputs:** episode plan written to `episodes/ep[XX]-[guest]-plan.md` - **next:** record → `/process-transcript` @@ -39,7 +39,7 @@ Each skill below is a self-contained instruction file in `commands/` (or `.claud - **role:** Content Analyst - **description:** Extract, score, classify best moments from a raw transcript -- **allowed-tools:** Read, Write +- **allowed-tools:** Read, Write, mcp__podcli__knowledge_base - **triggers:** transcript, process transcript, extract moments, podcast transcript - **outputs:** moment brief with timestamps, scores, titles, thumbnails, descriptions - **next:** `/generate-titles` or `/produce-shorts` @@ -48,7 +48,7 @@ Each skill below is a self-contained instruction file in `commands/` (or `.claud - **role:** Title Writer - **description:** Generate 8 verified title options for a clip or moment -- **allowed-tools:** Read +- **allowed-tools:** Read, mcp__podcli__knowledge_base - **triggers:** titles for, title options, write titles, generate titles - **outputs:** 8 titles + 2 top picks with rationale @@ -80,7 +80,7 @@ Each skill below is a self-contained instruction file in `commands/` (or `.claud - **role:** Producer (master orchestrator) - **description:** Full pipeline from transcript to publish-ready content package -- **allowed-tools:** Read, Write, Edit, Task +- **allowed-tools:** Read, Write, Edit, Task, mcp__podcli__knowledge_base - **triggers:** process episode, produce shorts, full pipeline, prep episode, make content package - **outputs:** complete content package in `episodes/ep[XX]-[guest]-content-package.md` - **orchestrates:** process-transcript → generate-titles → generate-descriptions → plan-thumbnails → review-content @@ -114,16 +114,16 @@ Skill files read the 14 knowledge files at `.podcli/knowledge/`; the full file t PodStack ships one source-of-truth (`commands/`) and installs to the right location for each tool: -| Host | Install location | Primary doc | -|------|-----------------|-------------| -| Claude Code | `.claude/commands/*.md` | `CLAUDE.md` | -| OpenAI Codex | `.codex/prompts/*.md` | `AGENTS.podstack.md` (this file) | -| Cursor | `.cursor/rules/*.mdc` | `AGENTS.podstack.md` | -| opencode | `.opencode/commands/*.md` | `AGENTS.podstack.md` | -| Generic | `commands/*.md` | `AGENTS.podstack.md` | +| Host | Install location | Installed by | Primary doc | +|------|-----------------|--------------|-------------| +| Claude Code | `.claude/commands/*.md` (per project) | `podcli auto` / any PodStack command | `CLAUDE.md` | +| OpenAI Codex | `~/.codex/skills//SKILL.md` (global) | `podcli auto` / any PodStack command, when the `codex` CLI is on PATH | `AGENTS.podstack.md` (this file) | +| Cursor | `.cursor/rules/*.mdc` | not automated yet, copy by hand | `AGENTS.podstack.md` | +| opencode | `.opencode/commands/*.md` | not automated yet, copy by hand | `AGENTS.podstack.md` | +| Generic | `commands/*.md` | not automated yet, copy by hand | `AGENTS.podstack.md` | -These command files ship with podcli; place the set for your tool (left column) in -its command dir. See `README.md` for per-host usage examples. +Claude and Codex installs are automatic and kept in sync on upgrade; see `README.md` +for per-host usage examples and manual steps for the other hosts. --- diff --git a/CLAUDE.md b/CLAUDE.md index a6332ba6..644a417d 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -31,7 +31,7 @@ Both share the same knowledge base at `.podcli/knowledge/`. ## MCP tools (podcli engine) -All 27 tools registered by the MCP server. +All 30 tools registered by the MCP server. **Transcription and input** @@ -43,6 +43,8 @@ All 27 tools registered by the MCP server. | `set_video` | Set the working video without transcribing | | `import_transcript` | Import an external transcript with word-level timestamps, skips Whisper | | `parse_transcript` | Parse a speaker-labeled plain text transcript into word-level timestamps | +| `compare_transcription_engines` | Transcribe one sample with two engines and write a side-by-side disagreement report | +| `mine_channel` | List a YouTube channel's uploads or mine one video's existing captions, without downloading the video | **Clip workflow** @@ -56,6 +58,7 @@ All 27 tools registered by the MCP server. | `batch_create_clips` | Render multiple clips in one batch | | `manage_reel` | Build a highlights reel: detect once, edit moments, rebuild without re-detecting | | `analyze_energy` | Analyze audio energy levels to find high-energy moments | +| `record_decisions` | Record per-episode decisions (clip count, duration range, captions, language, thumbnails, delivery target) so later runs never re-ask | **Full-episode editing** diff --git a/README.md b/README.md index 0dce7df5..1776c7ae 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ Clips land in `podcli-clips/` in the directory you ran it from, so each show kee **Shipping it** -- 27 MCP tools, so an agent can transcribe, score, render, and publish through conversation +- 30 MCP tools, so an agent can transcribe, score, render, and publish through conversation - YouTube publishing plus performance analytics to see which clips landed - DaVinci Resolve export as FCPXML when you want to finish by hand - Presets, clip history with duplicate detection, and a transcript cache @@ -108,7 +108,7 @@ If you are weighing podcli against the cloud clippers, this is the difference: - Runs locally. Transcription and rendering happen on your machine by default, so episodes stay there. Only the optional cloud engine (AssemblyAI) and publishing to YouTube send anything out. - Free and open source under AGPL-3.0. Exports are unlimited, full quality, and watermark-free. -- Agent-native. 27 MCP tools let Claude Code or Codex drive the whole flow, transcription through publishing. +- Agent-native. 30 MCP tools let Claude Code or Codex drive the whole flow, transcription through publishing. - A knowledge base keeps titles, captions, and descriptions in your show's voice, and stops the engine from resuggesting moments you already published. - DaVinci Resolve handoff. Export any clip as FCPXML when you want to finish the edit yourself. @@ -117,7 +117,7 @@ If you are weighing podcli against the cloud clippers, this is the difference: podcli is an [MCP](https://modelcontextprotocol.io) server, so an agent can transcribe, suggest clips, and render them through conversation. ```bash -podcli mcp install # registers it with Claude Code +podcli mcp install # registers it with Claude Code and Codex, whichever CLI is on PATH ``` Claude Desktop and Codex setup is in the [MCP docs](https://podcli.com/docs/mcp-server). diff --git a/backend/cli.py b/backend/cli.py index 2d05929e..c7ea09c1 100644 --- a/backend/cli.py +++ b/backend/cli.py @@ -124,16 +124,20 @@ def _json_object_arg(raw: str | None, name: str): return parsed -def _cached_face_map(video_path: str): - """Face maps are keyed by video content, not by transcript, so an imported - transcript can still borrow the map from an earlier run on the same file.""" +def _cached_transcript(video_path: str) -> dict: + """The cached transcript for this video, or {} when there is none.""" try: from services.transcript_packer import load_cached_transcript_for_video - cached = load_cached_transcript_for_video(video_path) + return load_cached_transcript_for_video(video_path) or {} except Exception: - return None - return (cached or {}).get("face_map") + return {} + + +def _cached_face_map(video_path: str): + """Face maps are keyed by video content, not by transcript, so an imported + transcript can still borrow the map from an earlier run on the same file.""" + return _cached_transcript(video_path).get("face_map") def _suggestions_session_path(cache_hash: str) -> str: @@ -557,6 +561,8 @@ def _pass_through(value): cmd += ["--progress-color", args.progress_color] if getattr(args, "cards", None): cmd += ["--cards", args.cards] + if getattr(args, "hook", None): + cmd += ["--hook", args.hook] if getattr(args, "brand", None): cmd += ["--brand", args.brand] if getattr(args, "style", None): @@ -798,7 +804,7 @@ def _open_multicam_session(target: str, people): session = mc.new_session(folder=target, people=people, progress_callback=report) done() else: - session = mc.MulticamSession.load(target) + session = mc.open_session(target) # Reopening keeps the saved names; --people on a re-run renames them. if people and people != [p.name for p in session.people]: session = mc.rename_people(session, people) @@ -994,10 +1000,12 @@ def _run_multicam(args, mc, target: str): def _render_multicam(args, mc, session): report, done = _multicam_progress("Rendering") - outputs = mc.render_session(session, stems=not args.no_stems, progress_callback=report) + outputs = mc.render_session(session, stems=not args.no_stems, validate=args.validate, progress_callback=report) done() for path in [outputs["video"], *(outputs.get("stems") or [])]: print(f" ✓ {path}") + for warning in (outputs.get("validation") or {}).get("warnings") or []: + print(f" ! {warning}", file=sys.stderr) def _pull_multicam(args, mc, target: str): @@ -1296,7 +1304,12 @@ def cmd_process(args): if skip_transcript: # Reuse an existing transcript so highlight boundaries snap to whole sentences; # only skip transcription outright when there is none (true no-dialogue footage). - cached = load_cached_transcript_for_video(video_path) + cached = load_cached_transcript_for_video( + video_path, + engine=os.environ.get("PODCLI_ENGINE"), + model=config.get("whisper_model", "base"), + language=config.get("language"), + ) if cached and not config.get("no_cache", False): words = cached["words"] segments = cached["segments"] @@ -1337,8 +1350,16 @@ def cmd_process(args): else: print(" No cached face map, crop falls back to per-clip face tracking") elif not skip_transcript: - # Check cache first - cached = load_cached_transcript_for_video(video_path) + # Check cache first, with the same (engine, model, language) the + # transcribe call below would run with, so a hit here is guaranteed + # to be the combo this invocation actually asked for, not a + # different one that happens to share the file. + cached = load_cached_transcript_for_video( + video_path, + engine=os.environ.get("PODCLI_ENGINE"), + model=config.get("whisper_model", "base"), + language=config.get("language"), + ) if cached and not config.get("no_cache", False): print(" [1/4] Loaded from cache (instant)") words = cached["words"] @@ -1398,8 +1419,16 @@ def _transcribe_progress(pct, msg): segments = result["segments"] print(f" Done: {len(segments)} segments, {len(words)} words") - # Save to cache for next run - save_cached_transcript_for_video(video_path, result) + # Save to cache for next run, under the same key the read above + # checked. result["engine"] is what actually ran, which can + # differ from the env var on a fallback (see transcribe_file). + save_cached_transcript_for_video( + video_path, + result, + engine=result.get("engine") or os.environ.get("PODCLI_ENGINE"), + model=config.get("whisper_model", "base"), + language=config.get("language"), + ) # Apply word corrections (Whisper misheard proper nouns, brand names) from services.corrections import apply_corrections @@ -1668,7 +1697,7 @@ def _transcribe_progress(pct, msg): _thumb_photo = None if _thumb_enabled: try: - from services.thumbnail_ai import generate_variations as _tv, thumbnail_to_video_frame as _ttv + from services.thumbnail_ai import render_variations as _tv, thumbnail_to_video_frame as _ttv _thumb_gen = _tv _thumb_to_video = _ttv _thumb_logo = config.get("logo_path") or None @@ -1718,6 +1747,7 @@ def _transcribe_progress(pct, msg): motion=config.get("motion"), bookend_fade=config.get("bookend_fade", 0.0), keep_segments=clip.get("segments"), + hook=_clip_hook(clip), face_map=face_map, allow_ass_fallback=config.get("allow_ass_fallback", False), use_ass_captions=config.get("use_ass_captions", False), @@ -1741,7 +1771,7 @@ def _transcribe_progress(pct, msg): output_path=os.path.join(clip_thumb_dir, "_lead_frame.jpg"), start_second=result.get("start_second", clip.get("start_second", 0)), ) - thumb_paths = _thumb_gen( + rendered_thumbs = _thumb_gen( title=clip.get("title", f"Clip {i+1}"), output_dir=clip_thumb_dir, photo_path=lead_frame or _thumb_photo, @@ -1750,7 +1780,14 @@ def _transcribe_progress(pct, msg): end_second=result.get("end_second", clip.get("end_second")), logo_path=_thumb_logo, config=_thumb_style, + grounding=_clip_grounding(clip), + face_map=face_map, + segments=segments, ) + thumb_paths = rendered_thumbs["paths"] + _pair = rendered_thumbs["pair"] + if _pair and _pair["layout"] != "pair": + print(f" ℹ {_pair['reason']}") if thumb_paths and _thumb_placement == "off": print(f" + {len(thumb_paths)} thumbnail(s) in " f"{os.path.basename(clip_thumb_dir)}/") @@ -1945,6 +1982,7 @@ def _transcribe_progress(pct, msg): outro_path=config.get("outro_path") or None, intro_path=config.get("intro_path") or None, keep_segments=clip.get("segments"), + hook=_clip_hook(clip), face_map=face_map, allow_ass_fallback=config.get("allow_ass_fallback", False), use_ass_captions=config.get("use_ass_captions", False), @@ -2196,6 +2234,24 @@ def _review_clips(clips: list, segments: list, energy_scores: list | None, confi print(f" No additional suggestions found.") +def _clip_hook(clip: dict) -> dict | None: + """The clip's opening hook when it still fits the clip, else None. + + Review can move a clip's edges after the hook was proposed. A hook left + outside the body would fail the whole render, so it is dropped with a note + and the clip renders without it. + """ + from services.opening_hook import validate_hook + + if not clip.get("hook"): + return None + try: + return validate_hook(clip["hook"], clip["start_second"], clip["end_second"], clip.get("segments")) + except ValueError as e: + print(f" Opening hook dropped: {e}") + return None + + def _filter_duplicate_clip_suggestions(candidates: list, existing: list, overlap_threshold: float = 5.0) -> list: """Drop suggestions that significantly overlap already-selected clips.""" filtered = [] @@ -2272,6 +2328,7 @@ def _rerender_clip(r): outro_path=config.get("outro_path") or None, intro_path=config.get("intro_path") or None, keep_segments=clip.get("segments"), + hook=_clip_hook(clip), face_map=face_map, allow_ass_fallback=config.get("allow_ass_fallback", False), use_ass_captions=config.get("use_ass_captions", False), @@ -2459,6 +2516,7 @@ def _rerender_clip(r): outro_path=config.get("outro_path") or None, intro_path=config.get("intro_path") or None, keep_segments=f_clip.get("segments"), + hook=_clip_hook(f_clip), face_map=face_map, allow_ass_fallback=config.get("allow_ass_fallback", False), use_ass_captions=config.get("use_ass_captions", False), @@ -2516,6 +2574,7 @@ def _rerender_clip(r): outro_path=config.get("outro_path") or None, intro_path=config.get("intro_path") or None, keep_segments=nc.get("segments"), + hook=_clip_hook(nc), face_map=face_map, allow_ass_fallback=config.get("allow_ass_fallback", False), use_ass_captions=config.get("use_ass_captions", False), @@ -3130,12 +3189,30 @@ def cmd_thumbnail_config(args): raise ValueError(f"unknown thumbnail-config action: {action}") +def _clip_grounding(clip: dict) -> dict | None: + """A suggestion's payoff, question and opening line, for grounding thumbnail copy.""" + grounding = {k: clip.get(k) for k in ("payoff", "context_line", "preview_text")} + return grounding if any(grounding.values()) else None + + +def _grounding_from_args(args) -> dict | None: + """Collect the clip's payoff/question/opening-line CLI flags into the dict + thumbnail_ai expects, or None if the caller passed none of them (e.g. a + bare title with no clip behind it, as in the standalone thumbnail studio).""" + grounding = { + "payoff": getattr(args, "payoff", None), + "context_line": getattr(args, "context_line", None), + "preview_text": getattr(args, "preview_text", None), + } + return grounding if any(grounding.values()) else None + + def cmd_thumbnail_options(args): """Emit candidate headline text pairs and face frames for the thumbnail picker.""" from services.thumbnail_ai import generate_headline_variations, extract_candidate_frames os.makedirs(args.output, exist_ok=True) - texts = generate_headline_variations(args.title, args.texts) or [] + texts = generate_headline_variations(args.title, args.texts, grounding=_grounding_from_args(args)) or [] frames = [] if args.video: frames = extract_candidate_frames( @@ -3180,34 +3257,80 @@ def _check_frame(path): sys.exit(1) +def _pair_from_args(args, output_dir: str) -> dict | None: + """The two-person panels the thumbnail flags ask for, or None for one face. + + The template's own layout applies when --layout is not given. Speaker + turns and the face map come from the cached transcript of --video. + """ + from services.thumbnail_ai import _load_brand_config, resolve_pair + + video = getattr(args, "video", None) + cached = _cached_transcript(video) if video else {} + try: + return resolve_pair( + _load_brand_config(), output_dir, video, + getattr(args, "start", None), getattr(args, "end", None), + layout=getattr(args, "layout", None), + left_image=getattr(args, "left_image", None), + right_image=getattr(args, "right_image", None), + swap=bool(getattr(args, "swap", False)), + face_map=cached.get("face_map"), segments=cached.get("segments"), + ) + except ValueError as err: + print(f"thumbnail failed: {err}", file=sys.stderr) + sys.exit(1) + + +def _layout_report(pair: dict | None) -> dict: + """What a thumbnail result says about its layout: which people, from which source seconds.""" + if not pair: + return {"layout": "single"} + if pair["layout"] != "pair": + return {"layout": "single", "note": pair["reason"]} + return {"layout": "pair", "roles": pair["roles"], "swapped": pair["swapped"], "people": pair["people"]} + + def cmd_thumbnail_render(args): """Render one final thumbnail from a chosen frame + headline. Empty line1/line2 let the AI write the text; a chosen frame is used as-is. + With the pair layout the two people come from --video between --start and + --end, or from --left-image and --right-image, and the frame is the + fallback when two people cannot be told apart. """ from services.thumbnail_ai import generate_thumbnail_with_template from services.asset_store import resolve_logo - _check_frame(args.frame) - frame_info = json.loads(args.frame_info) if args.frame_info else None + pair = _pair_from_args(args, os.path.dirname(os.path.abspath(args.output))) + people = pair["people"] if pair and pair["layout"] == "pair" else None + if not people: + if not args.frame: + reason = f" {pair['reason']}" if pair else "" + print(f"thumbnail render failed: no frame to fall back on.{reason}", file=sys.stderr) + sys.exit(1) + _check_frame(args.frame) + frame_info = json.loads(args.frame_info) if args.frame_info and not people else None out = generate_thumbnail_with_template( title=args.title, - frame_path=args.frame, + frame_path=None if people else args.frame, output_path=args.output, logo_path=resolve_logo(args.logo) if args.logo else None, frame_info=frame_info, line1_override=args.line1 or None, line2_override=args.line2 or None, + grounding=_grounding_from_args(args), + people=people, ) if not out: print("thumbnail render failed", file=sys.stderr) sys.exit(1) - print(json.dumps({"path": out})) + print(json.dumps({"path": out, **_layout_report(pair)})) def cmd_thumbnails(args): """Generate thumbnail variations for a title.""" - from services.thumbnail_ai import generate_variations + from services.thumbnail_ai import render_variations from services.asset_store import resolve as resolve_asset, resolve_logo accent = "\033[38;2;212;135;74m" @@ -3251,23 +3374,41 @@ def cmd_thumbnails(args): print(f"\n {bold}Generating {args.variations} thumbnail variations...{reset}") print(f" Title: {accent}{args.title}{reset}") - paths = generate_variations( - title=args.title, - output_dir=args.output, - photo_path=photo, - video_path=video, - start_second=getattr(args, "start", None), - end_second=getattr(args, "end", None), - logo_path=logo, - config={"variations": args.variations}, - line1=getattr(args, "line1", None), - line2=getattr(args, "line2", None), - ) + cached = _cached_transcript(video) if video else {} + try: + rendered = render_variations( + title=args.title, + output_dir=args.output, + photo_path=photo, + video_path=video, + start_second=getattr(args, "start", None), + end_second=getattr(args, "end", None), + logo_path=logo, + config={"variations": args.variations}, + line1=getattr(args, "line1", None), + line2=getattr(args, "line2", None), + layout=getattr(args, "layout", None), + left_image=getattr(args, "left_image", None), + right_image=getattr(args, "right_image", None), + swap=bool(getattr(args, "swap", False)), + face_map=cached.get("face_map"), + segments=cached.get("segments"), + ) + except ValueError as err: + print(f" {red}✗{reset} {err}", file=sys.stderr) + sys.exit(1) + paths = rendered["paths"] + report = _layout_report(rendered["pair"]) if as_json: - print(json.dumps({"paths": paths})) + print(json.dumps({"paths": paths, **report})) return + if report.get("note"): + print(f" {gray}{report['note']}{reset}") + for person in report.get("people", []): + when = f" at {person['source_time']:.1f}s" if person.get("source_time") is not None else "" + print(f" {gray}{person['side'].capitalize()}: {person.get('role') or 'person'}{when}{reset}") for p in paths: print(f" {green}✓{reset} {p}") print(f"\n {gray}Open the folder to preview and pick the best one.{reset}\n") @@ -4088,6 +4229,40 @@ def cmd_cache(args): print(f" {gray}Run {accent}podcli cache clear{reset} {gray}to delete all{reset}\n") +def cmd_compare_engines(args): + """Transcribe the same sample window with two engines and report where they disagree.""" + from services.engine_comparison import compare_engines + + accent = "\033[38;2;212;135;74m" + gray = "\033[38;5;245m" + reset = "\033[0m" + + if not os.path.exists(args.video): + print(f"podcli: file not found: {args.video}", file=sys.stderr) + sys.exit(1) + + print(f"Comparing {accent}{args.engine_a}{reset} vs {accent}{args.engine_b}{reset} " + f"on {args.duration:.0f}s starting at {args.start:.0f}s...") + + report = compare_engines( + args.video, + args.engine_a, + args.engine_b, + start_seconds=args.start, + duration_seconds=args.duration, + window_seconds=args.window, + model_size=args.model_size, + language=args.language, + output_dir=args.output, + ) + + print(f"\nOverall disagreement: {accent}{report['overall_disagreement']:.2f}{reset} " + f"(0 = identical output, 1 = completely different; not accuracy against a transcript)") + if report["both_empty_window_count"]: + print(f"{gray}{report['both_empty_window_count']} window(s) had no words from either engine{reset}") + print(f"{gray}Wrote {report['json_path']} and {report['html_path']}{reset}") + + def cmd_info(args): """Show system info.""" from services.encoder import get_encoder_info @@ -4820,6 +4995,15 @@ def _first_run_setup() -> bool: return True +def _add_pair_args(parser) -> None: + parser.add_argument("--layout", choices=["single", "pair"], + help="single: one face. pair: the guest left and the host right, from the clip's " + "footage. Defaults to the template's layout") + parser.add_argument("--left-image", dest="left_image", help="Image of the person on the left (pair layout)") + parser.add_argument("--right-image", dest="right_image", help="Image of the person on the right (pair layout)") + parser.add_argument("--swap", action="store_true", help="Swap the two people's sides (pair layout)") + + def main(): parser = argparse.ArgumentParser( prog="podcli", @@ -4856,7 +5040,7 @@ def main(): help="Render on podcli.com instead of this machine (needs `podcli login`)") proc.add_argument("--template-id", help="Cut in a saved cloud template, by id (with --cloud)") - proc.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai"], help="Transcription engine (default: whisper-py; whispercpp is local; assemblyai uses ASSEMBLYAI_API_KEY)") + proc.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai", "omnilingual"], help="Transcription engine (default: whisper-py; whispercpp is local; assemblyai uses ASSEMBLYAI_API_KEY)") proc.add_argument("--language", help="Language of the recording (e.g. es, pt-BR, ka). Auto-detect if omitted.") proc.add_argument("--assemblyai-api-key", help="AssemblyAI API key for --engine assemblyai. Prefer ASSEMBLYAI_API_KEY; command-line secrets can appear in process listings.") proc.add_argument("--fast", action="store_true", help="Draft mode: tiny Whisper, heuristic selection, center crop, low quality") @@ -4996,6 +5180,10 @@ def main(): "File > Export > Timeline > FCP 7 XML)") mc_p.add_argument("--no-render", action="store_true", dest="no_render", help="Skip the MP4 render (fast, export only)") mc_p.add_argument("--no-stems", action="store_true", dest="no_stems", help="Skip the per-person WAV files") + mc_p.add_argument("--validate", choices=["sample", "full"], default="sample", + help="How hard to check the rendered episode decodes cleanly: 'sample' (default) checks the " + "first and last 10s plus a few points in between; 'full' decodes the whole thing, which " + "costs minutes per hour of 1080p") mc_p.add_argument("--resync", action="store_true", help="Sync every file again, including offsets you set by hand") mc_p.add_argument("-y", "--yes", action="store_true", help="Don't stop to review guessed roles") @@ -5011,7 +5199,7 @@ def main(): mc_p.add_argument("--transcript", action="store_true", help="Transcribe the episode with each word credited to whoever's mic was speaking") mc_p.add_argument("--model", default="base", help="Whisper model for --transcript (default base)") - mc_p.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai"], + mc_p.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai", "omnilingual"], help="Transcription engine for --transcript (default: the one podcli is set up with)") mc_p.add_argument("--activity", action="store_true", help="Report who speaks when (talk time, or spans with --json)") mc_p.add_argument("--cloud", action="store_true", @@ -5028,7 +5216,7 @@ def main(): studio.add_argument("--end", type=float, help="Fragment end (seconds)") studio.add_argument("--paragraph", help="Find the fragment by matching this text in the transcript") studio.add_argument("--language", help="Transcription language (e.g. es). Auto-detect if omitted.") - studio.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai"], help="Transcription engine") + studio.add_argument("--engine", choices=["whisper-py", "whispercpp", "assemblyai", "omnilingual"], help="Transcription engine") studio.add_argument("--transcript", help="Word timings JSON for this video ({words:[...]} or a list); skips transcription") studio.add_argument("--assemblyai-api-key", help="AssemblyAI API key for --engine assemblyai. Prefer ASSEMBLYAI_API_KEY; command-line secrets can appear in process listings.") studio.add_argument("--caption-style", choices=["hormozi", "karaoke", "subtle", "branded", "outline"], default="hormozi") @@ -5056,6 +5244,10 @@ def main(): help="Draw how much of the clip is left along the bottom edge") studio.add_argument("--progress-color", dest="progress_color") studio.add_argument("--cards", help="On-screen cards as JSON, each with kind/start/end") + studio.add_argument("--hook", help='Opening hook as JSON, on the source clock: ' + '{"start":41.2,"end":45.8,"mode":"repeat"}. ' + "Plays a 1-15 s passage from inside the fragment first; " + '"move" lifts it out of the body') studio.add_argument("--brand", help="Show colours as JSON: " '{"accent":"#4C9DF5","ink":"#FFFFFF","surface":"#0A0D14"}') studio.add_argument("--style", help='Visual theme as JSON: ' @@ -5176,6 +5368,7 @@ def main(): thumb.add_argument("--line1", help="Explicit first thumbnail line (skips AI rewrite)") thumb.add_argument("--line2", help="Explicit second thumbnail line") thumb.add_argument("--json", action="store_true", help="Emit JSON {paths:[...]} to stdout") + _add_pair_args(thumb) # ── thumbnail-config ── tcfg = sub.add_parser("thumbnail-config", help="Show, export, import, or reset the thumbnail template") @@ -5196,16 +5389,27 @@ def main(): topt.add_argument("--end", type=float, help="Frame window end (seconds)") topt.add_argument("--texts", type=int, default=6, help="Number of headline options") topt.add_argument("--frames", type=int, default=6, help="Number of frame options") + topt.add_argument("--payoff", help="Clip's payoff line, so headline copy is grounded in it rather than the title alone") + topt.add_argument("--context-line", dest="context_line", help="The question this clip answers, if any") + topt.add_argument("--preview-text", dest="preview_text", help="Clip's verbatim opening line") # ── thumbnail-render (one final thumbnail from a chosen frame + headline) ── trnd = sub.add_parser("thumbnail-render", help="Render one thumbnail PNG from a chosen frame + headline") trnd.add_argument("title", help="Clip/episode title") - trnd.add_argument("--frame", required=True, help="Background frame image path") + trnd.add_argument("--frame", help="Background frame image path. Optional with the pair layout, " + "where it is the fallback when two people cannot be told apart") + trnd.add_argument("--video", help="Source video the pair layout takes both people from") + trnd.add_argument("--start", type=float, help="Clip start in --video (seconds)") + trnd.add_argument("--end", type=float, help="Clip end in --video (seconds)") + _add_pair_args(trnd) trnd.add_argument("-o", "--output", required=True, help="Destination PNG path") trnd.add_argument("--line1", help="Headline line 1 (empty = AI writes it)") trnd.add_argument("--line2", help="Headline line 2 (empty = AI writes it)") trnd.add_argument("--frame-info", dest="frame_info", help="JSON face metadata for the frame") trnd.add_argument("--logo", help="Logo (asset name or path)") + trnd.add_argument("--payoff", help="Clip's payoff line, so headline copy is grounded in it rather than the title alone") + trnd.add_argument("--context-line", dest="context_line", help="The question this clip answers, if any") + trnd.add_argument("--preview-text", dest="preview_text", help="Clip's verbatim opening line") # ── swap-thumbnail ── st = sub.add_parser("swap-thumbnail", help="Regenerate thumbnail on an existing clip") @@ -5324,6 +5528,21 @@ def main(): # ── info ── sub.add_parser("info", help="Show system info (encoder, etc.)") + # ── compare-engines ── + cmp_p = sub.add_parser( + "compare-engines", + help="Transcribe the same sample window with two engines and report where they disagree", + ) + cmp_p.add_argument("video", help="Path to podcast video/audio file") + cmp_p.add_argument("engine_a", choices=["whisper-py", "whispercpp", "assemblyai", "omnilingual"]) + cmp_p.add_argument("engine_b", choices=["whisper-py", "whispercpp", "assemblyai", "omnilingual"]) + cmp_p.add_argument("--start", type=float, default=0.0, help="Sample start, seconds into the source (default: 0)") + cmp_p.add_argument("--duration", type=float, default=120.0, help="Sample length in seconds (default: 120)") + cmp_p.add_argument("--window", type=float, default=20.0, help="Report window size in seconds (default: 20)") + cmp_p.add_argument("--model-size", default="base", help="Model size for engines that take one (default: base)") + cmp_p.add_argument("--language", help="ISO language code. Auto-detect if omitted.") + cmp_p.add_argument("-o", "--output", default="./engine-comparison", help="Output directory for comparison.json/.html") + init_thumb = sub.add_parser( "init-thumbnail", help="Scaffold .podcli/thumbnail-config.json so podcli generates thumbnails for you", @@ -5404,6 +5623,8 @@ def main(): cmd_cache(args) elif args.command == "info": cmd_info(args) + elif args.command == "compare-engines": + cmd_compare_engines(args) elif args.command == "init-thumbnail": cmd_init_thumbnail(args) elif args.command in ("ui", "webui"): diff --git a/backend/clip_studio.py b/backend/clip_studio.py index 02948df8..c4926dbe 100644 --- a/backend/clip_studio.py +++ b/backend/clip_studio.py @@ -196,6 +196,23 @@ def _json_file(raw, name): return None +def _hook_arg(raw, start, end): + """The --hook object checked against the fragment, or nothing. + + Checked as soon as the fragment's range is known, so a bad hook stops the + run before the face scan and the render. It stops rather than drops, + because someone asked for it. + """ + hook = _json_object_arg(raw, "--hook") + if hook is None: + return None + from services.opening_hook import validate_hook + try: + return validate_hook(hook, start, end) + except ValueError as exc: + raise SystemExit(f"Error: --hook {exc}") + + # The crops that cannot place a frame without knowing where the faces are. _WANTS_FACES = ("face", "speaker", "speaker-hardcut") @@ -249,7 +266,7 @@ def _render_fragment(video, start, end, words, style, crop, title, out_dir, fmt= logo=None, name_card=None, motion=None, caption_position="auto", caption_scale=1.0, logo_position="top-left", logo_scale=1.0, topic=None, progress=None, cards=None, brand=None, theme=None, font_family=None, - captions=True, face_map=None, crop_keyframes=None): + captions=True, face_map=None, crop_keyframes=None, hook=None): """Render the fragment with face-crop + captions via the existing engine.""" from services.clip_generator import generate_clip print(f" [fragment] rendering {start:.1f}s–{end:.1f}s ({style}, crop={crop}, {fmt})", flush=True) @@ -263,7 +280,7 @@ def _render_fragment(video, start, end, words, style, crop, title, out_dir, fmt= transcript_words=words, title=title, output_dir=out_dir, logo_path=logo, name_card=name_card, motion=motion, topic=topic, progress=progress, cards=cards, brand=brand, theme=theme, font_family=font_family, - captions=captions, + captions=captions, hook=hook, clean_fillers=True, allow_ass_fallback=True, progress_callback=lambda p, m: print(f" {p}% {m}", flush=True), ) @@ -388,6 +405,9 @@ def main(): help="Path to hand-placed crop positions as JSON. Used by --crop manual.") ap.add_argument("--cards", default=None, help="On-screen cards as JSON, each with kind/start/end") + ap.add_argument("--hook", default=None, + help='Opening hook as JSON on the source clock: {"start","end","mode"}. ' + 'mode is "repeat" or "move"') ap.add_argument("--brand", default=None, help='Show colours as JSON: {"accent":"#4C9DF5","ink":"#FFF","surface":"#000"}') ap.add_argument("--style", default=None, @@ -448,6 +468,8 @@ def main(): else: raise SystemExit("Provide either --start/--end or --paragraph") + hook = _hook_arg(args.hook, start, end) + # The canvas every part is rendered and stitched on. One lookup, so the # fragment, the bookends and the concat cannot disagree about the shape. from services.formats import get_format @@ -511,6 +533,7 @@ def media_path(value: str | None, kind: str) -> str | None: captions=not args.no_captions, face_map=face_map, crop_keyframes=_json_file(args.crop_keyframes, "--crop-keyframes"), + hook=hook, ) platforms = [p.strip() for p in platforms_str.split(",") if p.strip()] diff --git a/backend/main.py b/backend/main.py index a124f3b3..3b762165 100644 --- a/backend/main.py +++ b/backend/main.py @@ -67,15 +67,29 @@ def handle_ping(task_id: str, params: dict): emit_result(task_id, "success", data={"message": "pong", "version": VERSION}) +def handle_resolve_transcribe_engine(task_id: str, params: dict): + """Predict transcribe_file's engine resolution without transcribing, so + callers can build the right cache key before deciding whether to run it.""" + from services.transcription import resolve_engine_info + + emit_result( + task_id, + "success", + data=resolve_engine_info(params.get("engine"), params.get("model_size", "base")), + ) + + def handle_transcribe(task_id: str, params: dict): """Transcribe a podcast video/audio file with speaker detection.""" from services.transcription import transcribe_file from services.corrections import apply_corrections - from services.transcript_packer import compute_cache_hash, engine_cache_suffix, write_packed + from services.transcript_packer import compute_cache_hash, cache_key_suffix, write_packed emit_progress(task_id, "transcribing", 0, "Starting transcription...") file_path = params["file_path"] engine = params.get("engine") + model_size = params.get("model_size", "base") + language = params.get("language") previous_engine = os.environ.get("PODCLI_ENGINE") previous_assemblyai_key = os.environ.get("ASSEMBLYAI_API_KEY") if engine: @@ -83,14 +97,26 @@ def handle_transcribe(task_id: str, params: dict): if params.get("assemblyai_api_key"): os.environ["ASSEMBLYAI_API_KEY"] = params["assemblyai_api_key"] + start_seconds = params.get("start_seconds") + duration_seconds = params.get("duration_seconds") + # Matches src/handlers/transcribe.handler.ts and web-server.ts exactly: + # a sample is a *positive* window, not merely a present key. A caller + # that explicitly passes null (dropped by JSON on the TS side, but + # still visible here as None) or duration_seconds: 0 must get the full + # transcription path, the same as not passing the field at all. + is_sample = (duration_seconds or 0) > 0 or (start_seconds or 0) > 0 + # One shared 16 kHz mono wav feeds transcription, energy and reactions - # instead of decoding the source three times. + # instead of decoding the source three times. Skipped for a sample run: + # transcribe_file extracts its own trimmed window, and decoding the full + # source here would undo the whole point of a quick sample. shared_wav = None - try: - from services.audio_extract import extract_wav_16k_mono - shared_wav = extract_wav_16k_mono(file_path) - except Exception: - shared_wav = None + if not is_sample: + try: + from services.audio_extract import extract_wav_16k_mono + shared_wav = extract_wav_16k_mono(file_path) + except Exception: + shared_wav = None try: result = transcribe_file( @@ -100,46 +126,56 @@ def handle_transcribe(task_id: str, params: dict): language=params.get("language"), enable_diarization=params.get("enable_diarization", True), num_speakers=params.get("num_speakers"), + start_seconds=start_seconds, + duration_seconds=duration_seconds, progress_callback=lambda pct, msg: emit_progress(task_id, "transcribing", pct, msg), wav_path=shared_wav, ) # Apply word corrections (Whisper misheard proper nouns) apply_corrections(result.get("words", []), result.get("segments", [])) - energy_data = None - try: - from services.audio_analyzer import extract_audio_energy - energy_data = extract_audio_energy(file_path, wav_path=shared_wav) - except Exception: - pass # energy is a nice-to-have + # A sample is a throwaway language/quality check on a slice of the + # source. Energy/event signals and the packed view are keyed by the + # full file and meant to describe the whole episode, so skip them + # rather than caching partial (or source-wide-but-wrongly-expensive) + # data under those keys. + if not is_sample: + energy_data = None + try: + from services.audio_analyzer import extract_audio_energy + energy_data = extract_audio_energy(file_path, wav_path=shared_wav) + except Exception: + pass # energy is a nice-to-have - events_data = None - try: - from services.audio_events import extract_audio_events - events_data = extract_audio_events(file_path, wav_path=shared_wav) - except Exception: - pass # reactions are a nice-to-have + events_data = None + try: + from services.audio_events import extract_audio_events + events_data = extract_audio_events(file_path, wav_path=shared_wav) + except Exception: + pass # reactions are a nice-to-have - # Cached so clip suggestion reuses these instead of decoding the source again. - from services.signal_cache import save_signals - save_signals(file_path, energy_data=energy_data, events_data=events_data) + # Cached so clip suggestion reuses these instead of decoding the source again. + from services.signal_cache import save_signals + save_signals(file_path, energy_data=energy_data, events_data=events_data) - # Auto-pack: emit compact LLM-readable markdown alongside raw JSON. - # Pulls energy data so the packed view includes peak moments for clip reasoning. - try: - cache_hash = compute_cache_hash(file_path) + engine_cache_suffix(result.get("engine") or engine) - packed_path, packed_md = write_packed( - result, - cache_hash, - source_label=os.path.basename(file_path), - energy_data=energy_data, - events_data=events_data, - ) - result["packed_path"] = packed_path - result["packed_size_bytes"] = len(packed_md.encode("utf-8")) - except Exception as e: - # Non-fatal — transcription result is still useful without the packed view - emit_progress(task_id, "packing", 99, f"Packer skipped: {e}") + # Auto-pack: emit compact LLM-readable markdown alongside raw JSON. + # Pulls energy data so the packed view includes peak moments for clip reasoning. + try: + cache_hash = compute_cache_hash(file_path) + cache_key_suffix( + engine=result.get("engine") or engine, model=model_size, language=language + ) + packed_path, packed_md = write_packed( + result, + cache_hash, + source_label=os.path.basename(file_path), + energy_data=energy_data, + events_data=events_data, + ) + result["packed_path"] = packed_path + result["packed_size_bytes"] = len(packed_md.encode("utf-8")) + except Exception as e: + # Non-fatal: transcription result is still useful without the packed view + emit_progress(task_id, "packing", 99, f"Packer skipped: {e}") finally: if previous_engine is None: os.environ.pop("PODCLI_ENGINE", None) @@ -192,11 +228,13 @@ def handle_create_clip(task_id: str, params: dict): clean_fillers=params.get("clean_fillers", True), face_map=params.get("face_map"), keep_segments=params.get("keep_segments"), + hook=params.get("hook"), trim_opening=params.get("trim_opening"), preserve_timing=params.get("preserve_timing", False), allow_ass_fallback=params.get("allow_ass_fallback", False), use_ass_captions=params.get("use_ass_captions", False), keep_caption_overlay=params.get("keep_caption_overlay", False), + write_clean_variant=params.get("write_clean_variant", False), progress_callback=lambda pct, msg: emit_progress(task_id, "processing", pct, msg), ) emit_result(task_id, "success", data=result) @@ -258,11 +296,15 @@ def render_one(i: int, clip: dict) -> dict: clean_fillers=params.get("clean_fillers", True), face_map=params.get("face_map"), keep_segments=clip.get("keep_segments"), + hook=clip.get("hook"), allow_ass_fallback=clip.get("allow_ass_fallback", params.get("allow_ass_fallback", False)), use_ass_captions=clip.get("use_ass_captions", params.get("use_ass_captions", False)), keep_caption_overlay=clip.get( "keep_caption_overlay", params.get("keep_caption_overlay", False) ), + write_clean_variant=clip.get( + "write_clean_variant", params.get("write_clean_variant", False) + ), progress_callback=lambda pct, msg, _i=i: emit_progress( task_id, "batch", int((_i / total) * 100 + pct / total), msg ), @@ -331,13 +373,16 @@ def handle_parse_transcript(task_id: str, params: dict): raw_text = params.get("raw_text", "") total_duration = params.get("total_duration") time_adjust = params.get("time_adjust", 0.0) + language = params.get("language") if not raw_text: emit_result(task_id, "error", error="raw_text is required") return emit_progress(task_id, "parsing", 50, "Parsing transcript...") - result = detect_and_parse(raw_text, total_duration=total_duration, time_adjust=time_adjust) + result = detect_and_parse( + raw_text, total_duration=total_duration, time_adjust=time_adjust, language=language + ) if "error" in result: emit_result(task_id, "error", error=result["error"]) @@ -943,6 +988,40 @@ def handle_analyze_silence(task_id: str, params: dict): emit_result(task_id, "error", error=str(e)) +def handle_compare_engines(task_id: str, params: dict): + """Transcribe the same sample window with two engines and report where + they disagree. See services/engine_comparison.py for the scoring.""" + import time as _time + from config.paths import paths + from services.engine_comparison import compare_engines + + file_path = params.get("file_path", "") + if not file_path or not os.path.exists(file_path): + emit_result(task_id, "error", error=f"File not found: {file_path}") + return + + output_dir = params.get("output_dir") or os.path.join( + paths["output"], "engine-comparisons", f"{int(_time.time())}" + ) + try: + emit_progress(task_id, "comparing", 10, f"Transcribing with {params.get('engine_a')}...") + report = compare_engines( + file_path, + params.get("engine_a", "whispercpp"), + params.get("engine_b", "whisper-py"), + start_seconds=params.get("start_seconds", 0.0) or 0.0, + duration_seconds=params.get("duration_seconds", 120.0) or 120.0, + window_seconds=params.get("window_seconds", 20.0) or 20.0, + model_size=params.get("model_size", "base"), + language=params.get("language"), + output_dir=output_dir, + ) + emit_progress(task_id, "comparing", 100, "Comparison complete") + emit_result(task_id, "success", data=report) + except (FileNotFoundError, RuntimeError, ValueError) as e: + emit_result(task_id, "error", error=str(e)) + + def handle_render_silence_removed(task_id: str, params: dict): """Render the approved local cut plan and remap transcript timestamps.""" from config.paths import paths @@ -1006,7 +1085,7 @@ def progress(stage): emit_result(task_id, "success", data={"deleted": True, "session_id": session_id}) return - session = mc.MulticamSession.load(session_id) + session = mc.open_session(session_id) if action == "render" and any(k in params for k in MULTICAM_MAP_KEYS if k != "look"): # A mapping change can drop the cut, and a render needs one: map, then plan, then render. raise ValueError("render takes only look and stems. Change the mapping with 'map', then 'plan' again.") @@ -1036,7 +1115,10 @@ def progress(stage): stems = params.get("stems", True) if not isinstance(stems, bool): raise ValueError("stems is true or false") - mc.render_session(session, stems=stems, progress_callback=progress("rendering")) + validate = params.get("validate", "sample") + if validate not in ("sample", "full"): + raise ValueError("validate must be 'sample' or 'full'") + mc.render_session(session, stems=stems, validate=validate, progress_callback=progress("rendering")) elif action == "export": data["export_path"] = mc.export_xml(session, params.get("format", "premiere"), review=bool(params.get("review"))) elif action == "import_timeline": @@ -1098,7 +1180,9 @@ def handle_run_integration_tool(task_id: str, params: dict): TASK_HANDLERS = { "ping": handle_ping, + "resolve_transcribe_engine": handle_resolve_transcribe_engine, "transcribe": handle_transcribe, + "compare_engines": handle_compare_engines, "parse_transcript": handle_parse_transcript, "create_clip": handle_create_clip, "batch_clips": handle_batch_clips, diff --git a/backend/requirements-runtime.txt b/backend/requirements-runtime.txt index 8cd32528..de101619 100644 --- a/backend/requirements-runtime.txt +++ b/backend/requirements-runtime.txt @@ -7,6 +7,9 @@ numpy>=2.5.2 # Audio-event detection (YAMNet laughter/reaction channel) runs on ONNX Runtime — # no torch/TF, so it stays on the hermetic native path. onnxruntime>=1.28.0 +# --engine omnilingual: Meta's Omnilingual ASR CTC model via sherpa-onnx, for +# languages whisper.cpp handles poorly. No torch, stays on the hermetic path. +sherpa-onnx>=1.13.8 Pillow>=12.3.0 questionary>=2.1.1 python-dotenv>=1.2.2 diff --git a/backend/requirements.txt b/backend/requirements.txt index 546df293..96ededed 100644 --- a/backend/requirements.txt +++ b/backend/requirements.txt @@ -15,6 +15,10 @@ numpy>=2.5.2 # Audio-event detection (laughter/reaction channel via YAMNet ONNX — no torch/TF) onnxruntime>=1.28.0 +# --engine omnilingual: Meta's Omnilingual ASR CTC model via sherpa-onnx, for +# languages whisper.cpp handles poorly (e.g. Georgian). No torch/TF. +sherpa-onnx>=1.13.8 + # Thumbnails Pillow>=12.3.0 diff --git a/backend/services/audio_extract.py b/backend/services/audio_extract.py index f9a6883a..d96a2cec 100644 --- a/backend/services/audio_extract.py +++ b/backend/services/audio_extract.py @@ -17,18 +17,29 @@ def extract_wav_16k_mono( media_path: str, wav_path: Optional[str] = None, timeout: int = 1800, + start_seconds: Optional[float] = None, + duration_seconds: Optional[float] = None, ) -> str: """Extract audio as 16 kHz mono 16-bit PCM WAV. Returns the wav path. When wav_path is None a temp file is created; the caller owns cleanup. + start_seconds/duration_seconds trim the output to a window of the source, + used for sample-mode transcription (test a language on a short clip + instead of the full episode). """ owns_wav = wav_path is None if owns_wav: fd, wav_path = tempfile.mkstemp(prefix="podcli_audio_", suffix=".wav") os.close(fd) - cmd = [ - "ffmpeg", "-y", "-loglevel", "error", - "-i", media_path, + cmd = ["ffmpeg", "-y", "-loglevel", "error"] + if start_seconds: + # Before -i: fast (keyframe-seek) trim. Precision to the frame + # doesn't matter for a language-check sample. + cmd += ["-ss", str(start_seconds)] + cmd += ["-i", media_path] + if duration_seconds: + cmd += ["-t", str(duration_seconds)] + cmd += [ "-vn", "-acodec", "pcm_s16le", "-ar", "16000", diff --git a/backend/services/caption_renderer.py b/backend/services/caption_renderer.py index 8f44d3a9..f8bb5d44 100644 --- a/backend/services/caption_renderer.py +++ b/backend/services/caption_renderer.py @@ -12,6 +12,27 @@ from config.caption_styles import get_style from utils.timing_utils import seconds_to_ass +from utils.text import safe_upper + + +def _sanitize_ass_text(text: str) -> str: + """Neutralize characters that mean something to the ASS/libass parser + when they show up in transcribed word text instead of our own markup. + + A literal backslash isn't escapable in the plain-text part of a + Dialogue line: libass still reads \\N, \\n and \\h out of it (that's how + an uppercase transform can turn a stray "\\n" into a forced \\N + linebreak), and a literal "{" opens a new override block early, letting + anything after it (including a following "}") be read as ASS tags + instead of rendered as text. There's no escape sequence for any of + these in plain text, so swap in a full-width look-alike glyph that + renders the same character shape without being special to the parser. + """ + return ( + text.replace("\\", "\") + .replace("{", "{") + .replace("}", "}") + ) def generate_ass_header(style: dict, play_res_x: int = 1080, play_res_y: int = 1920) -> str: @@ -207,7 +228,7 @@ def _render_hormozi(words: list[dict], style: dict, offset: float) -> str: parts = [] for w in chunk: duration_cs = int((w["end"] - w["start"]) * 100) - text = w["word"].upper() if uppercase else w["word"] + text = _sanitize_ass_text(safe_upper(w["word"]) if uppercase else w["word"]) parts.append(f"{{\\kf{duration_cs}}}{text}") # \c = active (filled) color, \2c = inactive (unfilled) color @@ -246,7 +267,7 @@ def _render_karaoke(words: list[dict], style: dict, offset: float) -> str: parts = [] for w in sentence: duration_cs = int((w["end"] - w["start"]) * 100) - text = w["word"] + text = _sanitize_ass_text(w["word"]) parts.append(f"{{\\kf{duration_cs}}}{text}") line_text = " ".join(parts) @@ -279,7 +300,7 @@ def _render_subtle(words: list[dict], style: dict, offset: float) -> str: line_start = max(0, line_words[0]["start"] - offset) line_end = _hold_through_gap(lines, idx, max(0, line_words[-1]["end"] - offset), offset) - line_text = " ".join(w["word"] for w in line_words) + line_text = " ".join(_sanitize_ass_text(w["word"]) for w in line_words) start_ts = seconds_to_ass(line_start) end_ts = seconds_to_ass(line_end) @@ -544,9 +565,9 @@ def _render_branded(words: list[dict], style: dict, offset: float) -> str: # Normalize casing normalized = [] for j, w in enumerate(chunk): - text = _normalize_case(w["word"]) + text = _sanitize_ass_text(_normalize_case(w["word"])) if j == 0: - text = text[0].upper() + text[1:] if len(text) > 1 else text.upper() + text = safe_upper(text[0]) + text[1:] if len(text) > 1 else safe_upper(text) normalized.append(text) # Measure word widths for pill positioning. diff --git a/backend/services/claude_suggest.py b/backend/services/claude_suggest.py index 6585658c..280907b4 100644 --- a/backend/services/claude_suggest.py +++ b/backend/services/claude_suggest.py @@ -23,6 +23,7 @@ sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) from presets import DEFAULT_PRESET, MIN_CLIP_DURATION, MAX_CLIP_DURATION, TARGET_CLIP_DURATION_MIN, TARGET_CLIP_DURATION_MAX from services.formats import get_format +from services.opening_hook import HOOK_MODES, validate_hook from utils.text import clean_title from services import ai_provider, podcli_cloud from services.ai_cli import ( @@ -126,7 +127,7 @@ def keeps(self, seconds: float) -> bool: # Bump when the selection rules in _build_prompt change. Folded into the # suggestion cache key for the same reason the knowledge base is: replaying old # picks under new rules teaches people the rules do nothing. -PROMPT_VERSION = 5 +PROMPT_VERSION = 6 def kb_signature() -> str: @@ -245,6 +246,26 @@ def _total_score(clip: dict) -> float: MAX_REACTION_ANCHORS = 40 +def _suggested_hook(clip: dict, start: float, end: float, segments: list[dict]) -> Optional[dict]: + """The opening hook the model proposed for this clip, or None. + + A hook that is malformed or falls outside the clip it opens is dropped + rather than failing the clip: the moment is still worth rendering without it. + """ + raw = clip.get("hook") + if not isinstance(raw, dict) or raw.get("mode") not in HOOK_MODES: + return None + try: + hook = { + "start": round(float(raw.get("start")), 1), + "end": round(float(raw.get("end")), 1), + "mode": raw["mode"], + } + return validate_hook(hook, start, end, segments) + except (TypeError, ValueError): + return None + + def _format_reaction_anchors(reaction_times: list[float] | None) -> str: """Prompt block listing detected laughter/reaction peaks as anchor moments.""" if not reaction_times: @@ -401,6 +422,12 @@ def _build_prompt( - "start_second" / "end_second" = outer bounds (first segment start, last segment end) - Example: speaker makes great point (10s), rambles (8s), delivers punchline (12s) → 2 segments, 22s total +HOOK (optional): +- If the sharpest line sits mid-clip, add "hook": {{"start": X, "end": Y, "mode": "repeat"}} so it plays first. +- It must be 1-15 seconds of words actually spoken inside the clip's segments. Never invent text. +- "repeat" plays it again where it belongs. "move" lifts it out of the body. +- Leave "hook" out when the clip already opens on its strongest line. + Rules: - Final clip duration (sum of segments) MUST be {b.dur_min}-{b.dur_max} seconds (target {b.target_min}-{b.target_max}s) - Each segment must start and end on COMPLETE SENTENCES — never mid-thought @@ -576,6 +603,8 @@ def find_moments_from_text( - In "needs", name what the viewer must already know, or "nothing". If it is anything else, move start_second back until the clip covers it. "context_line" is a note for the editor, not a fix: nothing burns it into the video. +- Optional "hook": {{"start": X, "end": Y, "mode": "repeat" or "move"}}, 1-15 seconds of + words spoken inside the clip, played first. Leave it out when the clip opens strong. Return this JSON: {{ @@ -633,11 +662,14 @@ def announce(label: str) -> None: if not bounds.keeps(kept_duration): continue + clip_start = keep_segments[0]["start"] if keep_segments else start_sec + clip_end = keep_segments[-1]["end"] if keep_segments else end_sec found.append({ "title": clean_title(c.get("title", "Untitled")), - "start_second": keep_segments[0]["start"] if keep_segments else start_sec, - "end_second": keep_segments[-1]["end"] if keep_segments else end_sec, + "start_second": clip_start, + "end_second": clip_end, "segments": keep_segments, + "hook": _suggested_hook(c, clip_start, clip_end, keep_segments), "duration": round(kept_duration), "score": total, "content_type": c.get("content_type", "unknown"), @@ -839,11 +871,14 @@ def records(payload: object) -> list[dict]: if not bounds.keeps(kept_duration): continue + clip_start = keep_segments[0]["start"] if keep_segments else start_sec + clip_end = keep_segments[-1]["end"] if keep_segments else end_sec normalized.append({ "title": clean_title(c.get("title", "Untitled")), - "start_second": keep_segments[0]["start"] if keep_segments else start_sec, - "end_second": keep_segments[-1]["end"] if keep_segments else end_sec, + "start_second": clip_start, + "end_second": clip_end, "segments": keep_segments, + "hook": _suggested_hook(c, clip_start, clip_end, keep_segments), "duration": round(kept_duration), "score": total, "content_type": c.get("content_type", "unknown"), diff --git a/backend/services/clip_generator.py b/backend/services/clip_generator.py index 314f390d..a1cfc831 100644 --- a/backend/services/clip_generator.py +++ b/backend/services/clip_generator.py @@ -32,6 +32,15 @@ normalize_audio, concat_outro, ) +from services.video_cut import probe_has_audio_stream, verify_full_decode +from services.glyph_coverage import check_caption_font_coverage +from services.subtitle_export import write_sidecars +from services.opening_hook import ( + validate_hook, + snap_hook_to_words, + order_with_hook, + keyframes_to_playback, +) from config.caption_styles import get_style from services.formats import get_format @@ -131,6 +140,35 @@ def _get_media_duration(path: str) -> float: return 0.0 +def _bound_range_to_source( + start_second: float, + end_second: float, + keep_segments: Optional[list[dict]], + source_duration: float, +) -> tuple[float, Optional[list[dict]]]: + """Clamp the requested range (and any keep_segments) to the source's + actual duration, so a too-long end_second can't silently truncate the + render instead of failing. source_duration <= 0 means ffprobe couldn't + read it; skip bounding rather than clamp against an unknown length. + """ + if source_duration <= 0: + return end_second, keep_segments + if start_second >= source_duration: + raise ValueError( + f"start_second ({start_second}s) is at or past the end of the " + f"source video ({source_duration:.2f}s)." + ) + end_second = min(end_second, source_duration) + if keep_segments: + bounded_segments = [] + for seg in keep_segments: + if seg["start"] >= source_duration: + continue + bounded_segments.append({**seg, "end": min(seg["end"], source_duration)}) + keep_segments = bounded_segments or None + return end_second, keep_segments + + def _detect_scene_cuts(path: str, threshold: float = 0.22, max_cuts: int = 32) -> list[float]: """ Detect hard visual changes likely to feel like jump cuts. @@ -859,11 +897,13 @@ def generate_clip( bookend_fade: float = 0.0, clean_fillers: bool = True, keep_segments: list[dict] = None, + hook: Optional[dict] = None, trim_opening: Optional[bool] = None, preserve_timing: bool = False, allow_ass_fallback: bool = False, use_ass_captions: bool = False, keep_caption_overlay: bool = False, + write_clean_variant: bool = False, topic: Optional[dict] = None, progress: Optional[dict] = None, cards: Optional[list] = None, @@ -871,6 +911,7 @@ def generate_clip( font_family: Optional[str] = None, theme: Optional[dict] = None, captions: bool = True, + write_subtitles: Optional[bool] = None, progress_callback: Optional[Callable[[int, str], None]] = None, ) -> dict: """ @@ -887,6 +928,13 @@ def generate_clip( title: Clip title (used in filename) output_dir: Where to save the final clip (defaults to temp) logo_path: Path to logo image (PNG). Used with "branded" style. + hook: Optional {"start", "end", "mode"} on the source clock: a passage + from inside the body played first. "repeat" plays it again in + place, "move" removes it from the body. + captions: Whether to burn captions into the video. + write_subtitles: Whether to write .srt/.vtt sidecars. Defaults to + following `captions` (no burned captions, no sidecars either); + pass True to request them for a clean (caption-free) export. progress_callback: Optional (percent, message) callback Returns: @@ -912,6 +960,25 @@ def generate_clip( if end_second <= start_second: raise ValueError("end_second must be greater than start_second") + # A video-only source (e.g. a silent screen recording) never gets an + # audio stream from any of the steps below, so requiring one would + # reject every render of it as "broken" instead of a legitimately + # silent clip. + source_has_audio = probe_has_audio_stream(video_path) + + # ffmpeg's -ss/-t cut silently stops at EOF when end_second runs past the + # source, so without this the render "succeeds" with a shorter clip than + # requested while duration/end_second in the result still say the planned + # (too-long) window. Bound every range to what the source actually has. + source_duration = _get_media_duration(video_path) + end_second, keep_segments = _bound_range_to_source( + start_second, end_second, keep_segments, source_duration + ) + + # Checked against the requested body, before any tightening moves it. + hook = validate_hook(hook, start_second, end_second, keep_segments) + requested_start = start_second + spec = get_format(format) # An episode that was never filmed takes the other road entirely. Branching @@ -920,6 +987,11 @@ def generate_clip( # back is the same dict describing the same window. from services.audiogram import is_audio_only if is_audio_only(video_path): + if hook: + raise ValueError( + "An opening hook needs a video source. Audiogram clips play one " + "continuous range. Drop the hook to render this audio-only clip." + ) from services.audiogram import render_audiogram return render_audiogram( audio_path=video_path, @@ -1025,11 +1097,22 @@ def generate_clip( keep_segments = None duration = end_second - start_second + # The body is final here. The hook is cut as given (widened only to whole + # words) and put ahead of it, so every later step reads play_ranges, in + # playback order, never re-sorted. None means one continuous range. + play_ranges = keep_segments if keep_segments and len(keep_segments) > 1 else None + if hook: + hook = snap_hook_to_words(hook, transcript_words) + body = keep_segments or [{"start": start_second, "end": end_second}] + play_ranges = order_with_hook(body, hook) + duration = sum(r["end"] - r["start"] for r in play_ranges) + length_warning = None # A clip that asks for captions and renders without them used to be # indistinguishable from a clip that never wanted them. Two shipped clips # went out silent that way. caption_warning = None + glyph_warning = None if duration > spec.dur_max: length_warning = f"{duration:.0f}s, over the {spec.dur_max}s {spec.name} target" print(f" {length_warning}", file=sys.stderr, flush=True) @@ -1048,40 +1131,54 @@ def generate_clip( # Step 1: Cut the segment(s) from the source video if progress_callback: - n_segs = len(keep_segments) if keep_segments else 1 + n_segs = len(play_ranges) if play_ranges else 1 msg = f"Cutting {n_segs} segment{'s' if n_segs > 1 else ''} (1/{total_steps})" progress_callback(10, msg) segment_path = os.path.join(work_dir, "segment.mp4") - with timed("render", "cut", segments=len(keep_segments) if keep_segments else 1): - if keep_segments and len(keep_segments) > 1: - cut_multi_segment(video_path, segment_path, keep_segments) + part_durations: Optional[list[float]] = None + with timed("render", "cut", segments=len(play_ranges) if play_ranges else 1): + if play_ranges: + _, part_durations = cut_multi_segment(video_path, segment_path, play_ranges) else: cut_segment(video_path, segment_path, start_second, end_second) - # Remap transcript words for multi-segment clips. + # Remap transcript words for multi-segment clips, range by range in + # playback order, so a repeated hook's words appear twice. # Needed before crop (speaker detection) and captions. - if keep_segments and len(keep_segments) > 1 and transcript_words: - remapped_words = [] + remapped_words: list[dict] = [] + if play_ranges and transcript_words: cumulative_t = 0.0 - for seg in keep_segments: + for i, seg in enumerate(play_ranges): seg_words = [ w for w in transcript_words if w["end"] > seg["start"] and w["start"] < seg["end"] ] - seg_duration = seg["end"] - seg["start"] + # Each part is encoded separately before the concat, and an + # encoder snaps a cut to whole frames, so its real duration is + # typically a few ms off the requested end - start. Advancing + # cumulative_t by the planned length instead of the probed + # one drifts captions further out of sync with every segment + # concatenated in. Fall back to the planned length only if + # the part couldn't be probed. + requested_duration = seg["end"] - seg["start"] + actual_duration = ( + part_durations[i] + if part_durations and part_durations[i] > 0 + else requested_duration + ) for w in seg_words: # Clamp to segment bounds to avoid negative/overflow timestamps # for words that straddle a segment boundary - remapped_start = max(0, cumulative_t + (w["start"] - seg["start"])) - remapped_end = min(cumulative_t + seg_duration, cumulative_t + (w["end"] - seg["start"])) + remapped_start = max(cumulative_t, cumulative_t + (w["start"] - seg["start"])) + remapped_end = min(cumulative_t + actual_duration, cumulative_t + (w["end"] - seg["start"])) if remapped_end > remapped_start: remapped_words.append({ **w, "start": round(remapped_start, 3), "end": round(remapped_end, 3), }) - cumulative_t += seg_duration + cumulative_t += actual_duration crop_words = remapped_words crop_clip_start = 0 caption_time_offset = 0 @@ -1098,6 +1195,13 @@ def generate_clip( if progress_callback: progress_callback(30, f"Resizing for {spec.name} format (2/{total_steps})") + playback_keyframes = keyframes_to_playback( + crop_keyframes, + requested_start, + play_ranges or [{"start": start_second, "end": end_second}], + part_durations, + ) if crop_keyframes else crop_keyframes + cropped_path = os.path.join(work_dir, "cropped.mp4") with timed("render", "crop", strategy=crop_strategy if spec.reframe else "fit"): if spec.reframe: @@ -1107,7 +1211,7 @@ def generate_clip( transcript_words=crop_words, clip_start=crop_clip_start, face_map=face_map, - crop_keyframes=crop_keyframes, + crop_keyframes=playback_keyframes, target_dims=spec.dims, ) else: @@ -1140,6 +1244,9 @@ def generate_clip( logo_path or topic or progress or cards or theme or (name_card and name_card.get("title")) ) + # Default so the sidecar/clean-variant step below has something to + # check even when neither branch below runs (no transcript, no overlay). + clip_words = [] if not captions and wants_overlay: print(" drawing the overlay without captions", flush=True) @@ -1154,7 +1261,7 @@ def generate_clip( if progress_callback: progress_callback(50, f"Adding {caption_style} captions (3/{total_steps})") - if keep_segments and len(keep_segments) > 1: + if play_ranges: clip_words = remapped_words else: clip_words = [ @@ -1180,6 +1287,15 @@ def generate_clip( ) print(f" {caption_warning}", file=sys.stderr, flush=True) + if clip_words and captions: + caption_text = " ".join(w.get("word", "") for w in clip_words) + glyph_warning = check_caption_font_coverage( + caption_text, style_config["font_name"], bool(style_config["bold"]), + use_ass=use_ass_captions, + ) + if glyph_warning: + print(f" {glyph_warning}", file=sys.stderr, flush=True) + if (clip_words and captions) or wants_overlay: if progress_callback: progress_callback(65, f"Rendering captions into video (3/{total_steps})") @@ -1310,14 +1426,17 @@ def generate_clip( final_video_path = with_intro_path intro_offset = max(0.0, _get_media_duration(with_intro_path) - clip_duration) + outro_offset = 0.0 if outro_path and os.path.exists(outro_path): if progress_callback: progress_callback(85, f"Adding outro ({total_steps}/{total_steps})") + pre_outro_duration = _get_media_duration(final_video_path) with_outro_path = os.path.join(work_dir, "with_outro.mp4") concat_outro(final_video_path, outro_path, with_outro_path, crossfade_duration=bookend_fade) final_video_path = with_outro_path + outro_offset = max(0.0, _get_media_duration(with_outro_path) - pre_outro_duration) # Step 6: Move to output if progress_callback: @@ -1341,17 +1460,107 @@ def generate_clip( reframe=spec.reframe, crop_strategy=crop_strategy, crop_keyframes=crop_keyframes, - keep_segments=keep_segments, + keep_segments=play_ranges, ) if max_autofix_passes > 0: if progress_callback: progress_callback(97, "Quality gate: checking transitions...") + # The cut from the hook into the body is deliberate. Smoothing + # it would blur the first beat the hook exists to land. + hook_cut = [intro_offset + (part_durations[0] if part_durations else 0.0)] if hook else [] _auto_fix_transition_jumps( final_path, max_passes=max_autofix_passes, - designed=_designed_cuts(cards, offset=intro_offset), + designed=_designed_cuts(cards, offset=intro_offset) + hook_cut, ) + # A 0-exit ffmpeg run can still have written a short, silent, or + # corrupt file (source ran out under the requested range, a crop/ + # caption/normalize pass dropped the audio track, a concat mismatch). + # Catch that here instead of returning a "successful" result that + # describes a different clip than what landed on disk. + expected_duration = duration + intro_offset + outro_offset + actual_duration = _get_media_duration(final_path) + duration_tolerance = max(0.75, 0.05 * expected_duration) + if actual_duration <= 0 or abs(actual_duration - expected_duration) > duration_tolerance: + os.remove(final_path) + raise RuntimeError( + f"Render produced a {actual_duration:.2f}s file but expected " + f"~{expected_duration:.2f}s (body {duration:.2f}s" + f"{f' + intro {intro_offset:.2f}s' if intro_offset else ''}" + f"{f' + outro {outro_offset:.2f}s' if outro_offset else ''}). " + f"The source video likely ended before the requested range, or " + f"the render failed partway through." + ) + if source_has_audio and not probe_has_audio_stream(final_path): + os.remove(final_path) + raise RuntimeError( + "Render completed but the output has no audio stream. " + "Re-check the source video and caption/crop pipeline." + ) + decode_error = verify_full_decode(final_path) + if decode_error: + os.remove(final_path) + raise RuntimeError(f"Render produced an undecodable output: {decode_error}") + + # Sidecar subtitles: the same words already burned into the video, + # on the same playback clock as the delivered file (clip-local time, + # shifted past whatever intro got prepended). clip_words is populated + # for overlay-only renders too (logo/cards with captions off), so + # gate on whether captions were actually requested rather than just + # "were there words": otherwise a captions-off render still gets an + # .srt/.vtt describing dialogue nothing on screen shows. + want_subtitles = write_subtitles if write_subtitles is not None else captions + output_base, _ = os.path.splitext(final_path) + sidecar_paths: dict = {} + # A re-render claims the same output path as last time + # (_reserve_output_path), so a sidecar this pass doesn't produce can + # be a leftover from a previous render of this clip, now describing + # dialogue or a style that no longer matches. Clear it before + # (maybe) writing a fresh one. + for stale_ext in (".srt", ".vtt"): + stale_sidecar = f"{output_base}{stale_ext}" + if os.path.exists(stale_sidecar): + os.remove(stale_sidecar) + if want_subtitles: + retimed_words = [ + { + "word": w.get("word", ""), + "start": round(max(0.0, w["start"] - caption_time_offset + intro_offset), 3), + "end": round(max(0.0, w["end"] - caption_time_offset + intro_offset), 3), + } + for w in clip_words + ] + sidecar_paths = write_sidecars(retimed_words, output_base) + + # Optional clean variant: the same audio, loudness, and intro/outro, + # minus burned captions, built from the cropped (pre-caption) + # source with the identical normalize_audio/concat_outro calls the + # main render used, so the two files only differ in the overlay. + clean_output_path = None + stale_clean_path = f"{output_base}_clean.mp4" + if os.path.exists(stale_clean_path): + os.remove(stale_clean_path) + if write_clean_variant and os.path.exists(cropped_path): + clean_normalized_path = os.path.join(work_dir, "clean_normalized.mp4") + normalize_audio(cropped_path, clean_normalized_path) + clean_video_path = clean_normalized_path + + if intro_path and os.path.exists(intro_path): + clean_with_intro_path = os.path.join(work_dir, "clean_with_intro.mp4") + concat_outro(intro_scaled, clean_video_path, clean_with_intro_path, + crossfade_duration=bookend_fade) + clean_video_path = clean_with_intro_path + + if outro_path and os.path.exists(outro_path): + clean_with_outro_path = os.path.join(work_dir, "clean_with_outro.mp4") + concat_outro(clean_video_path, outro_path, clean_with_outro_path, + crossfade_duration=bookend_fade) + clean_video_path = clean_with_outro_path + + clean_output_path = f"{output_base}_clean.mp4" + shutil.copy2(clean_video_path, clean_output_path) + # Get file size file_size = os.path.getsize(final_path) file_size_mb = round(file_size / (1024 * 1024), 2) @@ -1359,9 +1568,21 @@ def generate_clip( if progress_callback: progress_callback(100, "Clip complete!") + # "duration" has always meant the content's own length: what + # clip_history and the learnings it feeds record, not however long + # the delivered file plays including intro/outro. Derive it from the + # probed actual_duration (not the planned `duration` window) so a + # source that ran out early still gets caught by the check above and + # reported as shorter, instead of silently reporting the plan. + content_duration = max(0.0, actual_duration - intro_offset - outro_offset) + out = { "output_path": final_path, - "duration": round(duration, 2), + "duration": round(content_duration, 2), + # The probed duration of the file actually on disk, including any + # intro/outro, for callers that need the delivered file's full + # playback length rather than just its content. + "output_duration": round(actual_duration, 2), "file_size_mb": file_size_mb, "title": title, "start_second": start_second, @@ -1369,8 +1590,13 @@ def generate_clip( "caption_style": caption_style, "crop_strategy": crop_strategy, "format": spec.name, + **sidecar_paths, } - warnings = [w for w in (caption_warning, length_warning) if w] + if clean_output_path: + out["clean_output_path"] = clean_output_path + if hook: + out["hook"] = hook + warnings = [w for w in (caption_warning, length_warning, glyph_warning) if w] if warnings: out["warning"] = "; ".join(warnings) if keep_caption_overlay and caption_overlay_path and os.path.exists(caption_overlay_path): diff --git a/backend/services/clips_history.py b/backend/services/clips_history.py index 4a5f8d91..7878a1d6 100644 --- a/backend/services/clips_history.py +++ b/backend/services/clips_history.py @@ -8,8 +8,8 @@ Entry shape (all fields beyond the core render record are optional): id, source_video, start_second, end_second, caption_style, crop_strategy, logo_path?, title, output_path, file_size_mb, duration, created_at, - content_type?, transcript_slice?, youtube_video_id?, metrics?, - generated_titles?, description?, tags?, hashtags? + content_type?, transcript_slice?, payoff?, context_line?, preview_text?, + youtube_video_id?, metrics?, generated_titles?, description?, tags?, hashtags? metrics? = {views?, retention?, ctr?, impressions?, fetched_at?} (Phase 2) diff --git a/backend/services/corrections.py b/backend/services/corrections.py index 959c2fc1..f1e5a310 100644 --- a/backend/services/corrections.py +++ b/backend/services/corrections.py @@ -73,6 +73,91 @@ def _replace_match(match: re.Match, corrections: dict[str, str]) -> str: return replacement if replacement is not None else matched +def _strip_for_match(text: str) -> str: + return text.strip(".,!?;:\"'()-") + + +_TRAILING_PUNCT_RE = re.compile(r"[.,!?;:\"'()\-]+$") + + +def _trailing_punct(text: str) -> str: + """Trailing punctuation only (not leading), e.g. 'AI.' -> '.'.""" + match = _TRAILING_PUNCT_RE.search(text) + return match.group(0) if match else "" + + +def _merge_multiword_corrections(words: list[dict], corrections: dict[str, str]) -> list[dict]: + """ + Merge consecutive words matching a multi-word correction key (e.g. + "open AI" -> "OpenAI") into the corrected word(s). + + Segment text gets multi-word corrections for free from the regex pass + below, but captions are burned from individual words. Left split, + "open" and "AI" render as two separate caption words instead of the + fix. The merged word(s) span from the start of the first matched word + to the end of the last one, so caption timing stays continuous. + """ + multiword = {k: v for k, v in corrections.items() if len(k.split()) > 1} + if not multiword or not words: + return words + + # Longest key first so a 3-word phrase wins over a 2-word prefix of it. + key_tokens = sorted( + ((key.split(), value) for key, value in multiword.items()), + key=lambda kv: len(kv[0]), + reverse=True, + ) + + merged: list[dict] = [] + i = 0 + n = len(words) + while i < n: + match = None + for tokens, replacement in key_tokens: + span = len(tokens) + if i + span > n: + continue + candidate = words[i : i + span] + candidate_norm = [_strip_for_match(w.get("word", "")).lower() for w in candidate] + if candidate_norm == [t.lower() for t in tokens]: + match = (candidate, replacement) + break + if match is None: + merged.append(words[i]) + i += 1 + continue + + candidate, replacement = match + first, last = candidate[0], candidate[-1] + repl_words = replacement.split() or [replacement] + span_start = first["start"] + span_end = last["end"] + span_dur = max(0.0, span_end - span_start) + per = span_dur / len(repl_words) + + # The matched words' punctuation and other fields (speaker, + # confidence, ...) would otherwise vanish behind the correction: + # "open AI." becomes "OpenAI" with no period and a dropped speaker. + # Carry the last word's trailing punctuation onto the final merged + # word, and the first word's other fields onto every merged word. + trailing_punct = _trailing_punct(last.get("word", "")) + confidences = [ + w["confidence"] for w in candidate if isinstance(w.get("confidence"), (int, float)) + ] + extra_fields = {k: v for k, v in first.items() if k not in ("word", "start", "end")} + if confidences: + extra_fields["confidence"] = min(confidences) + + for idx, rw in enumerate(repl_words): + w_start = span_start + per * idx + w_end = span_end if idx == len(repl_words) - 1 else span_start + per * (idx + 1) + word_text = rw + trailing_punct if idx == len(repl_words) - 1 else rw + merged.append({**extra_fields, "word": word_text, "start": w_start, "end": w_end}) + i += len(candidate) + + return merged + + def apply_corrections( words: list[dict], segments: list[dict], @@ -81,7 +166,11 @@ def apply_corrections( Apply corrections to transcript words and segments in-place. Modifies the 'word' field in each word dict and the 'text' field - in each segment dict. Returns the same lists (mutated). + in each segment dict. Multi-word corrections can change the number of + words (several words merge into the correction's word(s)), so the + `words` list itself is replaced in-place via slice assignment, so + callers that hold a reference to the original list still see the + update. Returns the same lists (mutated). """ corrections = _load_corrections() if not corrections: @@ -93,6 +182,12 @@ def apply_corrections( replacer = lambda m: _replace_match(m, corrections) + # Merge multi-word corrections first so captions (built from words) read + # the fix the same way the segment text already does. + merged_words = _merge_multiword_corrections(words, corrections) + if merged_words is not words: + words[:] = merged_words + # Fix individual words (strip punctuation for matching, preserve it in output) for w in words: word_text = w.get("word", "") diff --git a/backend/services/engine_comparison.py b/backend/services/engine_comparison.py new file mode 100644 index 00000000..8429d40f --- /dev/null +++ b/backend/services/engine_comparison.py @@ -0,0 +1,299 @@ +"""Compare two transcription engines over the same window of a file. + +Transcribes a sample range twice, once per engine, and reports where the two +transcripts disagree, windowed every 20s by default. This measures +disagreement between the two outputs, not accuracy against a ground truth. +Neither engine is assumed correct. +""" + +from __future__ import annotations + +import html +import json +import os +import shutil +import unicodedata +from typing import Any, Optional + +# Unicode categories kept when normalizing a word for comparison: letters (L*), +# marks (M*, e.g. combining diacritics), numbers (N*). Punctuation and symbols +# are dropped so "hello," and "hello" agree. +_KEEP_CATEGORY_PREFIXES = ("L", "M", "N") + + +def normalize_word(word: str) -> str: + """NFKC-normalize, casefold, and strip everything but letters/marks/numbers.""" + normalized = unicodedata.normalize("NFKC", word or "").casefold() + return "".join(ch for ch in normalized if unicodedata.category(ch)[0] in _KEEP_CATEGORY_PREFIXES) + + +def normalize_words(words: list[str]) -> list[str]: + return [w for w in (normalize_word(w) for w in words) if w] + + +def word_levenshtein(a: list[str], b: list[str]) -> int: + """Levenshtein distance over word sequences (insert/delete/substitute, cost 1 each).""" + if a == b: + return 0 + if not a: + return len(b) + if not b: + return len(a) + prev = list(range(len(b) + 1)) + for i, wa in enumerate(a, 1): + curr = [i] + [0] * len(b) + for j, wb in enumerate(b, 1): + cost = 0 if wa == wb else 1 + curr[j] = min( + prev[j] + 1, # delete from a + curr[j - 1] + 1, # insert into a + prev[j - 1] + cost, # substitute + ) + prev = curr + return prev[-1] + + +def disagreement_ratio(text_a: str, text_b: str) -> float: + """Levenshtein over normalized words, divided by the longer word count. + 0.0 = identical (after normalization), 1.0 = completely different. + Two empty texts disagree 0.0 (nothing to disagree about), so callers should + flag that case separately rather than reading it as agreement.""" + words_a = normalize_words(text_a.split()) + words_b = normalize_words(text_b.split()) + denom = max(len(words_a), len(words_b)) + if denom == 0: + return 0.0 + return word_levenshtein(words_a, words_b) / denom + + +def _words_in_window(words: list[dict], window_start: float, window_end: float) -> str: + """Join words whose midpoint falls in [window_start, window_end). Matches + the midpoint-ownership rule used elsewhere so a word isn't double-counted + in two adjacent windows.""" + picked = [] + for w in words: + try: + start, end = float(w.get("start", 0.0)), float(w.get("end", 0.0)) + except (TypeError, ValueError): + continue + midpoint = (start + end) / 2.0 + if window_start <= midpoint < window_end: + text = str(w.get("word", "")).strip() + if text: + picked.append(text) + return " ".join(picked) + + +def build_windows( + words_a: list[dict], + words_b: list[dict], + total_duration: float, + window_seconds: float = 20.0, +) -> list[dict]: + windows = [] + n_windows = max(1, int(total_duration // window_seconds) + (1 if total_duration % window_seconds else 0)) + for i in range(n_windows): + start = i * window_seconds + end = min(start + window_seconds, total_duration) + text_a = _words_in_window(words_a, start, end) + text_b = _words_in_window(words_b, start, end) + both_empty = not text_a and not text_b + windows.append({ + "start": round(start, 3), + "end": round(end, 3), + "text_a": text_a, + "text_b": text_b, + "both_empty": both_empty, + "disagreement": 0.0 if both_empty else round(disagreement_ratio(text_a, text_b), 4), + }) + return windows + + +def compare_engines( + file_path: str, + engine_a: str, + engine_b: str, + *, + start_seconds: float = 0.0, + duration_seconds: Optional[float] = None, + window_seconds: float = 20.0, + model_size: str = "base", + language: Optional[str] = None, + output_dir: Optional[str] = None, + transcribe_fn=None, +) -> dict: + """Transcribe the same [start_seconds, start_seconds + duration_seconds) + window with engine_a and engine_b, and report per-window disagreement. + + transcribe_fn defaults to services.transcription.transcribe_file; tests + inject a fake to avoid running real engines. + """ + if transcribe_fn is None: + from services.transcription import transcribe_file as transcribe_fn + + result_a = transcribe_fn( + file_path, model_size=model_size, engine=engine_a, language=language, + enable_diarization=False, start_seconds=start_seconds, duration_seconds=duration_seconds, + ) + result_b = transcribe_fn( + file_path, model_size=model_size, engine=engine_b, language=language, + enable_diarization=False, start_seconds=start_seconds, duration_seconds=duration_seconds, + ) + + duration = max( + float(result_a.get("duration", 0.0) or 0.0), + float(result_b.get("duration", 0.0) or 0.0), + duration_seconds or 0.0, + ) + windows = build_windows( + result_a.get("words") or [], result_b.get("words") or [], duration, window_seconds + ) + scored = [w["disagreement"] for w in windows if not w["both_empty"]] + overall_disagreement = round(sum(scored) / len(scored), 4) if scored else 0.0 + flagged_empty = sum(1 for w in windows if w["both_empty"]) + + report = { + "file_path": file_path, + "engine_a": engine_a, + "engine_b": engine_b, + "start_seconds": start_seconds, + "duration_seconds": duration, + "window_seconds": window_seconds, + "windows": windows, + "overall_disagreement": overall_disagreement, + "both_empty_window_count": flagged_empty, + "note": "disagreement between the two engines' output, not accuracy against a transcript", + } + + if output_dir: + os.makedirs(output_dir, exist_ok=True) + json_path = os.path.join(output_dir, "comparison.json") + with open(json_path, "w", encoding="utf-8") as f: + json.dump(report, f, indent=2, ensure_ascii=False) + report["json_path"] = json_path + + # A fresh extraction of the same window, purely for the report's + #