Repository navigation
Harden clips, transcription and multicam; add opening hooks and Omnilingual - #265
Conversation
…ls with pinned revisions - Pass -l auto instead of omitting -l so whisper-cli runs language detection on auto/unset instead of defaulting to English; label output from the detected result.language, not the (possibly auto) param. - --fast now provisions the multilingual tiny model unless --language en is explicit, since tiny.en silently mis-transcribes everything else. - Add tiny, medium and large-v3 pinned downloads so model_size medium/large actually resolve to a real ggml file instead of a 404. - Pin every whisper.cpp/VAD download to a commit revision instead of resolve/main so an upstream re-upload can't swap bytes behind an already-pinned hash.
- auto.md: replace nonexistent transcribe_status with job_status in allowed-tools and in the transcription poll loop - auto.md, bootstrap-knowledge.md: restore the dropped .podcli/knowledge/ path in the blank `` placeholders - auto.md: fallback instruction names both podcli studio and npm run ui instead of only the source-checkout command - generate-titles.md, plan-episode.md, process-transcript.md, produce-shorts.md: add the missing mcp__podcli__knowledge_base to allowed-tools since the body already calls it - server.ts/version.ts: add studioStartCommand(), which picks podcli studio vs npm run ui based on whether PODCLI_VERSION (set only by the Go launcher) is present, and use it everywhere the server tells the agent to start the Web UI - add a vitest contract test that parses every command's frontmatter and body against the server's actually-registered tool names
readUIState() now falls back to reading paths.uiState straight off disk when the Web UI's HTTP API is unreachable, instead of reporting state as unavailable. get_ui_state's handler now goes through readUIState() too (it previously duplicated the fetch and only reported the Web UI as down), so it only falls back to start-from-scratch guidance when there is truly no state anywhere.
A single viral or flop clip was swinging a small bucket's mean retention/CTR/ views far past what's typical. Switch _agg to statistics.median, and drop buckets under 4 clips entirely instead of reporting a number nobody should trust yet.
…rs in rendered captions
…e key
- Resolve the engine before reading the cache (new resolve_transcribe_engine
backend call) instead of reading with the raw unset/requested engine: an
unset request that falls back to whisper.cpp was written under
'whispercpp' but read under the bare key, so it never hit.
- web-server.ts's cache.set after a background transcribe job also wrote
under the request-time engine instead of what transcribe_file actually
ran with; it now uses the resolved engine from the result.
- Widen the raw JSON cache key to {engine, model, language}; base model and
auto/empty language contribute no suffix so existing cache entries keep
reading the same way they always have.
- Make cache writes atomic (write to a temp file, then rename).
…lease-hygiene check - ci.yml python job now installs ffmpeg on all three OSes (same per-OS steps nightly.yml already uses). Without it, every test that shells out to ffmpeg/ffprobe was silently skipped, about 22 tests on every run. - test_ai_fallback.py's shell-lookup-fallback test now redirects HOME to a temp dir before calling _find_cli: without it, a real ~/.local/bin/claude on the machine running the test wins before the shell-lookup mock is ever exercised, so the test could pass for the wrong reason. - vitest.config.ts excludes dist/, since `npm run build` compiles src/**/*.test.ts into dist/**/*.test.js and vitest would otherwise run every test twice against stale output. - add scripts/release-hygiene.mjs: fails on tracked audio/video files, tracked files over 1MB not on a short allowlist (the three that are already legitimately tracked), tracked .env files, and a few common secret patterns. Wired into CI as its own job.
… cuts - _map_range required overlap_end > overlap_start, which a zero-duration word (start == end) can never satisfy even when it sits inside a kept range; map it as a point instead of dropping it, and keep the 10ms-floor filter from also swallowing it. - A word that straddles a removed cut (part in one kept range, part in the gap or the next one) is now assigned wholly to whichever side holds its midpoint instead of being stitched across the cut, which silently absorbed the cut into the word's own duration. Segment-level stitching across a genuine gap (sentences spanning real pauses) is unchanged.
Covers multicam sync, auto-cut to speaker, transcription engines and diarization, caption styles and font coverage, thumbnail generation, and export targets, each with a plain can-do/cannot-do statement.
…ments drift - transcript_parser.py hardcoded language: en for every format (SRT, VTT, JSON, speaker-labeled). None of these formats carry language info of their own, so default to 'und' (undetermined) and accept an optional language override threaded through parse_transcript end to end (MCP schema, web-server route, Python handler). - speaker_segments used each block's raw start/end while words and segments applied time_adjust, so a non-zero adjust put speaker boundaries out of sync with the words and segments inside them. Apply time_adjust there too.
installCommands used to skip any command file that already existed, so a user never got command fixes/improvements after podcli update unless they deleted the file first. It now tracks a sha256 manifest (.claude/commands/.podcli-manifest.json) of what it last wrote per file: unmodified files refresh to the new embedded version, user-edited files are left alone and reported, and files that predate the manifest are adopted as a baseline rather than silently overwritten. Add the podstack package's first tests, covering first install, idempotent re-install, user-edit preservation, staleness updates, and pre-manifest adoption.
_voiced_intervals built an (nf, 400) int64 index matrix to gather every 25ms window at once for its RMS calculation -- about 1GB of indices alone per hour of 16kHz audio, before even touching the gathered samples. A cumulative sum of squared samples gives the identical per-window RMS in O(n) memory. Verified against the old implementation on synthetic mixed voiced/silent audio across several seeds.
podcli setup and podcli mcp install now also register the MCP server with Codex (codex mcp add podcli -- <self> mcp) when the codex CLI is on PATH, matching the existing Claude Code registration. installCodexSkills installs every PodStack command as a Codex skill under ~/.codex/skills/<name>/SKILL.md (confirmed against a local codex install — that's the real path it reads, not .codex/prompts), reusing the same install-manifest mechanism as .claude/commands so upgrades refresh unmodified skills and leave user edits alone. podstack.Run() no longer falls back from Claude to Codex on a nonzero exit code — only when Claude isn't on PATH at all. Running the same destructive workflow twice under two different agents because the first one's exit code was nonzero (e.g. the user cancelled) was never the intent. Add a root AGENTS.md pointing at CLAUDE.md and AGENTS.podstack.md, and correct AGENTS.podstack.md's host-compatibility table, which claimed a .codex/prompts install path nothing wrote. Add a Codex section to the Web UI's MCP setup page alongside Claude Desktop and Claude Code.
transcribe_podcast/transcribe_start take optional start_seconds and duration_seconds; transcribe_file extracts just that window (via a trimmed 16kHz mono wav) instead of decoding the whole file, skips diarization and face analysis, and marks the result complete: false with sample_offset_seconds carrying where in the source the window started. A sample never reads or writes the main transcript cache, the packed markdown view, the energy/event signal cache, or the UI/session transcript state — all of those are keyed by, and assumed to describe, the whole file. No CLI wiring: there is no standalone 'transcribe' subcommand in cli.py to attach --start-seconds/--duration-seconds to; 'process' runs the full clip pipeline, which a 40s sample isn't meant to feed. Left out rather than bolting the flags onto a command they don't fit.
…ines
audio_stream_index let a camera with one mono stream per mic pick which
stream carries that mic for sync, but _write_mix and _render_stems
hardcoded [{i}:a:0] regardless, so the rendered mix and stems always
pulled the first stream. The xmeml and FCPXML exports had the matching
bug: trackindex and srcCh were built from the channel within a stream
alone, so a mic on a later stream always mapped to track/channel 1 in the
editor.
cmd_process's skip_transcript and normal transcribe paths knew the exact (engine, model, language) they were about to run with, but still read and wrote the transcript cache with none of it, landing on the implicit base/auto key, or silently colliding with a different model's entry under the Python-only engine-only key this branch predates. Pass them through explicitly, matching backend/main.py's handle_transcribe. Readers that genuinely don't know the model/language (manage_reel, _cached_face_map) are left on the scanning fallback added for exactly that case.
A bare set_video (videoPath with no transcript in the same request) clears the server's transcript, suggestions and selections, since they describe a recording that's no longer loaded. The SSE broadcast only echoed back whichever fields the request body itself carried, so a request that sent only videoPath never told the studio those were cleared; the studio's stale in-memory copies then wrote themselves back on its next sync. Force transcript and suggestions/deselectedIndices onto the broadcast whenever the video change cleared them. The client already applies transcript:null and suggestions:[] correctly once they're present on the payload.
_TIMECODE_RE only matched ':' and ';' separators, so an embedded timecode written with '.' (some cameras and field recorders use it) failed to match at all and _timecode_seconds silently returned 0.0 instead of the real start time. Only the drop-frame marker before the frame field is special; allow '.' alongside ':' everywhere else.
…nd one The studio synced videoPath and transcript to the server in two separate POST /api/ui-state requests. After silence removal both change together in one render, but the requests could still land in either order: a videoPath-only request arriving after the transcript-only one looked exactly like a bare set_video (videoPath with no transcript), and the server cleared the transcript that had just arrived. Fold transcript into the same combined sync request whenever it changed, so the two fields can't be split across requests that race. Extracted the payload-building into buildUiStateSyncPayload (lib.ts) so the fix is unit-testable without a React/DOM harness, which this project doesn't have. The function is new, so there's no prior version to show the regression tests failing against; they verify the fixed behavior directly.
…erage avg_frame_rate is measured, not declared, and can read a steady 25 fps camera as 24.98 by noise alone. _ntsc_frame_duration decided pulldown by checking that measured float's closeness to a whole number, so that noise alone could misread a plain-integer-rate file as NTSC and shift its embedded start timecode by seconds. r_frame_rate is the stream's exact time base; every true pulldown rate reduces to a fraction with denominator 1001, so use that when it's available instead.
…rs it already dropped _sync_review_reasons flagged a fit for review whenever max(residual_ms, residual_all_ms) exceeded 30 ms. residual_all_ms includes every checkpoint, outliers included, so a single checkpoint correctly rejected as an outlier was enough to flag an otherwise clean fit for review on every sync. residual_ms is the residual the fit was actually built from; use that alone for the threshold and keep reporting residual_all_ms alongside it.
Style cleanup only, no behavior change: replace em dash punctuation in comments and docstrings added across the preceding commits with a comma, colon, or period, per the project's writing conventions.
These tests render with captions off and inspect the sidecar words via a spy on write_sidecars, which the captions-off-means-no-sidecars change (write_sidecars now follows the captions flag by default) stopped calling. Request subtitles explicitly, same as the other tests that need them without burned captions.
The 300M model takes no language hint and wrote accented English in Arabic script on a real podcast episode. The 1B v2 model transcribed the same audio correctly at about 8x realtime on the CPU. Existing 300M files fail the new hash check and get replaced on the next run.
# Conflicts: # src/ui/client/lib.test.ts
render_session swapped the video live, then the stems, as separate os.replace calls onto fixed filenames, contradicting its own comment: a crash between those renames left the new video next to old stems (or the reverse), and because the filenames never changed between renders, a save that failed after the renames succeeded left the session record pointing at paths whose bytes had already silently become the new render's. It also never removed a stem for a person dropped from the edit, so a stale file kept sitting at a name that looked current. Build the whole render in a fresh directory nothing points at yet, then flip the session to it with MulticamSession.save()'s own tmp-file-plus- os.replace write, the one atomic step available across several files. Drop the previous render's directory only after that save succeeds.
…anges An export started over MCP with no studio tab open lost its results, so the studio later showed every clip as unexported; the server now saves them when the job finishes. The studio engine picker offers Omnilingual, and the backend fetches its pinned model on first use from any entry point. record_decisions no longer stores max 0 when only a minimum was given.
_render_audio measured the episode's target gain on the premix before removals were spliced out, so a retake or a long silence that never airs still pulled the target gain toward its own loudness. Splice the premix first when there are removals, then measure and encode from what's actually kept, so the gain fits the episode anyone hears.
A LUT picked at map time can be moved or deleted on disk before the next render or preview. Nothing checked for that, so the failure surfaced from inside an ffmpeg worker partway through a shot or a still, as a bare 'No such file or directory' on the LUT path with no indication of which camera (or which tile in a split) it belonged to. Check every mapped LUT up front in both render_session and previews, and name the camera in the error.
_decode_errors fully decoded the finished episode on every render, costing minutes per hour of 1080p for a check that almost always finds nothing: a bad render shows up at a concat seam, a mux fault, or a codec error at the point it happened, not uniformly across the file. Default to decoding the first and last 10s plus three points through the middle, and keep the frame count and loudness checks exactly as they were. A full decode is still available: render_session(..., validate="full"), wired through the CLI's --validate and the MCP tool's render action.
Stills are keyed by filename on the timeline moment and a basis hash of offset, speed and LUT, so a nudge or re-sync changes the key and the old file never matches again. Nothing deleted it, so every nudge during a mapping session left one more frame and one more per-look still behind. Delete the stale ones for that camera (or that camera's look) whenever a fresh still actually needs to render.
previews(looks=True) rendered all four looks for every mapped camera, even one that never appears in the planned cut (an alternate angle left mapped but unused). Once a cut exists, filter to the cameras it actually uses; before any cut exists there's nothing to filter by, so every camera still gets looks at that stage, same as before. Cached stills continue to be reused exactly as they were.
…ne write" This reverts commit 269ebdd.
Rendering into a new folder each time moved episode.mp4 on every render, which breaks anyone who opens or scripts against the episode path. Keep the staged same-volume renames, say plainly what they guarantee, and remove stems for people no longer in the edit.
A caller-supplied model path made the first-use download land beside it, so two tests that swap in a temp model pulled the real 1 GB file and the suite stalled for most of an hour.
|
Important Review skippedToo many files! This PR contains 153 files, which is 53 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry. ⚙️ Run configuration
📒 Files selected for processing (153)
You can disable this status message by setting the
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
record_decisions used the zod 3 errorMap option; the locked zod 4 takes error. Tests: Windows forbids a colon in a folder name, a pair thumbnail URL is the path's file URI, and a three-part move hook can round each part up to a frame on another ffmpeg build.
Hardens the clip, transcription and multicam pipelines, adds an opening hook for clips, a local 1600-language engine, and agent session memory. Tested end to end on a full DeepTech Decoded episode (68 min, 4K).
Clips
repeat) or with that line removed from the body (move). Captions, subtitles and the clean copy follow playback order.modify_cliprange edits now reach the render, including widening a clip.{\an8}are stripped on import.get_ui_state.Transcription
omnilingualengine: Meta's Omnilingual ASR 1B CTC model on the CPU, pinned by revision and SHA-256, fetched on first use. Resumable per window.compare_transcription_engineswrites a side-by-side disagreement report for a sample.start_seconds,duration_seconds) for testing a language on 40 s.Multicam
color_handoff.json, audio stream selection, Premiere scale-to-fit, VFR warnings.Agents and install
record_decisionsstores per-episode answers;get_ui_statelists open questions and works with the studio closed.AGENTS.md.podcli doctorruns each binary, hash-checks models and exits nonzero on real failures.mine_channellists a channel's uploads and mines a video's captions (json3 word timing, no rolling repeats, URL validated,--before positional args)./autocalls real tool names.Verification
Known limits
omnilingualreturns lowercase text with no punctuation and downloads 1 GB on first use.suggest_clipsupdates the studio phase without awaiting it, so an immediateget_ui_statecan show the previous phase. Same on main.