Skip to content

About

Shot-by-shot video captioning CLI: scene detection, local VLM inference, timestamped one-line English descriptions.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

14 Commits

Folders and files

Repository files navigation

SceneDescriber

Smart scene captions generator (2025 edition) - transform any video into a concise, shot‑by‑shot English description in seconds.


Key Features

  • Automatic scene detection using PySceneDetect’s content detector
  • Vision‑language captions from two model families: Qwen2‑VL (tested on Qwen‑2 VL‑2B‑Instruct) and LLaVA‑style repos such as apple/FastVLM - the backend is picked from the model config
  • Contact‑sheet collage (optional) for multi‑frame scene context
  • Pure‑Python stack, CPU‑friendly, GPU/MPS acceleration when available
  • Clean one‑line English captions that keep names, actions & a mandatory shot size tag
  • CLI & Python API, configurable in a single config.py
  • Apache‑licensed

Table of Contents

  1. Quick start
  2. Installation
  3. CLI reference
  4. Python API
  5. Configuration
  6. How it works
  7. Performance tips
  8. License

Quick start

# 1. install dependencies (see below for details)
I've used conda
Python ≥3.10, Torch ≥2.3, Transformers ≥4.53, PySceneDetect ≥0.6.6.

# 2. run on the demo video
python3 src/main.py data/example.mp4 \
        --model Qwen/Qwen2-VL-2B-Instruct \
        --force-cpu \
        --thr 50 \
        --max-frames 2

Example output (truncated):

00:00:00 - In the frame, a river flows through a valley, with a town visible on its banks, medium
00:00:43 - In the frame, a large building with a courtyard and an ornate tower is depicted, medium
00:01:21 - In the frame, a medieval castle is situated on a hilltop, overlooking a town below, medium

Installation

Prerequisites

Requirement Tested version
Python 3.10 – 3.12
PyTorch 2.6
ffmpeg ≥4.3

SceneDescriber relies on opencv‑python, pillow, transformers, scenedetect and a few utilities:

pip install -r requirements.txt

GPU acceleration is detected automatically (CUDA, Apple M‑series, Intel XPU). Set --force-cpu to bypass.


CLI reference

usage: main.py [-h] [--model MODEL] [--thr THR] [--max-frames N] [--force-cpu]
               video

positional arguments:
  video                 path to the input video file (any format ffmpeg understands)

optional arguments:
  --model MODEL         Hugging Face ID of the vision‑language model (default: Qwen/…‑2B‑Instruct)
  --thr THR             scene‑change threshold on a 0‑100 scale; a 0‑1 fraction
                        (e.g. 0.27) is read as a percentage. Higher → fewer scenes
                        (default: 27.0)
  --max-frames N        max frames sampled per scene (speed/quality trade‑off)
  --force-cpu           disable any GPU/MPS device and run on CPU only

Common recipes

Goal Command
Faster captions on a short clip python3 src/main.py clip.mp4 --max-frames 1
More detailed captions (risk of repetition) python3 src/main.py clip.mp4 --max-frames 6 --thr 20
High‑end GPU, half‑precision USE_FP16=1 python3 src/main.py clip.mp4 --model <your‑HF‑id>
Apple FastVLM backend python3 src/main.py clip.mp4 --model apple/FastVLM-1.5B

Python API

from src.main import run

lines = run(
    video_path="data/example.mp4",
    model_id="Qwen/Qwen2-VL-2B-Instruct",
    threshold=27.0,          # 0-100 scale (0.27 works too)
    max_frames_per_scene=3,
    force_cpu=True,
)
print("\n".join(lines))

Each element in lines is a ready‑to‑print caption starting with a hh:mm:ss timestamp.


Configuration

Most knobs live in config.py :

Name Default Meaning
DEFAULT_SCENE_THRESHOLD 27.0 Scene detector aggressiveness, ContentDetector 0–100 scale
MAX_FRAMES_PER_SCENE 3 Hard cap per scene (speed)
CONTACT_SHEET True Create a collage to improve context
TILE 448 All frames are resized to TILE×TILEbefore inference
SHEET_TILE 896 Collage side - bigger, so individual tiles stay readable
USE_FP16 False Run the model in half‑precision on GPU/MPS (env USE_FP16=1)
LLAVA_MAX_NEW_TOKENS 96 Token budget for LLaVA‑style models, which ramble past it
TEMPERATURE / TOP_P 0.4 / 0.9 Sampling; low temperature keeps captions factual
REPETITION_PENALTY 1.1 Suppresses the "and … and …" loops of small models

--thr is normalised: a value ≤ 1 is read as a fraction (0.30 → 30) and 0 is clamped to 1 - a literal 0 threshold makes the detector cut on every frame.

Set environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 for smoother Apple Silicon experience (already applied by default).


How it works

  1. Scene detection – pyscenedetect finds hard cuts with a configurable content threshold.
  2. Uniform frame sampling – up to N frames picked per scene to represent motion.
  3. Pre‑processing – frames are down‑scaled, padded to square and optionally collaged.
  4. Caption generation – frames + the prompt go to the model. Qwen2‑VL takes them through AutoProcessor; LLaVA‑style repos (FastVLM) get one image, a hand‑spliced -200 image token and a second one‑word question for the shot size, since they answer one question per pass.
  5. Formatting – the reply is trimmed to its first sentence, the shot size is appended, and the line is prefixed with the scene’s start timestamp.

Performance tips

  • Threshold tuning – try --thr 20 ... 60; lower means more scenes (finer granularity). Avoid --thr 0.
  • Batch processing – this repo focuses on single‑video simplicity; wrap run() in your own loop for large datasets.
  • GPU half‑precision – export USE_FP16=1 to cut VRAM usage almost in half.
  • Collage off – set CONTACT_SHEET = False if you prefer random single frames (LLaVA‑style models always get a single frame: they have one image slot).

License

This project is released under the Apache License .

See LICENSE for details.


Acknowledgements


Made with passion by @Shel2123 - let the machines watch the video so you don’t have to!

About

Shot-by-shot video captioning CLI: scene detection, local VLM inference, timestamped one-line English descriptions.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages