Smart scene captions generator (2025 edition) - transform any video into a concise, shot‑by‑shot English description in seconds.
- Automatic scene detection using
PySceneDetect’s content detector - Vision‑language captions from two model families: Qwen2‑VL (tested on Qwen‑2 VL‑2B‑Instruct) and LLaVA‑style repos such as apple/FastVLM - the backend is picked from the model config
- Contact‑sheet collage (optional) for multi‑frame scene context
- Pure‑Python stack, CPU‑friendly, GPU/MPS acceleration when available
- Clean one‑line English captions that keep names, actions & a mandatory shot size tag
- CLI & Python API, configurable in a single
config.py - Apache‑licensed
- Quick start
- Installation
- CLI reference
- Python API
- Configuration
- How it works
- Performance tips
- License
# 1. install dependencies (see below for details)
I've used conda
Python ≥3.10, Torch ≥2.3, Transformers ≥4.53, PySceneDetect ≥0.6.6.
# 2. run on the demo video
python3 src/main.py data/example.mp4 \
--model Qwen/Qwen2-VL-2B-Instruct \
--force-cpu \
--thr 50 \
--max-frames 2Example output (truncated):
00:00:00 - In the frame, a river flows through a valley, with a town visible on its banks, medium
00:00:43 - In the frame, a large building with a courtyard and an ornate tower is depicted, medium
00:01:21 - In the frame, a medieval castle is situated on a hilltop, overlooking a town below, medium
| Requirement | Tested version |
|---|---|
| Python | 3.10 – 3.12 |
| PyTorch | 2.6 |
| ffmpeg | ≥4.3 |
SceneDescriber relies on opencv‑python, pillow, transformers, scenedetect and a few utilities:
pip install -r requirements.txtGPU acceleration is detected automatically (CUDA, Apple M‑series, Intel XPU). Set
--force-cputo bypass.
usage: main.py [-h] [--model MODEL] [--thr THR] [--max-frames N] [--force-cpu]
video
positional arguments:
video path to the input video file (any format ffmpeg understands)
optional arguments:
--model MODEL Hugging Face ID of the vision‑language model (default: Qwen/…‑2B‑Instruct)
--thr THR scene‑change threshold on a 0‑100 scale; a 0‑1 fraction
(e.g. 0.27) is read as a percentage. Higher → fewer scenes
(default: 27.0)
--max-frames N max frames sampled per scene (speed/quality trade‑off)
--force-cpu disable any GPU/MPS device and run on CPU only
| Goal | Command |
|---|---|
| Faster captions on a short clip | python3 src/main.py clip.mp4 --max-frames 1 |
| More detailed captions (risk of repetition) | python3 src/main.py clip.mp4 --max-frames 6 --thr 20 |
| High‑end GPU, half‑precision | USE_FP16=1 python3 src/main.py clip.mp4 --model <your‑HF‑id> |
| Apple FastVLM backend | python3 src/main.py clip.mp4 --model apple/FastVLM-1.5B |
from src.main import run
lines = run(
video_path="data/example.mp4",
model_id="Qwen/Qwen2-VL-2B-Instruct",
threshold=27.0, # 0-100 scale (0.27 works too)
max_frames_per_scene=3,
force_cpu=True,
)
print("\n".join(lines))Each element in lines is a ready‑to‑print caption starting with a hh:mm:ss timestamp.
Most knobs live in config.py :
| Name | Default | Meaning |
|---|---|---|
DEFAULT_SCENE_THRESHOLD |
27.0 |
Scene detector aggressiveness, ContentDetector 0–100 scale |
MAX_FRAMES_PER_SCENE |
3 |
Hard cap per scene (speed) |
CONTACT_SHEET |
True |
Create a collage to improve context |
TILE |
448 |
All frames are resized to TILE×TILEbefore inference |
SHEET_TILE |
896 |
Collage side - bigger, so individual tiles stay readable |
USE_FP16 |
False |
Run the model in half‑precision on GPU/MPS (env USE_FP16=1) |
LLAVA_MAX_NEW_TOKENS |
96 |
Token budget for LLaVA‑style models, which ramble past it |
TEMPERATURE / TOP_P |
0.4 / 0.9 |
Sampling; low temperature keeps captions factual |
REPETITION_PENALTY |
1.1 |
Suppresses the "and … and …" loops of small models |
--thris normalised: a value ≤ 1 is read as a fraction (0.30→30) and0is clamped to1- a literal0threshold makes the detector cut on every frame.
Set environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 for smoother Apple Silicon experience (already applied by default).
- Scene detection –
pyscenedetectfinds hard cuts with a configurable content threshold. - Uniform frame sampling – up to
Nframes picked per scene to represent motion. - Pre‑processing – frames are down‑scaled, padded to square and optionally collaged.
- Caption generation – frames + the prompt go to the model. Qwen2‑VL takes them through
AutoProcessor; LLaVA‑style repos (FastVLM) get one image, a hand‑spliced-200image token and a second one‑word question for the shot size, since they answer one question per pass. - Formatting – the reply is trimmed to its first sentence, the shot size is appended, and the line is prefixed with the scene’s start timestamp.
- Threshold tuning – try
--thr 20 ... 60; lower means more scenes (finer granularity). Avoid--thr 0. - Batch processing – this repo focuses on single‑video simplicity; wrap
run()in your own loop for large datasets. - GPU half‑precision – export
USE_FP16=1to cut VRAM usage almost in half. - Collage off – set
CONTACT_SHEET = Falseif you prefer random single frames (LLaVA‑style models always get a single frame: they have one image slot).
This project is released under the Apache License .
See LICENSE for details.
- PySceneDetect
- Transformers
- Qwen‑2 VL from Alibaba & ModelBest
Made with passion by @Shel2123 - let the machines watch the video so you don’t have to!