SWIFT performs normal and corrupted sliding-window reconstructions and uses their average loss ratio as the attribution signal. A model-specific threshold is estimated offline and used for online origin attribution.
SWIFT uses the original implementation and environment of each supported video generation model. Please follow the installation and model-weight instructions in the corresponding official repository.
| Target model | Official environment | SWIFT script |
|---|---|---|
| Wan2.1 | Wan-Video/Wan2.1 | Wan2.1.py |
| Wan2.2 | Wan-Video/Wan2.2 | Wan2.2.py |
| HunyuanVideo | Tencent-Hunyuan/HunyuanVideo | HunyuanVideo.py |
| LTX-Video | Lightricks/ltx-video | LTX-Video.py |
| EasyAnimate | aigc-apps/EasyAnimate | EasyAnimate.py |
Clone and configure the official model repositories listed above, including its environment and VAE checkpoint.
Place the matching script in the root directory of the target model repository. For example:
cp /path/to/SWIFT/Wan2.1.py /path/to/Wan2.1/
cd /path/to/Wan2.1Before running the script, update its path variables:
vae_path, ormodel_pathandconfig_pathfor EasyAnimate: local VAE/model files.video_folder: directory containing the videos to evaluate.output_file: file used to store the reconstruction losses.
Videos are processed in sorted filename order.
Run the script from the root of its corresponding official repository:
# Wan2.1
python Wan2.1.py
# Wan2.2
python Wan2.2.py
# HunyuanVideo
python HunyuanVideo.py
# LTX-Video
python LTX-Video.py
# EasyAnimate
python EasyAnimate.pyEach video produces two consecutive lines in the output file:
- Per-frame MSE values from the normal reconstruction.
- Per-frame MSE values from the corrupted reconstruction.
For each video, calculate the SWIFT attribution signal by averaging the ratio between the normal and corrupted reconstruction losses over their overlapping frames, as defined in Eq. (8) of the paper. The scripts use the normal window and a three-frame-shifted corrupted window; this corresponds to the selected window pair for all five released implementations, including the empirically selected W0 and W3 pair for LTX-Video.
Estimate a separate threshold for each target model from its belonging-video signals. The paper uses Gaussian kernel density estimation with Scott's bandwidth and the 95th percentile of the estimated distribution by default. Given an attribution signal t and threshold tau:
t < tau: the video belongs to the target model.t >= tau: the video does not belong to the target model.
The final threshold-estimation and accuracy-calculation stages follow the same separation used by AEDR: first determine a threshold on reference samples, then apply it to held-out signals and calculate attribution accuracy. See AEDR's cal_threshold.py and cal_acc.py as implementation references. SWIFT uses its loss-ratio signal directly and does not use AEDR's image-homogeneity/GLCM calibration.
We thank the authors of Wan2.1, Wan2.2, HunyuanVideo, LTX-Video, and EasyAnimate for releasing their work.
This project is released under the MIT License.
If you find this work useful, please consider citing our paper:
@inproceedings{wang2026swift,
title={SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution},
author={Wang, Chao and Yang, Zijin and Wang, Yaofei and Qi, Yuang and Zhang, Weiming and Yu, Nenghai and Chen, Kejiang},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={31725--31734},
year={2026}
}