Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Open In Colab

ReelTranscriber

A Python pipeline designed for qualitative social science research. It automatically scrapes audio from short-form social media videos (Instagram Reels, TikTok, YouTube Shorts), runs hardware-accelerated speech-to-text recognition via OpenAI Whisper, and exports sentence-indexed transcripts formatted into a PDF audit report.

Features

  • Automated Video Scraper: Ingests direct URLs from Instagram, TikTok, and YouTube via yt-dlp.
  • High-Accuracy GPU Transcription: Powered by OpenAI's Whisper model (medium model by default).
  • Sentence Indexing ([S1], [S2]): Automatically splits raw speech into discrete, numbered sentence units suitable for qualitative manual coding in software like NVivo or MAXQDA.
  • Publication-Ready PDF Export: Generates a PDF report containing video metadata, verbatim transcripts, and sentence index breakdowns.

Key Use Cases

ReelTranscriber is designed to streamline workflows across research, content creation, and data analysis:

  • Qualitative Academic Research: Instantly convert large batches of short-form social media videos (Instagram Reels, TikTok, YouTube Shorts) into sentence-indexed transcripts ([S1], [S2]) for deductive content analysis, thematic coding, or manual ingestion into qualitative software like NVivo and MAXQDA.
  • Content Creator Auditing: Extract spoken audio from high-performing viral Reels to analyze script structures, hook pacing, and engagement calls-to-action (CTAs).
  • Dataset Generation for NLP: Generate structured, sentence-segmented text datasets from spoken social media audio for natural language processing, sentiment analysis, or topic modeling tasks.
  • Accessibility & Repurposing: Quickly convert video speech into structured text documents to produce video captions, blog summaries, or downloadable PDF show notes.

Getting Started

Prerequisites

To run this pipeline locally or in Google Colab, ensure you have:

  • Python 3.9+
  • An active GPU environment (recommended: NVIDIA T4 or higher)
  • System package: ffmpeg (installed automatically in Google Colab)

Dependencies

Install required Python packages:

pip install openai-whisper yt-dlp reportlab torch pandas openpyxl

About

Extract audio from Instagram Reels, TikTok, and YouTube Shorts into sentence-indexed PDF transcripts for qualitative content analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages