Kinetic Video Typography with Python: Generating Word-Level Animated Subtitles using Whisper & Pillow

Kinetic Video Typography with Python: Generating Word-Level Animated Subtitles using Whisper & Pillow

(Updated: ) ๐Ÿ“– 1 min read

On TikTok, YouTube Shorts, and Instagram Reels, viewers scroll with sound muted over 60% of the time. The difference between a video that drops off in 2 seconds and one that goes viral is kinetic, word-by-word karaoke-style animated captions.

Here is how to extract word-level millisecond timestamps with Whisper and burn animated typography frames onto videos with Python.


1. Extracting Word Timestamps with Faster-Whisper

from faster_whisper import WhisperModel

model = WhisperModel("base", device="cpu", compute_type="int8")

segments, _ = model.transcribe("voiceover.mp3", word_timestamps=True)

words_timeline = []
for segment in segments:
    for word in segment.words:
        words_timeline.append({
            "word": word.word.strip(),
            "start": word.start,
            "end": word.end
        })

print(f"Extracted {len(words_timeline)} synchronized words.")

2. Dynamic Word Color Highlight Logic

To create the popular gold-highlight karaoke effect, group words into 4-word chunks. As the video clock matches a wordโ€™s [start, end] window, render that specific word in vibrant gold (#FBBF24) while holding surrounding words in clean white (#FFFFFF).

from PIL import Image, ImageDraw, ImageFont

def render_subtitle_card(chunk_words, current_time_sec):
    img = Image.new("RGBA", (1080, 200), (0, 0, 0, 0))
    draw = ImageDraw.Draw(img)
    font = ImageFont.truetype("Montserrat-Bold.ttf", size=54)

    x_cursor = 100
    y_pos = 60

    for item in chunk_words:
        word = item["word"]
        is_active = item["start"] <= current_time_sec <= item["end"]
        color = (251, 191, 36, 255) if is_active else (255, 255, 255, 255)

        draw.text((x_cursor, y_pos), word, font=font, fill=color)
        bbox = draw.textbbox((x_cursor, y_pos), word, font=font)
        x_cursor += (bbox[2] - bbox[0]) + 18 # Add space between words

    return img

3. Key Video Metrics

  • Average Retention Lift: +42% watch time on videos with word-level highlights.
  • Cognitive Ease: Synced reading reinforces voiceover Maqams and emotional cadence.
DEVELOPER TOOLKIT

Download the Headless Video Automation Pipeline Code

The Python/FFmpeg template with caption overlays and custom transitions to programmatically generate social media reels & shorts.

Professor XAI
Professor XAI ML Engineer passionate about advancing AI technologies and building intelligent systems.
comments powered by Disqus