Files
gui-video-clipper/chat-summaries/2026-09-22_00-21-rolling-karaoke-subtitles-summary.md
cottongin 8ad2f1c800 chore: stage all pending work — caption styling, media server, processing modal, docs, summaries
Includes:
- Extended caption styling (font, shadow, dimmed color, bg toggle)
- Media server, subtitle downloader, VTT parser, processing modal
- Waveform tiers, thumbnail/timeline improvements, transport controls
- Hybrid download model, dependency management, clip export enhancements
- 21 chat summaries, 2 implementation plans, 2 design specs

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-22 10:48:16 -04:00

3.9 KiB

Rolling Two-Line Karaoke Subtitle Display

Date: 2026-09-22 00:21 Task: Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering

Problem

The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard \N break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.

Prior Art

  • Sofronio/YouTubeVTT2ASS — C# tool that solves this exact problem using \move ASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines."
  • The key technique: for each spoken line, generate 3 ASS Dialogue entries with \move animations (Active → Context → Disappear).

Changes Made

File: src-tauri/src/services/clip_exporter.rs

New structs and functions:

  1. SpokenLine struct — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag.

  2. extract_spoken_lines() — Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation.

  3. build_karaoke_text() — Extracts word timings and builds \k karaoke tags for a single line.

  4. RollingLayout struct — Position parameters: cx (960), y_bottom (1040), line_height (font_size * 1.3), scroll_ms (350).

Rewritten vtt_to_ass_with_karaoke():

For each spoken line, generates up to 3 ASS Dialogue entries:

  • Phase 1 (Active/Karaoke): Line appears at bottom position, scrolls up one slot via \move(cx, y_bottom, cx, y_bottom-h, 0, 350). Has \k karaoke tags. Lasts from this line's start to the next line's start.

  • Phase 2 (Context/Static): Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.

  • Phase 3 (Disappear): Scrolls off-screen with \clip mask to cleanly cut off. Lasts 500ms.

Edge cases handled:

  • Long gaps (>2s): Context phase ends early; line disappears instead of lingering through silence/music.
  • Non-speech cues ([Music]): Single static Dialogue with \pos instead of rolling.
  • First line: No context above it — just starts normally.
  • Last line: Phase 1 uses line.end_time; Phase 2/3 use a hold + disappear.
  • Lines without karaoke data: Rendered as plain text with the same rolling behavior.

Style additions:

  • Added ScaledBorderAndShadow: yes to [Script Info] for proper scaling.
  • Each Dialogue line uses \an2 override for explicit bottom-center positioning.

Lessons Learned

  1. YouTube's VTT two-line pattern is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires \move animations, not just \N line breaks.

  2. ASS \move with \an2 — The alignment setting determines the anchor point for positioning. \an2 (bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas.

  3. Three-phase lifecycle per line is the key insight from YouTubeVTT2ASS. Phase 3 with \clip is important to cleanly mask the text as it scrolls off instead of having it abruptly disappear.

  4. Gap detection is essential — without it, stale context lines would linger through long silences or [Music] sections.

Build/Test Status

  • cargo build: Success (6 pre-existing warnings, 0 errors)
  • cargo test: 43 tests passed, 0 failed