Includes: - Extended caption styling (font, shadow, dimmed color, bg toggle) - Media server, subtitle downloader, VTT parser, processing modal - Waveform tiers, thumbnail/timeline improvements, transport controls - Hybrid download model, dependency management, clip export enhancements - 21 chat summaries, 2 implementation plans, 2 design specs Co-authored-by: Cursor <cursoragent@cursor.com>
3.9 KiB
Rolling Two-Line Karaoke Subtitle Display
Date: 2026-09-22 00:21 Task: Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering
Problem
The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard \N break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.
Prior Art
- Sofronio/YouTubeVTT2ASS — C# tool that solves this exact problem using
\moveASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines." - The key technique: for each spoken line, generate 3 ASS Dialogue entries with
\moveanimations (Active → Context → Disappear).
Changes Made
File: src-tauri/src/services/clip_exporter.rs
New structs and functions:
-
SpokenLinestruct — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag. -
extract_spoken_lines()— Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation. -
build_karaoke_text()— Extracts word timings and builds\kkaraoke tags for a single line. -
RollingLayoutstruct — Position parameters:cx(960),y_bottom(1040),line_height(font_size * 1.3),scroll_ms(350).
Rewritten vtt_to_ass_with_karaoke():
For each spoken line, generates up to 3 ASS Dialogue entries:
-
Phase 1 (Active/Karaoke): Line appears at bottom position, scrolls up one slot via
\move(cx, y_bottom, cx, y_bottom-h, 0, 350). Has\kkaraoke tags. Lasts from this line's start to the next line's start. -
Phase 2 (Context/Static): Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.
-
Phase 3 (Disappear): Scrolls off-screen with
\clipmask to cleanly cut off. Lasts 500ms.
Edge cases handled:
- Long gaps (>2s): Context phase ends early; line disappears instead of lingering through silence/music.
- Non-speech cues (
[Music]): Single static Dialogue with\posinstead of rolling. - First line: No context above it — just starts normally.
- Last line: Phase 1 uses
line.end_time; Phase 2/3 use a hold + disappear. - Lines without karaoke data: Rendered as plain text with the same rolling behavior.
Style additions:
- Added
ScaledBorderAndShadow: yesto[Script Info]for proper scaling. - Each Dialogue line uses
\an2override for explicit bottom-center positioning.
Lessons Learned
-
YouTube's VTT two-line pattern is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires
\moveanimations, not just\Nline breaks. -
ASS
\movewith\an2— The alignment setting determines the anchor point for positioning.\an2(bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas. -
Three-phase lifecycle per line is the key insight from YouTubeVTT2ASS. Phase 3 with
\clipis important to cleanly mask the text as it scrolls off instead of having it abruptly disappear. -
Gap detection is essential — without it, stale context lines would linger through long silences or
[Music]sections.
Build/Test Status
cargo build: Success (6 pre-existing warnings, 0 errors)cargo test: 43 tests passed, 0 failed