Files
gui-video-clipper/chat-summaries/2026-09-22_00-21-rolling-karaoke-subtitles-summary.md

66 lines
3.9 KiB
Markdown
Raw Normal View History

# Rolling Two-Line Karaoke Subtitle Display
**Date:** 2026-09-22 00:21
**Task:** Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering
## Problem
The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard `\N` break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.
## Prior Art
- [Sofronio/YouTubeVTT2ASS](https://github.com/Sofronio/YouTubeVTT2ASS) — C# tool that solves this exact problem using `\move` ASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines."
- The key technique: for each spoken line, generate 3 ASS Dialogue entries with `\move` animations (Active → Context → Disappear).
## Changes Made
**File:** `src-tauri/src/services/clip_exporter.rs`
### New structs and functions:
1. **`SpokenLine` struct** — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag.
2. **`extract_spoken_lines()`** — Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation.
3. **`build_karaoke_text()`** — Extracts word timings and builds `\k` karaoke tags for a single line.
4. **`RollingLayout` struct** — Position parameters: `cx` (960), `y_bottom` (1040), `line_height` (font_size * 1.3), `scroll_ms` (350).
### Rewritten `vtt_to_ass_with_karaoke()`:
For each spoken line, generates up to 3 ASS Dialogue entries:
- **Phase 1 (Active/Karaoke):** Line appears at bottom position, scrolls up one slot via `\move(cx, y_bottom, cx, y_bottom-h, 0, 350)`. Has `\k` karaoke tags. Lasts from this line's start to the next line's start.
- **Phase 2 (Context/Static):** Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.
- **Phase 3 (Disappear):** Scrolls off-screen with `\clip` mask to cleanly cut off. Lasts 500ms.
### Edge cases handled:
- **Long gaps (>2s):** Context phase ends early; line disappears instead of lingering through silence/music.
- **Non-speech cues (`[Music]`):** Single static Dialogue with `\pos` instead of rolling.
- **First line:** No context above it — just starts normally.
- **Last line:** Phase 1 uses `line.end_time`; Phase 2/3 use a hold + disappear.
- **Lines without karaoke data:** Rendered as plain text with the same rolling behavior.
### Style additions:
- Added `ScaledBorderAndShadow: yes` to `[Script Info]` for proper scaling.
- Each Dialogue line uses `\an2` override for explicit bottom-center positioning.
## Lessons Learned
1. **YouTube's VTT two-line pattern** is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires `\move` animations, not just `\N` line breaks.
2. **ASS `\move` with `\an2`** — The alignment setting determines the anchor point for positioning. `\an2` (bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas.
3. **Three-phase lifecycle per line** is the key insight from YouTubeVTT2ASS. Phase 3 with `\clip` is important to cleanly mask the text as it scrolls off instead of having it abruptly disappear.
4. **Gap detection** is essential — without it, stale context lines would linger through long silences or `[Music]` sections.
## Build/Test Status
- `cargo build`: Success (6 pre-existing warnings, 0 errors)
- `cargo test`: 43 tests passed, 0 failed