66 lines
3.9 KiB
Markdown
66 lines
3.9 KiB
Markdown
|
|
# Rolling Two-Line Karaoke Subtitle Display
|
||
|
|
|
||
|
|
**Date:** 2026-09-22 00:21
|
||
|
|
**Task:** Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering
|
||
|
|
|
||
|
|
## Problem
|
||
|
|
|
||
|
|
The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard `\N` break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.
|
||
|
|
|
||
|
|
## Prior Art
|
||
|
|
|
||
|
|
- [Sofronio/YouTubeVTT2ASS](https://github.com/Sofronio/YouTubeVTT2ASS) — C# tool that solves this exact problem using `\move` ASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines."
|
||
|
|
- The key technique: for each spoken line, generate 3 ASS Dialogue entries with `\move` animations (Active → Context → Disappear).
|
||
|
|
|
||
|
|
## Changes Made
|
||
|
|
|
||
|
|
**File:** `src-tauri/src/services/clip_exporter.rs`
|
||
|
|
|
||
|
|
### New structs and functions:
|
||
|
|
|
||
|
|
1. **`SpokenLine` struct** — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag.
|
||
|
|
|
||
|
|
2. **`extract_spoken_lines()`** — Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation.
|
||
|
|
|
||
|
|
3. **`build_karaoke_text()`** — Extracts word timings and builds `\k` karaoke tags for a single line.
|
||
|
|
|
||
|
|
4. **`RollingLayout` struct** — Position parameters: `cx` (960), `y_bottom` (1040), `line_height` (font_size * 1.3), `scroll_ms` (350).
|
||
|
|
|
||
|
|
### Rewritten `vtt_to_ass_with_karaoke()`:
|
||
|
|
|
||
|
|
For each spoken line, generates up to 3 ASS Dialogue entries:
|
||
|
|
|
||
|
|
- **Phase 1 (Active/Karaoke):** Line appears at bottom position, scrolls up one slot via `\move(cx, y_bottom, cx, y_bottom-h, 0, 350)`. Has `\k` karaoke tags. Lasts from this line's start to the next line's start.
|
||
|
|
|
||
|
|
- **Phase 2 (Context/Static):** Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.
|
||
|
|
|
||
|
|
- **Phase 3 (Disappear):** Scrolls off-screen with `\clip` mask to cleanly cut off. Lasts 500ms.
|
||
|
|
|
||
|
|
### Edge cases handled:
|
||
|
|
|
||
|
|
- **Long gaps (>2s):** Context phase ends early; line disappears instead of lingering through silence/music.
|
||
|
|
- **Non-speech cues (`[Music]`):** Single static Dialogue with `\pos` instead of rolling.
|
||
|
|
- **First line:** No context above it — just starts normally.
|
||
|
|
- **Last line:** Phase 1 uses `line.end_time`; Phase 2/3 use a hold + disappear.
|
||
|
|
- **Lines without karaoke data:** Rendered as plain text with the same rolling behavior.
|
||
|
|
|
||
|
|
### Style additions:
|
||
|
|
|
||
|
|
- Added `ScaledBorderAndShadow: yes` to `[Script Info]` for proper scaling.
|
||
|
|
- Each Dialogue line uses `\an2` override for explicit bottom-center positioning.
|
||
|
|
|
||
|
|
## Lessons Learned
|
||
|
|
|
||
|
|
1. **YouTube's VTT two-line pattern** is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires `\move` animations, not just `\N` line breaks.
|
||
|
|
|
||
|
|
2. **ASS `\move` with `\an2`** — The alignment setting determines the anchor point for positioning. `\an2` (bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas.
|
||
|
|
|
||
|
|
3. **Three-phase lifecycle per line** is the key insight from YouTubeVTT2ASS. Phase 3 with `\clip` is important to cleanly mask the text as it scrolls off instead of having it abruptly disappear.
|
||
|
|
|
||
|
|
4. **Gap detection** is essential — without it, stale context lines would linger through long silences or `[Music]` sections.
|
||
|
|
|
||
|
|
## Build/Test Status
|
||
|
|
|
||
|
|
- `cargo build`: Success (6 pre-existing warnings, 0 errors)
|
||
|
|
- `cargo test`: 43 tests passed, 0 failed
|