Includes: - Extended caption styling (font, shadow, dimmed color, bg toggle) - Media server, subtitle downloader, VTT parser, processing modal - Waveform tiers, thumbnail/timeline improvements, transport controls - Hybrid download model, dependency management, clip export enhancements - 21 chat summaries, 2 implementation plans, 2 design specs Co-authored-by: Cursor <cursoragent@cursor.com>
5.5 KiB
Paged Teleprompter Burn-In Subtitle Design
Date: 2026-09-22
Scope: Rewrite the ASS subtitle generation in clip_exporter.rs to use a paged teleprompter model instead of the current per-line rolling model.
Problem
The current rolling subtitle implementation generates 3 independent ASS Dialogue events per spoken line (Active → Context → Disappear), each with \move animations. When a new line arrives, both the new and previous lines scroll simultaneously, creating a jarring "2-line block jump" instead of a natural reading flow. The active line always resets to the bottom position, breaking the top-to-bottom reading direction users expect.
Design
Core Model: 2-Line Paged Blocks
Group spoken lines into pages of 2 lines each. Each page is a single ASS Dialogue event containing both lines separated by \N (ASS hard line break). Karaoke \k tags flow continuously from line 1 through line 2 within the same Dialogue event.
Reading flow: The viewer reads the top line (karaoke highlighting progresses left-to-right), then naturally drops to the bottom line (karaoke continues), then a cross-fade transitions to the next page.
Page Construction
Given a sequence of SpokenLine structs extracted from the VTT:
- Pair lines into consecutive groups of 2:
[line0, line1],[line2, line3], … - If the gap between two lines within a page exceeds 2 seconds, split them into separate pages instead of combining.
- If the total line count is odd, the final page contains a single line.
Karaoke Tag Stitching
For a 2-line page with lines A and B, the Dialogue text is:
{\k<dur>}word1_A {\k<dur>}word2_A ... {\k<dur>}lastword_A\N{\k<dur>}word1_B ... {\k<dur>}lastword_B
The \k duration for each word equals the time from that word's VTT start to the next word's VTT start. For the last word of line A, its \k duration extends to the start of line B's first word — this naturally covers any gap (silence/pause) between the two lines without any special gap-filler logic.
The \N is purely a visual line break and does not interrupt the karaoke timeline.
For lines without <c> word timing tags, the entire line text gets a single \k equal to the line's full spoken duration.
Cross-Fade Transitions
Pages transition via a 300ms cross-dissolve:
- Outgoing page: ASS end time extended by 300ms past its natural end. Uses
\fad(*, 300)for a 300ms fade-out. - Incoming page: ASS start time moved 300ms before its natural start. Uses
\fad(300, *)for a 300ms fade-in. - The 300ms overlap produces a smooth cross-dissolve between pages.
Specific \fad values:
| Page Position | \fad value |
|---|---|
| First page | \fad(0, 300) |
| Middle pages | \fad(300, 300) |
| Last page | \fad(300, 0) |
When pages are separated by a long gap (>2s), the outgoing page fades out and the incoming page fades in independently — no visual overlap, just a clean silence gap.
Positioning
Each Dialogue uses \an2\pos(cx, y_bottom) — bottom-center anchor. With \an2, the 2-line \N block renders with line 2 at the anchor point and line 1 stacked above it. No \move tags are used. Pages are static in position; transitions are opacity-only via \fad.
Layout constants (matching existing style):
cx = PlayResX / 2 = 960y_bottom = PlayResY - margin_v = 1080 - 40 = 1040
Edge Cases
| Case | Handling |
|---|---|
| Odd number of lines | Last page has 1 line (single-line Dialogue, no \N) |
| Gap >2s between consecutive lines within a pair | Split into separate single-line pages |
Non-speech cues ([Music], [Applause]) |
Standalone single-line Dialogue, no karaoke, just \pos |
Lines without <c> word timing |
Single \k tag covering the line's full spoken duration |
| Only 1 spoken line total | Single-line Dialogue with \fad(0, 0) |
ASS Header and Style
Unchanged from the current implementation:
PlayResX: 1920,PlayResY: 1080,ScaledBorderAndShadow: yes- Style parameters (font size, colours, border style, outline, margin) derived from user's
CaptionStylesettings Alignment: 2(bottom-center) in style definition, reinforced with\an2override in each Dialogue
What Gets Removed
The following components of the current implementation are replaced entirely:
RollingLayoutstruct (position calculation for\moveanimations)- Per-line 3-phase Dialogue generation (Active/Context/Disappear)
- All
\movetags - All
\cliptags - Phase-based timing calculations
What Gets Added
build_page_karaoke_text(lines: &[&SpokenLine]) -> String— stitches word timings from 1-2 lines into a single karaoke text with\Nseparator- Page grouping logic in
vtt_to_ass_with_karaoke()— pairs lines, handles gap-based splitting - Cross-fade timing logic — computes
\fadand adjusted start/end times per page
What Stays the Same
SpokenLinestruct andextract_spoken_lines()— line extraction from VTT is unchangedbuild_karaoke_text()— per-line word timing extraction (reused internally by the new page builder)parse_vtt_word_timings()— VTT<c>tag parser- ASS header/style generation
build_ffmpeg_args_burnin_subs()— FFmpeg argument construction (unchanged, still usesass=filter)- All other export logic (lossless, muxed subtitles, progress reporting)
Files Changed
| File | Change |
|---|---|
src-tauri/src/services/clip_exporter.rs |
Rewrite vtt_to_ass_with_karaoke() to use paged model. Remove RollingLayout. Add build_page_karaoke_text(). |
Single file change. No frontend, no new dependencies.