Files
gui-video-clipper/docs/superpowers/specs/2026-09-22-paged-teleprompter-subtitles-design.md
cottongin 8ad2f1c800 chore: stage all pending work — caption styling, media server, processing modal, docs, summaries
Includes:
- Extended caption styling (font, shadow, dimmed color, bg toggle)
- Media server, subtitle downloader, VTT parser, processing modal
- Waveform tiers, thumbnail/timeline improvements, transport controls
- Hybrid download model, dependency management, clip export enhancements
- 21 chat summaries, 2 implementation plans, 2 design specs

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-22 10:48:16 -04:00

5.5 KiB

Paged Teleprompter Burn-In Subtitle Design

Date: 2026-09-22
Scope: Rewrite the ASS subtitle generation in clip_exporter.rs to use a paged teleprompter model instead of the current per-line rolling model.

Problem

The current rolling subtitle implementation generates 3 independent ASS Dialogue events per spoken line (Active → Context → Disappear), each with \move animations. When a new line arrives, both the new and previous lines scroll simultaneously, creating a jarring "2-line block jump" instead of a natural reading flow. The active line always resets to the bottom position, breaking the top-to-bottom reading direction users expect.

Design

Core Model: 2-Line Paged Blocks

Group spoken lines into pages of 2 lines each. Each page is a single ASS Dialogue event containing both lines separated by \N (ASS hard line break). Karaoke \k tags flow continuously from line 1 through line 2 within the same Dialogue event.

Reading flow: The viewer reads the top line (karaoke highlighting progresses left-to-right), then naturally drops to the bottom line (karaoke continues), then a cross-fade transitions to the next page.

Page Construction

Given a sequence of SpokenLine structs extracted from the VTT:

  1. Pair lines into consecutive groups of 2: [line0, line1], [line2, line3], …
  2. If the gap between two lines within a page exceeds 2 seconds, split them into separate pages instead of combining.
  3. If the total line count is odd, the final page contains a single line.

Karaoke Tag Stitching

For a 2-line page with lines A and B, the Dialogue text is:

{\k<dur>}word1_A {\k<dur>}word2_A ... {\k<dur>}lastword_A\N{\k<dur>}word1_B ... {\k<dur>}lastword_B

The \k duration for each word equals the time from that word's VTT start to the next word's VTT start. For the last word of line A, its \k duration extends to the start of line B's first word — this naturally covers any gap (silence/pause) between the two lines without any special gap-filler logic.

The \N is purely a visual line break and does not interrupt the karaoke timeline.

For lines without <c> word timing tags, the entire line text gets a single \k equal to the line's full spoken duration.

Cross-Fade Transitions

Pages transition via a 300ms cross-dissolve:

  • Outgoing page: ASS end time extended by 300ms past its natural end. Uses \fad(*, 300) for a 300ms fade-out.
  • Incoming page: ASS start time moved 300ms before its natural start. Uses \fad(300, *) for a 300ms fade-in.
  • The 300ms overlap produces a smooth cross-dissolve between pages.

Specific \fad values:

Page Position \fad value
First page \fad(0, 300)
Middle pages \fad(300, 300)
Last page \fad(300, 0)

When pages are separated by a long gap (>2s), the outgoing page fades out and the incoming page fades in independently — no visual overlap, just a clean silence gap.

Positioning

Each Dialogue uses \an2\pos(cx, y_bottom) — bottom-center anchor. With \an2, the 2-line \N block renders with line 2 at the anchor point and line 1 stacked above it. No \move tags are used. Pages are static in position; transitions are opacity-only via \fad.

Layout constants (matching existing style):

  • cx = PlayResX / 2 = 960
  • y_bottom = PlayResY - margin_v = 1080 - 40 = 1040

Edge Cases

Case Handling
Odd number of lines Last page has 1 line (single-line Dialogue, no \N)
Gap >2s between consecutive lines within a pair Split into separate single-line pages
Non-speech cues ([Music], [Applause]) Standalone single-line Dialogue, no karaoke, just \pos
Lines without <c> word timing Single \k tag covering the line's full spoken duration
Only 1 spoken line total Single-line Dialogue with \fad(0, 0)

ASS Header and Style

Unchanged from the current implementation:

  • PlayResX: 1920, PlayResY: 1080, ScaledBorderAndShadow: yes
  • Style parameters (font size, colours, border style, outline, margin) derived from user's CaptionStyle settings
  • Alignment: 2 (bottom-center) in style definition, reinforced with \an2 override in each Dialogue

What Gets Removed

The following components of the current implementation are replaced entirely:

  • RollingLayout struct (position calculation for \move animations)
  • Per-line 3-phase Dialogue generation (Active/Context/Disappear)
  • All \move tags
  • All \clip tags
  • Phase-based timing calculations

What Gets Added

  • build_page_karaoke_text(lines: &[&SpokenLine]) -> String — stitches word timings from 1-2 lines into a single karaoke text with \N separator
  • Page grouping logic in vtt_to_ass_with_karaoke() — pairs lines, handles gap-based splitting
  • Cross-fade timing logic — computes \fad and adjusted start/end times per page

What Stays the Same

  • SpokenLine struct and extract_spoken_lines() — line extraction from VTT is unchanged
  • build_karaoke_text() — per-line word timing extraction (reused internally by the new page builder)
  • parse_vtt_word_timings() — VTT <c> tag parser
  • ASS header/style generation
  • build_ffmpeg_args_burnin_subs() — FFmpeg argument construction (unchanged, still uses ass= filter)
  • All other export logic (lossless, muxed subtitles, progress reporting)

Files Changed

File Change
src-tauri/src/services/clip_exporter.rs Rewrite vtt_to_ass_with_karaoke() to use paged model. Remove RollingLayout. Add build_page_karaoke_text().

Single file change. No frontend, no new dependencies.