Files
gui-video-clipper/chat-summaries/2026-09-21_23-22-fix-karaoke-and-progress-summary.md
cottongin 8ad2f1c800 chore: stage all pending work — caption styling, media server, processing modal, docs, summaries
Includes:
- Extended caption styling (font, shadow, dimmed color, bg toggle)
- Media server, subtitle downloader, VTT parser, processing modal
- Waveform tiers, thumbnail/timeline improvements, transport controls
- Hybrid download model, dependency management, clip export enhancements
- 21 chat summaries, 2 implementation plans, 2 design specs

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-22 10:48:16 -04:00

3.3 KiB

Fix Karaoke Burn-In Sync and Export Progress

Date: 2026-09-21 23:22
Task: Fix two bugs: karaoke ASS burn-in out of sync / skipping lines, and export progress stuck at 0% then jumping to 100%.

Changes Made

Bug 1: Karaoke ASS Burn-In — Two Root Causes Fixed

A. YouTube VTT two-line overlay pattern.

YouTube auto-generated VTT cues have a specific structure:

  • Many cues have TWO text lines: line 1 is static context (previous cue text, no <c> tags), line 2 is the karaoke line (with <timestamp><c> word</c> tags)
  • Zero-duration transition cues (e.g., 00:00:02.629 --> 00:00:02.639) are just visual transition frames

The old code joined all text lines and tried to parse word timings from the combined mess. Fix:

  • Skip cues where (end - start) < 0.05s — these are transition frames
  • For multi-line cues, only process the line that contains <c> tags for karaoke
  • Ignore the static context line (previous cue's text repeated)
  • Cues with no <c> tags emit as plain text (graceful fallback)

B. Wrong ffmpeg filter.

Per ffmpeg-micro.com: "The subtitles filter routes your file through libavformat's subtitle converter and applies force_style overrides, which can flatten karaoke timing. The ass filter hands the file straight to libass."

Fix: Changed build_ffmpeg_args_burnin_subs() to use -vf ass=<path> instead of -vf subtitles=<path>. Since style is already baked into the ASS [V4+ Styles] header, no force_style is needed. The _caption_style parameter is now ignored (prefixed with _).

Bug 2: Export Progress — Root Cause Fixed

ffmpeg writes its progress output to stderr using \r (carriage return) to overwrite the same line, NOT \n (newline). BufReader::lines() splits on \n only, so all progress was buffered as one giant "line" that only yielded when ffmpeg exited.

Fix: Use ffmpeg's -progress pipe:1 -nostats flag:

  • Adds -progress pipe:1 -nostats to ffmpeg args (before the output file)
  • Reads stdout (not stderr) — -progress outputs proper \n-delimited key=value pairs
  • Parses out_time_ms=<microseconds> lines and computes percent = out_time_ms / (duration * 1_000_000)
  • stderr is still piped but drained on a separate thread for error reporting on failure

Example -progress pipe:1 output (proper newlines):

frame=150
fps=45.2
out_time_ms=6250000
progress=continue

Files Modified

File Change
src-tauri/src/services/clip_exporter.rs Fixed VTT parsing, switched to ass= filter, rewrote progress to use -progress pipe:1

Build Verification

  • cargo test — all 43 tests pass
  • cargo check — passes (only pre-existing warnings)
  • svelte-check — passes (only pre-existing vite.config.ts errors)

Lessons Learned

  • YouTube VTT is NOT simple cue-per-line. It uses a two-line overlay pattern where line 1 is a static context repeat and line 2 has the karaoke data. Zero-duration cues are transition frames.
  • The ass filter and subtitles filter in ffmpeg both use libass, but subtitles applies format conversion and force_style that can destroy \kf karaoke timing. Always use ass for karaoke.
  • ffmpeg's stderr progress uses \r not \n. For Rust's BufReader::lines() to work, use -progress pipe:1 -nostats which outputs proper \n-delimited key=value pairs to stdout.