Includes: - Extended caption styling (font, shadow, dimmed color, bg toggle) - Media server, subtitle downloader, VTT parser, processing modal - Waveform tiers, thumbnail/timeline improvements, transport controls - Hybrid download model, dependency management, clip export enhancements - 21 chat summaries, 2 implementation plans, 2 design specs Co-authored-by: Cursor <cursoragent@cursor.com>
3.3 KiB
Fix Karaoke Burn-In Sync and Export Progress
Date: 2026-09-21 23:22
Task: Fix two bugs: karaoke ASS burn-in out of sync / skipping lines, and export progress stuck at 0% then jumping to 100%.
Changes Made
Bug 1: Karaoke ASS Burn-In — Two Root Causes Fixed
A. YouTube VTT two-line overlay pattern.
YouTube auto-generated VTT cues have a specific structure:
- Many cues have TWO text lines: line 1 is static context (previous cue text, no
<c>tags), line 2 is the karaoke line (with<timestamp><c> word</c>tags) - Zero-duration transition cues (e.g.,
00:00:02.629 --> 00:00:02.639) are just visual transition frames
The old code joined all text lines and tried to parse word timings from the combined mess. Fix:
- Skip cues where
(end - start) < 0.05s— these are transition frames - For multi-line cues, only process the line that contains
<c>tags for karaoke - Ignore the static context line (previous cue's text repeated)
- Cues with no
<c>tags emit as plain text (graceful fallback)
B. Wrong ffmpeg filter.
Per ffmpeg-micro.com: "The subtitles filter routes your file through libavformat's subtitle converter and applies force_style overrides, which can flatten karaoke timing. The ass filter hands the file straight to libass."
Fix: Changed build_ffmpeg_args_burnin_subs() to use -vf ass=<path> instead of -vf subtitles=<path>. Since style is already baked into the ASS [V4+ Styles] header, no force_style is needed. The _caption_style parameter is now ignored (prefixed with _).
Bug 2: Export Progress — Root Cause Fixed
ffmpeg writes its progress output to stderr using \r (carriage return) to overwrite the same line, NOT \n (newline). BufReader::lines() splits on \n only, so all progress was buffered as one giant "line" that only yielded when ffmpeg exited.
Fix: Use ffmpeg's -progress pipe:1 -nostats flag:
- Adds
-progress pipe:1 -nostatsto ffmpeg args (before the output file) - Reads stdout (not stderr) —
-progressoutputs proper\n-delimitedkey=valuepairs - Parses
out_time_ms=<microseconds>lines and computespercent = out_time_ms / (duration * 1_000_000) - stderr is still piped but drained on a separate thread for error reporting on failure
Example -progress pipe:1 output (proper newlines):
frame=150
fps=45.2
out_time_ms=6250000
progress=continue
Files Modified
| File | Change |
|---|---|
src-tauri/src/services/clip_exporter.rs |
Fixed VTT parsing, switched to ass= filter, rewrote progress to use -progress pipe:1 |
Build Verification
cargo test— all 43 tests passcargo check— passes (only pre-existing warnings)svelte-check— passes (only pre-existingvite.config.tserrors)
Lessons Learned
- YouTube VTT is NOT simple cue-per-line. It uses a two-line overlay pattern where line 1 is a static context repeat and line 2 has the karaoke data. Zero-duration cues are transition frames.
- The
assfilter andsubtitlesfilter in ffmpeg both use libass, butsubtitlesapplies format conversion andforce_stylethat can destroy\kfkaraoke timing. Always useassfor karaoke. - ffmpeg's stderr progress uses
\rnot\n. For Rust'sBufReader::lines()to work, use-progress pipe:1 -nostatswhich outputs proper\n-delimited key=value pairs to stdout.