Includes: - Extended caption styling (font, shadow, dimmed color, bg toggle) - Media server, subtitle downloader, VTT parser, processing modal - Waveform tiers, thumbnail/timeline improvements, transport controls - Hybrid download model, dependency management, clip export enhancements - 21 chat summaries, 2 implementation plans, 2 design specs Co-authored-by: Cursor <cursoragent@cursor.com>
605 lines
22 KiB
Markdown
605 lines
22 KiB
Markdown
# Paged Teleprompter Burn-In Subtitles Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
|
||
**Goal:** Replace the rolling per-line ASS subtitle generation with a paged teleprompter model where 2-line blocks display with continuous karaoke highlighting and cross-fade transitions.
|
||
|
||
**Architecture:** Rewrite `vtt_to_ass_with_karaoke()` in `clip_exporter.rs` to group `SpokenLine`s into 2-line pages, emit one ASS Dialogue per page with stitched `\k` tags across `\N` line breaks, and use `\fad` for cross-fade transitions. No `\move` or `\clip` tags. Single-file change.
|
||
|
||
**Tech Stack:** Rust, ASS subtitle format (libass), FFmpeg `ass=` filter
|
||
|
||
## Global Constraints
|
||
|
||
- All changes are in `src-tauri/src/services/clip_exporter.rs` only
|
||
- No new dependencies
|
||
- Existing tests must continue to pass (they test FFmpeg arg construction, not ASS content)
|
||
- ASS header/style generation stays unchanged
|
||
- `extract_spoken_lines()` stays unchanged
|
||
- `parse_vtt_word_timings()` stays unchanged
|
||
- The `ass=` filter usage in `build_ffmpeg_args_burnin_subs()` is unchanged
|
||
- Cross-fade duration: 300ms
|
||
- Gap threshold for page splitting: 2.0 seconds
|
||
- PlayRes: 1920×1080
|
||
|
||
---
|
||
|
||
### Task 1: Add `build_page_karaoke_text` and page-grouping logic
|
||
|
||
**Files:**
|
||
- Modify: `src-tauri/src/services/clip_exporter.rs:432-460` (replace `RollingLayout`, add new function)
|
||
|
||
**Interfaces:**
|
||
- Consumes: `SpokenLine` struct (unchanged, line 349), `parse_vtt_word_timings()` (unchanged, line 259), `strip_vtt_tags()` (unchanged, line 170)
|
||
- Produces: `fn build_page_karaoke_text(lines: &[&SpokenLine]) -> String` — returns ASS text with `\k` tags and `\N` separator for 1-2 line pages. `fn group_into_pages(spoken_lines: &[SpokenLine], gap_threshold: f64) -> Vec<Vec<usize>>` — returns groups of indices into spoken_lines.
|
||
|
||
- [ ] **Step 1: Write failing tests for `build_page_karaoke_text`**
|
||
|
||
Add these tests to the existing `#[cfg(test)] mod tests` block at line 1155:
|
||
|
||
```rust
|
||
#[test]
|
||
fn test_build_page_karaoke_single_line_with_karaoke() {
|
||
let line = SpokenLine {
|
||
plain_text: "hello world".to_string(),
|
||
raw_text: "hello<00:00:01.000><c> world</c>".to_string(),
|
||
start_time: 0.5,
|
||
end_time: 1.5,
|
||
has_karaoke: true,
|
||
is_non_speech: false,
|
||
};
|
||
let result = build_page_karaoke_text(&[&line]);
|
||
// "hello" starts at 0.5, "world" starts at 1.0, ends at 1.5
|
||
// hello duration = 1.0 - 0.5 = 0.5s = 50cs
|
||
// world duration = 1.5 - 1.0 = 0.5s = 50cs
|
||
assert!(result.contains("{\\k50}hello"));
|
||
assert!(result.contains("{\\k50}world"));
|
||
assert!(!result.contains("\\N"));
|
||
}
|
||
|
||
#[test]
|
||
fn test_build_page_karaoke_two_lines_stitched() {
|
||
let line1 = SpokenLine {
|
||
plain_text: "hello world".to_string(),
|
||
raw_text: "hello<00:00:01.000><c> world</c>".to_string(),
|
||
start_time: 0.5,
|
||
end_time: 1.5,
|
||
has_karaoke: true,
|
||
is_non_speech: false,
|
||
};
|
||
let line2 = SpokenLine {
|
||
plain_text: "foo bar".to_string(),
|
||
raw_text: "foo<00:00:02.500><c> bar</c>".to_string(),
|
||
start_time: 2.0,
|
||
end_time: 3.0,
|
||
has_karaoke: true,
|
||
is_non_speech: false,
|
||
};
|
||
let result = build_page_karaoke_text(&[&line1, &line2]);
|
||
// Last word of line1 ("world") should span from 1.0 to line2 first word (2.0) = 100cs
|
||
assert!(result.contains("\\N"));
|
||
assert!(result.contains("{\\k100}world"));
|
||
// line2: "foo" at 2.0, "bar" at 2.5, end at 3.0
|
||
assert!(result.contains("{\\k50}foo"));
|
||
assert!(result.contains("{\\k50}bar"));
|
||
}
|
||
|
||
#[test]
|
||
fn test_build_page_karaoke_no_karaoke_tags() {
|
||
let line = SpokenLine {
|
||
plain_text: "just plain text".to_string(),
|
||
raw_text: "just plain text".to_string(),
|
||
start_time: 1.0,
|
||
end_time: 3.0,
|
||
has_karaoke: false,
|
||
is_non_speech: false,
|
||
};
|
||
let result = build_page_karaoke_text(&[&line]);
|
||
// Single \k covering the full 2.0s = 200cs
|
||
assert!(result.contains("{\\k200}just plain text"));
|
||
}
|
||
|
||
#[test]
|
||
fn test_group_into_pages_even() {
|
||
let lines = vec![
|
||
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 1.0, end_time: 2.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "c".into(), raw_text: "c".into(), start_time: 2.0, end_time: 3.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "d".into(), raw_text: "d".into(), start_time: 3.0, end_time: 4.0, has_karaoke: false, is_non_speech: false },
|
||
];
|
||
let pages = group_into_pages(&lines, 2.0);
|
||
assert_eq!(pages.len(), 2);
|
||
assert_eq!(pages[0], vec![0, 1]);
|
||
assert_eq!(pages[1], vec![2, 3]);
|
||
}
|
||
|
||
#[test]
|
||
fn test_group_into_pages_gap_splits() {
|
||
let lines = vec![
|
||
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 5.0, end_time: 6.0, has_karaoke: false, is_non_speech: false },
|
||
];
|
||
// Gap between a (end=1.0) and b (start=5.0) is 4.0s > 2.0s threshold
|
||
let pages = group_into_pages(&lines, 2.0);
|
||
assert_eq!(pages.len(), 2);
|
||
assert_eq!(pages[0], vec![0]);
|
||
assert_eq!(pages[1], vec![1]);
|
||
}
|
||
|
||
#[test]
|
||
fn test_group_into_pages_odd_count() {
|
||
let lines = vec![
|
||
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 1.0, end_time: 2.0, has_karaoke: false, is_non_speech: false },
|
||
SpokenLine { plain_text: "c".into(), raw_text: "c".into(), start_time: 2.0, end_time: 3.0, has_karaoke: false, is_non_speech: false },
|
||
];
|
||
let pages = group_into_pages(&lines, 2.0);
|
||
assert_eq!(pages.len(), 2);
|
||
assert_eq!(pages[0], vec![0, 1]);
|
||
assert_eq!(pages[1], vec![2]);
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | head -40`
|
||
Expected: Compilation errors — `build_page_karaoke_text` and `group_into_pages` not found.
|
||
|
||
- [ ] **Step 3: Implement `group_into_pages`**
|
||
|
||
Add this function after `build_karaoke_text` (around line 450), replacing the `RollingLayout` struct (lines 452-458):
|
||
|
||
```rust
|
||
/// Group spoken lines into 2-line pages for the teleprompter display.
|
||
/// If the gap between two consecutive lines exceeds `gap_threshold` seconds,
|
||
/// the pair is split into separate single-line pages.
|
||
/// Non-speech lines (e.g., [Music]) always get their own page.
|
||
fn group_into_pages(spoken_lines: &[SpokenLine], gap_threshold: f64) -> Vec<Vec<usize>> {
|
||
let mut pages: Vec<Vec<usize>> = Vec::new();
|
||
let mut i = 0;
|
||
|
||
while i < spoken_lines.len() {
|
||
let line = &spoken_lines[i];
|
||
|
||
// Non-speech cues always get their own page
|
||
if line.is_non_speech {
|
||
pages.push(vec![i]);
|
||
i += 1;
|
||
continue;
|
||
}
|
||
|
||
// Try to pair with the next line
|
||
if i + 1 < spoken_lines.len() {
|
||
let next = &spoken_lines[i + 1];
|
||
let gap = next.start_time - line.end_time;
|
||
|
||
// Pair them if gap is small enough and next isn't non-speech
|
||
if gap <= gap_threshold && !next.is_non_speech {
|
||
pages.push(vec![i, i + 1]);
|
||
i += 2;
|
||
continue;
|
||
}
|
||
}
|
||
|
||
// Solo page (last line, or gap too large, or next is non-speech)
|
||
pages.push(vec![i]);
|
||
i += 1;
|
||
}
|
||
|
||
pages
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 4: Implement `build_page_karaoke_text`**
|
||
|
||
Add this function right after `group_into_pages`:
|
||
|
||
```rust
|
||
/// Build karaoke text for a 1-or-2-line page, stitching word timings
|
||
/// across lines with \N as the visual line break. The \k durations flow
|
||
/// continuously so karaoke highlighting progresses top-to-bottom.
|
||
fn build_page_karaoke_text(lines: &[&SpokenLine]) -> String {
|
||
// Collect all (word, absolute_start_time) across all lines in page order.
|
||
// For the cross-line stitch, the last word of line N extends to the first
|
||
// word of line N+1.
|
||
struct WordEntry {
|
||
text: String,
|
||
start: f64,
|
||
is_line_break_before: bool, // insert \N before this word
|
||
}
|
||
|
||
let mut entries: Vec<WordEntry> = Vec::new();
|
||
|
||
for (line_idx, line) in lines.iter().enumerate() {
|
||
let is_new_line = line_idx > 0;
|
||
|
||
if line.has_karaoke {
|
||
let words = parse_vtt_word_timings(&line.raw_text, line.start_time);
|
||
if words.len() >= 2 {
|
||
for (w_idx, (word, start)) in words.iter().enumerate() {
|
||
entries.push(WordEntry {
|
||
text: word.clone(),
|
||
start: *start,
|
||
is_line_break_before: is_new_line && w_idx == 0,
|
||
});
|
||
}
|
||
} else {
|
||
// Fallback: treat entire line as one word
|
||
entries.push(WordEntry {
|
||
text: strip_vtt_tags(&line.raw_text),
|
||
start: line.start_time,
|
||
is_line_break_before: is_new_line,
|
||
});
|
||
}
|
||
} else {
|
||
// No karaoke data — one entry for the whole line
|
||
entries.push(WordEntry {
|
||
text: line.plain_text.clone(),
|
||
start: line.start_time,
|
||
is_line_break_before: is_new_line,
|
||
});
|
||
}
|
||
}
|
||
|
||
if entries.is_empty() {
|
||
return String::new();
|
||
}
|
||
|
||
// The page's end time is the last line's end_time
|
||
let page_end = lines.last().unwrap().end_time;
|
||
|
||
// Build the output with \k tags
|
||
let mut parts: Vec<String> = Vec::new();
|
||
for (i, entry) in entries.iter().enumerate() {
|
||
let next_start = if i + 1 < entries.len() {
|
||
entries[i + 1].start
|
||
} else {
|
||
page_end
|
||
};
|
||
let duration_cs = ((next_start - entry.start) * 100.0).round().max(1.0) as u64;
|
||
|
||
let prefix = if entry.is_line_break_before { "\\N" } else if i > 0 { " " } else { "" };
|
||
parts.push(format!("{}{{\\k{}}}{}", prefix, duration_cs, entry.text));
|
||
}
|
||
|
||
parts.join("")
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 5: Run tests to verify they pass**
|
||
|
||
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -20`
|
||
Expected: All new tests pass. All existing tests still pass.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add src-tauri/src/services/clip_exporter.rs
|
||
git commit -m "feat(subtitles): add page grouping and cross-line karaoke stitching"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 2: Rewrite `vtt_to_ass_with_karaoke` to use the paged model
|
||
|
||
**Files:**
|
||
- Modify: `src-tauri/src/services/clip_exporter.rs:465-643` (the `vtt_to_ass_with_karaoke` function body)
|
||
|
||
**Interfaces:**
|
||
- Consumes: `build_page_karaoke_text()` and `group_into_pages()` from Task 1, `extract_spoken_lines()` (unchanged, line 362), `format_ass_timestamp()` (unchanged, line 244), style helper functions (unchanged)
|
||
- Produces: Same public signature `pub fn vtt_to_ass_with_karaoke(caption_path, temp_dir, style) -> Result<String, String>` — writes ASS file to temp_dir and returns its path. Existing callers are unchanged.
|
||
|
||
- [ ] **Step 1: Write a test for the full ASS generation**
|
||
|
||
Add to the test module:
|
||
|
||
```rust
|
||
#[test]
|
||
fn test_vtt_to_ass_paged_output() {
|
||
let vtt_content = "\
|
||
WEBVTT
|
||
|
||
00:00:00.500 --> 00:00:01.500
|
||
hello<00:00:01.000><c> world</c>
|
||
|
||
00:00:02.000 --> 00:00:03.000
|
||
foo<00:00:02.500><c> bar</c>
|
||
|
||
00:00:03.000 --> 00:00:04.000
|
||
baz<00:00:03.500><c> qux</c>
|
||
|
||
00:00:04.000 --> 00:00:05.000
|
||
last<00:00:04.500><c> line</c>
|
||
";
|
||
let temp = std::env::temp_dir().join("test-paged-ass");
|
||
let _ = std::fs::remove_dir_all(&temp);
|
||
std::fs::create_dir_all(&temp).unwrap();
|
||
|
||
let vtt_path = temp.join("test.vtt");
|
||
std::fs::write(&vtt_path, vtt_content).unwrap();
|
||
|
||
let result = vtt_to_ass_with_karaoke(
|
||
vtt_path.to_str().unwrap(),
|
||
temp.to_str().unwrap(),
|
||
None,
|
||
);
|
||
assert!(result.is_ok());
|
||
|
||
let ass_path = result.unwrap();
|
||
let ass_content = std::fs::read_to_string(&ass_path).unwrap();
|
||
|
||
// Should have ASS header
|
||
assert!(ass_content.contains("[Script Info]"));
|
||
assert!(ass_content.contains("PlayResX: 1920"));
|
||
assert!(ass_content.contains("[Events]"));
|
||
|
||
// Should use \an2\pos (static position), NOT \move
|
||
assert!(ass_content.contains("\\an2\\pos("));
|
||
assert!(!ass_content.contains("\\move("));
|
||
|
||
// Should have \N line breaks (paged 2-line blocks)
|
||
assert!(ass_content.contains("\\N"));
|
||
|
||
// Should have \fad for cross-fade transitions
|
||
assert!(ass_content.contains("\\fad("));
|
||
|
||
// Should have karaoke \k tags
|
||
assert!(ass_content.contains("\\k"));
|
||
|
||
// 4 lines → 2 pages → 2 Dialogue events (plus possible non-speech)
|
||
let dialogue_count = ass_content.matches("Dialogue:").count();
|
||
assert_eq!(dialogue_count, 2, "Expected 2 pages (4 lines / 2). Got {dialogue_count}.\nASS:\n{ass_content}");
|
||
|
||
let _ = std::fs::remove_dir_all(&temp);
|
||
}
|
||
|
||
#[test]
|
||
fn test_vtt_to_ass_paged_gap_splits_page() {
|
||
let vtt_content = "\
|
||
WEBVTT
|
||
|
||
00:00:00.500 --> 00:00:01.500
|
||
hello<00:00:01.000><c> world</c>
|
||
|
||
00:00:05.000 --> 00:00:06.000
|
||
far<00:00:05.500><c> away</c>
|
||
";
|
||
let temp = std::env::temp_dir().join("test-paged-gap-ass");
|
||
let _ = std::fs::remove_dir_all(&temp);
|
||
std::fs::create_dir_all(&temp).unwrap();
|
||
|
||
let vtt_path = temp.join("test.vtt");
|
||
std::fs::write(&vtt_path, vtt_content).unwrap();
|
||
|
||
let result = vtt_to_ass_with_karaoke(
|
||
vtt_path.to_str().unwrap(),
|
||
temp.to_str().unwrap(),
|
||
None,
|
||
);
|
||
assert!(result.is_ok());
|
||
|
||
let ass_content = std::fs::read_to_string(result.unwrap()).unwrap();
|
||
|
||
// Gap between lines is 3.5s > 2.0s threshold → should split into 2 single-line pages
|
||
let dialogue_count = ass_content.matches("Dialogue:").count();
|
||
assert_eq!(dialogue_count, 2, "Gap >2s should split into separate pages. Got {dialogue_count}.\nASS:\n{ass_content}");
|
||
|
||
// Single-line pages should NOT have \N
|
||
// (Each Dialogue has its own text without \N)
|
||
for line in ass_content.lines() {
|
||
if line.starts_with("Dialogue:") {
|
||
assert!(!line.contains("\\N"), "Single-line page should not contain \\N: {line}");
|
||
}
|
||
}
|
||
|
||
let _ = std::fs::remove_dir_all(&temp);
|
||
}
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify the new ones fail**
|
||
|
||
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests::test_vtt_to_ass_paged -- --nocapture 2>&1`
|
||
Expected: Failures — the current implementation uses `\move` and generates more Dialogue events per line.
|
||
|
||
- [ ] **Step 3: Rewrite `vtt_to_ass_with_karaoke`**
|
||
|
||
Replace the entire function body (lines 465-643) with the paged implementation. The function signature stays the same:
|
||
|
||
```rust
|
||
/// Convert a word-timed VTT file to an ASS file with paged teleprompter display.
|
||
/// Lines are grouped into 2-line pages. Karaoke \k tags flow continuously from
|
||
/// line 1 through line 2 within each page. Pages cross-fade with \fad transitions.
|
||
pub fn vtt_to_ass_with_karaoke(
|
||
caption_path: &str,
|
||
temp_dir: &str,
|
||
style: Option<&CaptionStyle>,
|
||
) -> Result<String, String> {
|
||
let content = std::fs::read_to_string(caption_path)
|
||
.map_err(|e| format!("Failed to read VTT file '{}': {e}", caption_path))?;
|
||
|
||
std::fs::create_dir_all(temp_dir)
|
||
.map_err(|e| format!("Failed to create temp dir for ASS: {e}"))?;
|
||
|
||
let out_path = Path::new(temp_dir).join("karaoke.ass");
|
||
|
||
// Build style values from CaptionStyle or use defaults
|
||
let font_size = style.map_or(45, |s| (s.font_size as f64 * 2.5).round() as u32);
|
||
let primary_colour = style.map_or("&H00FFFFFF".to_string(), |s| hex_to_ass_color(&s.text_color));
|
||
let secondary_colour = "&H73CCCCCC".to_string();
|
||
let outline_colour = "&H00000000".to_string();
|
||
let back_colour = style.map_or("&H80000000".to_string(), |s| opacity_to_ass_back_colour(s.background_opacity));
|
||
let text_outline = style.map_or(true, |s| s.text_outline);
|
||
let (border_style, outline_val, shadow_val) = if text_outline {
|
||
(4, 2, 0)
|
||
} else {
|
||
(3, 0, 0)
|
||
};
|
||
let margin_v: i32 = 40;
|
||
|
||
let mut output = String::new();
|
||
|
||
// ASS Script Info header
|
||
output.push_str("[Script Info]\n");
|
||
output.push_str("ScriptType: v4.00+\n");
|
||
output.push_str("PlayResX: 1920\n");
|
||
output.push_str("PlayResY: 1080\n");
|
||
output.push_str("WrapStyle: 0\n");
|
||
output.push_str("ScaledBorderAndShadow: yes\n");
|
||
output.push('\n');
|
||
|
||
// V4+ Styles
|
||
output.push_str("[V4+ Styles]\n");
|
||
output.push_str("Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding\n");
|
||
output.push_str(&format!(
|
||
"Style: Default,Arial,{},{},{},{},{},0,0,0,0,100,100,0,0,{},{},{},2,20,20,{},1\n",
|
||
font_size, primary_colour, secondary_colour, outline_colour, back_colour,
|
||
border_style, outline_val, shadow_val, margin_v
|
||
));
|
||
output.push('\n');
|
||
|
||
// Events
|
||
output.push_str("[Events]\n");
|
||
output.push_str("Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text\n");
|
||
|
||
let spoken_lines = extract_spoken_lines(&content);
|
||
let gap_threshold = 2.0;
|
||
let crossfade_ms = 300;
|
||
let crossfade_s = crossfade_ms as f64 / 1000.0;
|
||
|
||
let pages = group_into_pages(&spoken_lines, gap_threshold);
|
||
let page_count = pages.len();
|
||
|
||
let cx = 960;
|
||
let y_bottom = 1080 - margin_v;
|
||
|
||
for (page_idx, page_indices) in pages.iter().enumerate() {
|
||
let page_lines: Vec<&SpokenLine> = page_indices.iter().map(|&i| &spoken_lines[i]).collect();
|
||
let first_line = page_lines[0];
|
||
let last_line = *page_lines.last().unwrap();
|
||
|
||
// Non-speech cue: simple static display, no karaoke, no cross-fade
|
||
if page_lines.len() == 1 && first_line.is_non_speech {
|
||
output.push_str(&format!(
|
||
"Dialogue: 0,{},{},Default,,0,0,0,,{{\\an2\\pos({},{})}}{}\n",
|
||
format_ass_timestamp(first_line.start_time),
|
||
format_ass_timestamp(first_line.end_time),
|
||
cx, y_bottom,
|
||
first_line.plain_text
|
||
));
|
||
continue;
|
||
}
|
||
|
||
// Compute display timing with cross-fade overlap
|
||
let natural_start = first_line.start_time;
|
||
let natural_end = last_line.end_time;
|
||
|
||
let is_first = page_idx == 0;
|
||
let is_last = page_idx == page_count - 1;
|
||
|
||
let display_start = if is_first {
|
||
natural_start
|
||
} else {
|
||
(natural_start - crossfade_s).max(0.0)
|
||
};
|
||
let display_end = if is_last {
|
||
natural_end
|
||
} else {
|
||
natural_end + crossfade_s
|
||
};
|
||
|
||
let fade_in = if is_first { 0 } else { crossfade_ms };
|
||
let fade_out = if is_last { 0 } else { crossfade_ms };
|
||
|
||
let karaoke_text = build_page_karaoke_text(&page_lines);
|
||
|
||
output.push_str(&format!(
|
||
"Dialogue: 0,{},{},Default,,0,0,0,,{{\\an2\\pos({},{})\\fad({},{})}}{}\n",
|
||
format_ass_timestamp(display_start),
|
||
format_ass_timestamp(display_end),
|
||
cx, y_bottom,
|
||
fade_in, fade_out,
|
||
karaoke_text
|
||
));
|
||
}
|
||
|
||
std::fs::write(&out_path, &output)
|
||
.map_err(|e| format!("Failed to write ASS file: {e}"))?;
|
||
|
||
Ok(out_path.to_string_lossy().to_string())
|
||
}
|
||
```
|
||
|
||
Also remove the now-unused `RollingLayout` struct (around line 452-458). The old `build_karaoke_text` function (line 432-449) can remain — it's not called by the new code but doesn't hurt and could be useful for future single-line scenarios.
|
||
|
||
- [ ] **Step 4: Run all tests**
|
||
|
||
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -30`
|
||
Expected: All tests pass — both new paged tests and all existing tests.
|
||
|
||
- [ ] **Step 5: Build the full project**
|
||
|
||
Run: `cd src-tauri && cargo build 2>&1 | tail -10`
|
||
Expected: Clean build, no warnings about unused code (other than pre-existing ones).
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
git add src-tauri/src/services/clip_exporter.rs
|
||
git commit -m "feat(subtitles): rewrite ASS generation to paged teleprompter model
|
||
|
||
Replace per-line 3-phase rolling animation with 2-line paged blocks.
|
||
Karaoke \\k tags flow continuously across \\N line breaks within each page.
|
||
Pages cross-fade with \\fad transitions (300ms). No \\move or \\clip tags.
|
||
Reading flow is top-to-bottom within each page, matching teleprompter style."
|
||
```
|
||
|
||
---
|
||
|
||
### Task 3: Manual smoke test and cleanup
|
||
|
||
**Files:**
|
||
- Modify: `src-tauri/src/services/clip_exporter.rs` (only if issues found)
|
||
|
||
**Interfaces:**
|
||
- Consumes: The full paged ASS generation pipeline from Tasks 1-2
|
||
- Produces: Verified working burn-in subtitle export
|
||
|
||
- [ ] **Step 1: Run the app and test export**
|
||
|
||
Run: `cargo tauri dev`
|
||
|
||
Test flow:
|
||
1. Paste a YouTube URL with captions (e.g., `https://youtu.be/NruccMk0Jls`)
|
||
2. Wait for processing to complete
|
||
3. Mark a short clip (~10-15 seconds) containing speech
|
||
4. Open export dialog, enable captions with burn-in
|
||
5. Export the clip
|
||
6. Open the exported file and verify:
|
||
- Subtitles show as 2-line blocks
|
||
- Karaoke highlighting progresses top line → bottom line
|
||
- Page transitions cross-fade smoothly (no hard cuts or jumps)
|
||
- No position jumping or "2-line block scroll"
|
||
|
||
- [ ] **Step 2: Inspect the generated ASS file**
|
||
|
||
Check the temp ASS file to verify structure:
|
||
|
||
Run: `cat /tmp/video-clipper*/karaoke.ass | head -40`
|
||
|
||
Verify:
|
||
- Each `Dialogue:` line contains `\an2\pos(` (static position)
|
||
- No `\move(` tags anywhere
|
||
- Multi-line pages contain `\N` separator
|
||
- `\fad(` tags present on each Dialogue
|
||
- `\k` tags flow through both lines of each page
|
||
|
||
- [ ] **Step 3: Remove dead code if any**
|
||
|
||
If `RollingLayout` struct is still present, remove it. If `build_karaoke_text` is unused and triggers a warning, add `#[allow(dead_code)]` or remove it. Run `cargo build` to confirm clean.
|
||
|
||
- [ ] **Step 4: Final commit if cleanup was needed**
|
||
|
||
```bash
|
||
git add src-tauri/src/services/clip_exporter.rs
|
||
git commit -m "chore: remove dead rolling layout code"
|
||
```
|