chore: stage all pending work — caption styling, media server, processing modal, docs, summaries

Includes:
- Extended caption styling (font, shadow, dimmed color, bg toggle)
- Media server, subtitle downloader, VTT parser, processing modal
- Waveform tiers, thumbnail/timeline improvements, transport controls
- Hybrid download model, dependency management, clip export enhancements
- 21 chat summaries, 2 implementation plans, 2 design specs

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
2026-09-22 10:48:16 -04:00
parent fd05ae19b5
commit 8ad2f1c800
57 changed files with 4929 additions and 210 deletions

View File

@@ -0,0 +1,48 @@
# Fix: Video Playback, Waveform, and Thumbnails Not Working
**Date:** 2026-09-21 12:21
**Task:** Diagnose and fix three user-reported issues after downloading a video: no playback, waveform/thumbnails stuck at "Generating...", scrubbing doesn't update the player frame.
## Root Causes
### 1. `[Merger]` line path not parsed (Primary Bug)
When `yt-dlp` downloads video+audio separately and merges them, it outputs:
```
[download] Destination: /path/to/file.f399.mp4 ← intermediate
[download] Destination: /path/to/file.f251.webm ← intermediate
[Merger] Merging formats into "/path/to/file.webm" ← final merged file
```
The old parsing used `line.split(": ").nth(1)` for all three cases. This works for `[download] Destination:` lines (which have `: ` delimiter) but **fails for `[Merger]`** lines (which use `into "path"` format — no `: `). As a result, `last_file_path` pointed to a deleted intermediate file. Both `triggerPostDownloadProcessing(path)` and `session.localFilePath` received the wrong path, causing:
- All ffmpeg post-processing (waveform, thumbnails, keyframes) to fail with file-not-found
- The video player to reference a non-existent file
### 2. Stream URL doesn't work in Tauri webview
The `yt-dlp -g` stream URL (googlevideo.com) has anti-hotlinking protections (referer/IP checks) that block playback in a Tauri/WKWebView context. The `<video>` element was present but silently failed to load the source.
### 3. "Already downloaded" case not handled
When `yt-dlp` finds an existing file, it outputs `[download] /path/file.webm has already been downloaded` — a format that matches neither `Destination:` nor `Merger` patterns. So re-loading the same URL would result in `localFilePath = ""`.
## Changes Made
### `src-tauri/src/commands/video.rs`
- **Fixed `[Merger]` parsing:** Replaced the broken `split(": ")` approach with `trim_start_matches("[Merger] Merging formats into ").trim().trim_matches('"')`.
- **Added "already downloaded" parsing:** New branch for `has already been downloaded` using `strip_prefix`/`strip_suffix` to extract the path.
### `src/lib/components/VideoPlayer.svelte`
- **Switched to local file for playback:** Uses `convertFileSrc(session.localFilePath)` via Tauri's asset protocol once download is complete, instead of the unreliable stream URL.
- **Added download progress UI:** Shows a progress bar with percentage while downloading, instead of a broken black video area.
- **Added video error handling:** `onerror`/`onloadeddata` handlers with an overlay error message.
- **Import:** Added `convertFileSrc` from `@tauri-apps/api/core`.
## Lessons Learned
1. **`yt-dlp` stderr output formats are inconsistent.** `[download] Destination:` uses `: ` delimiter, `[Merger]` uses `into "path"`, and "already downloaded" embeds the path mid-sentence. Each needs its own parser.
2. **googlevideo.com stream URLs don't work in embedded webviews** due to anti-hotlinking. The design's "hybrid streaming preview" approach needs a local proxy or must fall back to showing download progress until the file is available locally.
3. **Always test with cached files.** The "already downloaded" edge case only surfaces on the second load of the same URL.
## Follow-up Items
- [ ] Add progress indicators for post-download processing (waveform/thumbnail/keyframe extraction can take minutes on long videos)
- [ ] Consider pre-download preview via local proxy or lower-quality quick download
- [ ] The waveform extraction uses `aresample=8000` (sample rate) which produces ~26M samples for a 54-min video — consider optimizing to reduce to target count directly
- [ ] Add error surfacing for failed post-download processing (currently silent `console.error`)

View File

@@ -0,0 +1,54 @@
# Fix: Video Codec Compatibility + Timeline Pan/Scroll + Minimap
**Date:** 2026-09-21 12:56
**Task:** Fix video playback (black screen despite play state), add timeline panning when zoomed, add minimap overview.
## Root Causes & Fixes
### 1. Video playback: WKWebView doesn't support VP9/WebM
The downloaded file was VP9+Opus in WebM container (yt-dlp format 399+251). Apple's WKWebView on macOS does NOT support VP9 codec. The `<video>` element accepted `.play()` but had no decodable frames.
**Fix:** Added `-f "bv*[vcodec^=avc1]+ba[acodec^=mp4a]/bv*[ext=mp4]+ba[ext=m4a]/b[ext=mp4]/b"` to the yt-dlp download command in `download_manager.rs`. This selects H.264 (avc1) video + AAC (mp4a) audio, producing an MP4 file playable by WKWebView. Deleted the cached WebM file.
### 2. Timeline: no scroll/pan when zoomed
Mouse wheel only triggered zoom. No way to scroll horizontally when zoomed in.
**Fix in `Timeline.svelte`:**
- **Horizontal scroll → pan:** `deltaX` from trackpad/shift+wheel now calls `panBy()` to shift the visible window
- **Vertical scroll → zoom** (unchanged)
**New `panBy()` in `interactions.ts`:** Computes new `visibleStart`/`visibleEnd` from pixel delta, clamped to `[0, duration]`.
### 3. Minimap overview bar
Added a minimap canvas that appears when zoomed in:
- Shows the full waveform at a glance
- Highlights the current viewport with a blue border
- Dims regions outside the viewport
- Click to center viewport, drag to pan
- Playhead indicator
### 4. Other fixes applied this session (cumulative)
- **stdout/stderr swap:** yt-dlp sends status messages to stdout, not stderr
- **`[Merger]` path parsing:** Correctly captures final merged file path
- **"Already downloaded" parsing:** Handles `has already been downloaded` message
- **Cache check:** `checkCachedDownload(title)` checks temp dir before invoking yt-dlp
- **Parallel processing:** Waveform/thumbnails/keyframes results applied individually as they complete
- **Asset protocol scope:** Added `/private/var/**` and `/var/**` for macOS symlink paths
## Files Changed
- `src-tauri/src/services/download_manager.rs` — H.264 format selection
- `src/lib/components/Timeline.svelte` — Minimap + horizontal pan support
- `src/lib/timeline/interactions.ts` — Added `panBy()` function
- `src-tauri/src/commands/video.rs` — Cache check, diagnostic logging
- `src-tauri/src/commands/media_analysis.rs` — Diagnostic logging
- `src-tauri/src/lib.rs` — Registered `check_cached_download` command
- `src-tauri/tauri.conf.json` — Expanded asset protocol scope
- `src/lib/stores/videoSession.svelte.ts` — Cache check before download, parallel results
- `src/lib/bindings/video.ts` — `checkCachedDownload` binding
- `src/lib/components/VideoPlayer.svelte` — Local file playback with progress UI
## Lessons Learned
1. **WKWebView codec support is limited.** VP9/AV1/WebM don't work. Must use H.264+AAC in MP4.
2. **macOS `/var` is a symlink to `/private/var`.** Asset protocol scope must cover both paths.
3. **`Promise.allSettled` blocks all results until the slowest completes.** Use individual `.then()` chains when you want progressive rendering.
4. **yt-dlp sends status to stdout, warnings to stderr.** Not the other way around.

View File

@@ -0,0 +1,60 @@
# Hybrid Download Model + Auto-Defocus URL Input
## Task Description
Two user-requested improvements:
1. **Hybrid download model**: Download a low-resolution (≤360p H.264+AAC) version for immediate preview/scrubbing/waveform/thumbnail generation, while downloading the best available quality in the background for export. Clips are cut from the best-quality file.
2. **Auto-defocus URL input**: After pasting a URL and triggering submit, blur the input field so keyboard shortcuts (I, O, etc.) don't accidentally modify the URL.
## Changes Made
### Session State Refactor (`src/lib/stores/videoSession.svelte.ts`)
- Replaced single `localFilePath`/`downloadProgress`/`downloadStatus` with dual-track state:
- `previewFilePath`/`previewProgress`/`previewStatus` — for the 360p preview
- `exportFilePath`/`exportProgress`/`exportStatus` — for the best-quality export
- Preview format: `bv*[vcodec^=avc1][height<=360]+ba[acodec^=mp4a]/b[ext=mp4][height<=360]/worst[ext=mp4]/worst`
- Export format: `bv*+ba/b` (best available, any codec since ffmpeg handles export)
- `beginDownload()` now:
1. Checks cache for preview → downloads if needed → triggers post-processing (waveform, thumbnails, keyframes)
2. Checks cache for export → downloads in background (fire-and-forget)
- Files stored in separate subdirectories: `video-clipper/preview/` and `video-clipper/export/`
### Rust Backend (`src-tauri/src/commands/video.rs`)
- `start_download` now accepts `format_spec` and `variant` parameters
- `check_cached_download` now accepts a `variant` parameter to check the correct subdirectory
- Helper `variant_dir()` builds `$TEMP/video-clipper/{variant}/` paths
- Diagnostic logging includes variant name for easier debugging
### Download Manager (`src-tauri/src/services/download_manager.rs`)
- `start_download` now accepts `format_spec: &str` parameter instead of hardcoding format selection
- Format string passed through from the frontend call
### Frontend Bindings (`src/lib/bindings/video.ts`)
- `startDownload` now takes `formatSpec` and `variant` parameters
- `checkCachedDownload` now takes a `variant` parameter
### VideoPlayer (`src/lib/components/VideoPlayer.svelte`)
- Uses `previewFilePath`/`previewStatus` instead of old single-track fields
- Download progress shows "Downloading preview…" label
### ExportDialog (`src/lib/components/ExportDialog.svelte`)
- Uses `exportFilePath`/`exportStatus`/`exportProgress` for export readiness
- Shows "Downloading best quality (X%)…" while export download is in progress
### StatusBar (`src/lib/components/StatusBar.svelte`)
- Shows dual progress: preview download → export download → ready to export
- Contextual status messages for each download phase
### UrlInput (`src/lib/components/UrlInput.svelte`)
- Added `bind:this={inputEl}` reference to input element
- `blurInput()` called at the start of `handleSubmit()` — defocuses immediately on paste/enter
- Keyboard shortcuts (I, O, etc.) now work immediately after submitting a URL
## Lessons Learned
- Storing preview and export files in separate subdirectories (`preview/`, `export/`) keeps cache management clean and avoids filename collisions between quality variants.
- Fire-and-forget pattern for background export download (`.catch()` at call site) keeps the preview flow responsive without blocking on the best-quality download.
- The `variant` parameter threading from frontend → Rust command → download manager keeps the API clean and extensible for future quality tiers.
## Follow-Up Items
- Consider showing export download progress in the timeline/player area as a subtle indicator
- The export format `bv*+ba/b` may download VP9/WebM — this is fine for ffmpeg export but won't play in WKWebView. The preview file handles playback.
- Old flat `video-clipper/` cache was cleared; users with existing caches in the old location won't get cache hits (harmless — just re-downloads)

View File

@@ -0,0 +1,72 @@
# Fix Audio Playback + Dynamic Waveform Resolution
## Task Description
1. **Fix missing audio on preview playback** — Video plays but with no sound in WKWebView.
2. **Dynamic waveform resolution** — Waveform should increase in detail as the user zooms in.
## Root Cause Analysis (Audio)
The preview file is correctly muxed (H.264 640x360 + AAC 128kbps, moov atom at offset 24 — fast-start). WKWebView config analysis:
- wry 0.55.1 defaults `autoplay: true` → sets `mediaTypesRequiringUserActionForPlayback = None`
- tauri-runtime-wry uses `WebViewBuilder::new_with_web_context()` which inherits this default
- Tauri v2.11.6 doesn't expose or override `autoplay`
- So WKWebView SHOULD be configured for audio playback
Despite correct configuration, WKWebView's Tauri asset protocol (`https://asset.localhost/...`) appears to silently drop audio tracks when streaming local files — possibly due to range request handling or MIME type issues in the custom protocol handler.
**Fix**: Load the video file as a `blob:` URL instead of using the asset protocol. This bypasses the asset protocol entirely and uses WKWebView's native blob URL handling, which reliably plays both audio and video tracks.
## Changes Made
### Audio Fix: Blob URL Video Loading (`src/lib/components/VideoPlayer.svelte`)
- Added a `$effect` that fetches the preview file via `convertFileSrc` URL, converts it to a Blob, then creates a `blob:` URL
- Video element now uses the blob URL instead of the asset protocol URL
- Added `playsinline` attribute to the video element
- Added explicit `volume = 1` and `muted = false` on `loadeddata`
- Added `.play()` Promise error handling (catches and logs rejections)
- Shows "Preparing video…" state while blob is loading
- Falls back to asset protocol URL if blob creation fails
- Properly revokes old blob URLs on session change
### Dynamic Waveform: Backend (`src-tauri/src/services/waveform_generator.rs`)
- Added `extract_waveform_range(file_path, start_time, end_time, peak_count)` function
- Uses `ffmpeg -ss {start} -t {duration}` with 44100 Hz sample rate for high-resolution extraction
- Computes peaks for just the specified time range
### Dynamic Waveform: Command (`src-tauri/src/commands/media_analysis.rs`)
- Added `extract_waveform_range` Tauri command
- Registered in `lib.rs`
### Dynamic Waveform: Frontend Binding (`src/lib/bindings/mediaAnalysis.ts`)
- Added `extractWaveformRange(filePath, startTime, endTime, peakCount)` binding
### Dynamic Waveform: Session State (`src/lib/stores/videoSession.svelte.ts`)
- Added `waveformDetailPeaks`, `waveformDetailStart`, `waveformDetailEnd` to session state
- Increased initial waveform extraction from 8,000 to 50,000 peaks (good for most zoom levels)
- Detail fields cleared on session reset
### Dynamic Waveform: Renderer (`src/lib/timeline/waveformRenderer.ts`)
- Refactored `drawWaveform` to accept a `WaveformData` object with both overview and detail peaks
- Renderer automatically uses detail peaks when they cover the visible viewport
- Falls back to overview peaks when detail is not available
### Dynamic Waveform: Timeline Component (`src/lib/components/Timeline.svelte`)
- Added auto-fetch `$effect` that monitors zoom level and viewport
- When overview peaks per pixel drops below 2, triggers a 300ms-debounced detail extraction
- Detail extraction pads visible range by 50% on each side to avoid re-fetching on small pans
- Requests 4 peaks per pixel for crisp detail
- Passes `WaveformData` to both main timeline and minimap renderers
### Renderer Types Updated (`src/lib/timeline/renderer.ts`)
- `drawTimeline` now accepts `WaveformData` instead of `number[]`
## Lessons Learned
- WKWebView's custom protocol handlers (like Tauri's `https://asset.localhost/`) can have subtle audio issues even when video plays fine. Blob URLs are a reliable workaround.
- For waveform LOD (level of detail), a two-tier approach (high-count overview + on-demand range extraction) provides the best UX: fast initial display with detail on demand.
- Debouncing the detail waveform fetch prevents spamming ffmpeg during rapid zoom/pan.
- Padding the extraction range by 50% on each side significantly reduces re-fetch frequency during small viewport adjustments.
## Follow-Up Items
- Investigate if Tauri's asset protocol can be configured to properly serve audio (may be a wry bug)
- Consider pre-computing multiple LOD levels for the waveform instead of on-demand extraction
- The blob approach loads the entire preview file into memory (~97MB for a 54-min video at 360p) — acceptable for desktop but may need optimization for very long videos

View File

@@ -0,0 +1,116 @@
# Fix Multi-Clip Workflow + Caption Support
**Date:** 2026-09-21 14:09
**Task:** Fix multi-clip creation and add time-synced caption support
## Task Description
Two main issues addressed:
1. **Multi-clip bug**: Pressing I/O keys always edited the selected clip instead of creating new clips when the playhead was outside the clip region.
2. **Caption support**: Full pipeline for downloading, displaying, and exporting time-synced captions (embedded > external, English preferred, auto-generated deprioritized).
## Changes Made
### 1. Fix Multi-Clip Workflow
**`src/lib/stores/clips.svelte.ts`**:
- Added `isTimeInsideClip(time, clipId)` helper with 0.5s tolerance
- `markInPoint(time)`: Now checks if playhead is inside the selected clip's range. If outside, deselects and sets `pendingInPoint` (starts new clip). If inside, edits the existing clip.
- `markOutPoint(time)`: Same logic — edits selected clip if inside, creates new clip from pending in-point if outside.
**`src/lib/components/Timeline.svelte`**:
- Added `selectClip(null)` call when clicking empty timeline space (no `hitTestClip` result), so the next I/O presses create a new clip.
### 2. Caption Metadata Parsing
**`src-tauri/src/services/video_resolver.rs`**:
- Extended `YtDlpJson` struct with `subtitles` and `automatic_captions` fields (both `HashMap<String, Vec<SubtitleFormat>>`)
- Added `detect_captions()` function: checks for English subtitles (en, en-US, en-GB variants), prioritizes manual over auto-generated
- Added `SubtitleFormat` struct for deserialization
- Added 4 new unit tests for caption detection
**`src-tauri/src/models.rs`**:
- Added `has_captions: bool` and `captions_are_auto: bool` to `VideoMetadata`
**`src/lib/bindings/video.ts`**:
- Added `hasCaptions` and `captionsAreAuto` to `VideoMetadata` interface
### 3. Caption Download
**`src-tauri/src/services/subtitle_downloader.rs`** (new file):
- Downloads English VTT subtitles via `yt-dlp --write-subs` (manual) or `--write-auto-subs` (auto-generated)
- Uses `--sub-langs en.*,en --sub-format vtt --convert-subs vtt`
- Searches output directory for `.en.vtt` files, falls back to any `.vtt`
**`src-tauri/src/commands/video.rs`**:
- Added `download_subtitles` Tauri command
**`src/lib/bindings/video.ts`**:
- Added `downloadSubtitles()` binding
### 4. Embedded Subtitle Check
**`src-tauri/src/commands/media_analysis.rs`**:
- Added `check_embedded_subtitles` command that uses `ffprobe -show_streams -select_streams s` to detect subtitle streams, then extracts as VTT via `ffmpeg -map 0:s:0 -f webvtt`
**`src/lib/bindings/mediaAnalysis.ts`**:
- Added `checkEmbeddedSubtitles()` binding
### 5. Caption Display in Player
**`src/lib/stores/videoSession.svelte.ts`**:
- Added `hasCaptions`, `captionsAreAuto`, `captionFilePath` to session state
- Added `loadCaptions()` function that checks embedded subs first, then downloads external
- Integrated into `triggerPostDownloadProcessing()`
**`src/lib/components/VideoPlayer.svelte`**:
- Loads VTT file as blob URL via `captionBlobUrl`
- Renders `<track>` element with proper `srclang`, `label` (with "(auto)" suffix), and `default` attribute
- Added CC toggle button (bottom-right overlay) with active/inactive styling
- Syncs `textTracks[].mode` with `captionsEnabled` state
### 6. Caption Export
**`src-tauri/src/models.rs`**:
- Added `include_captions: bool` and `caption_file_path: Option<String>` to `ExportConfig`
**`src-tauri/src/services/clip_exporter.rs`**:
- Added `build_ffmpeg_args_with_subs()`: lossless → mux as `mov_text`/`srt`; precise → burn-in via `subtitles=` filter
- Added `export_single_clip_with_subs()` with fallback to no-subs on failure
- Added `export_merged_with_subs()`
- Added 2 new unit tests for subtitle-aware arg building
**`src-tauri/src/commands/export.rs`**:
- Updated to conditionally use `_with_subs` variants based on `include_captions`
**`src/lib/bindings/export.ts`**:
- Added `includeCaptions` and `captionFilePath` to `ExportConfig` interface
**`src/lib/components/ExportDialog.svelte`**:
- Added "Captions" section with checkbox toggle (only shown when captions are available)
- Shows "(auto-generated)" label when applicable
- Shows mux vs. burn-in note based on cut mode
### Infrastructure
**`src-tauri/src/services/mod.rs`**: Registered `subtitle_downloader` module
**`src-tauri/src/lib.rs`**: Registered `download_subtitles` and `check_embedded_subtitles` commands
## Build & Test Results
- **Rust**: 34 tests pass (including 4 new caption tests + 2 new export tests)
- **Frontend**: Builds cleanly (0 errors)
- **Rust build**: Compiles with only pre-existing unused variant warnings
## Lessons Learned
- `yt-dlp --dump-json` includes `subtitles` and `automatic_captions` maps — both map language codes to arrays of `{ext, url}` objects
- WKWebView `<track>` elements work with blob URLs but need `textTracks[].mode` managed manually to sync with a toggle
- For ffmpeg subtitle burn-in, paths with colons need escaping (`\:`) in the `subtitles=` filter
- Subtitle muxing in lossless mode requires separate `-i` for the subtitle file and explicit stream mapping (`-map 0:v -map 0:a -map 1:s`)
## Follow-up Items
- The `SubtitleFormat.ext` and `SubtitleFormat.url` fields generate "never read" warnings — they exist for serde deserialization but could be suppressed with `#[allow(dead_code)]`
- Caption time offset accuracy: when using `-ss` before `-i`, subtitle timestamps may need adjustment for precise alignment in edge cases
- Sprite sheet optimization (noted from prior session) is still pending

View File

@@ -0,0 +1,69 @@
# Performance, Captions, Clip UX, and Resizable Panels
**Date:** 2026-09-21 14:39
**Task:** Fix timeline performance, caption positioning, clip creation workflow, resizable panels, and deselect UX
## Changes Made
### 1. Performance Optimization (3 files)
**`src/lib/components/VideoPlayer.svelte`**:
- Throttled `handleTimeUpdate` to ~15fps (66ms interval) using `performance.now()` check. Previously every `ontimeupdate` event triggered a reactive cascade through `session.currentTime`.
**`src/lib/components/Timeline.svelte`**:
- **In-flight guard**: Added `detailFetchInFlight` flag to prevent overlapping waveform range extraction subprocess calls. Only one ffmpeg call runs at a time.
- **Increased debounce**: Waveform detail fetch debounce increased from 300ms to 500ms.
- **Playhead-only redraws**: Split the monolithic redraw `$effect` into two:
- Structural changes (clips, zoom, waveform data, thumbnails) trigger a full `drawMainCanvas()` which saves an `ImageData` snapshot.
- `session.currentTime` changes trigger a lightweight `drawPlayheadOnly()` that restores the snapshot and draws only the playhead — avoiding the expensive thumbnail/waveform/clip rendering on every time update.
- **Thumbnail load callback**: Registered a callback via `setThumbnailLoadCallback` so thumbnail image loads trigger a targeted redraw instead of relying on the next reactive cycle.
**`src/lib/timeline/thumbnailRenderer.ts`**:
- Added `pendingLoads` Set to prevent creating duplicate `Image` objects for the same path across rapid redraws.
- Added `setThumbnailLoadCallback` API so the Timeline can request a redraw when thumbnails finish loading asynchronously.
- Image `onerror` handler cleans up the pending state.
### 2. Caption Positioning (1 file)
**`src/lib/components/VideoPlayer.svelte`**:
- Added `:global(video::cue)` CSS to center captions at the bottom of the video with a semi-transparent black background, white text, and 16px font size. Used `:global()` to bypass Svelte scoping since `::cue` is a browser-level pseudo-element.
- Added `position: relative` to the `<video>` element to scope cue rendering.
### 3. Pending In-Point Preservation (1 file)
**`src/lib/stores/clips.svelte.ts`**:
- Modified `selectClip()` to only clear `pendingInPoint` when `id !== null` (selecting a specific clip). When `id === null` (deselecting via empty timeline click), the pending in-point is preserved. This fixes the workflow: press I → click timeline to seek → press O to complete the clip.
### 4. Resizable Panels (3 files)
**`src/App.svelte`**:
- Wrapped `<Timeline />` and `<ClipList />` in a new `.timeline-clip-area` container with flex column layout.
- Added a 5px `.resize-handle` divider between them with `cursor: ns-resize` and accent color on hover/active.
- Implemented `handleResizeStart/Move/End` with `mousedown`/`mousemove`/`mouseup` on `<svelte:window>` to drag-resize. Timeline height is clamped between 80px and (total - 60px).
- `timelineHeight` stored in local `$state` (defaults to 180px).
**`src/lib/components/Timeline.svelte`**:
- Changed `.timeline-container` from `height: 160px` to `flex: 1; min-height: 80px` so it fills the space given by the parent.
- `.timeline-wrapper` set to `height: 100%`.
**`src/lib/components/ClipList.svelte`**:
- Changed from `max-height: 200px` to `height: 100%; min-height: 40px` with `box-sizing: border-box`.
### 5. Deselect UX (1 file)
**`src/lib/components/ClipList.svelte`**:
- Added `handleContainerClick` on the `.clip-list` div that calls `selectClip(null)` when clicking empty space (not a child clip row).
- Replaced the single "Export All" button with an `.action-buttons` row containing a conditional "Deselect" button (shown only when a clip is selected) alongside "Export All".
- Changed `role="button"` to `role="option"` on clip rows and `role="listbox"` on the container for better semantics.
## Build Status
- Frontend: builds cleanly (0 errors)
- Rust: compiles with only pre-existing unused variant warnings
## Lessons Learned
- Svelte's scoped CSS doesn't apply to `::cue` pseudo-elements since they're browser-level shadow constructs — `:global()` wrapper is needed.
- Canvas `getImageData`/`putImageData` is an effective technique for "snapshot and overlay" rendering patterns, avoiding expensive full redraws when only one layer (like a playhead) changes.
- `ontimeupdate` fires at varying rates across browsers/engines; throttling to a target frame rate prevents reactive cascades from dominating the UI thread.
- The `pendingInPoint` clearing bug was a subtle consequence of `selectClip(null)` being used for two different intents (deselecting a clip vs. resetting all clip state).

View File

@@ -0,0 +1,51 @@
# Fix Fundamental Performance & Audio Regression
**Date:** 2026-09-21 14:51
**Task:** Fix root causes of sluggish UI and broken audio on long videos
## Root Causes Identified
### 1. Blob URL for Video (CRITICAL)
The `VideoPlayer` was fetching the ENTIRE preview MP4 file into JavaScript memory via `fetch()` → `blob()` → `URL.createObjectURL()`. For a 54-minute 360p video (~200-500MB), this:
- Blocks the UI while the entire file is fetched into memory
- Doubles memory usage temporarily (file buffer + blob)
- Makes seeks slow because the browser must parse the blob
- Was the root cause of sluggish playback, delayed seek response, and overall UI lag
The blob URL was originally a workaround for WKWebView audio issues with Tauri's asset protocol, but the tradeoff is unacceptable for any video longer than a few minutes.
### 2. getImageData/putImageData Overhead
The "playhead-only redraw" optimization was using `getImageData()` to snapshot the canvas and `putImageData()` to restore it. On a Retina display, this copies 20+ MB of pixel data per frame — worse than just redrawing the canvas from scratch.
### 3. Parallel ffmpeg Subprocesses
`triggerPostDownloadProcessing` fired all three analysis tasks (waveform, keyframes, thumbnails) simultaneously. Each spawns an ffmpeg subprocess that decodes the full file. Three concurrent ffmpeg processes on a 54-minute video saturate the CPU.
### 4. Excessive Waveform Peaks
Hardcoded at 50,000 peaks regardless of duration. For a 54-minute video, this requires decoding the entire audio track at high resolution. Most of these peaks are never visible at the default zoom level.
## Changes Made
### `src/lib/components/VideoPlayer.svelte`
- **Removed blob URL entirely** — now uses `convertFileSrc()` directly (Tauri asset protocol). This streams from disk with zero memory overhead, enabling instant seek on any video length.
- Changed `preload="auto"` to `preload="metadata"` — only loads metadata and first frames, not the entire file.
- Removed `loadingBlob` state and associated "Preparing video…" UI state.
- Audio should work with asset protocol for H.264+AAC in MP4 (the preview format). If it doesn't, we'll investigate the specific WKWebView config rather than working around it with blob URLs.
### `src/lib/components/Timeline.svelte`
- **Removed getImageData/putImageData** snapshot mechanism entirely. All redraws go through a single `drawMainCanvas()` call, coalesced by `requestAnimationFrame`.
- Removed `lastDrawnState`, `lastFullDrawTime`, `pendingPlayheadDraw`, `drawPlayheadOnly()`.
- Merged the separate "structural" and "playhead" `$effect`s back into one — `requestAnimationFrame` already coalesces multiple calls per frame.
- Changed `pendingDraw` and `pendingMinimapDraw` from `$state` to plain `let` — they don't need reactivity and were causing unnecessary tracking overhead.
- Changed `detailFetchTimer` and `detailFetchInFlight` from `$state` to plain `let` — same reason.
- Reduced waveform detail peak count from `pw * 4` to `pw * 2` (2 peaks per pixel is sufficient).
### `src/lib/stores/videoSession.svelte.ts`
- **Sequential processing**: Changed `triggerPostDownloadProcessing` from parallel fire-and-forget to `async` sequential execution: waveform → keyframes → thumbnails → captions. Only one ffmpeg process runs at a time.
- **Scaled waveform peaks**: Changed from hardcoded 50,000 to `Math.min(10000, Math.max(2000, duration * 10))`. A 54-minute video gets 10,000 peaks (vs. 50,000 before). A 2-minute video gets 2,000.
## Lessons Learned
- **Blob URLs are not a scalable workaround** — they work for small files but are catastrophic for anything over a few minutes. Always use the asset protocol (streaming from disk) for video playback.
- **getImageData/putImageData is expensive on Retina displays** — the data transfer cost (20+ MB per snapshot) exceeds the cost of just redrawing the canvas from primitives.
- **Sequential ffmpeg is faster than parallel** for analysis tasks on the same file — the file is read from disk for each, and concurrent processes compete for I/O and CPU. Sequential processing also leaves CPU available for the UI thread.
- **$state variables for non-reactive bookkeeping** (requestAnimationFrame IDs, timers) add unnecessary tracking overhead.

View File

@@ -0,0 +1,47 @@
# Pre-Process Waveform Tiers + Processing Modal
**Date:** 2026-09-21 15:10
**Task:** Replace on-demand waveform range extraction with one-time multi-tier pre-processing pipeline, add a processing modal, eliminate the reactive `$effect` loop bug.
## Problem
The `$effect` in `Timeline.svelte` that fetched waveform ranges on demand created an infinite reactive loop: writing to `session.waveformDetailPeaks` re-triggered the same `$effect`, which re-evaluated `needsDetail`, which fired again. The terminal showed `waveform-range` calls repeating 40+ times for the same range, killing performance.
## Changes Made
### 1. Rust Backend — Waveform Tiers (`waveform-tiers-rust`)
- **`src-tauri/src/models.rs`**: Added `WaveformTiers` struct with `tier0` (~2K peaks), `tier1` (~10K peaks), `tier2` (all raw peaks, 50K–200K depending on duration).
- **`src-tauri/src/services/waveform_generator.rs`**: Rewrote entirely. New `extract_waveform_tiers(file_path, duration)` function runs ffmpeg once at a sample rate yielding ~100 peaks/sec (capped 50K–200K), then downsamples into 3 tiers in Rust. Added `downsample()` helper (max-of-chunk). Removed `extract_waveform_range` function. Added 6 unit tests for downsample and compute_peaks.
- **`src-tauri/src/commands/media_analysis.rs`**: Replaced `extract_waveform` and `extract_waveform_range` commands with single `extract_waveform_tiers` command.
- **`src-tauri/src/lib.rs`**: Updated handler registration — `extract_waveform_tiers` replaces `extract_waveform` + `extract_waveform_range`.
### 2. Frontend — Tier-Based Rendering (`waveform-tiers-frontend`)
- **`src/lib/bindings/mediaAnalysis.ts`**: Added `WaveformTiers` interface and `extractWaveformTiers` binding. Removed `extractWaveformRange` binding.
- **`src/lib/timeline/waveformRenderer.ts`**: Rewrote. `WaveformData` is now `{ tiers: WaveformTiers | null }`. `pickTier()` selects the lowest-resolution tier with ≥1 peak per pixel for the visible range — pure arithmetic, zero IPC.
- **`src/lib/timeline/renderer.ts`**: Updated `drawTimeline` signature and waveform condition to use `waveform.tiers`.
- **`src/lib/stores/videoSession.svelte.ts`**: Replaced `waveformPeaks`, `waveformDetailPeaks`, `waveformDetailStart`, `waveformDetailEnd` with `waveformTiers: WaveformTiers | null`. Added `processingStep` state. Processing pipeline now calls `extractWaveformTiers` instead of `extractWaveform`.
- **`src/lib/components/Timeline.svelte`**: **Deleted the entire `$effect` for waveform range fetching** — the infinite loop bug is eliminated. Removed `detailFetchTimer`, `detailFetchInFlight`, `needsDetail`, `extractWaveformRange` import. `waveformData` is now simply `{ tiers: session.waveformTiers }`.
### 3. Processing Modal (`processing-modal`)
- **`src/lib/components/ProcessingModal.svelte`** (new): Modal overlay showing processing steps with icons (✓ done, ⏳ current, ○ pending). Steps: downloading → waveform → keyframes → thumbnails → captions → done. Includes download progress bar.
- **`src/App.svelte`**: Imported and renders `<ProcessingModal />` when `session.processingStep` is neither `'idle'` nor `'done'`.
- **`src/lib/stores/videoSession.svelte.ts`**: `processingStep` is set at each stage of `triggerPostDownloadProcessing` and `beginDownload`.
## Verification
- **Rust**: `cargo build` succeeds (3 pre-existing warnings). `cargo test` passes all 37 tests (including 6 new waveform generator tests).
- **Frontend**: `svelte-check` reports only pre-existing vite.config.ts errors and a11y warnings — no new issues.
## Lessons Learned
- The on-demand `$effect` → async IPC → write reactive state → re-trigger `$effect` pattern is fundamentally broken in Svelte 5. The correct approach is pre-computation: do all CPU work up front, store the results, and let the renderer pick the right data with pure arithmetic.
- For waveform at typical zoom levels (even 100x on an 800px-wide canvas), 50K peaks gives ~6 peaks/pixel, which is plenty. No need for on-demand range extraction.
- Downsampling by max-of-chunk preserves peak amplitude fidelity while drastically reducing array size.
## Follow-up Items
- Audio playback may need verification after the session store changes (the `previewFilePath` → `convertFileSrc` pipeline is unchanged, but should be tested).
- Consider persisting waveform tiers to disk cache for instant reload on re-open.

View File

@@ -0,0 +1,82 @@
# Progress, Resize, Captions, Audio, Preview Upgrade
**Date:** 2026-09-21 19:31
**Task:** Implement 6 items from the "Progress Resize Captions Audio" plan
## Changes Made
### 1. Waveform Speed Fix + Granular Progress (waveform-speed-fix)
**Files:** `src-tauri/src/services/waveform_generator.rs`, `src-tauri/src/commands/media_analysis.rs`, `src/lib/bindings/mediaAnalysis.ts`, `src/lib/stores/videoSession.svelte.ts`, `src/lib/components/ProcessingModal.svelte`
- **Critical bug fix:** `aresample=N` was setting the output sample rate to N Hz (e.g., 200,000 Hz) instead of producing N total samples. For a 54-min video, this generated ~650M samples. Fixed by computing actual rate as `raw_count / duration` and using `-ar` flag instead. ~3000x speedup for long videos.
- Changed `waveform_generator.rs` to use child process spawning with piped stdout/stderr, parse `out_time_us=` from ffmpeg progress output, and report progress via a callback.
- Added `Channel<f64>` progress parameter to `extract_waveform_tiers` Tauri command.
- Added `processingProgress` to session store; reset in `setMetadata`, `clearMediaFields`.
- Updated `ProcessingModal` to show progress bars for every step (not just download).
### 2. Fix Resize Handles (fix-resize)
**Files:** `src/lib/components/VideoPlayer.svelte`, `src/App.svelte`
- Changed `VideoPlayer` min-height from 200px to 0.
- Removed `max-height: 60vh` from timeline-clip-area.
- Lowered min-height constraints: timeline-clip-area 140→80, timeline-pane 80→40, cliplist-pane 40→20.
- Updated resize handler min/max constraints to match.
### 3. Custom Caption Rendering (fix-captions)
**Files:** `src/lib/utils/vttParser.ts` (new), `src/lib/stores/preferences.svelte.ts`, `src/lib/components/VideoPlayer.svelte`
- Created `vttParser.ts` with `parseVtt()` and `getActiveCues()` for custom VTT parsing.
- Replaced `<track>` element with custom HTML overlay positioned absolutely in the video player.
- Captions now rendered bottom-center (or top, configurable) with single background layer (no double-background).
- Added `CaptionSettings` interface to preferences store with font size, text color, background opacity, text outline, and position.
### 4. Caption Settings Panel (caption-settings-ui)
**Files:** `src/lib/components/CaptionSettingsPanel.svelte` (new), `src/lib/components/VideoPlayer.svelte`
- Created settings sub-menu with range sliders, color picker, checkbox, radio buttons.
- Accessible via ⚙ button next to CC toggle.
- Settings persist via preferences store.
- "Reset to defaults" button included.
### 5. Local HTTP Media Server for Audio (fix-audio)
**Files:** `src-tauri/src/services/media_server.rs` (new), `src-tauri/src/services/mod.rs`, `src-tauri/src/lib.rs`, `src-tauri/Cargo.toml`, `src/lib/bindings/video.ts`, `src/lib/components/VideoPlayer.svelte`
- Added `axum` + `tower-http` (with fs feature) dependencies.
- Created media server that starts on a random available port at app launch, serving files from filesystem root with full range-request support via `ServeDir`.
- Exposed port to frontend via `get_media_server_port` Tauri command.
- VideoPlayer now constructs video URLs as `http://127.0.0.1:{port}/path/to/file.mp4` instead of using `convertFileSrc()`.
- This fixes WKWebView's audio issues with the asset protocol.
### 6. Preview Upgrade Toast (preview-upgrade)
**Files:** `src/lib/stores/videoSession.svelte.ts`, `src/lib/components/VideoPlayer.svelte`, `src/lib/components/StatusBar.svelte`
- Added `activeVideoPath` (initially set to preview, upgradeable to export), `showUpgradeToast`, `upgradePreview()`, `dismissUpgradeToast()`.
- Toast appears in top-right of video player when export download completes and processing is done.
- "Reload Preview (HQ)" button added to StatusBar.
- Video source now driven by `activeVideoPath` instead of `previewFilePath`.
## Verification
- `cargo build`: ✅ (only pre-existing warnings)
- `cargo test`: ✅ 37 tests pass
- `npm run check`: ✅ (only pre-existing vite.config.ts errors)
## Lessons Learned
- `aresample=N` in ffmpeg sets the **output sample rate in Hz**, not a total count. Must compute rate = count/duration.
- Tauri's asset protocol (`https://asset.localhost/...`) has issues with audio track streaming in WKWebView — a local HTTP server with range request support via axum/tower-http is a clean fix.
- TypeScript's control flow analysis can over-narrow `$state` values across async callback boundaries, requiring explicit cast (`as ProcessingStep`) to work around.
- Custom caption rendering (parsing VTT + HTML overlay) gives full styling control vs. browser `::cue` pseudo-element limitations.
## Follow-up Items
- Smoke test the full flow end-to-end with a real YouTube video.
- Verify audio playback works correctly via the local HTTP server.
- Test caption settings persistence across app restarts.
- Consider adding progress reporting to keyframe extraction and thumbnail generation (currently only waveform has it).

View File

@@ -0,0 +1,50 @@
# Fix HQ Upgrade, Audio Playback, and Waveform
**Date:** 2026-09-21 20:01
**Task:** Fix three bugs reported after the previous batch of changes
## Changes Made
### 1. Waveform Truncation Fix
**File:** `src-tauri/src/services/waveform_generator.rs`
**Root cause:** `compute_peaks` used `peaks.resize(target_count, 0.0)` to pad the peaks array to exactly `target_count` entries. But when `chunk_size = ceil(samples / target_count)` rounds up (e.g., 201439 samples / 200000 target = chunk_size 2), the actual number of chunks produced is far less than `target_count` (100720 vs 200000). The resize pads ~50% of the array with zeros, compressing the real waveform into the first half and leaving the rest flat.
**Fix:** Removed the `peaks.resize(target_count, 0.0)` call. The downstream rendering code already handles arrays of any length — the peak count from chunking is the correct count.
### 2. Export Format Fix (HQ Upgrade Black Screen)
**File:** `src/lib/stores/videoSession.svelte.ts`
**Root cause:** `EXPORT_FORMAT = 'bv*+ba/b'` allowed yt-dlp to pick the best format regardless of codec — which resulted in AV1+Opus in WebM. WKWebView on macOS cannot decode AV1 or Opus, so upgrading to the HQ preview produced a black screen with no playback.
**Fix:** Changed `EXPORT_FORMAT` to `'bv*[vcodec^=avc1]+ba[acodec^=mp4a]/b[ext=mp4]/best[ext=mp4]'` to force H.264+AAC in MP4, matching what WKWebView can play. Deleted the cached `.webm` export file so the next run re-downloads in the correct format.
### 3. Volume Controls + Audio Diagnostics
**File:** `src/lib/components/VideoPlayer.svelte`
- Added `volume` and `isMuted` state variables with reactive sync to the video element
- Added mute toggle button (speaker emoji) and volume slider in a new bottom controls bar
- Added diagnostic `console.log` in `handleLoadedData` to print muted, volume, readyState, audioTracks length, and dimensions — this will help trace the audio issue if it persists
- Reorganized the bottom overlay: volume controls on the left, CC/settings on the right
## Lessons Learned
- When `samples.chunks(chunk_size)` is used with a `chunk_size > 1`, the number of resulting chunks is `ceil(samples / chunk_size)`, which is strictly less than `target_count` when `chunk_size = ceil(samples / target_count)`. Never pad with zeros — it corrupts the waveform mapping to time.
- yt-dlp's `bv*+ba/b` format spec can pick any codec. On macOS, always constrain to H.264+AAC for WKWebView compatibility.
- The audio issue may be related to mixed content (HTTPS frontend loading HTTP media) — the diagnostic logging added here will help isolate whether the issue is mute state, volume, or something lower-level in WebKit.
## Verification
- `cargo build`: OK (only pre-existing warnings)
- `cargo test`: 37 tests pass
- `npm run check`: OK (only pre-existing vite.config.ts errors)
## Follow-up
- Run `npx tauri dev` and verify waveform renders across the full timeline
- Verify HQ upgrade no longer produces a black screen
- Check console output from `handleLoadedData` diagnostic to trace audio issue
- If audio still doesn't play, investigate mixed content blocking in WKWebView

View File

@@ -0,0 +1,20 @@
# Fix Captions Missing Due to CORS
## Task
Captions/subtitles were not appearing at all - the CC button and settings button were invisible because `parsedCues` was always empty.
## Root Cause
The local media server (axum on `http://127.0.0.1:<random_port>`) had **no CORS headers**. The frontend WebView runs at a different origin (`http://localhost:1420` in dev, `tauri://localhost` in prod). While `<video>` elements can load cross-origin media without CORS (they use "no-cors" mode), the `fetch()` call used to load the VTT caption file was blocked by the browser's same-origin policy.
The `fetch()` silently failed (caught by `.catch()`), setting `parsedCues = []`, which meant the `{#if parsedCues.length > 0}` conditional in `VideoPlayer.svelte` never rendered the CC controls.
## Changes Made
1. **`src-tauri/Cargo.toml`**: Added `"cors"` feature to `tower-http` dependency.
2. **`src-tauri/src/services/media_server.rs`**: Added `CorsLayer::permissive()` to the axum router, enabling cross-origin `fetch()` from the WebView.
## Lessons Learned
- `<video src="...">` does NOT require CORS for basic playback - the browser loads media in "no-cors" mode. But `fetch()` to the same URL WILL be blocked without CORS headers. This discrepancy is why video/audio played fine but caption loading silently failed.
- When a `fetch()` fails silently in a `.catch()` handler, there's no visible error in the UI — only in the browser DevTools console. Adding CORS from the start would have prevented this class of issues.
## Follow-up
None - this was a targeted 2-file fix.

View File

@@ -0,0 +1,47 @@
# Fix Caption Export and Add Burn-In Option
## Task
Exported clips were not including captions despite the "include captions" option being checked. Additionally, the user requested a "burn-in" option to render subtitles directly into the video frames.
## Root Causes
1. **Precise mode burn-in (timestamp offset)**: `-ss` was placed before `-i` (input seeking), shifting output PTS to 0. The `subtitles` filter reads the original VTT with absolute timestamps (e.g., cues at 300s), but the output video starts at 0s — no cues matched.
2. **Lossless mode muxing (VTT formatting)**: YouTube auto-generated VTT has karaoke-style `<c>` tags, inline timestamps (`<00:00:00.599>`), and positioning metadata (`align:start position:0%`) that confuse ffmpeg's VTT parser when converting to `mov_text`.
3. **Silent fallback**: When the ffmpeg subtitle export failed, the code silently fell back to exporting without subtitles, making the failure invisible to the user.
## Changes Made
### `src-tauri/src/services/clip_exporter.rs` (major rewrite)
- **Added `sanitize_vtt_for_ffmpeg()`**: Strips karaoke `<c>` tags, inline timestamps, and positioning metadata from VTT files. Writes a cleaned temp file.
- **Added `strip_vtt_tags()`**: Helper to remove all HTML-like tags from VTT cue text lines.
- **Replaced `build_ffmpeg_args_with_subs()`** with two separate functions:
- `build_ffmpeg_args_mux_subs()`: Muxes subtitles as a track. Uses output seeking (`-ss`/`-to` after `-i`) for correct timestamp alignment. Works with both lossless and precise cut modes.
- `build_ffmpeg_args_burnin_subs()`: Burns subtitles into video using the `subtitles` filter. Uses output seeking so the filter reads correct VTT timestamps. Always re-encodes (overrides to H.264/AAC if lossless mode selected).
- **Updated `export_single_clip_with_subs()`**: Now takes `burn_in: bool`, sanitizes VTT before use, dispatches to mux or burn-in builder. **Removed the silent fallback** — errors now surface to the user.
- **Updated `export_merged_with_subs()`**: Takes `burn_in: bool`, passes through to per-segment export.
- **Updated tests**: Replaced old tests for removed function, added tests for both new builders, added a VTT tag stripping test. All 40 tests pass.
### `src-tauri/src/models.rs`
- Added `pub burn_in_captions: bool` to `ExportConfig`.
### `src-tauri/src/commands/export.rs`
- Threads `config.burn_in_captions` through to `export_single_clip_with_subs` and `export_merged_with_subs`.
### `src/lib/bindings/export.ts`
- Added `burnInCaptions: boolean` to the TypeScript `ExportConfig` interface.
### `src/lib/components/ExportDialog.svelte`
- Added `burnInCaptions` state.
- Added "Burn into video" sub-checkbox under the "Include captions" checkbox.
- Contextual notes: "Captions will be muxed as a subtitle track" vs "Captions will be burned into the video" vs "Burn-in requires re-encoding (precise mode will be used)".
- Threads `burnInCaptions` through to the export config.
## Lessons Learned
- ffmpeg's `-ss` before `-i` (input seeking) adjusts output PTS to start at 0, but the `subtitles` filter reads timestamps from the original VTT file — they must match. Output seeking (`-ss` after `-i`) preserves original timestamps.
- YouTube auto-generated VTT files contain karaoke formatting that ffmpeg's VTT-to-mov_text converter can't handle cleanly. Sanitizing before use is essential.
- Silent fallbacks that hide errors waste debugging time. Surface errors to the user.
## Follow-up
- The CORS fix from the prior session (media_server.rs adding `CorsLayer::permissive()`) is also needed for captions to display in the preview player.

View File

@@ -0,0 +1,41 @@
# Fix Caption Export Bugs (Round 2)
## Task
Three bugs in the caption export implementation needed fixing after the initial burn-in/mux feature was added.
## Bug 1: Burn-in filter syntax error
**Symptom**: ffmpeg error `No option name near '/var/.../sanitized.vtt'`
**Root cause**: `format!("subtitles='{}'", path)` passed literal single quotes to ffmpeg via `Command::args()`. Since there's no shell to strip them, ffmpeg tried to open `'sanitized.vtt'` (with quotes in the filename).
**Fix**: Changed to `format!("subtitles={}", path)` — no quotes needed when passing args directly.
## Bug 2: Wrong output duration (1:43 instead of 32s)
**Symptom**: Exported clip was 103 seconds instead of 32 seconds.
**Root cause**: Both subtitle functions used `-to end_time` with output seeking (`-ss`/`-to` after `-i`). After `-ss` discards initial frames and the muxer resets PTS to 0, `-to 103` means "output 103 seconds" rather than "stop at input timestamp 103".
**Fix**: Replaced `-to end_time` with `-t duration` (`clip.end_time - clip.start_time`) in both `build_ffmpeg_args_mux_subs` and `build_ffmpeg_args_burnin_subs`. `-t` is unambiguous.
## Bug 3: Muxed subtitles invisible in VLC
**Symptom**: Subtitle track exists in the file but selecting it in VLC shows nothing.
**Root cause**: With output seeking, video PTS was reset to 0 by the muxer, but subtitle cues retained their original absolute timestamps (e.g., 71s-103s). The player sees subtitle cues at 71s but the video is only 32s long.
**Fix**: Two-part strategy change:
1. Added `trim_and_sanitize_vtt()` function that extracts cues within the clip's time range and shifts all timestamps to start from 0.
2. Switched `build_ffmpeg_args_mux_subs` to use **input seeking** (`-ss`/`-t` before `-i`) for speed. The pre-trimmed VTT timestamps already start from 0, matching the video output.
3. Burn-in mode (`build_ffmpeg_args_burnin_subs`) retains **output seeking** since the `subtitles` filter needs to see the original absolute timestamps from the VTT.
## Changes Made
### `src-tauri/src/services/clip_exporter.rs`
- Removed single quotes from `subtitles=` filter value
- Changed `-to end` to `-t duration` in both subtitle builder functions
- Added `trim_and_sanitize_vtt()` with VTT timestamp parsing/shifting
- Added helper functions: `parse_vtt_timestamp_line()`, `parse_vtt_ts()`, `format_vtt_timestamp()`
- Switched `build_ffmpeg_args_mux_subs` to input seeking with pre-trimmed VTT
- Updated `export_single_clip_with_subs` to use trimmed VTT for mux, sanitized VTT for burn-in
- Updated all related tests (43 tests pass)
## Lessons Learned
- `Command::args()` in Rust bypasses the shell entirely — shell-style quoting (single quotes) becomes literal characters in the argument. Never quote paths for ffmpeg filters when using `Command::args()`.
- ffmpeg's `-to` is ambiguous with output seeking — it may refer to post-PTS-reset time rather than input time. `-t duration` is always safe.
- When using input seeking on video (`-ss` before `-i`), subtitle PTS from a separate input must match the reset PTS (~0). Pre-trimming the VTT with shifted timestamps is the reliable approach.
## Follow-up
None — all three user-reported bugs are addressed.

View File

@@ -0,0 +1,74 @@
# Fix Burn-In Export and Re-Introduce Karaoke Preview
**Date:** 2026-09-21 22:06
**Task:** Implement graceful burn-in fallback and re-introduce karaoke word-by-word highlighting in preview captions.
## Changes Made
### 1. Detect `subtitles` Filter Availability (Burn-In Guard)
**Problem:** The user's ffmpeg build (Homebrew 9.0.1_1) lacks `libass`, so the `subtitles` filter is unavailable. Attempting burn-in export caused ffmpeg to crash with a filter parse error.
**Fix:**
- **`src-tauri/src/services/dependency_manager.rs`** — Added `check_subtitles_filter()` that runs `ffmpeg -filters` and checks for `subtitles V->V` in the output.
- **`src-tauri/src/commands/dependencies.rs`** — Added `check_subtitles_filter_available()` Tauri command.
- **`src-tauri/src/lib.rs`** — Registered the new command.
- **`src/lib/bindings/dependencies.ts`** — Added `checkSubtitlesFilterAvailable()` frontend binding.
- **`src/lib/components/ExportDialog.svelte`** — On mount, checks filter availability via the new command. If unavailable, the "Burn into video" checkbox is disabled with a help message telling the user how to install libass.
### 2. Parse Word-Level Timings from VTT
**Problem:** YouTube auto-generated VTT files contain word-level timestamps in `<timestamp><c> word</c>` patterns. The old parser stripped ALL tags, losing this timing data.
**Fix:**
- **`src/lib/utils/vttParser.ts`** — Complete rewrite:
- Added `WordSegment` interface (`{ text: string; startTime: number }`).
- Added optional `words?: WordSegment[]` to `VttCue`.
- Added `parseWordTimings()` that extracts `<HH:MM:SS.mmm><c> word</c>` patterns.
- Added `getActiveWords()` helper that returns each word marked as `spoken` or upcoming based on `currentTime`.
- The plain `text` field is preserved (stripped) for backward compat.
### 3. Render Karaoke Word-by-Word Highlights
**Fix:**
- **`src/lib/components/VideoPlayer.svelte`** — Updated caption rendering:
- When `wordHighlight` is enabled and the cue has word segments, each word is rendered as a separate `<span>`.
- Spoken words display at full opacity; upcoming words display at 40% opacity.
- CSS transition smooths the opacity change.
- When disabled or no word data available, falls back to plain text rendering.
### 4. Add Word Highlight Toggle
**Fix:**
- **`src/lib/stores/preferences.svelte.ts`** — Added `wordHighlight: boolean` to `CaptionSettings` interface (default: `true`).
- **`src/lib/components/CaptionSettingsPanel.svelte`** — Added "Word-by-word highlight" checkbox toggle.
## Files Modified
| File | Change |
|------|--------|
| `src-tauri/src/services/dependency_manager.rs` | Added `check_subtitles_filter()` |
| `src-tauri/src/commands/dependencies.rs` | Added `check_subtitles_filter_available` command |
| `src-tauri/src/lib.rs` | Registered new command |
| `src/lib/bindings/dependencies.ts` | Added `checkSubtitlesFilterAvailable()` binding |
| `src/lib/components/ExportDialog.svelte` | Conditional burn-in disable + help text |
| `src/lib/utils/vttParser.ts` | Word-level timing parser + `getActiveWords()` |
| `src/lib/components/VideoPlayer.svelte` | Karaoke word rendering |
| `src/lib/stores/preferences.svelte.ts` | Added `wordHighlight` to `CaptionSettings` |
| `src/lib/components/CaptionSettingsPanel.svelte` | Added word highlight toggle |
## Build Verification
- `cargo check` — passes (only pre-existing warnings)
- `svelte-check` — passes (only pre-existing `vite.config.ts` type errors)
## Lessons Learned
- YouTube VTT word-timing format: leading text has no `<c>` wrapper, only subsequent words do. The first word inherits the cue's start time.
- The `subtitles` filter requires libass at ffmpeg compile time — it's not enough to have ffmpeg installed. Homebrew's default build may or may not include it depending on the formula version.
## Follow-Up Items
- Manual smoke test: Load a video with captions, verify karaoke highlighting works during playback.
- Test export with captions enabled (muxed track) to confirm no regressions.
- If user installs libass and reinstalls ffmpeg, burn-in should automatically become available next time ExportDialog opens.

View File

@@ -0,0 +1,64 @@
# Styled Burn-In and Granular Export Progress
**Date:** 2026-09-21 23:00
**Task:** Apply user's caption style settings to burned-in subtitles via ffmpeg's ASS force_style, and add granular per-clip progress reporting during export.
## Changes Made
### 1. CaptionStyle Model (Rust + TS)
- **`src-tauri/src/models.rs`** — Added `CaptionStyle` struct with `font_size`, `text_color`, `background_opacity`, `text_outline`, `position`. Added `caption_style: Option<CaptionStyle>` to `ExportConfig`.
- **`src/lib/bindings/export.ts`** — Added matching `CaptionStyle` interface and `captionStyle: CaptionStyle | null` to `ExportConfig`. Added `clipProgress` event variant to the discriminated union. Added optional `onClipProgress` callback parameter to `exportClips()`.
### 2. ASS force_style in clip_exporter.rs
- **`src-tauri/src/services/clip_exporter.rs`**:
- `hex_to_ass_color()` — Converts CSS hex `#RRGGBB` to ASS `&H00BBGGRR` format.
- `opacity_to_ass_back_colour()` — Converts 0-1 opacity to ASS alpha-prefixed BackColour.
- `caption_style_to_force_style()` — Builds the full ASS force_style string from CaptionStyle (Fontsize, PrimaryColour, BackColour, Outline, Shadow, Alignment, MarginV, BorderStyle).
- `build_ffmpeg_args_burnin_subs()` now accepts `Option<&CaptionStyle>` and appends `:force_style='...'` to the subtitles filter when provided.
- `parse_ffmpeg_time()` — Extracts `time=HH:MM:SS.mm` from ffmpeg stderr progress output.
- `run_ffmpeg_with_progress()` — Runs ffmpeg with piped stderr, parses progress lines, and invokes a callback with 0.0-1.0 percent values.
### 3. Progress Callbacks Throughout Export Pipeline
All four export functions (`export_single_clip`, `export_single_clip_with_subs`, `export_merged`, `export_merged_with_subs`) now accept progress callbacks. They use `run_ffmpeg_with_progress()` instead of `.output()` for the encode step, streaming real-time progress.
### 4. export.rs Command Handler
- **`src-tauri/src/commands/export.rs`** — Added `ClipProgress` event variant with `current`, `total`, `label`, `percent`. Wired progress callbacks for both Individual and Merged export paths. Passes `caption_style` through to the exporter.
### 5. ExportDialog Frontend
- **`src/lib/components/ExportDialog.svelte`**:
- Populates `captionStyle` in config from `preferences.captionSettings` when burn-in is enabled.
- Shows a progress bar with percentage during each clip export.
- Displays note that karaoke word highlighting is preview-only for burn-in.
## Files Modified
| File | Change |
|------|--------|
| `src-tauri/src/models.rs` | Added `CaptionStyle` struct, added field to `ExportConfig` |
| `src-tauri/src/services/clip_exporter.rs` | ASS helpers, force_style, ffmpeg progress streaming, progress callbacks |
| `src-tauri/src/commands/export.rs` | `ClipProgress` event, caption_style passthrough, progress wiring |
| `src/lib/bindings/export.ts` | `CaptionStyle` interface, `clipProgress` event, `onClipProgress` callback |
| `src/lib/components/ExportDialog.svelte` | captionStyle in config, progress bar UI, karaoke note |
## Build Verification
- `cargo check` — passes (only pre-existing warnings)
- `cargo test` — all 43 tests pass
- `svelte-check` — passes (only pre-existing `vite.config.ts` errors)
## Lessons Learned
- ASS color format uses BGR byte order with alpha prefix (`&HAA_BB_GG_RR`), where alpha `00` = opaque and `FF` = transparent (inverted from CSS).
- `BorderStyle=4` in ASS gives an opaque background box behind text (like CSS background), vs `BorderStyle=1` which uses outline+shadow only.
- ffmpeg progress is on stderr, not stdout. The `time=` field is the key metric for computing encode progress percentage.
- For merged exports, progress is reported per-segment during encoding, plus a concat step at the end.
## Follow-Up Items
- Smoke test burn-in export to verify styled subtitles render correctly.
- Karaoke word-by-word animation in burn-in would require converting VTT word timestamps to ASS `\k` override tags — complex and fragile, noted as preview-only limitation.

View File

@@ -0,0 +1,60 @@
# Karaoke Burn-In for Exported Subtitles
**Date:** 2026-09-21 23:11
**Task:** Convert word-timed YouTube VTT captions to ASS format with karaoke `\kf` tags so burned-in subtitles have the same word-by-word highlight effect as the preview.
## Changes Made
### 1. VTT-to-ASS Converter with Karaoke Tags
**`src-tauri/src/services/clip_exporter.rs`** — Added several new functions:
- `vtt_to_ass_with_karaoke()` — Main converter. Reads raw VTT, generates a complete ASS file with:
- `[Script Info]` header (1920x1080 play resolution)
- `[V4+ Styles]` section with the user's CaptionStyle baked in (PrimaryColour, SecondaryColour at 60% alpha for the "dim/upcoming" look, BackColour, Outline, Shadow, Alignment, BorderStyle=4)
- `[Events]` section where each cue is a Dialogue line with `\kf<centiseconds>` tags for word-level karaoke fill animation
- Graceful fallback: cues without word-level timestamps render as plain text (no `\kf`)
- `parse_vtt_word_timings()` — Rust equivalent of the TypeScript `parseWordTimings()`. Extracts `(word, start_time)` pairs from YouTube's `<timestamp><c> word</c>` format.
- `format_ass_timestamp()` — Formats seconds as ASS timestamp `H:MM:SS.cc` (centiseconds).
- `hex_to_ass_color_with_alpha()` — Like `hex_to_ass_color()` but with a specific alpha byte.
- `find_timestamp_tag_pos()` / `find_next_timestamp()` — Helpers for parsing VTT timestamp tags without regex.
### 2. Updated Burn-In Export Path
In `export_single_clip_with_subs()`, the burn-in path now uses:
```
VTT → vtt_to_ass_with_karaoke() → .ass file → subtitles= filter
```
Instead of the previous:
```
VTT → sanitize_vtt_for_ffmpeg() → plain VTT → subtitles= filter
```
Since the style is now embedded in the ASS file header, `caption_style` is passed as `None` to `build_ffmpeg_args_burnin_subs()` (no `force_style` needed — avoids potential conflicts between the ASS header and force_style overrides).
### 3. Updated UI Note
**`src/lib/components/ExportDialog.svelte`** — Changed the burn-in note from "word-by-word highlighting is preview-only" to "Captions will be burned into the video with your style settings".
## Files Modified
| File | Change |
|------|--------|
| `src-tauri/src/services/clip_exporter.rs` | Added VTT-to-ASS converter, word timing parser, ASS timestamp formatter, updated burn-in path |
| `src/lib/components/ExportDialog.svelte` | Updated burn-in note text |
## Build Verification
- `cargo check` — passes (only pre-existing warnings)
- `cargo test` — all 43 tests pass
## Technical Notes
- ASS karaoke `\kf<N>` means "fill this word over N centiseconds", transitioning from SecondaryColour to PrimaryColour. This matches the preview behavior where upcoming words are dimmed and spoken words are bright.
- The ASS SecondaryColour is set to the same text color with 0x99 alpha (~60% transparent), matching the preview's `opacity: 0.4` for upcoming words.
- No regex crate was needed — the VTT word timing pattern is simple enough to parse with manual string operations.
- The `subtitles` ffmpeg filter handles ASS files natively (same libass backend), so no filter change was needed.

View File

@@ -0,0 +1,62 @@
# Fix Karaoke Burn-In Sync and Export Progress
**Date:** 2026-09-21 23:22
**Task:** Fix two bugs: karaoke ASS burn-in out of sync / skipping lines, and export progress stuck at 0% then jumping to 100%.
## Changes Made
### Bug 1: Karaoke ASS Burn-In — Two Root Causes Fixed
**A. YouTube VTT two-line overlay pattern.**
YouTube auto-generated VTT cues have a specific structure:
- Many cues have TWO text lines: line 1 is static context (previous cue text, no `<c>` tags), line 2 is the karaoke line (with `<timestamp><c> word</c>` tags)
- Zero-duration transition cues (e.g., `00:00:02.629 --> 00:00:02.639`) are just visual transition frames
The old code joined all text lines and tried to parse word timings from the combined mess. Fix:
- Skip cues where `(end - start) < 0.05s` — these are transition frames
- For multi-line cues, only process the line that contains `<c>` tags for karaoke
- Ignore the static context line (previous cue's text repeated)
- Cues with no `<c>` tags emit as plain text (graceful fallback)
**B. Wrong ffmpeg filter.**
Per ffmpeg-micro.com: "The `subtitles` filter routes your file through libavformat's subtitle converter and applies `force_style` overrides, which can flatten karaoke timing. The `ass` filter hands the file straight to libass."
Fix: Changed `build_ffmpeg_args_burnin_subs()` to use `-vf ass=<path>` instead of `-vf subtitles=<path>`. Since style is already baked into the ASS `[V4+ Styles]` header, no `force_style` is needed. The `_caption_style` parameter is now ignored (prefixed with `_`).
### Bug 2: Export Progress — Root Cause Fixed
ffmpeg writes its progress output to stderr using `\r` (carriage return) to overwrite the same line, NOT `\n` (newline). `BufReader::lines()` splits on `\n` only, so all progress was buffered as one giant "line" that only yielded when ffmpeg exited.
Fix: Use ffmpeg's `-progress pipe:1 -nostats` flag:
- Adds `-progress pipe:1 -nostats` to ffmpeg args (before the output file)
- Reads stdout (not stderr) — `-progress` outputs proper `\n`-delimited `key=value` pairs
- Parses `out_time_ms=<microseconds>` lines and computes `percent = out_time_ms / (duration * 1_000_000)`
- stderr is still piped but drained on a separate thread for error reporting on failure
Example `-progress pipe:1` output (proper newlines):
```
frame=150
fps=45.2
out_time_ms=6250000
progress=continue
```
## Files Modified
| File | Change |
|------|--------|
| `src-tauri/src/services/clip_exporter.rs` | Fixed VTT parsing, switched to `ass=` filter, rewrote progress to use `-progress pipe:1` |
## Build Verification
- `cargo test` — all 43 tests pass
- `cargo check` — passes (only pre-existing warnings)
- `svelte-check` — passes (only pre-existing `vite.config.ts` errors)
## Lessons Learned
- YouTube VTT is NOT simple cue-per-line. It uses a two-line overlay pattern where line 1 is a static context repeat and line 2 has the karaoke data. Zero-duration cues are transition frames.
- The `ass` filter and `subtitles` filter in ffmpeg both use libass, but `subtitles` applies format conversion and `force_style` that can destroy `\kf` karaoke timing. Always use `ass` for karaoke.
- ffmpeg's stderr progress uses `\r` not `\n`. For Rust's `BufReader::lines()` to work, use `-progress pipe:1 -nostats` which outputs proper `\n`-delimited key=value pairs to stdout.

View File

@@ -0,0 +1,27 @@
# Fix Burned-In Caption Styling
**Date:** 2026-09-21 23:37
**Task:** Fix ASS styling in burn-in export to be readable and match the preview better.
## Changes Made
**`src-tauri/src/services/clip_exporter.rs`** — Updated `vtt_to_ass_with_karaoke()` ASS style generation:
| ASS Field | Before (broken) | After (fixed) |
|-----------|-----------------|---------------|
| OutlineColour | `primary_colour` (white) | `&H00000000` (black) |
| SecondaryColour | Same color with alpha `&H99FFFFFF` | Distinct grey `&H99AAAAAA` |
| BorderStyle | `4` (opaque box) always | `1` (outline+shadow) when textOutline=true; `3` (box only) when false |
| Outline | 2 always (even with box) | 2 with BorderStyle=1; 0 with BorderStyle=3 |
### Why each change matters:
- **OutlineColour=black**: The preview uses `text-shadow: 1px 1px 2px rgba(0,0,0,0.9)` for readability. Setting OutlineColour to white made the outline invisible against white text and washed out the karaoke fill effect.
- **SecondaryColour=grey**: For `\kf` karaoke, libass sweeps from SecondaryColour to PrimaryColour. With both being variants of white (just different alpha), the sweep was barely visible. Using a distinct grey (`&H99AAAAAA`) makes upcoming words clearly dimmed grey, sweeping to bright white when spoken.
- **BorderStyle=1 vs 3**: `BorderStyle=4` draws an opaque background box AND renders the outline, creating a "double background" artifact. `BorderStyle=1` (outline+shadow, no box) matches the preview's outlined text look. `BorderStyle=3` (opaque box, no outline) is used when the user disables text outline, giving a clean box background.
## Build Verification
- `cargo test` — all 43 tests pass

View File

@@ -0,0 +1,65 @@
# Rolling Two-Line Karaoke Subtitle Display
**Date:** 2026-09-22 00:21
**Task:** Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering
## Problem
The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard `\N` break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.
## Prior Art
- [Sofronio/YouTubeVTT2ASS](https://github.com/Sofronio/YouTubeVTT2ASS) — C# tool that solves this exact problem using `\move` ASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines."
- The key technique: for each spoken line, generate 3 ASS Dialogue entries with `\move` animations (Active → Context → Disappear).
## Changes Made
**File:** `src-tauri/src/services/clip_exporter.rs`
### New structs and functions:
1. **`SpokenLine` struct** — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag.
2. **`extract_spoken_lines()`** — Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation.
3. **`build_karaoke_text()`** — Extracts word timings and builds `\k` karaoke tags for a single line.
4. **`RollingLayout` struct** — Position parameters: `cx` (960), `y_bottom` (1040), `line_height` (font_size * 1.3), `scroll_ms` (350).
### Rewritten `vtt_to_ass_with_karaoke()`:
For each spoken line, generates up to 3 ASS Dialogue entries:
- **Phase 1 (Active/Karaoke):** Line appears at bottom position, scrolls up one slot via `\move(cx, y_bottom, cx, y_bottom-h, 0, 350)`. Has `\k` karaoke tags. Lasts from this line's start to the next line's start.
- **Phase 2 (Context/Static):** Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.
- **Phase 3 (Disappear):** Scrolls off-screen with `\clip` mask to cleanly cut off. Lasts 500ms.
### Edge cases handled:
- **Long gaps (>2s):** Context phase ends early; line disappears instead of lingering through silence/music.
- **Non-speech cues (`[Music]`):** Single static Dialogue with `\pos` instead of rolling.
- **First line:** No context above it — just starts normally.
- **Last line:** Phase 1 uses `line.end_time`; Phase 2/3 use a hold + disappear.
- **Lines without karaoke data:** Rendered as plain text with the same rolling behavior.
### Style additions:
- Added `ScaledBorderAndShadow: yes` to `[Script Info]` for proper scaling.
- Each Dialogue line uses `\an2` override for explicit bottom-center positioning.
## Lessons Learned
1. **YouTube's VTT two-line pattern** is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires `\move` animations, not just `\N` line breaks.
2. **ASS `\move` with `\an2`** — The alignment setting determines the anchor point for positioning. `\an2` (bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas.
3. **Three-phase lifecycle per line** is the key insight from YouTubeVTT2ASS. Phase 3 with `\clip` is important to cleanly mask the text as it scrolls off instead of having it abruptly disappear.
4. **Gap detection** is essential — without it, stale context lines would linger through long silences or `[Music]` sections.
## Build/Test Status
- `cargo build`: Success (6 pre-existing warnings, 0 errors)
- `cargo test`: 43 tests passed, 0 failed

View File

@@ -0,0 +1,62 @@
# Paged Teleprompter Burn-In Subtitles
**Date:** 2026-09-22
**Task:** Replace rolling per-line ASS subtitle generation with a paged teleprompter model for burn-in captions.
## Problem
The previous rolling subtitle implementation generated 3 independent ASS Dialogue events per spoken line (`Active → Context → Disappear`) with `\move` animations. When a new line arrived, both the new and previous lines scrolled simultaneously, creating a "2-line block jump" rather than a natural top-to-bottom reading flow. The active line always reset to the bottom position.
## Solution
Rewrote the ASS generation to use **2-line paged blocks**:
- **Paged model:** Spoken lines are grouped into consecutive pairs. Each page is a single ASS Dialogue event with `\N` (hard line break) between lines.
- **Continuous karaoke:** `\k` tags flow from line 1 through line 2 within the same Dialogue, creating a natural top-to-bottom reading experience (teleprompter style).
- **Cross-fade transitions:** Pages transition via 300ms cross-dissolve using `\fad` tags — no `\move` or `\clip` tags needed.
- **Static positioning:** `\an2\pos(960, 1040)` anchors each page at bottom-center. No animation on position.
- **Gap-based splitting:** If two consecutive lines are >2s apart, they're split into separate single-line pages.
## Changes Made
| File | Change |
|------|--------|
| `src-tauri/src/services/clip_exporter.rs` | Added `group_into_pages()`, `build_page_karaoke_text()`. Rewrote `vtt_to_ass_with_karaoke()`. Removed `RollingLayout` struct and dead `build_karaoke_text()`. Fixed stale comments. |
| `docs/superpowers/specs/2026-09-22-paged-teleprompter-subtitles-design.md` | Design spec |
| `docs/superpowers/plans/2026-09-22-paged-teleprompter-subtitles.md` | Implementation plan |
## Key Implementation Details
### Karaoke Stitching Across Lines
The `build_page_karaoke_text()` function collects word timings from both lines into a flat sequence. The last word of line 1 gets a `\k` duration that extends to line 2's first word start time, naturally covering any gap (silence) between lines. The `\N` is purely visual and doesn't interrupt the karaoke timeline.
### Cross-Fade Timing
| Page Position | `\fad` value |
|---|---|
| First page | `\fad(0, 300)` |
| Middle pages | `\fad(300, 300)` |
| Last page | `\fad(300, 0)` |
Display times are extended/preponed by 300ms to create overlap.
## Commits
- `934c270` — feat(subtitles): add page grouping and cross-line karaoke stitching
- `d13ba27` — feat(subtitles): rewrite ASS generation to paged teleprompter model
- `2b95d8a` — chore: remove dead build_karaoke_text, fix stale comments
## Tests
23 tests passing (8 new + 15 existing).
## Lessons Learned
1. **ASS `\N` doesn't interrupt `\k` flow** — Hard line breaks within a single Dialogue event are purely visual; karaoke timing continues seamlessly across them.
2. **`\fad` > `\move` for transitions** — Static positioning with opacity-only transitions (`\fad`) is much simpler and cleaner than coordinate-based animations, especially for avoiding background-box artifacts.
3. **Trailing `\N` bug recurrence** — The trailing `\\N` in Dialogue format strings was a recurring bug from previous iterations. Must always check for this when writing ASS output.
## Follow-Up
- Manual smoke test needed: export a clip with burn-in captions and verify the teleprompter reading flow.

View File

@@ -0,0 +1,66 @@
# Extended Caption Styling Implementation Summary
**Date:** 2026-09-22 10:37
**Task:** Implement extended caption styling (font selection, drop shadow, dimmed text color, background toggle)
## Changes Made
### Task 1: Data model (commit `a4a1376`)
- **`src/lib/stores/preferences.svelte.ts`**: Extended `CaptionSettings` interface with 8 new fields: `fontFamily`, `backgroundEnabled`, `shadowEnabled`, `shadowDepth`, `shadowColor`, `dimmedColorMode`, `dimmedOpacity`, `dimmedColor`
- **`src-tauri/src/models.rs`**: Extended `CaptionStyle` Rust struct with matching fields (snake_case, serde camelCase)
- **`src/lib/bindings/export.ts`**: Extended TS `CaptionStyle` binding interface
- **`src/lib/components/ExportDialog.svelte`**: Updated captionStyle passthrough to include all 17 fields
### Task 2: System font detection (commit `1447071`)
- **`src-tauri/src/commands/media_analysis.rs`**: Added `list_system_fonts` command using `fc-list :lang=en family` with hardcoded fallback (8 common fonts)
- **`src-tauri/src/lib.rs`**: Registered new command in `generate_handler!`
- **`src/lib/bindings/mediaAnalysis.ts`**: Added `listSystemFonts()` TS binding
### Task 3: ASS generation wiring (commit `2cf8cdc`)
- **`src-tauri/src/services/clip_exporter.rs`**: Rewrote style extraction block in `vtt_to_ass_with_karaoke()`:
- Uses `font_family` instead of hardcoded `Arial`
- Dimmed color: auto mode derives from `textColor` + `dimmedOpacity`, custom mode uses `dimmedColor` directly
- Background/shadow mutual exclusivity: background mode (BorderStyle 3/4 + BackColour for box), shadow mode (BorderStyle 1 + BackColour for shadow + Shadow depth), or neither (transparent BackColour)
### Task 4: Preview caption rendering (commit `bc39c2a`)
- **`src/lib/components/VideoPlayer.svelte`**: Updated `captionStyle` derived to use `$derived.by()`, computing:
- `fontFamily`, `fontWeight` from settings
- Background conditional on `backgroundEnabled` with `rgba()` using `hexToRgb` helper
- Shadow conditional on `shadowEnabled` with depth/color
- Dimmed word colors: auto mode uses opacity, custom mode uses direct color
### Task 5: Settings panel UI (commit `5c47e38`)
- **`src/lib/components/CaptionSettingsPanel.svelte`**: Full rewrite with:
- System font dropdown (loaded on mount via `listSystemFonts()`, each option styled in its own font)
- Font size slider, bold checkbox
- Text color picker
- Dimmed text: auto/custom radio toggle with opacity slider or color picker
- Outline: checkbox + color picker
- Background: checkbox + color picker + opacity slider (disables shadow when enabled)
- Drop shadow: checkbox + depth slider + color picker (disables background when enabled)
- Position: bottom/top radio
- Word highlight checkbox
- Reset to defaults button
- 320px width, scrollable, dark theme
### Task 6: Verification
- `cargo build`: Clean (6 pre-existing warnings)
- `cargo test --lib services::clip_exporter::tests`: 23/23 passing
- `npm run check`: Only 4 pre-existing errors (node:path/process/url, overload)
## Approach
- Used Subagent-Driven Development: fresh subagent per task + task reviewer per task
- 5 implementer dispatches + 5 reviewer dispatches + 1 verification pass
- All reviews approved on first pass (no fix cycles needed)
## Minor Items for Future
- Font-family CSS quoting for multi-word names (e.g., `Times New Roman`) in dropdown options
- No defensive mutual exclusivity check in preview rendering (relies on settings panel enforcement)
- `hexToRgb` has no malformed-hex guard (UI-controlled values only)
- `dimmedColorMode` is `string` in Rust/TS binding vs `'auto' | 'custom'` union in preferences
- No ASS Style-line unit tests for new field combinations
## Follow-up Items
- Manual smoke test with `cargo tauri dev` recommended
- Test font selection with burn-in export
- Verify dimmed color auto vs custom modes in both preview and export

View File

@@ -0,0 +1,858 @@
# Extended Caption Styling Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add system font selection, drop shadow controls, configurable dimmed text color, and background toggle to the shared caption settings, wired through to both preview and burn-in export.
**Architecture:** Extend `CaptionSettings` (frontend) and `CaptionStyle` (backend) with new fields. Add a `list_system_fonts` Tauri command using `fc-list`. Update the settings panel UI, the ASS generation code, and the preview caption rendering.
**Tech Stack:** Svelte 5, Rust/Tauri v2, ASS subtitle format, fontconfig (`fc-list`)
## Global Constraints
- Background and shadow are mutually exclusive: enabling one disables the other
- `fc-list` fallback: if unavailable, use preset list: Arial, Helvetica, Verdana, Georgia, Times New Roman, Courier New, Impact
- Dimmed color auto mode derives from `textColor` + `dimmedOpacity`; custom mode uses `dimmedColor` directly
- ASS `BackColour` serves as background color (when background ON) or shadow color (when shadow ON)
- All new fields must have backwards-compatible defaults (existing saved preferences without new fields should merge cleanly via `{ ...DEFAULT_CAPTION_SETTINGS, ...savedCaptions }`)
- No new crate dependencies
---
### Task 1: Data model — add new fields to frontend and backend
**Files:**
- Modify: `src/lib/stores/preferences.svelte.ts:6-27`
- Modify: `src-tauri/src/models.rs:91-108`
- Modify: `src/lib/bindings/export.ts:3-10`
- Modify: `src/lib/components/ExportDialog.svelte:86-94`
**Interfaces:**
- Consumes: Existing `CaptionSettings`, `CaptionStyle`, `CaptionStyle` (TS binding)
- Produces: Extended versions of all three with new fields. All later tasks depend on these types.
- [ ] **Step 1: Update `CaptionSettings` interface and defaults**
In `src/lib/stores/preferences.svelte.ts`, replace the interface and defaults:
```typescript
export interface CaptionSettings {
fontFamily: string;
fontSize: number;
textColor: string;
bold: boolean;
outlineColor: string;
backgroundEnabled: boolean;
backgroundOpacity: number;
backgroundColor: string;
textOutline: boolean;
shadowEnabled: boolean;
shadowDepth: number;
shadowColor: string;
dimmedColorMode: 'auto' | 'custom';
dimmedOpacity: number;
dimmedColor: string;
position: 'bottom' | 'top';
wordHighlight: boolean;
}
export const DEFAULT_CAPTION_SETTINGS: CaptionSettings = {
fontFamily: 'Arial',
fontSize: 18,
textColor: '#ffffff',
bold: true,
outlineColor: '#000000',
backgroundEnabled: true,
backgroundOpacity: 0.75,
backgroundColor: '#000000',
textOutline: true,
shadowEnabled: false,
shadowDepth: 2,
shadowColor: '#000000',
dimmedColorMode: 'auto',
dimmedOpacity: 0.4,
dimmedColor: '#999999',
position: 'bottom',
wordHighlight: true,
};
```
- [ ] **Step 2: Update `CaptionStyle` struct in Rust**
In `src-tauri/src/models.rs`, replace the struct:
```rust
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
pub struct CaptionStyle {
pub font_family: String,
pub font_size: u32,
pub text_color: String,
pub bold: bool,
pub outline_color: String,
pub background_enabled: bool,
pub background_opacity: f64,
pub background_color: String,
pub text_outline: bool,
pub shadow_enabled: bool,
pub shadow_depth: u32,
pub shadow_color: String,
pub dimmed_color_mode: String,
pub dimmed_opacity: f64,
pub dimmed_color: String,
pub position: String,
pub word_highlight: bool,
}
```
- [ ] **Step 3: Update TypeScript `CaptionStyle` binding**
In `src/lib/bindings/export.ts`, replace the interface:
```typescript
export interface CaptionStyle {
fontFamily: string;
fontSize: number;
textColor: string;
bold: boolean;
outlineColor: string;
backgroundEnabled: boolean;
backgroundOpacity: number;
backgroundColor: string;
textOutline: boolean;
shadowEnabled: boolean;
shadowDepth: number;
shadowColor: string;
dimmedColorMode: string;
dimmedOpacity: number;
dimmedColor: string;
position: string;
wordHighlight: boolean;
}
```
- [ ] **Step 4: Update ExportDialog passthrough**
In `src/lib/components/ExportDialog.svelte`, replace the `captionStyle` construction:
```typescript
captionStyle: (burnInCaptions && includeCaptions && hasCaptions)
? {
fontFamily: preferences.captionSettings.fontFamily,
fontSize: preferences.captionSettings.fontSize,
textColor: preferences.captionSettings.textColor,
bold: preferences.captionSettings.bold,
outlineColor: preferences.captionSettings.outlineColor,
backgroundEnabled: preferences.captionSettings.backgroundEnabled,
backgroundOpacity: preferences.captionSettings.backgroundOpacity,
backgroundColor: preferences.captionSettings.backgroundColor,
textOutline: preferences.captionSettings.textOutline,
shadowEnabled: preferences.captionSettings.shadowEnabled,
shadowDepth: preferences.captionSettings.shadowDepth,
shadowColor: preferences.captionSettings.shadowColor,
dimmedColorMode: preferences.captionSettings.dimmedColorMode,
dimmedOpacity: preferences.captionSettings.dimmedOpacity,
dimmedColor: preferences.captionSettings.dimmedColor,
position: preferences.captionSettings.position,
wordHighlight: preferences.captionSettings.wordHighlight,
}
: null,
```
- [ ] **Step 5: Build and test**
Run: `cd src-tauri && cargo build 2>&1 | tail -5`
Expected: Clean build (warnings OK if pre-existing).
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -5`
Expected: All 23 tests pass (they don't construct CaptionStyle directly).
- [ ] **Step 6: Commit**
```bash
git add src/lib/stores/preferences.svelte.ts src-tauri/src/models.rs src/lib/bindings/export.ts src/lib/components/ExportDialog.svelte
git commit -m "feat(captions): extend data model with font, shadow, dimmed color, bg toggle fields"
```
---
### Task 2: System font detection — add `list_system_fonts` command
**Files:**
- Modify: `src-tauri/src/commands/media_analysis.rs` (add new command)
- Modify: `src-tauri/src/lib.rs:31-46` (register command)
- Modify: `src/lib/bindings/mediaAnalysis.ts` (add TS binding)
**Interfaces:**
- Consumes: `fc-list` CLI tool (from fontconfig)
- Produces: `list_system_fonts() -> Result<Vec<String>, String>` (Rust), `listSystemFonts(): Promise<string[]>` (TS)
- [ ] **Step 1: Add the Rust command**
Append to `src-tauri/src/commands/media_analysis.rs`:
```rust
#[tauri::command]
pub async fn list_system_fonts() -> Result<Vec<String>, String> {
tokio::task::spawn_blocking(|| {
// Try fc-list first (available if fontconfig is installed, which ffmpeg depends on)
let output = std::process::Command::new("fc-list")
.args([":lang=en", "family"])
.output();
match output {
Ok(out) if out.status.success() => {
let text = String::from_utf8_lossy(&out.stdout);
let mut families: Vec<String> = text
.lines()
.flat_map(|line| line.split(','))
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
.collect();
families.sort_unstable();
families.dedup();
Ok(families)
}
_ => {
// Fallback: common cross-platform fonts
Ok(vec![
"Arial".to_string(),
"Comic Sans MS".to_string(),
"Courier New".to_string(),
"Georgia".to_string(),
"Helvetica".to_string(),
"Impact".to_string(),
"Times New Roman".to_string(),
"Verdana".to_string(),
])
}
}
})
.await
.map_err(|e| format!("Font detection failed: {e}"))?
}
```
- [ ] **Step 2: Register the command in `lib.rs`**
In `src-tauri/src/lib.rs`, add `media_analysis::list_system_fonts` to the `generate_handler!` macro, after `media_analysis::check_embedded_subtitles`:
```rust
media_analysis::check_embedded_subtitles,
media_analysis::list_system_fonts,
```
- [ ] **Step 3: Add the TS binding**
Append to `src/lib/bindings/mediaAnalysis.ts`:
```typescript
export async function listSystemFonts(): Promise<string[]> {
return invoke<string[]>('list_system_fonts');
}
```
- [ ] **Step 4: Build and test**
Run: `cd src-tauri && cargo build 2>&1 | tail -5`
Expected: Clean build.
- [ ] **Step 5: Commit**
```bash
git add src-tauri/src/commands/media_analysis.rs src-tauri/src/lib.rs src/lib/bindings/mediaAnalysis.ts
git commit -m "feat(captions): add list_system_fonts command with fc-list + fallback"
```
---
### Task 3: ASS generation — wire all new fields into burn-in subtitle output
**Files:**
- Modify: `src-tauri/src/services/clip_exporter.rs:549-584`
**Interfaces:**
- Consumes: `CaptionStyle` struct from Task 1 (with all new fields)
- Produces: Updated `vtt_to_ass_with_karaoke()` that uses font family, background/shadow mutual exclusivity, and dimmed color settings
- [ ] **Step 1: Replace the style extraction block in `vtt_to_ass_with_karaoke`**
In `src-tauri/src/services/clip_exporter.rs`, replace lines 550-566 (the style extraction block) with:
```rust
// Font
let font_name = style.map_or("Arial".to_string(), |s| s.font_family.clone());
let font_size = style.map_or(50, |s| (s.font_size as f64 * 2.75).round() as u32);
let bold_flag: i32 = if style.map_or(true, |s| s.bold) { -1 } else { 0 };
// Colors
let primary_colour = style.map_or("&H00FFFFFF".to_string(), |s| hex_to_ass_color(&s.text_color));
let outline_colour = style.map_or("&H00000000".to_string(), |s| hex_to_ass_color(&s.outline_color));
// Dimmed color (SecondaryColour) — auto derives from text color, custom uses direct color
let secondary_colour = style.map_or("&H73CCCCCC".to_string(), |s| {
if s.dimmed_color_mode == "custom" {
hex_to_ass_color(&s.dimmed_color)
} else {
let alpha = ((1.0 - s.dimmed_opacity.clamp(0.0, 1.0)) * 255.0).round() as u8;
hex_to_ass_color_with_alpha(&s.text_color, alpha)
}
});
// Background/Shadow mutual exclusivity
let bg_enabled = style.map_or(true, |s| s.background_enabled);
let shadow_enabled = style.map_or(false, |s| s.shadow_enabled);
let text_outline = style.map_or(true, |s| s.text_outline);
let (border_style, outline_val, shadow_val, back_colour) = if bg_enabled {
// Background mode: box + optional outline, no shadow
let bc = style.map_or("&H80000000".to_string(), |s| {
let alpha = ((1.0 - s.background_opacity.clamp(0.0, 1.0)) * 255.0).round() as u8;
hex_to_ass_color_with_alpha(&s.background_color, alpha)
});
let bs = if text_outline { 4 } else { 3 };
let ol = if text_outline { 2 } else { 0 };
(bs, ol, 0, bc)
} else if shadow_enabled {
// Shadow mode: outline + shadow, no box
let sc = style.map_or("&H00000000".to_string(), |s| hex_to_ass_color(&s.shadow_color));
let sd = style.map_or(2, |s| s.shadow_depth) as i32;
let ol = if text_outline { 2 } else { 0 };
(1, ol, sd, sc)
} else {
// Neither: just outline if enabled
let ol = if text_outline { 2 } else { 0 };
(1, ol, 0, "&HFF000000".to_string()) // fully transparent back
};
let word_highlight = style.map_or(true, |s| s.word_highlight);
let margin_v: i32 = 40;
```
- [ ] **Step 2: Update the Style format string to use `font_name`**
Replace the Style line (currently around line 580-584):
```rust
output.push_str(&format!(
"Style: Default,{},{},{},{},{},{},{},0,0,0,100,100,0,0,{},{},{},2,20,20,{},1\n",
font_name, font_size, primary_colour, secondary_colour, outline_colour, back_colour,
bold_flag, border_style, outline_val, shadow_val, margin_v
));
```
- [ ] **Step 3: Run tests**
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -10`
Expected: All 23 tests pass.
Run: `cd src-tauri && cargo build 2>&1 | tail -5`
Expected: Clean build.
- [ ] **Step 4: Commit**
```bash
git add src-tauri/src/services/clip_exporter.rs
git commit -m "feat(captions): wire font, shadow, dimmed color, bg toggle into ASS generation"
```
---
### Task 4: Preview captions — use new settings in VideoPlayer
**Files:**
- Modify: `src/lib/components/VideoPlayer.svelte:29-40, 184-189`
**Interfaces:**
- Consumes: `CaptionSettings` from Task 1 (via `preferences.captionSettings`)
- Produces: Updated preview caption rendering using font, background toggle, shadow, and dimmed color settings
- [ ] **Step 1: Update the `captionStyle` derived**
In `src/lib/components/VideoPlayer.svelte`, replace the `captionStyle` derived block (lines 29-36):
```typescript
let captionStyle = $derived.by(() => {
const s = preferences.captionSettings;
const bg = s.backgroundEnabled
? `rgba(${hexToRgb(s.backgroundColor)}, ${s.backgroundOpacity})`
: 'transparent';
const shadow = s.shadowEnabled
? `${s.shadowDepth}px ${s.shadowDepth}px ${s.shadowDepth * 2}px ${s.shadowColor}`
: s.textOutline
? '1px 1px 2px rgba(0,0,0,0.9), -1px -1px 2px rgba(0,0,0,0.9)'
: 'none';
return {
fontFamily: s.fontFamily,
fontSize: `${s.fontSize}px`,
fontWeight: s.bold ? 'bold' : 'normal',
color: s.textColor,
background: bg,
textShadow: shadow,
};
});
```
- [ ] **Step 2: Add the `hexToRgb` helper**
Add this helper function at the top of the `<script>` block (after imports):
```typescript
function hexToRgb(hex: string): string {
const h = hex.replace('#', '');
const r = parseInt(h.substring(0, 2), 16);
const g = parseInt(h.substring(2, 4), 16);
const b = parseInt(h.substring(4, 6), 16);
return `${r}, ${g}, ${b}`;
}
```
- [ ] **Step 3: Update the dimmed word color in the template**
In the caption rendering section (around line 184-189), update the unspoken word styling:
Replace:
```svelte
style="color: {word.spoken ? captionStyle.color : captionStyle.color}; opacity: {word.spoken ? 1 : 0.4};"
```
With:
```svelte
style="color: {word.spoken
? captionStyle.color
: preferences.captionSettings.dimmedColorMode === 'custom'
? preferences.captionSettings.dimmedColor
: captionStyle.color}; opacity: {word.spoken ? 1 : (preferences.captionSettings.dimmedColorMode === 'custom' ? 1 : preferences.captionSettings.dimmedOpacity)};"
```
- [ ] **Step 4: Update template references to use new properties**
`$derived.by` still produces a reactive value, so all template references remain `captionStyle.X` (unchanged from before). Add `font-family` and `font-weight` to the inline styles on `.caption-text`:
Replace the `style` attribute on `.caption-text`:
```svelte
style="font-family: {captionStyle.fontFamily}; font-size: {captionStyle.fontSize}; font-weight: {captionStyle.fontWeight}; background: {captionStyle.background}; text-shadow: {captionStyle.textShadow};"
```
And the non-karaoke span:
```svelte
<span style="color: {captionStyle.color};">{cue.text}</span>
```
- [ ] **Step 5: Commit**
```bash
git add src/lib/components/VideoPlayer.svelte
git commit -m "feat(captions): use font, shadow, dimmed color, bg toggle in preview captions"
```
---
### Task 5: Settings panel UI — add all new controls
**Files:**
- Modify: `src/lib/components/CaptionSettingsPanel.svelte`
**Interfaces:**
- Consumes: `CaptionSettings` from Task 1, `listSystemFonts()` from Task 2
- Produces: Full settings panel with all controls
- [ ] **Step 1: Rewrite the panel component**
Replace the entire contents of `src/lib/components/CaptionSettingsPanel.svelte` with the expanded panel. The component should:
1. Import `listSystemFonts` from `$lib/bindings/mediaAnalysis`
2. Call `listSystemFonts()` on mount, store result in `let systemFonts = $state<string[]>(['Arial'])`
3. Use `onMount` to trigger the font loading
4. Layout the controls in this order:
- **Font**: `<select>` dropdown with system fonts, each `<option>` styled with `font-family: {font}`. Font size slider. Bold checkbox.
- **Text Color**: Color picker
- **Dimmed Text**: Auto/Custom toggle (radio). When auto: opacity slider (0.1-0.9). When custom: color picker.
- **Outline**: Checkbox + color picker (picker disabled when unchecked)
- **Background**: Checkbox + color picker + opacity slider. When enabled, sets `shadowEnabled: false`.
- **Shadow**: Checkbox + depth slider (1-5) + color picker. When enabled, sets `backgroundEnabled: false`.
- **Position**: Bottom/Top radio
- **Word Highlight**: Checkbox
- **Reset to defaults** button
The panel width stays at 320px. Max-height with overflow-y scroll for smaller screens.
Full component code:
```svelte
<script lang="ts">
import { onMount } from 'svelte';
import {
preferences,
setCaptionSettings,
DEFAULT_CAPTION_SETTINGS,
type CaptionSettings,
} from '$lib/stores/preferences.svelte';
import { listSystemFonts } from '$lib/bindings/mediaAnalysis';
let { onClose }: { onClose: () => void } = $props();
let settings = $derived(preferences.captionSettings);
let systemFonts = $state<string[]>(['Arial']);
onMount(async () => {
try {
systemFonts = await listSystemFonts();
} catch {
systemFonts = ['Arial', 'Helvetica', 'Verdana', 'Georgia', 'Times New Roman', 'Courier New', 'Impact'];
}
});
function update(partial: Partial<CaptionSettings>) {
// Enforce mutual exclusivity
if (partial.backgroundEnabled) {
partial.shadowEnabled = false;
} else if (partial.shadowEnabled) {
partial.backgroundEnabled = false;
}
setCaptionSettings(partial);
}
function resetToDefaults() {
setCaptionSettings({ ...DEFAULT_CAPTION_SETTINGS });
}
</script>
<div class="panel">
<div class="panel-header">
<h3>Caption Appearance</h3>
<button class="close-btn" onclick={onClose} title="Close">x</button>
</div>
<div class="scroll-area">
<!-- Font Family -->
<div class="control">
<label>Font</label>
<select
value={settings.fontFamily}
onchange={(e) => update({ fontFamily: (e.target as HTMLSelectElement).value })}
>
{#each systemFonts as font}
<option value={font} style="font-family: {font}">{font}</option>
{/each}
</select>
</div>
<!-- Font Size -->
<div class="control">
<label>
Font Size
<span class="value">{settings.fontSize}px</span>
</label>
<input
type="range" min="12" max="36" step="1"
value={settings.fontSize}
oninput={(e) => update({ fontSize: parseInt((e.target as HTMLInputElement).value, 10) })}
/>
</div>
<!-- Text Color + Bold -->
<div class="control row">
<label>Text Color</label>
<div class="row-controls">
<input
type="color" value={settings.textColor}
oninput={(e) => update({ textColor: (e.target as HTMLInputElement).value })}
/>
<label class="inline-check">
<input type="checkbox" checked={settings.bold}
onchange={(e) => update({ bold: (e.target as HTMLInputElement).checked })} />
Bold
</label>
</div>
</div>
<!-- Dimmed Text Color -->
<div class="control">
<label>Dimmed Text</label>
<div class="radio-row">
<label>
<input type="radio" name="dimmed-mode" value="auto"
checked={settings.dimmedColorMode === 'auto'}
onchange={() => update({ dimmedColorMode: 'auto' })} />
Auto
</label>
<label>
<input type="radio" name="dimmed-mode" value="custom"
checked={settings.dimmedColorMode === 'custom'}
onchange={() => update({ dimmedColorMode: 'custom' })} />
Custom
</label>
</div>
</div>
{#if settings.dimmedColorMode === 'auto'}
<div class="control">
<label>
Dim Opacity
<span class="value">{Math.round(settings.dimmedOpacity * 100)}%</span>
</label>
<input
type="range" min="10" max="90" step="5"
value={Math.round(settings.dimmedOpacity * 100)}
oninput={(e) => update({ dimmedOpacity: parseInt((e.target as HTMLInputElement).value, 10) / 100 })}
/>
</div>
{:else}
<div class="control row">
<label>Dimmed Color</label>
<div class="row-controls">
<input
type="color" value={settings.dimmedColor}
oninput={(e) => update({ dimmedColor: (e.target as HTMLInputElement).value })}
/>
</div>
</div>
{/if}
<!-- Outline -->
<div class="control row">
<label>
<input type="checkbox" checked={settings.textOutline}
onchange={(e) => update({ textOutline: (e.target as HTMLInputElement).checked })} />
Outline
</label>
<div class="row-controls">
<input
type="color" value={settings.outlineColor}
disabled={!settings.textOutline}
class:disabled={!settings.textOutline}
oninput={(e) => update({ outlineColor: (e.target as HTMLInputElement).value })}
/>
</div>
</div>
<!-- Background (mutually exclusive with Shadow) -->
<div class="control row">
<label>
<input type="checkbox" checked={settings.backgroundEnabled}
onchange={(e) => update({ backgroundEnabled: (e.target as HTMLInputElement).checked })} />
Background
</label>
<div class="row-controls">
<input
type="color" value={settings.backgroundColor}
disabled={!settings.backgroundEnabled}
class:disabled={!settings.backgroundEnabled}
oninput={(e) => update({ backgroundColor: (e.target as HTMLInputElement).value })}
/>
</div>
</div>
{#if settings.backgroundEnabled}
<div class="control">
<label>
BG Opacity
<span class="value">{Math.round(settings.backgroundOpacity * 100)}%</span>
</label>
<input
type="range" min="0" max="100" step="5"
value={Math.round(settings.backgroundOpacity * 100)}
oninput={(e) => update({ backgroundOpacity: parseInt((e.target as HTMLInputElement).value, 10) / 100 })}
/>
</div>
{/if}
<!-- Shadow (mutually exclusive with Background) -->
<div class="control row">
<label>
<input type="checkbox" checked={settings.shadowEnabled}
onchange={(e) => update({ shadowEnabled: (e.target as HTMLInputElement).checked })} />
Drop Shadow
</label>
<div class="row-controls">
<input
type="color" value={settings.shadowColor}
disabled={!settings.shadowEnabled}
class:disabled={!settings.shadowEnabled}
oninput={(e) => update({ shadowColor: (e.target as HTMLInputElement).value })}
/>
</div>
</div>
{#if settings.shadowEnabled}
<div class="control">
<label>
Shadow Depth
<span class="value">{settings.shadowDepth}px</span>
</label>
<input
type="range" min="1" max="5" step="1"
value={settings.shadowDepth}
oninput={(e) => update({ shadowDepth: parseInt((e.target as HTMLInputElement).value, 10) })}
/>
</div>
{/if}
<!-- Position -->
<div class="control">
<label>Position</label>
<div class="radio-row">
<label>
<input type="radio" name="caption-position" value="bottom"
checked={settings.position === 'bottom'}
onchange={() => update({ position: 'bottom' })} />
Bottom
</label>
<label>
<input type="radio" name="caption-position" value="top"
checked={settings.position === 'top'}
onchange={() => update({ position: 'top' })} />
Top
</label>
</div>
</div>
<!-- Word Highlight -->
<div class="control">
<label>
<input type="checkbox" checked={settings.wordHighlight}
onchange={(e) => update({ wordHighlight: (e.target as HTMLInputElement).checked })} />
Word-by-word highlight
</label>
</div>
</div>
<div class="panel-footer">
<button class="reset-btn" onclick={resetToDefaults}>Reset to defaults</button>
</div>
</div>
<style>
.panel {
position: absolute;
bottom: 44px;
right: 12px;
background: var(--bg-primary, #1e1e2e);
border: 1px solid var(--border, #353550);
border-radius: 8px;
padding: 12px 16px;
width: 320px;
max-height: 480px;
display: flex;
flex-direction: column;
z-index: 20;
box-shadow: 0 8px 24px rgba(0, 0, 0, 0.5);
}
.panel-header {
display: flex;
justify-content: space-between;
align-items: center;
margin-bottom: 8px;
flex-shrink: 0;
}
h3 { margin: 0; font-size: 13px; font-weight: 600; color: var(--text-primary, #cdd6f4); }
.close-btn {
background: none; border: none; color: var(--text-secondary, #6c7086);
cursor: pointer; font-size: 14px; padding: 2px 6px; border-radius: 3px;
}
.close-btn:hover { background: var(--bg-tertiary, #2a2a3e); color: var(--text-primary, #cdd6f4); }
.scroll-area {
overflow-y: auto;
flex: 1;
min-height: 0;
padding-right: 4px;
}
.control { margin-bottom: 10px; }
.control label {
display: flex; align-items: center; gap: 6px;
font-size: 12px; color: var(--text-secondary, #6c7086); margin-bottom: 4px;
}
.control.row { display: flex; align-items: center; justify-content: space-between; }
.control.row > label { margin-bottom: 0; }
.row-controls { display: flex; align-items: center; gap: 8px; }
.inline-check {
display: flex; align-items: center; gap: 4px;
font-size: 12px; color: var(--text-secondary, #6c7086);
margin-bottom: 0 !important; cursor: pointer;
}
.value {
margin-left: auto; font-size: 11px;
color: var(--text-primary, #cdd6f4); font-variant-numeric: tabular-nums;
}
select {
width: 100%; padding: 4px 6px; font-size: 12px;
background: var(--bg-secondary, #252535); color: var(--text-primary, #cdd6f4);
border: 1px solid var(--border, #353550); border-radius: 4px;
}
input[type='range'] { width: 100%; accent-color: #89b4fa; }
input[type='color'] {
width: 32px; height: 24px; padding: 0;
border: 1px solid var(--border, #353550); border-radius: 3px;
cursor: pointer; background: none;
}
input[type='color'].disabled { opacity: 0.3; cursor: not-allowed; }
input[type='checkbox'] { accent-color: #89b4fa; }
input[type='radio'] { accent-color: #89b4fa; }
.radio-row { display: flex; gap: 16px; }
.radio-row label { font-size: 12px; color: var(--text-secondary, #6c7086); cursor: pointer; }
.panel-footer {
margin-top: 8px; padding-top: 8px; flex-shrink: 0;
border-top: 1px solid var(--border, #353550);
}
.reset-btn {
font-size: 11px; color: var(--text-secondary, #6c7086);
background: none; border: none; cursor: pointer; padding: 2px 0;
}
.reset-btn:hover { color: var(--text-primary, #cdd6f4); }
</style>
```
- [ ] **Step 2: Build check**
Run: `npm run check 2>&1 | grep "Error:" | head -10`
Expected: Only pre-existing errors (node:path, node:process, node:url, overload). No new errors.
- [ ] **Step 3: Commit**
```bash
git add src/lib/components/CaptionSettingsPanel.svelte
git commit -m "feat(captions): expand settings panel with font, shadow, dimmed color, bg toggle controls"
```
---
### Task 6: Build, test, and verify
**Files:**
- No new changes (verification only)
**Interfaces:**
- Consumes: All changes from Tasks 1-5
- [ ] **Step 1: Full Rust build and test**
Run: `cd src-tauri && cargo build 2>&1 | tail -5`
Expected: Clean build.
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -10`
Expected: All tests pass.
- [ ] **Step 2: Frontend type check**
Run: `npm run check 2>&1 | grep "Error:" | head -10`
Expected: Only pre-existing errors. No new errors from caption changes.
- [ ] **Step 3: Manual smoke test**
Run: `cargo tauri dev`
Verify:
1. Caption settings panel opens, shows font dropdown (populated with system fonts)
2. Selecting a different font updates the preview captions
3. Enabling shadow disables background, and vice versa
4. Dimmed color auto mode shows opacity slider; custom mode shows color picker
5. Export a short clip with burn-in and verify the ASS uses the selected font and style settings

View File

@@ -0,0 +1,604 @@
# Paged Teleprompter Burn-In Subtitles Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Replace the rolling per-line ASS subtitle generation with a paged teleprompter model where 2-line blocks display with continuous karaoke highlighting and cross-fade transitions.
**Architecture:** Rewrite `vtt_to_ass_with_karaoke()` in `clip_exporter.rs` to group `SpokenLine`s into 2-line pages, emit one ASS Dialogue per page with stitched `\k` tags across `\N` line breaks, and use `\fad` for cross-fade transitions. No `\move` or `\clip` tags. Single-file change.
**Tech Stack:** Rust, ASS subtitle format (libass), FFmpeg `ass=` filter
## Global Constraints
- All changes are in `src-tauri/src/services/clip_exporter.rs` only
- No new dependencies
- Existing tests must continue to pass (they test FFmpeg arg construction, not ASS content)
- ASS header/style generation stays unchanged
- `extract_spoken_lines()` stays unchanged
- `parse_vtt_word_timings()` stays unchanged
- The `ass=` filter usage in `build_ffmpeg_args_burnin_subs()` is unchanged
- Cross-fade duration: 300ms
- Gap threshold for page splitting: 2.0 seconds
- PlayRes: 1920×1080
---
### Task 1: Add `build_page_karaoke_text` and page-grouping logic
**Files:**
- Modify: `src-tauri/src/services/clip_exporter.rs:432-460` (replace `RollingLayout`, add new function)
**Interfaces:**
- Consumes: `SpokenLine` struct (unchanged, line 349), `parse_vtt_word_timings()` (unchanged, line 259), `strip_vtt_tags()` (unchanged, line 170)
- Produces: `fn build_page_karaoke_text(lines: &[&SpokenLine]) -> String` — returns ASS text with `\k` tags and `\N` separator for 1-2 line pages. `fn group_into_pages(spoken_lines: &[SpokenLine], gap_threshold: f64) -> Vec<Vec<usize>>` — returns groups of indices into spoken_lines.
- [ ] **Step 1: Write failing tests for `build_page_karaoke_text`**
Add these tests to the existing `#[cfg(test)] mod tests` block at line 1155:
```rust
#[test]
fn test_build_page_karaoke_single_line_with_karaoke() {
let line = SpokenLine {
plain_text: "hello world".to_string(),
raw_text: "hello<00:00:01.000><c> world</c>".to_string(),
start_time: 0.5,
end_time: 1.5,
has_karaoke: true,
is_non_speech: false,
};
let result = build_page_karaoke_text(&[&line]);
// "hello" starts at 0.5, "world" starts at 1.0, ends at 1.5
// hello duration = 1.0 - 0.5 = 0.5s = 50cs
// world duration = 1.5 - 1.0 = 0.5s = 50cs
assert!(result.contains("{\\k50}hello"));
assert!(result.contains("{\\k50}world"));
assert!(!result.contains("\\N"));
}
#[test]
fn test_build_page_karaoke_two_lines_stitched() {
let line1 = SpokenLine {
plain_text: "hello world".to_string(),
raw_text: "hello<00:00:01.000><c> world</c>".to_string(),
start_time: 0.5,
end_time: 1.5,
has_karaoke: true,
is_non_speech: false,
};
let line2 = SpokenLine {
plain_text: "foo bar".to_string(),
raw_text: "foo<00:00:02.500><c> bar</c>".to_string(),
start_time: 2.0,
end_time: 3.0,
has_karaoke: true,
is_non_speech: false,
};
let result = build_page_karaoke_text(&[&line1, &line2]);
// Last word of line1 ("world") should span from 1.0 to line2 first word (2.0) = 100cs
assert!(result.contains("\\N"));
assert!(result.contains("{\\k100}world"));
// line2: "foo" at 2.0, "bar" at 2.5, end at 3.0
assert!(result.contains("{\\k50}foo"));
assert!(result.contains("{\\k50}bar"));
}
#[test]
fn test_build_page_karaoke_no_karaoke_tags() {
let line = SpokenLine {
plain_text: "just plain text".to_string(),
raw_text: "just plain text".to_string(),
start_time: 1.0,
end_time: 3.0,
has_karaoke: false,
is_non_speech: false,
};
let result = build_page_karaoke_text(&[&line]);
// Single \k covering the full 2.0s = 200cs
assert!(result.contains("{\\k200}just plain text"));
}
#[test]
fn test_group_into_pages_even() {
let lines = vec![
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 1.0, end_time: 2.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "c".into(), raw_text: "c".into(), start_time: 2.0, end_time: 3.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "d".into(), raw_text: "d".into(), start_time: 3.0, end_time: 4.0, has_karaoke: false, is_non_speech: false },
];
let pages = group_into_pages(&lines, 2.0);
assert_eq!(pages.len(), 2);
assert_eq!(pages[0], vec![0, 1]);
assert_eq!(pages[1], vec![2, 3]);
}
#[test]
fn test_group_into_pages_gap_splits() {
let lines = vec![
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 5.0, end_time: 6.0, has_karaoke: false, is_non_speech: false },
];
// Gap between a (end=1.0) and b (start=5.0) is 4.0s > 2.0s threshold
let pages = group_into_pages(&lines, 2.0);
assert_eq!(pages.len(), 2);
assert_eq!(pages[0], vec![0]);
assert_eq!(pages[1], vec![1]);
}
#[test]
fn test_group_into_pages_odd_count() {
let lines = vec![
SpokenLine { plain_text: "a".into(), raw_text: "a".into(), start_time: 0.0, end_time: 1.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "b".into(), raw_text: "b".into(), start_time: 1.0, end_time: 2.0, has_karaoke: false, is_non_speech: false },
SpokenLine { plain_text: "c".into(), raw_text: "c".into(), start_time: 2.0, end_time: 3.0, has_karaoke: false, is_non_speech: false },
];
let pages = group_into_pages(&lines, 2.0);
assert_eq!(pages.len(), 2);
assert_eq!(pages[0], vec![0, 1]);
assert_eq!(pages[1], vec![2]);
}
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | head -40`
Expected: Compilation errors — `build_page_karaoke_text` and `group_into_pages` not found.
- [ ] **Step 3: Implement `group_into_pages`**
Add this function after `build_karaoke_text` (around line 450), replacing the `RollingLayout` struct (lines 452-458):
```rust
/// Group spoken lines into 2-line pages for the teleprompter display.
/// If the gap between two consecutive lines exceeds `gap_threshold` seconds,
/// the pair is split into separate single-line pages.
/// Non-speech lines (e.g., [Music]) always get their own page.
fn group_into_pages(spoken_lines: &[SpokenLine], gap_threshold: f64) -> Vec<Vec<usize>> {
let mut pages: Vec<Vec<usize>> = Vec::new();
let mut i = 0;
while i < spoken_lines.len() {
let line = &spoken_lines[i];
// Non-speech cues always get their own page
if line.is_non_speech {
pages.push(vec![i]);
i += 1;
continue;
}
// Try to pair with the next line
if i + 1 < spoken_lines.len() {
let next = &spoken_lines[i + 1];
let gap = next.start_time - line.end_time;
// Pair them if gap is small enough and next isn't non-speech
if gap <= gap_threshold && !next.is_non_speech {
pages.push(vec![i, i + 1]);
i += 2;
continue;
}
}
// Solo page (last line, or gap too large, or next is non-speech)
pages.push(vec![i]);
i += 1;
}
pages
}
```
- [ ] **Step 4: Implement `build_page_karaoke_text`**
Add this function right after `group_into_pages`:
```rust
/// Build karaoke text for a 1-or-2-line page, stitching word timings
/// across lines with \N as the visual line break. The \k durations flow
/// continuously so karaoke highlighting progresses top-to-bottom.
fn build_page_karaoke_text(lines: &[&SpokenLine]) -> String {
// Collect all (word, absolute_start_time) across all lines in page order.
// For the cross-line stitch, the last word of line N extends to the first
// word of line N+1.
struct WordEntry {
text: String,
start: f64,
is_line_break_before: bool, // insert \N before this word
}
let mut entries: Vec<WordEntry> = Vec::new();
for (line_idx, line) in lines.iter().enumerate() {
let is_new_line = line_idx > 0;
if line.has_karaoke {
let words = parse_vtt_word_timings(&line.raw_text, line.start_time);
if words.len() >= 2 {
for (w_idx, (word, start)) in words.iter().enumerate() {
entries.push(WordEntry {
text: word.clone(),
start: *start,
is_line_break_before: is_new_line && w_idx == 0,
});
}
} else {
// Fallback: treat entire line as one word
entries.push(WordEntry {
text: strip_vtt_tags(&line.raw_text),
start: line.start_time,
is_line_break_before: is_new_line,
});
}
} else {
// No karaoke data — one entry for the whole line
entries.push(WordEntry {
text: line.plain_text.clone(),
start: line.start_time,
is_line_break_before: is_new_line,
});
}
}
if entries.is_empty() {
return String::new();
}
// The page's end time is the last line's end_time
let page_end = lines.last().unwrap().end_time;
// Build the output with \k tags
let mut parts: Vec<String> = Vec::new();
for (i, entry) in entries.iter().enumerate() {
let next_start = if i + 1 < entries.len() {
entries[i + 1].start
} else {
page_end
};
let duration_cs = ((next_start - entry.start) * 100.0).round().max(1.0) as u64;
let prefix = if entry.is_line_break_before { "\\N" } else if i > 0 { " " } else { "" };
parts.push(format!("{}{{\\k{}}}{}", prefix, duration_cs, entry.text));
}
parts.join("")
}
```
- [ ] **Step 5: Run tests to verify they pass**
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -20`
Expected: All new tests pass. All existing tests still pass.
- [ ] **Step 6: Commit**
```bash
git add src-tauri/src/services/clip_exporter.rs
git commit -m "feat(subtitles): add page grouping and cross-line karaoke stitching"
```
---
### Task 2: Rewrite `vtt_to_ass_with_karaoke` to use the paged model
**Files:**
- Modify: `src-tauri/src/services/clip_exporter.rs:465-643` (the `vtt_to_ass_with_karaoke` function body)
**Interfaces:**
- Consumes: `build_page_karaoke_text()` and `group_into_pages()` from Task 1, `extract_spoken_lines()` (unchanged, line 362), `format_ass_timestamp()` (unchanged, line 244), style helper functions (unchanged)
- Produces: Same public signature `pub fn vtt_to_ass_with_karaoke(caption_path, temp_dir, style) -> Result<String, String>` — writes ASS file to temp_dir and returns its path. Existing callers are unchanged.
- [ ] **Step 1: Write a test for the full ASS generation**
Add to the test module:
```rust
#[test]
fn test_vtt_to_ass_paged_output() {
let vtt_content = "\
WEBVTT
00:00:00.500 --> 00:00:01.500
hello<00:00:01.000><c> world</c>
00:00:02.000 --> 00:00:03.000
foo<00:00:02.500><c> bar</c>
00:00:03.000 --> 00:00:04.000
baz<00:00:03.500><c> qux</c>
00:00:04.000 --> 00:00:05.000
last<00:00:04.500><c> line</c>
";
let temp = std::env::temp_dir().join("test-paged-ass");
let _ = std::fs::remove_dir_all(&temp);
std::fs::create_dir_all(&temp).unwrap();
let vtt_path = temp.join("test.vtt");
std::fs::write(&vtt_path, vtt_content).unwrap();
let result = vtt_to_ass_with_karaoke(
vtt_path.to_str().unwrap(),
temp.to_str().unwrap(),
None,
);
assert!(result.is_ok());
let ass_path = result.unwrap();
let ass_content = std::fs::read_to_string(&ass_path).unwrap();
// Should have ASS header
assert!(ass_content.contains("[Script Info]"));
assert!(ass_content.contains("PlayResX: 1920"));
assert!(ass_content.contains("[Events]"));
// Should use \an2\pos (static position), NOT \move
assert!(ass_content.contains("\\an2\\pos("));
assert!(!ass_content.contains("\\move("));
// Should have \N line breaks (paged 2-line blocks)
assert!(ass_content.contains("\\N"));
// Should have \fad for cross-fade transitions
assert!(ass_content.contains("\\fad("));
// Should have karaoke \k tags
assert!(ass_content.contains("\\k"));
// 4 lines → 2 pages → 2 Dialogue events (plus possible non-speech)
let dialogue_count = ass_content.matches("Dialogue:").count();
assert_eq!(dialogue_count, 2, "Expected 2 pages (4 lines / 2). Got {dialogue_count}.\nASS:\n{ass_content}");
let _ = std::fs::remove_dir_all(&temp);
}
#[test]
fn test_vtt_to_ass_paged_gap_splits_page() {
let vtt_content = "\
WEBVTT
00:00:00.500 --> 00:00:01.500
hello<00:00:01.000><c> world</c>
00:00:05.000 --> 00:00:06.000
far<00:00:05.500><c> away</c>
";
let temp = std::env::temp_dir().join("test-paged-gap-ass");
let _ = std::fs::remove_dir_all(&temp);
std::fs::create_dir_all(&temp).unwrap();
let vtt_path = temp.join("test.vtt");
std::fs::write(&vtt_path, vtt_content).unwrap();
let result = vtt_to_ass_with_karaoke(
vtt_path.to_str().unwrap(),
temp.to_str().unwrap(),
None,
);
assert!(result.is_ok());
let ass_content = std::fs::read_to_string(result.unwrap()).unwrap();
// Gap between lines is 3.5s > 2.0s threshold → should split into 2 single-line pages
let dialogue_count = ass_content.matches("Dialogue:").count();
assert_eq!(dialogue_count, 2, "Gap >2s should split into separate pages. Got {dialogue_count}.\nASS:\n{ass_content}");
// Single-line pages should NOT have \N
// (Each Dialogue has its own text without \N)
for line in ass_content.lines() {
if line.starts_with("Dialogue:") {
assert!(!line.contains("\\N"), "Single-line page should not contain \\N: {line}");
}
}
let _ = std::fs::remove_dir_all(&temp);
}
```
- [ ] **Step 2: Run tests to verify the new ones fail**
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests::test_vtt_to_ass_paged -- --nocapture 2>&1`
Expected: Failures — the current implementation uses `\move` and generates more Dialogue events per line.
- [ ] **Step 3: Rewrite `vtt_to_ass_with_karaoke`**
Replace the entire function body (lines 465-643) with the paged implementation. The function signature stays the same:
```rust
/// Convert a word-timed VTT file to an ASS file with paged teleprompter display.
/// Lines are grouped into 2-line pages. Karaoke \k tags flow continuously from
/// line 1 through line 2 within each page. Pages cross-fade with \fad transitions.
pub fn vtt_to_ass_with_karaoke(
caption_path: &str,
temp_dir: &str,
style: Option<&CaptionStyle>,
) -> Result<String, String> {
let content = std::fs::read_to_string(caption_path)
.map_err(|e| format!("Failed to read VTT file '{}': {e}", caption_path))?;
std::fs::create_dir_all(temp_dir)
.map_err(|e| format!("Failed to create temp dir for ASS: {e}"))?;
let out_path = Path::new(temp_dir).join("karaoke.ass");
// Build style values from CaptionStyle or use defaults
let font_size = style.map_or(45, |s| (s.font_size as f64 * 2.5).round() as u32);
let primary_colour = style.map_or("&H00FFFFFF".to_string(), |s| hex_to_ass_color(&s.text_color));
let secondary_colour = "&H73CCCCCC".to_string();
let outline_colour = "&H00000000".to_string();
let back_colour = style.map_or("&H80000000".to_string(), |s| opacity_to_ass_back_colour(s.background_opacity));
let text_outline = style.map_or(true, |s| s.text_outline);
let (border_style, outline_val, shadow_val) = if text_outline {
(4, 2, 0)
} else {
(3, 0, 0)
};
let margin_v: i32 = 40;
let mut output = String::new();
// ASS Script Info header
output.push_str("[Script Info]\n");
output.push_str("ScriptType: v4.00+\n");
output.push_str("PlayResX: 1920\n");
output.push_str("PlayResY: 1080\n");
output.push_str("WrapStyle: 0\n");
output.push_str("ScaledBorderAndShadow: yes\n");
output.push('\n');
// V4+ Styles
output.push_str("[V4+ Styles]\n");
output.push_str("Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding\n");
output.push_str(&format!(
"Style: Default,Arial,{},{},{},{},{},0,0,0,0,100,100,0,0,{},{},{},2,20,20,{},1\n",
font_size, primary_colour, secondary_colour, outline_colour, back_colour,
border_style, outline_val, shadow_val, margin_v
));
output.push('\n');
// Events
output.push_str("[Events]\n");
output.push_str("Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text\n");
let spoken_lines = extract_spoken_lines(&content);
let gap_threshold = 2.0;
let crossfade_ms = 300;
let crossfade_s = crossfade_ms as f64 / 1000.0;
let pages = group_into_pages(&spoken_lines, gap_threshold);
let page_count = pages.len();
let cx = 960;
let y_bottom = 1080 - margin_v;
for (page_idx, page_indices) in pages.iter().enumerate() {
let page_lines: Vec<&SpokenLine> = page_indices.iter().map(|&i| &spoken_lines[i]).collect();
let first_line = page_lines[0];
let last_line = *page_lines.last().unwrap();
// Non-speech cue: simple static display, no karaoke, no cross-fade
if page_lines.len() == 1 && first_line.is_non_speech {
output.push_str(&format!(
"Dialogue: 0,{},{},Default,,0,0,0,,{{\\an2\\pos({},{})}}{}\n",
format_ass_timestamp(first_line.start_time),
format_ass_timestamp(first_line.end_time),
cx, y_bottom,
first_line.plain_text
));
continue;
}
// Compute display timing with cross-fade overlap
let natural_start = first_line.start_time;
let natural_end = last_line.end_time;
let is_first = page_idx == 0;
let is_last = page_idx == page_count - 1;
let display_start = if is_first {
natural_start
} else {
(natural_start - crossfade_s).max(0.0)
};
let display_end = if is_last {
natural_end
} else {
natural_end + crossfade_s
};
let fade_in = if is_first { 0 } else { crossfade_ms };
let fade_out = if is_last { 0 } else { crossfade_ms };
let karaoke_text = build_page_karaoke_text(&page_lines);
output.push_str(&format!(
"Dialogue: 0,{},{},Default,,0,0,0,,{{\\an2\\pos({},{})\\fad({},{})}}{}\n",
format_ass_timestamp(display_start),
format_ass_timestamp(display_end),
cx, y_bottom,
fade_in, fade_out,
karaoke_text
));
}
std::fs::write(&out_path, &output)
.map_err(|e| format!("Failed to write ASS file: {e}"))?;
Ok(out_path.to_string_lossy().to_string())
}
```
Also remove the now-unused `RollingLayout` struct (around line 452-458). The old `build_karaoke_text` function (line 432-449) can remain — it's not called by the new code but doesn't hurt and could be useful for future single-line scenarios.
- [ ] **Step 4: Run all tests**
Run: `cd src-tauri && cargo test --lib services::clip_exporter::tests -- --nocapture 2>&1 | tail -30`
Expected: All tests pass — both new paged tests and all existing tests.
- [ ] **Step 5: Build the full project**
Run: `cd src-tauri && cargo build 2>&1 | tail -10`
Expected: Clean build, no warnings about unused code (other than pre-existing ones).
- [ ] **Step 6: Commit**
```bash
git add src-tauri/src/services/clip_exporter.rs
git commit -m "feat(subtitles): rewrite ASS generation to paged teleprompter model
Replace per-line 3-phase rolling animation with 2-line paged blocks.
Karaoke \\k tags flow continuously across \\N line breaks within each page.
Pages cross-fade with \\fad transitions (300ms). No \\move or \\clip tags.
Reading flow is top-to-bottom within each page, matching teleprompter style."
```
---
### Task 3: Manual smoke test and cleanup
**Files:**
- Modify: `src-tauri/src/services/clip_exporter.rs` (only if issues found)
**Interfaces:**
- Consumes: The full paged ASS generation pipeline from Tasks 1-2
- Produces: Verified working burn-in subtitle export
- [ ] **Step 1: Run the app and test export**
Run: `cargo tauri dev`
Test flow:
1. Paste a YouTube URL with captions (e.g., `https://youtu.be/NruccMk0Jls`)
2. Wait for processing to complete
3. Mark a short clip (~10-15 seconds) containing speech
4. Open export dialog, enable captions with burn-in
5. Export the clip
6. Open the exported file and verify:
- Subtitles show as 2-line blocks
- Karaoke highlighting progresses top line → bottom line
- Page transitions cross-fade smoothly (no hard cuts or jumps)
- No position jumping or "2-line block scroll"
- [ ] **Step 2: Inspect the generated ASS file**
Check the temp ASS file to verify structure:
Run: `cat /tmp/video-clipper*/karaoke.ass | head -40`
Verify:
- Each `Dialogue:` line contains `\an2\pos(` (static position)
- No `\move(` tags anywhere
- Multi-line pages contain `\N` separator
- `\fad(` tags present on each Dialogue
- `\k` tags flow through both lines of each page
- [ ] **Step 3: Remove dead code if any**
If `RollingLayout` struct is still present, remove it. If `build_karaoke_text` is unused and triggers a warning, add `#[allow(dead_code)]` or remove it. Run `cargo build` to confirm clean.
- [ ] **Step 4: Final commit if cleanup was needed**
```bash
git add src-tauri/src/services/clip_exporter.rs
git commit -m "chore: remove dead rolling layout code"
```

View File

@@ -0,0 +1,95 @@
# Extended Caption Styling: Font, Shadow, Dimmed Color
**Date:** 2026-09-22
**Scope:** Add font selection (system font detection), drop shadow controls, dimmed/unspoken text color customization, and background toggle to the shared caption settings panel.
## Problem
Several burn-in caption style properties are hardcoded: font is always Arial, shadow is always off, and the dimmed (unspoken) text color in word-by-word mode is a hardcoded grey (`&H73CCCCCC`) that doesn't relate to the user's chosen text color. The preview uses a different approach (CSS opacity on the main color), creating a visual mismatch between preview and exported captions.
## New Settings Fields
Added to `CaptionSettings` (frontend) and `CaptionStyle` (backend):
| Field | Type | Default | Description |
|---|---|---|---|
| `fontFamily` | string | `'Arial'` | Font family name from system fonts |
| `backgroundEnabled` | boolean | `true` | Explicit toggle for the background box (replaces relying on opacity=0) |
| `shadowEnabled` | boolean | `false` | Drop shadow toggle (mutually exclusive with background) |
| `shadowDepth` | number (1-5) | `2` | Shadow offset in ASS units |
| `shadowColor` | string | `'#000000'` | Shadow color (hex) |
| `dimmedColorMode` | `'auto' \| 'custom'` | `'auto'` | How unspoken word color is derived |
| `dimmedOpacity` | number (0.1-0.9) | `0.4` | Opacity for auto-derived dimmed color |
| `dimmedColor` | string | `'#999999'` | Custom dimmed color (used in custom mode) |
## Background / Shadow Mutual Exclusivity
Background and shadow are **mutually exclusive**. Enabling one disables the other. This is required because ASS uses `BackColour` for both the background box and shadow color — they cannot be independent simultaneously.
### ASS BorderStyle Mapping
| Background | Outline | Shadow | ASS BorderStyle | ASS Outline | ASS Shadow |
|---|---|---|---|---|---|
| ON | ON | - | 4 | 2 | 0 |
| ON | OFF | - | 3 | 0 | 0 |
| - | ON | ON | 1 | 2 | depth |
| - | OFF | ON | 1 | 0 | depth |
| OFF | ON | OFF | 1 | 2 | 0 |
| OFF | OFF | OFF | 1 | 0 | 0 |
When **background is ON**: `BackColour` = `backgroundColor` + `backgroundOpacity` alpha.
When **shadow is ON**: `BackColour` = `shadowColor` (fully opaque).
When **neither**: `BackColour` = transparent (`&HFF000000`).
## Dimmed Text Color
In word-by-word (karaoke) mode, unspoken words appear in a dimmed color (ASS `SecondaryColour`).
### Auto Mode (default)
The dimmed color is derived from the main `textColor` at `dimmedOpacity`. Both preview and burn-in use the same approach:
- **Preview CSS:** `opacity: {dimmedOpacity}` on unspoken word spans (current behavior, but using the configurable value instead of hardcoded 0.4)
- **Burn-in ASS:** `SecondaryColour = hex_to_ass_color_with_alpha(textColor, (1 - dimmedOpacity) * 255)`
### Custom Mode
The user picks a specific `dimmedColor`. Both preview and burn-in use it directly:
- **Preview CSS:** `color: {dimmedColor}; opacity: 1` on unspoken word spans
- **Burn-in ASS:** `SecondaryColour = hex_to_ass_color(dimmedColor)`
## System Font Detection
A new Tauri command `list_system_fonts` runs `fc-list : family` (fontconfig, available because ffmpeg/libass depend on it). Output is parsed into a sorted, deduplicated list of font family names.
- Called once when the caption settings panel opens; result is cached in component state.
- If `fc-list` is not available, falls back to a hardcoded preset list: Arial, Helvetica, Verdana, Georgia, Times New Roman, Courier New, Impact.
- The dropdown renders each option with `font-family` set to that font's name, providing a live preview of each font.
## Panel UI Layout
The settings panel (320px wide) is organized into grouped sections:
1. **Font**: Dropdown (system fonts) + Size slider + Bold checkbox
2. **Text Color**: Color picker
3. **Dimmed Text**: Auto/Custom toggle. Auto: opacity slider. Custom: color picker.
4. **Outline**: Checkbox + color picker (disabled when unchecked)
5. **Background** (mutually exclusive with shadow): Checkbox + color picker + opacity slider
6. **Shadow** (mutually exclusive with background): Checkbox + depth slider + color picker
7. **Position**: Bottom/Top radio
8. **Word Highlight**: Checkbox (karaoke on/off)
9. **Reset to defaults** button
## Files Changed
| File | Change |
|---|---|
| `src/lib/stores/preferences.svelte.ts` | Add new fields to `CaptionSettings` interface and defaults |
| `src-tauri/src/models.rs` | Add new fields to `CaptionStyle` struct |
| `src/lib/bindings/export.ts` | Update TS `CaptionStyle` interface |
| `src/lib/components/ExportDialog.svelte` | Pass new fields through to export config |
| `src-tauri/src/services/clip_exporter.rs` | Use font, shadow, dimmed color, background toggle in ASS generation |
| `src/lib/components/CaptionSettingsPanel.svelte` | Add font dropdown, shadow controls, dimmed color controls, background toggle |
| `src/lib/components/VideoPlayer.svelte` | Use dimmed color settings + font in preview CSS |
| `src-tauri/src/commands/media_analysis.rs` (or new file) | Add `list_system_fonts` command |
| `src-tauri/src/lib.rs` | Register new command |
| `src/lib/bindings/mediaAnalysis.ts` (or new binding) | TS binding for `list_system_fonts` |

View File

@@ -0,0 +1,112 @@
# Paged Teleprompter Burn-In Subtitle Design
**Date:** 2026-09-22
**Scope:** Rewrite the ASS subtitle generation in `clip_exporter.rs` to use a paged teleprompter model instead of the current per-line rolling model.
## Problem
The current rolling subtitle implementation generates 3 independent ASS Dialogue events per spoken line (Active → Context → Disappear), each with `\move` animations. When a new line arrives, both the new and previous lines scroll simultaneously, creating a jarring "2-line block jump" instead of a natural reading flow. The active line always resets to the bottom position, breaking the top-to-bottom reading direction users expect.
## Design
### Core Model: 2-Line Paged Blocks
Group spoken lines into **pages of 2 lines each**. Each page is a **single ASS Dialogue event** containing both lines separated by `\N` (ASS hard line break). Karaoke `\k` tags flow continuously from line 1 through line 2 within the same Dialogue event.
**Reading flow:** The viewer reads the top line (karaoke highlighting progresses left-to-right), then naturally drops to the bottom line (karaoke continues), then a cross-fade transitions to the next page.
### Page Construction
Given a sequence of `SpokenLine` structs extracted from the VTT:
1. Pair lines into consecutive groups of 2: `[line0, line1]`, `[line2, line3]`, …
2. If the gap between two lines within a page exceeds 2 seconds, split them into separate pages instead of combining.
3. If the total line count is odd, the final page contains a single line.
### Karaoke Tag Stitching
For a 2-line page with lines A and B, the Dialogue text is:
```
{\k<dur>}word1_A {\k<dur>}word2_A ... {\k<dur>}lastword_A\N{\k<dur>}word1_B ... {\k<dur>}lastword_B
```
The `\k` duration for each word equals the time from that word's VTT start to the next word's VTT start. For the **last word of line A**, its `\k` duration extends to the start of line B's first word — this naturally covers any gap (silence/pause) between the two lines without any special gap-filler logic.
The `\N` is purely a visual line break and does not interrupt the karaoke timeline.
For lines without `<c>` word timing tags, the entire line text gets a single `\k` equal to the line's full spoken duration.
### Cross-Fade Transitions
Pages transition via a 300ms cross-dissolve:
- **Outgoing page:** ASS end time extended by 300ms past its natural end. Uses `\fad(*, 300)` for a 300ms fade-out.
- **Incoming page:** ASS start time moved 300ms before its natural start. Uses `\fad(300, *)` for a 300ms fade-in.
- The 300ms overlap produces a smooth cross-dissolve between pages.
Specific `\fad` values:
| Page Position | `\fad` value |
|---|---|
| First page | `\fad(0, 300)` |
| Middle pages | `\fad(300, 300)` |
| Last page | `\fad(300, 0)` |
When pages are separated by a long gap (>2s), the outgoing page fades out and the incoming page fades in independently — no visual overlap, just a clean silence gap.
### Positioning
Each Dialogue uses `\an2\pos(cx, y_bottom)` — bottom-center anchor. With `\an2`, the 2-line `\N` block renders with line 2 at the anchor point and line 1 stacked above it. **No `\move` tags are used.** Pages are static in position; transitions are opacity-only via `\fad`.
Layout constants (matching existing style):
- `cx = PlayResX / 2 = 960`
- `y_bottom = PlayResY - margin_v = 1080 - 40 = 1040`
### Edge Cases
| Case | Handling |
|---|---|
| Odd number of lines | Last page has 1 line (single-line Dialogue, no `\N`) |
| Gap >2s between consecutive lines within a pair | Split into separate single-line pages |
| Non-speech cues (`[Music]`, `[Applause]`) | Standalone single-line Dialogue, no karaoke, just `\pos` |
| Lines without `<c>` word timing | Single `\k` tag covering the line's full spoken duration |
| Only 1 spoken line total | Single-line Dialogue with `\fad(0, 0)` |
### ASS Header and Style
Unchanged from the current implementation:
- `PlayResX: 1920`, `PlayResY: 1080`, `ScaledBorderAndShadow: yes`
- Style parameters (font size, colours, border style, outline, margin) derived from user's `CaptionStyle` settings
- `Alignment: 2` (bottom-center) in style definition, reinforced with `\an2` override in each Dialogue
### What Gets Removed
The following components of the current implementation are replaced entirely:
- `RollingLayout` struct (position calculation for `\move` animations)
- Per-line 3-phase Dialogue generation (Active/Context/Disappear)
- All `\move` tags
- All `\clip` tags
- Phase-based timing calculations
### What Gets Added
- `build_page_karaoke_text(lines: &[&SpokenLine]) -> String` — stitches word timings from 1-2 lines into a single karaoke text with `\N` separator
- Page grouping logic in `vtt_to_ass_with_karaoke()` — pairs lines, handles gap-based splitting
- Cross-fade timing logic — computes `\fad` and adjusted start/end times per page
### What Stays the Same
- `SpokenLine` struct and `extract_spoken_lines()` — line extraction from VTT is unchanged
- `build_karaoke_text()` — per-line word timing extraction (reused internally by the new page builder)
- `parse_vtt_word_timings()` — VTT `<c>` tag parser
- ASS header/style generation
- `build_ffmpeg_args_burnin_subs()` — FFmpeg argument construction (unchanged, still uses `ass=` filter)
- All other export logic (lossless, muxed subtitles, progress reporting)
## Files Changed
| File | Change |
|---|---|
| `src-tauri/src/services/clip_exporter.rs` | Rewrite `vtt_to_ass_with_karaoke()` to use paged model. Remove `RollingLayout`. Add `build_page_karaoke_text()`. |
Single file change. No frontend, no new dependencies.

129
src-tauri/Cargo.lock generated
View File

@@ -82,6 +82,58 @@ version = "1.5.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f2032f911046de80f0a198e0901378627c33f59ea0ac00e363d481118bd70a53"
[[package]]
name = "axum"
version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90"
dependencies = [
"axum-core",
"bytes",
"form_urlencoded",
"futures-util",
"http",
"http-body",
"http-body-util",
"hyper",
"hyper-util",
"itoa",
"matchit",
"memchr",
"mime",
"percent-encoding",
"pin-project-lite",
"serde_core",
"serde_json",
"serde_path_to_error",
"serde_urlencoded",
"sync_wrapper",
"tokio",
"tower",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
name = "axum-core"
version = "0.5.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1"
dependencies = [
"bytes",
"futures-core",
"http",
"http-body",
"http-body-util",
"mime",
"pin-project-lite",
"sync_wrapper",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
name = "base64"
version = "0.21.7"
@@ -1306,12 +1358,24 @@ version = "0.1.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "21dec9db110f5f872ed9699c3ecf50cf16f423502706ba5c72462e28d3157573"
[[package]]
name = "http-range-header"
version = "0.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9171a2ea8a68358193d15dd5d70c1c10a2afc3e7e4c5bc92bc9f025cebd7359c"
[[package]]
name = "httparse"
version = "1.10.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6dbf3de79e51f3d586ab4cb9d5c3e2c14aa28ed23d180cf89b4df0454a69cc87"
[[package]]
name = "httpdate"
version = "1.0.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "df3b46402a9d5adb4c86a0cf463f42e19994e3ee891101b1841f30a545cb49a9"
[[package]]
name = "hyper"
version = "1.11.1"
@@ -1325,6 +1389,7 @@ dependencies = [
"http",
"http-body",
"httparse",
"httpdate",
"itoa",
"pin-project-lite",
"smallvec",
@@ -1823,6 +1888,12 @@ dependencies = [
"web_atoms",
]
[[package]]
name = "matchit"
version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3"
[[package]]
name = "memchr"
version = "2.8.3"
@@ -1844,6 +1915,16 @@ version = "0.3.17"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6877bb514081ee2a7ff5ef9de3281f14a4dd4bceac4c09388074a6b5df8a139a"
[[package]]
name = "mime_guess"
version = "2.0.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f7c44f8e672c00fe5308fa235f821cb4198414e1c77935c1ab6948d3fd78550e"
dependencies = [
"mime",
"unicase",
]
[[package]]
name = "miniz_oxide"
version = "0.8.9"
@@ -2671,6 +2752,12 @@ version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f"
[[package]]
name = "ryu"
version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f"
[[package]]
name = "same-file"
version = "1.0.6"
@@ -2832,6 +2919,17 @@ dependencies = [
"zmij",
]
[[package]]
name = "serde_path_to_error"
version = "0.1.20"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "10a9ff822e371bb5403e391ecd83e182e0e77ba7f6fe0160b795797109d1b457"
dependencies = [
"itoa",
"serde",
"serde_core",
]
[[package]]
name = "serde_repr"
version = "0.1.21"
@@ -2861,6 +2959,18 @@ dependencies = [
"serde_core",
]
[[package]]
name = "serde_urlencoded"
version = "0.7.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d3491c14715ca2294c4d6a88f15e84739788c1d030eed8c110436aafdaa2f3fd"
dependencies = [
"form_urlencoded",
"itoa",
"ryu",
"serde",
]
[[package]]
name = "serde_with"
version = "3.23.0"
@@ -3297,6 +3407,7 @@ dependencies = [
name = "tauri-app"
version = "0.1.0"
dependencies = [
"axum",
"dirs",
"serde",
"serde_json",
@@ -3309,6 +3420,7 @@ dependencies = [
"tempfile",
"thiserror 2.0.20",
"tokio",
"tower-http",
"uuid",
]
@@ -3861,6 +3973,7 @@ dependencies = [
"tokio",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
@@ -3871,10 +3984,19 @@ checksum = "4cfcf7e2740e6fc6d4d688b4ef00650406bb94adf4731e43c096c3a19fe40840"
dependencies = [
"bitflags 2.13.2",
"bytes",
"futures-core",
"futures-util",
"http",
"http-body",
"http-body-util",
"http-range-header",
"httpdate",
"mime",
"mime_guess",
"percent-encoding",
"pin-project-lite",
"tokio",
"tokio-util",
"tower",
"tower-layer",
"tower-service",
@@ -3899,6 +4021,7 @@ version = "0.1.44"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "63e71662fa4b2a2c3a26f570f037eb95bb1f85397f3cd8076caed2f026a6d100"
dependencies = [
"log",
"pin-project-lite",
"tracing-attributes",
"tracing-core",
@@ -4005,6 +4128,12 @@ dependencies = [
"unic-common",
]
[[package]]
name = "unicase"
version = "2.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dbc4bc3a9f746d862c45cb89d705aa10f187bb96c76001afab07a0d35ce60142"
[[package]]
name = "unicode-ident"
version = "1.0.26"

View File

@@ -30,6 +30,8 @@ thiserror = "2.0.20"
uuid = { version = "1.26.1", features = ["v4"] }
tempfile = "3.27.0"
dirs = "6.0.0"
axum = "0.8"
tower-http = { version = "0.6", features = ["fs", "cors"] }
# Read the optimization guideline for more details: https://tauri.app/concept/size/#cargo-configuration

View File

@@ -12,6 +12,11 @@ pub fn check_dependencies(cookie_source: CookieSource) -> Vec<DependencyStatus>
dependency_manager::get_all_dependency_statuses(&cookie_source)
}
#[tauri::command]
pub fn check_subtitles_filter_available() -> bool {
dependency_manager::check_subtitles_filter()
}
#[derive(Clone, Serialize)]
#[serde(rename_all = "camelCase", tag = "event", content = "data")]
pub enum InstallEvent {

View File

@@ -11,6 +11,12 @@ pub enum ExportEvent {
total: usize,
label: String,
},
ClipProgress {
current: usize,
total: usize,
label: String,
percent: f64,
},
Finished {
paths: Vec<String>,
},
@@ -27,6 +33,10 @@ pub async fn export_clips(
tokio::task::spawn_blocking(move || {
let output_directory = clip_exporter::expand_tilde_path(&config.output_directory);
let ext = clip_exporter::get_extension(&config.source_file_path, &config.cut_mode);
let use_subs = config.include_captions && config.caption_file_path.is_some();
let burn_in = config.burn_in_captions;
let caption_path = config.caption_file_path.as_deref().unwrap_or("");
let caption_style = config.caption_style.as_ref();
match config.export_scope {
ExportScope::Individual => {
@@ -34,10 +44,13 @@ pub async fn export_clips(
let mut output_paths = Vec::new();
for (i, clip) in config.clips.iter().enumerate() {
let clip_num = i + 1;
let label = clip.label.clone();
let _ = on_event.send(ExportEvent::Progress {
current: i + 1,
current: clip_num,
total,
label: clip.label.clone(),
label: label.clone(),
});
let output = clip_exporter::generate_output_path(
@@ -47,12 +60,35 @@ pub async fn export_clips(
&ext,
);
let progress_cb = |pct: f64| {
let _ = on_event.send(ExportEvent::ClipProgress {
current: clip_num,
total,
label: label.clone(),
percent: pct,
});
};
if use_subs {
clip_exporter::export_single_clip_with_subs(
clip,
&config.source_file_path,
&output,
&config.cut_mode,
caption_path,
burn_in,
caption_style,
&progress_cb,
)?;
} else {
clip_exporter::export_single_clip(
clip,
&config.source_file_path,
&output,
&config.cut_mode,
&progress_cb,
)?;
}
output_paths.push(output);
}
@@ -63,6 +99,7 @@ pub async fn export_clips(
Ok(output_paths)
}
ExportScope::Merged => {
let total_clips = config.clips.len();
let _ = on_event.send(ExportEvent::Progress {
current: 1,
total: 1,
@@ -89,14 +126,45 @@ pub async fn export_clips(
.unwrap_or(std::cmp::Ordering::Equal)
});
let clip_labels: Vec<String> = sorted_clips.iter().map(|c| c.label.clone()).collect();
let temp_dir_str = temp_dir.to_string_lossy().to_string();
let merge_result = clip_exporter::export_merged(
let merge_progress = |clip_idx: usize, pct: f64| {
let label = if clip_idx < clip_labels.len() {
format!("Encoding {}", clip_labels[clip_idx])
} else {
"Concatenating…".to_string()
};
let _ = on_event.send(ExportEvent::ClipProgress {
current: clip_idx + 1,
total: total_clips,
label,
percent: pct,
});
};
let merge_result = if use_subs {
clip_exporter::export_merged_with_subs(
&sorted_clips,
&config.source_file_path,
&output,
&config.cut_mode,
&temp_dir_str,
);
caption_path,
burn_in,
caption_style,
&merge_progress,
)
} else {
clip_exporter::export_merged(
&sorted_clips,
&config.source_file_path,
&output,
&config.cut_mode,
&temp_dir_str,
&merge_progress,
)
};
let _ = std::fs::remove_dir_all(&temp_dir);
merge_result?;

View File

@@ -1,5 +1,5 @@
use crate::models::{CookieSource, VideoMetadata};
use crate::services::{download_manager, video_resolver};
use crate::services::{download_manager, subtitle_downloader, video_resolver};
use serde::Serialize;
use std::io::{BufRead, BufReader};
use tauri::ipc::Channel;
@@ -23,24 +23,36 @@ pub enum DownloadEvent {
Error { message: String },
}
fn variant_dir(variant: &str) -> String {
std::env::temp_dir()
.join("video-clipper")
.join(variant)
.to_string_lossy()
.to_string()
}
#[tauri::command]
pub async fn start_download(
url: String,
cookie_source: CookieSource,
format_spec: String,
variant: String,
on_event: Channel<DownloadEvent>,
) -> Result<(), String> {
tokio::task::spawn_blocking(move || {
let temp_dir = std::env::temp_dir()
.join("video-clipper")
.to_string_lossy()
.to_string();
std::fs::create_dir_all(&temp_dir)
.map_err(|e| format!("Failed to create temp dir: {e}"))?;
let output_dir = variant_dir(&variant);
std::fs::create_dir_all(&output_dir)
.map_err(|e| format!("Failed to create dir: {e}"))?;
eprintln!(
"[video-clipper:download:{}] starting — format='{}' dir='{}'",
variant, format_spec, output_dir
);
let (mut child, _template) =
download_manager::start_download(&url, &cookie_source, &temp_dir)?;
download_manager::start_download(&url, &cookie_source, &output_dir, &format_spec)?;
let Some(stderr) = child.stderr.take() else {
let Some(stdout) = child.stdout.take() else {
let status = child
.wait()
.map_err(|e| format!("Download failed: {e}"))?;
@@ -51,7 +63,7 @@ pub async fn start_download(
return Ok(());
};
let reader = BufReader::new(stderr);
let reader = BufReader::new(stdout);
let mut last_file_path: Option<String> = None;
for line in reader.lines().map_while(Result::ok) {
@@ -59,9 +71,7 @@ pub async fn start_download(
let _ = on_event.send(DownloadEvent::Progress { percent: pct });
}
if line.contains("[download] Destination:")
|| line.contains("[Merger] Merging formats into")
{
if line.contains("[download] Destination:") {
if let Some(path) = line.split(": ").nth(1) {
let trimmed = path.trim().trim_matches('"').to_string();
let _ = on_event.send(DownloadEvent::FilePath {
@@ -69,6 +79,30 @@ pub async fn start_download(
});
last_file_path = Some(trimmed);
}
} else if line.contains("[Merger] Merging formats into") {
let trimmed = line
.trim_start_matches("[Merger] Merging formats into ")
.trim()
.trim_matches('"')
.to_string();
if !trimmed.is_empty() {
let _ = on_event.send(DownloadEvent::FilePath {
path: trimmed.clone(),
});
last_file_path = Some(trimmed);
}
} else if line.contains("has already been downloaded") {
if let Some(rest) = line.strip_prefix("[download] ") {
if let Some(path) = rest.strip_suffix(" has already been downloaded") {
let trimmed = path.trim().to_string();
if !trimmed.is_empty() {
let _ = on_event.send(DownloadEvent::FilePath {
path: trimmed.clone(),
});
last_file_path = Some(trimmed);
}
}
}
}
}
@@ -76,6 +110,12 @@ pub async fn start_download(
.wait()
.map_err(|e| format!("Download failed: {e}"))?;
let path = last_file_path.unwrap_or_default();
eprintln!(
"[video-clipper:download:{}] finished — success={}, path='{}'",
variant,
status.success(),
path
);
let _ = on_event.send(DownloadEvent::Finished {
success: status.success(),
path,
@@ -86,3 +126,67 @@ pub async fn start_download(
.await
.map_err(|e| format!("Task failed: {e}"))?
}
/// Check if a video with the given title exists in the variant's download dir.
#[tauri::command]
pub async fn check_cached_download(title: String, variant: String) -> Option<String> {
let dir = variant_dir(&variant);
let dir_path = std::path::Path::new(&dir);
if !dir_path.is_dir() {
return None;
}
let entries = std::fs::read_dir(dir_path).ok()?;
for entry in entries.flatten() {
let path = entry.path();
if path.is_file() {
if let Some(stem) = path.file_stem().and_then(|s| s.to_str()) {
if stem == title {
let path_str = path.to_string_lossy().to_string();
if let Ok(meta) = std::fs::metadata(&path_str) {
if meta.len() > 0 {
eprintln!(
"[video-clipper:cache:{}] found: '{}'",
variant, path_str
);
return Some(path_str);
}
}
}
}
}
}
None
}
/// Download English VTT subtitles for a video URL.
/// Returns the path to the .vtt file, or an error string.
#[tauri::command]
pub async fn download_subtitles(
url: String,
cookie_source: CookieSource,
is_auto: bool,
) -> Result<String, String> {
tokio::task::spawn_blocking(move || {
let output_dir = std::env::temp_dir()
.join("video-clipper")
.join("subtitles")
.to_string_lossy()
.to_string();
eprintln!(
"[video-clipper:subtitles] downloading (auto={}) to '{}'",
is_auto, output_dir
);
let result =
subtitle_downloader::download_subtitles(&url, &cookie_source, &output_dir, is_auto);
match &result {
Ok(path) => eprintln!("[video-clipper:subtitles] done: '{}'", path),
Err(e) => eprintln!("[video-clipper:subtitles] FAILED: {}", e),
}
result
})
.await
.map_err(|e| format!("Task failed: {e}"))?
}

View File

@@ -1,5 +1,37 @@
use crate::models::{CookieSource, DependencyStatus};
use std::path::PathBuf;
use std::process::Command;
use std::sync::OnceLock;
/// Cached resolved paths for ffmpeg and ffprobe binaries.
/// Prefers ffmpeg-full (Homebrew keg) if available, otherwise falls back to PATH.
static FFMPEG_PATH: OnceLock<String> = OnceLock::new();
static FFPROBE_PATH: OnceLock<String> = OnceLock::new();
/// Returns the best available ffmpeg binary path.
/// Checks for Homebrew's ffmpeg-full keg first (has libass), then falls back to PATH ffmpeg.
pub fn ffmpeg_bin() -> &'static str {
FFMPEG_PATH.get_or_init(|| {
let keg = PathBuf::from("/opt/homebrew/opt/ffmpeg-full/bin/ffmpeg");
if keg.exists() {
keg.to_string_lossy().into_owned()
} else {
"ffmpeg".to_string()
}
})
}
/// Returns the best available ffprobe binary path.
pub fn ffprobe_bin() -> &'static str {
FFPROBE_PATH.get_or_init(|| {
let keg = PathBuf::from("/opt/homebrew/opt/ffmpeg-full/bin/ffprobe");
if keg.exists() {
keg.to_string_lossy().into_owned()
} else {
"ffprobe".to_string()
}
})
}
pub fn check_tool_exists(name: &str) -> Option<String> {
let output = Command::new("which").arg(name).output().ok()?;
@@ -80,6 +112,26 @@ pub fn get_all_dependency_statuses(cookie_source: &CookieSource) -> Vec<Dependen
]
}
/// Check whether ffmpeg's `subtitles` video filter is available.
/// Requires ffmpeg to be compiled with libass support.
pub fn check_subtitles_filter() -> bool {
let output = Command::new(ffmpeg_bin())
.arg("-filters")
.output();
match output {
Ok(out) => {
let stdout = String::from_utf8_lossy(&out.stdout);
stdout.lines().any(|line| {
let trimmed = line.trim();
// Filter list lines look like: " T. subtitles V->V ..."
trimmed.contains("subtitles") && trimmed.contains("V->V")
})
}
Err(_) => false,
}
}
pub fn get_install_args(dep_name: &str) -> Result<Vec<String>, String> {
match dep_name {
"ffmpeg" => Ok(vec![

View File

@@ -16,9 +16,10 @@ pub fn parse_progress_line(line: &str) -> Option<f64> {
pub fn start_download(
url: &str,
cookie_source: &CookieSource,
temp_dir: &str,
output_dir: &str,
format_spec: &str,
) -> Result<(Child, String), String> {
let output_template = format!("{}/%(title)s.%(ext)s", temp_dir);
let output_template = format!("{}/%(title)s.%(ext)s", output_dir);
let mut cmd = Command::new("yt-dlp");
cmd.arg("-o")
@@ -27,15 +28,17 @@ pub fn start_download(
.arg("--no-part")
.arg("-c")
.arg("--remote-components")
.arg("ejs:github");
.arg("ejs:github")
.arg("-f")
.arg(format_spec);
for arg in cookie_source.to_ytdlp_args() {
cmd.arg(arg);
}
cmd.arg(url)
.stdout(Stdio::null())
.stderr(Stdio::piped());
.stdout(Stdio::piped())
.stderr(Stdio::inherit());
let child = cmd
.spawn()

View File

@@ -1,3 +1,4 @@
use crate::services::dependency_manager::ffprobe_bin;
use std::process::Command;
pub fn parse_ffprobe_output(output: &str) -> Vec<f64> {
@@ -17,7 +18,7 @@ pub fn parse_ffprobe_output(output: &str) -> Vec<f64> {
}
pub fn extract_keyframes(file_path: &str) -> Result<Vec<f64>, String> {
let output = Command::new("ffprobe")
let output = Command::new(ffprobe_bin())
.args([
"-select_streams",
"v",

View File

@@ -0,0 +1,44 @@
use std::net::TcpListener;
use tower_http::cors::CorsLayer;
use tower_http::services::ServeDir;
/// Bind a local HTTP file server to a random port and return the port + std listener.
/// The caller is responsible for spawning the server on a Tokio runtime.
///
/// WKWebView plays audio correctly from http://localhost URLs with proper
/// range-request support (provided by tower-http's ServeDir).
pub fn bind_media_server() -> Result<(u16, TcpListener), String> {
let listener = TcpListener::bind("127.0.0.1:0")
.map_err(|e| format!("Failed to bind media server: {e}"))?;
listener
.set_nonblocking(true)
.map_err(|e| format!("Failed to set non-blocking: {e}"))?;
let port = listener
.local_addr()
.map_err(|e| format!("Failed to get local addr: {e}"))?
.port();
Ok((port, listener))
}
/// Spawn the axum server on the Tauri async runtime.
/// Must be called after a Tokio runtime is available.
pub fn spawn_media_server(listener: TcpListener) {
tauri::async_runtime::spawn(async move {
let tokio_listener = tokio::net::TcpListener::from_std(listener)
.expect("Failed to convert std listener to tokio");
let service = ServeDir::new("/");
let app = axum::Router::new()
.fallback_service(service)
.layer(CorsLayer::permissive());
eprintln!(
"[video-clipper:media-server] serving on {:?}",
tokio_listener.local_addr()
);
if let Err(e) = axum::serve(tokio_listener, app).await {
eprintln!("[video-clipper:media-server] server error: {e}");
}
});
}

View File

@@ -2,6 +2,8 @@ pub mod clip_exporter;
pub mod dependency_manager;
pub mod download_manager;
pub mod keyframe_index;
pub mod media_server;
pub mod subtitle_downloader;
pub mod thumbnail_extractor;
pub mod video_resolver;
pub mod waveform_generator;

View File

@@ -0,0 +1,76 @@
use crate::models::CookieSource;
use std::process::Command;
/// Download English VTT subtitles for a URL.
/// `is_auto` controls whether to use --write-subs (manual) or --write-auto-subs (auto-generated).
/// Returns the path to the downloaded .vtt file, or an error.
pub fn download_subtitles(
url: &str,
cookie_source: &CookieSource,
output_dir: &str,
is_auto: bool,
) -> Result<String, String> {
std::fs::create_dir_all(output_dir)
.map_err(|e| format!("Failed to create subtitle dir: {e}"))?;
let output_template = format!("{}/%(title)s.%(ext)s", output_dir);
let mut cmd = Command::new("yt-dlp");
cmd.arg("--skip-download")
.arg("--remote-components")
.arg("ejs:github");
if is_auto {
cmd.arg("--write-auto-subs");
} else {
cmd.arg("--write-subs");
}
cmd.arg("--sub-langs").arg("en.*,en");
cmd.arg("--sub-format").arg("vtt");
cmd.arg("--convert-subs").arg("vtt");
cmd.arg("-o").arg(&output_template);
for arg in cookie_source.to_ytdlp_args() {
cmd.arg(arg);
}
cmd.arg(url);
let output = cmd
.output()
.map_err(|e| format!("Failed to run yt-dlp for subtitles: {e}"))?;
if !output.status.success() {
let stderr = String::from_utf8_lossy(&output.stderr);
return Err(format!("yt-dlp subtitle download failed: {stderr}"));
}
// Find the downloaded .vtt file in the output directory
find_vtt_file(output_dir)
}
fn find_vtt_file(dir: &str) -> Result<String, String> {
let entries = std::fs::read_dir(dir)
.map_err(|e| format!("Failed to read subtitle dir: {e}"))?;
// Prefer .en.vtt, then any .vtt
let mut vtt_files: Vec<String> = Vec::new();
for entry in entries.flatten() {
let path = entry.path();
if path.extension().and_then(|e| e.to_str()) == Some("vtt") {
let path_str = path.to_string_lossy().to_string();
vtt_files.push(path_str);
}
}
// Prefer files with ".en." in the name
if let Some(en_file) = vtt_files.iter().find(|f| f.contains(".en.")) {
return Ok(en_file.clone());
}
// Fall back to any VTT file
vtt_files
.into_iter()
.next()
.ok_or_else(|| "No .vtt subtitle file found after download".to_string())
}

View File

@@ -1,4 +1,5 @@
use crate::models::ThumbnailSpritesheet;
use crate::services::dependency_manager::ffmpeg_bin;
use std::process::Command;
pub fn compute_interval(duration: f64) -> f64 {
@@ -27,7 +28,7 @@ pub fn extract_thumbnails(
let frames_per_sheet = (columns * rows_per_sheet) as usize;
let frames_pattern = format!("{output_dir}/frame_%06d.jpg");
let status = Command::new("ffmpeg")
let status = Command::new(ffmpeg_bin())
.args([
"-i",
file_path,

View File

@@ -1,13 +1,56 @@
use crate::models::{CookieSource, VideoMetadata};
use serde::Deserialize;
use std::collections::HashMap;
use std::process::Command;
#[derive(Debug, Deserialize)]
pub struct SubtitleFormat {
pub ext: Option<String>,
pub url: Option<String>,
}
#[derive(Debug, Deserialize)]
pub struct YtDlpJson {
pub title: String,
pub duration: Option<f64>,
pub fps: Option<f64>,
pub thumbnail: Option<String>,
/// Manually authored subtitles: language_code -> [formats]
#[serde(default)]
pub subtitles: HashMap<String, Vec<SubtitleFormat>>,
/// Auto-generated subtitles: language_code -> [formats]
#[serde(default)]
pub automatic_captions: HashMap<String, Vec<SubtitleFormat>>,
}
/// Check if subtitles have a VTT format for English (en, en-US, en-GB, en-orig).
fn find_english_key(map: &HashMap<String, Vec<SubtitleFormat>>) -> Option<&str> {
// Prefer exact "en", then any en-* variant
for key in ["en", "en-US", "en-GB", "en-orig"] {
if map.contains_key(key) {
return Some(key);
}
}
// Try any key starting with "en"
for key in map.keys() {
if key.starts_with("en") {
return Some(key.as_str());
}
}
None
}
/// Determine caption availability: (has_captions, is_auto)
pub fn detect_captions(meta: &YtDlpJson) -> (bool, bool) {
// Priority 1: manual English subtitles
if find_english_key(&meta.subtitles).is_some() {
return (true, false);
}
// Priority 2: auto-generated English captions
if find_english_key(&meta.automatic_captions).is_some() {
return (true, true);
}
(false, false)
}
pub fn parse_ytdlp_json(json_str: &str) -> Result<YtDlpJson, String> {
@@ -41,6 +84,8 @@ pub fn resolve(url: &str, cookie_source: &CookieSource) -> Result<VideoMetadata,
let json_str = String::from_utf8_lossy(&meta_output.stdout);
let meta = parse_ytdlp_json(&json_str)?;
let (has_captions, captions_are_auto) = detect_captions(&meta);
let mut stream_cmd = Command::new("yt-dlp");
stream_cmd
.arg("-g")
@@ -79,6 +124,8 @@ pub fn resolve(url: &str, cookie_source: &CookieSource) -> Result<VideoMetadata,
fps: meta.fps.unwrap_or(30.0),
thumbnail_url: meta.thumbnail,
stream_url,
has_captions,
captions_are_auto,
})
}
@@ -107,6 +154,8 @@ mod tests {
assert!(result.duration.is_none());
assert!(result.fps.is_none());
assert!(result.thumbnail.is_none());
assert!(result.subtitles.is_empty());
assert!(result.automatic_captions.is_empty());
}
#[test]
@@ -114,4 +163,39 @@ mod tests {
let result = parse_ytdlp_json("not json");
assert!(result.is_err());
}
#[test]
fn test_detect_captions_manual() {
let json = r#"{"title": "T", "subtitles": {"en": [{"ext": "vtt", "url": "http://x"}]}, "automatic_captions": {}}"#;
let meta = parse_ytdlp_json(json).unwrap();
let (has, is_auto) = detect_captions(&meta);
assert!(has);
assert!(!is_auto);
}
#[test]
fn test_detect_captions_auto_only() {
let json = r#"{"title": "T", "subtitles": {}, "automatic_captions": {"en": [{"ext": "vtt", "url": "http://x"}]}}"#;
let meta = parse_ytdlp_json(json).unwrap();
let (has, is_auto) = detect_captions(&meta);
assert!(has);
assert!(is_auto);
}
#[test]
fn test_detect_captions_none() {
let json = r#"{"title": "T"}"#;
let meta = parse_ytdlp_json(json).unwrap();
let (has, _) = detect_captions(&meta);
assert!(!has);
}
#[test]
fn test_detect_captions_en_variant() {
let json = r#"{"title": "T", "subtitles": {"en-US": [{"ext": "vtt"}]}}"#;
let meta = parse_ytdlp_json(json).unwrap();
let (has, is_auto) = detect_captions(&meta);
assert!(has);
assert!(!is_auto);
}
}

View File

@@ -1,30 +1,110 @@
use std::process::Command;
use crate::models::WaveformTiers;
use crate::services::dependency_manager::ffmpeg_bin;
use std::io::{BufRead, BufReader, Read as _};
use std::process::{Command, Stdio};
pub fn extract_waveform(file_path: &str, sample_count: usize) -> Result<Vec<f64>, String> {
let output = Command::new("ffmpeg")
const TIER0_SIZE: usize = 2_000;
const TIER1_SIZE: usize = 10_000;
/// Extract waveform peaks and build all zoom tiers in a single ffmpeg call.
/// The `on_progress` callback receives values from 0.0 to 1.0.
pub fn extract_waveform_tiers(
file_path: &str,
duration: f64,
on_progress: impl Fn(f64) + Send + 'static,
) -> Result<WaveformTiers, String> {
// ~100 peaks/sec, capped at 200K, floored at 50K
let raw_count = (duration * 100.0).round() as usize;
let raw_count = raw_count.clamp(50_000, 200_000);
// Compute the actual output sample rate in Hz.
// aresample=N sets rate to N Hz (NOT total count), so we must divide by duration.
let sample_rate = ((raw_count as f64) / duration).ceil() as usize;
let sample_rate = sample_rate.max(8);
eprintln!(
"[waveform_generator] extracting ~{} peaks ({}Hz x {:.0}s) for '{}'",
raw_count, sample_rate, duration, file_path
);
let mut child = Command::new(ffmpeg_bin())
.args([
"-i",
file_path,
"-ac",
"1",
"-filter:a",
&format!("aresample={sample_count}"),
"-f",
"f32le",
"-i", file_path,
"-ac", "1",
"-ar", &sample_rate.to_string(),
"-f", "f32le",
"-vn",
"-progress", "pipe:2",
"-",
])
.output()
.stdout(Stdio::piped())
.stderr(Stdio::piped())
.spawn()
.map_err(|e| format!("Failed to run ffmpeg: {e}"))?;
if !output.status.success() {
let stderr = String::from_utf8_lossy(&output.stderr);
return Err(format!("ffmpeg waveform extraction failed: {stderr}"));
let stderr = child.stderr.take().ok_or("No stderr")?;
let mut stdout = child.stdout.take().ok_or("No stdout")?;
// Parse progress from stderr in a background thread
let duration_us = duration * 1_000_000.0;
let progress_thread = std::thread::spawn(move || {
let reader = BufReader::new(stderr);
for line in reader.lines().map_while(Result::ok) {
if let Some(us_str) = line.strip_prefix("out_time_us=") {
if let Ok(us) = us_str.parse::<f64>() {
if duration_us > 0.0 {
on_progress((us / duration_us).clamp(0.0, 1.0));
}
}
}
}
});
// Read all waveform data from stdout
let mut waveform_data = Vec::new();
stdout
.read_to_end(&mut waveform_data)
.map_err(|e| format!("Failed to read ffmpeg stdout: {e}"))?;
progress_thread.join().ok();
let status = child.wait().map_err(|e| format!("ffmpeg wait failed: {e}"))?;
if !status.success() {
return Err("ffmpeg waveform extraction failed".to_string());
}
let samples = parse_f32_samples(&output.stdout);
let peaks = compute_peaks(&samples, sample_count);
Ok(peaks)
let samples = parse_f32_samples(&waveform_data);
let tier2 = compute_peaks(&samples, raw_count);
let tier1 = downsample(&tier2, TIER1_SIZE);
let tier0 = downsample(&tier2, TIER0_SIZE);
eprintln!(
"[waveform_generator] tiers: {} / {} / {} peaks",
tier0.len(),
tier1.len(),
tier2.len()
);
Ok(WaveformTiers { tier0, tier1, tier2 })
}
/// Downsample a peaks array to a smaller size by taking the max of each chunk.
fn downsample(peaks: &[f64], target_count: usize) -> Vec<f64> {
if peaks.is_empty() || target_count == 0 {
return vec![0.0; target_count];
}
if target_count >= peaks.len() {
return peaks.to_vec();
}
let chunk_size = (peaks.len() as f64 / target_count as f64).ceil() as usize;
let chunk_size = chunk_size.max(1);
peaks
.chunks(chunk_size)
.map(|chunk| chunk.iter().cloned().fold(0.0_f64, f64::max))
.take(target_count)
.collect()
}
fn parse_f32_samples(data: &[u8]) -> Vec<f32> {
@@ -53,7 +133,8 @@ fn compute_peaks(samples: &[f32], target_count: usize) -> Vec<f64> {
}
}
peaks.resize(target_count, 0.0);
// Do NOT resize/pad — the actual peak count from chunking is correct.
// Padding with zeros causes the waveform to appear truncated.
peaks
}
@@ -87,4 +168,28 @@ mod tests {
assert_eq!(samples.len(), 1);
assert!((samples[0] - 0.5).abs() < 0.001);
}
#[test]
fn test_downsample() {
let peaks = vec![0.1, 0.5, 0.3, 0.8, 0.2, 0.9];
let down = downsample(&peaks, 3);
assert_eq!(down.len(), 3);
assert!((down[0] - 0.5).abs() < 0.001);
assert!((down[1] - 0.8).abs() < 0.001);
assert!((down[2] - 0.9).abs() < 0.001);
}
#[test]
fn test_downsample_identity() {
let peaks = vec![0.1, 0.5, 0.3];
let down = downsample(&peaks, 5);
assert_eq!(down, peaks);
}
#[test]
fn test_downsample_empty() {
let down = downsample(&[], 3);
assert_eq!(down.len(), 3);
assert!(down.iter().all(|p| *p == 0.0));
}
}

View File

@@ -21,7 +21,14 @@
"csp": null,
"assetProtocol": {
"enable": true,
"scope": ["$TEMP/**", "$TMP/**", "/tmp/**"]
"scope": [
"$TEMP/**",
"$TMP/**",
"/tmp/**",
"/private/tmp/**",
"/private/var/**",
"/var/**"
]
}
}
},

View File

@@ -8,6 +8,7 @@
import SetupWizard from '$lib/components/SetupWizard.svelte';
import PreferencesPanel from '$lib/components/PreferencesPanel.svelte';
import ExportDialog from '$lib/components/ExportDialog.svelte';
import ProcessingModal from '$lib/components/ProcessingModal.svelte';
import { loadPreferences } from '$lib/stores/preferences.svelte';
import { getSelectedClipId, removeClip } from '$lib/stores/clips.svelte';
import { adjustShuttle, resetShuttleRate } from '$lib/transport/playback';
@@ -21,6 +22,13 @@
let transportControls = $state<TransportControls | null>(null);
// Resizable split: timeline height in px (clip list gets the rest)
let timelineHeight = $state(180);
let isResizing = $state(false);
let resizeStartY = $state(0);
let resizeStartHeight = $state(0);
let splitAreaEl = $state<HTMLDivElement | null>(null);
$effect(() => {
loadPreferences();
});
@@ -36,6 +44,25 @@
transportControls?.handleKeyAction(action);
}
function handleResizeStart(e: MouseEvent) {
e.preventDefault();
isResizing = true;
resizeStartY = e.clientY;
resizeStartHeight = timelineHeight;
}
function handleResizeMove(e: MouseEvent) {
if (!isResizing || !splitAreaEl) return;
const delta = e.clientY - resizeStartY;
const totalHeight = splitAreaEl.clientHeight;
const newHeight = Math.max(40, Math.min(totalHeight - 20, resizeStartHeight + delta));
timelineHeight = newHeight;
}
function handleResizeEnd() {
isResizing = false;
}
function handleGlobalKeydown(e: KeyboardEvent) {
const target = e.target as HTMLElement;
if (target.tagName === 'INPUT' || target.tagName === 'TEXTAREA') return;
@@ -110,7 +137,11 @@
}
</script>
<svelte:window onkeydown={handleGlobalKeydown} />
<svelte:window
onkeydown={handleGlobalKeydown}
onmousemove={handleResizeMove}
onmouseup={handleResizeEnd}
/>
{#if showSetupWizard}
<SetupWizard onComplete={() => (showSetupWizard = false)} />
@@ -124,6 +155,10 @@
<ExportDialog onClose={() => (showExportDialog = false)} />
{/if}
{#if session.processingStep !== 'idle' && session.processingStep !== 'done'}
<ProcessingModal />
{/if}
<div class="app-shell">
<header class="toolbar">
<UrlInput />
@@ -133,8 +168,21 @@
<main class="content">
<VideoPlayer />
<TransportControls bind:this={transportControls} />
<div class="timeline-clip-area" bind:this={splitAreaEl}>
<div class="timeline-pane" style="height: {timelineHeight}px">
<Timeline />
</div>
<div
class="resize-handle"
class:active={isResizing}
role="separator"
aria-orientation="horizontal"
onmousedown={handleResizeStart}
></div>
<div class="cliplist-pane">
<ClipList onExport={() => (showExportDialog = true)} />
</div>
</div>
</main>
<StatusBar />
@@ -163,6 +211,38 @@
overflow: hidden;
}
.timeline-clip-area {
display: flex;
flex-direction: column;
flex-shrink: 0;
min-height: 80px;
}
.timeline-pane {
flex-shrink: 0;
min-height: 40px;
overflow: hidden;
}
.resize-handle {
height: 5px;
background: var(--border, #353550);
cursor: ns-resize;
flex-shrink: 0;
transition: background 0.1s;
}
.resize-handle:hover,
.resize-handle.active {
background: var(--accent, #89b4fa);
}
.cliplist-pane {
flex: 1;
min-height: 20px;
overflow: hidden;
}
.prefs-btn {
font-size: 18px;
padding: 4px 8px;

View File

@@ -27,6 +27,10 @@ export async function checkDependencies(
return invoke<DependencyStatus[]>('check_dependencies', { cookieSource });
}
export async function checkSubtitlesFilterAvailable(): Promise<boolean> {
return invoke<boolean>('check_subtitles_filter_available');
}
export async function installDependency(
name: string,
onOutput: (line: string) => void,

View File

@@ -8,6 +8,8 @@ export interface VideoMetadata {
fps: number;
thumbnailUrl: string | null;
streamUrl: string;
hasCaptions: boolean;
captionsAreAuto: boolean;
}
type DownloadEvent =
@@ -16,6 +18,17 @@ type DownloadEvent =
| { event: 'finished'; data: { success: boolean; path: string } }
| { event: 'error'; data: { message: string } };
export async function getMediaServerPort(): Promise<number> {
return invoke<number>('get_media_server_port');
}
export async function checkCachedDownload(
title: string,
variant: string
): Promise<string | null> {
return invoke<string | null>('check_cached_download', { title, variant });
}
export async function resolveUrl(url: string): Promise<VideoMetadata> {
return invoke<VideoMetadata>('resolve_url', {
url,
@@ -23,8 +36,21 @@ export async function resolveUrl(url: string): Promise<VideoMetadata> {
});
}
export async function downloadSubtitles(
url: string,
isAuto: boolean
): Promise<string> {
return invoke<string>('download_subtitles', {
url,
cookieSource: preferences.cookieSource,
isAuto,
});
}
export async function startDownload(
url: string,
formatSpec: string,
variant: string,
onProgress: (percent: number) => void,
onFilePath: (path: string) => void,
onFinished: (success: boolean, path: string) => void,
@@ -54,6 +80,8 @@ export async function startDownload(
await invoke('start_download', {
url,
cookieSource: preferences.cookieSource,
formatSpec,
variant,
onEvent,
});
}

View File

@@ -33,9 +33,17 @@
function handleLabelEdit(clipId: string, value: string) {
updateClip(clipId, { label: value });
}
function handleContainerClick(e: MouseEvent) {
const target = e.target as HTMLElement;
// Only deselect if the click landed on the container itself, not a child
if (target.classList.contains('clip-list')) {
selectClip(null);
}
}
</script>
<div class="clip-list">
<div class="clip-list" role="listbox" onclick={handleContainerClick}>
{#if clips.length === 0}
<div class="empty">No clips yet — press I to mark in-point, O to mark out-point</div>
{:else}
@@ -46,7 +54,8 @@
<div
class="clip-row"
class:selected={clip.id === selectedId}
role="button"
role="option"
aria-selected={clip.id === selectedId}
tabindex="0"
onclick={() => handleSelect(clip.id)}
onkeydown={(e) => {
@@ -91,11 +100,14 @@
>✕</button>
</div>
{/each}
{#if onExport}
<div class="export-buttons">
<button type="button" onclick={() => onExport()}>Export All</button>
</div>
<div class="action-buttons">
{#if selectedId}
<button type="button" class="deselect-btn" onclick={() => selectClip(null)}>Deselect</button>
{/if}
{#if onExport}
<button type="button" onclick={() => onExport()}>Export All</button>
{/if}
</div>
{/if}
</div>
@@ -103,10 +115,11 @@
.clip-list {
padding: 8px 12px;
background: var(--bg-secondary);
min-height: 60px;
max-height: 200px;
height: 100%;
min-height: 40px;
overflow-y: auto;
border-top: 1px solid var(--border);
box-sizing: border-box;
}
.empty {
@@ -123,17 +136,24 @@
margin-bottom: 4px;
}
.export-buttons {
.action-buttons {
margin-top: 8px;
padding-top: 8px;
border-top: 1px solid var(--border);
display: flex;
gap: 8px;
}
.export-buttons button {
.action-buttons button {
font-size: 13px;
padding: 6px 12px;
}
.deselect-btn {
color: var(--text-secondary);
background: var(--bg-tertiary);
}
.clip-row {
display: flex;
align-items: center;

View File

@@ -0,0 +1,183 @@
<script lang="ts">
import { session, type ProcessingStep } from '$lib/stores/videoSession.svelte';
const STEP_LABELS: Record<ProcessingStep, string> = {
idle: '',
downloading: 'Downloading preview…',
waveform: 'Generating waveform…',
keyframes: 'Extracting keyframes…',
thumbnails: 'Generating thumbnails…',
captions: 'Downloading captions…',
done: 'Ready!',
};
const STEP_ORDER: ProcessingStep[] = [
'downloading',
'waveform',
'keyframes',
'thumbnails',
'captions',
'done',
];
let currentStepIndex = $derived(STEP_ORDER.indexOf(session.processingStep));
// Pick progress source based on current step
let stepProgress = $derived(() => {
switch (session.processingStep) {
case 'downloading':
return session.previewProgress;
case 'waveform':
case 'keyframes':
case 'thumbnails':
case 'captions':
return session.processingProgress;
case 'done':
return 1;
default:
return 0;
}
});
let progressPercent = $derived(Math.round(stepProgress() * 100));
let showProgressBar = $derived(
session.processingStep !== 'idle' && session.processingStep !== 'done'
);
</script>
<div class="modal-overlay">
<div class="modal">
<h2>Processing Video</h2>
<p class="title">{session.title}</p>
<div class="steps">
{#each STEP_ORDER as step, i}
{@const stepLabel = STEP_LABELS[step]}
{@const isDone = i < currentStepIndex}
{@const isCurrent = step === session.processingStep}
<div class="step" class:done={isDone} class:current={isCurrent}>
<span class="icon">
{#if isDone}
✓
{:else if isCurrent}
⏳
{:else}
○
{/if}
</span>
<span class="label">{stepLabel}</span>
{#if isCurrent && progressPercent > 0}
<span class="step-pct">{progressPercent}%</span>
{/if}
</div>
{/each}
</div>
{#if showProgressBar}
<div class="progress-bar">
<div class="progress-fill" style="width: {progressPercent}%"></div>
</div>
<p class="progress-text">{progressPercent}%</p>
{/if}
</div>
</div>
<style>
.modal-overlay {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.7);
display: flex;
align-items: center;
justify-content: center;
z-index: 1000;
}
.modal {
background: var(--bg-primary, #1e1e2e);
border: 1px solid var(--border, #353550);
border-radius: 12px;
padding: 32px 40px;
min-width: 360px;
max-width: 480px;
text-align: center;
box-shadow: 0 20px 60px rgba(0, 0, 0, 0.5);
}
h2 {
margin: 0 0 4px;
font-size: 18px;
font-weight: 600;
color: var(--text-primary, #cdd6f4);
}
.title {
margin: 0 0 24px;
font-size: 13px;
color: var(--text-secondary, #6c7086);
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.steps {
display: flex;
flex-direction: column;
gap: 8px;
text-align: left;
margin-bottom: 20px;
}
.step {
display: flex;
align-items: center;
gap: 10px;
font-size: 13px;
color: var(--text-secondary, #6c7086);
transition: color 0.2s;
}
.step.done {
color: #a6e3a1;
}
.step.current {
color: var(--text-primary, #cdd6f4);
font-weight: 500;
}
.icon {
width: 18px;
text-align: center;
flex-shrink: 0;
font-size: 14px;
}
.step-pct {
margin-left: auto;
font-size: 12px;
color: var(--text-secondary, #6c7086);
font-weight: 400;
}
.progress-bar {
height: 6px;
background: var(--bg-tertiary, #2a2a3e);
border-radius: 3px;
overflow: hidden;
margin-bottom: 6px;
}
.progress-fill {
height: 100%;
background: #89b4fa;
border-radius: 3px;
transition: width 0.3s ease;
}
.progress-text {
margin: 0;
font-size: 12px;
color: var(--text-secondary, #6c7086);
}
</style>

View File

@@ -1,32 +1,43 @@
<script lang="ts">
import { session } from '$lib/stores/videoSession.svelte';
import { session, upgradePreview } from '$lib/stores/videoSession.svelte';
let progressPercent = $derived(Math.round(session.downloadProgress * 100));
let isDownloading = $derived(
let previewPercent = $derived(Math.round(session.previewProgress * 100));
let exportPercent = $derived(Math.round(session.exportProgress * 100));
let isPreviewDownloading = $derived(
session.status === 'ready' &&
session.downloadStatus === 'downloading' &&
session.downloadProgress > 0 &&
session.downloadProgress < 1
);
let isDownloaded = $derived(
session.localFilePath !== null &&
(session.downloadStatus === 'complete' || session.downloadProgress >= 1)
session.previewStatus === 'downloading' &&
session.previewProgress > 0 &&
session.previewProgress < 1
);
let isPreviewComplete = $derived(session.previewStatus === 'complete');
let isExportDownloading = $derived(session.exportStatus === 'downloading');
let isExportComplete = $derived(session.exportStatus === 'complete');
</script>
<div class="status-bar">
{#if isDownloading}
{#if isPreviewDownloading}
<div class="progress-bar">
<div class="progress-fill" style="width: {progressPercent}%"></div>
<div class="progress-fill" style="width: {previewPercent}%"></div>
</div>
<span class="status-text">Downloading: {progressPercent}%</span>
{:else if isDownloaded}
<span class="status-text success">✓ Downloaded: {session.title}</span>
<span class="status-text">Preview: {previewPercent}%</span>
{:else if isPreviewComplete && isExportDownloading}
<span class="status-text success">✓ Preview ready</span>
<div class="progress-bar">
<div class="progress-fill" style="width: {exportPercent}%"></div>
</div>
<span class="status-text">Best quality: {exportPercent}%</span>
{:else if isPreviewComplete && isExportComplete}
<span class="status-text success">✓ {session.title} — ready to export</span>
{#if session.showUpgradeToast}
<button class="upgrade-link" onclick={upgradePreview}>Reload Preview (HQ)</button>
{/if}
{:else if session.status === 'resolving'}
<span class="status-text">Resolving URL…</span>
{:else if session.status === 'ready' && session.downloadStatus === 'downloading'}
<span class="status-text">Starting download…</span>
{:else if session.downloadStatus === 'failed'}
{:else if session.status === 'ready' && session.previewStatus === 'idle'}
<span class="status-text">Preparing download…</span>
{:else if session.previewStatus === 'failed' || session.exportStatus === 'failed'}
<span class="status-text">Download failed</span>
{:else}
<span class="status-text">Ready</span>
@@ -64,4 +75,20 @@
.success {
color: var(--success);
}
.upgrade-link {
background: none;
border: 1px solid var(--accent, #89b4fa);
color: var(--accent, #89b4fa);
font-size: 11px;
padding: 2px 8px;
border-radius: 3px;
cursor: pointer;
white-space: nowrap;
}
.upgrade-link:hover {
background: var(--accent, #89b4fa);
color: var(--bg-primary, #1e1e2e);
}
</style>

View File

@@ -8,15 +8,20 @@
import { drawTimeline, type TimelineState } from '$lib/timeline/renderer';
import {
computeZoom,
panBy,
handleClick as computeClickTime,
} from '$lib/timeline/interactions';
import { hitTestClip } from '$lib/timeline/clipRenderer';
import { drawWaveform, type WaveformData } from '$lib/timeline/waveformRenderer';
import { setThumbnailLoadCallback } from '$lib/timeline/thumbnailRenderer';
type ClipDragTarget = { clipId: string; edge: 'start' | 'end' };
let canvas = $state<HTMLCanvasElement | null>(null);
let minimapCanvas = $state<HTMLCanvasElement | null>(null);
let containerEl = $state<HTMLDivElement | null>(null);
let animFrameId = $state(0);
let pendingDraw = 0;
let pendingMinimapDraw = 0;
let timelineState = $state<TimelineState>({
visibleStart: 0,
@@ -28,6 +33,139 @@
let isDragging = $state(false);
let dragTarget = $state<ClipDragTarget | null>(null);
let isMinimapDragging = $state(false);
let minimapDragStartX = $state(0);
let minimapDragStartVisibleStart = $state(0);
let isZoomed = $derived(timelineState.zoom > 1.01);
// Construct waveform data from pre-computed tiers — pure reactive, no IPC
let waveformData = $derived<WaveformData>({
tiers: session.waveformTiers,
});
// Register thumbnail load callback
$effect(() => {
setThumbnailLoadCallback(() => requestDraw());
return () => setThumbnailLoadCallback(null);
});
function requestDraw() {
if (!pendingDraw) {
pendingDraw = requestAnimationFrame(() => {
drawMainCanvas();
pendingDraw = 0;
});
}
}
function requestMinimapDraw() {
if (!pendingMinimapDraw) {
pendingMinimapDraw = requestAnimationFrame(() => {
drawMinimapCanvas();
pendingMinimapDraw = 0;
});
}
}
function drawMainCanvas() {
if (!canvas) return;
const ctx = canvas.getContext('2d');
if (!ctx) return;
ctx.save();
ctx.scale(window.devicePixelRatio, window.devicePixelRatio);
drawTimeline(
ctx,
timelineState,
session.currentTime,
session.duration,
clipStore.clips,
clipStore.selectedClipId,
clipStore.pendingInPoint,
waveformData,
session.thumbnailSpritesheets
);
ctx.restore();
}
function drawMinimapCanvas() {
if (!minimapCanvas || !isZoomed || session.duration <= 0) return;
const ctx = minimapCanvas.getContext('2d');
if (!ctx) return;
const w = minimapCanvas.width / window.devicePixelRatio;
const h = minimapCanvas.height / window.devicePixelRatio;
ctx.save();
ctx.scale(window.devicePixelRatio, window.devicePixelRatio);
ctx.clearRect(0, 0, w, h);
ctx.fillStyle = '#12121e';
ctx.fillRect(0, 0, w, h);
// Minimap always shows the overview tier at full duration
if (session.waveformTiers) {
const fullState: TimelineState = {
visibleStart: 0,
visibleEnd: session.duration,
zoom: 1,
width: w,
height: h,
};
const minimapWaveform: WaveformData = {
tiers: session.waveformTiers,
};
drawWaveform(ctx, fullState, minimapWaveform, session.duration, 0, h);
}
const vpLeft = (timelineState.visibleStart / session.duration) * w;
const vpRight = (timelineState.visibleEnd / session.duration) * w;
ctx.fillStyle = 'rgba(0, 0, 0, 0.5)';
ctx.fillRect(0, 0, vpLeft, h);
ctx.fillRect(vpRight, 0, w - vpRight, h);
ctx.strokeStyle = '#89b4fa';
ctx.lineWidth = 1.5;
ctx.strokeRect(vpLeft, 0, vpRight - vpLeft, h);
const phX = (session.currentTime / session.duration) * w;
ctx.strokeStyle = '#ffffff';
ctx.lineWidth = 1;
ctx.beginPath();
ctx.moveTo(phX, 0);
ctx.lineTo(phX, h);
ctx.stroke();
ctx.restore();
}
// Single unified redraw effect — coalesced by requestAnimationFrame
$effect(() => {
void session.currentTime;
void session.duration;
void session.waveformTiers;
void session.thumbnailSpritesheets;
void clipStore.clips;
void clipStore.selectedClipId;
void clipStore.pendingInPoint;
void timelineState.visibleStart;
void timelineState.visibleEnd;
void timelineState.width;
void timelineState.height;
void canvas;
requestDraw();
});
// Minimap redraws
$effect(() => {
void session.currentTime;
void session.waveformTiers;
void timelineState.visibleStart;
void timelineState.visibleEnd;
void isZoomed;
void minimapCanvas;
requestMinimapDraw();
});
$effect(() => {
if (session.duration > 0) {
@@ -56,31 +194,21 @@
});
$effect(() => {
if (!canvas) return;
const ctx = canvas.getContext('2d');
if (!ctx) return;
function draw() {
if (!ctx || !canvas) return;
ctx.save();
ctx.scale(window.devicePixelRatio, window.devicePixelRatio);
drawTimeline(
ctx,
timelineState,
session.currentTime,
session.duration,
clipStore.clips,
clipStore.selectedClipId,
clipStore.pendingInPoint,
session.waveformPeaks,
session.thumbnailSpritesheets
);
ctx.restore();
animFrameId = requestAnimationFrame(draw);
if (!minimapCanvas || !containerEl) return;
const observer = new ResizeObserver((entries) => {
for (const entry of entries) {
if (minimapCanvas) {
const w = entry.contentRect.width;
const h = 28;
minimapCanvas.width = w * window.devicePixelRatio;
minimapCanvas.height = h * window.devicePixelRatio;
minimapCanvas.style.width = `${w}px`;
minimapCanvas.style.height = `${h}px`;
}
draw();
return () => cancelAnimationFrame(animFrameId);
}
});
observer.observe(containerEl);
return () => observer.disconnect();
});
function getCanvasX(e: MouseEvent): number {
@@ -119,6 +247,7 @@
return;
}
selectClip(null);
isDragging = true;
const targetTime = computeClickTime(x, timelineState, session.duration);
seekTo(targetTime);
@@ -156,6 +285,14 @@
function handleWheel(e: WheelEvent) {
e.preventDefault();
if (e.deltaX !== 0) {
const result = panBy(timelineState, e.deltaX, session.duration);
timelineState.visibleStart = result.visibleStart;
timelineState.visibleEnd = result.visibleEnd;
return;
}
const x = getCanvasX(e);
const anchorTime = computeClickTime(x, timelineState, session.duration);
const factor = e.deltaY < 0 ? 1.15 : 0.87;
@@ -164,8 +301,76 @@
timelineState.visibleEnd = result.visibleEnd;
timelineState.zoom = result.zoom;
}
function handleMinimapMouseDown(e: MouseEvent) {
if (!minimapCanvas || session.duration <= 0) return;
const rect = minimapCanvas.getBoundingClientRect();
const x = e.clientX - rect.left;
isMinimapDragging = true;
minimapDragStartX = x;
minimapDragStartVisibleStart = timelineState.visibleStart;
}
function handleMinimapMouseMove(e: MouseEvent) {
if (!isMinimapDragging || !minimapCanvas || session.duration <= 0) return;
const rect = minimapCanvas.getBoundingClientRect();
const x = e.clientX - rect.left;
const deltaX = x - minimapDragStartX;
const deltaTime = (deltaX / rect.width) * session.duration;
const range = timelineState.visibleEnd - timelineState.visibleStart;
let newStart = minimapDragStartVisibleStart + deltaTime;
let newEnd = newStart + range;
if (newStart < 0) {
newStart = 0;
newEnd = range;
}
if (newEnd > session.duration) {
newEnd = session.duration;
newStart = Math.max(0, newEnd - range);
}
timelineState.visibleStart = newStart;
timelineState.visibleEnd = newEnd;
}
function handleMinimapMouseUp() {
isMinimapDragging = false;
}
function handleMinimapClick(e: MouseEvent) {
if (!minimapCanvas || session.duration <= 0) return;
const rect = minimapCanvas.getBoundingClientRect();
const x = e.clientX - rect.left;
const clickTime = (x / rect.width) * session.duration;
const range = timelineState.visibleEnd - timelineState.visibleStart;
let newStart = clickTime - range / 2;
let newEnd = newStart + range;
if (newStart < 0) {
newStart = 0;
newEnd = range;
}
if (newEnd > session.duration) {
newEnd = session.duration;
newStart = Math.max(0, newEnd - range);
}
timelineState.visibleStart = newStart;
timelineState.visibleEnd = newEnd;
}
</script>
<div class="timeline-wrapper">
{#if isZoomed}
<div class="minimap">
<canvas
bind:this={minimapCanvas}
onmousedown={handleMinimapMouseDown}
onclick={handleMinimapClick}
></canvas>
</div>
{/if}
<div class="timeline-container" bind:this={containerEl}>
<canvas
bind:this={canvas}
@@ -176,10 +381,41 @@
onwheel={handleWheel}
></canvas>
</div>
</div>
<svelte:window
onmousemove={handleMinimapMouseMove}
onmouseup={handleMinimapMouseUp}
/>
<style>
.timeline-wrapper {
display: flex;
flex-direction: column;
height: 100%;
}
.minimap {
height: 28px;
background: #12121e;
border-bottom: 1px solid var(--border, #353550);
cursor: grab;
flex-shrink: 0;
}
.minimap:active {
cursor: grabbing;
}
.minimap canvas {
display: block;
width: 100%;
height: 100%;
}
.timeline-container {
height: 160px;
flex: 1;
min-height: 80px;
background: var(--timeline-bg);
position: relative;
cursor: crosshair;

View File

@@ -8,11 +8,35 @@
seekBy,
stepFrame,
togglePlayPause,
setVolume,
setMuted,
getVolume,
getMuted,
type TransportKeyAction,
} from '$lib/transport/playback';
let hasKeyframes = $derived(session.keyframePositions.length > 0);
let volume = $state(1);
let isMuted = $state(false);
function toggleMute() {
isMuted = !isMuted;
setMuted(isMuted);
}
function handleVolumeChange(e: Event) {
const val = parseFloat((e.target as HTMLInputElement).value);
volume = val;
if (val > 0 && isMuted) {
isMuted = false;
}
setVolume(val);
if (val > 0) {
setMuted(false);
}
}
export function handleKeyAction(action: TransportKeyAction) {
runTransportAction(action, markInPoint, markOutPoint);
}
@@ -53,6 +77,32 @@
<span class="total">{formatTime(session.duration)}</span>
</div>
<div class="volume-controls">
<button
class="vol-btn"
onclick={toggleMute}
title={isMuted ? 'Unmute' : 'Mute'}
type="button"
>
{#if isMuted || volume === 0}
🔇
{:else if volume < 0.5}
🔉
{:else}
🔊
{/if}
</button>
<input
type="range"
class="vol-slider"
min="0"
max="1"
step="0.05"
value={isMuted ? 0 : volume}
oninput={handleVolumeChange}
/>
</div>
<div class="mark-buttons">
<button
class="mark-btn"
@@ -114,6 +164,34 @@
color: var(--text-secondary);
}
.volume-controls {
display: flex;
align-items: center;
gap: 4px;
}
.vol-btn {
background: var(--bg-tertiary);
border: 1px solid var(--border);
border-radius: 4px;
padding: 3px 6px;
font-size: 14px;
cursor: pointer;
line-height: 1;
color: var(--text-secondary);
}
.vol-btn:hover {
color: var(--text-primary);
}
.vol-slider {
width: 80px;
height: 4px;
accent-color: #89b4fa;
cursor: pointer;
}
.mark-buttons {
margin-left: auto;
display: flex;

View File

@@ -9,11 +9,17 @@
} from '$lib/stores/videoSession.svelte';
let inputValue = $state('');
let inputEl = $state<HTMLInputElement | null>(null);
function blurInput() {
inputEl?.blur();
}
async function handleSubmit() {
const trimmed = inputValue.trim();
if (!trimmed) return;
blurInput();
setResolving();
try {
const meta = await resolveUrl(trimmed);
@@ -43,6 +49,7 @@
<div class="url-input-wrapper">
<div class="url-input">
<input
bind:this={inputEl}
type="text"
placeholder="Paste a video URL and press Enter…"
bind:value={inputValue}

View File

@@ -8,6 +8,8 @@ export interface Clip {
color: string;
}
const EDIT_TOLERANCE = 0.5; // seconds — how close the playhead must be to a clip edge to count as "inside"
export const clipStore = $state({
clips: [] as Clip[],
selectedClipId: null as string | null,
@@ -27,10 +29,21 @@ export function getPendingInPoint(): number | null {
return clipStore.pendingInPoint;
}
function isTimeInsideClip(time: number, clipId: string): boolean {
const clip = clipStore.clips.find((c) => c.id === clipId);
if (!clip) return false;
return time >= clip.startTime - EDIT_TOLERANCE && time <= clip.endTime + EDIT_TOLERANCE;
}
export function selectClip(id: string | null) {
clipStore.selectedClipId = id;
// Only clear pendingInPoint when selecting a specific clip.
// When deselecting (id=null), preserve it so clicking the timeline
// to seek after pressing I doesn't lose the pending in-point.
if (id !== null) {
clipStore.pendingInPoint = null;
}
}
export function addClip(startTime: number, endTime: number): Clip {
const start = Math.min(startTime, endTime);
@@ -71,20 +84,32 @@ export function updateClip(
}
export function markInPoint(time: number) {
if (clipStore.selectedClipId) {
if (
clipStore.selectedClipId &&
isTimeInsideClip(time, clipStore.selectedClipId)
) {
// Playhead is inside the selected clip — edit its in-point
updateClip(clipStore.selectedClipId, { startTime: time });
} else {
// Playhead is outside — deselect and begin a new clip
clipStore.selectedClipId = null;
clipStore.pendingInPoint = time;
}
}
export function markOutPoint(time: number) {
if (clipStore.selectedClipId) {
if (
clipStore.selectedClipId &&
isTimeInsideClip(time, clipStore.selectedClipId)
) {
// Playhead is inside the selected clip — edit its out-point
updateClip(clipStore.selectedClipId, { endTime: time });
} else if (clipStore.pendingInPoint !== null) {
// We have a pending in-point — create a new clip
addClip(clipStore.pendingInPoint, time);
clipStore.pendingInPoint = null;
}
// Otherwise (no selection, no pending in-point) — no-op
}
export function clearAll() {

View File

@@ -1,17 +1,33 @@
import type { VideoMetadata } from '$lib/bindings/video';
import { startDownload } from '$lib/bindings/video';
import type { ThumbnailSpritesheet } from '$lib/bindings/mediaAnalysis';
import { checkCachedDownload, downloadSubtitles, startDownload } from '$lib/bindings/video';
import type { ThumbnailSpritesheet, WaveformTiers } from '$lib/bindings/mediaAnalysis';
import {
checkEmbeddedSubtitles,
cleanupThumbnails,
extractKeyframes,
extractThumbnails,
extractWaveform,
extractWaveformTiers,
} from '$lib/bindings/mediaAnalysis';
import { clearThumbnailCache } from '$lib/timeline/thumbnailRenderer';
import { clearAll as clearAllClips } from '$lib/stores/clips.svelte';
export type SessionStatus = 'idle' | 'resolving' | 'ready' | 'error';
export type DownloadStatus = 'idle' | 'downloading' | 'complete' | 'failed';
export type ProcessingStep =
| 'idle'
| 'downloading'
| 'waveform'
| 'keyframes'
| 'thumbnails'
| 'captions'
| 'done';
// H.264+AAC ≤360p for fast preview/scrubbing in WKWebView
const PREVIEW_FORMAT =
'bv*[vcodec^=avc1][height<=360]+ba[acodec^=mp4a]/b[ext=mp4][height<=360]/worst[ext=mp4]/worst';
// Best available H.264+AAC in MP4 for export — WKWebView can't play AV1/VP9/WebM
const EXPORT_FORMAT =
'bv*[vcodec^=avc1]+ba[acodec^=mp4a]/b[ext=mp4]/best[ext=mp4]';
export const session = $state({
status: 'idle' as SessionStatus,
@@ -24,12 +40,35 @@ export const session = $state({
streamUrl: '',
currentTime: 0,
isPlaying: false,
localFilePath: null as string | null,
downloadProgress: 0,
downloadStatus: 'idle' as DownloadStatus,
// Preview (low-res for playback/scrubbing/waveform/thumbnails)
previewFilePath: null as string | null,
previewProgress: 0,
previewStatus: 'idle' as DownloadStatus,
// Export (best quality for clip extraction)
exportFilePath: null as string | null,
exportProgress: 0,
exportStatus: 'idle' as DownloadStatus,
keyframePositions: [] as number[],
waveformPeaks: [] as number[],
waveformTiers: null as WaveformTiers | null,
thumbnailSpritesheets: [] as ThumbnailSpritesheet[],
// Captions
hasCaptions: false,
captionsAreAuto: false,
captionFilePath: null as string | null,
// The path used for the <video> element (may be preview or export quality)
activeVideoPath: null as string | null,
// Preview upgrade toast
showUpgradeToast: false,
// Processing modal state
processingStep: 'idle' as ProcessingStep,
processingProgress: 0,
});
export function setMetadata(meta: VideoMetadata) {
@@ -40,15 +79,25 @@ export function setMetadata(meta: VideoMetadata) {
session.fps = meta.fps;
session.thumbnailUrl = meta.thumbnailUrl;
session.streamUrl = meta.streamUrl;
session.hasCaptions = meta.hasCaptions;
session.captionsAreAuto = meta.captionsAreAuto;
session.status = 'ready';
session.error = null;
session.localFilePath = null;
session.downloadProgress = 0;
session.downloadStatus = 'idle';
session.captionFilePath = null;
session.previewFilePath = null;
session.previewProgress = 0;
session.previewStatus = 'idle';
session.exportFilePath = null;
session.exportProgress = 0;
session.exportStatus = 'idle';
session.keyframePositions = [];
session.waveformPeaks = [];
session.waveformTiers = null;
session.thumbnailSpritesheets = [];
session.activeVideoPath = null;
session.showUpgradeToast = false;
session.processingStep = 'idle';
session.processingProgress = 0;
}
function clearMediaFields() {
@@ -60,12 +109,22 @@ function clearMediaFields() {
session.streamUrl = '';
session.currentTime = 0;
session.isPlaying = false;
session.localFilePath = null;
session.downloadProgress = 0;
session.downloadStatus = 'idle';
session.previewFilePath = null;
session.previewProgress = 0;
session.previewStatus = 'idle';
session.exportFilePath = null;
session.exportProgress = 0;
session.exportStatus = 'idle';
session.keyframePositions = [];
session.waveformPeaks = [];
session.waveformTiers = null;
session.thumbnailSpritesheets = [];
session.hasCaptions = false;
session.captionsAreAuto = false;
session.captionFilePath = null;
session.activeVideoPath = null;
session.showUpgradeToast = false;
session.processingStep = 'idle';
session.processingProgress = 0;
}
export function setError(msg: string) {
@@ -87,45 +146,112 @@ export function reset() {
session.error = null;
}
/** Switch the preview video to the full-quality export version. */
export function upgradePreview() {
if (session.exportFilePath) {
console.log('[upgradePreview] switching from', session.activeVideoPath, 'to', session.exportFilePath);
session.activeVideoPath = session.exportFilePath;
session.showUpgradeToast = false;
}
}
export function dismissUpgradeToast() {
session.showUpgradeToast = false;
}
async function loadCaptions(filePath: string) {
session.processingStep = 'captions';
// Priority 1: Check for embedded subtitle streams in the file
try {
const embeddedPath = await checkEmbeddedSubtitles(filePath);
if (embeddedPath) {
session.captionFilePath = embeddedPath;
return;
}
} catch (err) {
console.error('Embedded subtitle check failed:', err);
}
// Priority 2: Download external subtitles via yt-dlp (manual > auto-generated)
if (session.hasCaptions && session.url) {
try {
const path = await downloadSubtitles(session.url, session.captionsAreAuto);
session.captionFilePath = path;
} catch (err) {
console.error('Subtitle download failed:', err);
}
}
}
async function triggerPostDownloadProcessing(filePath: string) {
const duration = session.duration;
const [keyframes, waveform, thumbnails] = await Promise.allSettled([
extractKeyframes(filePath),
extractWaveform(filePath, 8000),
extractThumbnails(filePath, duration),
]);
if (keyframes.status === 'fulfilled') {
session.keyframePositions = keyframes.value;
} else {
console.error('Keyframe extraction failed:', keyframes.reason);
// Run processing SEQUENTIALLY to avoid overloading the CPU with
// concurrent ffmpeg subprocesses. Each one decodes the full file.
// Order: waveform first (most useful for editing), then keyframes, then thumbnails.
session.processingStep = 'waveform';
session.processingProgress = 0;
try {
const tiers = await extractWaveformTiers(filePath, duration, (pct) => {
session.processingProgress = pct;
});
session.waveformTiers = tiers;
} catch (err) {
console.error('Waveform extraction failed:', err);
}
if (waveform.status === 'fulfilled') {
session.waveformPeaks = waveform.value;
} else {
console.error('Waveform extraction failed:', waveform.reason);
session.processingStep = 'keyframes';
session.processingProgress = 0;
try {
const kf = await extractKeyframes(filePath);
session.keyframePositions = kf;
} catch (err) {
console.error('Keyframe extraction failed:', err);
}
if (thumbnails.status === 'fulfilled') {
session.thumbnailSpritesheets = thumbnails.value;
} else {
console.error('Thumbnail extraction failed:', thumbnails.reason);
session.processingStep = 'thumbnails';
session.processingProgress = 0;
try {
const sheets = await extractThumbnails(filePath, duration);
session.thumbnailSpritesheets = sheets;
} catch (err) {
console.error('Thumbnail extraction failed:', err);
}
// Captions can run after analysis since they use yt-dlp, not ffmpeg
await loadCaptions(filePath);
session.processingStep = 'done';
// If export download already completed while we were processing, show upgrade toast
if (session.exportStatus === 'complete' && session.exportFilePath) {
session.showUpgradeToast = true;
}
}
async function runDownload(
url: string,
formatSpec: string,
outputDir: string,
onProgress: (pct: number) => void,
onFilePath: (path: string) => void,
onFinished: (success: boolean, path: string) => void,
onError: (message: string) => void,
): Promise<void> {
await startDownload(url, formatSpec, outputDir, onProgress, onFilePath, onFinished, onError);
}
export async function beginDownload() {
if (!session.url) return;
if (!session.url || !session.title) return;
const oldThumbnailDirs = [
...new Set(session.thumbnailSpritesheets.map((sheet) => sheet.filePath)),
];
clearThumbnailCache();
session.downloadStatus = 'downloading';
session.downloadProgress = 0;
session.localFilePath = null;
session.keyframePositions = [];
session.waveformPeaks = [];
session.waveformTiers = null;
session.thumbnailSpritesheets = [];
if (oldThumbnailDirs.length > 0) {
@@ -134,32 +260,90 @@ export async function beginDownload() {
});
}
session.processingStep = 'downloading';
// --- Preview download (low-res, fast) ---
const previewCached = await checkCachedDownload(session.title, 'preview').catch(() => null);
if (previewCached) {
session.previewFilePath = previewCached;
session.activeVideoPath = previewCached;
session.previewProgress = 1.0;
session.previewStatus = 'complete';
triggerPostDownloadProcessing(previewCached);
} else {
session.previewStatus = 'downloading';
session.previewProgress = 0;
session.previewFilePath = null;
try {
await startDownload(
await runDownload(
session.url,
(percent) => {
session.downloadProgress = percent;
},
(path) => {
session.localFilePath = path;
},
PREVIEW_FORMAT,
'preview',
(pct) => { session.previewProgress = pct; },
(path) => { session.previewFilePath = path; },
(success, path) => {
if (success) {
session.downloadProgress = 1.0;
session.localFilePath = path;
session.downloadStatus = 'complete';
void triggerPostDownloadProcessing(path);
session.previewProgress = 1.0;
session.previewFilePath = path;
session.activeVideoPath = path;
session.previewStatus = 'complete';
triggerPostDownloadProcessing(path);
} else {
session.downloadStatus = 'failed';
session.previewStatus = 'failed';
}
},
(message) => {
session.downloadStatus = 'failed';
console.error('Download error:', message);
session.previewStatus = 'failed';
console.error('Preview download error:', message);
}
);
} catch (e) {
session.downloadStatus = 'failed';
console.error('Download failed:', e);
session.previewStatus = 'failed';
console.error('Preview download failed:', e);
}
}
// --- Export download (best quality, background) ---
const exportCached = await checkCachedDownload(session.title, 'export').catch(() => null);
if (exportCached) {
session.exportFilePath = exportCached;
session.exportProgress = 1.0;
session.exportStatus = 'complete';
if ((session.processingStep as ProcessingStep) === 'done') {
session.showUpgradeToast = true;
}
} else {
session.exportStatus = 'downloading';
session.exportProgress = 0;
session.exportFilePath = null;
// Fire-and-forget — runs in background
runDownload(
session.url,
EXPORT_FORMAT,
'export',
(pct) => { session.exportProgress = pct; },
(path) => { session.exportFilePath = path; },
(success, path) => {
if (success) {
session.exportProgress = 1.0;
session.exportFilePath = path;
session.exportStatus = 'complete';
if ((session.processingStep as ProcessingStep) === 'done') {
session.showUpgradeToast = true;
}
} else {
session.exportStatus = 'failed';
}
},
(message) => {
session.exportStatus = 'failed';
console.error('Export download error:', message);
}
).catch((e) => {
session.exportStatus = 'failed';
console.error('Export download failed:', e);
});
}
}

View File

@@ -45,6 +45,28 @@ export function computeZoom(
return { visibleStart: newStart, visibleEnd: newEnd, zoom: newZoom };
}
export function panBy(
state: TimelineState,
deltaPixels: number,
duration: number
): { visibleStart: number; visibleEnd: number } {
const range = state.visibleEnd - state.visibleStart;
const deltaTime = (deltaPixels / state.width) * range;
let newStart = state.visibleStart + deltaTime;
let newEnd = newStart + range;
if (newStart < 0) {
newStart = 0;
newEnd = range;
}
if (newEnd > duration) {
newEnd = duration;
newStart = Math.max(0, newEnd - range);
}
return { visibleStart: newStart, visibleEnd: newEnd };
}
export function handleClick(
x: number,
state: TimelineState,

View File

@@ -1,6 +1,6 @@
import { drawClips } from './clipRenderer';
import type { Clip } from '$lib/stores/clips.svelte';
import { drawWaveform } from './waveformRenderer';
import { drawWaveform, type WaveformData } from './waveformRenderer';
import { drawThumbnails } from './thumbnailRenderer';
import type { ThumbnailSpritesheet } from '$lib/bindings/mediaAnalysis';
@@ -36,7 +36,7 @@ export function drawTimeline(
clips: Clip[] = [],
selectedClipId: string | null = null,
pendingInPoint: number | null = null,
waveformPeaks: number[] = [],
waveform: WaveformData = { tiers: null },
thumbnailSpritesheets: ThumbnailSpritesheet[] = []
): void {
const { width, height } = state;
@@ -53,8 +53,8 @@ export function drawTimeline(
drawPlaceholderLane(ctx, state, 0, THUMB_LANE_HEIGHT, 'Thumbnails');
}
if (waveformPeaks.length > 0) {
drawWaveform(ctx, state, waveformPeaks, duration, THUMB_LANE_HEIGHT, WAVEFORM_LANE_HEIGHT);
if (waveform.tiers) {
drawWaveform(ctx, state, waveform, duration, THUMB_LANE_HEIGHT, WAVEFORM_LANE_HEIGHT);
} else {
drawPlaceholderLane(ctx, state, THUMB_LANE_HEIGHT, WAVEFORM_LANE_HEIGHT, 'Waveform');
}

View File

@@ -3,9 +3,17 @@ import { convertFileSrc } from '@tauri-apps/api/core';
import { timeToX, type TimelineState } from './renderer';
const imageCache = new Map<string, HTMLImageElement>();
const pendingLoads = new Set<string>();
let onLoadCallback: (() => void) | null = null;
export function clearThumbnailCache(): void {
imageCache.clear();
pendingLoads.clear();
}
/** Register a callback to be invoked when any thumbnail finishes loading */
export function setThumbnailLoadCallback(cb: (() => void) | null): void {
onLoadCallback = cb;
}
const TARGET_SPACING_PX = 100;
@@ -16,14 +24,21 @@ function getImagePath(dir: string, frameIndex: number): string {
}
function loadImage(path: string): void {
if (imageCache.has(path)) {
if (imageCache.has(path) || pendingLoads.has(path)) {
return;
}
pendingLoads.add(path);
const img = new Image();
img.src = convertFileSrc(path);
img.onload = () => imageCache.set(path, img);
img.onload = () => {
imageCache.set(path, img);
pendingLoads.delete(path);
onLoadCallback?.();
};
img.onerror = () => {
pendingLoads.delete(path);
};
}
function findSheetForFrame(

View File

@@ -1,17 +1,64 @@
import type { WaveformTiers } from '$lib/bindings/mediaAnalysis';
import type { TimelineState } from './renderer';
export interface WaveformData {
tiers: WaveformTiers | null;
}
/**
* Pick the best tier for the current viewport.
* Returns the tier array that gives at least 1 peak per pixel,
* preferring the lowest-resolution tier that meets that threshold.
*/
function pickTier(
tiers: WaveformTiers,
visibleStart: number,
visibleEnd: number,
duration: number,
widthPx: number
): number[] {
const visibleRange = visibleEnd - visibleStart;
const visibleFraction = visibleRange / duration;
const candidates: [number[], number][] = [
[tiers.tier0, (tiers.tier0.length * visibleFraction) / widthPx],
[tiers.tier1, (tiers.tier1.length * visibleFraction) / widthPx],
[tiers.tier2, (tiers.tier2.length * visibleFraction) / widthPx],
];
for (const [tier, peaksPerPx] of candidates) {
if (peaksPerPx >= 1) return tier;
}
return tiers.tier2;
}
export function drawWaveform(
ctx: CanvasRenderingContext2D,
state: TimelineState,
peaks: number[],
waveform: WaveformData,
duration: number,
y: number,
height: number
): void {
if (peaks.length === 0 || duration <= 0) return;
if (duration <= 0 || !waveform.tiers) return;
const { width, visibleStart, visibleEnd } = state;
const visibleRange = visibleEnd - visibleStart;
const peaks = pickTier(waveform.tiers, visibleStart, visibleEnd, duration, width);
const peakCount = peaks.length;
// Find local maximum across the visible region for dynamic scaling.
// This makes the waveform fill the available height when zoomed in,
// rather than staying relative to the global max.
const visStartIdx = Math.max(0, Math.floor((visibleStart / duration) * peakCount));
const visEndIdx = Math.min(peakCount - 1, Math.ceil((visibleEnd / duration) * peakCount));
let localMax = 0;
for (let i = visStartIdx; i <= visEndIdx; i++) {
localMax = Math.max(localMax, peaks[i]);
}
// Avoid division by zero; use 1.0 if all peaks are silent
const scale = localMax > 0 ? 1 / localMax : 1;
ctx.fillStyle = '#89b4fa44';
ctx.strokeStyle = '#89b4fa88';
@@ -21,19 +68,18 @@ export function drawWaveform(
ctx.moveTo(0, y + height);
for (let px = 0; px < width; px++) {
const startSample = Math.floor(
((visibleStart + ((px - 0.5) / width) * visibleRange) / duration) * peaks.length
);
const endSample = Math.floor(
((visibleStart + ((px + 0.5) / width) * visibleRange) / duration) * peaks.length
);
const timeStart = visibleStart + ((px - 0.5) / width) * visibleRange;
const timeEnd = visibleStart + ((px + 0.5) / width) * visibleRange;
const s0 = Math.floor((timeStart / duration) * peakCount);
const s1 = Math.floor((timeEnd / duration) * peakCount);
let maxPeak = 0;
for (let i = Math.max(0, startSample); i <= Math.min(endSample, peaks.length - 1); i++) {
for (let i = Math.max(0, s0); i <= Math.min(s1, peakCount - 1); i++) {
maxPeak = Math.max(maxPeak, peaks[i]);
}
const barHeight = maxPeak * height;
const barHeight = maxPeak * scale * height;
ctx.lineTo(px, y + height - barHeight);
}

View File

@@ -1,9 +1,19 @@
import { session } from '$lib/stores/videoSession.svelte';
let _videoEl: HTMLVideoElement | null = null;
export function setVideoElement(el: HTMLVideoElement | null) {
_videoEl = el;
}
function getVideo(): HTMLVideoElement | null {
return _videoEl;
}
export function seekTo(time: number) {
const clamped = Math.max(0, Math.min(time, session.duration));
session.currentTime = clamped;
const videoEl = document.querySelector('video');
const videoEl = getVideo();
if (videoEl) {
videoEl.currentTime = clamped;
}
@@ -14,7 +24,7 @@ export function seekBy(seconds: number) {
}
export function togglePlayPause() {
const videoEl = document.querySelector('video');
const videoEl = getVideo();
if (!videoEl) return;
if (videoEl.paused) {
void videoEl.play();
@@ -24,7 +34,7 @@ export function togglePlayPause() {
}
export function stepFrame(direction: 1 | -1) {
const videoEl = document.querySelector('video');
const videoEl = getVideo();
if (videoEl && !videoEl.paused) {
videoEl.pause();
}
@@ -44,15 +54,38 @@ export function jumpKeyframe(direction: 1 | -1) {
}
}
export function setVolume(v: number) {
const videoEl = getVideo();
if (videoEl) {
videoEl.volume = v;
videoEl.muted = false;
}
}
export function setMuted(m: boolean) {
const videoEl = getVideo();
if (videoEl) {
videoEl.muted = m;
}
}
export function getVolume(): number {
return getVideo()?.volume ?? 1;
}
export function getMuted(): boolean {
return getVideo()?.muted ?? false;
}
export function resetShuttleRate() {
const videoEl = document.querySelector('video');
const videoEl = getVideo();
if (videoEl) {
videoEl.playbackRate = 1;
}
}
export function adjustShuttle(dir: 1 | -1, shuttleRate: number): number {
const videoEl = document.querySelector('video');
const videoEl = getVideo();
if (!videoEl) return shuttleRate;
const nextRate = Math.max(0.25, Math.min(4, shuttleRate + dir * 0.5));

142
src/lib/utils/vttParser.ts Normal file
View File

@@ -0,0 +1,142 @@
export interface WordSegment {
text: string;
startTime: number;
}
export interface VttCue {
startTime: number;
endTime: number;
text: string;
words?: WordSegment[];
}
/**
* Parse a WebVTT string into an array of cues.
* Handles both standard VTT and common variations.
* Extracts word-level timing data from YouTube-style <c> tags when present.
*/
export function parseVtt(vttContent: string): VttCue[] {
const cues: VttCue[] = [];
const blocks = vttContent.replace(/\r\n/g, '\n').split(/\n\n+/);
for (const block of blocks) {
const lines = block.trim().split('\n');
let timestampLineIdx = -1;
for (let i = 0; i < lines.length; i++) {
if (lines[i].includes('-->')) {
timestampLineIdx = i;
break;
}
}
if (timestampLineIdx === -1) continue;
const timestampLine = lines[timestampLineIdx];
const match = timestampLine.match(
/(\d{1,2}:?\d{2}:\d{2}[.,]\d{3})\s*-->\s*(\d{1,2}:?\d{2}:\d{2}[.,]\d{3})/
);
if (!match) continue;
const startTime = parseTimestamp(match[1]);
const endTime = parseTimestamp(match[2]);
const rawText = lines
.slice(timestampLineIdx + 1)
.join('\n')
.trim();
const words = parseWordTimings(rawText, startTime);
const text = rawText.replace(/<[^>]+>/g, '').trim();
if (text && startTime < endTime) {
cues.push({ startTime, endTime, text, words: words.length > 0 ? words : undefined });
}
}
return cues;
}
/**
* Parse word-level timing data from YouTube-style VTT cue text.
*
* YouTube format: `the<00:00:00.599><c> gym</c><00:00:00.840><c> I</c>`
*
* The first word(s) before any timestamp tag start at `cueStartTime`.
* Each `<HH:MM:SS.mmm>` sets the start time for the text in the following `<c>...</c>`.
*/
function parseWordTimings(rawText: string, cueStartTime: number): WordSegment[] {
// Only parse if the text contains the YouTube-style <c> word timing pattern
if (!rawText.includes('<c>')) return [];
const segments: WordSegment[] = [];
// Pattern to match: optional leading text, then repeated <timestamp><c> word</c> groups
// We process the raw text character by character using regex matches
// Match all timestamp + <c>word</c> pairs, plus any leading text
const timestampCPattern = /<(\d{1,2}:?\d{2}:\d{2}[.,]\d{3})><c>(.*?)<\/c>/g;
// Get any leading text before the first timestamp
const firstTimestampMatch = rawText.match(/<\d{1,2}:?\d{2}:\d{2}[.,]\d{3}>/);
if (firstTimestampMatch && firstTimestampMatch.index !== undefined && firstTimestampMatch.index > 0) {
const leadingText = rawText.substring(0, firstTimestampMatch.index).replace(/<[^>]+>/g, '').trim();
if (leadingText) {
segments.push({ text: leadingText, startTime: cueStartTime });
}
}
let m: RegExpExecArray | null;
while ((m = timestampCPattern.exec(rawText)) !== null) {
const wordTime = parseTimestamp(m[1]);
const wordText = m[2].replace(/<[^>]+>/g, '').trim();
if (wordText) {
segments.push({ text: wordText, startTime: wordTime });
}
}
return segments;
}
/**
* Parse a VTT/SRT timestamp like "00:01:23.456" or "1:23.456" into seconds.
*/
function parseTimestamp(ts: string): number {
const normalized = ts.replace(',', '.');
const parts = normalized.split(':');
if (parts.length === 3) {
return (
parseInt(parts[0], 10) * 3600 +
parseInt(parts[1], 10) * 60 +
parseFloat(parts[2])
);
} else if (parts.length === 2) {
return parseInt(parts[0], 10) * 60 + parseFloat(parts[1]);
}
return parseFloat(normalized);
}
/**
* Find all active cues at the given time.
* Assumes cues are sorted by startTime (VTT spec requires this).
*/
export function getActiveCues(cues: VttCue[], currentTime: number): VttCue[] {
return cues.filter(
(cue) => currentTime >= cue.startTime && currentTime < cue.endTime
);
}
/**
* Given a cue with word segments and the current playback time,
* return which words are "spoken" (their startTime <= currentTime)
* and which are "upcoming" (their startTime > currentTime).
*/
export function getActiveWords(
cue: VttCue,
currentTime: number
): { text: string; spoken: boolean }[] {
if (!cue.words || cue.words.length === 0) {
return [{ text: cue.text, spoken: true }];
}
return cue.words.map((word) => ({
text: word.text,
spoken: currentTime >= word.startTime,
}));
}