chore: stage all pending work — caption styling, media server, processing modal, docs, summaries
Includes: - Extended caption styling (font, shadow, dimmed color, bg toggle) - Media server, subtitle downloader, VTT parser, processing modal - Waveform tiers, thumbnail/timeline improvements, transport controls - Hybrid download model, dependency management, clip export enhancements - 21 chat summaries, 2 implementation plans, 2 design specs Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# Fix: Video Playback, Waveform, and Thumbnails Not Working
|
||||
|
||||
**Date:** 2026-09-21 12:21
|
||||
**Task:** Diagnose and fix three user-reported issues after downloading a video: no playback, waveform/thumbnails stuck at "Generating...", scrubbing doesn't update the player frame.
|
||||
|
||||
## Root Causes
|
||||
|
||||
### 1. `[Merger]` line path not parsed (Primary Bug)
|
||||
When `yt-dlp` downloads video+audio separately and merges them, it outputs:
|
||||
```
|
||||
[download] Destination: /path/to/file.f399.mp4 ← intermediate
|
||||
[download] Destination: /path/to/file.f251.webm ← intermediate
|
||||
[Merger] Merging formats into "/path/to/file.webm" ← final merged file
|
||||
```
|
||||
The old parsing used `line.split(": ").nth(1)` for all three cases. This works for `[download] Destination:` lines (which have `: ` delimiter) but **fails for `[Merger]`** lines (which use `into "path"` format — no `: `). As a result, `last_file_path` pointed to a deleted intermediate file. Both `triggerPostDownloadProcessing(path)` and `session.localFilePath` received the wrong path, causing:
|
||||
- All ffmpeg post-processing (waveform, thumbnails, keyframes) to fail with file-not-found
|
||||
- The video player to reference a non-existent file
|
||||
|
||||
### 2. Stream URL doesn't work in Tauri webview
|
||||
The `yt-dlp -g` stream URL (googlevideo.com) has anti-hotlinking protections (referer/IP checks) that block playback in a Tauri/WKWebView context. The `<video>` element was present but silently failed to load the source.
|
||||
|
||||
### 3. "Already downloaded" case not handled
|
||||
When `yt-dlp` finds an existing file, it outputs `[download] /path/file.webm has already been downloaded` — a format that matches neither `Destination:` nor `Merger` patterns. So re-loading the same URL would result in `localFilePath = ""`.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### `src-tauri/src/commands/video.rs`
|
||||
- **Fixed `[Merger]` parsing:** Replaced the broken `split(": ")` approach with `trim_start_matches("[Merger] Merging formats into ").trim().trim_matches('"')`.
|
||||
- **Added "already downloaded" parsing:** New branch for `has already been downloaded` using `strip_prefix`/`strip_suffix` to extract the path.
|
||||
|
||||
### `src/lib/components/VideoPlayer.svelte`
|
||||
- **Switched to local file for playback:** Uses `convertFileSrc(session.localFilePath)` via Tauri's asset protocol once download is complete, instead of the unreliable stream URL.
|
||||
- **Added download progress UI:** Shows a progress bar with percentage while downloading, instead of a broken black video area.
|
||||
- **Added video error handling:** `onerror`/`onloadeddata` handlers with an overlay error message.
|
||||
- **Import:** Added `convertFileSrc` from `@tauri-apps/api/core`.
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
1. **`yt-dlp` stderr output formats are inconsistent.** `[download] Destination:` uses `: ` delimiter, `[Merger]` uses `into "path"`, and "already downloaded" embeds the path mid-sentence. Each needs its own parser.
|
||||
2. **googlevideo.com stream URLs don't work in embedded webviews** due to anti-hotlinking. The design's "hybrid streaming preview" approach needs a local proxy or must fall back to showing download progress until the file is available locally.
|
||||
3. **Always test with cached files.** The "already downloaded" edge case only surfaces on the second load of the same URL.
|
||||
|
||||
## Follow-up Items
|
||||
|
||||
- [ ] Add progress indicators for post-download processing (waveform/thumbnail/keyframe extraction can take minutes on long videos)
|
||||
- [ ] Consider pre-download preview via local proxy or lower-quality quick download
|
||||
- [ ] The waveform extraction uses `aresample=8000` (sample rate) which produces ~26M samples for a 54-min video — consider optimizing to reduce to target count directly
|
||||
- [ ] Add error surfacing for failed post-download processing (currently silent `console.error`)
|
||||
@@ -0,0 +1,54 @@
|
||||
# Fix: Video Codec Compatibility + Timeline Pan/Scroll + Minimap
|
||||
|
||||
**Date:** 2026-09-21 12:56
|
||||
**Task:** Fix video playback (black screen despite play state), add timeline panning when zoomed, add minimap overview.
|
||||
|
||||
## Root Causes & Fixes
|
||||
|
||||
### 1. Video playback: WKWebView doesn't support VP9/WebM
|
||||
The downloaded file was VP9+Opus in WebM container (yt-dlp format 399+251). Apple's WKWebView on macOS does NOT support VP9 codec. The `<video>` element accepted `.play()` but had no decodable frames.
|
||||
|
||||
**Fix:** Added `-f "bv*[vcodec^=avc1]+ba[acodec^=mp4a]/bv*[ext=mp4]+ba[ext=m4a]/b[ext=mp4]/b"` to the yt-dlp download command in `download_manager.rs`. This selects H.264 (avc1) video + AAC (mp4a) audio, producing an MP4 file playable by WKWebView. Deleted the cached WebM file.
|
||||
|
||||
### 2. Timeline: no scroll/pan when zoomed
|
||||
Mouse wheel only triggered zoom. No way to scroll horizontally when zoomed in.
|
||||
|
||||
**Fix in `Timeline.svelte`:**
|
||||
- **Horizontal scroll → pan:** `deltaX` from trackpad/shift+wheel now calls `panBy()` to shift the visible window
|
||||
- **Vertical scroll → zoom** (unchanged)
|
||||
|
||||
**New `panBy()` in `interactions.ts`:** Computes new `visibleStart`/`visibleEnd` from pixel delta, clamped to `[0, duration]`.
|
||||
|
||||
### 3. Minimap overview bar
|
||||
Added a minimap canvas that appears when zoomed in:
|
||||
- Shows the full waveform at a glance
|
||||
- Highlights the current viewport with a blue border
|
||||
- Dims regions outside the viewport
|
||||
- Click to center viewport, drag to pan
|
||||
- Playhead indicator
|
||||
|
||||
### 4. Other fixes applied this session (cumulative)
|
||||
- **stdout/stderr swap:** yt-dlp sends status messages to stdout, not stderr
|
||||
- **`[Merger]` path parsing:** Correctly captures final merged file path
|
||||
- **"Already downloaded" parsing:** Handles `has already been downloaded` message
|
||||
- **Cache check:** `checkCachedDownload(title)` checks temp dir before invoking yt-dlp
|
||||
- **Parallel processing:** Waveform/thumbnails/keyframes results applied individually as they complete
|
||||
- **Asset protocol scope:** Added `/private/var/**` and `/var/**` for macOS symlink paths
|
||||
|
||||
## Files Changed
|
||||
- `src-tauri/src/services/download_manager.rs` — H.264 format selection
|
||||
- `src/lib/components/Timeline.svelte` — Minimap + horizontal pan support
|
||||
- `src/lib/timeline/interactions.ts` — Added `panBy()` function
|
||||
- `src-tauri/src/commands/video.rs` — Cache check, diagnostic logging
|
||||
- `src-tauri/src/commands/media_analysis.rs` — Diagnostic logging
|
||||
- `src-tauri/src/lib.rs` — Registered `check_cached_download` command
|
||||
- `src-tauri/tauri.conf.json` — Expanded asset protocol scope
|
||||
- `src/lib/stores/videoSession.svelte.ts` — Cache check before download, parallel results
|
||||
- `src/lib/bindings/video.ts` — `checkCachedDownload` binding
|
||||
- `src/lib/components/VideoPlayer.svelte` — Local file playback with progress UI
|
||||
|
||||
## Lessons Learned
|
||||
1. **WKWebView codec support is limited.** VP9/AV1/WebM don't work. Must use H.264+AAC in MP4.
|
||||
2. **macOS `/var` is a symlink to `/private/var`.** Asset protocol scope must cover both paths.
|
||||
3. **`Promise.allSettled` blocks all results until the slowest completes.** Use individual `.then()` chains when you want progressive rendering.
|
||||
4. **yt-dlp sends status to stdout, warnings to stderr.** Not the other way around.
|
||||
@@ -0,0 +1,60 @@
|
||||
# Hybrid Download Model + Auto-Defocus URL Input
|
||||
|
||||
## Task Description
|
||||
Two user-requested improvements:
|
||||
1. **Hybrid download model**: Download a low-resolution (≤360p H.264+AAC) version for immediate preview/scrubbing/waveform/thumbnail generation, while downloading the best available quality in the background for export. Clips are cut from the best-quality file.
|
||||
2. **Auto-defocus URL input**: After pasting a URL and triggering submit, blur the input field so keyboard shortcuts (I, O, etc.) don't accidentally modify the URL.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### Session State Refactor (`src/lib/stores/videoSession.svelte.ts`)
|
||||
- Replaced single `localFilePath`/`downloadProgress`/`downloadStatus` with dual-track state:
|
||||
- `previewFilePath`/`previewProgress`/`previewStatus` — for the 360p preview
|
||||
- `exportFilePath`/`exportProgress`/`exportStatus` — for the best-quality export
|
||||
- Preview format: `bv*[vcodec^=avc1][height<=360]+ba[acodec^=mp4a]/b[ext=mp4][height<=360]/worst[ext=mp4]/worst`
|
||||
- Export format: `bv*+ba/b` (best available, any codec since ffmpeg handles export)
|
||||
- `beginDownload()` now:
|
||||
1. Checks cache for preview → downloads if needed → triggers post-processing (waveform, thumbnails, keyframes)
|
||||
2. Checks cache for export → downloads in background (fire-and-forget)
|
||||
- Files stored in separate subdirectories: `video-clipper/preview/` and `video-clipper/export/`
|
||||
|
||||
### Rust Backend (`src-tauri/src/commands/video.rs`)
|
||||
- `start_download` now accepts `format_spec` and `variant` parameters
|
||||
- `check_cached_download` now accepts a `variant` parameter to check the correct subdirectory
|
||||
- Helper `variant_dir()` builds `$TEMP/video-clipper/{variant}/` paths
|
||||
- Diagnostic logging includes variant name for easier debugging
|
||||
|
||||
### Download Manager (`src-tauri/src/services/download_manager.rs`)
|
||||
- `start_download` now accepts `format_spec: &str` parameter instead of hardcoding format selection
|
||||
- Format string passed through from the frontend call
|
||||
|
||||
### Frontend Bindings (`src/lib/bindings/video.ts`)
|
||||
- `startDownload` now takes `formatSpec` and `variant` parameters
|
||||
- `checkCachedDownload` now takes a `variant` parameter
|
||||
|
||||
### VideoPlayer (`src/lib/components/VideoPlayer.svelte`)
|
||||
- Uses `previewFilePath`/`previewStatus` instead of old single-track fields
|
||||
- Download progress shows "Downloading preview…" label
|
||||
|
||||
### ExportDialog (`src/lib/components/ExportDialog.svelte`)
|
||||
- Uses `exportFilePath`/`exportStatus`/`exportProgress` for export readiness
|
||||
- Shows "Downloading best quality (X%)…" while export download is in progress
|
||||
|
||||
### StatusBar (`src/lib/components/StatusBar.svelte`)
|
||||
- Shows dual progress: preview download → export download → ready to export
|
||||
- Contextual status messages for each download phase
|
||||
|
||||
### UrlInput (`src/lib/components/UrlInput.svelte`)
|
||||
- Added `bind:this={inputEl}` reference to input element
|
||||
- `blurInput()` called at the start of `handleSubmit()` — defocuses immediately on paste/enter
|
||||
- Keyboard shortcuts (I, O, etc.) now work immediately after submitting a URL
|
||||
|
||||
## Lessons Learned
|
||||
- Storing preview and export files in separate subdirectories (`preview/`, `export/`) keeps cache management clean and avoids filename collisions between quality variants.
|
||||
- Fire-and-forget pattern for background export download (`.catch()` at call site) keeps the preview flow responsive without blocking on the best-quality download.
|
||||
- The `variant` parameter threading from frontend → Rust command → download manager keeps the API clean and extensible for future quality tiers.
|
||||
|
||||
## Follow-Up Items
|
||||
- Consider showing export download progress in the timeline/player area as a subtle indicator
|
||||
- The export format `bv*+ba/b` may download VP9/WebM — this is fine for ffmpeg export but won't play in WKWebView. The preview file handles playback.
|
||||
- Old flat `video-clipper/` cache was cleared; users with existing caches in the old location won't get cache hits (harmless — just re-downloads)
|
||||
@@ -0,0 +1,72 @@
|
||||
# Fix Audio Playback + Dynamic Waveform Resolution
|
||||
|
||||
## Task Description
|
||||
1. **Fix missing audio on preview playback** — Video plays but with no sound in WKWebView.
|
||||
2. **Dynamic waveform resolution** — Waveform should increase in detail as the user zooms in.
|
||||
|
||||
## Root Cause Analysis (Audio)
|
||||
|
||||
The preview file is correctly muxed (H.264 640x360 + AAC 128kbps, moov atom at offset 24 — fast-start). WKWebView config analysis:
|
||||
- wry 0.55.1 defaults `autoplay: true` → sets `mediaTypesRequiringUserActionForPlayback = None`
|
||||
- tauri-runtime-wry uses `WebViewBuilder::new_with_web_context()` which inherits this default
|
||||
- Tauri v2.11.6 doesn't expose or override `autoplay`
|
||||
- So WKWebView SHOULD be configured for audio playback
|
||||
|
||||
Despite correct configuration, WKWebView's Tauri asset protocol (`https://asset.localhost/...`) appears to silently drop audio tracks when streaming local files — possibly due to range request handling or MIME type issues in the custom protocol handler.
|
||||
|
||||
**Fix**: Load the video file as a `blob:` URL instead of using the asset protocol. This bypasses the asset protocol entirely and uses WKWebView's native blob URL handling, which reliably plays both audio and video tracks.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### Audio Fix: Blob URL Video Loading (`src/lib/components/VideoPlayer.svelte`)
|
||||
- Added a `$effect` that fetches the preview file via `convertFileSrc` URL, converts it to a Blob, then creates a `blob:` URL
|
||||
- Video element now uses the blob URL instead of the asset protocol URL
|
||||
- Added `playsinline` attribute to the video element
|
||||
- Added explicit `volume = 1` and `muted = false` on `loadeddata`
|
||||
- Added `.play()` Promise error handling (catches and logs rejections)
|
||||
- Shows "Preparing video…" state while blob is loading
|
||||
- Falls back to asset protocol URL if blob creation fails
|
||||
- Properly revokes old blob URLs on session change
|
||||
|
||||
### Dynamic Waveform: Backend (`src-tauri/src/services/waveform_generator.rs`)
|
||||
- Added `extract_waveform_range(file_path, start_time, end_time, peak_count)` function
|
||||
- Uses `ffmpeg -ss {start} -t {duration}` with 44100 Hz sample rate for high-resolution extraction
|
||||
- Computes peaks for just the specified time range
|
||||
|
||||
### Dynamic Waveform: Command (`src-tauri/src/commands/media_analysis.rs`)
|
||||
- Added `extract_waveform_range` Tauri command
|
||||
- Registered in `lib.rs`
|
||||
|
||||
### Dynamic Waveform: Frontend Binding (`src/lib/bindings/mediaAnalysis.ts`)
|
||||
- Added `extractWaveformRange(filePath, startTime, endTime, peakCount)` binding
|
||||
|
||||
### Dynamic Waveform: Session State (`src/lib/stores/videoSession.svelte.ts`)
|
||||
- Added `waveformDetailPeaks`, `waveformDetailStart`, `waveformDetailEnd` to session state
|
||||
- Increased initial waveform extraction from 8,000 to 50,000 peaks (good for most zoom levels)
|
||||
- Detail fields cleared on session reset
|
||||
|
||||
### Dynamic Waveform: Renderer (`src/lib/timeline/waveformRenderer.ts`)
|
||||
- Refactored `drawWaveform` to accept a `WaveformData` object with both overview and detail peaks
|
||||
- Renderer automatically uses detail peaks when they cover the visible viewport
|
||||
- Falls back to overview peaks when detail is not available
|
||||
|
||||
### Dynamic Waveform: Timeline Component (`src/lib/components/Timeline.svelte`)
|
||||
- Added auto-fetch `$effect` that monitors zoom level and viewport
|
||||
- When overview peaks per pixel drops below 2, triggers a 300ms-debounced detail extraction
|
||||
- Detail extraction pads visible range by 50% on each side to avoid re-fetching on small pans
|
||||
- Requests 4 peaks per pixel for crisp detail
|
||||
- Passes `WaveformData` to both main timeline and minimap renderers
|
||||
|
||||
### Renderer Types Updated (`src/lib/timeline/renderer.ts`)
|
||||
- `drawTimeline` now accepts `WaveformData` instead of `number[]`
|
||||
|
||||
## Lessons Learned
|
||||
- WKWebView's custom protocol handlers (like Tauri's `https://asset.localhost/`) can have subtle audio issues even when video plays fine. Blob URLs are a reliable workaround.
|
||||
- For waveform LOD (level of detail), a two-tier approach (high-count overview + on-demand range extraction) provides the best UX: fast initial display with detail on demand.
|
||||
- Debouncing the detail waveform fetch prevents spamming ffmpeg during rapid zoom/pan.
|
||||
- Padding the extraction range by 50% on each side significantly reduces re-fetch frequency during small viewport adjustments.
|
||||
|
||||
## Follow-Up Items
|
||||
- Investigate if Tauri's asset protocol can be configured to properly serve audio (may be a wry bug)
|
||||
- Consider pre-computing multiple LOD levels for the waveform instead of on-demand extraction
|
||||
- The blob approach loads the entire preview file into memory (~97MB for a 54-min video at 360p) — acceptable for desktop but may need optimization for very long videos
|
||||
@@ -0,0 +1,116 @@
|
||||
# Fix Multi-Clip Workflow + Caption Support
|
||||
|
||||
**Date:** 2026-09-21 14:09
|
||||
**Task:** Fix multi-clip creation and add time-synced caption support
|
||||
|
||||
## Task Description
|
||||
|
||||
Two main issues addressed:
|
||||
1. **Multi-clip bug**: Pressing I/O keys always edited the selected clip instead of creating new clips when the playhead was outside the clip region.
|
||||
2. **Caption support**: Full pipeline for downloading, displaying, and exporting time-synced captions (embedded > external, English preferred, auto-generated deprioritized).
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Fix Multi-Clip Workflow
|
||||
|
||||
**`src/lib/stores/clips.svelte.ts`**:
|
||||
- Added `isTimeInsideClip(time, clipId)` helper with 0.5s tolerance
|
||||
- `markInPoint(time)`: Now checks if playhead is inside the selected clip's range. If outside, deselects and sets `pendingInPoint` (starts new clip). If inside, edits the existing clip.
|
||||
- `markOutPoint(time)`: Same logic — edits selected clip if inside, creates new clip from pending in-point if outside.
|
||||
|
||||
**`src/lib/components/Timeline.svelte`**:
|
||||
- Added `selectClip(null)` call when clicking empty timeline space (no `hitTestClip` result), so the next I/O presses create a new clip.
|
||||
|
||||
### 2. Caption Metadata Parsing
|
||||
|
||||
**`src-tauri/src/services/video_resolver.rs`**:
|
||||
- Extended `YtDlpJson` struct with `subtitles` and `automatic_captions` fields (both `HashMap<String, Vec<SubtitleFormat>>`)
|
||||
- Added `detect_captions()` function: checks for English subtitles (en, en-US, en-GB variants), prioritizes manual over auto-generated
|
||||
- Added `SubtitleFormat` struct for deserialization
|
||||
- Added 4 new unit tests for caption detection
|
||||
|
||||
**`src-tauri/src/models.rs`**:
|
||||
- Added `has_captions: bool` and `captions_are_auto: bool` to `VideoMetadata`
|
||||
|
||||
**`src/lib/bindings/video.ts`**:
|
||||
- Added `hasCaptions` and `captionsAreAuto` to `VideoMetadata` interface
|
||||
|
||||
### 3. Caption Download
|
||||
|
||||
**`src-tauri/src/services/subtitle_downloader.rs`** (new file):
|
||||
- Downloads English VTT subtitles via `yt-dlp --write-subs` (manual) or `--write-auto-subs` (auto-generated)
|
||||
- Uses `--sub-langs en.*,en --sub-format vtt --convert-subs vtt`
|
||||
- Searches output directory for `.en.vtt` files, falls back to any `.vtt`
|
||||
|
||||
**`src-tauri/src/commands/video.rs`**:
|
||||
- Added `download_subtitles` Tauri command
|
||||
|
||||
**`src/lib/bindings/video.ts`**:
|
||||
- Added `downloadSubtitles()` binding
|
||||
|
||||
### 4. Embedded Subtitle Check
|
||||
|
||||
**`src-tauri/src/commands/media_analysis.rs`**:
|
||||
- Added `check_embedded_subtitles` command that uses `ffprobe -show_streams -select_streams s` to detect subtitle streams, then extracts as VTT via `ffmpeg -map 0:s:0 -f webvtt`
|
||||
|
||||
**`src/lib/bindings/mediaAnalysis.ts`**:
|
||||
- Added `checkEmbeddedSubtitles()` binding
|
||||
|
||||
### 5. Caption Display in Player
|
||||
|
||||
**`src/lib/stores/videoSession.svelte.ts`**:
|
||||
- Added `hasCaptions`, `captionsAreAuto`, `captionFilePath` to session state
|
||||
- Added `loadCaptions()` function that checks embedded subs first, then downloads external
|
||||
- Integrated into `triggerPostDownloadProcessing()`
|
||||
|
||||
**`src/lib/components/VideoPlayer.svelte`**:
|
||||
- Loads VTT file as blob URL via `captionBlobUrl`
|
||||
- Renders `<track>` element with proper `srclang`, `label` (with "(auto)" suffix), and `default` attribute
|
||||
- Added CC toggle button (bottom-right overlay) with active/inactive styling
|
||||
- Syncs `textTracks[].mode` with `captionsEnabled` state
|
||||
|
||||
### 6. Caption Export
|
||||
|
||||
**`src-tauri/src/models.rs`**:
|
||||
- Added `include_captions: bool` and `caption_file_path: Option<String>` to `ExportConfig`
|
||||
|
||||
**`src-tauri/src/services/clip_exporter.rs`**:
|
||||
- Added `build_ffmpeg_args_with_subs()`: lossless → mux as `mov_text`/`srt`; precise → burn-in via `subtitles=` filter
|
||||
- Added `export_single_clip_with_subs()` with fallback to no-subs on failure
|
||||
- Added `export_merged_with_subs()`
|
||||
- Added 2 new unit tests for subtitle-aware arg building
|
||||
|
||||
**`src-tauri/src/commands/export.rs`**:
|
||||
- Updated to conditionally use `_with_subs` variants based on `include_captions`
|
||||
|
||||
**`src/lib/bindings/export.ts`**:
|
||||
- Added `includeCaptions` and `captionFilePath` to `ExportConfig` interface
|
||||
|
||||
**`src/lib/components/ExportDialog.svelte`**:
|
||||
- Added "Captions" section with checkbox toggle (only shown when captions are available)
|
||||
- Shows "(auto-generated)" label when applicable
|
||||
- Shows mux vs. burn-in note based on cut mode
|
||||
|
||||
### Infrastructure
|
||||
|
||||
**`src-tauri/src/services/mod.rs`**: Registered `subtitle_downloader` module
|
||||
**`src-tauri/src/lib.rs`**: Registered `download_subtitles` and `check_embedded_subtitles` commands
|
||||
|
||||
## Build & Test Results
|
||||
|
||||
- **Rust**: 34 tests pass (including 4 new caption tests + 2 new export tests)
|
||||
- **Frontend**: Builds cleanly (0 errors)
|
||||
- **Rust build**: Compiles with only pre-existing unused variant warnings
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- `yt-dlp --dump-json` includes `subtitles` and `automatic_captions` maps — both map language codes to arrays of `{ext, url}` objects
|
||||
- WKWebView `<track>` elements work with blob URLs but need `textTracks[].mode` managed manually to sync with a toggle
|
||||
- For ffmpeg subtitle burn-in, paths with colons need escaping (`\:`) in the `subtitles=` filter
|
||||
- Subtitle muxing in lossless mode requires separate `-i` for the subtitle file and explicit stream mapping (`-map 0:v -map 0:a -map 1:s`)
|
||||
|
||||
## Follow-up Items
|
||||
|
||||
- The `SubtitleFormat.ext` and `SubtitleFormat.url` fields generate "never read" warnings — they exist for serde deserialization but could be suppressed with `#[allow(dead_code)]`
|
||||
- Caption time offset accuracy: when using `-ss` before `-i`, subtitle timestamps may need adjustment for precise alignment in edge cases
|
||||
- Sprite sheet optimization (noted from prior session) is still pending
|
||||
@@ -0,0 +1,69 @@
|
||||
# Performance, Captions, Clip UX, and Resizable Panels
|
||||
|
||||
**Date:** 2026-09-21 14:39
|
||||
**Task:** Fix timeline performance, caption positioning, clip creation workflow, resizable panels, and deselect UX
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Performance Optimization (3 files)
|
||||
|
||||
**`src/lib/components/VideoPlayer.svelte`**:
|
||||
- Throttled `handleTimeUpdate` to ~15fps (66ms interval) using `performance.now()` check. Previously every `ontimeupdate` event triggered a reactive cascade through `session.currentTime`.
|
||||
|
||||
**`src/lib/components/Timeline.svelte`**:
|
||||
- **In-flight guard**: Added `detailFetchInFlight` flag to prevent overlapping waveform range extraction subprocess calls. Only one ffmpeg call runs at a time.
|
||||
- **Increased debounce**: Waveform detail fetch debounce increased from 300ms to 500ms.
|
||||
- **Playhead-only redraws**: Split the monolithic redraw `$effect` into two:
|
||||
- Structural changes (clips, zoom, waveform data, thumbnails) trigger a full `drawMainCanvas()` which saves an `ImageData` snapshot.
|
||||
- `session.currentTime` changes trigger a lightweight `drawPlayheadOnly()` that restores the snapshot and draws only the playhead — avoiding the expensive thumbnail/waveform/clip rendering on every time update.
|
||||
- **Thumbnail load callback**: Registered a callback via `setThumbnailLoadCallback` so thumbnail image loads trigger a targeted redraw instead of relying on the next reactive cycle.
|
||||
|
||||
**`src/lib/timeline/thumbnailRenderer.ts`**:
|
||||
- Added `pendingLoads` Set to prevent creating duplicate `Image` objects for the same path across rapid redraws.
|
||||
- Added `setThumbnailLoadCallback` API so the Timeline can request a redraw when thumbnails finish loading asynchronously.
|
||||
- Image `onerror` handler cleans up the pending state.
|
||||
|
||||
### 2. Caption Positioning (1 file)
|
||||
|
||||
**`src/lib/components/VideoPlayer.svelte`**:
|
||||
- Added `:global(video::cue)` CSS to center captions at the bottom of the video with a semi-transparent black background, white text, and 16px font size. Used `:global()` to bypass Svelte scoping since `::cue` is a browser-level pseudo-element.
|
||||
- Added `position: relative` to the `<video>` element to scope cue rendering.
|
||||
|
||||
### 3. Pending In-Point Preservation (1 file)
|
||||
|
||||
**`src/lib/stores/clips.svelte.ts`**:
|
||||
- Modified `selectClip()` to only clear `pendingInPoint` when `id !== null` (selecting a specific clip). When `id === null` (deselecting via empty timeline click), the pending in-point is preserved. This fixes the workflow: press I → click timeline to seek → press O to complete the clip.
|
||||
|
||||
### 4. Resizable Panels (3 files)
|
||||
|
||||
**`src/App.svelte`**:
|
||||
- Wrapped `<Timeline />` and `<ClipList />` in a new `.timeline-clip-area` container with flex column layout.
|
||||
- Added a 5px `.resize-handle` divider between them with `cursor: ns-resize` and accent color on hover/active.
|
||||
- Implemented `handleResizeStart/Move/End` with `mousedown`/`mousemove`/`mouseup` on `<svelte:window>` to drag-resize. Timeline height is clamped between 80px and (total - 60px).
|
||||
- `timelineHeight` stored in local `$state` (defaults to 180px).
|
||||
|
||||
**`src/lib/components/Timeline.svelte`**:
|
||||
- Changed `.timeline-container` from `height: 160px` to `flex: 1; min-height: 80px` so it fills the space given by the parent.
|
||||
- `.timeline-wrapper` set to `height: 100%`.
|
||||
|
||||
**`src/lib/components/ClipList.svelte`**:
|
||||
- Changed from `max-height: 200px` to `height: 100%; min-height: 40px` with `box-sizing: border-box`.
|
||||
|
||||
### 5. Deselect UX (1 file)
|
||||
|
||||
**`src/lib/components/ClipList.svelte`**:
|
||||
- Added `handleContainerClick` on the `.clip-list` div that calls `selectClip(null)` when clicking empty space (not a child clip row).
|
||||
- Replaced the single "Export All" button with an `.action-buttons` row containing a conditional "Deselect" button (shown only when a clip is selected) alongside "Export All".
|
||||
- Changed `role="button"` to `role="option"` on clip rows and `role="listbox"` on the container for better semantics.
|
||||
|
||||
## Build Status
|
||||
|
||||
- Frontend: builds cleanly (0 errors)
|
||||
- Rust: compiles with only pre-existing unused variant warnings
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- Svelte's scoped CSS doesn't apply to `::cue` pseudo-elements since they're browser-level shadow constructs — `:global()` wrapper is needed.
|
||||
- Canvas `getImageData`/`putImageData` is an effective technique for "snapshot and overlay" rendering patterns, avoiding expensive full redraws when only one layer (like a playhead) changes.
|
||||
- `ontimeupdate` fires at varying rates across browsers/engines; throttling to a target frame rate prevents reactive cascades from dominating the UI thread.
|
||||
- The `pendingInPoint` clearing bug was a subtle consequence of `selectClip(null)` being used for two different intents (deselecting a clip vs. resetting all clip state).
|
||||
@@ -0,0 +1,51 @@
|
||||
# Fix Fundamental Performance & Audio Regression
|
||||
|
||||
**Date:** 2026-09-21 14:51
|
||||
**Task:** Fix root causes of sluggish UI and broken audio on long videos
|
||||
|
||||
## Root Causes Identified
|
||||
|
||||
### 1. Blob URL for Video (CRITICAL)
|
||||
The `VideoPlayer` was fetching the ENTIRE preview MP4 file into JavaScript memory via `fetch()` → `blob()` → `URL.createObjectURL()`. For a 54-minute 360p video (~200-500MB), this:
|
||||
- Blocks the UI while the entire file is fetched into memory
|
||||
- Doubles memory usage temporarily (file buffer + blob)
|
||||
- Makes seeks slow because the browser must parse the blob
|
||||
- Was the root cause of sluggish playback, delayed seek response, and overall UI lag
|
||||
|
||||
The blob URL was originally a workaround for WKWebView audio issues with Tauri's asset protocol, but the tradeoff is unacceptable for any video longer than a few minutes.
|
||||
|
||||
### 2. getImageData/putImageData Overhead
|
||||
The "playhead-only redraw" optimization was using `getImageData()` to snapshot the canvas and `putImageData()` to restore it. On a Retina display, this copies 20+ MB of pixel data per frame — worse than just redrawing the canvas from scratch.
|
||||
|
||||
### 3. Parallel ffmpeg Subprocesses
|
||||
`triggerPostDownloadProcessing` fired all three analysis tasks (waveform, keyframes, thumbnails) simultaneously. Each spawns an ffmpeg subprocess that decodes the full file. Three concurrent ffmpeg processes on a 54-minute video saturate the CPU.
|
||||
|
||||
### 4. Excessive Waveform Peaks
|
||||
Hardcoded at 50,000 peaks regardless of duration. For a 54-minute video, this requires decoding the entire audio track at high resolution. Most of these peaks are never visible at the default zoom level.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### `src/lib/components/VideoPlayer.svelte`
|
||||
- **Removed blob URL entirely** — now uses `convertFileSrc()` directly (Tauri asset protocol). This streams from disk with zero memory overhead, enabling instant seek on any video length.
|
||||
- Changed `preload="auto"` to `preload="metadata"` — only loads metadata and first frames, not the entire file.
|
||||
- Removed `loadingBlob` state and associated "Preparing video…" UI state.
|
||||
- Audio should work with asset protocol for H.264+AAC in MP4 (the preview format). If it doesn't, we'll investigate the specific WKWebView config rather than working around it with blob URLs.
|
||||
|
||||
### `src/lib/components/Timeline.svelte`
|
||||
- **Removed getImageData/putImageData** snapshot mechanism entirely. All redraws go through a single `drawMainCanvas()` call, coalesced by `requestAnimationFrame`.
|
||||
- Removed `lastDrawnState`, `lastFullDrawTime`, `pendingPlayheadDraw`, `drawPlayheadOnly()`.
|
||||
- Merged the separate "structural" and "playhead" `$effect`s back into one — `requestAnimationFrame` already coalesces multiple calls per frame.
|
||||
- Changed `pendingDraw` and `pendingMinimapDraw` from `$state` to plain `let` — they don't need reactivity and were causing unnecessary tracking overhead.
|
||||
- Changed `detailFetchTimer` and `detailFetchInFlight` from `$state` to plain `let` — same reason.
|
||||
- Reduced waveform detail peak count from `pw * 4` to `pw * 2` (2 peaks per pixel is sufficient).
|
||||
|
||||
### `src/lib/stores/videoSession.svelte.ts`
|
||||
- **Sequential processing**: Changed `triggerPostDownloadProcessing` from parallel fire-and-forget to `async` sequential execution: waveform → keyframes → thumbnails → captions. Only one ffmpeg process runs at a time.
|
||||
- **Scaled waveform peaks**: Changed from hardcoded 50,000 to `Math.min(10000, Math.max(2000, duration * 10))`. A 54-minute video gets 10,000 peaks (vs. 50,000 before). A 2-minute video gets 2,000.
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- **Blob URLs are not a scalable workaround** — they work for small files but are catastrophic for anything over a few minutes. Always use the asset protocol (streaming from disk) for video playback.
|
||||
- **getImageData/putImageData is expensive on Retina displays** — the data transfer cost (20+ MB per snapshot) exceeds the cost of just redrawing the canvas from primitives.
|
||||
- **Sequential ffmpeg is faster than parallel** for analysis tasks on the same file — the file is read from disk for each, and concurrent processes compete for I/O and CPU. Sequential processing also leaves CPU available for the UI thread.
|
||||
- **$state variables for non-reactive bookkeeping** (requestAnimationFrame IDs, timers) add unnecessary tracking overhead.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Pre-Process Waveform Tiers + Processing Modal
|
||||
|
||||
**Date:** 2026-09-21 15:10
|
||||
**Task:** Replace on-demand waveform range extraction with one-time multi-tier pre-processing pipeline, add a processing modal, eliminate the reactive `$effect` loop bug.
|
||||
|
||||
## Problem
|
||||
|
||||
The `$effect` in `Timeline.svelte` that fetched waveform ranges on demand created an infinite reactive loop: writing to `session.waveformDetailPeaks` re-triggered the same `$effect`, which re-evaluated `needsDetail`, which fired again. The terminal showed `waveform-range` calls repeating 40+ times for the same range, killing performance.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Rust Backend — Waveform Tiers (`waveform-tiers-rust`)
|
||||
|
||||
- **`src-tauri/src/models.rs`**: Added `WaveformTiers` struct with `tier0` (~2K peaks), `tier1` (~10K peaks), `tier2` (all raw peaks, 50K–200K depending on duration).
|
||||
- **`src-tauri/src/services/waveform_generator.rs`**: Rewrote entirely. New `extract_waveform_tiers(file_path, duration)` function runs ffmpeg once at a sample rate yielding ~100 peaks/sec (capped 50K–200K), then downsamples into 3 tiers in Rust. Added `downsample()` helper (max-of-chunk). Removed `extract_waveform_range` function. Added 6 unit tests for downsample and compute_peaks.
|
||||
- **`src-tauri/src/commands/media_analysis.rs`**: Replaced `extract_waveform` and `extract_waveform_range` commands with single `extract_waveform_tiers` command.
|
||||
- **`src-tauri/src/lib.rs`**: Updated handler registration — `extract_waveform_tiers` replaces `extract_waveform` + `extract_waveform_range`.
|
||||
|
||||
### 2. Frontend — Tier-Based Rendering (`waveform-tiers-frontend`)
|
||||
|
||||
- **`src/lib/bindings/mediaAnalysis.ts`**: Added `WaveformTiers` interface and `extractWaveformTiers` binding. Removed `extractWaveformRange` binding.
|
||||
- **`src/lib/timeline/waveformRenderer.ts`**: Rewrote. `WaveformData` is now `{ tiers: WaveformTiers | null }`. `pickTier()` selects the lowest-resolution tier with ≥1 peak per pixel for the visible range — pure arithmetic, zero IPC.
|
||||
- **`src/lib/timeline/renderer.ts`**: Updated `drawTimeline` signature and waveform condition to use `waveform.tiers`.
|
||||
- **`src/lib/stores/videoSession.svelte.ts`**: Replaced `waveformPeaks`, `waveformDetailPeaks`, `waveformDetailStart`, `waveformDetailEnd` with `waveformTiers: WaveformTiers | null`. Added `processingStep` state. Processing pipeline now calls `extractWaveformTiers` instead of `extractWaveform`.
|
||||
- **`src/lib/components/Timeline.svelte`**: **Deleted the entire `$effect` for waveform range fetching** — the infinite loop bug is eliminated. Removed `detailFetchTimer`, `detailFetchInFlight`, `needsDetail`, `extractWaveformRange` import. `waveformData` is now simply `{ tiers: session.waveformTiers }`.
|
||||
|
||||
### 3. Processing Modal (`processing-modal`)
|
||||
|
||||
- **`src/lib/components/ProcessingModal.svelte`** (new): Modal overlay showing processing steps with icons (✓ done, ⏳ current, ○ pending). Steps: downloading → waveform → keyframes → thumbnails → captions → done. Includes download progress bar.
|
||||
- **`src/App.svelte`**: Imported and renders `<ProcessingModal />` when `session.processingStep` is neither `'idle'` nor `'done'`.
|
||||
- **`src/lib/stores/videoSession.svelte.ts`**: `processingStep` is set at each stage of `triggerPostDownloadProcessing` and `beginDownload`.
|
||||
|
||||
## Verification
|
||||
|
||||
- **Rust**: `cargo build` succeeds (3 pre-existing warnings). `cargo test` passes all 37 tests (including 6 new waveform generator tests).
|
||||
- **Frontend**: `svelte-check` reports only pre-existing vite.config.ts errors and a11y warnings — no new issues.
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- The on-demand `$effect` → async IPC → write reactive state → re-trigger `$effect` pattern is fundamentally broken in Svelte 5. The correct approach is pre-computation: do all CPU work up front, store the results, and let the renderer pick the right data with pure arithmetic.
|
||||
- For waveform at typical zoom levels (even 100x on an 800px-wide canvas), 50K peaks gives ~6 peaks/pixel, which is plenty. No need for on-demand range extraction.
|
||||
- Downsampling by max-of-chunk preserves peak amplitude fidelity while drastically reducing array size.
|
||||
|
||||
## Follow-up Items
|
||||
|
||||
- Audio playback may need verification after the session store changes (the `previewFilePath` → `convertFileSrc` pipeline is unchanged, but should be tested).
|
||||
- Consider persisting waveform tiers to disk cache for instant reload on re-open.
|
||||
@@ -0,0 +1,82 @@
|
||||
# Progress, Resize, Captions, Audio, Preview Upgrade
|
||||
|
||||
**Date:** 2026-09-21 19:31
|
||||
**Task:** Implement 6 items from the "Progress Resize Captions Audio" plan
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Waveform Speed Fix + Granular Progress (waveform-speed-fix)
|
||||
|
||||
**Files:** `src-tauri/src/services/waveform_generator.rs`, `src-tauri/src/commands/media_analysis.rs`, `src/lib/bindings/mediaAnalysis.ts`, `src/lib/stores/videoSession.svelte.ts`, `src/lib/components/ProcessingModal.svelte`
|
||||
|
||||
- **Critical bug fix:** `aresample=N` was setting the output sample rate to N Hz (e.g., 200,000 Hz) instead of producing N total samples. For a 54-min video, this generated ~650M samples. Fixed by computing actual rate as `raw_count / duration` and using `-ar` flag instead. ~3000x speedup for long videos.
|
||||
- Changed `waveform_generator.rs` to use child process spawning with piped stdout/stderr, parse `out_time_us=` from ffmpeg progress output, and report progress via a callback.
|
||||
- Added `Channel<f64>` progress parameter to `extract_waveform_tiers` Tauri command.
|
||||
- Added `processingProgress` to session store; reset in `setMetadata`, `clearMediaFields`.
|
||||
- Updated `ProcessingModal` to show progress bars for every step (not just download).
|
||||
|
||||
### 2. Fix Resize Handles (fix-resize)
|
||||
|
||||
**Files:** `src/lib/components/VideoPlayer.svelte`, `src/App.svelte`
|
||||
|
||||
- Changed `VideoPlayer` min-height from 200px to 0.
|
||||
- Removed `max-height: 60vh` from timeline-clip-area.
|
||||
- Lowered min-height constraints: timeline-clip-area 140→80, timeline-pane 80→40, cliplist-pane 40→20.
|
||||
- Updated resize handler min/max constraints to match.
|
||||
|
||||
### 3. Custom Caption Rendering (fix-captions)
|
||||
|
||||
**Files:** `src/lib/utils/vttParser.ts` (new), `src/lib/stores/preferences.svelte.ts`, `src/lib/components/VideoPlayer.svelte`
|
||||
|
||||
- Created `vttParser.ts` with `parseVtt()` and `getActiveCues()` for custom VTT parsing.
|
||||
- Replaced `<track>` element with custom HTML overlay positioned absolutely in the video player.
|
||||
- Captions now rendered bottom-center (or top, configurable) with single background layer (no double-background).
|
||||
- Added `CaptionSettings` interface to preferences store with font size, text color, background opacity, text outline, and position.
|
||||
|
||||
### 4. Caption Settings Panel (caption-settings-ui)
|
||||
|
||||
**Files:** `src/lib/components/CaptionSettingsPanel.svelte` (new), `src/lib/components/VideoPlayer.svelte`
|
||||
|
||||
- Created settings sub-menu with range sliders, color picker, checkbox, radio buttons.
|
||||
- Accessible via ⚙ button next to CC toggle.
|
||||
- Settings persist via preferences store.
|
||||
- "Reset to defaults" button included.
|
||||
|
||||
### 5. Local HTTP Media Server for Audio (fix-audio)
|
||||
|
||||
**Files:** `src-tauri/src/services/media_server.rs` (new), `src-tauri/src/services/mod.rs`, `src-tauri/src/lib.rs`, `src-tauri/Cargo.toml`, `src/lib/bindings/video.ts`, `src/lib/components/VideoPlayer.svelte`
|
||||
|
||||
- Added `axum` + `tower-http` (with fs feature) dependencies.
|
||||
- Created media server that starts on a random available port at app launch, serving files from filesystem root with full range-request support via `ServeDir`.
|
||||
- Exposed port to frontend via `get_media_server_port` Tauri command.
|
||||
- VideoPlayer now constructs video URLs as `http://127.0.0.1:{port}/path/to/file.mp4` instead of using `convertFileSrc()`.
|
||||
- This fixes WKWebView's audio issues with the asset protocol.
|
||||
|
||||
### 6. Preview Upgrade Toast (preview-upgrade)
|
||||
|
||||
**Files:** `src/lib/stores/videoSession.svelte.ts`, `src/lib/components/VideoPlayer.svelte`, `src/lib/components/StatusBar.svelte`
|
||||
|
||||
- Added `activeVideoPath` (initially set to preview, upgradeable to export), `showUpgradeToast`, `upgradePreview()`, `dismissUpgradeToast()`.
|
||||
- Toast appears in top-right of video player when export download completes and processing is done.
|
||||
- "Reload Preview (HQ)" button added to StatusBar.
|
||||
- Video source now driven by `activeVideoPath` instead of `previewFilePath`.
|
||||
|
||||
## Verification
|
||||
|
||||
- `cargo build`: ✅ (only pre-existing warnings)
|
||||
- `cargo test`: ✅ 37 tests pass
|
||||
- `npm run check`: ✅ (only pre-existing vite.config.ts errors)
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- `aresample=N` in ffmpeg sets the **output sample rate in Hz**, not a total count. Must compute rate = count/duration.
|
||||
- Tauri's asset protocol (`https://asset.localhost/...`) has issues with audio track streaming in WKWebView — a local HTTP server with range request support via axum/tower-http is a clean fix.
|
||||
- TypeScript's control flow analysis can over-narrow `$state` values across async callback boundaries, requiring explicit cast (`as ProcessingStep`) to work around.
|
||||
- Custom caption rendering (parsing VTT + HTML overlay) gives full styling control vs. browser `::cue` pseudo-element limitations.
|
||||
|
||||
## Follow-up Items
|
||||
|
||||
- Smoke test the full flow end-to-end with a real YouTube video.
|
||||
- Verify audio playback works correctly via the local HTTP server.
|
||||
- Test caption settings persistence across app restarts.
|
||||
- Consider adding progress reporting to keyframe extraction and thumbnail generation (currently only waveform has it).
|
||||
@@ -0,0 +1,50 @@
|
||||
# Fix HQ Upgrade, Audio Playback, and Waveform
|
||||
|
||||
**Date:** 2026-09-21 20:01
|
||||
**Task:** Fix three bugs reported after the previous batch of changes
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Waveform Truncation Fix
|
||||
|
||||
**File:** `src-tauri/src/services/waveform_generator.rs`
|
||||
|
||||
**Root cause:** `compute_peaks` used `peaks.resize(target_count, 0.0)` to pad the peaks array to exactly `target_count` entries. But when `chunk_size = ceil(samples / target_count)` rounds up (e.g., 201439 samples / 200000 target = chunk_size 2), the actual number of chunks produced is far less than `target_count` (100720 vs 200000). The resize pads ~50% of the array with zeros, compressing the real waveform into the first half and leaving the rest flat.
|
||||
|
||||
**Fix:** Removed the `peaks.resize(target_count, 0.0)` call. The downstream rendering code already handles arrays of any length — the peak count from chunking is the correct count.
|
||||
|
||||
### 2. Export Format Fix (HQ Upgrade Black Screen)
|
||||
|
||||
**File:** `src/lib/stores/videoSession.svelte.ts`
|
||||
|
||||
**Root cause:** `EXPORT_FORMAT = 'bv*+ba/b'` allowed yt-dlp to pick the best format regardless of codec — which resulted in AV1+Opus in WebM. WKWebView on macOS cannot decode AV1 or Opus, so upgrading to the HQ preview produced a black screen with no playback.
|
||||
|
||||
**Fix:** Changed `EXPORT_FORMAT` to `'bv*[vcodec^=avc1]+ba[acodec^=mp4a]/b[ext=mp4]/best[ext=mp4]'` to force H.264+AAC in MP4, matching what WKWebView can play. Deleted the cached `.webm` export file so the next run re-downloads in the correct format.
|
||||
|
||||
### 3. Volume Controls + Audio Diagnostics
|
||||
|
||||
**File:** `src/lib/components/VideoPlayer.svelte`
|
||||
|
||||
- Added `volume` and `isMuted` state variables with reactive sync to the video element
|
||||
- Added mute toggle button (speaker emoji) and volume slider in a new bottom controls bar
|
||||
- Added diagnostic `console.log` in `handleLoadedData` to print muted, volume, readyState, audioTracks length, and dimensions — this will help trace the audio issue if it persists
|
||||
- Reorganized the bottom overlay: volume controls on the left, CC/settings on the right
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- When `samples.chunks(chunk_size)` is used with a `chunk_size > 1`, the number of resulting chunks is `ceil(samples / chunk_size)`, which is strictly less than `target_count` when `chunk_size = ceil(samples / target_count)`. Never pad with zeros — it corrupts the waveform mapping to time.
|
||||
- yt-dlp's `bv*+ba/b` format spec can pick any codec. On macOS, always constrain to H.264+AAC for WKWebView compatibility.
|
||||
- The audio issue may be related to mixed content (HTTPS frontend loading HTTP media) — the diagnostic logging added here will help isolate whether the issue is mute state, volume, or something lower-level in WebKit.
|
||||
|
||||
## Verification
|
||||
|
||||
- `cargo build`: OK (only pre-existing warnings)
|
||||
- `cargo test`: 37 tests pass
|
||||
- `npm run check`: OK (only pre-existing vite.config.ts errors)
|
||||
|
||||
## Follow-up
|
||||
|
||||
- Run `npx tauri dev` and verify waveform renders across the full timeline
|
||||
- Verify HQ upgrade no longer produces a black screen
|
||||
- Check console output from `handleLoadedData` diagnostic to trace audio issue
|
||||
- If audio still doesn't play, investigate mixed content blocking in WKWebView
|
||||
20
chat-summaries/2026-09-21_20-25-fix-captions-cors-summary.md
Normal file
20
chat-summaries/2026-09-21_20-25-fix-captions-cors-summary.md
Normal file
@@ -0,0 +1,20 @@
|
||||
# Fix Captions Missing Due to CORS
|
||||
|
||||
## Task
|
||||
Captions/subtitles were not appearing at all - the CC button and settings button were invisible because `parsedCues` was always empty.
|
||||
|
||||
## Root Cause
|
||||
The local media server (axum on `http://127.0.0.1:<random_port>`) had **no CORS headers**. The frontend WebView runs at a different origin (`http://localhost:1420` in dev, `tauri://localhost` in prod). While `<video>` elements can load cross-origin media without CORS (they use "no-cors" mode), the `fetch()` call used to load the VTT caption file was blocked by the browser's same-origin policy.
|
||||
|
||||
The `fetch()` silently failed (caught by `.catch()`), setting `parsedCues = []`, which meant the `{#if parsedCues.length > 0}` conditional in `VideoPlayer.svelte` never rendered the CC controls.
|
||||
|
||||
## Changes Made
|
||||
1. **`src-tauri/Cargo.toml`**: Added `"cors"` feature to `tower-http` dependency.
|
||||
2. **`src-tauri/src/services/media_server.rs`**: Added `CorsLayer::permissive()` to the axum router, enabling cross-origin `fetch()` from the WebView.
|
||||
|
||||
## Lessons Learned
|
||||
- `<video src="...">` does NOT require CORS for basic playback - the browser loads media in "no-cors" mode. But `fetch()` to the same URL WILL be blocked without CORS headers. This discrepancy is why video/audio played fine but caption loading silently failed.
|
||||
- When a `fetch()` fails silently in a `.catch()` handler, there's no visible error in the UI — only in the browser DevTools console. Adding CORS from the start would have prevented this class of issues.
|
||||
|
||||
## Follow-up
|
||||
None - this was a targeted 2-file fix.
|
||||
@@ -0,0 +1,47 @@
|
||||
# Fix Caption Export and Add Burn-In Option
|
||||
|
||||
## Task
|
||||
Exported clips were not including captions despite the "include captions" option being checked. Additionally, the user requested a "burn-in" option to render subtitles directly into the video frames.
|
||||
|
||||
## Root Causes
|
||||
|
||||
1. **Precise mode burn-in (timestamp offset)**: `-ss` was placed before `-i` (input seeking), shifting output PTS to 0. The `subtitles` filter reads the original VTT with absolute timestamps (e.g., cues at 300s), but the output video starts at 0s — no cues matched.
|
||||
|
||||
2. **Lossless mode muxing (VTT formatting)**: YouTube auto-generated VTT has karaoke-style `<c>` tags, inline timestamps (`<00:00:00.599>`), and positioning metadata (`align:start position:0%`) that confuse ffmpeg's VTT parser when converting to `mov_text`.
|
||||
|
||||
3. **Silent fallback**: When the ffmpeg subtitle export failed, the code silently fell back to exporting without subtitles, making the failure invisible to the user.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### `src-tauri/src/services/clip_exporter.rs` (major rewrite)
|
||||
- **Added `sanitize_vtt_for_ffmpeg()`**: Strips karaoke `<c>` tags, inline timestamps, and positioning metadata from VTT files. Writes a cleaned temp file.
|
||||
- **Added `strip_vtt_tags()`**: Helper to remove all HTML-like tags from VTT cue text lines.
|
||||
- **Replaced `build_ffmpeg_args_with_subs()`** with two separate functions:
|
||||
- `build_ffmpeg_args_mux_subs()`: Muxes subtitles as a track. Uses output seeking (`-ss`/`-to` after `-i`) for correct timestamp alignment. Works with both lossless and precise cut modes.
|
||||
- `build_ffmpeg_args_burnin_subs()`: Burns subtitles into video using the `subtitles` filter. Uses output seeking so the filter reads correct VTT timestamps. Always re-encodes (overrides to H.264/AAC if lossless mode selected).
|
||||
- **Updated `export_single_clip_with_subs()`**: Now takes `burn_in: bool`, sanitizes VTT before use, dispatches to mux or burn-in builder. **Removed the silent fallback** — errors now surface to the user.
|
||||
- **Updated `export_merged_with_subs()`**: Takes `burn_in: bool`, passes through to per-segment export.
|
||||
- **Updated tests**: Replaced old tests for removed function, added tests for both new builders, added a VTT tag stripping test. All 40 tests pass.
|
||||
|
||||
### `src-tauri/src/models.rs`
|
||||
- Added `pub burn_in_captions: bool` to `ExportConfig`.
|
||||
|
||||
### `src-tauri/src/commands/export.rs`
|
||||
- Threads `config.burn_in_captions` through to `export_single_clip_with_subs` and `export_merged_with_subs`.
|
||||
|
||||
### `src/lib/bindings/export.ts`
|
||||
- Added `burnInCaptions: boolean` to the TypeScript `ExportConfig` interface.
|
||||
|
||||
### `src/lib/components/ExportDialog.svelte`
|
||||
- Added `burnInCaptions` state.
|
||||
- Added "Burn into video" sub-checkbox under the "Include captions" checkbox.
|
||||
- Contextual notes: "Captions will be muxed as a subtitle track" vs "Captions will be burned into the video" vs "Burn-in requires re-encoding (precise mode will be used)".
|
||||
- Threads `burnInCaptions` through to the export config.
|
||||
|
||||
## Lessons Learned
|
||||
- ffmpeg's `-ss` before `-i` (input seeking) adjusts output PTS to start at 0, but the `subtitles` filter reads timestamps from the original VTT file — they must match. Output seeking (`-ss` after `-i`) preserves original timestamps.
|
||||
- YouTube auto-generated VTT files contain karaoke formatting that ffmpeg's VTT-to-mov_text converter can't handle cleanly. Sanitizing before use is essential.
|
||||
- Silent fallbacks that hide errors waste debugging time. Surface errors to the user.
|
||||
|
||||
## Follow-up
|
||||
- The CORS fix from the prior session (media_server.rs adding `CorsLayer::permissive()`) is also needed for captions to display in the preview player.
|
||||
@@ -0,0 +1,41 @@
|
||||
# Fix Caption Export Bugs (Round 2)
|
||||
|
||||
## Task
|
||||
Three bugs in the caption export implementation needed fixing after the initial burn-in/mux feature was added.
|
||||
|
||||
## Bug 1: Burn-in filter syntax error
|
||||
**Symptom**: ffmpeg error `No option name near '/var/.../sanitized.vtt'`
|
||||
**Root cause**: `format!("subtitles='{}'", path)` passed literal single quotes to ffmpeg via `Command::args()`. Since there's no shell to strip them, ffmpeg tried to open `'sanitized.vtt'` (with quotes in the filename).
|
||||
**Fix**: Changed to `format!("subtitles={}", path)` — no quotes needed when passing args directly.
|
||||
|
||||
## Bug 2: Wrong output duration (1:43 instead of 32s)
|
||||
**Symptom**: Exported clip was 103 seconds instead of 32 seconds.
|
||||
**Root cause**: Both subtitle functions used `-to end_time` with output seeking (`-ss`/`-to` after `-i`). After `-ss` discards initial frames and the muxer resets PTS to 0, `-to 103` means "output 103 seconds" rather than "stop at input timestamp 103".
|
||||
**Fix**: Replaced `-to end_time` with `-t duration` (`clip.end_time - clip.start_time`) in both `build_ffmpeg_args_mux_subs` and `build_ffmpeg_args_burnin_subs`. `-t` is unambiguous.
|
||||
|
||||
## Bug 3: Muxed subtitles invisible in VLC
|
||||
**Symptom**: Subtitle track exists in the file but selecting it in VLC shows nothing.
|
||||
**Root cause**: With output seeking, video PTS was reset to 0 by the muxer, but subtitle cues retained their original absolute timestamps (e.g., 71s-103s). The player sees subtitle cues at 71s but the video is only 32s long.
|
||||
**Fix**: Two-part strategy change:
|
||||
1. Added `trim_and_sanitize_vtt()` function that extracts cues within the clip's time range and shifts all timestamps to start from 0.
|
||||
2. Switched `build_ffmpeg_args_mux_subs` to use **input seeking** (`-ss`/`-t` before `-i`) for speed. The pre-trimmed VTT timestamps already start from 0, matching the video output.
|
||||
3. Burn-in mode (`build_ffmpeg_args_burnin_subs`) retains **output seeking** since the `subtitles` filter needs to see the original absolute timestamps from the VTT.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### `src-tauri/src/services/clip_exporter.rs`
|
||||
- Removed single quotes from `subtitles=` filter value
|
||||
- Changed `-to end` to `-t duration` in both subtitle builder functions
|
||||
- Added `trim_and_sanitize_vtt()` with VTT timestamp parsing/shifting
|
||||
- Added helper functions: `parse_vtt_timestamp_line()`, `parse_vtt_ts()`, `format_vtt_timestamp()`
|
||||
- Switched `build_ffmpeg_args_mux_subs` to input seeking with pre-trimmed VTT
|
||||
- Updated `export_single_clip_with_subs` to use trimmed VTT for mux, sanitized VTT for burn-in
|
||||
- Updated all related tests (43 tests pass)
|
||||
|
||||
## Lessons Learned
|
||||
- `Command::args()` in Rust bypasses the shell entirely — shell-style quoting (single quotes) becomes literal characters in the argument. Never quote paths for ffmpeg filters when using `Command::args()`.
|
||||
- ffmpeg's `-to` is ambiguous with output seeking — it may refer to post-PTS-reset time rather than input time. `-t duration` is always safe.
|
||||
- When using input seeking on video (`-ss` before `-i`), subtitle PTS from a separate input must match the reset PTS (~0). Pre-trimming the VTT with shifted timestamps is the reliable approach.
|
||||
|
||||
## Follow-up
|
||||
None — all three user-reported bugs are addressed.
|
||||
@@ -0,0 +1,74 @@
|
||||
# Fix Burn-In Export and Re-Introduce Karaoke Preview
|
||||
|
||||
**Date:** 2026-09-21 22:06
|
||||
**Task:** Implement graceful burn-in fallback and re-introduce karaoke word-by-word highlighting in preview captions.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. Detect `subtitles` Filter Availability (Burn-In Guard)
|
||||
|
||||
**Problem:** The user's ffmpeg build (Homebrew 9.0.1_1) lacks `libass`, so the `subtitles` filter is unavailable. Attempting burn-in export caused ffmpeg to crash with a filter parse error.
|
||||
|
||||
**Fix:**
|
||||
- **`src-tauri/src/services/dependency_manager.rs`** — Added `check_subtitles_filter()` that runs `ffmpeg -filters` and checks for `subtitles V->V` in the output.
|
||||
- **`src-tauri/src/commands/dependencies.rs`** — Added `check_subtitles_filter_available()` Tauri command.
|
||||
- **`src-tauri/src/lib.rs`** — Registered the new command.
|
||||
- **`src/lib/bindings/dependencies.ts`** — Added `checkSubtitlesFilterAvailable()` frontend binding.
|
||||
- **`src/lib/components/ExportDialog.svelte`** — On mount, checks filter availability via the new command. If unavailable, the "Burn into video" checkbox is disabled with a help message telling the user how to install libass.
|
||||
|
||||
### 2. Parse Word-Level Timings from VTT
|
||||
|
||||
**Problem:** YouTube auto-generated VTT files contain word-level timestamps in `<timestamp><c> word</c>` patterns. The old parser stripped ALL tags, losing this timing data.
|
||||
|
||||
**Fix:**
|
||||
- **`src/lib/utils/vttParser.ts`** — Complete rewrite:
|
||||
- Added `WordSegment` interface (`{ text: string; startTime: number }`).
|
||||
- Added optional `words?: WordSegment[]` to `VttCue`.
|
||||
- Added `parseWordTimings()` that extracts `<HH:MM:SS.mmm><c> word</c>` patterns.
|
||||
- Added `getActiveWords()` helper that returns each word marked as `spoken` or upcoming based on `currentTime`.
|
||||
- The plain `text` field is preserved (stripped) for backward compat.
|
||||
|
||||
### 3. Render Karaoke Word-by-Word Highlights
|
||||
|
||||
**Fix:**
|
||||
- **`src/lib/components/VideoPlayer.svelte`** — Updated caption rendering:
|
||||
- When `wordHighlight` is enabled and the cue has word segments, each word is rendered as a separate `<span>`.
|
||||
- Spoken words display at full opacity; upcoming words display at 40% opacity.
|
||||
- CSS transition smooths the opacity change.
|
||||
- When disabled or no word data available, falls back to plain text rendering.
|
||||
|
||||
### 4. Add Word Highlight Toggle
|
||||
|
||||
**Fix:**
|
||||
- **`src/lib/stores/preferences.svelte.ts`** — Added `wordHighlight: boolean` to `CaptionSettings` interface (default: `true`).
|
||||
- **`src/lib/components/CaptionSettingsPanel.svelte`** — Added "Word-by-word highlight" checkbox toggle.
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src-tauri/src/services/dependency_manager.rs` | Added `check_subtitles_filter()` |
|
||||
| `src-tauri/src/commands/dependencies.rs` | Added `check_subtitles_filter_available` command |
|
||||
| `src-tauri/src/lib.rs` | Registered new command |
|
||||
| `src/lib/bindings/dependencies.ts` | Added `checkSubtitlesFilterAvailable()` binding |
|
||||
| `src/lib/components/ExportDialog.svelte` | Conditional burn-in disable + help text |
|
||||
| `src/lib/utils/vttParser.ts` | Word-level timing parser + `getActiveWords()` |
|
||||
| `src/lib/components/VideoPlayer.svelte` | Karaoke word rendering |
|
||||
| `src/lib/stores/preferences.svelte.ts` | Added `wordHighlight` to `CaptionSettings` |
|
||||
| `src/lib/components/CaptionSettingsPanel.svelte` | Added word highlight toggle |
|
||||
|
||||
## Build Verification
|
||||
|
||||
- `cargo check` — passes (only pre-existing warnings)
|
||||
- `svelte-check` — passes (only pre-existing `vite.config.ts` type errors)
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- YouTube VTT word-timing format: leading text has no `<c>` wrapper, only subsequent words do. The first word inherits the cue's start time.
|
||||
- The `subtitles` filter requires libass at ffmpeg compile time — it's not enough to have ffmpeg installed. Homebrew's default build may or may not include it depending on the formula version.
|
||||
|
||||
## Follow-Up Items
|
||||
|
||||
- Manual smoke test: Load a video with captions, verify karaoke highlighting works during playback.
|
||||
- Test export with captions enabled (muxed track) to confirm no regressions.
|
||||
- If user installs libass and reinstalls ffmpeg, burn-in should automatically become available next time ExportDialog opens.
|
||||
@@ -0,0 +1,64 @@
|
||||
# Styled Burn-In and Granular Export Progress
|
||||
|
||||
**Date:** 2026-09-21 23:00
|
||||
**Task:** Apply user's caption style settings to burned-in subtitles via ffmpeg's ASS force_style, and add granular per-clip progress reporting during export.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. CaptionStyle Model (Rust + TS)
|
||||
|
||||
- **`src-tauri/src/models.rs`** — Added `CaptionStyle` struct with `font_size`, `text_color`, `background_opacity`, `text_outline`, `position`. Added `caption_style: Option<CaptionStyle>` to `ExportConfig`.
|
||||
- **`src/lib/bindings/export.ts`** — Added matching `CaptionStyle` interface and `captionStyle: CaptionStyle | null` to `ExportConfig`. Added `clipProgress` event variant to the discriminated union. Added optional `onClipProgress` callback parameter to `exportClips()`.
|
||||
|
||||
### 2. ASS force_style in clip_exporter.rs
|
||||
|
||||
- **`src-tauri/src/services/clip_exporter.rs`**:
|
||||
- `hex_to_ass_color()` — Converts CSS hex `#RRGGBB` to ASS `&H00BBGGRR` format.
|
||||
- `opacity_to_ass_back_colour()` — Converts 0-1 opacity to ASS alpha-prefixed BackColour.
|
||||
- `caption_style_to_force_style()` — Builds the full ASS force_style string from CaptionStyle (Fontsize, PrimaryColour, BackColour, Outline, Shadow, Alignment, MarginV, BorderStyle).
|
||||
- `build_ffmpeg_args_burnin_subs()` now accepts `Option<&CaptionStyle>` and appends `:force_style='...'` to the subtitles filter when provided.
|
||||
- `parse_ffmpeg_time()` — Extracts `time=HH:MM:SS.mm` from ffmpeg stderr progress output.
|
||||
- `run_ffmpeg_with_progress()` — Runs ffmpeg with piped stderr, parses progress lines, and invokes a callback with 0.0-1.0 percent values.
|
||||
|
||||
### 3. Progress Callbacks Throughout Export Pipeline
|
||||
|
||||
All four export functions (`export_single_clip`, `export_single_clip_with_subs`, `export_merged`, `export_merged_with_subs`) now accept progress callbacks. They use `run_ffmpeg_with_progress()` instead of `.output()` for the encode step, streaming real-time progress.
|
||||
|
||||
### 4. export.rs Command Handler
|
||||
|
||||
- **`src-tauri/src/commands/export.rs`** — Added `ClipProgress` event variant with `current`, `total`, `label`, `percent`. Wired progress callbacks for both Individual and Merged export paths. Passes `caption_style` through to the exporter.
|
||||
|
||||
### 5. ExportDialog Frontend
|
||||
|
||||
- **`src/lib/components/ExportDialog.svelte`**:
|
||||
- Populates `captionStyle` in config from `preferences.captionSettings` when burn-in is enabled.
|
||||
- Shows a progress bar with percentage during each clip export.
|
||||
- Displays note that karaoke word highlighting is preview-only for burn-in.
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src-tauri/src/models.rs` | Added `CaptionStyle` struct, added field to `ExportConfig` |
|
||||
| `src-tauri/src/services/clip_exporter.rs` | ASS helpers, force_style, ffmpeg progress streaming, progress callbacks |
|
||||
| `src-tauri/src/commands/export.rs` | `ClipProgress` event, caption_style passthrough, progress wiring |
|
||||
| `src/lib/bindings/export.ts` | `CaptionStyle` interface, `clipProgress` event, `onClipProgress` callback |
|
||||
| `src/lib/components/ExportDialog.svelte` | captionStyle in config, progress bar UI, karaoke note |
|
||||
|
||||
## Build Verification
|
||||
|
||||
- `cargo check` — passes (only pre-existing warnings)
|
||||
- `cargo test` — all 43 tests pass
|
||||
- `svelte-check` — passes (only pre-existing `vite.config.ts` errors)
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- ASS color format uses BGR byte order with alpha prefix (`&HAA_BB_GG_RR`), where alpha `00` = opaque and `FF` = transparent (inverted from CSS).
|
||||
- `BorderStyle=4` in ASS gives an opaque background box behind text (like CSS background), vs `BorderStyle=1` which uses outline+shadow only.
|
||||
- ffmpeg progress is on stderr, not stdout. The `time=` field is the key metric for computing encode progress percentage.
|
||||
- For merged exports, progress is reported per-segment during encoding, plus a concat step at the end.
|
||||
|
||||
## Follow-Up Items
|
||||
|
||||
- Smoke test burn-in export to verify styled subtitles render correctly.
|
||||
- Karaoke word-by-word animation in burn-in would require converting VTT word timestamps to ASS `\k` override tags — complex and fragile, noted as preview-only limitation.
|
||||
@@ -0,0 +1,60 @@
|
||||
# Karaoke Burn-In for Exported Subtitles
|
||||
|
||||
**Date:** 2026-09-21 23:11
|
||||
**Task:** Convert word-timed YouTube VTT captions to ASS format with karaoke `\kf` tags so burned-in subtitles have the same word-by-word highlight effect as the preview.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### 1. VTT-to-ASS Converter with Karaoke Tags
|
||||
|
||||
**`src-tauri/src/services/clip_exporter.rs`** — Added several new functions:
|
||||
|
||||
- `vtt_to_ass_with_karaoke()` — Main converter. Reads raw VTT, generates a complete ASS file with:
|
||||
- `[Script Info]` header (1920x1080 play resolution)
|
||||
- `[V4+ Styles]` section with the user's CaptionStyle baked in (PrimaryColour, SecondaryColour at 60% alpha for the "dim/upcoming" look, BackColour, Outline, Shadow, Alignment, BorderStyle=4)
|
||||
- `[Events]` section where each cue is a Dialogue line with `\kf<centiseconds>` tags for word-level karaoke fill animation
|
||||
- Graceful fallback: cues without word-level timestamps render as plain text (no `\kf`)
|
||||
|
||||
- `parse_vtt_word_timings()` — Rust equivalent of the TypeScript `parseWordTimings()`. Extracts `(word, start_time)` pairs from YouTube's `<timestamp><c> word</c>` format.
|
||||
|
||||
- `format_ass_timestamp()` — Formats seconds as ASS timestamp `H:MM:SS.cc` (centiseconds).
|
||||
|
||||
- `hex_to_ass_color_with_alpha()` — Like `hex_to_ass_color()` but with a specific alpha byte.
|
||||
|
||||
- `find_timestamp_tag_pos()` / `find_next_timestamp()` — Helpers for parsing VTT timestamp tags without regex.
|
||||
|
||||
### 2. Updated Burn-In Export Path
|
||||
|
||||
In `export_single_clip_with_subs()`, the burn-in path now uses:
|
||||
```
|
||||
VTT → vtt_to_ass_with_karaoke() → .ass file → subtitles= filter
|
||||
```
|
||||
Instead of the previous:
|
||||
```
|
||||
VTT → sanitize_vtt_for_ffmpeg() → plain VTT → subtitles= filter
|
||||
```
|
||||
|
||||
Since the style is now embedded in the ASS file header, `caption_style` is passed as `None` to `build_ffmpeg_args_burnin_subs()` (no `force_style` needed — avoids potential conflicts between the ASS header and force_style overrides).
|
||||
|
||||
### 3. Updated UI Note
|
||||
|
||||
**`src/lib/components/ExportDialog.svelte`** — Changed the burn-in note from "word-by-word highlighting is preview-only" to "Captions will be burned into the video with your style settings".
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src-tauri/src/services/clip_exporter.rs` | Added VTT-to-ASS converter, word timing parser, ASS timestamp formatter, updated burn-in path |
|
||||
| `src/lib/components/ExportDialog.svelte` | Updated burn-in note text |
|
||||
|
||||
## Build Verification
|
||||
|
||||
- `cargo check` — passes (only pre-existing warnings)
|
||||
- `cargo test` — all 43 tests pass
|
||||
|
||||
## Technical Notes
|
||||
|
||||
- ASS karaoke `\kf<N>` means "fill this word over N centiseconds", transitioning from SecondaryColour to PrimaryColour. This matches the preview behavior where upcoming words are dimmed and spoken words are bright.
|
||||
- The ASS SecondaryColour is set to the same text color with 0x99 alpha (~60% transparent), matching the preview's `opacity: 0.4` for upcoming words.
|
||||
- No regex crate was needed — the VTT word timing pattern is simple enough to parse with manual string operations.
|
||||
- The `subtitles` ffmpeg filter handles ASS files natively (same libass backend), so no filter change was needed.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Fix Karaoke Burn-In Sync and Export Progress
|
||||
|
||||
**Date:** 2026-09-21 23:22
|
||||
**Task:** Fix two bugs: karaoke ASS burn-in out of sync / skipping lines, and export progress stuck at 0% then jumping to 100%.
|
||||
|
||||
## Changes Made
|
||||
|
||||
### Bug 1: Karaoke ASS Burn-In — Two Root Causes Fixed
|
||||
|
||||
**A. YouTube VTT two-line overlay pattern.**
|
||||
|
||||
YouTube auto-generated VTT cues have a specific structure:
|
||||
- Many cues have TWO text lines: line 1 is static context (previous cue text, no `<c>` tags), line 2 is the karaoke line (with `<timestamp><c> word</c>` tags)
|
||||
- Zero-duration transition cues (e.g., `00:00:02.629 --> 00:00:02.639`) are just visual transition frames
|
||||
|
||||
The old code joined all text lines and tried to parse word timings from the combined mess. Fix:
|
||||
- Skip cues where `(end - start) < 0.05s` — these are transition frames
|
||||
- For multi-line cues, only process the line that contains `<c>` tags for karaoke
|
||||
- Ignore the static context line (previous cue's text repeated)
|
||||
- Cues with no `<c>` tags emit as plain text (graceful fallback)
|
||||
|
||||
**B. Wrong ffmpeg filter.**
|
||||
|
||||
Per ffmpeg-micro.com: "The `subtitles` filter routes your file through libavformat's subtitle converter and applies `force_style` overrides, which can flatten karaoke timing. The `ass` filter hands the file straight to libass."
|
||||
|
||||
Fix: Changed `build_ffmpeg_args_burnin_subs()` to use `-vf ass=<path>` instead of `-vf subtitles=<path>`. Since style is already baked into the ASS `[V4+ Styles]` header, no `force_style` is needed. The `_caption_style` parameter is now ignored (prefixed with `_`).
|
||||
|
||||
### Bug 2: Export Progress — Root Cause Fixed
|
||||
|
||||
ffmpeg writes its progress output to stderr using `\r` (carriage return) to overwrite the same line, NOT `\n` (newline). `BufReader::lines()` splits on `\n` only, so all progress was buffered as one giant "line" that only yielded when ffmpeg exited.
|
||||
|
||||
Fix: Use ffmpeg's `-progress pipe:1 -nostats` flag:
|
||||
- Adds `-progress pipe:1 -nostats` to ffmpeg args (before the output file)
|
||||
- Reads stdout (not stderr) — `-progress` outputs proper `\n`-delimited `key=value` pairs
|
||||
- Parses `out_time_ms=<microseconds>` lines and computes `percent = out_time_ms / (duration * 1_000_000)`
|
||||
- stderr is still piped but drained on a separate thread for error reporting on failure
|
||||
|
||||
Example `-progress pipe:1` output (proper newlines):
|
||||
```
|
||||
frame=150
|
||||
fps=45.2
|
||||
out_time_ms=6250000
|
||||
progress=continue
|
||||
```
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src-tauri/src/services/clip_exporter.rs` | Fixed VTT parsing, switched to `ass=` filter, rewrote progress to use `-progress pipe:1` |
|
||||
|
||||
## Build Verification
|
||||
|
||||
- `cargo test` — all 43 tests pass
|
||||
- `cargo check` — passes (only pre-existing warnings)
|
||||
- `svelte-check` — passes (only pre-existing `vite.config.ts` errors)
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
- YouTube VTT is NOT simple cue-per-line. It uses a two-line overlay pattern where line 1 is a static context repeat and line 2 has the karaoke data. Zero-duration cues are transition frames.
|
||||
- The `ass` filter and `subtitles` filter in ffmpeg both use libass, but `subtitles` applies format conversion and `force_style` that can destroy `\kf` karaoke timing. Always use `ass` for karaoke.
|
||||
- ffmpeg's stderr progress uses `\r` not `\n`. For Rust's `BufReader::lines()` to work, use `-progress pipe:1 -nostats` which outputs proper `\n`-delimited key=value pairs to stdout.
|
||||
@@ -0,0 +1,27 @@
|
||||
# Fix Burned-In Caption Styling
|
||||
|
||||
**Date:** 2026-09-21 23:37
|
||||
**Task:** Fix ASS styling in burn-in export to be readable and match the preview better.
|
||||
|
||||
## Changes Made
|
||||
|
||||
**`src-tauri/src/services/clip_exporter.rs`** — Updated `vtt_to_ass_with_karaoke()` ASS style generation:
|
||||
|
||||
| ASS Field | Before (broken) | After (fixed) |
|
||||
|-----------|-----------------|---------------|
|
||||
| OutlineColour | `primary_colour` (white) | `&H00000000` (black) |
|
||||
| SecondaryColour | Same color with alpha `&H99FFFFFF` | Distinct grey `&H99AAAAAA` |
|
||||
| BorderStyle | `4` (opaque box) always | `1` (outline+shadow) when textOutline=true; `3` (box only) when false |
|
||||
| Outline | 2 always (even with box) | 2 with BorderStyle=1; 0 with BorderStyle=3 |
|
||||
|
||||
### Why each change matters:
|
||||
|
||||
- **OutlineColour=black**: The preview uses `text-shadow: 1px 1px 2px rgba(0,0,0,0.9)` for readability. Setting OutlineColour to white made the outline invisible against white text and washed out the karaoke fill effect.
|
||||
|
||||
- **SecondaryColour=grey**: For `\kf` karaoke, libass sweeps from SecondaryColour to PrimaryColour. With both being variants of white (just different alpha), the sweep was barely visible. Using a distinct grey (`&H99AAAAAA`) makes upcoming words clearly dimmed grey, sweeping to bright white when spoken.
|
||||
|
||||
- **BorderStyle=1 vs 3**: `BorderStyle=4` draws an opaque background box AND renders the outline, creating a "double background" artifact. `BorderStyle=1` (outline+shadow, no box) matches the preview's outlined text look. `BorderStyle=3` (opaque box, no outline) is used when the user disables text outline, giving a clean box background.
|
||||
|
||||
## Build Verification
|
||||
|
||||
- `cargo test` — all 43 tests pass
|
||||
@@ -0,0 +1,65 @@
|
||||
# Rolling Two-Line Karaoke Subtitle Display
|
||||
|
||||
**Date:** 2026-09-22 00:21
|
||||
**Task:** Rewrite burn-in subtitle export to use smooth scrolling two-line display matching YouTube's native caption rendering
|
||||
|
||||
## Problem
|
||||
|
||||
The previous burn-in subtitle approach generated one ASS Dialogue entry per VTT cue, combining context + karaoke text with a hard `\N` break. When one cue ended and the next began, the display hard-cut — the previous karaoke line instantly became static context at the top, and a new karaoke line appeared at the bottom. This felt "disjointed" compared to YouTube's smooth scrolling behavior.
|
||||
|
||||
## Prior Art
|
||||
|
||||
- [Sofronio/YouTubeVTT2ASS](https://github.com/Sofronio/YouTubeVTT2ASS) — C# tool that solves this exact problem using `\move` ASS tags to create a rolling/scrolling effect. Their v0.0.3 specifically notes "Smooth rolling effect, no intervals between lines."
|
||||
- The key technique: for each spoken line, generate 3 ASS Dialogue entries with `\move` animations (Active → Context → Disappear).
|
||||
|
||||
## Changes Made
|
||||
|
||||
**File:** `src-tauri/src/services/clip_exporter.rs`
|
||||
|
||||
### New structs and functions:
|
||||
|
||||
1. **`SpokenLine` struct** — Represents a single spoken line extracted from VTT, with fields for plain text, raw VTT text, timing, karaoke flag, and non-speech flag.
|
||||
|
||||
2. **`extract_spoken_lines()`** — Pre-pass parser that converts YouTube VTT into a flat sequence of spoken lines. Skips zero-duration transition cues. For two-line cues, extracts only the active (karaoke) line — context is reconstructed from the previous SpokenLine during generation.
|
||||
|
||||
3. **`build_karaoke_text()`** — Extracts word timings and builds `\k` karaoke tags for a single line.
|
||||
|
||||
4. **`RollingLayout` struct** — Position parameters: `cx` (960), `y_bottom` (1040), `line_height` (font_size * 1.3), `scroll_ms` (350).
|
||||
|
||||
### Rewritten `vtt_to_ass_with_karaoke()`:
|
||||
|
||||
For each spoken line, generates up to 3 ASS Dialogue entries:
|
||||
|
||||
- **Phase 1 (Active/Karaoke):** Line appears at bottom position, scrolls up one slot via `\move(cx, y_bottom, cx, y_bottom-h, 0, 350)`. Has `\k` karaoke tags. Lasts from this line's start to the next line's start.
|
||||
|
||||
- **Phase 2 (Context/Static):** Same text (plain), scrolls up another slot. Lasts from next line's start to the line after that.
|
||||
|
||||
- **Phase 3 (Disappear):** Scrolls off-screen with `\clip` mask to cleanly cut off. Lasts 500ms.
|
||||
|
||||
### Edge cases handled:
|
||||
|
||||
- **Long gaps (>2s):** Context phase ends early; line disappears instead of lingering through silence/music.
|
||||
- **Non-speech cues (`[Music]`):** Single static Dialogue with `\pos` instead of rolling.
|
||||
- **First line:** No context above it — just starts normally.
|
||||
- **Last line:** Phase 1 uses `line.end_time`; Phase 2/3 use a hold + disappear.
|
||||
- **Lines without karaoke data:** Rendered as plain text with the same rolling behavior.
|
||||
|
||||
### Style additions:
|
||||
|
||||
- Added `ScaledBorderAndShadow: yes` to `[Script Info]` for proper scaling.
|
||||
- Each Dialogue line uses `\an2` override for explicit bottom-center positioning.
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
1. **YouTube's VTT two-line pattern** is inherently a "teleprompter" — the bottom line fills with karaoke words, then scrolls up to become context while a new line appears below. Reproducing this requires `\move` animations, not just `\N` line breaks.
|
||||
|
||||
2. **ASS `\move` with `\an2`** — The alignment setting determines the anchor point for positioning. `\an2` (bottom-center) means Y coordinates refer to the bottom edge of the text, and X=960 centers it horizontally on a 1920-wide canvas.
|
||||
|
||||
3. **Three-phase lifecycle per line** is the key insight from YouTubeVTT2ASS. Phase 3 with `\clip` is important to cleanly mask the text as it scrolls off instead of having it abruptly disappear.
|
||||
|
||||
4. **Gap detection** is essential — without it, stale context lines would linger through long silences or `[Music]` sections.
|
||||
|
||||
## Build/Test Status
|
||||
|
||||
- `cargo build`: Success (6 pre-existing warnings, 0 errors)
|
||||
- `cargo test`: 43 tests passed, 0 failed
|
||||
@@ -0,0 +1,62 @@
|
||||
# Paged Teleprompter Burn-In Subtitles
|
||||
|
||||
**Date:** 2026-09-22
|
||||
**Task:** Replace rolling per-line ASS subtitle generation with a paged teleprompter model for burn-in captions.
|
||||
|
||||
## Problem
|
||||
|
||||
The previous rolling subtitle implementation generated 3 independent ASS Dialogue events per spoken line (`Active → Context → Disappear`) with `\move` animations. When a new line arrived, both the new and previous lines scrolled simultaneously, creating a "2-line block jump" rather than a natural top-to-bottom reading flow. The active line always reset to the bottom position.
|
||||
|
||||
## Solution
|
||||
|
||||
Rewrote the ASS generation to use **2-line paged blocks**:
|
||||
|
||||
- **Paged model:** Spoken lines are grouped into consecutive pairs. Each page is a single ASS Dialogue event with `\N` (hard line break) between lines.
|
||||
- **Continuous karaoke:** `\k` tags flow from line 1 through line 2 within the same Dialogue, creating a natural top-to-bottom reading experience (teleprompter style).
|
||||
- **Cross-fade transitions:** Pages transition via 300ms cross-dissolve using `\fad` tags — no `\move` or `\clip` tags needed.
|
||||
- **Static positioning:** `\an2\pos(960, 1040)` anchors each page at bottom-center. No animation on position.
|
||||
- **Gap-based splitting:** If two consecutive lines are >2s apart, they're split into separate single-line pages.
|
||||
|
||||
## Changes Made
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src-tauri/src/services/clip_exporter.rs` | Added `group_into_pages()`, `build_page_karaoke_text()`. Rewrote `vtt_to_ass_with_karaoke()`. Removed `RollingLayout` struct and dead `build_karaoke_text()`. Fixed stale comments. |
|
||||
| `docs/superpowers/specs/2026-09-22-paged-teleprompter-subtitles-design.md` | Design spec |
|
||||
| `docs/superpowers/plans/2026-09-22-paged-teleprompter-subtitles.md` | Implementation plan |
|
||||
|
||||
## Key Implementation Details
|
||||
|
||||
### Karaoke Stitching Across Lines
|
||||
|
||||
The `build_page_karaoke_text()` function collects word timings from both lines into a flat sequence. The last word of line 1 gets a `\k` duration that extends to line 2's first word start time, naturally covering any gap (silence) between lines. The `\N` is purely visual and doesn't interrupt the karaoke timeline.
|
||||
|
||||
### Cross-Fade Timing
|
||||
|
||||
| Page Position | `\fad` value |
|
||||
|---|---|
|
||||
| First page | `\fad(0, 300)` |
|
||||
| Middle pages | `\fad(300, 300)` |
|
||||
| Last page | `\fad(300, 0)` |
|
||||
|
||||
Display times are extended/preponed by 300ms to create overlap.
|
||||
|
||||
## Commits
|
||||
|
||||
- `934c270` — feat(subtitles): add page grouping and cross-line karaoke stitching
|
||||
- `d13ba27` — feat(subtitles): rewrite ASS generation to paged teleprompter model
|
||||
- `2b95d8a` — chore: remove dead build_karaoke_text, fix stale comments
|
||||
|
||||
## Tests
|
||||
|
||||
23 tests passing (8 new + 15 existing).
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
1. **ASS `\N` doesn't interrupt `\k` flow** — Hard line breaks within a single Dialogue event are purely visual; karaoke timing continues seamlessly across them.
|
||||
2. **`\fad` > `\move` for transitions** — Static positioning with opacity-only transitions (`\fad`) is much simpler and cleaner than coordinate-based animations, especially for avoiding background-box artifacts.
|
||||
3. **Trailing `\N` bug recurrence** — The trailing `\\N` in Dialogue format strings was a recurring bug from previous iterations. Must always check for this when writing ASS output.
|
||||
|
||||
## Follow-Up
|
||||
|
||||
- Manual smoke test needed: export a clip with burn-in captions and verify the teleprompter reading flow.
|
||||
@@ -0,0 +1,66 @@
|
||||
# Extended Caption Styling Implementation Summary
|
||||
|
||||
**Date:** 2026-09-22 10:37
|
||||
**Task:** Implement extended caption styling (font selection, drop shadow, dimmed text color, background toggle)
|
||||
|
||||
## Changes Made
|
||||
|
||||
### Task 1: Data model (commit `a4a1376`)
|
||||
- **`src/lib/stores/preferences.svelte.ts`**: Extended `CaptionSettings` interface with 8 new fields: `fontFamily`, `backgroundEnabled`, `shadowEnabled`, `shadowDepth`, `shadowColor`, `dimmedColorMode`, `dimmedOpacity`, `dimmedColor`
|
||||
- **`src-tauri/src/models.rs`**: Extended `CaptionStyle` Rust struct with matching fields (snake_case, serde camelCase)
|
||||
- **`src/lib/bindings/export.ts`**: Extended TS `CaptionStyle` binding interface
|
||||
- **`src/lib/components/ExportDialog.svelte`**: Updated captionStyle passthrough to include all 17 fields
|
||||
|
||||
### Task 2: System font detection (commit `1447071`)
|
||||
- **`src-tauri/src/commands/media_analysis.rs`**: Added `list_system_fonts` command using `fc-list :lang=en family` with hardcoded fallback (8 common fonts)
|
||||
- **`src-tauri/src/lib.rs`**: Registered new command in `generate_handler!`
|
||||
- **`src/lib/bindings/mediaAnalysis.ts`**: Added `listSystemFonts()` TS binding
|
||||
|
||||
### Task 3: ASS generation wiring (commit `2cf8cdc`)
|
||||
- **`src-tauri/src/services/clip_exporter.rs`**: Rewrote style extraction block in `vtt_to_ass_with_karaoke()`:
|
||||
- Uses `font_family` instead of hardcoded `Arial`
|
||||
- Dimmed color: auto mode derives from `textColor` + `dimmedOpacity`, custom mode uses `dimmedColor` directly
|
||||
- Background/shadow mutual exclusivity: background mode (BorderStyle 3/4 + BackColour for box), shadow mode (BorderStyle 1 + BackColour for shadow + Shadow depth), or neither (transparent BackColour)
|
||||
|
||||
### Task 4: Preview caption rendering (commit `bc39c2a`)
|
||||
- **`src/lib/components/VideoPlayer.svelte`**: Updated `captionStyle` derived to use `$derived.by()`, computing:
|
||||
- `fontFamily`, `fontWeight` from settings
|
||||
- Background conditional on `backgroundEnabled` with `rgba()` using `hexToRgb` helper
|
||||
- Shadow conditional on `shadowEnabled` with depth/color
|
||||
- Dimmed word colors: auto mode uses opacity, custom mode uses direct color
|
||||
|
||||
### Task 5: Settings panel UI (commit `5c47e38`)
|
||||
- **`src/lib/components/CaptionSettingsPanel.svelte`**: Full rewrite with:
|
||||
- System font dropdown (loaded on mount via `listSystemFonts()`, each option styled in its own font)
|
||||
- Font size slider, bold checkbox
|
||||
- Text color picker
|
||||
- Dimmed text: auto/custom radio toggle with opacity slider or color picker
|
||||
- Outline: checkbox + color picker
|
||||
- Background: checkbox + color picker + opacity slider (disables shadow when enabled)
|
||||
- Drop shadow: checkbox + depth slider + color picker (disables background when enabled)
|
||||
- Position: bottom/top radio
|
||||
- Word highlight checkbox
|
||||
- Reset to defaults button
|
||||
- 320px width, scrollable, dark theme
|
||||
|
||||
### Task 6: Verification
|
||||
- `cargo build`: Clean (6 pre-existing warnings)
|
||||
- `cargo test --lib services::clip_exporter::tests`: 23/23 passing
|
||||
- `npm run check`: Only 4 pre-existing errors (node:path/process/url, overload)
|
||||
|
||||
## Approach
|
||||
- Used Subagent-Driven Development: fresh subagent per task + task reviewer per task
|
||||
- 5 implementer dispatches + 5 reviewer dispatches + 1 verification pass
|
||||
- All reviews approved on first pass (no fix cycles needed)
|
||||
|
||||
## Minor Items for Future
|
||||
- Font-family CSS quoting for multi-word names (e.g., `Times New Roman`) in dropdown options
|
||||
- No defensive mutual exclusivity check in preview rendering (relies on settings panel enforcement)
|
||||
- `hexToRgb` has no malformed-hex guard (UI-controlled values only)
|
||||
- `dimmedColorMode` is `string` in Rust/TS binding vs `'auto' | 'custom'` union in preferences
|
||||
- No ASS Style-line unit tests for new field combinations
|
||||
|
||||
## Follow-up Items
|
||||
- Manual smoke test with `cargo tauri dev` recommended
|
||||
- Test font selection with burn-in export
|
||||
- Verify dimmed color auto vs custom modes in both preview and export
|
||||
Reference in New Issue
Block a user