Compare commits

..

20 Commits

Author SHA1 Message Date
sudacode 18790481ff fix(stats): preserve manual episode assignments across reparsing 2026-08-14 01:51:35 -07:00
sudacode 8b71b38a26 test(stats): seed lifetime media rows in legacy redistribution fixture
The reassignAnimeAnilist redistribution test seeded raw sessions and
telemetry but not the imm_lifetime_media rows a finalized session leaves
behind. The old destructive lifetime rebuild re-derived summaries from
raw sessions and hid that gap; the non-destructive recompute reads
per-video lifetime rows, so the fixture must contain them.
2026-08-14 00:26:07 -07:00
sudacode 6311838ef6 fix(stats): preserve lifetime history across anime merges and moves
- Recompute anime aggregates from retained lifetime media summaries
- Prevent sequel title evidence and AniList conflicts from misattributing entries
2026-08-14 00:17:14 -07:00
sudacode 93bdff1ca2 fix(stats): fail closed when queued writes cannot drain
- Guard delete maintenance and AniList reassignment
- Add regression coverage for undrained write queues
2026-08-13 23:32:40 -07:00
sudacode 23e382dff1 fix(stats): preserve season boundaries during AniList merge resolution
- Allow manual cross-season reassignment without merging entries
- Require validated title matches for automatic merges
- Share title normalization across AniList and stats flows
2026-08-13 23:12:08 -07:00
sudacode 528cc798ec fix(stats): review fuzzy AniList duplicates before merging
- Preserve merged title aliases for future episodes
- Fail closed when queued writes cannot drain
2026-08-13 23:12:08 -07:00
sudacode 874fb22dc3 test(stats): assert lifetime rebuild reflects drained telemetry
- queue a telemetry write behind subtitle lines so it lands past the first batch
- assert imm_lifetime_anime picks up linesSeen/activeMs/cards only after the full queue drains
2026-08-13 23:10:15 -07:00
sudacode 8af4570c07 test(stats): move write-queue drain test to its own file
- Extract the mergeAnime/moveVideoToAnime queue-drain test into immersion-tracker-write-queue.test.ts
- Factor shared setup (seedTwoEntries, queueSubtitleLines, countLinesForAnime) into helpers and add a moveVideoToAnime coverage case
2026-08-13 23:10:15 -07:00
sudacode 2ea95c7b13 fix(stats): drain full write queue before anime merge/move rebuilds
- Replace single flushNow() with drainWriteQueue loop so forced telemetry appended after a full batch isn't left unwritten before merge/move/rebuild summaries recompute
- Add dialog a11y to AnimeMergeDialog/LibraryEntryPicker: aria-modal, labelled headings, alert roles for errors, labelled search input, close button labels
2026-08-13 23:10:15 -07:00
sudacode 5abc3b84f4 fix(stats): fix data loss and error handling in anime merge/move
- flush pending telemetry before merge/move so in-progress session time isn't dropped from lifetime totals
- repoint subtitle lines by video_id instead of anime_id so lines recorded before the async title parse assigns a link aren't stranded
- stop absorbing metadata into the target when moving an episode out of an emptied entry (a move isn't a same-show claim)
- return 404 only for missing episode/target, not storage failures; return 404 when a merge folds nothing
- keep AnimeMergeDialog open while a merge is in flight instead of letting dismiss race the request
- surface library load failures in LibraryEntryPicker instead of showing an empty list
2026-08-13 23:10:15 -07:00
sudacode 9c503fc67d feat(stats): add library entry merge and episode move
- Add multi-select "Merge Selected" flow to fold duplicate library cards into one, preserving sessions, mined cards, and watch time
- Add per-episode "move to another entry" action for reassigning stray episodes, pruning the source entry when emptied
- Auto-merge library entries that resolve to the same AniList id when their parsed seasons are compatible
- Add mergeAnime/moveVideoToAnime service methods, stats-server routes, and HTTP contract types
2026-08-13 23:10:15 -07:00
sudacode 47b5903392 chore(release): prepare v0.19.3 2026-08-13 23:02:57 -07:00
sudacode bf85554d1e fix(stats): batch deletes off the main thread (#194) 2026-08-13 22:31:52 -07:00
sudacode d74c7e1235 chore: add Claude instructions symlink and update js-yaml 2026-08-11 22:31:26 -07:00
sudacode 57ddd19953 fix(overlay): handle X11 display scaling across monitors (#193) 2026-08-11 22:24:01 -07:00
sudacode ee25536d90 fix(dictionary): stop large character dictionaries from timing out (#189) 2026-08-11 18:40:18 -07:00
sudacode 7b0fbdf254 fix(subtitles): collapse duplicate ASS events and decode text once (#186) 2026-08-10 22:21:44 -07:00
sudacode 2fefc83e3f fix(playback): stop forcing legacy OpenGL renderer on X11 mpv backend (#188) 2026-08-06 23:52:35 -07:00
sudacode dbdf578c68 perf(tokenizer): single-pass Yomitan scan with cross-line caching and prefetch fixes (#185) 2026-08-06 21:44:09 -07:00
sudacode 441ecf3c04 feat(overlay): add in-app changelog modal (#187) 2026-08-05 22:19:13 -07:00
184 changed files with 16555 additions and 2317 deletions
+24
View File
@@ -1,5 +1,29 @@
# Changelog
## v0.19.3 (2026-08-13)
### Added
- Changelog Modal: Adds an in-app changelog you can open from the tray ("View Changelog") or the "What's New" button on the update notification, so the notification stays reachable while you read. It shows the newest published release notes (falling back to the bundled changelog if that fetch fails), folds older versions while keeping the current one expanded, and supports keyboard navigation (`J`/`K`/arrows, `Enter`, `R`, `Esc`).
### Changed
- Subtitle Tokenization Performance: Reworks subtitle dictionary lookups to cut per-line work roughly in half, cache repeated lookups across lines, and stop tokenization from competing with on-screen subtitle prefetching. Also fixes several accuracy issues along the way: dropped readings on trailing kana, character names being skipped after a dictionary sync, annotations not refreshing after mining a card, and halfwidth katakana character names losing their reading or being swallowed by other words.
### Fixed
- Character Dictionary Large Imports: Large character dictionaries (e.g. One Piece) no longer fail to install from a fixed timeout budget; the import now scales its time budget to dictionary size and reports detailed progress (page/character counts, image download progress, elapsed time) instead of one static message.
- Stats Delete Responsiveness: Deleting sessions, episodes, or library entries no longer freezes the stats page or an active video player; deletes are now batched into a single transaction.
- Styled Subtitle Cue Parsing: Heavily typeset subtitles (karaoke, signs) no longer flood the subtitle sidebar with garbage; vector drawing commands are no longer shown as text, and duplicate/animation-burst cues now collapse into one.
- X11 mpv Renderer: Fixes an mpv crash on the first fullscreen toggle for X11/XWayland users with `gpu-next` shaders (e.g. ArtCNN), which was caused by X11 mode forcing the legacy OpenGL renderer.
- X11 Overlay Display Scaling: Fixes the overlay appearing oversized and offset from mpv on X11/XWayland under fractional or mixed-monitor display scaling.
<details>
<summary>Internal changes</summary>
### Internal
- Subtitle text is now decoded from ASS exactly once at ingest, so the renderer, timing tracker, and tokenizer all share one decoded value instead of each re-deriving it.
- Added per-stage debug timings (`scanMs`, `mecabMs`, `frequencyMs`, `annotateMs`) to the subtitle tokenization pipeline log.
</details>
## v0.19.2 (2026-08-04)
### Changed
Symlink
+1
View File
@@ -0,0 +1 @@
AGENTS.md
+2 -2
View File
@@ -41,7 +41,7 @@
"fast-uri": "3.1.5",
"form-data": "4.0.6",
"ip-address": "10.2.0",
"js-yaml": "4.3.0",
"js-yaml": "4.3.1",
"lodash": "4.18.0",
"minimatch": "10.2.5",
"picomatch": "4.0.4",
@@ -498,7 +498,7 @@
"jiti": ["jiti@2.6.1", "", { "bin": { "jiti": "lib/jiti-cli.mjs" } }, "sha512-ekilCSN1jwRvIbgeg/57YFh8qQDNbwDb9xT/qu2DAHbFFZUicIl4ygVaAvzveMhMVr3LnpSKTNnwt8PoOfmKhQ=="],
"js-yaml": ["js-yaml@4.3.0", "", { "dependencies": { "argparse": "^2.0.1" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-1td788aAnnZ5qs7V2QIRl1owjtYpbKt749Y3xauqQgwIIGF/xXWz1wMTEBx5O3LK3lXLVuqXPdPxj2BoFHaW9Q=="],
"js-yaml": ["js-yaml@4.3.1", "", { "dependencies": { "argparse": "^2.0.1" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-CY6crGq313MX8GkwvB7tzgp99vjQxY1++5y10/BKN/GUfHqWaOGQMNZkBvqSzsZKWk/ijwHlWzzkLulsGHhjWQ=="],
"json-buffer": ["json-buffer@3.0.1", "", {}, "sha512-4bV5BfR2mqfQTJm+V5tPPdf+ZpuhiIvTuAB5g8kcrXOZpTT/QwwVRWBywX1ozr6lEuPdbHxwaJlm9G6mI2sfSQ=="],
+6
View File
@@ -0,0 +1,6 @@
type: added
area: stats
- Library: duplicate cards for the same show can now be combined. Press "Select" above the library grid, tick the cards, and use "Merge Selected"; the dialog picks which entry to keep and moves every episode onto it. Sessions, mined cards, and watch time are preserved, the emptied entries disappear, and remembered title aliases keep future episodes on the merged card.
- Library: episodes can be reassigned to another library entry from the "→" button on an episode row, which is the fix when one file lands under a stray title (e.g. an episode name parsed as the series). Manual assignments now survive later filename parsing, Jellyfin refreshes, and season repair. Compatible local episodes in the same directory reuse a uniquely corrected destination, while conflicting seasons or manual destinations are not forced together. Emptying an entry this way removes it and returns to the grid.
- Library: exact AniList title matches with compatible seasons fold duplicate cards automatically. Fuzzy same-AniList matches appear as dismissible "Possible duplicate" reviews instead of changing the library without confirmation; conflicting explicit seasons are left alone.
+2 -2
View File
@@ -75,8 +75,8 @@ src/
renderer/ # Overlay renderer (modularized UI/runtime)
handlers/ # Keyboard/mouse/gamepad interaction modules
modals/ # Modal flows (Jimaku, Kiku, subsync, runtime options, session help,
# character dictionary, playlist browser, subtitle sidebar,
# YouTube track picker, controller config/debug/select)
# changelog, character dictionary, playlist browser, subtitle
# sidebar, YouTube track picker, controller config/debug/select)
positioning/ # Subtitle position controller (drag-to-reposition)
settings/ # Settings window UI (model, controls, markup)
types/ # Domain type modules (anki, config, integrations, ...)
+24
View File
@@ -1,5 +1,29 @@
# Changelog
## v0.19.3 (2026-08-13)
**Added**
- Changelog Modal: Adds an in-app changelog you can open from the tray ("View Changelog") or the "What's New" button on the update notification, so the notification stays reachable while you read. It shows the newest published release notes (falling back to the bundled changelog if that fetch fails), folds older versions while keeping the current one expanded, and supports keyboard navigation (`J`/`K`/arrows, `Enter`, `R`, `Esc`).
**Changed**
- Subtitle Tokenization Performance: Reworks subtitle dictionary lookups to cut per-line work roughly in half, cache repeated lookups across lines, and stop tokenization from competing with on-screen subtitle prefetching. Also fixes several accuracy issues along the way: dropped readings on trailing kana, character names being skipped after a dictionary sync, annotations not refreshing after mining a card, and halfwidth katakana character names losing their reading or being swallowed by other words.
**Fixed**
- Character Dictionary Large Imports: Large character dictionaries (e.g. One Piece) no longer fail to install from a fixed timeout budget; the import now scales its time budget to dictionary size and reports detailed progress (page/character counts, image download progress, elapsed time) instead of one static message.
- Stats Delete Responsiveness: Deleting sessions, episodes, or library entries no longer freezes the stats page or an active video player; deletes are now batched into a single transaction.
- Styled Subtitle Cue Parsing: Heavily typeset subtitles (karaoke, signs) no longer flood the subtitle sidebar with garbage; vector drawing commands are no longer shown as text, and duplicate/animation-burst cues now collapse into one.
- X11 mpv Renderer: Fixes an mpv crash on the first fullscreen toggle for X11/XWayland users with `gpu-next` shaders (e.g. ArtCNN), which was caused by X11 mode forcing the legacy OpenGL renderer.
- X11 Overlay Display Scaling: Fixes the overlay appearing oversized and offset from mpv on X11/XWayland under fractional or mixed-monitor display scaling.
<details>
<summary>Internal changes</summary>
**Internal**
- Subtitle text is now decoded from ASS exactly once at ingest, so the renderer, timing tracker, and tokenizer all share one decoded value instead of each re-deriving it.
- Added per-stage debug timings (`scanMs`, `mecabMs`, `frequencyMs`, `annotateMs`) to the subtitle tokenization pipeline log.
</details>
## v0.19.2 (2026-08-04)
**Changed**
+7
View File
@@ -57,6 +57,13 @@ Jellyfin stream URLs are normalized to stable item links before stats titles are
When YouTube channel metadata is available, the Library tab groups videos by creator/channel and treats each tracked video as an episode-like entry inside that channel section.
A library entry is identified by its parsed title plus any detected season, so the same show can end up on several cards when releases disagree about the title or omit the season tag. Two fixes are available:
- **Merge duplicates.** Hit **Select** above the grid, tick the cards that are the same show, and choose **Merge Selected**. Pick which entry to keep in the dialog; every episode moves onto it and the other cards are removed. Nothing is deleted, so sessions, mined cards and watch time all carry over. SubMiner remembers the merged title variants, so future episodes parsed with one of those names join the kept entry instead of recreating a duplicate card.
- **Move a single episode.** Hover an episode row in a title's episode list and use the **→** button to reassign it to another library entry. The correction is remembered, so later filename parsing or Jellyfin metadata cannot move that episode back. For local files, later episodes in the same directory inherit the correction when their detected seasons are compatible and every manual correction there points to the same entry. Conflicting seasons or manual destinations are left for review. If the move empties the old entry, that card is removed and you are returned to the grid.
Once cover art resolves a series to an AniList entry, cards with compatible seasons are folded together automatically only when the searched title exactly matches an AniList title or synonym. A fuzzy result that points at an AniList entry already used by another card appears as a **Possible duplicate** review above the Library grid instead. Choose **Review merge** to compare the cards and pick which one to keep, or **Not duplicates** to dismiss that suggestion permanently. Entries with conflicting explicit season numbers are left alone rather than merged or suggested.
Open a title and use **Delete Entry** in its header to remove a mistakenly tracked show outright. This deletes every episode of that title along with their sessions, subtitle lines, rollups and cover art, drops the words and kanji that were only seen there, and removes the card from the Library grid. Individual episodes and sessions can still be deleted on their own from the episode list and session rows. Entry deletion is refused while that title is the one currently playing.
![Stats Library](/screenshots/stats-library.png)
+4 -3
View File
@@ -405,8 +405,9 @@ On any Wayland session that is not Hyprland or Sway (KDE Plasma, GNOME, and othe
SubMiner handles this automatically:
- It launches its own window under XWayland (it sets `--ozone-platform-hint=x11`).
- Every mpv it launches (via the `subminer` launcher, Jellyfin, or YouTube) is pinned to XWayland too - Wayland environment hints are stripped and an X11 GPU context (`--gpu-context=x11egl,x11`) is applied.
- It launches its own window under XWayland (it sets `--ozone-platform=x11`).
- Every mpv it launches (via the `subminer` launcher, Jellyfin, or YouTube) is pinned to XWayland too - Wayland environment hints are stripped and an X11 GPU context (`--gpu-context=x11vk,x11egl,x11`) is applied. Only the window context is overridden; your `vo`/`gpu-api` and user shaders are left alone.
- Fractional and mixed-monitor display scaling is handled per screen when SubMiner maps XWayland mpv coordinates to the overlay.
- While mpv is windowed, the overlay is a managed X11 window owned by the tracked mpv window (`WM_TRANSIENT_FOR`), so it stays above mpv while other foreground X11/Xwayland apps can still cover both windows.
- While tracked mpv is fullscreen, SubMiner swaps the visible overlay to a focusable-false X11 override-redirect window. That path can stay above the active fullscreen mpv window without requiring a KDE/KWin-specific rule, and SubMiner hides/releases it when mpv is no longer the active X11/Xwayland window.
- The visible overlay is shown inactive on Linux, so normal hover should not steal keyboard focus from mpv.
@@ -420,7 +421,7 @@ Requirements: `xdotool`, `xprop`, and `xwininfo` must be installed. SubMiner use
This almost always means mpv came up as a **native Wayland** window that the XWayland overlay cannot cover. It happens when mpv is launched **manually** (your own command), because SubMiner can only force XWayland on the mpv processes it launches itself. Fix it one of these ways:
- Launch playback through SubMiner (the `subminer` launcher or the tray), which forces XWayland for you, or
- Force XWayland in your own mpv invocation, e.g. `mpv --gpu-context=x11egl …`, or launch with `WAYLAND_DISPLAY= mpv …`, or set `gpu-context=x11egl` in your `mpv.conf`.
- Force XWayland in your own mpv invocation, e.g. `mpv --gpu-context=x11vk,x11egl,x11 …`, or launch with `WAYLAND_DISPLAY= mpv …`, or set `gpu-context=x11vk` (Vulkan) / `gpu-context=x11egl` (OpenGL) in your `mpv.conf`.
To confirm mpv is on XWayland, `xdotool search --class mpv` should return a window id (a native Wayland mpv returns nothing).
+4
View File
@@ -145,6 +145,8 @@ The tray menu includes `Export Logs`, which creates the same sanitized local-dat
Once Jellyfin is configured, the tray menu includes `Jellyfin Discovery` for starting or stopping cast discovery in the current app session without changing config.
The tray menu also includes `View Changelog`, which opens the in-app changelog modal. It fetches the changelog from the newest published release, so you see release notes for versions newer than the one you run; if the download fails it falls back to the changelog bundled with your install and says so. Versions in the current `0.x` line are expanded by default and older lines are folded, matching this site's [Changelog](/changelog). A badge marks the version you have installed, and newer versions are tagged `New`. The same modal opens from the `What's New` button on the update-available overlay notification.
### Logging and App Mode
- `--log-level` controls logger verbosity.
@@ -368,6 +370,8 @@ Press `V` to cycle the primary SubMiner subtitle bar through hidden → visible
`Ctrl/Cmd+/` opens the session help modal with the current overlay and mpv keybindings. The same help view is also available through the `y-h` chord in mpv.
The changelog modal (tray > `View Changelog`) works the same way: it renders over mpv when a video is playing and in its own window otherwise. Use `J`/`K` or the arrow keys to move between versions, `Enter` to fold or unfold one, `R` to refetch, and `Esc` to close.
Hovering over subtitle text pauses mpv by default; leaving resumes it. Yomitan popups also pause playback by default. Set `subtitleStyle.autoPauseVideoOnHover: false` or `subtitleStyle.autoPauseVideoOnYomitanPopup: false` to disable either behavior.
### Drag-and-Drop
@@ -64,18 +64,23 @@ External subtitle files only (SRT, VTT, ASS). Embedded subtitle tracks are out o
A cue parser extracts both timing and text content from subtitle files for prefetching.
**Parsed cue structure:**
```typescript
interface SubtitleCue {
startTime: number; // seconds
endTime: number; // seconds
text: string; // raw subtitle text
text: string; // plain text, decoded from the source format
}
```
**Supported formats:**
- SRT/VTT: Regex-based parsing of timing lines + text content between timing blocks.
- ASS: Parse `[Events]` section, extract `Dialogue:` lines, split on the first 9 commas only (ASS v4+ has 10 fields; the last field is Text which can itself contain commas). Strip ASS override tags (`{\...}`) from the text before storing.
ASS text fields contain inline override tags like `{\b1}`, `{\an8}`, `{\fad(200,300)}`. The cue parser strips these during extraction so the tokenizer receives clean text.
- ASS: Parse `[Events]` section, extract `Dialogue:` lines, read the field order from the `Format:` row, and take everything after the Text field index as the text (Text can itself contain commas).
**ASS decoding.** The parser is where ASS text is decoded, once, via `assToPlainText()` in `src/core/services/ass-text.ts`. That decoder mirrors mpv's `ass_to_plaintext` so a cue read from a file reads identically to the same line arriving live on `sub-text`: `{...}` override blocks are markup, `\pN … \p0` vector drawing runs are dropped rather than shown as text, `\N`/`\n`/`\h` are the only escapes (`\{`, `\}` and `\\` are not), and an unclosed `{` is rendered verbatim. Every layer downstream — renderer, timing tracker, tokenizer, tokenization cache keys — receives plain text and uses `normalizePlainSubtitleText()` for whitespace only, so nothing decodes the same string twice and one authored line always maps to one cache key.
**Duplicate collapsing.** Typeset scripts emit one `Dialogue:` event per animation frame, plus layered copies of the same line. The parser collapses identical text over an identical span unconditionally, and collapses contiguous same-text runs of at least three events when the run looks like an animation. For ASS that means shared style and actor plus authoring evidence: a temporal tag (`\t`, `\move`, `\k`/`\kf`/`\ko`/`\K`, or anything wrapped in `\t(...)`), an animated `Effect` column (`Karaoke`, `Banner`, `Scroll`), or override values that change across the run. Static tags shared by every event (`\pos`, an identical `\clip`) are not evidence. SRT/VTT carry no such metadata, so there collapsing needs at least five contiguous events all under 0.1s — the frame timing left behind by ASS-to-SRT conversion. The parser keeps this authoring metadata (style, actor, layer, `Effect`, parsed override commands, source order) private; `parseSubtitleCues()` returns only `SubtitleCue`.
#### Prefetch Service Lifecycle
@@ -153,6 +158,7 @@ tokens (already have frequencyRank values from parser-level applyFrequencyRanks)
### Dependency Analysis
All annotations either depend on MeCab POS data or benefit from running after it:
- **Known word marking:** Needs base tokens (surface/headword). No POS dependency, but no reason to run separately.
- **Frequency filtering:** Uses `pos1Exclusions` and `pos2Exclusions` to clear frequency ranks on excluded tokens (particles, noise). Depends on MeCab POS data.
- **JLPT marking:** Uses `shouldIgnoreJlptForMecabPos1` to filter. Depends on MeCab POS data.
@@ -169,18 +175,14 @@ function annotateTokens(tokens, deps, options): MergedToken[] {
// Single pass: known word + frequency filtering + JLPT computed together
const annotated = tokens.map((token) => {
const isKnown = nPlusOneEnabled
? token.isKnown || computeIsKnown(token, deps)
: false;
const isKnown = nPlusOneEnabled ? token.isKnown || computeIsKnown(token, deps) : false;
// Filter frequency rank using POS exclusions (rank values already set at parser level)
const frequencyRank = frequencyEnabled
? filterFrequencyRank(token, pos1Exclusions, pos2Exclusions)
: undefined;
const jlptLevel = jlptEnabled
? computeJlptLevel(token, deps.getJlptLevel)
: undefined;
const jlptLevel = jlptEnabled ? computeJlptLevel(token, deps.getJlptLevel) : undefined;
return { ...token, isKnown, frequencyRank, jlptLevel };
});
@@ -221,6 +223,7 @@ Replace `document.createElement('span')` calls in the renderer with `templateSpa
### Current Behavior
In `renderWithTokens` (`subtitle-render.ts`), each render cycle:
1. Clears DOM with `innerHTML = ''`
2. Creates a `DocumentFragment`
3. Calls `document.createElement('span')` for each token (~10-15 per subtitle)
@@ -257,7 +260,7 @@ Full recycling (collecting old nodes, clearing attributes, reusing them) require
## Combined Impact Summary
| Scenario | Before | After | Improvement |
|----------|--------|-------|-------------|
| --------------------------------- | ---------- | ---------- | ----------- |
| Normal playback (prefetch-warmed) | ~200-320ms | ~30-50ms | ~80-85% |
| Cache hit (repeated subtitle) | ~72ms | ~55-65ms | ~10-20% |
| Cache miss (immediate seek) | ~200-320ms | ~150-260ms | ~20-25% |
@@ -267,16 +270,19 @@ Full recycling (collecting old nodes, clearing attributes, reusing them) require
## Files Summary
### New Files
- `src/core/services/subtitle-prefetch.ts`
- `src/core/services/subtitle-cue-parser.ts`
### Modified Files
- `src/core/services/subtitle-processing-controller.ts` (expose `preCacheTokenization`)
- `src/core/services/tokenizer/annotation-stage.ts` (batched single-pass)
- `src/renderer/subtitle-render.ts` (template cloneNode)
- `src/main.ts` (wire up prefetch service)
### Test Files
- New tests for subtitle cue parser (SRT, VTT, ASS formats)
- New tests for subtitle prefetch service (priority window, seek, pause/resume)
- Updated tests for annotation stage (same behavior, new implementation)
+2
View File
@@ -25,6 +25,8 @@ Read when: you need to find the owner module for a behavior or test surface
- Anki workflow: `src/anki-integration/`, `src/core/services/anki-jimaku*.ts`
- Immersion tracking: `src/core/services/immersion-tracker/`
Includes stats storage/query schema such as `imm_videos`, `imm_media_art`, and `imm_youtube_videos` for per-video and YouTube-specific library metadata.
Library-entry identity aliases and merge recommendations are persisted alongside this schema; the stats HTTP and SPA layers only expose and present those domain decisions.
`delete-maintenance-scheduler.ts` coalesces and serializes stats deletes; expensive deletion and summary rebuilds run in `delete-maintenance-worker-thread.ts` while the tracker queues playback writes. Each batch uses one transaction, lexical update, rollup refresh, and lifetime rebuild.
- AniList tracking + character dictionary: `src/core/services/anilist/`, `src/main/runtime/composers/anilist-*`, `src/main/character-dictionary-runtime.ts`, `src/main/character-dictionary-runtime/`
- Jellyfin integration: `src/core/services/jellyfin*.ts`, `src/main/runtime/composers/jellyfin-*`
- Window trackers: `src/window-trackers/`
+16 -4
View File
@@ -50,10 +50,22 @@ subtitles do not draw.
7. Cache miss: call `refreshCurrentSubtitle(text)`. Normal processing emits a plain payload
synchronously, then replaces it with the tokenized payload when ready.
In `src/main.ts`, both `onSubtitleChange` and `refreshCurrentSubtitle` pause
`subtitlePrefetchService`, notify it with `onSeek(lastObservedTimePos)`, and then call the matching
`subtitleProcessingController` method. This gives the visible overlay priority over background
prefetch work and re-centers prefetch around the live playback time.
Both `onSubtitleChange` and `refreshCurrentSubtitle` pause `subtitlePrefetchService` and then call
the matching `subtitleProcessingController` method, giving the visible overlay priority over
background prefetch work. Prefetch is not re-centered here: restarting the run per line
(`onSeek`) discarded the in-flight tokenization every time the subtitle changed, so only real
seeks restart it (see `onTimePosUpdate` in `src/main.ts`).
On an uncached autoplay prime the raw payload is emitted here and reported to the controller with
`notePlainSubtitleEmitted`, so the controller skips its own plain emit for that line and the
overlay receives one plain payload followed by the annotated one.
The pause is released by the controller's `onProcessingSettled` callback, which fires once it has
no work left. Emits do not release it: the first emit for an uncached line is the plain payload
that precedes tokenization, and a run can finish without emitting at all (a suppressed duplicate,
a failed tokenization). Both controller methods return whether processing is now pending, and the
caller resumes immediately when it is not — a repeated subtitle schedules no work, so no settle is
coming and prefetching would otherwise idle for the rest of the cue.
## Live Cue Delivery
+5 -7
View File
@@ -222,7 +222,7 @@ test('buildMpvEnv preserves native Wayland env for supported Hyprland and Sway a
});
});
test('buildMpvBackendArgs forces an explicit X11 renderer stack when backend resolves to x11', () => {
test('buildMpvBackendArgs pins the X11 window context when backend resolves to x11', () => {
withPlatform('linux', () => {
assert.deepEqual(
buildMpvBackendArgs(makeArgs({ backend: 'x11' }), {
@@ -230,12 +230,12 @@ test('buildMpvBackendArgs forces an explicit X11 renderer stack when backend res
WAYLAND_DISPLAY: 'wayland-0',
XDG_SESSION_TYPE: 'wayland',
}),
['--vo=gpu', '--gpu-api=opengl', '--gpu-context=x11egl,x11'],
['--gpu-context=x11vk,x11egl,x11'],
);
});
});
test('buildMpvBackendArgs forces the same X11 renderer stack for unsupported Wayland auto fallback', () => {
test('buildMpvBackendArgs pins the same X11 window context for unsupported Wayland auto fallback', () => {
withPlatform('linux', () => {
assert.deepEqual(
buildMpvBackendArgs(makeArgs({ backend: 'auto' }), {
@@ -245,7 +245,7 @@ test('buildMpvBackendArgs forces the same X11 renderer stack for unsupported Way
XDG_CURRENT_DESKTOP: 'KDE',
XDG_SESSION_DESKTOP: 'plasma',
}),
['--vo=gpu', '--gpu-api=opengl', '--gpu-context=x11egl,x11'],
['--gpu-context=x11vk,x11egl,x11'],
);
});
});
@@ -292,9 +292,7 @@ test('buildConfiguredMpvDefaultArgs appends maximized launch mode to configured
'--secondary-sub-visibility=no',
'--alang=ja,jp,jpn,japanese,en,eng,english,enus,en-us',
'--slang=ja,jp,jpn,japanese,en,eng,english,enus,en-us',
'--vo=gpu',
'--gpu-api=opengl',
'--gpu-context=x11egl,x11',
'--gpu-context=x11vk,x11egl,x11',
'--window-maximized=yes',
],
);
+13
View File
@@ -80,6 +80,11 @@ test('merges remote-only sessions with catalog, lifetime, and rollups', () => {
{ headword: '食べる', word: '食べた', reading: 'たべた', count: 1 },
],
});
withWritableDb(remotePath, (db) => {
db.prepare(
`UPDATE imm_videos SET anime_assignment_locked = 1 WHERE video_key = 'showb-e1'`,
).run();
});
const summary = mergeSnapshotIntoDb(localPath, remotePath);
assert.equal(summary.sessionsMerged, 1);
@@ -126,6 +131,14 @@ test('merges remote-only sessions with catalog, lifetime, and rollups', () => {
`SELECT video_id FROM imm_videos WHERE video_key = 'showb-e1'`,
)?.video_id,
);
assert.equal(
queryOne<{ locked: number }>(
localPath,
'SELECT anime_assignment_locked AS locked FROM imm_videos WHERE video_id = ?',
[mergedVideoId],
)?.locked,
1,
);
assert.equal(
count(localPath, 'SELECT COUNT(*) AS n FROM imm_daily_rollups WHERE video_id = ?', [
mergedVideoId,
+2 -1
View File
@@ -1,4 +1,4 @@
// Schema-version-18 shape of the tables the sync merge touches (plus the
// Current schema shape of the tables the sync merge touches (plus the
// app's indexes), mirroring ensureSchema / ensureLifetimeSummaryTables /
// ensureStatsExcludedWordsTable in src/core/services/immersion-tracker/storage.ts.
export const IMMERSION_DB_FIXTURE_DDL = `
@@ -39,6 +39,7 @@ export const IMMERSION_DB_FIXTURE_DDL = `
parser_source TEXT,
parser_confidence REAL,
parse_metadata_json TEXT,
anime_assignment_locked INTEGER NOT NULL DEFAULT 0 CHECK(anime_assignment_locked IN (0, 1)),
watched INTEGER NOT NULL DEFAULT 0,
duration_ms INTEGER NOT NULL CHECK(duration_ms>=0),
file_size_bytes INTEGER CHECK(file_size_bytes>=0),
+6 -2
View File
@@ -2,7 +2,7 @@
"name": "subminer",
"productName": "SubMiner",
"desktopName": "SubMiner.desktop",
"version": "0.19.2",
"version": "0.19.3",
"description": "All-in-one sentence mining overlay with AnkiConnect and dictionary integration",
"packageManager": "bun@1.3.5",
"main": "dist/main-entry.js",
@@ -89,7 +89,7 @@
"fast-uri": "3.1.5",
"form-data": "4.0.6",
"ip-address": "10.2.0",
"js-yaml": "4.3.0",
"js-yaml": "4.3.1",
"lodash": "4.18.0",
"minimatch": "10.2.5",
"picomatch": "4.0.4",
@@ -260,6 +260,10 @@
{
"from": "dist/launcher/subminer",
"to": "launcher/subminer"
},
{
"from": "CHANGELOG.md",
"to": "CHANGELOG.md"
}
]
},
+28 -20
View File
@@ -1,30 +1,38 @@
## Highlights
### Changed
### Added
- **In-App Changelog**
- View release notes without leaving the app, from the tray menu ("View Changelog") or the "What's New" button on update notifications.
- Shows notes for the latest published release even when it's newer than your installed build, and falls back to the notes bundled with your install if the download fails.
- Older versions fold automatically, your installed version is badged, and newer ones are tagged "New"; navigate with `J`/`K` or the arrow keys, `Enter` to expand/collapse, `R` to refresh, and `Esc` to close.
- **Subsync Reference & Target Picker**
- You can now choose both sides of a sync run: which subtitle is the timing reference and which one gets retimed.
- The video file itself can be used as the reference for local files (audio-based sync), though a subtitle track stays the default.
- Works for both alass and ffsubsync, and retiming the secondary subtitle track no longer overwrites your primary one.
### Changed
- **Faster Subtitle Tokenization**
- Subtitle lines are parsed and looked up roughly twice as efficiently, with results cached across lines so repeated words and grammar no longer re-query the dictionary.
- Enabling a character dictionary no longer slows subtitle scanning as much, since name lookups now only check positions where a known name can actually start.
- Fixed related accuracy issues along the way: readings that could go missing on certain word endings, subtitle text that stayed unannotated after mining a card, character names that could drop out of disambiguation rules, and halfwidth-katakana character names that weren't recognized or read correctly.
### Fixed
- **Startup Logging**
- Background startup now respects your configured log level even when no `--log-level` flag is passed.
- **Streaming Subtitle Tokenization**
- Jellyfin playback now seeds tokenization straight from the downloaded subtitle file, so episodes no longer fall back to slow, line-by-line tokenizing while waiting on playback events.
- Subtitle cues are no longer dropped when switching to a subtitle track embedded in the stream.
- Prefetching now runs through the whole episode instead of stopping once the cache filled, and the cache clears between episodes so slowdowns don't carry over to later titles.
- The tokenization cache was expanded from 256 to 2,500 lines, leaving more room for repeated lines (like openings and endings) to stay cached across episodes.
- **Subtitle Line Display**
- Subtitle lines now appear immediately at their cue time even if tokenization hasn't finished, upgrading in place with annotations once ready.
- A failed tokenization attempt is no longer cached as plain text, so the line gets another chance at full annotations later.
- **Large Character Dictionary Generation**
- Big character dictionaries (long-running series like One Piece) no longer fail to install with a timeout error; the import time budget now scales with dictionary size instead of using a fixed 7-second limit.
- The "Generating character dictionary" notification now shows real progress (character/page counts, image download progress with an ETA, name-processing progress) and an elapsed-time clock, so a long-running import no longer looks frozen.
- **Stats Deletion Responsiveness**
- Deleting sessions, episodes, or library entries on the stats page no longer freezes the page or an active video player; deletes are now batched into a single transaction.
- **Subtitle Sidebar Clutter from Styled Subtitles**
- Heavily typeset subtitles (karaoke openings/endings, stylized signs) no longer flood the subtitle sidebar with garbled vector-drawing text or duplicate "shadow" copies of the same line.
- Subtitle text is now decoded consistently in one place, matching what mpv actually renders on screen, so it can no longer diverge or get cached inconsistently.
- **X11/XWayland Playback and Overlay Fixes**
- Fixed a crash on the first fullscreen toggle when using an mpv `gpu-next` shader (e.g. ArtCNN) in X11/XWayland mode; SubMiner no longer forces mpv onto its older OpenGL renderer.
- Fixed the overlay appearing oversized and offset from the video under fractional or mixed-monitor display scaling in X11/XWayland mode.
## What's Changed
- feat(subsync): add reference and target subtitle track picker by @ksyasuda in #181
- fix(logging): surface subtitle processing debug/warn logs by @ksyasuda in #182
- fix(streaming): keep subtitle tokenization prefetch warm for full episodes by @ksyasuda in #183
- fix(overlay): show plain subtitle line immediately on tokenization cache miss by @ksyasuda in #184
- perf(tokenizer): single-pass Yomitan scan with cross-line caching and prefetch fixes by @ksyasuda in #185
- fix(subtitles): collapse duplicate ASS events and decode text once by @ksyasuda in #186
- feat(overlay): add in-app changelog modal by @ksyasuda in #187
- fix(playback): stop forcing legacy OpenGL renderer on X11 mpv backend by @ksyasuda in #188
- fix(dictionary): stop large character dictionaries from timing out by @ksyasuda in #189
- fix(overlay): handle X11 display scaling across monitors by @ksyasuda in #193
- fix(stats): batch deletes off the main thread by @ksyasuda in #194
## Installation
@@ -0,0 +1,322 @@
import test from 'node:test';
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import type { DatabaseSync } from '../immersion-tracker/sqlite';
type ImmersionTrackerService = import('../immersion-tracker-service').ImmersionTrackerService;
type ImmersionTrackerServiceCtor =
typeof import('../immersion-tracker-service').ImmersionTrackerService;
let trackerCtor: ImmersionTrackerServiceCtor | null = null;
async function loadTrackerCtor(): Promise<ImmersionTrackerServiceCtor> {
if (trackerCtor) return trackerCtor;
const mod = await import('../immersion-tracker-service');
trackerCtor = mod.ImmersionTrackerService;
return trackerCtor;
}
function makeDbPath(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-write-queue-test-'));
return path.join(dir, 'immersion.sqlite');
}
function cleanupDbPath(dbPath: string): void {
const dir = path.dirname(dbPath);
if (!fs.existsSync(dir)) return;
fs.rmSync(dir, { recursive: true, force: true });
}
interface TrackerInternals {
db: DatabaseSync;
queue: unknown[];
recordWrite: (write: Record<string, unknown>) => void;
deleteSession: (sessionId: number) => Promise<void>;
mergeAnime: (targetAnimeId: number, sourceAnimeIds: number[]) => Promise<unknown>;
moveVideoToAnime: (videoId: number, targetAnimeId: number) => Promise<unknown>;
rebuildLifetimeSummaries: () => Promise<unknown>;
reassignAnimeAnilist: (animeId: number, info: { anilistId: number }) => Promise<void>;
flushNow: () => void;
writeLock: { locked: boolean };
}
test('delete maintenance fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
let deleteRunnerCalls = 0;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor(
{ dbPath, policy: { batchSize: 2 } },
{
runDeleteMaintenanceTask: async () => {
deleteRunnerCalls += 1;
},
},
);
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
let flushCalls = 0;
internals.flushNow = () => {
flushCalls += 1;
if (flushCalls > 1) throw new Error('bounded no-progress sentinel');
};
await assert.rejects(internals.deleteSession(1), /queue did not drain/i);
assert.equal(flushCalls, 1);
assert.equal(deleteRunnerCalls, 0);
assert.equal(internals.writeLock.locked, false);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('reassignAnimeAnilist fails closed before resolving a conflict when writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
internals.db.prepare('UPDATE imm_anime SET anilist_id = 123 WHERE anime_id = 2').run();
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(
internals.reassignAnimeAnilist(1, { anilistId: 123 }),
/queue did not drain/i,
);
assert.deepEqual(
internals.db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all(),
[
{ animeId: 1, anilistId: null },
{ animeId: 2, anilistId: 123 },
],
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('mergeAnime fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.mergeAnime(1, [2]), /queue did not drain/i);
assert.deepEqual(
internals.db
.prepare('SELECT anime_id AS animeId FROM imm_anime ORDER BY anime_id')
.all()
.map((row) => (row as { animeId: number }).animeId),
[1, 2],
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('moveVideoToAnime fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.moveVideoToAnime(2, 1), /queue did not drain/i);
assert.equal(
(
internals.db
.prepare('SELECT anime_id AS animeId FROM imm_videos WHERE video_id = 2')
.get() as {
animeId: number;
}
).animeId,
2,
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('rebuildLifetimeSummaries fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.rebuildLifetimeSummaries(), /queue did not drain/i);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
function seedTwoEntries(db: DatabaseSync): void {
db.exec(`
INSERT INTO imm_anime (anime_id, normalized_title_key, canonical_title, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'show', 'Show', 1000, 1000), (2, 'show season 1', 'Show Season 1', 1000, 1000);
INSERT INTO imm_videos (video_id, video_key, canonical_title, anime_id, source_type, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'local:/tmp/a.mkv', 'A', 1, 1, 0, 1440000, 1000, 1000),
(2, 'local:/tmp/b.mkv', 'B', 2, 1, 0, 1440000, 1000, 1000);
INSERT INTO imm_sessions (session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, active_watched_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'drain-session', 2, '1000', '2000', 2, 1000, 1000, 2000);
`);
}
function queueSubtitleLines(tracker: TrackerInternals, count: number): void {
for (let index = 0; index < count; index += 1) {
tracker.recordWrite({
kind: 'subtitleLine',
sessionId: 1,
videoId: 2,
lineIndex: index,
segmentStartMs: index * 1000,
segmentEndMs: index * 1000 + 900,
text: `line ${index}`,
wordOccurrences: [],
kanjiOccurrences: [],
firstSeen: 1000,
lastSeen: 2000,
});
}
}
/**
* Queued last so it sits past the first batch. Lifetime `total_lines_seen`
* reads this counter, not a COUNT over imm_subtitle_lines, so the rebuilt
* summary only reflects the session once the queue is drained all the way.
*/
function queueTelemetry(tracker: TrackerInternals, linesSeen: number): void {
tracker.recordWrite({
kind: 'telemetry',
sessionId: 1,
sampleMs: 3000,
lastMediaMs: 3000,
totalWatchedMs: 4000,
activeWatchedMs: 3500,
linesSeen,
tokensSeen: linesSeen * 5,
cardsMined: 2,
lookupCount: 0,
lookupHits: 0,
yomitanLookupCount: 0,
pauseCount: 0,
pauseMs: 0,
seekForwardCount: 0,
seekBackwardCount: 0,
mediaBufferEvents: 0,
});
}
/** The queued telemetry sample only exists in the database once the queue drained fully. */
function latestTelemetryLinesSeen(db: DatabaseSync, sessionId: number): number | null {
const row = db
.prepare(
`SELECT lines_seen AS linesSeen
FROM imm_session_telemetry
WHERE session_id = ?
ORDER BY sample_ms DESC, telemetry_id DESC
LIMIT 1`,
)
.get(sessionId) as { linesSeen: number } | undefined;
return row ? Number(row.linesSeen) : null;
}
function countLinesForAnime(db: DatabaseSync, animeId: number): number {
const row = db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id = ?')
.get(animeId) as { total: number };
return Number(row.total);
}
/**
* Both entry points must see a settled database before changing episode
* ownership. A single flushNow() only writes one batch off the front of the
* queue, so anything past `batchSize` would still be unwritten when the merge
* repoints rows.
*/
test('mergeAnime drains a queue larger than one batch before repointing rows', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 8);
queueTelemetry(internals, 8);
assert.ok(internals.queue.length > 2, 'expected more queued writes than one batch');
await internals.mergeAnime(1, [2]);
assert.equal(internals.queue.length, 0);
// Every queued line landed, attributed to the surviving entry.
assert.equal(countLinesForAnime(internals.db, 1), 8);
assert.equal(latestTelemetryLinesSeen(internals.db, 1), 8);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('moveVideoToAnime drains a queue larger than one batch before repointing rows', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 8);
queueTelemetry(internals, 8);
await internals.moveVideoToAnime(2, 1);
assert.equal(internals.queue.length, 0);
assert.equal(countLinesForAnime(internals.db, 1), 8);
assert.equal(latestTelemetryLinesSeen(internals.db, 1), 8);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
@@ -1053,6 +1053,55 @@ describe('stats server API routes', () => {
assert.equal(body[0].canonicalTitle, 'Little Witch Academia');
});
it('GET /api/stats/anime/merge-recommendations returns pending duplicate pairs', async () => {
const app = createStatsApp(
createMockTracker({
getAnimeMergeRecommendations: async () => [{ recommendationId: 4, animeIds: [1, 2] }],
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations');
assert.equal(res.status, 200);
assert.deepEqual(await res.json(), {
recommendations: [{ recommendationId: 4, animeIds: [1, 2] }],
});
});
it('DELETE /api/stats/anime/merge-recommendations/:id dismisses a pending pair', async () => {
let dismissedId: number | null = null;
const app = createStatsApp(
createMockTracker({
dismissAnimeMergeRecommendation: async (recommendationId: number) => {
dismissedId = recommendationId;
return true;
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations/4', {
method: 'DELETE',
});
assert.equal(res.status, 200);
assert.equal(dismissedId, 4);
assert.deepEqual(await res.json(), { ok: true });
});
it('DELETE /api/stats/anime/merge-recommendations/:id reports missing recommendations', async () => {
const app = createStatsApp(
createMockTracker({
dismissAnimeMergeRecommendation: async () => false,
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations/99', {
method: 'DELETE',
});
assert.equal(res.status, 404);
});
it('GET /api/stats/anime/:animeId returns anime detail with episodes', async () => {
const app = createStatsApp(createMockTracker());
const res = await app.request('/api/stats/anime/1');
@@ -3024,6 +3073,148 @@ Aligned English subtitle
assert.equal(deleteCalls, 0);
});
it('POST /api/stats/anime/:animeId/merge folds the given entries into the target', async () => {
let merged: { targetAnimeId: number; sourceAnimeIds: number[] } | null = null;
const app = createStatsApp(
createMockTracker({
mergeAnime: async (targetAnimeId: number, sourceAnimeIds: number[]) => {
merged = { targetAnimeId, sourceAnimeIds };
return {
survivingAnimeId: targetAnimeId,
mergedAnimeIds: sourceAnimeIds,
movedVideos: 3,
};
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
// The target repeated in the sources must not delete the entry we keep.
body: '{"sourceAnimeIds":[8,9,8,7]}',
});
assert.equal(res.status, 200);
assert.deepEqual(merged, { targetAnimeId: 7, sourceAnimeIds: [8, 9] });
assert.deepEqual(await res.json(), {
ok: true,
animeId: 7,
mergedAnimeIds: [8, 9],
movedVideos: 3,
});
});
it('POST /api/stats/anime/:animeId/merge rejects an empty or malformed source list', async () => {
let mergeCalls = 0;
const app = createStatsApp(
createMockTracker({
mergeAnime: async () => {
mergeCalls += 1;
return { survivingAnimeId: 7, mergedAnimeIds: [], movedVideos: 0 };
},
} as Partial<ImmersionTrackerService>),
);
for (const body of [
'{"sourceAnimeIds":[]}',
'{"sourceAnimeIds":[7]}',
'{"sourceAnimeIds":0}',
]) {
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body,
});
assert.equal(res.status, 400);
}
assert.equal(mergeCalls, 0);
});
it('PATCH /api/stats/media/:videoId/anime moves the episode to another entry', async () => {
let moved: { videoId: number; animeId: number } | null = null;
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async (videoId: number, animeId: number) => {
moved = { videoId, animeId };
return { targetAnimeId: animeId, previousAnimeId: 4, removedPreviousAnime: true };
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":7}',
});
assert.equal(res.status, 200);
assert.deepEqual(moved, { videoId: 12, animeId: 7 });
assert.deepEqual(await res.json(), {
ok: true,
animeId: 7,
previousAnimeId: 4,
removedPreviousAnime: true,
});
});
it('POST /api/stats/anime/:animeId/merge reports a merge that folded nothing as 404', async () => {
const app = createStatsApp(
createMockTracker({
mergeAnime: async (targetAnimeId: number) => ({
survivingAnimeId: targetAnimeId,
mergedAnimeIds: [],
movedVideos: 0,
}),
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: '{"sourceAnimeIds":[8]}',
});
assert.equal(res.status, 404);
});
it('PATCH /api/stats/media/:videoId/anime reports an unknown target as 404', async () => {
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async () => {
throw new Error('Unknown episode or target library entry');
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":99}',
});
assert.equal(res.status, 404);
});
it('PATCH /api/stats/media/:videoId/anime does not disguise storage failures as 404', async () => {
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async () => {
throw new Error('database is locked');
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":7}',
});
assert.notEqual(res.status, 404);
assert.equal(res.status >= 500, true);
});
it('POST /api/stats/anki/browse returns 400 for missing noteId', async () => {
const app = createStatsApp(createMockTracker());
const res = await app.request('/api/stats/anki/browse', { method: 'POST' });
@@ -327,6 +327,7 @@ export function createCoverArtFetcher(
titleEnglish: selected.title?.english ?? null,
titleNative: selected.title?.native ?? null,
episodesTotal: selected.episodes ?? null,
exactTitleMatch: resolution?.exactTitleMatch ?? false,
});
logger.info(
@@ -156,6 +156,79 @@ test('season 1 resolves to the anchor without relation lookups', async () => {
assert.deepEqual(relationLookups, []);
});
test('a sequel resolution is not certified by the anchor exact-title evidence', async () => {
// The anchor matched the search title exactly, but the hopped-to entry is a
// different inference (a split-cour chain can land one season short), so the
// sequel result must report its own title evidence, not the anchor's.
const { execute } = createExecutor(OREGAIRU_SEARCH, OREGAIRU_RELATIONS);
const result = await resolveAnilistSeasonMedia(
{ title: 'My Teen Romantic Comedy SNAFU', season: 2, episode: 1 },
{ execute },
);
assert.equal(result?.id, 20698);
assert.equal(result?.via, 'sequel-chain');
assert.equal(result?.exactTitleMatch, false);
});
test('a sequel resolution whose own title matches the parsed title stays exact', async () => {
const anchor: AnilistSeasonMedia = {
id: 1,
episodes: 12,
format: 'TV',
title: { english: 'Show' },
};
const sequel: AnilistSeasonMedia = {
id: 2,
episodes: 12,
format: 'TV',
title: { english: 'Show 2nd Season' },
};
const { execute } = createExecutor([anchor], {
1: [{ relationType: 'SEQUEL', node: sequel }],
});
const result = await resolveAnilistSeasonMedia(
{ title: 'Show 2nd Season', season: 2, episode: 1 },
{ execute },
);
assert.equal(result?.id, 2);
assert.equal(result?.via, 'sequel-chain');
assert.equal(result?.exactTitleMatch, true);
});
test('reports an exact normalized synonym match as strong evidence', async () => {
const { execute } = createExecutor([
{
id: 1,
episodes: 12,
format: 'TV',
title: { english: 'Hitori Gotoh Story' },
synonyms: ['BOCCHI THE ROCK'],
},
]);
const result = await resolveAnilistSeasonMedia({ title: 'Bocchi the Rock!' }, { execute });
assert.equal(result?.exactTitleMatch, true);
});
test('reports a fuzzy-only search result as weak evidence', async () => {
const { execute } = createExecutor([
{
id: 1,
episodes: 12,
format: 'TV',
title: { english: 'Actual Show' },
},
]);
const result = await resolveAnilistSeasonMedia({ title: 'Unrelated Release' }, { execute });
assert.equal(result?.exactTitleMatch, false);
});
test('strips a season marker already present in the parsed title', async () => {
const { execute, searches } = createExecutor(OREGAIRU_SEARCH, OREGAIRU_RELATIONS);
const result = await resolveAnilistSeasonMedia(
+40 -12
View File
@@ -9,6 +9,8 @@
* reports `seasonResolved: false` so callers can refuse to act instead of guessing.
*/
import { normalizeTitleIdentity } from '../../utils/title-normalization';
export interface AnilistSeasonMediaTitle {
romaji?: string | null;
english?: string | null;
@@ -42,6 +44,8 @@ export interface AnilistSeasonResolution {
seasonResolved: boolean;
requestedSeason: number | null;
via: AnilistSeasonResolutionVia;
/** Exact normalized match against an AniList title or synonym. */
exactTitleMatch: boolean;
}
export interface ResolveAnilistSeasonMediaInput {
@@ -115,10 +119,6 @@ const SEASONAL_FORMAT_PRIORITY = ['TV', 'TV_SHORT', 'ONA'];
const MAX_SEQUEL_HOPS = 12;
function normalizeTitle(value: string): string {
return value.trim().toLowerCase().replace(/\s+/g, ' ');
}
/**
* Drops season markers a release name carries but AniList titles never do,
* so "Some Show Season 3" and "Some Show S3" both search as "Some Show".
@@ -136,7 +136,7 @@ function mediaTitles(media: AnilistSeasonMedia): string[] {
const synonyms = Array.isArray(media.synonyms) ? media.synonyms : [];
return [media.title?.english, media.title?.romaji, media.title?.native, ...synonyms]
.filter((value): value is string => typeof value === 'string' && value.trim().length > 0)
.map((value) => normalizeTitle(value));
.map((value) => normalizeTitleIdentity(value));
}
function displayTitle(media: AnilistSeasonMedia, fallback: string): string {
@@ -176,6 +176,7 @@ function toResolution(
season: number | null,
via: AnilistSeasonResolutionVia,
seasonResolved: boolean,
exactTitleMatch: boolean,
): AnilistSeasonResolution {
return {
id: media.id,
@@ -185,6 +186,7 @@ function toResolution(
seasonResolved,
requestedSeason: season,
via,
exactTitleMatch,
};
}
@@ -209,9 +211,10 @@ export function pickAnchorMedia(
: media;
const pool = episodeFiltered.length > 0 ? episodeFiltered : media;
const targets = [normalizeTitle(title), normalizeTitle(stripSeasonSuffix(title))].filter(
(value, index, all) => value.length > 0 && all.indexOf(value) === index,
);
const targets = [
normalizeTitleIdentity(title),
normalizeTitleIdentity(stripSeasonSuffix(title)),
].filter((value, index, all) => value.length > 0 && all.indexOf(value) === index);
const scored = pool.map((entry, index) => {
const candidateTitles = mediaTitles(entry);
@@ -367,9 +370,20 @@ export async function resolveAnilistSeasonMedia(
episode: season === null || season <= 1 ? input.episode : null,
});
if (!anchor) return null;
// Certifies the media actually returned, never the anchor on its behalf: a
// sequel-chain hop can land one season short (split-cour entries) while the
// anchor title still matches perfectly, and that certainty must not carry
// over to the hopped-to entry.
const exactMatchFor = (candidate: AnilistSeasonMedia): boolean => {
const titles = mediaTitles(candidate);
return (
titles.includes(normalizeTitleIdentity(searchTitle)) ||
titles.includes(normalizeTitleIdentity(input.title))
);
};
if (season === null || season <= 1) {
return toResolution(anchor, searchTitle, season, 'anchor', true);
return toResolution(anchor, searchTitle, season, 'anchor', true, exactMatchFor(anchor));
}
let chainError: unknown = null;
@@ -383,7 +397,14 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] season ${season} of "${searchTitle}" resolved via sequel chain: ${displayTitle(viaChain, searchTitle)} (${viaChain.id})`,
);
return toResolution(viaChain, searchTitle, season, 'sequel-chain', true);
return toResolution(
viaChain,
searchTitle,
season,
'sequel-chain',
true,
exactMatchFor(viaChain),
);
}
const viaAirOrder = pickByAirOrder(anchor, season, media);
@@ -391,7 +412,14 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] season ${season} of "${searchTitle}" resolved via air order: ${displayTitle(viaAirOrder, searchTitle)} (${viaAirOrder.id})`,
);
return toResolution(viaAirOrder, searchTitle, season, 'air-order', true);
return toResolution(
viaAirOrder,
searchTitle,
season,
'air-order',
true,
exactMatchFor(viaAirOrder),
);
}
// The chain failed for transport reasons rather than because the season is absent;
@@ -403,5 +431,5 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] could not resolve season ${season} of "${searchTitle}"; falling back to ${displayTitle(anchor, searchTitle)} (${anchor.id})`,
);
return toResolution(anchor, searchTitle, season, 'anchor', false);
return toResolution(anchor, searchTitle, season, 'anchor', false, exactMatchFor(anchor));
}
+195
View File
@@ -0,0 +1,195 @@
import { test } from 'node:test';
import assert from 'node:assert/strict';
import {
assOverrideSignature,
assToPlainText,
collectAssOverrideCommands,
extractAssOverrideBlocks,
hasAssTemporalOverride,
isAnimatedAssEffectKind,
isAssTemporalCommand,
normalizePlainSubtitleText,
parseAssEffectField,
} from './ass-text';
test('assToPlainText drops vector drawing runs', () => {
assert.equal(
assToPlainText(
'{\\an5\\pos(730,1042)\\p1\\blur1}m 20 0 b 10 0 0 10 0 20 b 0 31 10 40 20 40 {\\p0}',
),
'',
);
});
test('assToPlainText keeps text around drawing runs on the same event', () => {
assert.equal(
assToPlainText('{\\p1}m 0 0 l 10 10{\\p0}本文{\\p1}m 5 5 l 6 6{\\p0}続き'),
'本文続き',
);
});
test('assToPlainText leaves \\pos alone when no drawing mode is active', () => {
assert.equal(assToPlainText('{\\pos(960,1068)\\bord3}位置指定'), '位置指定');
});
test('assToPlainText does not read \\pos as a drawing tag', () => {
assert.equal(assToPlainText('{\\p1\\pos(1,2)}m 0 0 l 5 5'), '');
});
test('assToPlainText resolves line-break and space escapes', () => {
assert.equal(assToPlainText('一行目\\N二行目'), '一行目\n二行目');
assert.equal(assToPlainText('一行目\\n二行目'), '一行目\n二行目');
assert.equal(assToPlainText('一行目\\N二行目', ' '), '一行目 二行目');
assert.equal(assToPlainText('間\\h隔'), '間 隔');
});
test('assToPlainText matches mpv on brace and backslash sequences', () => {
// mpv has no `\{` / `\}` / `\\` escapes: the backslashes are literal text and the
// braces still open and close an override block.
assert.equal(assToPlainText('\\{注\\}'), '\\');
assert.equal(assToPlainText('\\\\N'), '\\\n');
});
test('assToPlainText renders an unclosed override block verbatim', () => {
// mpv shows the stray brace; guessing where the block ended can eat a whole line.
assert.equal(assToPlainText('本文{\\pos(1,2)'), '本文{\\pos(1,2)');
});
test('assToPlainText is idempotent', () => {
const samples = [
'{\\an5\\p1}m 0 0 l 5 5{\\p0}本文',
'\\{注\\}',
'\\\\N',
'本文{\\pos(1,2)',
'一行目\\N二行目\\h終わり',
];
for (const sample of samples) {
const once = assToPlainText(sample);
assert.equal(assToPlainText(once), once, sample);
}
});
test('assToPlainText normalizes CRLF before converting', () => {
assert.equal(assToPlainText('一行目\r\n二行目'), '一行目\n二行目');
});
test('normalizePlainSubtitleText settles whitespace without decoding ASS', () => {
// A brace reaching this layer is literal text mpv chose to show, not markup.
assert.equal(normalizePlainSubtitleText('本文{\\pos(1,2)'), '本文{\\pos(1,2)');
assert.equal(normalizePlainSubtitleText('一行目\\N二行目'), '一行目\n二行目');
assert.equal(
normalizePlainSubtitleText('一行目\\N二行目', { collapseLineBreaks: true }),
'一行目 二行目',
);
assert.equal(normalizePlainSubtitleText(' 余白 ', { trim: false }), ' 余白 ');
});
test('normalizePlainSubtitleText is idempotent', () => {
for (const sample of ['一行目\\N二行目', '間\\h隔', '本文{\\pos(1,2)', ' 余白 ']) {
const once = normalizePlainSubtitleText(sample);
assert.equal(normalizePlainSubtitleText(once), once, sample);
}
});
test('extractAssOverrideBlocks returns block contents', () => {
assert.deepEqual(extractAssOverrideBlocks('{\\an8}上{\\fad(200,200)}下'), [
'\\an8',
'\\fad(200,200)',
]);
assert.deepEqual(extractAssOverrideBlocks('括弧なし'), []);
});
test('collectAssOverrideCommands captures names and arguments from blocks only', () => {
const commands = collectAssOverrideCommands('{\\pos(1,2)\\1c&HFFFFFF&\\kf30}歌詞');
assert.deepEqual(commands, [
{ name: 'pos', args: '1,2', animated: false },
{ name: '1c', args: '&HFFFFFF&', animated: false },
{ name: 'kf', args: '30', animated: false },
]);
// A `\pos(...)` sitting in visible text is not typesetting markup.
assert.deepEqual(collectAssOverrideCommands('\\pos(730,1042) と書いてある'), []);
});
test('collectAssOverrideCommands marks tags animated by a wrapping \\t', () => {
const commands = collectAssOverrideCommands('{\\clip(0,0,10,10)\\t(0,500,\\frz30)}文字');
assert.deepEqual(
commands.map((command) => [command.name, command.animated]),
[
['clip', false],
['t', false],
['frz', true],
],
);
assert.equal(hasAssTemporalOverride(commands), true);
});
test('collectAssOverrideCommands stops descending into deeply nested \\t tags', () => {
// Nested far past the recursion cap. Uncapped, this recurses once per level, and a
// pathological line (real files reach one or two levels) overflows the stack.
const nesting = 32;
const block = `{${'\\t(0,500,'.repeat(nesting)}\\frz30${')'.repeat(nesting)}}文字`;
const commands = collectAssOverrideCommands(block);
// The outer `\t` plus one per allowed recursion level, and nothing from below the cap.
assert.equal(commands.length, 9);
assert.deepEqual(new Set(commands.map((command) => command.name)), new Set(['t']));
assert.equal(hasAssTemporalOverride(commands), true);
});
test('hasAssTemporalOverride ignores static placement and shape tags', () => {
assert.equal(
hasAssTemporalOverride(collectAssOverrideCommands('{\\pos(1,2)\\clip(m 1 1)\\blur2}文字')),
false,
);
assert.equal(hasAssTemporalOverride(collectAssOverrideCommands('{\\move(1,2,3,4)}文字')), true);
});
test('isAssTemporalCommand covers only intrinsically animated tags', () => {
for (const command of ['t', 'move', 'k', 'kf', 'ko', 'K']) {
assert.equal(isAssTemporalCommand(command), true, command);
}
for (const command of ['clip', 'iclip', 'frz', 'fscx', 'blur', 'be', 'pos', 'fad']) {
assert.equal(isAssTemporalCommand(command), false, command);
}
});
test('assOverrideSignature distinguishes events by their override values', () => {
const first = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 1 1)}歌詞'));
const second = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 2 2)}歌詞'));
const repeat = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 1 1)}別の行'));
assert.notEqual(first, second);
assert.equal(first, repeat);
});
test('parseAssEffectField classifies the event-level Effect column', () => {
assert.equal(parseAssEffectField(''), 'none');
assert.equal(parseAssEffectField(' '), 'none');
assert.equal(parseAssEffectField('Banner;20;1;0'), 'banner');
assert.equal(parseAssEffectField('Scroll up;0;0;30;10'), 'scroll');
assert.equal(parseAssEffectField('Scroll down;0;0;30;10'), 'scroll');
assert.equal(parseAssEffectField('Karaoke'), 'karaoke');
assert.equal(parseAssEffectField('fx-template'), 'other');
});
test('parseAssEffectField matches stock effect names exactly', () => {
// Custom effect names that merely start with a stock name are not stock effects.
assert.equal(parseAssEffectField('scrolling-credit'), 'other');
assert.equal(parseAssEffectField('bannerfx;1'), 'other');
assert.equal(parseAssEffectField('karaoke-template'), 'other');
assert.equal(parseAssEffectField('Scroll'), 'other');
});
test('isAnimatedAssEffectKind covers the stock animated effects only', () => {
assert.equal(isAnimatedAssEffectKind('karaoke'), true);
assert.equal(isAnimatedAssEffectKind('banner'), true);
assert.equal(isAnimatedAssEffectKind('scroll'), true);
// Typesetting groups put static template names in this column too.
assert.equal(isAnimatedAssEffectKind('other'), false);
assert.equal(isAnimatedAssEffectKind('none'), false);
});
+280
View File
@@ -0,0 +1,280 @@
/*
* ASS/SSA text handling, split into two deliberately distinct contracts:
*
* assToPlainText() raw ASS event text -> plain text. Ingestion only.
* normalizePlainSubtitleText() already-decoded text -> display/lookup form.
*
* Subtitle text is decoded from ASS exactly once, at the point it enters the app: the
* file cue parser does it for sidecar/embedded scripts, and mpv does it for live text
* (`sub-text` is already run through mpv's own `ass_to_plaintext`). Everything
* downstream -- renderer, timing tracker, tokenizer, tokenization cache keys -- gets
* plain text and only normalizes whitespace, so no layer decodes the same string twice.
*
* assToPlainText mirrors mpv's `ass_to_plaintext` rather than inventing its own rules,
* so a cue parsed from a file reads the same as the same line arriving live:
* - `{...}` override blocks are markup
* - `\pN ... \p0` runs are vector paths, not dialogue
* - `\N`, `\n` and `\h` are the only escapes; `\{`, `\}` and `\\` are NOT escapes,
* so `\{注\}` decodes to a lone backslash exactly as mpv renders it
* - an unclosed `{` is rendered verbatim instead of swallowing the rest of the line
* Because the decoder never emits an escape or a closed brace, running it twice is a
* no-op -- but downstream code should still use normalizePlainSubtitleText.
*/
/** What `\N` and `\n` become. */
export type AssLineBreak = '\n' | ' ';
// `\p<n>` with n > 0 switches libass into vector-drawing mode: everything until the
// next `\p0` is a path (`m 20 0 b 10 0 ...`), not dialogue. The negative lookahead keeps
// `\pos(...)` from being read as a drawing tag.
const ASS_DRAWING_SCALE_PATTERN = /\\p(?![a-zA-Z])(\d*)/g;
function readDrawingScale(block: string): number | null {
ASS_DRAWING_SCALE_PATTERN.lastIndex = 0;
let scale: number | null = null;
let match: RegExpExecArray | null;
// Drawing mode is whatever the last `\p` tag in this block set it to.
while ((match = ASS_DRAWING_SCALE_PATTERN.exec(block)) !== null) {
scale = match[1] ? Number(match[1]) : 0;
}
return scale;
}
/** Resolve `\N`, `\n` and `\h`. The only text-level escapes libass recognises. */
function resolveWhitespaceEscapes(text: string, lineBreak: AssLineBreak): string {
return text.replace(/\\([Nnh])/g, (_match, escaped: string) =>
escaped === 'h' ? ' ' : lineBreak,
);
}
/** Strip `{...}` override blocks and the drawing runs they enable. */
function stripAssMarkup(raw: string): string {
let out = '';
let cursor = 0;
let drawing = false;
while (cursor < raw.length) {
if (raw[cursor] !== '{') {
if (!drawing) {
out += raw[cursor];
}
cursor += 1;
continue;
}
const close = raw.indexOf('}', cursor + 1);
if (close === -1) {
// mpv shows an unclosed `{` and everything after it. Guessing where the block was
// meant to end can eat a whole line of dialogue.
if (!drawing) {
out += raw.slice(cursor);
}
break;
}
const scale = readDrawingScale(raw.slice(cursor, close + 1));
if (scale !== null) {
drawing = scale > 0;
}
cursor = close + 1;
}
return out;
}
/**
* Decode a raw ASS/SSA event text field. Call this once, where the text enters the app;
* downstream layers take the result as plain text.
*/
export function assToPlainText(text: string, lineBreak: AssLineBreak = '\n'): string {
if (!text) return '';
return resolveWhitespaceEscapes(stripAssMarkup(text.replace(/\r\n/g, '\n')), lineBreak);
}
export interface NormalizePlainSubtitleTextOptions {
/** Fold every line break into a single space. */
collapseLineBreaks?: boolean;
trim?: boolean;
}
/**
* Whitespace normalization for text that has already been decoded -- by mpv for live
* subtitles, by the cue parser for files. Override blocks and drawing runs are none of
* this function's business; a `{` that reaches here is literal text mpv chose to show.
*
* `\N`/`\n`/`\h` are still folded, because subtitle sources outside the ASS path (asbplayer
* and other websocket clients) forward them raw and the display layer has to cope.
*/
export function normalizePlainSubtitleText(
text: string,
options: NormalizePlainSubtitleTextOptions = {},
): string {
if (!text) return '';
const { collapseLineBreaks = false, trim = true } = options;
let normalized = resolveWhitespaceEscapes(
text.replace(/\r\n/g, '\n'),
collapseLineBreaks ? ' ' : '\n',
);
if (collapseLineBreaks) {
normalized = normalized.replace(/\n/g, ' ').replace(/\s+/g, ' ');
}
return trim ? normalized.trim() : normalized;
}
/** The contents of each `{...}` block, without the braces. */
export function extractAssOverrideBlocks(text: string): string[] {
const blocks: string[] = [];
let cursor = 0;
while (cursor < text.length) {
const open = text.indexOf('{', cursor);
if (open === -1) {
break;
}
const close = text.indexOf('}', open + 1);
if (close === -1) {
break;
}
blocks.push(text.slice(open + 1, close));
cursor = close + 1;
}
return blocks;
}
export interface AssOverrideCommand {
/** Tag name without the backslash, e.g. `pos`, `kf`, `1c`. */
name: string;
/** Everything the tag was given, e.g. `960,1068` for `\pos(960,1068)`. */
args: string;
/** Nested inside a `\t(...)` argument, so its value is animated over the event. */
animated: boolean;
}
const ASS_OVERRIDE_NAME_PATTERN = /[1-4]?[a-zA-Z]+/y;
function readCommandArgs(block: string, start: number): { args: string; next: number } {
if (block[start] === '(') {
let depth = 0;
for (let i = start; i < block.length; i += 1) {
if (block[i] === '(') depth += 1;
else if (block[i] === ')') {
depth -= 1;
if (depth === 0) {
return { args: block.slice(start + 1, i), next: i + 1 };
}
}
}
return { args: block.slice(start + 1), next: block.length };
}
const nextTag = block.indexOf('\\', start);
const end = nextTag === -1 ? block.length : nextTag;
return { args: block.slice(start, end), next: end };
}
// `\t(...)` can wrap another `\t(...)`, and nothing in the format stops an author (or a
// malformed file) from nesting them thousands deep. Real typesetting never goes past one
// or two levels, so stop recursing well before the call stack is at risk.
const MAX_ANIMATION_NESTING_DEPTH = 8;
function parseOverrideBlock(
block: string,
animated: boolean,
into: AssOverrideCommand[],
depth = 0,
): void {
let cursor = 0;
while (cursor < block.length) {
if (block[cursor] !== '\\') {
cursor += 1;
continue;
}
ASS_OVERRIDE_NAME_PATTERN.lastIndex = cursor + 1;
const nameMatch = ASS_OVERRIDE_NAME_PATTERN.exec(block);
if (!nameMatch) {
cursor += 1;
continue;
}
const name = nameMatch[0];
const { args, next } = readCommandArgs(block, cursor + 1 + name.length);
into.push({ name, args: args.trim(), animated });
// `\t(0,500,\frz30)` animates whatever it wraps, so record the inner tags too.
if (name === 't' && args.includes('\\') && depth < MAX_ANIMATION_NESTING_DEPTH) {
parseOverrideBlock(args, true, into, depth + 1);
}
cursor = next;
}
}
/**
* Override commands with their arguments, in source order. Only `{...}` blocks are
* inspected, so a `\pos(...)` sitting in visible text is never mistaken for markup.
*/
export function collectAssOverrideCommands(text: string): AssOverrideCommand[] {
const commands: AssOverrideCommand[] = [];
for (const block of extractAssOverrideBlocks(text)) {
parseOverrideBlock(block, false, commands);
}
return commands;
}
// Tags that are animated by definition: `\t` interpolates, `\move` travels, and the
// karaoke tags advance a highlight across the event's own duration. Everything else --
// `\pos`, `\clip`, `\frz`, `\blur`, `\fad` -- is a static value for the event, so its
// presence says nothing about whether neighbouring events form one animation.
const ASS_TEMPORAL_COMMANDS = new Set(['t', 'move', 'k', 'kf', 'ko', 'K']);
export function isAssTemporalCommand(name: string): boolean {
return ASS_TEMPORAL_COMMANDS.has(name);
}
/** True when the event animates on its own, or animates a static tag through `\t(...)`. */
export function hasAssTemporalOverride(commands: readonly AssOverrideCommand[]): boolean {
return commands.some((command) => command.animated || isAssTemporalCommand(command.name));
}
/**
* Canonical form of an event's override values, for comparing consecutive events. Two
* events with the same signature were typeset identically, so neither is a frame of an
* animation the other belongs to.
*/
export function assOverrideSignature(commands: readonly AssOverrideCommand[]): string {
return commands.map((command) => `${command.name}(${command.args})`).join('|');
}
export type AssEffectKind = 'none' | 'banner' | 'scroll' | 'karaoke' | 'other';
// The stock effects, matched exactly. Typesetting groups put their own template names in
// this column -- `scrolling-credit` is a static sign, not libass's `Scroll up` -- so a
// prefix match would hand out animation evidence to arbitrary custom effects.
const STOCK_ASS_EFFECTS = new Map<string, AssEffectKind>([
['banner', 'banner'],
['scroll up', 'scroll'],
['scroll down', 'scroll'],
['karaoke', 'karaoke'],
]);
/**
* The event-level `Effect` column. The stock values (`Banner;...`, `Scroll up;...`,
* `Scroll down;...`, `Karaoke`) all animate; anything else is a custom name and lands in
* `other`.
*/
export function parseAssEffectField(raw: string): AssEffectKind {
const value = raw.trim().toLowerCase();
if (!value) return 'none';
const name = value.split(';', 1)[0]!.trim();
return STOCK_ASS_EFFECTS.get(name) ?? 'other';
}
const ANIMATED_ASS_EFFECT_KINDS = new Set<AssEffectKind>(['banner', 'scroll', 'karaoke']);
export function isAnimatedAssEffectKind(kind: AssEffectKind): boolean {
return ANIMATED_ASS_EFFECT_KINDS.has(kind);
}
@@ -1414,6 +1414,353 @@ test('deleteSession ignores the currently active session and keeps new writes fl
}
});
test('deleteSession yields the main event loop while delete maintenance is pending', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const deleteGate: { release?: () => void } = {};
let deleteRunnerCalled = false;
let bufferedWritesAtDeleteStart = -1;
try {
const Ctor = await loadTrackerCtor();
const createdTracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async () => {
deleteRunnerCalled = true;
bufferedWritesAtDeleteStart = (tracker as unknown as { queue: unknown[] }).queue.length;
await new Promise<void>((resolve) => {
deleteGate.release = resolve;
});
},
},
);
tracker = createdTracker;
createdTracker.handleMediaChange('/tmp/delete-yield-first.mkv', 'Delete Yield First');
createdTracker.handleMediaChange('/tmp/delete-yield-active.mkv', 'Delete Yield Active');
const privateApi = createdTracker as unknown as {
db: DatabaseSync;
queue: unknown[];
flushNow: () => void;
};
const sessionId = (
privateApi.db
.prepare(
`SELECT session_id AS sessionId
FROM imm_sessions
WHERE ended_at_ms IS NOT NULL
ORDER BY session_id
LIMIT 1`,
)
.get() as { sessionId: number } | null
)?.sessionId;
assert.ok(sessionId);
const deletePromise = createdTracker.deleteSession(sessionId);
let timerAdvanced = false;
setTimeout(() => {
timerAdvanced = true;
}, 0);
await waitForCondition(() => deleteRunnerCalled);
assert.equal(deleteRunnerCalled, true, 'delete should be dispatched to the maintenance runner');
assert.equal(
bufferedWritesAtDeleteStart,
0,
'writes buffered before delete should flush first',
);
await waitForCondition(() => timerAdvanced);
createdTracker.recordSubtitleLine('queued during delete', 0, 1);
privateApi.flushNow();
assert.ok(privateApi.queue.length > 0, 'tracking writes should wait for delete maintenance');
assert.ok(deleteGate.release);
deleteGate.release();
await deletePromise;
} finally {
deleteGate.release?.();
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('delete maintenance flushes the entire write queue before locking writes', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const deleteGate: { release?: () => void } = {};
let queuedWritesAtDeleteStart = -1;
let writeLockedAtDeleteStart = false;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async () => {
const privateApi = tracker as unknown as {
queue: unknown[];
writeLock: { locked: boolean };
};
queuedWritesAtDeleteStart = privateApi.queue.length;
writeLockedAtDeleteStart = privateApi.writeLock.locked;
await new Promise<void>((resolve) => {
deleteGate.release = resolve;
});
},
},
);
const privateApi = tracker as unknown as {
batchSize: number;
flushNow: () => void;
queue: unknown[];
};
privateApi.batchSize = 1;
privateApi.queue.push({}, {}, {});
privateApi.flushNow = () => {
privateApi.queue.shift();
};
const deletePromise = tracker.deleteSession(101);
await waitForCondition(() => deleteGate.release !== undefined);
assert.equal(queuedWritesAtDeleteStart, 0);
assert.equal(writeLockedAtDeleteStart, true);
deleteGate.release?.();
await deletePromise;
} finally {
deleteGate.release?.();
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('delete maintenance tasks stay serialized under concurrent requests', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const releases: Array<() => void> = [];
let activeTasks = 0;
let maxActiveTasks = 0;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async () => {
activeTasks += 1;
maxActiveTasks = Math.max(maxActiveTasks, activeTasks);
await new Promise<void>((resolve) => {
releases.push(resolve);
});
activeTasks -= 1;
},
},
);
const firstDelete = tracker.deleteSession(101);
await waitForCondition(() => releases.length === 1);
assert.equal(maxActiveTasks, 1);
const secondDelete = tracker.deleteSession(102);
releases[0]?.();
await waitForCondition(() => releases.length === 2);
assert.equal(maxActiveTasks, 1);
releases[1]?.();
await Promise.all([firstDelete, secondDelete]);
} finally {
for (const release of releases) release();
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('concurrent delete requests share one maintenance worker batch', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const tasks: unknown[] = [];
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async (_path, task) => {
tasks.push(task);
},
},
);
const firstDelete = tracker.deleteSession(201);
const secondDelete = tracker.deleteSessions([202, 203]);
const thirdDelete = tracker.deleteVideo(204);
await Promise.all([firstDelete, secondDelete, thirdDelete]);
assert.equal(tasks.length, 1, 'concurrent deletes should use one maintenance pass');
assert.deepEqual(tasks[0], {
kind: 'batch',
tasks: [
{ kind: 'session', sessionId: 201 },
{ kind: 'sessions', sessionIds: [202, 203] },
{ kind: 'video', videoId: 204 },
],
});
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('destroy rejects delete requests waiting behind active maintenance', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
let releaseFirstTask: () => void = () => {};
try {
const Ctor = await loadTrackerCtor();
let markFirstTaskStarted: () => void = () => {};
const firstTaskStarted = new Promise<void>((resolve) => {
markFirstTaskStarted = resolve;
});
tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async () => {
markFirstTaskStarted();
await new Promise<void>((resolve) => {
releaseFirstTask = resolve;
});
},
},
);
const firstDelete = tracker.deleteSession(301);
await firstTaskStarted;
const queuedDelete = tracker.deleteSession(302);
tracker.destroy();
const queuedOutcome = await Promise.race([
queuedDelete.then(
() => 'resolved',
(error: unknown) =>
error instanceof Error && /shutting down/.test(error.message)
? 'rejected'
: 'wrong-error',
),
new Promise<'pending'>((resolve) => setTimeout(() => resolve('pending'), 25)),
]);
assert.equal(queuedOutcome, 'rejected');
releaseFirstTask();
await firstDelete;
} finally {
releaseFirstTask();
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('delete requested after destroy rejects without running maintenance', async () => {
const dbPath = makeDbPath();
let maintenanceCalls = 0;
const Ctor = await loadTrackerCtor();
const tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async () => {
maintenanceCalls += 1;
},
},
);
tracker.destroy();
await assert.rejects(tracker.deleteSession(303), /shutting down/);
assert.equal(maintenanceCalls, 0);
cleanupDbPath(dbPath);
});
test('deleteSessions skips maintenance when no sessions are deletable', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const tasks: unknown[] = [];
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async (_path, task) => {
tasks.push(task);
},
},
);
await tracker.deleteSessions([]);
assert.deepEqual(tasks, []);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('queued video delete is skipped when that video becomes active before dispatch', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
const tasks: Array<{ kind: string }> = [];
let releaseFirstTask: () => void = () => {};
try {
const Ctor = await loadTrackerCtor();
const createdTracker = new Ctor(
{ dbPath },
{
runDeleteMaintenanceTask: async (_path, task) => {
tasks.push(task);
if (tasks.length === 1) {
await new Promise<void>((resolve) => {
releaseFirstTask = resolve;
});
}
},
},
);
tracker = createdTracker;
createdTracker.handleMediaChange('/tmp/delete-race-target.mkv', 'Delete Race Target');
createdTracker.handleMediaChange('/tmp/delete-race-other.mkv', 'Delete Race Other');
const privateApi = createdTracker as unknown as { db: DatabaseSync };
const targetVideoId = (
privateApi.db
.prepare(`SELECT video_id AS videoId FROM imm_videos WHERE video_key LIKE '%target.mkv'`)
.get() as { videoId: number } | null
)?.videoId;
assert.ok(targetVideoId);
const firstDelete = createdTracker.deleteSession(999_001);
await waitForCondition(() => tasks.length === 1);
const queuedVideoDelete = createdTracker.deleteVideo(targetVideoId);
createdTracker.handleMediaChange('/tmp/delete-race-target.mkv', 'Delete Race Target');
releaseFirstTask();
await Promise.all([firstDelete, queuedVideoDelete]);
assert.deepEqual(
tasks.map((task) => task.kind),
['session'],
);
} finally {
releaseFirstTask();
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('deleteVideo ignores the currently active video and keeps new writes flushable', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
@@ -1785,6 +2132,78 @@ test('handleMediaChange reuses the same provisional anime row across matching fi
}
});
test('local parsing reuses a unique compatible manual assignment from the same directory', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath });
const anchorPath = '/tmp/grouped/Incorrect Name S01E01.mkv';
tracker.handleMediaChange(anchorPath, 'Episode 1');
await waitForPendingAnimeMetadata(tracker);
const privateApi = tracker as unknown as {
db: DatabaseSync;
sessionState: { videoId: number } | null;
};
const anchorVideoId = privateApi.sessionState?.videoId;
assert.ok(anchorVideoId);
tracker.handleMediaChange(null, null);
const timestamp = toDbTimestamp(trackerNowMs());
const target = privateApi.db
.prepare(
`
INSERT INTO imm_anime (
normalized_title_key,
canonical_title,
CREATED_DATE,
LAST_UPDATE_DATE
) VALUES ('correct show season 1', 'Correct Show Season 1', ?, ?)
RETURNING anime_id AS animeId
`,
)
.get(timestamp, timestamp) as { animeId: number };
await tracker.moveVideoToAnime(anchorVideoId, target.animeId);
tracker.handleMediaChange(anchorPath, 'Episode 1');
await waitForPendingAnimeMetadata(tracker);
tracker.handleMediaChange('/tmp/grouped/Another Wrong Name S01E02.mkv', 'Episode 2');
await waitForPendingAnimeMetadata(tracker);
tracker.handleMediaChange('/tmp/grouped/Another Wrong Name S02E01.mkv', 'Episode 1');
await waitForPendingAnimeMetadata(tracker);
const rows = privateApi.db
.prepare(
`
SELECT source_path AS sourcePath, anime_id AS animeId, anime_assignment_locked AS locked
FROM imm_videos
WHERE source_path LIKE '/tmp/grouped/%'
ORDER BY source_path
`,
)
.all() as Array<{ sourcePath: string; animeId: number; locked: number }>;
const assignments = new Map(rows.map((row) => [row.sourcePath, row]));
assert.deepEqual(assignments.get(anchorPath), {
sourcePath: anchorPath,
animeId: target.animeId,
locked: 1,
});
assert.deepEqual(assignments.get('/tmp/grouped/Another Wrong Name S01E02.mkv'), {
sourcePath: '/tmp/grouped/Another Wrong Name S01E02.mkv',
animeId: target.animeId,
locked: 0,
});
assert.notEqual(
assignments.get('/tmp/grouped/Another Wrong Name S02E01.mkv')?.animeId,
target.animeId,
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('handleMediaChange splits matching parsed titles across distinct seasons', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
@@ -2271,6 +2690,67 @@ test('Jellyfin playback metadata links stream videos to existing series title',
}
});
test('Jellyfin metadata refresh preserves a manual episode assignment', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath });
const metadata = {
mediaPath: 'http://jellyfin.local/Videos/item-locked/stream?api_key=token',
displayTitle: 'Parsed Show S01E01',
itemTitle: 'Episode 1',
seriesTitle: 'Parsed Show',
seasonNumber: 1,
episodeNumber: 1,
itemId: 'item-locked',
};
tracker.recordJellyfinPlaybackMetadata(metadata);
const privateApi = tracker as unknown as { db: DatabaseSync };
const video = privateApi.db.prepare('SELECT video_id AS videoId FROM imm_videos').get() as {
videoId: number;
};
const timestamp = toDbTimestamp(trackerNowMs());
const target = privateApi.db
.prepare(
`
INSERT INTO imm_anime (
normalized_title_key,
canonical_title,
CREATED_DATE,
LAST_UPDATE_DATE
) VALUES ('correct show', 'Correct Show', ?, ?)
RETURNING anime_id AS animeId
`,
)
.get(timestamp, timestamp) as { animeId: number };
await tracker.moveVideoToAnime(video.videoId, target.animeId);
tracker.recordJellyfinPlaybackMetadata(metadata);
const assignment = privateApi.db
.prepare(
`
SELECT anime_id AS animeId, anime_assignment_locked AS locked
FROM imm_videos
WHERE video_id = ?
`,
)
.get(video.videoId) as { animeId: number; locked: number };
assert.equal(assignment.animeId, target.animeId);
assert.equal(assignment.locked, 1);
const animeCount = privateApi.db.prepare('SELECT COUNT(*) AS count FROM imm_anime').get() as {
count: number;
};
assert.equal(animeCount.count, 1);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('startup repairs existing Jellyfin stream video links to metadata rows', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
@@ -3457,6 +3937,22 @@ test('reassignAnimeAnilist redistributes conflicting legacy combined row before
(1, 2000, 1000, 1000, 1, 10, 0, 0, 0, 0, 0, 0, 0, 0),
(2, 4000, 2000, 2000, 2, 20, 0, 0, 0, 0, 0, 0, 0, 0),
(3, 6000, 3000, 3000, 3, 30, 0, 0, 0, 0, 0, 0, 0, 0);
-- The per-video lifetime rows those finalized sessions would have left
-- behind; redistributing videos re-derives imm_lifetime_anime from these.
INSERT INTO imm_lifetime_media (
video_id,
total_sessions,
total_active_ms,
completed,
first_watched_ms,
last_watched_ms,
CREATED_DATE,
LAST_UPDATE_DATE
) VALUES
(1, 1, 1000, 0, '1000', '2000', 1000, 2000),
(2, 1, 2000, 0, '3000', '4000', 3000, 4000),
(3, 1, 3000, 0, '5000', '6000', 5000, 6000);
`);
await tracker.reassignAnimeAnilist(2, {
+163 -24
View File
@@ -16,6 +16,8 @@ import {
applyPragmas,
createTrackerPreparedStatements,
ensureSchema,
findManualDirectoryAnimeAssignment,
getManualAnimeAssignment,
executeQueuedWrite,
getOrCreateAnimeRecord,
getOrCreateVideoRecord,
@@ -28,6 +30,7 @@ import {
} from './immersion-tracker/storage';
import {
applySessionLifetimeSummary,
recomputeLifetimeAnimeAggregates,
reconcileStaleActiveSessions,
rebuildLifetimeSummaries as rebuildLifetimeSummaryTables,
shouldBackfillLifetimeSummaries,
@@ -83,19 +86,29 @@ import {
} from './immersion-tracker/query-library';
import {
cleanupVocabularyStats,
deleteAnime as deleteAnimeQuery,
deleteSession as deleteSessionQuery,
deleteSessions as deleteSessionsQuery,
deleteVideo as deleteVideoQuery,
getVideoDurationMs,
markVideoWatched,
upsertCoverArt,
} from './immersion-tracker/query-maintenance';
import {
DeleteMaintenanceWorkerRuntime,
type RunDeleteMaintenanceTask,
} from './immersion-tracker/delete-maintenance-worker-runtime';
import { DeleteMaintenanceScheduler } from './immersion-tracker/delete-maintenance-scheduler';
import { repairJellyfinStreamVideoLinks } from './immersion-tracker/jellyfin-link-repair';
import {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
repairLegacySeasonlessAnimeRows,
resolveAnimeAnilistConflict,
type AnimeMergeRecommendation,
} from './immersion-tracker/anime-season-repair';
import {
mergeAnimeRecords,
moveVideoToAnime as moveVideoToAnimeQuery,
type AnimeMergeSummary,
type VideoMoveSummary,
} from './immersion-tracker/anime-merge';
import {
buildVideoKey,
deriveCanonicalTitle,
@@ -182,6 +195,7 @@ const YOUTUBE_SCREENSHOT_MAX_SECONDS = 120;
const YOUTUBE_OEMBED_ENDPOINT = 'https://www.youtube.com/oembed';
const YOUTUBE_ID_PATTERN = /^[A-Za-z0-9_-]{6,}$/;
const YOUTUBE_METADATA_REFRESH_MS = 24 * 60 * 60 * 1000;
const DELETE_MAINTENANCE_BATCH_WINDOW_MS = 10;
function isValidYouTubeVideoId(value: string | null): boolean {
return Boolean(value && YOUTUBE_ID_PATTERN.test(value));
@@ -385,6 +399,8 @@ export class ImmersionTrackerService {
private readonly vacuumIntervalMs: number;
private readonly dbPath: string;
private readonly writeLock = { locked: false };
private readonly destroyDeleteMaintenanceRunner: () => void;
private readonly deleteMaintenanceScheduler: DeleteMaintenanceScheduler;
private flushTimer: ReturnType<typeof setTimeout> | null = null;
private maintenanceTimer: ReturnType<typeof setInterval> | null = null;
private flushScheduled = false;
@@ -406,9 +422,37 @@ export class ImmersionTrackerService {
| ((row: LegacyVocabularyPosRow) => Promise<LegacyVocabularyPosResolution | null>)
| undefined;
constructor(options: ImmersionTrackerOptions) {
constructor(
options: ImmersionTrackerOptions,
dependencies: {
runDeleteMaintenanceTask?: RunDeleteMaintenanceTask;
destroyDeleteMaintenanceRunner?: () => void;
} = {},
) {
this.dbPath = options.dbPath;
this.resolveLegacyVocabularyPos = options.resolveLegacyVocabularyPos;
let runDeleteMaintenanceTask: RunDeleteMaintenanceTask;
if (dependencies.runDeleteMaintenanceTask) {
runDeleteMaintenanceTask = dependencies.runDeleteMaintenanceTask;
this.destroyDeleteMaintenanceRunner =
dependencies.destroyDeleteMaintenanceRunner ?? (() => {});
} else {
const deleteMaintenanceRuntime = new DeleteMaintenanceWorkerRuntime();
runDeleteMaintenanceTask = (dbPath, task) => deleteMaintenanceRuntime.run(dbPath, task);
this.destroyDeleteMaintenanceRunner = () => deleteMaintenanceRuntime.destroy();
}
this.deleteMaintenanceScheduler = new DeleteMaintenanceScheduler({
batchWindowMs: DELETE_MAINTENANCE_BATCH_WINDOW_MS,
runTask: (task) => runDeleteMaintenanceTask(this.dbPath, task),
onBusy: () => {
this.requireWriteQueueDrained('delete maintenance');
this.writeLock.locked = true;
},
onIdle: () => {
this.writeLock.locked = false;
if (!this.isDestroyed && this.queue.length > 0) this.scheduleFlush(0);
},
});
const parentDir = path.dirname(this.dbPath);
if (!fs.existsSync(parentDir)) {
fs.mkdirSync(parentDir, { recursive: true });
@@ -485,7 +529,7 @@ export class ImmersionTrackerService {
this.logger.info(
`Repaired season-scoped stats links on startup: scanned=${seasonRepair.scanned} movedVideos=${seasonRepair.movedVideos} deletedAnimeRows=${seasonRepair.deletedAnimeRows}`,
);
rebuildLifetimeSummaryTables(this.db);
recomputeLifetimeAnimeAggregates(this.db);
}
if (shouldBackfillLifetimeSummaries(this.db)) {
const result = rebuildLifetimeSummaryTables(this.db);
@@ -512,6 +556,8 @@ export class ImmersionTrackerService {
}
this.finalizeActiveSession();
this.isDestroyed = true;
this.deleteMaintenanceScheduler.destroy();
this.destroyDeleteMaintenanceRunner();
this.db.close();
}
@@ -596,8 +642,7 @@ export class ImmersionTrackerService {
}
async rebuildLifetimeSummaries(): Promise<LifetimeRebuildSummary> {
this.flushTelemetry(true);
this.flushNow();
this.requireWriteQueueDrained('rebuilding lifetime summaries');
return rebuildLifetimeSummaryTables(this.db);
}
@@ -664,6 +709,14 @@ export class ImmersionTrackerService {
return getAnimeLibrary(this.db);
}
async getAnimeMergeRecommendations(): Promise<AnimeMergeRecommendation[]> {
return getAnimeMergeRecommendations(this.db);
}
async dismissAnimeMergeRecommendation(recommendationId: number): Promise<boolean> {
return dismissAnimeMergeRecommendation(this.db, recommendationId);
}
async getAnimeDetail(animeId: number): Promise<AnimeDetailRow | null> {
this.relinkYoutubeAnimeLibrary();
return getAnimeDetail(this.db, animeId);
@@ -709,10 +762,11 @@ export class ImmersionTrackerService {
this.logger.warn(`Ignoring delete request for active immersion session ${sessionId}`);
return;
}
deleteSessionQuery(this.db, sessionId);
await this.enqueueDeleteMaintenanceTask(() => ({ kind: 'session', sessionId }));
}
async deleteSessions(sessionIds: number[]): Promise<void> {
await this.enqueueDeleteMaintenanceTask(() => {
const activeSessionId = this.sessionState?.sessionId;
const deletableSessionIds =
activeSessionId === undefined
@@ -723,21 +777,25 @@ export class ImmersionTrackerService {
`Ignoring bulk delete request for active immersion session ${activeSessionId}`,
);
}
deleteSessionsQuery(this.db, deletableSessionIds);
if (deletableSessionIds.length === 0) return null;
return { kind: 'sessions', sessionIds: deletableSessionIds };
});
}
async deleteVideo(videoId: number): Promise<void> {
await this.enqueueDeleteMaintenanceTask(() => {
if (this.sessionState?.videoId === videoId) {
this.logger.warn(`Ignoring delete request for active immersion video ${videoId}`);
return;
return null;
}
deleteVideoQuery(this.db, videoId);
return { kind: 'video', videoId };
});
}
async deleteAnime(animeId: number): Promise<void> {
// The active video's anime link is assigned asynchronously after the title
// is parsed, so a guard reading imm_videos too early sees a null and lets
// the delete through — then the late update recreates the anime row.
await this.enqueueDeleteMaintenanceTask(async () => {
// Resolve this at dispatch time because another queued delete can leave
// enough time for playback to switch to an episode of this anime.
const pendingVideoId = this.sessionState?.videoId;
if (pendingVideoId !== undefined) {
await this.pendingAnimeMetadataUpdates.get(pendingVideoId);
@@ -750,10 +808,77 @@ export class ImmersionTrackerService {
.get(activeVideoId) as { anime_id: number | null } | null;
if (activeAnime?.anime_id === animeId) {
this.logger.warn(`Ignoring delete request for active immersion anime ${animeId}`);
return;
return null;
}
}
deleteAnimeQuery(this.db, animeId);
return { kind: 'anime', animeId };
});
}
private enqueueDeleteMaintenanceTask(
resolveTask: Parameters<DeleteMaintenanceScheduler['enqueue']>[0],
): Promise<void> {
if (this.isDestroyed) {
return Promise.reject(new Error('Immersion tracker is shutting down'));
}
return this.deleteMaintenanceScheduler.enqueue(resolveTask);
}
/**
* Fold duplicate library entries into one. Sources that hold the currently
* playing episode are fine: the videos move, nothing is deleted out from
* under the active session.
*/
async mergeAnime(targetAnimeId: number, sourceAnimeIds: number[]): Promise<AnimeMergeSummary> {
const pendingVideoId = this.sessionState?.videoId;
if (pendingVideoId !== undefined) {
await this.pendingAnimeMetadataUpdates.get(pendingVideoId);
}
// This rebuilds the lifetime summaries, which recompute from the database:
// queued writes have to land first or the active session is dropped from
// the merged totals.
this.requireWriteQueueDrained('merging library entries');
return mergeAnimeRecords(this.db, targetAnimeId, sourceAnimeIds);
}
async moveVideoToAnime(videoId: number, targetAnimeId: number): Promise<VideoMoveSummary> {
await this.pendingAnimeMetadataUpdates.get(videoId);
this.requireWriteQueueDrained('moving an episode');
return moveVideoToAnimeQuery(this.db, videoId, targetAnimeId);
}
/**
* Persist every queued write before a caller recomputes summaries from the
* database.
*
* A single `flushNow()` is not enough: forced telemetry is appended to the
* back of the queue while `flushNow()` writes at most `batchSize` entries off
* the front, so a busy session leaves the newest sample unwritten. Stops as
* soon as a pass makes no progress a rolled-back batch is pushed back onto
* the queue, and looping on that would spin forever.
*
* Returns false when the queue could not be emptied. Summary-rebuilding
* callers fail closed in that case.
*/
private drainWriteQueue(context: string): boolean {
this.flushTelemetry(true);
while (this.queue.length > 0) {
const pending = this.queue.length;
this.flushNow();
if (this.queue.length >= pending) {
this.logger.warn(
`Immersion tracker queue did not drain before ${context}; summaries may lag by ${this.queue.length} writes`,
);
return false;
}
}
return true;
}
private requireWriteQueueDrained(context: string): void {
if (!this.drainWriteQueue(context)) {
throw new Error(`Immersion tracker queue did not drain before ${context}`);
}
}
async reassignAnimeAnilist(
@@ -768,7 +893,14 @@ export class ImmersionTrackerService {
coverUrl?: string | null;
},
): Promise<void> {
const repair = resolveAnimeAnilistConflict(this.db, animeId, info.anilistId);
this.requireWriteQueueDrained('reassigning an AniList entry');
// The user is acting on this entry, so it is the one that survives when
// another row already claims the same AniList id.
const repair = resolveAnimeAnilistConflict(this.db, animeId, info.anilistId, {
survivor: 'target',
matchConfidence: 'manual',
});
if (repair.anilistAssignmentBlocked) return;
this.db
.prepare(
`
@@ -795,7 +927,7 @@ export class ImmersionTrackerService {
animeId,
);
if (repair.movedVideos > 0 || repair.deletedAnimeRows > 0) {
rebuildLifetimeSummaryTables(this.db);
recomputeLifetimeAnimeAggregates(this.db);
}
// Update cover art for all videos in this anime
@@ -1243,7 +1375,7 @@ export class ImmersionTrackerService {
metadataJson: candidate.metadataJson,
});
}
rebuildLifetimeSummaryTables(this.db);
recomputeLifetimeAnimeAggregates(this.db);
}
recordJellyfinPlaybackMetadata(metadata: JellyfinPlaybackMetadataInput): void {
@@ -1291,7 +1423,9 @@ export class ImmersionTrackerService {
seasonNumber,
episodeNumber,
});
const animeId = getOrCreateAnimeRecord(this.db, {
const animeId =
getManualAnimeAssignment(this.db, videoId) ??
getOrCreateAnimeRecord(this.db, {
parsedTitle: libraryTitle,
canonicalTitle: libraryTitle,
seasonScope: seasonNumber,
@@ -1316,7 +1450,7 @@ export class ImmersionTrackerService {
this.db.prepare('SELECT 1 FROM imm_lifetime_media WHERE video_id = ?').get(videoId),
);
if (hasLifetimeMedia || (previousLink && previousLink.animeId !== animeId)) {
rebuildLifetimeSummaryTables(this.db);
recomputeLifetimeAnimeAggregates(this.db);
}
}
@@ -1811,7 +1945,7 @@ export class ImmersionTrackerService {
}
private runMaintenance(): void {
if (this.isDestroyed) return;
if (this.isDestroyed || this.writeLock.locked) return;
try {
this.flushTelemetry(true);
this.flushNow();
@@ -1937,7 +2071,12 @@ export class ImmersionTrackerService {
return;
}
const animeId = getOrCreateAnimeRecord(this.db, {
const animeId =
getManualAnimeAssignment(this.db, videoId) ??
(mediaPath && !isRemoteSource(mediaPath)
? findManualDirectoryAnimeAssignment(this.db, videoId, mediaPath, parsed.parsedSeason)
: null) ??
getOrCreateAnimeRecord(this.db, {
parsedTitle: parsed.parsedTitle,
canonicalTitle: parsed.parsedTitle,
seasonScope: parsed.parsedSeason,
@@ -0,0 +1,896 @@
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import test from 'node:test';
import { Database } from '../sqlite.js';
import type { DatabaseSync } from '../sqlite.js';
import {
applyPragmas,
ensureSchema,
findManualDirectoryAnimeAssignment,
getManualAnimeAssignment,
getOrCreateAnimeRecord,
linkVideoToAnimeRecord,
} from '../storage.js';
import { mergeAnimeRecords, moveVideoToAnime } from '../anime-merge.js';
import {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
resolveAnimeAnilistConflict,
} from '../anime-season-repair.js';
import { updateAnimeAnilistInfo } from '../query-maintenance.js';
const BASE_MS = 1_700_000_000_000;
function makeDbPath(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-anime-merge-test-'));
return path.join(dir, 'immersion.sqlite');
}
function cleanupDbPath(dbPath: string): void {
const dir = path.dirname(dbPath);
if (!fs.existsSync(dir)) return;
fs.rmSync(dir, { recursive: true, force: true });
}
function withDb(work: (db: DatabaseSync) => void): void {
const dbPath = makeDbPath();
const db = new Database(dbPath);
try {
applyPragmas(db);
ensureSchema(db);
work(db);
} finally {
db.close();
cleanupDbPath(dbPath);
}
}
interface AnimeSeed {
animeId: number;
key: string;
title: string;
anilistId?: number | null;
titleRomaji?: string | null;
}
function insertAnime(db: DatabaseSync, seed: AnimeSeed): void {
db.prepare(
`INSERT INTO imm_anime(anime_id, normalized_title_key, canonical_title, anilist_id, title_romaji, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, ?, ?, ?)`,
).run(
seed.animeId,
seed.key,
seed.title,
seed.anilistId ?? null,
seed.titleRomaji ?? null,
BASE_MS,
BASE_MS,
);
}
interface EpisodeSeed {
videoId: number;
animeId: number;
season?: number | null;
episode?: number;
activeMs?: number;
cards?: number;
}
/**
* One episode with one ended session, plus the imm_lifetime_media row the
* session would have left behind, so lifetime aggregates have something to sum.
*/
function insertEpisode(db: DatabaseSync, seed: EpisodeSeed): void {
const activeMs = seed.activeMs ?? 1000;
const cards = seed.cards ?? 1;
db.prepare(
`INSERT INTO imm_videos(video_id, video_key, anime_id, canonical_title, source_type, parsed_title, parsed_season, parsed_episode, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, 1, 'Show', ?, ?, 1, 1440000, ?, ?)`,
).run(
seed.videoId,
`local:/tmp/show-${seed.videoId}.mkv`,
seed.animeId,
`Show ${seed.videoId}`,
seed.season ?? null,
seed.episode ?? seed.videoId,
BASE_MS,
BASE_MS,
);
db.prepare(
`INSERT INTO imm_sessions(session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, active_watched_ms, cards_mined, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, ?, 2, ?, ?, ?, ?)`,
).run(
seed.videoId,
`session-${seed.videoId}`,
seed.videoId,
String(BASE_MS),
String(BASE_MS + activeMs),
activeMs,
cards,
BASE_MS,
BASE_MS,
);
db.prepare(
`INSERT INTO imm_subtitle_lines(session_id, video_id, anime_id, line_index, text, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, 1, ?, ?, ?)`,
).run(seed.videoId, seed.videoId, seed.animeId, `line ${seed.videoId}`, BASE_MS, BASE_MS);
db.prepare(
`INSERT INTO imm_lifetime_media(video_id, total_sessions, total_active_ms, total_cards, completed, first_watched_ms, last_watched_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, 1, ?, ?, 1, ?, ?, ?, ?)`,
).run(
seed.videoId,
activeMs,
cards,
String(BASE_MS),
String(BASE_MS + activeMs),
BASE_MS,
BASE_MS,
);
}
function animeIds(db: DatabaseSync): number[] {
return (
db.prepare('SELECT anime_id AS id FROM imm_anime ORDER BY anime_id').all() as Array<{
id: number;
}>
).map((row) => row.id);
}
function videoAnimeId(db: DatabaseSync, videoId: number): number | null {
return (
db.prepare('SELECT anime_id AS id FROM imm_videos WHERE video_id = ?').get(videoId) as {
id: number | null;
}
).id;
}
function assignmentLocked(db: DatabaseSync, videoId: number): number {
return (
db
.prepare('SELECT anime_assignment_locked AS locked FROM imm_videos WHERE video_id = ?')
.get(videoId) as { locked: number }
).locked;
}
function lineAnimeIds(db: DatabaseSync, animeId: number): number {
return Number(
(
db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id = ?')
.get(animeId) as { total: number }
).total,
);
}
test('mergeAnimeRecords folds episodes, lines and lifetime totals into the target', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1', anilistId: 555 });
insertEpisode(db, { videoId: 1, animeId: 1, activeMs: 1000, cards: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1, activeMs: 2000, cards: 3 });
const summary = mergeAnimeRecords(db, 1, [2]);
assert.equal(summary.survivingAnimeId, 1);
assert.deepEqual(summary.mergedAnimeIds, [2]);
assert.equal(summary.movedVideos, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.equal(lineAnimeIds(db, 1), 2);
const lifetime = db
.prepare(
'SELECT total_active_ms AS activeMs, total_cards AS cards, episodes_started AS episodes FROM imm_lifetime_anime WHERE anime_id = 1',
)
.get() as { activeMs: number; cards: number; episodes: number };
assert.equal(lifetime.activeMs, 3000);
assert.equal(lifetime.cards, 4);
assert.equal(lifetime.episodes, 2);
});
});
test('merge and move preserve lifetime history whose raw sessions were pruned', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertAnime(db, { animeId: 3, key: 'other show', title: 'Other Show' });
insertEpisode(db, { videoId: 1, animeId: 1, activeMs: 1000, cards: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1, activeMs: 2000, cards: 3 });
insertEpisode(db, { videoId: 3, animeId: 3, activeMs: 4000, cards: 5 });
// Retention pruned every raw session; only the lifetime summaries remain.
db.exec('DELETE FROM imm_sessions');
db.prepare(
`UPDATE imm_lifetime_global
SET total_sessions = 200, total_active_ms = 360000000, total_cards = 500, active_days = 90
WHERE global_id = 1`,
).run();
mergeAnimeRecords(db, 1, [2]);
moveVideoToAnime(db, 3, 1);
const globalRow = db
.prepare(
`SELECT total_sessions AS sessions, total_active_ms AS activeMs, total_cards AS cards, active_days AS days
FROM imm_lifetime_global WHERE global_id = 1`,
)
.get() as { sessions: number; activeMs: number; cards: number; days: number };
assert.equal(globalRow.sessions, 200);
assert.equal(globalRow.activeMs, 360000000);
assert.equal(globalRow.cards, 500);
assert.equal(globalRow.days, 90);
const survivor = db
.prepare(
`SELECT total_active_ms AS activeMs, total_cards AS cards, episodes_started AS episodes
FROM imm_lifetime_anime WHERE anime_id = 1`,
)
.get() as { activeMs: number; cards: number; episodes: number };
assert.equal(survivor.activeMs, 7000);
assert.equal(survivor.cards, 9);
assert.equal(survivor.episodes, 3);
assert.equal(
db.prepare('SELECT 1 FROM imm_lifetime_anime WHERE anime_id = 3').get(),
undefined,
);
});
});
test('mergeAnimeRecords repoints subtitle lines recorded before the anime link landed', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
// Lines are written with the video's anime_id at the time, which is NULL
// until the async title parse assigns one.
db.prepare(
`INSERT INTO imm_subtitle_lines(session_id, video_id, anime_id, line_index, text, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (2, 2, NULL, 2, 'unlinked line', ?, ?)`,
).run(BASE_MS, BASE_MS);
mergeAnimeRecords(db, 1, [2]);
assert.equal(lineAnimeIds(db, 1), 3);
const orphaned = Number(
(
db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id IS NULL')
.get() as { total: number }
).total,
);
assert.equal(orphaned, 0);
});
});
test('mergeAnimeRecords inherits metadata the target is missing without clobbering its own', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', titleRomaji: 'Shou' });
insertAnime(db, {
animeId: 2,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 555,
titleRomaji: 'Show Romaji',
});
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
mergeAnimeRecords(db, 1, [2]);
const row = db
.prepare(
'SELECT canonical_title AS title, anilist_id AS anilistId, title_romaji AS romaji FROM imm_anime WHERE anime_id = 1',
)
.get() as { title: string; anilistId: number | null; romaji: string | null };
assert.equal(row.title, 'Show');
// anilist_id is UNIQUE, so inheriting it proves the source row was gone first.
assert.equal(row.anilistId, 555);
assert.equal(row.romaji, 'Shou');
});
});
test('mergeAnimeRecords preserves source title identities as aliases of the survivor', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
db.prepare(
`INSERT INTO imm_anime_title_aliases(normalized_title_key, anime_id, CREATED_DATE, LAST_UPDATE_DATE)
VALUES ('show s01', 2, ?, ?)`,
).run(BASE_MS, BASE_MS);
mergeAnimeRecords(db, 1, [2]);
const fromSourceTitle = getOrCreateAnimeRecord(db, {
parsedTitle: 'Show Season 1',
canonicalTitle: 'Show Season 1',
seasonScope: 1,
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
const fromTransferredAlias = getOrCreateAnimeRecord(db, {
parsedTitle: 'Show S01',
canonicalTitle: 'Show S01',
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
assert.equal(fromSourceTitle, 1);
assert.equal(fromTransferredAlias, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(
(
db.prepare('SELECT canonical_title AS title FROM imm_anime WHERE anime_id = 1').get() as {
title: string;
}
).title,
'Show',
);
});
});
test('mergeAnimeRecords ignores unknown targets and self-merges', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
assert.deepEqual(mergeAnimeRecords(db, 99, [1]).mergedAnimeIds, []);
assert.deepEqual(mergeAnimeRecords(db, 1, [1]).mergedAnimeIds, []);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 1), 1);
});
});
test('moveVideoToAnime moves one episode and prunes the emptied entry', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'stray', title: 'Stray Episode Title', anilistId: 777 });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, activeMs: 5000, cards: 2 });
const summary = moveVideoToAnime(db, 2, 1);
assert.equal(summary.targetAnimeId, 1);
assert.equal(summary.previousAnimeId, 2);
assert.equal(summary.removedPreviousAnime, true);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.equal(assignmentLocked(db, 2), 1);
assert.equal(getManualAnimeAssignment(db, 2), 1);
assert.equal(lineAnimeIds(db, 1), 2);
const lifetime = db
.prepare('SELECT total_active_ms AS activeMs FROM imm_lifetime_anime WHERE anime_id = 1')
.get() as { activeMs: number };
assert.equal(lifetime.activeMs, 6000);
// The stray entry's AniList link is dropped, not inherited: a move makes no
// claim that the two entries are the same show.
const target = db
.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 1')
.get() as { anilistId: number | null };
assert.equal(target.anilistId, null);
});
});
test('moveVideoToAnime is a no-op when the episode is already in the target entry', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
const summary = moveVideoToAnime(db, 1, 1);
assert.equal(summary.targetAnimeId, 1);
assert.equal(summary.previousAnimeId, 1);
assert.equal(summary.removedPreviousAnime, false);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(assignmentLocked(db, 1), 1);
});
});
test('automatic metadata cannot overwrite a manual episode assignment', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'stray', title: 'Stray' });
insertAnime(db, { animeId: 3, key: 'parser result', title: 'Parser Result' });
insertEpisode(db, { videoId: 1, animeId: 2, season: 1 });
moveVideoToAnime(db, 1, 1);
linkVideoToAnimeRecord(db, 1, {
animeId: 3,
parsedBasename: 'Parser Result S01E01.mkv',
parsedTitle: 'Parser Result',
parsedSeason: 1,
parsedEpisode: 1,
parserSource: 'guessit',
parserConfidence: 1,
parseMetadataJson: null,
});
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(getManualAnimeAssignment(db, 1), 1);
const parsedTitle = db
.prepare('SELECT parsed_title AS parsedTitle FROM imm_videos WHERE video_id = 1')
.get() as { parsedTitle: string | null };
assert.equal(parsedTitle.parsedTitle, 'Parser Result');
});
});
test('directory grouping requires one season-compatible manual destination', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'stray', title: 'Stray' });
insertAnime(db, { animeId: 3, key: 'other', title: 'Other' });
insertEpisode(db, { videoId: 1, animeId: 2, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 3, season: 1 });
insertEpisode(db, { videoId: 3, animeId: 3, season: 1 });
db.prepare('UPDATE imm_videos SET source_path = ? WHERE video_id = ?').run(
'/library/show/Show S01E01.mkv',
1,
);
db.prepare('UPDATE imm_videos SET source_path = ? WHERE video_id = ?').run(
'/library/show/Stray S01E02.mkv',
2,
);
db.prepare('UPDATE imm_videos SET source_path = ? WHERE video_id = ?').run(
'/library/show/Other S01E03.mkv',
3,
);
moveVideoToAnime(db, 1, 1);
assert.equal(findManualDirectoryAnimeAssignment(db, 2, '/library/show/Stray S01E02.mkv', 1), 1);
assert.equal(
findManualDirectoryAnimeAssignment(db, 2, '/library/show/Stray S02E02.mkv', 2),
null,
);
assert.equal(
findManualDirectoryAnimeAssignment(db, 2, '/library/other/Stray S01E02.mkv', 1),
null,
);
moveVideoToAnime(db, 3, 3);
assert.equal(
findManualDirectoryAnimeAssignment(db, 2, '/library/show/Stray S01E02.mkv', 1),
null,
);
});
});
test('moveVideoToAnime keeps the source entry when other episodes remain', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'other', title: 'Other' });
insertEpisode(db, { videoId: 1, animeId: 2 });
insertEpisode(db, { videoId: 2, animeId: 2 });
const summary = moveVideoToAnime(db, 2, 1);
assert.equal(summary.removedPreviousAnime, false);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 1), 2);
assert.equal(videoAnimeId(db, 2), 1);
});
});
test('moveVideoToAnime rejects unknown episodes and targets', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
assert.throws(() => moveVideoToAnime(db, 99, 1));
assert.throws(() => moveVideoToAnime(db, 1, 99));
assert.equal(videoAnimeId(db, 1), 1);
});
});
test('resolveAnimeAnilistConflict folds a seasonless duplicate into the entry that owns the id', () => {
withDb((db) => {
// Same show, split because one release tagged S01 and the other did not.
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.survivingAnimeId, 1);
assert.equal(summary.movedVideos, 1);
assert.equal(summary.deletedAnimeRows, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
});
});
test('resolveAnimeAnilistConflict recommends a weak title collision instead of merging it', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.repaired, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('automatic AniList update leaves a weak collision unassigned for user review', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Actual Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
exactTitleMatch: false,
});
const target = db
.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 2')
.get() as {
anilistId: number | null;
};
assert.equal(target.anilistId, null);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('dismissed weak collision stays dismissed when automatic resolution repeats', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(dismissAnimeMergeRecommendation(db, 1), true);
resolveAnimeAnilistConflict(db, 2, 163132);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('dismissed recommendation prevents a later exact automatic merge of the pair', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'weak' });
assert.equal(dismissAnimeMergeRecommendation(db, 1), true);
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'exact' });
assert.equal(summary.repaired, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('manual merge clears recommendations involving the absorbed entry', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
mergeAnimeRecords(db, 1, [2]);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('resolveAnimeAnilistConflict keeps the target entry when the user drove the change', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { survivor: 'target' });
assert.equal(summary.survivingAnimeId, 2);
assert.deepEqual(animeIds(db), [2]);
assert.equal(videoAnimeId(db, 1), 2);
const row = db.prepare('SELECT anilist_id AS id FROM imm_anime WHERE anime_id = 2').get() as {
id: number | null;
};
assert.equal(row.id, 163132);
});
});
test('resolveAnimeAnilistConflict falls back to season redistribution for multi-season rows', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 1, season: 2 });
insertEpisode(db, { videoId: 3, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
// The mixed row is split by season instead of being poured onto one card.
const titles = (
db.prepare('SELECT canonical_title AS title FROM imm_anime ORDER BY title').all() as Array<{
title: string;
}>
).map((row) => row.title);
assert.deepEqual(titles, ['Show Season 1', 'Show Season 2']);
assert.equal(videoAnimeId(db, 1), 2);
assert.equal(videoAnimeId(db, 3), 2);
assert.notEqual(videoAnimeId(db, 2), 2);
});
});
test('season redistribution leaves manually assigned episodes in place', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 1, season: 2 });
insertEpisode(db, { videoId: 3, animeId: 2, season: 1 });
moveVideoToAnime(db, 1, 1);
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(assignmentLocked(db, 1), 1);
assert.notEqual(videoAnimeId(db, 2), 1);
assert.equal(summary.movedVideos, 1);
});
});
test('resolveAnimeAnilistConflict leaves explicit incompatible seasons and assignments unchanged', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'exact' });
assert.equal(summary.repaired, 0);
assert.equal(summary.movedVideos, 0);
assert.equal(summary.deletedAnimeRows, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(videoAnimeId(db, 2), 2);
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: 163132 },
{ animeId: 2, anilistId: null },
]);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('manual AniList resolution reassigns across explicit seasons without merging them', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { survivor: 'target' });
assert.equal(summary.anilistAssignmentBlocked, false);
assert.deepEqual(animeIds(db), [1, 2]);
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: null },
{ animeId: 2, anilistId: 163132 },
]);
});
});
test('automatic AniList update does not transfer an assignment across explicit seasons', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
exactTitleMatch: true,
});
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: 163132 },
{ animeId: 2, anilistId: null },
]);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(videoAnimeId(db, 2), 2);
});
});
test('automatic AniList update with unknown match confidence validates stored titles', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Actual Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
});
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('stored AniList titles ignore season suffixes when validating an automatic merge', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'legacy show',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show Season 1',
});
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.deletedAnimeRows, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('resolveAnimeAnilistConflict leaves an entry that already links elsewhere alone', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show s2', title: 'Show Season 2', anilistId: 999 });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(videoAnimeId(db, 2), 2);
assert.ok(animeIds(db).includes(2));
assert.equal(
(
db.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 2').get() as {
anilistId: number;
}
).anilistId,
999,
);
assert.equal(summary.repaired, 0);
assert.equal(summary.movedVideos, 0);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('automatic AniList update onto an entry that already links elsewhere does not throw', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show s2', title: 'Show Season 2', anilistId: 999 });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
// Entry 2 explicitly links to 999; a later video re-resolving to entry 1's
// id must be refused, not written over the UNIQUE anilist_id column.
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
exactTitleMatch: true,
});
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(
(
db.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 2').get() as {
anilistId: number;
}
).anilistId,
999,
);
});
});
@@ -50,6 +50,7 @@ import {
updateAnimeAnilistInfo,
upsertCoverArt,
} from '../query-maintenance.js';
import { deleteMaintenanceBatch } from '../query-delete-maintenance.js';
import { getLocalEpochDay } from '../query-shared.js';
import { EVENT_CARD_MINED, EVENT_SUBTITLE_LINE, SOURCE_TYPE_LOCAL } from '../types.js';
@@ -985,3 +986,197 @@ test('split maintenance helpers delete multiple sessions and whole videos with d
cleanupDbPath(dbPath);
}
});
test('delete maintenance batch preserves retained data across overlapping session, video, and anime targets', () => {
const { db, dbPath, stmts } = createDb();
try {
const retainedAnimeId = getOrCreateAnimeRecord(db, {
parsedTitle: 'Retained Anime',
canonicalTitle: 'Retained Anime',
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
const deletedAnimeId = getOrCreateAnimeRecord(db, {
parsedTitle: 'Deleted Anime',
canonicalTitle: 'Deleted Anime',
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
const retainedVideoId = getOrCreateVideoRecord(db, 'local:/tmp/batch-retain.mkv', {
canonicalTitle: 'Batch Retain',
sourcePath: '/tmp/batch-retain.mkv',
sourceUrl: null,
sourceType: SOURCE_TYPE_LOCAL,
});
const deletedVideoId = getOrCreateVideoRecord(db, 'local:/tmp/batch-video.mkv', {
canonicalTitle: 'Batch Video',
sourcePath: '/tmp/batch-video.mkv',
sourceUrl: null,
sourceType: SOURCE_TYPE_LOCAL,
});
const animeVideoId = getOrCreateVideoRecord(db, 'local:/tmp/batch-anime.mkv', {
canonicalTitle: 'Batch Anime',
sourcePath: '/tmp/batch-anime.mkv',
sourceUrl: null,
sourceType: SOURCE_TYPE_LOCAL,
});
for (const [videoId, animeId, episode] of [
[retainedVideoId, retainedAnimeId, 1],
[deletedVideoId, retainedAnimeId, 2],
[animeVideoId, deletedAnimeId, 1],
] as const) {
linkVideoToAnimeRecord(db, videoId, {
animeId,
parsedBasename: `batch-${episode}.mkv`,
parsedTitle: animeId === retainedAnimeId ? 'Retained Anime' : 'Deleted Anime',
parsedSeason: 1,
parsedEpisode: episode,
parserSource: 'test',
parserConfidence: 1,
parseMetadataJson: null,
});
}
const startedAtMs = 1_700_000_000_000;
const deletedSessionId = startSessionRecord(db, retainedVideoId, startedAtMs).sessionId;
const retainedSessionId = startSessionRecord(
db,
retainedVideoId,
startedAtMs + 1_000,
).sessionId;
const videoSessionId = startSessionRecord(db, deletedVideoId, startedAtMs + 2_000).sessionId;
const animeSessionId = startSessionRecord(db, animeVideoId, startedAtMs + 3_000).sessionId;
for (const [sessionId, sessionStartedAtMs] of [
[deletedSessionId, startedAtMs],
[retainedSessionId, startedAtMs + 1_000],
[videoSessionId, startedAtMs + 2_000],
[animeSessionId, startedAtMs + 3_000],
] as const) {
finalizeSessionMetrics(db, sessionId, sessionStartedAtMs);
}
for (const [index, sessionId, videoId, animeId] of [
[1, deletedSessionId, retainedVideoId, retainedAnimeId],
[2, retainedSessionId, retainedVideoId, retainedAnimeId],
[3, videoSessionId, deletedVideoId, retainedAnimeId],
[4, animeSessionId, animeVideoId, deletedAnimeId],
] as const) {
insertWordOccurrence(db, stmts, {
sessionId,
videoId,
animeId,
lineIndex: index,
text: '猫日',
word: { headword: '猫', word: '猫', reading: 'ねこ' },
});
insertKanjiOccurrence(db, stmts, {
sessionId,
videoId,
animeId,
lineIndex: index + 10,
text: '猫日',
kanji: '日',
});
}
const rollupDay = getLocalEpochDay(db, startedAtMs);
const rollupMonth = (
db
.prepare(
`SELECT CAST(strftime('%Y%m', CAST(? AS REAL) / 1000, 'unixepoch', 'localtime') AS INTEGER) AS rollupMonth`,
)
.get(startedAtMs) as { rollupMonth: number }
).rollupMonth;
for (const videoId of [retainedVideoId, deletedVideoId, animeVideoId]) {
db.prepare(
`INSERT INTO imm_daily_rollups (
rollup_day, video_id, total_sessions, total_active_min, total_lines_seen,
total_tokens_seen, total_cards, CREATED_DATE, LAST_UPDATE_DATE
) VALUES (?, ?, 99, 99, 99, 99, 99, ?, ?)`,
).run(rollupDay, videoId, startedAtMs, startedAtMs);
db.prepare(
`INSERT INTO imm_monthly_rollups (
rollup_month, video_id, total_sessions, total_active_min, total_lines_seen,
total_tokens_seen, total_cards, CREATED_DATE, LAST_UPDATE_DATE
) VALUES (?, ?, 99, 99, 99, 99, 99, ?, ?)`,
).run(rollupMonth, videoId, startedAtMs, startedAtMs);
}
deleteMaintenanceBatch(db, [
{ kind: 'session', sessionId: deletedSessionId },
{ kind: 'session', sessionId: videoSessionId },
{ kind: 'video', videoId: deletedVideoId },
{ kind: 'video', videoId: animeVideoId },
{ kind: 'anime', animeId: deletedAnimeId },
]);
assert.deepEqual(db.prepare('SELECT session_id FROM imm_sessions').all(), [
{ session_id: retainedSessionId },
]);
assert.deepEqual(db.prepare('SELECT video_id FROM imm_videos').all(), [
{ video_id: retainedVideoId },
]);
assert.deepEqual(db.prepare('SELECT anime_id FROM imm_anime').all(), [
{ anime_id: retainedAnimeId },
]);
assert.equal(
(
db.prepare(`SELECT frequency FROM imm_words WHERE headword = '猫'`).get() as {
frequency: number;
}
).frequency,
1,
);
assert.equal(
(
db.prepare(`SELECT frequency FROM imm_kanji WHERE kanji = '日'`).get() as {
frequency: number;
}
).frequency,
1,
);
assert.deepEqual(
db.prepare('SELECT video_id, total_sessions FROM imm_daily_rollups').all() as Array<{
video_id: number;
total_sessions: number;
}>,
[{ video_id: retainedVideoId, total_sessions: 1 }],
);
assert.deepEqual(
db.prepare('SELECT video_id, total_sessions FROM imm_monthly_rollups').all() as Array<{
video_id: number;
total_sessions: number;
}>,
[{ video_id: retainedVideoId, total_sessions: 1 }],
);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('delete maintenance batch chunks id lists below the SQLite variable limit', () => {
const { db, dbPath } = createDb();
try {
const ids = Array.from({ length: 32_767 }, (_, index) => index + 1);
assert.doesNotThrow(() => {
deleteMaintenanceBatch(db, [
{ kind: 'sessions', sessionIds: ids },
...ids.map((videoId) => ({ kind: 'video' as const, videoId })),
...ids.map((animeId) => ({ kind: 'anime' as const, animeId })),
]);
});
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
@@ -0,0 +1,169 @@
import type { DatabaseSync } from './sqlite';
import { animeSeasonsAreMergeCompatible, getParsedSeasonsForAnime } from './anime-merge';
import { toDbTimestamp } from './query-shared';
import { normalizeAnimeIdentityKey } from './storage';
import { nowMs } from './time';
export interface AnimeMergeRecommendation {
recommendationId: number;
animeIds: [number, number];
}
export interface AnimeConflictRecommendationOptions {
survivor?: 'target' | 'existing';
/** Automatic matches must be exact; manual assignment is authoritative. */
matchConfidence?: 'exact' | 'weak' | 'manual';
}
interface AnimeTitleRow {
canonical_title: string;
title_romaji: string | null;
title_english: string | null;
title_native: string | null;
}
function getAnimeTitles(db: DatabaseSync, animeId: number): AnimeTitleRow | null {
return db
.prepare(
`SELECT canonical_title, title_romaji, title_english, title_native
FROM imm_anime
WHERE anime_id = ?`,
)
.get(animeId) as AnimeTitleRow | null;
}
function getParsedTitles(db: DatabaseSync, animeId: number): Array<string | null> {
return (
db.prepare('SELECT parsed_title FROM imm_videos WHERE anime_id = ?').all(animeId) as Array<{
parsed_title: string | null;
}>
).map((row) => row.parsed_title);
}
function stripSeasonIdentitySuffix(title: string): string {
return title
.replace(/\bseason\s*\d{1,2}\b/gi, ' ')
.replace(/\b\d{1,2}(?:st|nd|rd|th)\s+season\b/gi, ' ')
.replace(/\bs\d{1,2}\b/gi, ' ');
}
export function hasExactStoredTitleMatch(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
): boolean {
const target = getAnimeTitles(db, targetAnimeId);
const conflict = getAnimeTitles(db, conflictAnimeId);
if (!target || !conflict) return false;
const targetKeys = [target.canonical_title, ...getParsedTitles(db, targetAnimeId)]
.filter((title): title is string => Boolean(title?.trim()))
.map((title) => normalizeAnimeIdentityKey(stripSeasonIdentitySuffix(title)))
.filter(Boolean);
const anilistTitleKeys = [
conflict.title_romaji,
conflict.title_english,
conflict.title_native,
conflict.canonical_title,
]
.filter((title): title is string => Boolean(title?.trim()))
.map((title) => normalizeAnimeIdentityKey(stripSeasonIdentitySuffix(title)))
.filter(Boolean);
return targetKeys.some((key) => anilistTitleKeys.includes(key));
}
export function shouldRecommendAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
options: AnimeConflictRecommendationOptions,
): boolean {
if (options.survivor === 'target' || options.matchConfidence === 'manual') return false;
if (
!animeSeasonsAreMergeCompatible(
getParsedSeasonsForAnime(db, targetAnimeId),
getParsedSeasonsForAnime(db, conflictAnimeId),
)
) {
return false;
}
return (
options.matchConfidence === 'weak' ||
(options.matchConfidence === undefined &&
!hasExactStoredTitleMatch(db, targetAnimeId, conflictAnimeId))
);
}
export function recordAnimeMergeRecommendation(
db: DatabaseSync,
firstCandidateAnimeId: number,
secondCandidateAnimeId: number,
anilistId: number,
): void {
const firstAnimeId = Math.min(firstCandidateAnimeId, secondCandidateAnimeId);
const secondAnimeId = Math.max(firstCandidateAnimeId, secondCandidateAnimeId);
const timestamp = toDbTimestamp(nowMs());
db.prepare(
`INSERT INTO imm_anime_merge_recommendations(
first_anime_id, second_anime_id, anilist_id, status, CREATED_DATE, LAST_UPDATE_DATE
) VALUES (?, ?, ?, 'pending', ?, ?)
ON CONFLICT(first_anime_id, second_anime_id, anilist_id) DO UPDATE SET
LAST_UPDATE_DATE = excluded.LAST_UPDATE_DATE`,
).run(firstAnimeId, secondAnimeId, anilistId, timestamp, timestamp);
}
export function hasDismissedAnimeMergeRecommendation(
db: DatabaseSync,
firstCandidateAnimeId: number,
secondCandidateAnimeId: number,
): boolean {
const firstAnimeId = Math.min(firstCandidateAnimeId, secondCandidateAnimeId);
const secondAnimeId = Math.max(firstCandidateAnimeId, secondCandidateAnimeId);
return Boolean(
db
.prepare(
`SELECT 1
FROM imm_anime_merge_recommendations
WHERE first_anime_id = ?
AND second_anime_id = ?
AND status = 'dismissed'
LIMIT 1`,
)
.get(firstAnimeId, secondAnimeId),
);
}
export function getAnimeMergeRecommendations(db: DatabaseSync): AnimeMergeRecommendation[] {
return (
db
.prepare(
`SELECT recommendation_id AS recommendationId,
first_anime_id AS firstAnimeId,
second_anime_id AS secondAnimeId
FROM imm_anime_merge_recommendations
WHERE status = 'pending'
ORDER BY recommendation_id ASC`,
)
.all() as Array<{
recommendationId: number;
firstAnimeId: number;
secondAnimeId: number;
}>
).map((row) => ({
recommendationId: row.recommendationId,
animeIds: [row.firstAnimeId, row.secondAnimeId],
}));
}
export function dismissAnimeMergeRecommendation(
db: DatabaseSync,
recommendationId: number,
): boolean {
const result = db
.prepare(
`UPDATE imm_anime_merge_recommendations
SET status = 'dismissed', LAST_UPDATE_DATE = ?
WHERE recommendation_id = ? AND status = 'pending'`,
)
.run(toDbTimestamp(nowMs()), recommendationId) as { changes: number };
return result.changes > 0;
}
@@ -0,0 +1,296 @@
import type { DatabaseSync } from './sqlite';
import { recomputeLifetimeAnimeAggregatesInTransaction } from './lifetime';
import { toDbTimestamp } from './query-shared';
import { nowMs } from './time';
/** Thrown when a move names an episode or destination entry that is not there. */
export const UNKNOWN_MOVE_TARGET_MESSAGE = 'Unknown episode or target library entry';
export interface AnimeMergeSummary {
/** Library entry that owns every moved episode once the merge finishes. */
survivingAnimeId: number;
/** Entries that were folded into the survivor and deleted. */
mergedAnimeIds: number[];
movedVideos: number;
}
export interface VideoMoveSummary {
targetAnimeId: number;
/** Previous owner, or null when the episode had no library entry yet. */
previousAnimeId: number | null;
/** True when the previous owner was left empty and pruned. */
removedPreviousAnime: boolean;
}
interface AnimeMetadataRow {
normalized_title_key: string;
anilist_id: number | null;
title_romaji: string | null;
title_english: string | null;
title_native: string | null;
episodes_total: number | null;
description: string | null;
}
function emptyMergeSummary(survivingAnimeId: number): AnimeMergeSummary {
return { survivingAnimeId, mergedAnimeIds: [], movedVideos: 0 };
}
function runInTransaction<T>(db: DatabaseSync, work: () => T): T {
db.exec('BEGIN IMMEDIATE');
try {
const result = work();
db.exec('COMMIT');
return result;
} catch (error) {
db.exec('ROLLBACK');
throw error;
}
}
function readAnimeMetadata(db: DatabaseSync, animeId: number): AnimeMetadataRow | null {
return (db
.prepare(
`
SELECT normalized_title_key, anilist_id, title_romaji, title_english, title_native, episodes_total, description
FROM imm_anime
WHERE anime_id = ?
`,
)
.get(animeId) ?? null) as AnimeMetadataRow | null;
}
function animeExists(db: DatabaseSync, animeId: number): boolean {
return Boolean(db.prepare('SELECT 1 FROM imm_anime WHERE anime_id = ?').get(animeId));
}
function hasAnimeReferences(db: DatabaseSync, animeId: number): boolean {
const row = db
.prepare(
`
SELECT 1 AS found
WHERE EXISTS (SELECT 1 FROM imm_videos WHERE anime_id = ?)
OR EXISTS (SELECT 1 FROM imm_subtitle_lines WHERE anime_id = ?)
`,
)
.get(animeId, animeId) as { found: number } | null;
return Boolean(row);
}
/**
* Distinct explicit seasons behind a library entry. Videos with no parsed
* season are ignored, so an entry built from `Show - 03.mkv` style filenames
* reports an empty set rather than a bogus season.
*/
export function getParsedSeasonsForAnime(db: DatabaseSync, animeId: number): Set<number> {
const rows = db
.prepare(
`
SELECT DISTINCT parsed_season AS season
FROM imm_videos
WHERE anime_id = ?
AND parsed_season IS NOT NULL
AND parsed_season > 0
`,
)
.all(animeId) as Array<{ season: number }>;
return new Set(rows.map((row) => row.season));
}
/**
* Two entries are safe to fold together when neither spans more than one
* explicit season and they do not disagree about which season that is. A
* seasonless entry is compatible with anything single-season: those are the
* `Show - 03.mkv` vs `Show.S01E03.mkv` splits that produce duplicate cards.
*/
export function animeSeasonsAreMergeCompatible(a: Set<number>, b: Set<number>): boolean {
if (a.size > 1 || b.size > 1) return false;
if (a.size === 0 || b.size === 0) return true;
return [...a][0] === [...b][0];
}
/**
* Fill in whatever the target is missing from a source row that is on its way
* out. Must run after the source row is deleted: imm_anime.anilist_id is
* UNIQUE, so the two rows cannot hold the same id at once.
*/
function absorbAnimeMetadata(
db: DatabaseSync,
targetAnimeId: number,
source: AnimeMetadataRow | null,
updatedAt: string,
): void {
if (!source) return;
db.prepare(
`
UPDATE imm_anime
SET
anilist_id = COALESCE(anilist_id, ?),
title_romaji = COALESCE(title_romaji, ?),
title_english = COALESCE(title_english, ?),
title_native = COALESCE(title_native, ?),
episodes_total = COALESCE(episodes_total, ?),
description = COALESCE(description, ?),
LAST_UPDATE_DATE = ?
WHERE anime_id = ?
`,
).run(
source.anilist_id,
source.title_romaji,
source.title_english,
source.title_native,
source.episodes_total,
source.description,
updatedAt,
targetAnimeId,
);
}
/**
* Fold `sourceAnimeIds` into `targetAnimeId`: every episode and subtitle line
* is repointed, metadata the target is missing is inherited from the sources,
* and the emptied source rows are deleted.
*
* Assumes the caller already holds a write transaction and refreshes the
* per-anime lifetime aggregates afterwards; use {@link mergeAnimeRecords}
* otherwise.
*/
export function mergeAnimeRecordsInTransaction(
db: DatabaseSync,
targetAnimeId: number,
sourceAnimeIds: number[],
): AnimeMergeSummary {
const summary = emptyMergeSummary(targetAnimeId);
if (!animeExists(db, targetAnimeId)) {
return summary;
}
const updatedAt = toDbTimestamp(nowMs());
const sourceVideosStmt = db.prepare(
'SELECT video_id AS videoId FROM imm_videos WHERE anime_id = ?',
);
const moveVideosStmt = db.prepare(
'UPDATE imm_videos SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE anime_id = ?',
);
// Repointed per video rather than by anime_id: lines recorded before the
// async title parse assigns the link are stored with a NULL anime_id, and
// matching on the source id would strand them unattributed.
const moveLinesStmt = db.prepare(
'UPDATE imm_subtitle_lines SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE video_id = ?',
);
const dropLifetimeStmt = db.prepare('DELETE FROM imm_lifetime_anime WHERE anime_id = ?');
const sourceAliasesStmt = db.prepare(
'SELECT normalized_title_key AS normalizedTitleKey FROM imm_anime_title_aliases WHERE anime_id = ?',
);
const upsertAliasStmt = db.prepare(
`INSERT INTO imm_anime_title_aliases(normalized_title_key, anime_id, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?)
ON CONFLICT(normalized_title_key) DO UPDATE SET
anime_id = excluded.anime_id,
LAST_UPDATE_DATE = excluded.LAST_UPDATE_DATE`,
);
const dropSourceAliasesStmt = db.prepare(
'DELETE FROM imm_anime_title_aliases WHERE anime_id = ?',
);
const dropAnimeStmt = db.prepare('DELETE FROM imm_anime WHERE anime_id = ?');
for (const sourceAnimeId of new Set(sourceAnimeIds)) {
if (sourceAnimeId === targetAnimeId || !animeExists(db, sourceAnimeId)) {
continue;
}
const sourceMetadata = readAnimeMetadata(db, sourceAnimeId);
const sourceAliases = sourceAliasesStmt.all(sourceAnimeId) as Array<{
normalizedTitleKey: string;
}>;
const sourceVideoIds = (sourceVideosStmt.all(sourceAnimeId) as Array<{ videoId: number }>).map(
(row) => row.videoId,
);
const moved = moveVideosStmt.run(targetAnimeId, updatedAt, sourceAnimeId) as {
changes: number;
};
for (const videoId of sourceVideoIds) {
moveLinesStmt.run(targetAnimeId, updatedAt, videoId);
}
dropSourceAliasesStmt.run(sourceAnimeId);
for (const alias of [
...(sourceMetadata ? [sourceMetadata.normalized_title_key] : []),
...sourceAliases.map((row) => row.normalizedTitleKey),
]) {
upsertAliasStmt.run(alias, targetAnimeId, updatedAt, updatedAt);
}
dropLifetimeStmt.run(sourceAnimeId);
dropAnimeStmt.run(sourceAnimeId);
absorbAnimeMetadata(db, targetAnimeId, sourceMetadata, updatedAt);
summary.mergedAnimeIds.push(sourceAnimeId);
summary.movedVideos += moved.changes;
}
return summary;
}
export function mergeAnimeRecords(
db: DatabaseSync,
targetAnimeId: number,
sourceAnimeIds: number[],
): AnimeMergeSummary {
return runInTransaction(db, () => {
const summary = mergeAnimeRecordsInTransaction(db, targetAnimeId, sourceAnimeIds);
if (summary.mergedAnimeIds.length > 0) {
recomputeLifetimeAnimeAggregatesInTransaction(db);
}
return summary;
});
}
/**
* Move a single episode to another library entry, pruning the previous owner
* when it is left with nothing.
*/
export function moveVideoToAnime(
db: DatabaseSync,
videoId: number,
targetAnimeId: number,
): VideoMoveSummary {
return runInTransaction(db, () => {
const videoRow = db
.prepare('SELECT anime_id AS animeId FROM imm_videos WHERE video_id = ?')
.get(videoId) as { animeId: number | null } | null;
if (!videoRow || !animeExists(db, targetAnimeId)) {
throw new Error(UNKNOWN_MOVE_TARGET_MESSAGE);
}
const previousAnimeId = videoRow.animeId;
if (previousAnimeId === targetAnimeId) {
db.prepare(
'UPDATE imm_videos SET anime_assignment_locked = 1, LAST_UPDATE_DATE = ? WHERE video_id = ?',
).run(toDbTimestamp(nowMs()), videoId);
return { targetAnimeId, previousAnimeId, removedPreviousAnime: false };
}
const updatedAt = toDbTimestamp(nowMs());
db.prepare(
`UPDATE imm_videos
SET anime_id = ?, anime_assignment_locked = 1, LAST_UPDATE_DATE = ?
WHERE video_id = ?`,
).run(targetAnimeId, updatedAt, videoId);
db.prepare(
'UPDATE imm_subtitle_lines SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE video_id = ?',
).run(targetAnimeId, updatedAt, videoId);
let removedPreviousAnime = false;
if (previousAnimeId !== null && !hasAnimeReferences(db, previousAnimeId)) {
// The emptied entry's metadata is deliberately dropped rather than
// absorbed. A move says "this episode belongs elsewhere", not "these are
// the same show", and the entry being emptied is usually a mis-parse
// whose AniList link would be wrong for the target.
db.prepare('DELETE FROM imm_lifetime_anime WHERE anime_id = ?').run(previousAnimeId);
db.prepare('DELETE FROM imm_anime WHERE anime_id = ?').run(previousAnimeId);
removedPreviousAnime = true;
}
recomputeLifetimeAnimeAggregatesInTransaction(db);
return { targetAnimeId, previousAnimeId, removedPreviousAnime };
});
}
@@ -1,4 +1,16 @@
import type { DatabaseSync } from './sqlite';
import {
animeSeasonsAreMergeCompatible,
getParsedSeasonsForAnime,
mergeAnimeRecordsInTransaction,
} from './anime-merge';
import {
hasExactStoredTitleMatch,
hasDismissedAnimeMergeRecommendation,
recordAnimeMergeRecommendation,
shouldRecommendAnilistConflict,
type AnimeConflictRecommendationOptions,
} from './anime-merge-recommendations';
import { getOrCreateAnimeRecord } from './storage';
import { toDbTimestamp } from './query-shared';
import { nowMs } from './time';
@@ -8,8 +20,33 @@ export interface AnimeSeasonRepairSummary {
repaired: number;
movedVideos: number;
deletedAnimeRows: number;
/**
* Entry that owns the videos afterwards when two rows were folded together,
* so callers can keep pointing at a row that still exists.
*/
survivingAnimeId: number | null;
/** True when an ambiguous AniList collision was saved for user review. */
mergeRecommended: boolean;
/** True when automatic metadata must not assign the colliding AniList id. */
anilistAssignmentBlocked: boolean;
}
export interface AnimeAnilistConflictOptions extends AnimeConflictRecommendationOptions {
/**
* Which row keeps its identity when two entries claim the same AniList id.
* `existing` (the default) keeps the row that already held the id, so
* automatic cover-art resolution does not rename a card under the user;
* `target` keeps the row the user is acting on.
*/
survivor?: 'target' | 'existing';
}
export {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
type AnimeMergeRecommendation,
} from './anime-merge-recommendations';
interface AnimeRow {
anime_id: number;
anilist_id: number | null;
@@ -24,6 +61,7 @@ interface ParsedVideoRow {
video_id: number;
parsed_title: string | null;
parsed_season: number | null;
anime_assignment_locked: number;
}
interface RedistributeOptions {
@@ -38,6 +76,9 @@ function emptySummary(scanned = 0): AnimeSeasonRepairSummary {
repaired: 0,
movedVideos: 0,
deletedAnimeRows: 0,
survivingAnimeId: null,
mergeRecommended: false,
anilistAssignmentBlocked: false,
};
}
@@ -49,6 +90,9 @@ function mergeSummary(
target.repaired += source.repaired;
target.movedVideos += source.movedVideos;
target.deletedAnimeRows += source.deletedAnimeRows;
target.survivingAnimeId = source.survivingAnimeId ?? target.survivingAnimeId;
target.mergeRecommended ||= source.mergeRecommended;
target.anilistAssignmentBlocked ||= source.anilistAssignmentBlocked;
return target;
}
@@ -94,7 +138,7 @@ function getParsedVideos(db: DatabaseSync, animeId: number): ParsedVideoRow[] {
return db
.prepare(
`
SELECT video_id, parsed_title, parsed_season
SELECT video_id, parsed_title, parsed_season, anime_assignment_locked
FROM imm_videos
WHERE anime_id = ?
ORDER BY video_id ASC
@@ -188,6 +232,9 @@ function redistributeAnimeRowByParsedSeasonsInTransaction(
const targetBySeason = new Map<number, number>();
for (const video of videos) {
if (video.anime_assignment_locked === 1) {
continue;
}
const parsedTitle = video.parsed_title?.trim();
const season = normalizeSeason(video.parsed_season);
if (!parsedTitle || season === null) {
@@ -301,10 +348,18 @@ export function repairLegacySeasonlessAnimeRows(db: DatabaseSync): AnimeSeasonRe
});
}
/**
* Two library entries cannot both hold the same AniList id
* (`imm_anime.anilist_id` is UNIQUE). Fold an automatic collision only when
* exact title evidence and compatible parsed seasons make it safe. Persist a
* review recommendation for compatible weak matches. Fall back to legacy
* season redistribution when the conflicting row spans several seasons.
*/
export function resolveAnimeAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
anilistId: number,
options: AnimeAnilistConflictOptions = {},
): AnimeSeasonRepairSummary {
const conflict = db
.prepare(
@@ -321,10 +376,100 @@ export function resolveAnimeAnilistConflict(
return emptySummary();
}
return runInTransaction(db, () =>
redistributeAnimeRowByParsedSeasonsInTransaction(db, conflict.animeId, {
return runInTransaction(db, () => {
const targetRow = getAnimeRow(db, targetAnimeId);
if (
options.survivor !== 'target' &&
targetRow?.anilist_id != null &&
targetRow.anilist_id !== anilistId
) {
// An automatic lookup disagreeing with an existing explicit link is a
// mis-resolution, not evidence that either row should move or merge. The
// colliding id must not be assigned either: another row owns it and
// imm_anime.anilist_id is UNIQUE.
const summary = emptySummary(1);
summary.anilistAssignmentBlocked = true;
return summary;
}
const isManual = options.survivor === 'target' || options.matchConfidence === 'manual';
if (!isManual && hasDismissedAnimeMergeRecommendation(db, targetAnimeId, conflict.animeId)) {
const summary = emptySummary(1);
summary.anilistAssignmentBlocked = true;
return summary;
}
const targetSeasons = getParsedSeasonsForAnime(db, targetAnimeId);
const conflictSeasons = getParsedSeasonsForAnime(db, conflict.animeId);
if (
!isManual &&
targetSeasons.size === 1 &&
conflictSeasons.size === 1 &&
[...targetSeasons][0] !== [...conflictSeasons][0]
) {
const summary = emptySummary(1);
summary.anilistAssignmentBlocked = true;
return summary;
}
if (canMergeAnilistConflict(db, targetAnimeId, conflict.animeId, anilistId, options)) {
const survivingAnimeId = options.survivor === 'target' ? targetAnimeId : conflict.animeId;
const absorbedAnimeId = survivingAnimeId === targetAnimeId ? conflict.animeId : targetAnimeId;
const merge = mergeAnimeRecordsInTransaction(db, survivingAnimeId, [absorbedAnimeId]);
const summary = emptySummary(1);
summary.movedVideos = merge.movedVideos;
summary.deletedAnimeRows = merge.mergedAnimeIds.length;
if (merge.mergedAnimeIds.length > 0) {
summary.repaired = 1;
// Only reported once a row really absorbed the other, so callers never
// follow this to an anime id that was never written.
summary.survivingAnimeId = survivingAnimeId;
}
// Lifetime summaries are rebuilt by the caller off this summary, the same
// as the redistribution path below.
return summary;
}
if (shouldRecommendAnilistConflict(db, targetAnimeId, conflict.animeId, options)) {
recordAnimeMergeRecommendation(db, targetAnimeId, conflict.animeId, anilistId);
const summary = emptySummary(1);
summary.mergeRecommended = true;
return summary;
}
return redistributeAnimeRowByParsedSeasonsInTransaction(db, conflict.animeId, {
transferAnilistToAnimeId: targetAnimeId,
overwriteTargetAnilist: true,
}),
});
});
}
function canMergeAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
anilistId: number,
options: AnimeAnilistConflictOptions,
): boolean {
const targetRow = getAnimeRow(db, targetAnimeId);
if (!targetRow) {
// Nothing to merge with a row that no longer exists (a stale id from the
// caller); fall through to the redistribution path.
return false;
}
if (options.survivor !== 'target') {
// The target is the row about to disappear here, so an existing link of its
// own means this is a mis-resolution rather than a duplicate: leave it be.
if (targetRow.anilist_id != null && targetRow.anilist_id !== anilistId) {
return false;
}
}
if (
options.matchConfidence === 'weak' ||
(options.matchConfidence === undefined &&
!hasExactStoredTitleMatch(db, targetAnimeId, conflictAnimeId))
) {
return false;
}
return animeSeasonsAreMergeCompatible(
getParsedSeasonsForAnime(db, targetAnimeId),
getParsedSeasonsForAnime(db, conflictAnimeId),
);
}
@@ -0,0 +1,160 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { DeleteMaintenanceScheduler } from './delete-maintenance-scheduler';
import type { DeleteMaintenanceTask } from './delete-maintenance';
test('scheduler batches same-turn requests and balances busy state', async () => {
const tasks: DeleteMaintenanceTask[] = [];
const states: string[] = [];
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async (task) => {
tasks.push(task);
},
onBusy: () => states.push('busy'),
onIdle: () => states.push('idle'),
});
const first = scheduler.enqueue(() => ({ kind: 'session', sessionId: 1 }));
const second = scheduler.enqueue(() => ({ kind: 'sessions', sessionIds: [2, 3] }));
const third = scheduler.enqueue(() => null);
await Promise.all([first, second, third]);
assert.deepEqual(tasks, [
{
kind: 'batch',
tasks: [
{ kind: 'session', sessionId: 1 },
{ kind: 'sessions', sessionIds: [2, 3] },
],
},
]);
assert.deepEqual(states, ['busy', 'idle']);
});
test('scheduler rejects enqueue after destruction without entering busy state', async () => {
let busyCalls = 0;
let runCalls = 0;
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async () => {
runCalls += 1;
},
onBusy: () => {
busyCalls += 1;
},
onIdle: () => {},
});
scheduler.destroy();
await assert.rejects(
scheduler.enqueue(() => ({ kind: 'session', sessionId: 1 })),
/shutting down/,
);
assert.equal(busyCalls, 0);
assert.equal(runCalls, 0);
});
test('scheduler rejects every request in a batch when the maintenance task fails', async () => {
const failure = new Error('maintenance failed');
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async () => {
throw failure;
},
onBusy: () => {},
onIdle: () => {},
});
const first = scheduler.enqueue(() => ({ kind: 'session', sessionId: 1 }));
const second = scheduler.enqueue(() => ({ kind: 'session', sessionId: 2 }));
const results = await Promise.allSettled([first, second]);
assert.deepEqual(
results.map((result) => (result.status === 'rejected' ? result.reason : null)),
[failure, failure],
);
});
test('scheduler rejects only the request whose task resolution fails', async () => {
const failure = new Error('resolution failed');
const tasks: DeleteMaintenanceTask[] = [];
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async (task) => {
tasks.push(task);
},
onBusy: () => {},
onIdle: () => {},
});
const failed = scheduler.enqueue(() => {
throw failure;
});
const succeeded = scheduler.enqueue(() => ({ kind: 'session', sessionId: 2 }));
const results = await Promise.allSettled([failed, succeeded]);
assert.equal(results[0]?.status, 'rejected');
assert.equal(results[0]?.status === 'rejected' ? results[0].reason : null, failure);
assert.equal(results[1]?.status, 'fulfilled');
assert.deepEqual(tasks, [{ kind: 'session', sessionId: 2 }]);
});
test('scheduler does not schedule another drain when the queue is empty', async () => {
const originalSetTimeout = globalThis.setTimeout;
let timerCalls = 0;
globalThis.setTimeout = ((handler: TimerHandler, timeout?: number, ...args: unknown[]) => {
timerCalls += 1;
return originalSetTimeout(handler, timeout, ...args);
}) as typeof setTimeout;
try {
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async () => {},
onBusy: () => {},
onIdle: () => {},
});
await scheduler.enqueue(() => ({ kind: 'session', sessionId: 1 }));
assert.equal(timerCalls, 1);
} finally {
globalThis.setTimeout = originalSetTimeout;
}
});
test('scheduler serializes batches and rejects requests queued at destruction', async () => {
const releases: Array<() => void> = [];
let activeTasks = 0;
let maxActiveTasks = 0;
const scheduler = new DeleteMaintenanceScheduler({
batchWindowMs: 0,
runTask: async () => {
activeTasks += 1;
maxActiveTasks = Math.max(maxActiveTasks, activeTasks);
await new Promise<void>((resolve) => releases.push(resolve));
activeTasks -= 1;
},
onBusy: () => {},
onIdle: () => {},
});
const first = scheduler.enqueue(() => ({ kind: 'session', sessionId: 1 }));
const maxPollAttempts = 100;
let pollAttempts = 0;
while (releases.length === 0 && pollAttempts < maxPollAttempts) {
pollAttempts += 1;
await new Promise<void>((resolve) => setTimeout(resolve, 0));
}
assert.ok(
releases.length > 0,
`runTask did not produce a release after ${maxPollAttempts} polling attempts`,
);
const queued = scheduler.enqueue(() => ({ kind: 'session', sessionId: 2 }));
scheduler.destroy();
await assert.rejects(queued, /shutting down/);
releases[0]?.();
await first;
assert.equal(maxActiveTasks, 1);
});
@@ -0,0 +1,105 @@
import type { DeleteMaintenanceOperation, DeleteMaintenanceTask } from './delete-maintenance';
type ResolveDeleteMaintenanceOperation = () =>
| DeleteMaintenanceOperation
| null
| Promise<DeleteMaintenanceOperation | null>;
interface PendingDeleteMaintenanceRequest {
resolveTask: ResolveDeleteMaintenanceOperation;
resolve: () => void;
reject: (error: unknown) => void;
}
interface DeleteMaintenanceSchedulerOptions {
batchWindowMs: number;
runTask: (task: DeleteMaintenanceTask) => Promise<void>;
onBusy: () => void;
onIdle: () => void;
}
export class DeleteMaintenanceScheduler {
private readonly pendingRequests: PendingDeleteMaintenanceRequest[] = [];
private running = false;
private drainTimer: ReturnType<typeof setTimeout> | null = null;
private pendingTaskCount = 0;
private destroyed = false;
constructor(private readonly options: DeleteMaintenanceSchedulerOptions) {}
enqueue(resolveTask: ResolveDeleteMaintenanceOperation): Promise<void> {
if (this.destroyed) {
return Promise.reject(new Error('Immersion tracker is shutting down'));
}
if (this.pendingTaskCount === 0) this.options.onBusy();
this.pendingTaskCount += 1;
const result = new Promise<void>((resolve, reject) => {
this.pendingRequests.push({ resolveTask, resolve, reject });
this.scheduleDrain();
});
return result.finally(() => {
this.pendingTaskCount -= 1;
if (this.pendingTaskCount === 0) this.options.onIdle();
});
}
destroy(): void {
if (this.destroyed) return;
this.destroyed = true;
if (this.drainTimer) {
clearTimeout(this.drainTimer);
this.drainTimer = null;
}
const error = new Error('Immersion tracker is shutting down');
for (const request of this.pendingRequests.splice(0)) request.reject(error);
}
private scheduleDrain(): void {
if (this.destroyed || this.running || this.drainTimer || this.pendingRequests.length === 0) {
return;
}
this.drainTimer = setTimeout(() => {
this.drainTimer = null;
void this.drain();
}, this.options.batchWindowMs);
}
private async drain(): Promise<void> {
if (this.running || this.pendingRequests.length === 0) return;
this.running = true;
const requests = this.pendingRequests.splice(0);
const runnable: Array<{
request: PendingDeleteMaintenanceRequest;
task: DeleteMaintenanceOperation;
}> = [];
for (const request of requests) {
try {
const task = await request.resolveTask();
if (task) runnable.push({ request, task });
else request.resolve();
} catch (error) {
request.reject(error);
}
}
if (runnable.length > 0) {
const task: DeleteMaintenanceTask =
runnable.length === 1
? runnable[0]!.task
: { kind: 'batch', tasks: runnable.map((entry) => entry.task) };
try {
await this.options.runTask(task);
for (const { request } of runnable) request.resolve();
} catch (error) {
for (const { request } of runnable) request.reject(error);
}
}
this.running = false;
this.scheduleDrain();
}
}
@@ -0,0 +1,239 @@
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import test from 'node:test';
import {
DeleteMaintenanceWorkerRuntime,
resolveDeleteMaintenanceWorkerPath,
} from './delete-maintenance-worker-runtime';
import { executeDeleteMaintenanceTask } from './delete-maintenance';
import { startSessionRecord } from './session';
import { Database } from './sqlite';
import { applyPragmas, ensureSchema, getOrCreateVideoRecord } from './storage';
type FakeWorkerListener = (value: never) => void;
function createFakeWorker() {
const listeners = new Map<string, FakeWorkerListener>();
const terminationState = { calls: 0 };
const worker = {
once(event: string, listener: FakeWorkerListener) {
listeners.set(event, listener);
return this;
},
terminate: async () => {
terminationState.calls += 1;
return 0;
},
};
return { worker, listeners, terminationState };
}
type FakeWorker = ReturnType<typeof createFakeWorker>['worker'];
test('a delete batch rebuilds lifetime summaries once', () => {
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-delete-batch-test-'));
const dbPath = path.join(tempDir, 'immersion.sqlite');
let db = new Database(dbPath);
try {
applyPragmas(db);
ensureSchema(db);
const videoId = getOrCreateVideoRecord(db, 'local:/tmp/batch-delete.mkv', {
canonicalTitle: 'Batch Delete',
sourcePath: '/tmp/batch-delete.mkv',
sourceUrl: null,
sourceType: 1,
});
const firstSessionId = startSessionRecord(db, videoId, 1_000).sessionId;
const secondSessionId = startSessionRecord(db, videoId, 2_000).sessionId;
const deletedVideoId = getOrCreateVideoRecord(db, 'local:/tmp/batch-delete-video.mkv', {
canonicalTitle: 'Batch Delete Video',
sourcePath: '/tmp/batch-delete-video.mkv',
sourceUrl: null,
sourceType: 1,
});
startSessionRecord(db, deletedVideoId, 3_000);
db.exec(`
CREATE TABLE delete_rebuild_audit (id INTEGER PRIMARY KEY);
CREATE TRIGGER count_delete_lifetime_rebuild
AFTER UPDATE OF last_rebuilt_ms ON imm_lifetime_global
BEGIN
INSERT INTO delete_rebuild_audit (id) VALUES (NULL);
END;
`);
db.close();
executeDeleteMaintenanceTask(dbPath, {
kind: 'batch',
tasks: [
{ kind: 'session', sessionId: firstSessionId },
{ kind: 'video', videoId: deletedVideoId },
],
});
db = new Database(dbPath);
const audit = db.prepare('SELECT COUNT(*) AS total FROM delete_rebuild_audit').get() as {
total: number;
};
const retainedSession = db
.prepare('SELECT session_id AS sessionId FROM imm_sessions WHERE video_id = ?')
.get(videoId) as { sessionId: number } | null;
const deletedVideo = db
.prepare('SELECT video_id AS videoId FROM imm_videos WHERE video_id = ?')
.get(deletedVideoId) as { videoId: number } | null;
assert.equal(retainedSession?.sessionId, secondSessionId);
assert.equal(deletedVideo, undefined);
assert.equal(
audit.total,
2,
'one rebuild performs exactly its reset and final global summary writes',
);
} finally {
try {
db.close();
} catch {
// The setup connection closes before maintenance runs.
}
fs.rmSync(tempDir, { recursive: true, force: true });
}
});
test(
'compiled delete worker removes data through its separate database connection',
{ skip: resolveDeleteMaintenanceWorkerPath() === null },
async () => {
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-delete-worker-test-'));
const dbPath = path.join(tempDir, 'immersion.sqlite');
const runtime = new DeleteMaintenanceWorkerRuntime();
let db = new Database(dbPath);
try {
applyPragmas(db);
ensureSchema(db);
const videoId = getOrCreateVideoRecord(db, 'local:/tmp/worker-delete.mkv', {
canonicalTitle: 'Worker Delete',
sourcePath: '/tmp/worker-delete.mkv',
sourceUrl: null,
sourceType: 1,
});
const firstSessionId = startSessionRecord(db, videoId, 1_000).sessionId;
const secondSessionId = startSessionRecord(db, videoId, 2_000).sessionId;
db.close();
await runtime.run(dbPath, {
kind: 'batch',
tasks: [
{ kind: 'session', sessionId: firstSessionId },
{ kind: 'session', sessionId: secondSessionId },
],
});
db = new Database(dbPath);
const row = db
.prepare('SELECT COUNT(*) AS total FROM imm_sessions WHERE video_id = ?')
.get(videoId) as { total: number };
assert.equal(row.total, 0);
} finally {
runtime.destroy();
try {
db.close();
} catch {
// The setup connection is already closed before the worker starts.
}
fs.rmSync(tempDir, { recursive: true, force: true });
}
},
);
test('worker runtime warns before falling back when no emitted worker is available', async () => {
const warnings: unknown[][] = [];
const fallbackTasks: unknown[] = [];
const runtime = new DeleteMaintenanceWorkerRuntime({
resolveWorkerPath: () => null,
warn: (...args) => warnings.push(args),
executeFallback: (_dbPath, task) => fallbackTasks.push(task),
});
await runtime.run('/tmp/fallback.sqlite', { kind: 'session', sessionId: 1 });
assert.equal(warnings.length, 1);
assert.match(String(warnings[0]?.[0]), /worker unavailable/i);
assert.deepEqual(fallbackTasks, [{ kind: 'session', sessionId: 1 }]);
});
test('worker runtime terminates a worker after successful settlement', async () => {
const { worker, listeners, terminationState } = createFakeWorker();
const runtime = new DeleteMaintenanceWorkerRuntime({
resolveWorkerPath: () => '/tmp/delete-worker.js',
createWorker: async () => worker,
});
const result = runtime.run('/tmp/test.sqlite', { kind: 'session', sessionId: 1 });
await new Promise<void>((resolve) => setTimeout(resolve, 0));
listeners.get('message')?.({ ok: true } as never);
await result;
assert.equal(terminationState.calls, 1);
});
test('worker runtime terminates a worker after failed settlement', async () => {
const { worker, listeners, terminationState } = createFakeWorker();
const runtime = new DeleteMaintenanceWorkerRuntime({
resolveWorkerPath: () => '/tmp/delete-worker.js',
createWorker: async () => worker,
});
const result = runtime.run('/tmp/test.sqlite', { kind: 'session', sessionId: 1 });
await new Promise<void>((resolve) => setTimeout(resolve, 0));
listeners.get('error')?.(new Error('worker failed') as never);
await assert.rejects(result, /worker failed/);
assert.equal(terminationState.calls, 1);
});
test('worker runtime terminates a worker created after shutdown begins', async () => {
const { worker, listeners, terminationState } = createFakeWorker();
const createGate: { resolve?: (worker: FakeWorker) => void } = {};
const fallbackTasks: unknown[] = [];
const runtime = new DeleteMaintenanceWorkerRuntime({
resolveWorkerPath: () => '/tmp/delete-worker.js',
createWorker: () =>
new Promise((resolve) => {
createGate.resolve = resolve;
}),
executeFallback: (_dbPath, task) => fallbackTasks.push(task),
});
const result = runtime.run('/tmp/test.sqlite', { kind: 'session', sessionId: 1 });
await new Promise<void>((resolve) => setTimeout(resolve, 0));
runtime.destroy();
createGate.resolve?.(worker);
await assert.rejects(result, /shut down/);
assert.equal(terminationState.calls, 1);
assert.equal(listeners.size, 0);
assert.deepEqual(fallbackTasks, []);
});
test('worker runtime does not fall back when worker creation fails during shutdown', async () => {
const createGate: { reject?: (error: Error) => void } = {};
const fallbackTasks: unknown[] = [];
const runtime = new DeleteMaintenanceWorkerRuntime({
resolveWorkerPath: () => '/tmp/delete-worker.js',
createWorker: () =>
new Promise((_resolve, reject) => {
createGate.reject = reject;
}),
executeFallback: (_dbPath, task) => fallbackTasks.push(task),
});
const result = runtime.run('/tmp/test.sqlite', { kind: 'session', sessionId: 1 });
await new Promise<void>((resolve) => setTimeout(resolve, 0));
runtime.destroy();
createGate.reject?.(new Error('creation failed'));
await assert.rejects(result, /shut down/);
assert.deepEqual(fallbackTasks, []);
});
@@ -0,0 +1,121 @@
import fs from 'node:fs';
import path from 'node:path';
import { createLogger } from '../../../logger';
import { executeDeleteMaintenanceTask, type DeleteMaintenanceTask } from './delete-maintenance';
interface DeleteMaintenanceWorkerResponse {
ok?: unknown;
error?: unknown;
}
export type RunDeleteMaintenanceTask = (
dbPath: string,
task: DeleteMaintenanceTask,
) => Promise<void>;
interface DeleteMaintenanceWorkerHandle {
once(event: 'message', listener: (message: DeleteMaintenanceWorkerResponse) => void): this;
once(event: 'error', listener: (error: Error) => void): this;
once(event: 'exit', listener: (code: number) => void): this;
terminate(): Promise<number>;
}
interface DeleteMaintenanceWorkerRuntimeOptions {
resolveWorkerPath?: () => string | null;
createWorker?: (
workerPath: string,
workerData: { dbPath: string; task: DeleteMaintenanceTask },
) => Promise<DeleteMaintenanceWorkerHandle>;
executeFallback?: typeof executeDeleteMaintenanceTask;
warn?: (message: string, ...meta: unknown[]) => void;
}
export function resolveDeleteMaintenanceWorkerPath(): string | null {
const workerPath = path.join(__dirname, 'delete-maintenance-worker-thread.js');
return fs.existsSync(workerPath) ? workerPath : null;
}
const logger = createLogger('main:immersion-tracker:delete-worker');
export class DeleteMaintenanceWorkerRuntime {
private readonly activeWorkers = new Set<DeleteMaintenanceWorkerHandle>();
private destroyed = false;
constructor(private readonly options: DeleteMaintenanceWorkerRuntimeOptions = {}) {}
async run(dbPath: string, task: DeleteMaintenanceTask): Promise<void> {
if (this.destroyed) {
throw new Error('Delete maintenance worker is shut down');
}
let worker: DeleteMaintenanceWorkerHandle;
try {
const workerPath = (this.options.resolveWorkerPath ?? resolveDeleteMaintenanceWorkerPath)();
if (!workerPath) throw new Error('Emitted delete-maintenance worker module was not found');
const createWorker =
this.options.createWorker ??
(async (resolvedPath, workerData) => {
const { Worker } = await import('node:worker_threads');
return new Worker(resolvedPath, { workerData });
});
worker = await createWorker(workerPath, { dbPath, task });
} catch (error) {
if (this.destroyed) {
throw new Error('Delete maintenance worker is shut down');
}
(this.options.warn ?? logger.warn)(
'Delete maintenance worker unavailable; running maintenance on the current thread',
error,
);
(this.options.executeFallback ?? executeDeleteMaintenanceTask)(dbPath, task);
return;
}
if (this.destroyed) {
await worker.terminate().catch(() => undefined);
throw new Error('Delete maintenance worker is shut down');
}
await new Promise<void>((resolve, reject) => {
let settled = false;
this.activeWorkers.add(worker);
const settle = (error?: Error) => {
if (settled) return;
settled = true;
this.activeWorkers.delete(worker);
if (error) reject(error);
else resolve();
void worker.terminate();
};
worker.once('message', (message: DeleteMaintenanceWorkerResponse) => {
if (message.ok === true) {
settle();
return;
}
const detail = typeof message.error === 'string' ? message.error : 'unknown worker error';
settle(new Error(`Delete maintenance failed: ${detail}`));
});
worker.once('error', (error) => settle(error));
worker.once('exit', (code) => {
settle(
new Error(
code === 0
? 'Delete maintenance worker exited without a response'
: `Delete maintenance worker exited with code ${code}`,
),
);
});
});
}
destroy(): void {
if (this.destroyed) return;
this.destroyed = true;
for (const worker of this.activeWorkers) {
void worker.terminate();
}
this.activeWorkers.clear();
}
}
@@ -0,0 +1,22 @@
import { parentPort, workerData } from 'node:worker_threads';
import { executeDeleteMaintenanceTask, type DeleteMaintenanceTask } from './delete-maintenance';
interface DeleteMaintenanceWorkerData {
dbPath: string;
task: DeleteMaintenanceTask;
}
if (!parentPort) {
throw new Error('delete maintenance worker missing parent port');
}
const port = parentPort;
const request = workerData as DeleteMaintenanceWorkerData;
try {
executeDeleteMaintenanceTask(request.dbPath, request.task);
port.postMessage({ ok: true });
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
port.postMessage({ ok: false, error: message });
}
@@ -0,0 +1,47 @@
import { Database } from './sqlite';
import { applyPragmas } from './storage';
import { deleteAnime, deleteSession, deleteSessions, deleteVideo } from './query-maintenance';
import {
deleteMaintenanceBatch,
type DeleteMaintenanceOperation,
} from './query-delete-maintenance';
export type { DeleteMaintenanceOperation } from './query-delete-maintenance';
export type DeleteMaintenanceTask =
| DeleteMaintenanceOperation
| { kind: 'batch'; tasks: DeleteMaintenanceOperation[] };
function executeDeleteMaintenanceOperation(
db: InstanceType<typeof Database>,
task: DeleteMaintenanceOperation,
): void {
switch (task.kind) {
case 'session':
deleteSession(db, task.sessionId);
return;
case 'sessions':
deleteSessions(db, task.sessionIds);
return;
case 'video':
deleteVideo(db, task.videoId);
return;
case 'anime':
deleteAnime(db, task.animeId);
return;
}
}
export function executeDeleteMaintenanceTask(dbPath: string, task: DeleteMaintenanceTask): void {
const db = new Database(dbPath);
try {
applyPragmas(db);
if (task.kind === 'batch') {
deleteMaintenanceBatch(db, task.tasks);
return;
}
executeDeleteMaintenanceOperation(db, task);
} finally {
db.close();
}
}
@@ -8,6 +8,8 @@ import type { JellyfinLinkRepairSummary } from './types';
type LegacyJellyfinVideoRow = {
video_id: number;
video_key: string;
anime_id: number | null;
anime_assignment_locked: number;
source_url: string | null;
canonical_title: string;
};
@@ -15,6 +17,7 @@ type LegacyJellyfinVideoRow = {
type JellyfinTargetVideoRow = {
video_id: number;
anime_id: number | null;
anime_assignment_locked: number;
canonical_title: string;
parsed_basename: string | null;
parsed_title: string | null;
@@ -258,7 +261,13 @@ export function repairJellyfinStreamVideoLinks(db: DatabaseSync): JellyfinLinkRe
const candidates = db
.prepare(
`
SELECT video_id, video_key, source_url, canonical_title
SELECT
video_id,
video_key,
anime_id,
anime_assignment_locked,
source_url,
canonical_title
FROM imm_videos
WHERE source_type = 2
AND (
@@ -310,6 +319,7 @@ export function repairJellyfinStreamVideoLinks(db: DatabaseSync): JellyfinLinkRe
SELECT
video_id,
anime_id,
anime_assignment_locked,
canonical_title,
parsed_basename,
parsed_title,
@@ -357,12 +367,17 @@ export function repairJellyfinStreamVideoLinks(db: DatabaseSync): JellyfinLinkRe
continue;
}
const assignmentAnimeId =
candidate.anime_assignment_locked === 1 ? candidate.anime_id : target.anime_id;
const assignmentLocked =
candidate.anime_assignment_locked === 1 || target.anime_assignment_locked === 1 ? 1 : 0;
db.prepare(
`
UPDATE imm_videos
SET
video_key = ?,
anime_id = ?,
anime_assignment_locked = ?,
canonical_title = ?,
source_url = ?,
parsed_basename = ?,
@@ -377,7 +392,8 @@ export function repairJellyfinStreamVideoLinks(db: DatabaseSync): JellyfinLinkRe
`,
).run(
sanitizedVideoKey,
target.anime_id,
assignmentAnimeId,
assignmentLocked,
target.canonical_title,
statsUrl,
target.parsed_basename,
@@ -390,14 +406,14 @@ export function repairJellyfinStreamVideoLinks(db: DatabaseSync): JellyfinLinkRe
currentTimestamp,
candidate.video_id,
);
if (target.anime_id !== null) {
if (assignmentAnimeId !== null) {
db.prepare(
`
UPDATE imm_subtitle_lines
SET anime_id = ?, LAST_UPDATE_DATE = ?
WHERE video_id = ?
`,
).run(target.anime_id, currentTimestamp, candidate.video_id);
).run(assignmentAnimeId, currentTimestamp, candidate.video_id);
}
summary.repaired += 1;
}
@@ -708,6 +708,87 @@ export function rebuildLifetimeSummariesInTransaction(
return rebuildLifetimeSummariesInternal(db, rebuiltAtMs);
}
/**
* Re-derive every per-anime lifetime row from the per-video summaries after
* episodes changed owners (merge, move, season repair).
*
* Deliberately NOT a full rebuild: {@link rebuildLifetimeSummariesInTransaction}
* recomputes from raw sessions, which are pruned after the retention window, so
* it silently truncates lifetime history. `imm_lifetime_media` is keyed by
* video and survives repointing, so aggregating it preserves all-time totals;
* `imm_lifetime_global` only needs `anime_completed` refreshed because moving
* attribution between entries cannot change the global counters.
*
* Assumes the caller holds a write transaction; use
* {@link recomputeLifetimeAnimeAggregates} otherwise.
*/
export function recomputeLifetimeAnimeAggregatesInTransaction(db: DatabaseSync): void {
const updatedAt = toDbTimestamp(nowMs());
db.exec('DELETE FROM imm_lifetime_anime');
db.prepare(
`
INSERT INTO imm_lifetime_anime (
anime_id,
total_sessions,
total_active_ms,
total_cards,
total_lines_seen,
total_tokens_seen,
episodes_started,
episodes_completed,
first_watched_ms,
last_watched_ms,
CREATED_DATE,
LAST_UPDATE_DATE
)
SELECT
v.anime_id,
COALESCE(SUM(m.total_sessions), 0),
COALESCE(SUM(m.total_active_ms), 0),
COALESCE(SUM(m.total_cards), 0),
COALESCE(SUM(m.total_lines_seen), 0),
COALESCE(SUM(m.total_tokens_seen), 0),
COUNT(*),
COUNT(CASE WHEN m.completed > 0 THEN 1 END),
MIN(m.first_watched_ms),
MAX(m.last_watched_ms),
?,
?
FROM imm_lifetime_media m
JOIN imm_videos v ON v.video_id = m.video_id
WHERE v.anime_id IS NOT NULL
GROUP BY v.anime_id
`,
).run(updatedAt, updatedAt);
db.prepare(
`
UPDATE imm_lifetime_global
SET
anime_completed = (
SELECT COUNT(*)
FROM imm_lifetime_anime la
JOIN imm_anime a ON a.anime_id = la.anime_id
WHERE a.episodes_total IS NOT NULL
AND a.episodes_total > 0
AND la.episodes_completed >= a.episodes_total
),
LAST_UPDATE_DATE = ?
WHERE global_id = 1
`,
).run(updatedAt);
}
export function recomputeLifetimeAnimeAggregates(db: DatabaseSync): void {
db.exec('BEGIN IMMEDIATE');
try {
recomputeLifetimeAnimeAggregatesInTransaction(db);
db.exec('COMMIT');
} catch (error) {
db.exec('ROLLBACK');
throw error;
}
}
export function reconcileStaleActiveSessions(db: DatabaseSync): number {
const sessions = getRetainedStaleActiveSessions(db);
if (sessions.length === 0) {
@@ -0,0 +1,196 @@
import type { DatabaseSync } from './sqlite';
import { rebuildLifetimeSummariesInTransaction } from './lifetime';
import { getRollupGroupsForSessions, refreshRollupsForGroupsInTransaction } from './maintenance';
import {
applyLexicalRemovals,
cleanupUnusedCoverArtBlobHash,
deleteSessionsByIds,
forEachIdChunk,
makePlaceholders,
planLexicalRemovalsForSessions,
SQLITE_ID_CHUNK_SIZE,
type LexicalRemovalPlan,
} from './query-shared';
export type DeleteMaintenanceOperation =
| { kind: 'session'; sessionId: number }
| { kind: 'sessions'; sessionIds: number[] }
| { kind: 'video'; videoId: number }
| { kind: 'anime'; animeId: number };
function addOperationTargets(
operations: DeleteMaintenanceOperation[],
sessionIds: Set<number>,
videoIds: Set<number>,
animeIds: Set<number>,
): void {
for (const operation of operations) {
switch (operation.kind) {
case 'session':
sessionIds.add(operation.sessionId);
break;
case 'sessions':
for (const sessionId of operation.sessionIds) sessionIds.add(sessionId);
break;
case 'video':
videoIds.add(operation.videoId);
break;
case 'anime':
animeIds.add(operation.animeId);
break;
}
}
}
function selectIds(
db: DatabaseSync,
buildSql: (placeholders: string) => string,
params: number[],
column: string,
): number[] {
if (params.length === 0) return [];
const ids: number[] = [];
forEachIdChunk(params, (chunk) => {
const rows = db.prepare(buildSql(makePlaceholders(chunk))).all(...chunk) as Array<
Record<string, number>
>;
for (const row of rows) ids.push(row[column]!);
});
return ids;
}
function planLexicalRemovalsInChunks(db: DatabaseSync, sessionIds: number[]): LexicalRemovalPlan {
const combined: LexicalRemovalPlan = { words: [], kanji: [] };
const merge = (target: LexicalRemovalPlan['words'], source: LexicalRemovalPlan['words']) => {
const byId = new Map(target.map((entry) => [entry.id, entry]));
for (const entry of source) {
const existing = byId.get(entry.id);
if (!existing) {
const added = { ...entry };
target.push(added);
byId.set(entry.id, added);
continue;
}
existing.removedFrequency += entry.removedFrequency;
if (
entry.removedFirstSeenMs !== null &&
(existing.removedFirstSeenMs === null ||
entry.removedFirstSeenMs < existing.removedFirstSeenMs)
) {
existing.removedFirstSeenMs = entry.removedFirstSeenMs;
}
if (
entry.removedLastSeenMs !== null &&
(existing.removedLastSeenMs === null ||
entry.removedLastSeenMs > existing.removedLastSeenMs)
) {
existing.removedLastSeenMs = entry.removedLastSeenMs;
}
}
};
forEachIdChunk(sessionIds, (chunk) => {
const plan = planLexicalRemovalsForSessions(db, chunk);
merge(combined.words, plan.words);
merge(combined.kanji, plan.kanji);
});
return combined;
}
export function deleteMaintenanceBatch(
db: DatabaseSync,
operations: DeleteMaintenanceOperation[],
): void {
if (operations.length === 0) return;
db.exec('BEGIN IMMEDIATE');
try {
const sessionIds = new Set<number>();
const videoIds = new Set<number>();
const animeIds = new Set<number>();
addOperationTargets(operations, sessionIds, videoIds, animeIds);
const animeIdList = [...animeIds];
for (const videoId of selectIds(
db,
(placeholders) => `SELECT video_id FROM imm_videos WHERE anime_id IN (${placeholders})`,
animeIdList,
'video_id',
)) {
videoIds.add(videoId);
}
const videoIdList = [...videoIds];
for (const sessionId of selectIds(
db,
(placeholders) => `SELECT session_id FROM imm_sessions WHERE video_id IN (${placeholders})`,
videoIdList,
'session_id',
)) {
sessionIds.add(sessionId);
}
const sessionIdList = [...sessionIds];
const lexicalRemovals = planLexicalRemovalsInChunks(db, sessionIdList);
const affectedRollupGroups = sessionIdList
.flatMap((_, index) =>
index % SQLITE_ID_CHUNK_SIZE === 0
? getRollupGroupsForSessions(db, sessionIdList.slice(index, index + SQLITE_ID_CHUNK_SIZE))
: [],
)
.filter((group) => !videoIds.has(group.videoId));
const coverBlobHashes = new Set<string>();
if (videoIdList.length > 0) {
forEachIdChunk(videoIdList, (chunk) => {
const placeholders = makePlaceholders(chunk);
const artRows = db
.prepare(
`SELECT cover_blob_hash AS coverBlobHash
FROM imm_media_art
WHERE video_id IN (${placeholders}) AND cover_blob_hash IS NOT NULL`,
)
.all(...chunk) as Array<{ coverBlobHash: string }>;
for (const row of artRows) coverBlobHashes.add(row.coverBlobHash);
});
deleteSessionsByIds(db, sessionIdList);
forEachIdChunk(videoIdList, (chunk) => {
const placeholders = makePlaceholders(chunk);
db.prepare(`DELETE FROM imm_subtitle_lines WHERE video_id IN (${placeholders})`).run(
...chunk,
);
db.prepare(`DELETE FROM imm_daily_rollups WHERE video_id IN (${placeholders})`).run(
...chunk,
);
db.prepare(`DELETE FROM imm_monthly_rollups WHERE video_id IN (${placeholders})`).run(
...chunk,
);
db.prepare(`DELETE FROM imm_media_art WHERE video_id IN (${placeholders})`).run(...chunk);
db.prepare(`DELETE FROM imm_videos WHERE video_id IN (${placeholders})`).run(...chunk);
});
} else {
deleteSessionsByIds(db, sessionIdList);
}
for (const coverBlobHash of coverBlobHashes) {
cleanupUnusedCoverArtBlobHash(db, coverBlobHash);
}
if (animeIdList.length > 0) {
forEachIdChunk(animeIdList, (chunk) => {
const placeholders = makePlaceholders(chunk);
db.prepare(`DELETE FROM imm_lifetime_anime WHERE anime_id IN (${placeholders})`).run(
...chunk,
);
db.prepare(`DELETE FROM imm_anime WHERE anime_id IN (${placeholders})`).run(...chunk);
});
}
applyLexicalRemovals(db, lexicalRemovals);
rebuildLifetimeSummariesInTransaction(db);
refreshRollupsForGroupsInTransaction(db, affectedRollupGroups);
db.exec('COMMIT');
} catch (error) {
db.exec('ROLLBACK');
throw error;
}
}
@@ -1,7 +1,10 @@
import { createHash } from 'node:crypto';
import type { DatabaseSync } from './sqlite';
import { buildCoverBlobReference, normalizeCoverBlobBytes } from './storage';
import { rebuildLifetimeSummaries, rebuildLifetimeSummariesInTransaction } from './lifetime';
import {
recomputeLifetimeAnimeAggregates,
rebuildLifetimeSummariesInTransaction,
} from './lifetime';
import { getRollupGroupsForSessions, refreshRollupsForGroupsInTransaction } from './maintenance';
import { nowMs } from './time';
import { resolveAnimeAnilistConflict } from './anime-season-repair';
@@ -418,6 +421,7 @@ export function updateAnimeAnilistInfo(
titleEnglish: string | null;
titleNative: string | null;
episodesTotal: number | null;
exactTitleMatch?: boolean;
},
): void {
const row = db.prepare('SELECT anime_id FROM imm_videos WHERE video_id = ?').get(videoId) as {
@@ -425,7 +429,11 @@ export function updateAnimeAnilistInfo(
} | null;
if (!row?.anime_id) return;
const repair = resolveAnimeAnilistConflict(db, row.anime_id, info.anilistId);
const repair = resolveAnimeAnilistConflict(db, row.anime_id, info.anilistId, {
matchConfidence:
info.exactTitleMatch === true ? 'exact' : info.exactTitleMatch === false ? 'weak' : undefined,
});
if (repair.mergeRecommended || repair.anilistAssignmentBlocked) return;
const targetRow = db
.prepare('SELECT anime_id FROM imm_videos WHERE video_id = ?')
.get(videoId) as {
@@ -455,7 +463,7 @@ export function updateAnimeAnilistInfo(
targetRow.anime_id,
);
if (repair.movedVideos > 0 || repair.deletedAnimeRows > 0) {
rebuildLifetimeSummaries(db);
recomputeLifetimeAnimeAggregates(db);
}
}
@@ -80,6 +80,14 @@ export function makePlaceholders(values: number[]): string {
return values.map(() => '?').join(',');
}
export const SQLITE_ID_CHUNK_SIZE = 1_000;
export function forEachIdChunk(ids: number[], callback: (chunk: number[]) => void): void {
for (let start = 0; start < ids.length; start += SQLITE_ID_CHUNK_SIZE) {
callback(ids.slice(start, start + SQLITE_ID_CHUNK_SIZE));
}
}
export function resolvedCoverBlobExpr(mediaAlias: string, blobStoreAlias: string): string {
return `COALESCE(${blobStoreAlias}.cover_blob, CASE WHEN ${mediaAlias}.cover_blob_hash IS NULL THEN ${mediaAlias}.cover_blob ELSE NULL END)`;
}
@@ -490,17 +498,19 @@ export function deleteSessionsByIds(db: DatabaseSync, sessionIds: number[]): voi
return;
}
const placeholders = makePlaceholders(sessionIds);
forEachIdChunk(sessionIds, (chunk) => {
const placeholders = makePlaceholders(chunk);
db.prepare(`DELETE FROM imm_subtitle_lines WHERE session_id IN (${placeholders})`).run(
...sessionIds,
...chunk,
);
db.prepare(`DELETE FROM imm_session_telemetry WHERE session_id IN (${placeholders})`).run(
...sessionIds,
...chunk,
);
db.prepare(`DELETE FROM imm_session_events WHERE session_id IN (${placeholders})`).run(
...sessionIds,
...chunk,
);
db.prepare(`DELETE FROM imm_sessions WHERE session_id IN (${placeholders})`).run(...sessionIds);
db.prepare(`DELETE FROM imm_sessions WHERE session_id IN (${placeholders})`).run(...chunk);
});
}
export function toDbMs(ms: number | bigint): bigint {
@@ -20,6 +20,7 @@ import {
} from './storage';
import {
EVENT_SUBTITLE_LINE,
SCHEMA_VERSION,
SESSION_STATUS_ENDED,
SOURCE_TYPE_LOCAL,
SOURCE_TYPE_REMOTE,
@@ -132,6 +133,7 @@ test('ensureSchema creates immersion core tables', () => {
assert.ok(videoColumns.has('parser_source'));
assert.ok(videoColumns.has('parser_confidence'));
assert.ok(videoColumns.has('parse_metadata_json'));
assert.ok(videoColumns.has('anime_assignment_locked'));
const mediaArtColumns = new Set(
(
@@ -155,6 +157,33 @@ test('ensureSchema creates immersion core tables', () => {
}
});
test('ensureSchema adds manual assignment locks when upgrading the previous schema', () => {
const dbPath = makeDbPath();
const db = new Database(dbPath);
try {
ensureSchema(db);
db.exec('ALTER TABLE imm_videos DROP COLUMN anime_assignment_locked');
db.prepare('UPDATE imm_schema_version SET schema_version = ?').run(SCHEMA_VERSION - 1);
ensureSchema(db);
const columns = new Set(
(db.prepare('PRAGMA table_info(imm_videos)').all() as Array<{ name: string }>).map(
(row) => row.name,
),
);
assert.ok(columns.has('anime_assignment_locked'));
const version = db
.prepare('SELECT MAX(schema_version) AS version FROM imm_schema_version')
.get() as { version: number };
assert.equal(version.version, SCHEMA_VERSION);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('stats excluded words are replaced and read from sqlite storage', () => {
const dbPath = makeDbPath();
const db = new Database(dbPath);
@@ -807,6 +836,7 @@ test('ensureSchema migrates legacy videos and backfills anime metadata from file
assert.ok(videoColumns.has('parser_source'));
assert.ok(videoColumns.has('parser_confidence'));
assert.ok(videoColumns.has('parse_metadata_json'));
assert.ok(videoColumns.has('anime_assignment_locked'));
const animeRows = db
.prepare('SELECT canonical_title FROM imm_anime ORDER BY canonical_title')
+115 -11
View File
@@ -1,5 +1,7 @@
import { createHash } from 'node:crypto';
import path from 'node:path';
import { parseMediaInfo } from '../../../jimaku/utils';
import { normalizeTitleIdentity } from '../../utils/title-normalization';
import type { DatabaseSync } from './sqlite';
import { nowMs } from './time';
import { SCHEMA_VERSION } from './types';
@@ -319,14 +321,7 @@ export function applyPragmas(db: DatabaseSync): void {
db.exec(`PRAGMA journal_size_limit = ${WAL_JOURNAL_SIZE_LIMIT_BYTES}`);
}
export function normalizeAnimeIdentityKey(title: string): string {
return title
.normalize('NFKC')
.toLowerCase()
.replace(/[^\p{L}\p{N}]+/gu, ' ')
.trim()
.replace(/\s+/g, ' ');
}
export const normalizeAnimeIdentityKey = normalizeTitleIdentity;
function normalizeSeasonScope(value: number | null | undefined): number | null {
if (typeof value !== 'number' || !Number.isSafeInteger(value) || value <= 0) {
@@ -530,6 +525,36 @@ function ensureStatsExcludedWordsTable(db: DatabaseSync): void {
`);
}
function ensureAnimeMergeTables(db: DatabaseSync): void {
db.exec(`
CREATE TABLE IF NOT EXISTS imm_anime_title_aliases(
normalized_title_key TEXT PRIMARY KEY,
anime_id INTEGER NOT NULL,
CREATED_DATE TEXT,
LAST_UPDATE_DATE TEXT,
FOREIGN KEY(anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE
);
CREATE INDEX IF NOT EXISTS idx_anime_title_aliases_anime_id
ON imm_anime_title_aliases(anime_id);
CREATE TABLE IF NOT EXISTS imm_anime_merge_recommendations(
recommendation_id INTEGER PRIMARY KEY AUTOINCREMENT,
first_anime_id INTEGER NOT NULL,
second_anime_id INTEGER NOT NULL,
anilist_id INTEGER NOT NULL,
status TEXT NOT NULL DEFAULT 'pending' CHECK(status IN ('pending', 'dismissed')),
CREATED_DATE TEXT,
LAST_UPDATE_DATE TEXT,
CHECK(first_anime_id < second_anime_id),
UNIQUE(first_anime_id, second_anime_id, anilist_id),
FOREIGN KEY(first_anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE,
FOREIGN KEY(second_anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE
);
CREATE INDEX IF NOT EXISTS idx_anime_merge_recommendations_status
ON imm_anime_merge_recommendations(status, recommendation_id);
`);
}
export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput): number {
const seasonScope = normalizeSeasonScope(input.seasonScope);
const identityTitle = buildSeasonScopedAnimeTitle(input.parsedTitle, seasonScope);
@@ -550,8 +575,14 @@ export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput
const byNormalizedTitle = db
.prepare('SELECT anime_id FROM imm_anime WHERE normalized_title_key = ?')
.get(normalizedTitleKey) as { anime_id: number } | null;
const existing = byAnilistId ?? byNormalizedTitle;
const byTitleAlias = db
.prepare('SELECT anime_id FROM imm_anime_title_aliases WHERE normalized_title_key = ?')
.get(normalizedTitleKey) as { anime_id: number } | null;
const existing = byAnilistId ?? byNormalizedTitle ?? byTitleAlias;
if (existing?.anime_id) {
// An alias remembers an intentionally merged-away spelling. Reusing it
// must not rename the survivor back to that discarded display title.
const canonicalTitleUpdate = byAnilistId || byNormalizedTitle ? canonicalTitle : null;
db.prepare(
`
UPDATE imm_anime
@@ -566,7 +597,7 @@ export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput
WHERE anime_id = ?
`,
).run(
canonicalTitle,
canonicalTitleUpdate,
input.anilistId,
input.titleRomaji,
input.titleEnglish,
@@ -618,7 +649,10 @@ export function linkVideoToAnimeRecord(
`
UPDATE imm_videos
SET
anime_id = ?,
anime_id = CASE
WHEN anime_assignment_locked = 1 THEN anime_id
ELSE ?
END,
parsed_basename = ?,
parsed_title = ?,
parsed_season = ?,
@@ -643,6 +677,67 @@ export function linkVideoToAnimeRecord(
);
}
export function getManualAnimeAssignment(db: DatabaseSync, videoId: number): number | null {
const row = db
.prepare(
`
SELECT anime_id AS animeId
FROM imm_videos
WHERE video_id = ?
AND anime_assignment_locked = 1
`,
)
.get(videoId) as { animeId: number | null } | null;
return row?.animeId ?? null;
}
/**
* A manual correction in the same folder is a useful grouping hint, but only
* when every season-compatible correction agrees on the destination.
*/
export function findManualDirectoryAnimeAssignment(
db: DatabaseSync,
videoId: number,
mediaPath: string,
parsedSeason: number | null,
): number | null {
const directory = path.dirname(path.resolve(mediaPath));
const rows = db
.prepare(
`
SELECT
anime_id AS animeId,
source_path AS sourcePath,
parsed_season AS parsedSeason
FROM imm_videos
WHERE video_id != ?
AND anime_assignment_locked = 1
AND anime_id IS NOT NULL
AND source_path IS NOT NULL
`,
)
.all(videoId) as Array<{
animeId: number;
sourcePath: string;
parsedSeason: number | null;
}>;
const candidates = new Set<number>();
for (const row of rows) {
if (path.dirname(path.resolve(row.sourcePath)) !== directory) {
continue;
}
if (parsedSeason !== null && row.parsedSeason !== null && parsedSeason !== row.parsedSeason) {
continue;
}
candidates.add(row.animeId);
if (candidates.size > 1) {
return null;
}
}
return candidates.values().next().value ?? null;
}
export function linkYoutubeVideoToAnimeRecord(
db: DatabaseSync,
videoId: number,
@@ -751,6 +846,7 @@ export function ensureSchema(db: DatabaseSync): void {
if (currentVersion?.schema_version === SCHEMA_VERSION) {
ensureLifetimeSummaryTables(db);
ensureStatsExcludedWordsTable(db);
ensureAnimeMergeTables(db);
return;
}
@@ -786,6 +882,7 @@ export function ensureSchema(db: DatabaseSync): void {
parser_source TEXT,
parser_confidence REAL,
parse_metadata_json TEXT,
anime_assignment_locked INTEGER NOT NULL DEFAULT 0 CHECK(anime_assignment_locked IN (0, 1)),
watched INTEGER NOT NULL DEFAULT 0,
duration_ms INTEGER NOT NULL CHECK(duration_ms>=0),
file_size_bytes INTEGER CHECK(file_size_bytes>=0),
@@ -799,6 +896,13 @@ export function ensureSchema(db: DatabaseSync): void {
FOREIGN KEY(anime_id) REFERENCES imm_anime(anime_id) ON DELETE SET NULL
);
`);
addColumnIfMissing(
db,
'imm_videos',
'anime_assignment_locked',
'INTEGER NOT NULL DEFAULT 0 CHECK(anime_assignment_locked IN (0, 1))',
);
ensureAnimeMergeTables(db);
db.exec(`
CREATE TABLE IF NOT EXISTS imm_sessions(
session_id INTEGER PRIMARY KEY AUTOINCREMENT,
+1 -1
View File
@@ -1,4 +1,4 @@
export const SCHEMA_VERSION = 19;
export const SCHEMA_VERSION = 21;
export const DEFAULT_QUEUE_CAP = 1_000;
export const DEFAULT_BATCH_SIZE = 25;
export const DEFAULT_FLUSH_INTERVAL_MS = 500;
+15
View File
@@ -1,6 +1,7 @@
import electron from 'electron';
import type { BrowserWindow as ElectronBrowserWindow, IpcMainEvent } from 'electron';
import type {
ChangelogSnapshot,
CompiledSessionBinding,
ControllerConfigUpdate,
PlaylistBrowserMutationResult,
@@ -122,6 +123,7 @@ export interface IpcServiceDeps {
removeCharacterDictionaryManagedEntry?: (mediaId: number) => Promise<unknown>;
moveCharacterDictionaryManagedEntry?: (mediaId: number, direction: 1 | -1) => Promise<unknown>;
appendClipboardVideoToQueue: () => { ok: boolean; message: string };
getChangelogSnapshot?: (options?: { refresh?: boolean }) => Promise<ChangelogSnapshot>;
getPlaylistBrowserSnapshot: () => Promise<PlaylistBrowserSnapshot>;
appendPlaylistBrowserFile: (filePath: string) => Promise<PlaylistBrowserMutationResult>;
playPlaylistBrowserIndex: (index: number) => Promise<PlaylistBrowserMutationResult>;
@@ -297,6 +299,7 @@ export interface IpcDepsRuntimeOptions {
removeCharacterDictionaryManagedEntry?: (mediaId: number) => Promise<unknown>;
moveCharacterDictionaryManagedEntry?: (mediaId: number, direction: 1 | -1) => Promise<unknown>;
appendClipboardVideoToQueue: () => { ok: boolean; message: string };
getChangelogSnapshot?: (options?: { refresh?: boolean }) => Promise<ChangelogSnapshot>;
getPlaylistBrowserSnapshot: () => Promise<PlaylistBrowserSnapshot>;
appendPlaylistBrowserFile: (filePath: string) => Promise<PlaylistBrowserMutationResult>;
playPlaylistBrowserIndex: (index: number) => Promise<PlaylistBrowserMutationResult>;
@@ -418,6 +421,7 @@ export function createIpcDepsRuntime(options: IpcDepsRuntimeOptions): IpcService
entries: [],
})),
appendClipboardVideoToQueue: options.appendClipboardVideoToQueue,
getChangelogSnapshot: options.getChangelogSnapshot,
getPlaylistBrowserSnapshot: options.getPlaylistBrowserSnapshot,
appendPlaylistBrowserFile: options.appendPlaylistBrowserFile,
playPlaylistBrowserIndex: options.playPlaylistBrowserIndex,
@@ -820,6 +824,17 @@ export function registerIpcHandlers(deps: IpcServiceDeps, ipc: IpcMainRegistrar
return deps.appendClipboardVideoToQueue();
});
ipc.handle(IPC_CHANNELS.request.getChangelogSnapshot, async (_event, payload: unknown) => {
const refresh =
typeof payload === 'object' && payload !== null && 'refresh' in payload
? (payload as { refresh?: unknown }).refresh === true
: false;
if (!deps.getChangelogSnapshot) {
throw new Error('Changelog service is unavailable.');
}
return await deps.getChangelogSnapshot({ refresh });
});
ipc.handle(IPC_CHANNELS.request.getPlaylistBrowserSnapshot, async () => {
return await deps.getPlaylistBrowserSnapshot();
});
@@ -1,5 +1,6 @@
import type { Hono } from 'hono';
import { statsJson } from '../../../types/stats-http-contract.js';
import { UNKNOWN_MOVE_TARGET_MESSAGE } from '../immersion-tracker/anime-merge.js';
import type { ImmersionTrackerService } from '../immersion-tracker-service.js';
import {
buildSentenceSearchOptions,
@@ -7,6 +8,7 @@ import {
parseBooleanQuery,
parseExcludedWordsBody,
parseIntQuery,
parsePositiveIdList,
} from './route-support.js';
export function registerStatsLibraryRoutes(
@@ -130,6 +132,19 @@ export function registerStatsLibraryRoutes(
return c.json(statsJson('animeLibrary', rows));
});
app.get('/api/stats/anime/merge-recommendations', async (c) => {
const recommendations = await tracker.getAnimeMergeRecommendations();
return c.json(statsJson('animeMergeRecommendations', { recommendations }));
});
app.delete('/api/stats/anime/merge-recommendations/:recommendationId', async (c) => {
const recommendationId = parseIntQuery(c.req.param('recommendationId'), 0);
if (recommendationId <= 0) return c.body(null, 400);
const dismissed = await tracker.dismissAnimeMergeRecommendation(recommendationId);
if (!dismissed) return c.body(null, 404);
return c.json(statsJson('dismissAnimeMergeRecommendation', { ok: true }));
});
app.get('/api/stats/anime/:animeId', async (c) => {
const animeId = parseIntQuery(c.req.param('animeId'), 0);
if (animeId <= 0) return c.body(null, 400);
@@ -197,4 +212,50 @@ export function registerStatsLibraryRoutes(
await tracker.deleteAnime(animeId);
return c.json(statsJson('deleteAnime', { ok: true }));
});
app.post('/api/stats/anime/:animeId/merge', async (c) => {
const animeId = parseIntQuery(c.req.param('animeId'), 0);
if (animeId <= 0) return c.body(null, 400);
const body = await c.req.json().catch(() => null);
const sourceAnimeIds = parsePositiveIdList(body?.sourceAnimeIds).filter((id) => id !== animeId);
if (sourceAnimeIds.length === 0) return c.body(null, 400);
const summary = await tracker.mergeAnime(animeId, sourceAnimeIds);
// Nothing folded means the target or every source was already gone, so the
// caller should not be told the merge succeeded.
if (summary.mergedAnimeIds.length === 0) return c.body(null, 404);
return c.json(
statsJson('mergeAnime', {
ok: true,
animeId: summary.survivingAnimeId,
mergedAnimeIds: summary.mergedAnimeIds,
movedVideos: summary.movedVideos,
}),
);
});
app.patch('/api/stats/media/:videoId/anime', async (c) => {
const videoId = parseIntQuery(c.req.param('videoId'), 0);
if (videoId <= 0) return c.body(null, 400);
const body = await c.req.json().catch(() => null);
const animeId = Number.isSafeInteger(body?.animeId) ? (body.animeId as number) : 0;
if (animeId <= 0) return c.body(null, 400);
try {
const summary = await tracker.moveVideoToAnime(videoId, animeId);
return c.json(
statsJson('moveVideoToAnime', {
ok: true,
animeId: summary.targetAnimeId,
previousAnimeId: summary.previousAnimeId,
removedPreviousAnime: summary.removedPreviousAnime,
}),
);
} catch (error) {
// Only a missing episode or entry is a 404; storage failures must not be
// reported to the caller as "not found".
if (error instanceof Error && error.message === UNKNOWN_MOVE_TARGET_MESSAGE) {
return c.body(null, 404);
}
throw error;
}
});
}
@@ -170,6 +170,18 @@ export async function enrichSessionsWithKnownWordMetrics<
);
}
/** Deduplicated positive integer ids from an untrusted JSON body field. */
export function parsePositiveIdList(raw: unknown): number[] {
if (!Array.isArray(raw)) return [];
const ids = new Set<number>();
for (const value of raw) {
if (Number.isSafeInteger(value) && (value as number) > 0) {
ids.add(value as number);
}
}
return [...ids];
}
export function parseBooleanQuery(raw: string | undefined, fallback: boolean): boolean {
if (raw === undefined) return fallback;
const normalized = raw.trim().toLowerCase();
@@ -28,6 +28,7 @@ const VIDEO_COPY_COLUMNS = [
'parser_source',
'parser_confidence',
'parse_metadata_json',
'anime_assignment_locked',
'watched',
'duration_ms',
'file_size_bytes',
+191
View File
@@ -0,0 +1,191 @@
/*
* Duplicate/animation-burst collapsing for parsed subtitle cues.
*
* Split out of the cue parser so the parsing rules and the "is this run one animation?"
* heuristics can be read -- and tested -- on their own. The parser owns the cue shape;
* this module only decides which cues survive.
*/
import { hasAssTemporalOverride, isAnimatedAssEffectKind } from './ass-text';
import type {
AnnotatedSubtitleCue,
SubtitleCue,
SubtitleSourceFormat,
} from './subtitle-cue-parser';
// Back-to-back frames of the same animation are authored flush against each other; a
// tiny tolerance absorbs the centisecond rounding of the ASS timestamp format.
const DUPLICATE_CUE_GAP_TOLERANCE_SECONDS = 0.05;
// A burst is a *sequence*. Two adjacent events are two events, not an animation --
// characters do repeat each other, and a repeated line can legitimately be short.
const MIN_BURST_EVENTS = 3;
// Real dialogue holds on screen for about a second, so a run with a couple of much
// shorter events among them looks like frames. Used only alongside authoring evidence.
const ANIMATION_FRAME_MAX_SECONDS = 0.3;
// A karaoke run usually ends on a long "hold" frame, so not every event is short.
const MIN_TAGGED_BURST_FRAMES = 2;
// SRT and VTT carry no authoring metadata at all, so timing is the only signal available
// -- which makes it the easiest one to get wrong. ASS->SRT conversion leaves frames at
// ~0.04s, well under any real utterance, and a burst leaves many of them behind. Both
// bounds are deliberately far stricter than the ASS path: a run of ordinary short lines
// (`えっ` traded between characters) must not clear them.
const TIMING_ONLY_FRAME_MAX_SECONDS = 0.1;
const MIN_TIMING_ONLY_FRAMES = 5;
function cueKey(cue: SubtitleCue): string {
return `${cue.startTime}|${cue.endTime}|${cue.text}`;
}
/**
* Identical text over an identical span is redundant however it was authored -- most
* often a layered ASS event stacking a shadow copy under the visible one.
*/
function collapseExactDuplicates(cues: AnnotatedSubtitleCue[]): AnnotatedSubtitleCue[] {
const seen = new Set<string>();
return cues.filter((cue) => {
const key = cueKey(cue);
if (seen.has(key)) {
return false;
}
seen.add(key);
return true;
});
}
function countFramesShorterThan(run: AnnotatedSubtitleCue[], maxSeconds: number): number {
return run.filter((cue) => cue.endTime - cue.startTime < maxSeconds).length;
}
/**
* Evidence that a run of ASS events is one animation rather than several authored lines.
* A static tag says nothing on its own -- three events sharing one `\clip(...)` are three
* signs -- so the tag has to be temporal by nature (`\t`, `\move`, karaoke timing, or
* anything wrapped in `\t(...)`), an animated `Effect` column, or a value that actually
* changes from event to event, which is how per-frame typesetting is authored.
*/
export function hasAssAnimationEvidence(run: AnnotatedSubtitleCue[]): boolean {
if (run.every((cue) => hasAssTemporalOverride(cue.overrides))) {
return true;
}
if (run.every((cue) => isAnimatedAssEffectKind(cue.effectKind))) {
return true;
}
const [first] = run;
const everyEventTypeset = run.every((cue) => cue.overrides.length > 0);
const signatureChanges = run.some((cue) => cue.overrideSignature !== first!.overrideSignature);
return everyEventTypeset && signatureChanges;
}
export function isAnimationBurst(
run: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): boolean {
if (run.length < MIN_BURST_EVENTS) {
return false;
}
if (format === 'srt') {
return (
run.length >= MIN_TIMING_ONLY_FRAMES &&
countFramesShorterThan(run, TIMING_ONLY_FRAME_MAX_SECONDS) === run.length
);
}
if (countFramesShorterThan(run, ANIMATION_FRAME_MAX_SECONDS) < MIN_TAGGED_BURST_FRAMES) {
return false;
}
// One animation belongs to one styled, one named source line. Two characters trading
// the same short word are two styles or two actors, and never merge.
const [first] = run;
if (run.some((cue) => cue.style !== first!.style || cue.name !== first!.name)) {
return false;
}
return hasAssAnimationEvidence(run);
}
/**
* Karaoke and sign typesetting emits one Dialogue event per animation frame, all carrying
* the same visible text over a contiguous span. Collapse each such run into a single cue.
*
* Only runs that look like animation collapse. Two ordinary lines that happen to repeat
* -- several characters each saying `おはよう` in turn, a positioned sign redrawn with a
* different fade -- stay separate, because merging them would destroy real mineable lines.
*/
function collapseAnimationBursts(
cues: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): AnnotatedSubtitleCue[] {
const indicesByText = new Map<string, number[]>();
cues.forEach((cue, index) => {
const bucket = indicesByText.get(cue.text);
if (bucket) {
bucket.push(index);
} else {
indicesByText.set(cue.text, [index]);
}
});
const dropped = new Set<number>();
const extendedEnd = new Map<number, number>();
for (const indices of indicesByText.values()) {
if (indices.length < MIN_BURST_EVENTS) {
continue;
}
let runStart = 0;
while (runStart < indices.length) {
let runEnd = runStart;
let chainEnd = cues[indices[runStart]!]!.endTime;
while (runEnd + 1 < indices.length) {
const next = cues[indices[runEnd + 1]!]!;
if (next.startTime > chainEnd + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS) {
break;
}
chainEnd = Math.max(chainEnd, next.endTime);
runEnd += 1;
}
const run = indices.slice(runStart, runEnd + 1).map((index) => cues[index]!);
if (isAnimationBurst(run, format)) {
for (let i = runStart + 1; i <= runEnd; i += 1) {
dropped.add(indices[i]!);
}
extendedEnd.set(indices[runStart]!, chainEnd);
}
runStart = runEnd + 1;
}
}
if (dropped.size === 0) {
return cues;
}
const merged: AnnotatedSubtitleCue[] = [];
cues.forEach((cue, index) => {
if (dropped.has(index)) {
return;
}
const end = extendedEnd.get(index);
merged.push(end !== undefined && end > cue.endTime ? { ...cue, endTime: end } : cue);
});
return merged;
}
/**
* Collapse redundant cues. Input must already be sorted by non-decreasing `startTime`,
* ties broken by `endTime` then source `order` -- burst detection chains events by
* comparing each one against the running end of the events before it, so an unsorted
* list breaks runs apart and leaves the frames behind.
*/
export function mergeDuplicateCues(
cues: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): AnnotatedSubtitleCue[] {
return collapseAnimationBursts(collapseExactDuplicates(cues), format);
}
+393 -3
View File
@@ -91,6 +91,17 @@ test('parseSrtCues skips malformed timing lines gracefully', () => {
assert.equal(cues[0]!.text, '有効');
});
test('parseSubtitleCues strips complete brace blocks from SRT and VTT text', () => {
const content = ['1', '00:00:01,000 --> 00:00:02,000', '彼は{謎}と言った', ''].join('\n');
for (const filename of ['test.srt', 'test.vtt']) {
const cues = parseSubtitleCues(content, filename);
assert.equal(cues.length, 1, filename);
assert.equal(cues[0]!.text, '彼はと言った', filename);
}
});
test('parseAssCues parses basic ASS dialogue lines', () => {
const content = [
'[Script Info]',
@@ -137,7 +148,9 @@ test('parseAssCues handles text containing commas', () => {
assert.equal(cues[0]!.text, 'はい、そうです、ね');
});
test('parseAssCues handles \\N line breaks', () => {
test('parseAssCues decodes \\N line breaks into real newlines', () => {
// ASS is decoded once, here at ingestion, so cue text matches what mpv hands over for
// the same line played live.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
@@ -146,7 +159,7 @@ test('parseAssCues handles \\N line breaks', () => {
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '一行目\\N二行目');
assert.equal(cues[0]!.text, '一行目\n二行目');
});
test('parseAssCues strips HTML-like markup while preserving ASS line breaks', () => {
@@ -158,7 +171,46 @@ test('parseAssCues strips HTML-like markup while preserving ASS line breaks', ()
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '一行目\\N二行目');
assert.equal(cues[0]!.text, '一行目\n二行目');
});
test('parseAssCues drops vector drawing runs enabled by \\p', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 1,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\an5\\pos(730,1042)\\p1\\blur1}m 20 0 b 10 0 0 10 0 20 b 0 31 10 40 20 40 {\\p0}',
'Dialogue: 0,0:00:05.00,0:00:08.00,Default,,0,0,0,,これは字幕',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues.length, 1);
assert.equal(cues[0]!.text, 'これは字幕');
});
test('parseAssCues keeps text that follows a \\p0 reset on the same line', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\p1}m 0 0 l 10 10{\\p0}本文{\\p1}m 5 5 l 6 6{\\p0}続き',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues.length, 1);
assert.equal(cues[0]!.text, '本文続き');
});
test('parseAssCues leaves \\pos untouched when no drawing mode is active', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\pos(960,1068)\\bord3}位置指定',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '位置指定');
});
test('parseAssCues returns empty for content without Events section', () => {
@@ -258,6 +310,344 @@ test('parseSubtitleCues returns cues sorted by start time', () => {
assert.equal(cues[1]!.text, '二番目');
});
test('parseSubtitleCues collapses per-frame karaoke duplicates into one cue', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}過ぎ去ってしまう瞬間を',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}過ぎ去ってしまう瞬間を',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,,{\\clip(m 3 3)}過ぎ去ってしまう瞬間を',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[0]!.endTime, 3.55);
assert.equal(cues[0]!.text, '過ぎ去ってしまう瞬間を');
});
test('parseSubtitleCues keeps back-to-back plain dialogue repeats separate', () => {
// Several characters greeting in turn: distinct utterances that happen to abut.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:04:05.67,0:04:06.82,Dial_JP,,0,0,0,,おはよう',
'Dialogue: 0,0:04:06.82,0:04:07.56,Dial_JP,,0,0,0,,おはよう',
'Dialogue: 0,0:04:07.56,0:04:08.78,Dial_JP,,0,0,0,,おはよう',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
assert.equal(cues[0]!.endTime, 246.82);
assert.equal(cues[2]!.startTime, 247.56);
});
test('parseSubtitleCues collapses exact duplicate cues even without effect tags', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,重なった行',
'Dialogue: 1,0:00:01.00,0:00:04.00,Default,,0,0,0,,重なった行',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
});
test('parseSubtitleCues collapses tag-less animation frames in converted SRT', () => {
// ASS -> SRT conversion drops override tags, so only the ~0.04s frame timing remains.
const lines = ['1', '00:00:07,870 --> 00:00:07,910', 'Kaguya Wants to be Confessed to', ''];
for (let i = 1; i < 8; i++) {
const start = 7910 + (i - 1) * 40;
const end = start + 40;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(end)}`, 'Kaguya Wants to be Confessed to', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 7.87);
});
test('parseSubtitleCues keeps identical lines that recur far apart', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:02.00,Default,,0,0,0,,なんで',
'Dialogue: 0,0:05:00.00,0:05:01.00,Default,,0,0,0,,なんで',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[1]!.startTime, 300.0);
});
test('parseSubtitleCues keeps two positioned signs that repeat the same text', () => {
// Both carry override tags, but `\pos` and `\fad` are static placement, not animation.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:01:00.00,0:01:03.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(200,200)}第一話',
'Dialogue: 0,0:01:03.00,0:01:06.00,Sign,,0,0,0,,{\\pos(960,900)\\fad(200,200)}第一話',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[1]!.startTime, 63.0);
});
test('parseSubtitleCues keeps a run of ordinary positioned lines separate', () => {
// Three events is a sequence, but none of them runs at animation-frame speed.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:01:00.00,0:01:02.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
'Dialogue: 0,0:01:02.00,0:01:04.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
'Dialogue: 0,0:01:04.00,0:01:06.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps a short repeated SRT pair without burst evidence', () => {
const content = [
'1',
'00:00:01,000 --> 00:00:01,200',
'えっ',
'',
'2',
'00:00:01,200 --> 00:00:01,400',
'えっ',
'',
].join('\n');
const cues = parseSubtitleCues(content, 'test.srt');
assert.equal(cues.length, 2);
});
test('parseSubtitleCues collapses a burst marked only by the Effect column', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,Karaoke,歌詞',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,Karaoke,歌詞',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,Karaoke,歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.endTime, 3.55);
});
test('parseSubtitleCues keeps a second karaoke burst that starts after a gap', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}リフレイン',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}リフレイン',
'Dialogue: 0,0:00:01.09,0:00:03.00,OP_JP,,0,0,0,,{\\clip(m 3 3)}リフレイン',
'Dialogue: 0,0:00:20.00,0:00:20.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}リフレイン',
'Dialogue: 0,0:00:20.05,0:00:20.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}リフレイン',
'Dialogue: 0,0:00:20.09,0:00:22.00,OP_JP,,0,0,0,,{\\clip(m 3 3)}リフレイン',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[0]!.endTime, 3.0);
assert.equal(cues[1]!.startTime, 20.0);
assert.equal(cues[1]!.endTime, 22.0);
});
test('parseSubtitleCues does not merge a burst into unrelated dialogue between frames', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}歌詞',
'Dialogue: 0,0:00:01.02,0:00:03.00,Dial_JP,,0,0,0,,別のセリフ',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}歌詞',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,,{\\clip(m 3 3)}歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.deepEqual(
cues.map((cue) => cue.text),
['歌詞', '別のセリフ'],
);
assert.equal(cues[0]!.endTime, 3.55);
});
test('parseSubtitleCues keeps rapid ASS lines from different actors separate', () => {
// Three 200ms `えっ` reactions traded between characters. Fast, adjacent and identical,
// but authored as three lines: different styles and different actors.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_A,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_B,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_C,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues reads the speaker column when it is spelled Actor', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Actor, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_JP,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_JP,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_JP,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues does not treat a custom Effect name as animation', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,scrolling-credit,制作',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,scrolling-credit,制作',
'Dialogue: 0,0:00:01.40,0:00:01.60,Sign,,0,0,0,scrolling-credit,制作',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps rapid ASS lines that share a style but not an actor', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_JP,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_JP,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_JP,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps untagged rapid ASS repeats separate', () => {
// No overrides at all: timing-only evidence is an SRT/VTT fallback and must not apply
// to ASS, where the absence of typesetting is itself evidence of plain dialogue.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.05,0:00:01.10,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.10,0:00:01.15,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.15,0:00:01.20,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.25,Dial_JP,,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 5);
});
test('parseSubtitleCues keeps repeated signs sharing one static clip', () => {
// `\clip` is a static shape for the event. Three events with the identical clip were
// typeset the same way, so none of them is a frame of the others.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
'Dialogue: 0,0:00:01.40,0:00:01.60,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues collapses a sign animated through \\t', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
'Dialogue: 0,0:00:01.40,0:00:03.00,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.endTime, 3.0);
});
test('parseSubtitleCues keeps a short repeated SRT run above the frame threshold', () => {
// Five contiguous 200ms cues: a sequence, but nowhere near animation-frame speed.
const lines: string[] = [];
for (let i = 0; i < 5; i++) {
const start = 1000 + i * 200;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(start + 200)}`, 'えっ', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 5);
});
test('parseSubtitleCues keeps a short SRT frame run below the minimum length', () => {
// Four 40ms frames: frame-speed, but too few to tell an animation from an artefact.
const lines: string[] = [];
for (let i = 0; i < 4; i++) {
const start = 7870 + i * 40;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(start + 40)}`, 'タイトル', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 4);
});
test('parseSubtitleCues applies ASS burst rules to ASS content behind an .srt filename', () => {
// The extension lies, so the SRT parser finds nothing and the content-sniffing fallback
// takes over -- which has to carry the `ass` source format with it, or the far stricter
// timing-only thresholds would let this karaoke burst through as three cues.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Karaoke,,0,0,0,,{\\k20}歌詞',
'Dialogue: 0,0:00:01.20,0:00:01.40,Karaoke,,0,0,0,,{\\k20}歌詞',
'Dialogue: 0,0:00:01.40,0:00:03.00,Karaoke,,0,0,0,,{\\k20}歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.srt');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[0]!.endTime, 3.0);
assert.equal(cues[0]!.text, '歌詞');
});
test('parseSubtitleCues detects subtitle formats from remote URLs', () => {
const assContent = [
'[Events]',
+160 -34
View File
@@ -1,9 +1,46 @@
import {
assOverrideSignature,
assToPlainText,
collectAssOverrideCommands,
parseAssEffectField,
type AssEffectKind,
type AssOverrideCommand,
} from './ass-text';
import { mergeDuplicateCues } from './subtitle-cue-dedup';
export interface SubtitleCue {
startTime: number;
endTime: number;
text: string;
}
/**
* Everything the parser knows about a source event, shared only with the dedup engine.
* Deduplication needs the authoring context -- which style the line belongs to, which
* override commands it carries, whether the `Effect` column was set -- to tell a karaoke
* burst apart from two characters saying the same word in turn. None of it is meaningful
* outside the parser, so the public API stays `{startTime, endTime, text}`.
*/
export interface AnnotatedSubtitleCue extends SubtitleCue {
/** Text exactly as authored, override blocks and all. */
rawText: string;
style: string;
layer: number;
/** ASS `Name`/`Actor` column. */
name: string;
/** ASS `Effect` column, verbatim. */
effect: string;
effectKind: AssEffectKind;
/** Override commands found in `{...}` blocks, with their arguments. */
overrides: readonly AssOverrideCommand[];
/** Canonical form of `overrides`, for spotting values that change across a run. */
overrideSignature: string;
/** Position in the source file, so sorting by time stays deterministic across layers. */
order: number;
}
export type SubtitleSourceFormat = 'ass' | 'srt';
const HTML_SUBTITLE_TAG_PATTERN = /<\/?[A-Za-z][^>\n]*>/g;
const SRT_TIMING_PATTERN =
@@ -23,12 +60,21 @@ function parseTimestamp(
);
}
/**
* The single ASS decode for the file path: cues leave the parser as plain text with real
* line breaks, matching what mpv hands over for the same line played live. No layer
* downstream decodes ASS again.
*/
function sanitizeSubtitleCueText(text: string): string {
return text.replace(ASS_OVERRIDE_TAG_PATTERN, '').replace(HTML_SUBTITLE_TAG_PATTERN, '').trim();
return assToPlainText(text, '\n').replace(HTML_SUBTITLE_TAG_PATTERN, '').trim();
}
export function parseSrtCues(content: string): SubtitleCue[] {
const cues: SubtitleCue[] = [];
function toPublicCues(cues: AnnotatedSubtitleCue[]): SubtitleCue[] {
return cues.map(({ startTime, endTime, text }) => ({ startTime, endTime, text }));
}
function parseAnnotatedSrtCues(content: string): AnnotatedSubtitleCue[] {
const cues: AnnotatedSubtitleCue[] = [];
const lines = content.split(/\r?\n/);
let i = 0;
@@ -60,20 +106,39 @@ export function parseSrtCues(content: string): SubtitleCue[] {
i += 1;
}
const text = sanitizeSubtitleCueText(textLines.join('\n'));
const rawText = textLines.join('\n');
const text = sanitizeSubtitleCueText(rawText);
if (text) {
cues.push({ startTime, endTime, text });
cues.push({
startTime,
endTime,
text,
rawText,
style: '',
layer: 0,
name: '',
effect: '',
effectKind: 'none',
// SRT and VTT carry no authoring metadata, and the dedup engine never reads
// overrides for those formats -- collecting them would be parsing for nobody.
overrides: [],
overrideSignature: '',
order: cues.length,
});
}
}
return cues;
}
const ASS_OVERRIDE_TAG_PATTERN = /\{[^}]*\}/g;
export function parseSrtCues(content: string): SubtitleCue[] {
return toPublicCues(parseAnnotatedSrtCues(content));
}
const ASS_TIMING_PATTERN = /^(\d+):(\d{2}):(\d{2})\.(\d{1,2})$/;
const ASS_FORMAT_PREFIX = 'Format:';
const ASS_DIALOGUE_PREFIX = 'Dialogue:';
const ASS_NAME_FIELD_ALIASES = ['name', 'actor'];
function parseAssTimestamp(raw: string): number | null {
const match = ASS_TIMING_PATTERN.exec(raw.trim());
@@ -87,13 +152,43 @@ function parseAssTimestamp(raw: string): number | null {
return hours * 3600 + minutes * 60 + seconds + centiseconds / 100;
}
export function parseAssCues(content: string): SubtitleCue[] {
const cues: SubtitleCue[] = [];
function readField(fields: string[], index: number): string {
return index >= 0 && index < fields.length ? fields[index]!.trim() : '';
}
function findFieldIndex(formatFields: string[], aliases: string[]): number {
for (const alias of aliases) {
const index = formatFields.indexOf(alias);
if (index >= 0) {
return index;
}
}
return -1;
}
function parseAnnotatedAssCues(content: string): AnnotatedSubtitleCue[] {
const cues: AnnotatedSubtitleCue[] = [];
const lines = content.split(/\r?\n/);
let inEventsSection = false;
let startFieldIndex = -1;
let endFieldIndex = -1;
let textFieldIndex = -1;
const fieldIndex = {
start: -1,
end: -1,
text: -1,
style: -1,
layer: -1,
name: -1,
effect: -1,
};
const resetFieldIndex = () => {
fieldIndex.start = -1;
fieldIndex.end = -1;
fieldIndex.text = -1;
fieldIndex.style = -1;
fieldIndex.layer = -1;
fieldIndex.name = -1;
fieldIndex.effect = -1;
};
for (const line of lines) {
const trimmed = line.trim();
@@ -101,9 +196,7 @@ export function parseAssCues(content: string): SubtitleCue[] {
if (trimmed.startsWith('[') && trimmed.endsWith(']')) {
inEventsSection = trimmed.toLowerCase() === '[events]';
if (!inEventsSection) {
startFieldIndex = -1;
endFieldIndex = -1;
textFieldIndex = -1;
resetFieldIndex();
}
continue;
}
@@ -117,9 +210,15 @@ export function parseAssCues(content: string): SubtitleCue[] {
.slice(ASS_FORMAT_PREFIX.length)
.split(',')
.map((field) => field.trim().toLowerCase());
startFieldIndex = formatFields.indexOf('start');
endFieldIndex = formatFields.indexOf('end');
textFieldIndex = formatFields.indexOf('text');
fieldIndex.start = formatFields.indexOf('start');
fieldIndex.end = formatFields.indexOf('end');
fieldIndex.text = formatFields.indexOf('text');
fieldIndex.style = formatFields.indexOf('style');
fieldIndex.layer = formatFields.indexOf('layer');
// Aegisub writes the speaker column as `Actor`; the v4+ spec calls it `Name`.
// Missing it costs the burst check its speaker guard, so both spellings count.
fieldIndex.name = findFieldIndex(formatFields, ASS_NAME_FIELD_ALIASES);
fieldIndex.effect = formatFields.indexOf('effect');
continue;
}
@@ -127,34 +226,57 @@ export function parseAssCues(content: string): SubtitleCue[] {
continue;
}
if (startFieldIndex < 0 || endFieldIndex < 0 || textFieldIndex < 0) {
if (fieldIndex.start < 0 || fieldIndex.end < 0 || fieldIndex.text < 0) {
continue;
}
const fields = trimmed.slice(ASS_DIALOGUE_PREFIX.length).split(',');
if (
startFieldIndex >= fields.length ||
endFieldIndex >= fields.length ||
textFieldIndex >= fields.length
fieldIndex.start >= fields.length ||
fieldIndex.end >= fields.length ||
fieldIndex.text >= fields.length
) {
continue;
}
const startTime = parseAssTimestamp(fields[startFieldIndex]!);
const endTime = parseAssTimestamp(fields[endFieldIndex]!);
const startTime = parseAssTimestamp(fields[fieldIndex.start]!);
const endTime = parseAssTimestamp(fields[fieldIndex.end]!);
if (startTime === null || endTime === null) {
continue;
}
const text = sanitizeSubtitleCueText(fields.slice(textFieldIndex).join(','));
if (text) {
cues.push({ startTime, endTime, text });
const rawText = fields.slice(fieldIndex.text).join(',');
const text = sanitizeSubtitleCueText(rawText);
if (!text) {
continue;
}
const effect = readField(fields, fieldIndex.effect);
const layer = Number(readField(fields, fieldIndex.layer));
const overrides = collectAssOverrideCommands(rawText);
cues.push({
startTime,
endTime,
text,
rawText,
style: readField(fields, fieldIndex.style),
layer: Number.isFinite(layer) ? layer : 0,
name: readField(fields, fieldIndex.name),
effect,
effectKind: parseAssEffectField(effect),
overrides,
overrideSignature: assOverrideSignature(overrides),
order: cues.length,
});
}
return cues;
}
export function parseAssCues(content: string): SubtitleCue[] {
return toPublicCues(parseAnnotatedAssCues(content));
}
function detectSubtitleFormat(source: string): 'srt' | 'vtt' | 'ass' | 'ssa' | null {
const [normalizedSource = source] =
(() => {
@@ -173,27 +295,31 @@ function detectSubtitleFormat(source: string): 'srt' | 'vtt' | 'ass' | 'ssa' | n
export function parseSubtitleCues(content: string, filename: string): SubtitleCue[] {
const format = detectSubtitleFormat(filename);
let cues: SubtitleCue[];
let cues: AnnotatedSubtitleCue[];
let sourceFormat: SubtitleSourceFormat = 'srt';
switch (format) {
case 'srt':
case 'vtt':
cues = parseSrtCues(content);
cues = parseAnnotatedSrtCues(content);
break;
case 'ass':
case 'ssa':
cues = parseAssCues(content);
cues = parseAnnotatedAssCues(content);
sourceFormat = 'ass';
break;
default:
cues = [];
}
if (cues.length === 0) {
const assCues = parseAssCues(content);
const srtCues = parseSrtCues(content);
cues = assCues.length >= srtCues.length ? assCues : srtCues;
const assCues = parseAnnotatedAssCues(content);
const srtCues = parseAnnotatedSrtCues(content);
const preferAss = assCues.length >= srtCues.length;
cues = preferAss ? assCues : srtCues;
sourceFormat = preferAss && assCues.length > 0 ? 'ass' : 'srt';
}
cues.sort((a, b) => a.startTime - b.startTime);
return cues;
cues.sort((a, b) => a.startTime - b.startTime || a.endTime - b.endTime || a.order - b.order);
return toPublicCues(mergeDuplicateCues(cues, sourceFormat));
}
@@ -115,6 +115,21 @@ test('subtitle processing does not emit plain payload for cached lines', async (
assert.deepEqual(emitted, [{ text: '字幕', tokens: [] }]);
});
test('text that normalizes to nothing is never cached', () => {
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => ({ text, tokens: [] }),
emitSubtitle: () => {},
});
// Two different inputs both reduce to an empty key; sharing one entry would serve the
// first one's tokens for the second.
controller.preCacheTokenization(' ', { text: ' ', tokens: [] });
assert.equal(controller.hasCachedSubtitle(' '), false);
assert.equal(controller.hasCachedSubtitle('\\n'), false);
assert.equal(controller.consumeCachedSubtitle('\\n'), null);
});
test('subtitle processing shows plain line while tokenization is still pending', async () => {
const emitted: SubtitleData[] = [];
let resolveTokenization: ((value: SubtitleData | null) => void) | undefined;
@@ -539,3 +554,125 @@ test('default cache limit covers a full-length title without evicting', () => {
assert.equal(controller.hasCachedSubtitle('line-0'), true);
assert.equal(controller.hasCachedSubtitle('line-1999'), true);
});
test('onSubtitleChange reports whether processing was scheduled', async () => {
const emitted: SubtitleData[] = [];
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => ({ text, tokens: [] }),
emitSubtitle: (payload) => emitted.push(payload),
});
// New text schedules work, so an emit (and anything gated on it) will follow.
assert.equal(controller.onSubtitleChange('字幕'), true);
await flushMicrotasks();
// A repeat emits nothing, so callers must not wait on an emit that is never
// coming (subtitle prefetching would stay paused for the rest of the cue).
const emittedCount = emitted.length;
assert.equal(controller.onSubtitleChange('字幕'), false);
await flushMicrotasks();
assert.equal(emitted.length, emittedCount);
});
test('refreshCurrentSubtitle reports the empty-text emit that an in-flight run will deliver', async () => {
const emitted: SubtitleData[] = [];
let resolveFirst: ((value: SubtitleData | null) => void) | undefined;
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => {
if (text === '字幕') {
return await new Promise<SubtitleData | null>((resolve) => {
resolveFirst = resolve;
});
}
return { text, tokens: [] };
},
emitSubtitle: (payload) => emitted.push(payload),
});
controller.onSubtitleChange('字幕');
await flushMicrotasks();
// Clearing the subtitle while tokenization is in flight: the running loop
// picks the empty text up and emits it, so callers gated on that emit (the
// prefetch pause) must be told one is coming.
assert.equal(controller.refreshCurrentSubtitle(''), true);
resolveFirst?.({ text: '字幕', tokens: [] });
await flushMicrotasks();
await flushMicrotasks();
// '字幕' is the provisional plain emit the in-flight run already made before
// the refresh; '' is the emit the refresh promised.
assert.deepEqual(
emitted.map((payload) => payload.text),
['字幕', ''],
);
});
test('onProcessingSettled fires once after the queue drains, including runs that emit nothing', async () => {
const events: string[] = [];
let resolveFirst: ((value: SubtitleData | null) => void) | undefined;
let tokenizationFails = false;
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => {
if (tokenizationFails) {
return null;
}
if (text === '一行目') {
return await new Promise<SubtitleData | null>((resolve) => {
resolveFirst = resolve;
});
}
return { text, tokens: [] };
},
emitSubtitle: (payload) => events.push(`emit:${payload.text}`),
onProcessingSettled: () => events.push('settled'),
});
controller.onSubtitleChange('一行目');
await flushMicrotasks();
// A second line arrives before the first finishes: the controller still has
// work, so it must not report itself settled between the two.
controller.onSubtitleChange('二行目');
resolveFirst?.({ text: '一行目', tokens: [] });
await flushMicrotasks();
await flushMicrotasks();
assert.deepEqual(events, ['emit:一行目', 'emit:二行目', 'emit:二行目', 'settled']);
// Tokenization failure on a line already shown plain: nothing is emitted, and
// the settle signal is the only way a caller learns the work is over.
events.length = 0;
tokenizationFails = true;
controller.invalidateTokenizationCache();
assert.equal(controller.refreshCurrentSubtitle('二行目'), true);
await flushMicrotasks();
await flushMicrotasks();
assert.deepEqual(events, ['settled']);
});
test('notePlainSubtitleEmitted suppresses the controller repeat of a payload already shown', async () => {
const emitted: SubtitleData[] = [];
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => ({ text, tokens: [] }),
emitSubtitle: (payload) => emitted.push(payload),
});
// Autoplay priming paints the plain line itself, then asks for tokenization.
controller.notePlainSubtitleEmitted('字幕');
controller.refreshCurrentSubtitle('字幕');
await flushMicrotasks();
assert.deepEqual(emitted, [{ text: '字幕', tokens: [] }]);
});
test('refreshCurrentSubtitle reports no emit for empty text when nothing is running', async () => {
const emitted: SubtitleData[] = [];
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => ({ text, tokens: [] }),
emitSubtitle: (payload) => emitted.push(payload),
});
assert.equal(controller.refreshCurrentSubtitle(''), false);
await flushMicrotasks();
assert.deepEqual(emitted, []);
});
@@ -1,8 +1,17 @@
import type { SubtitleData } from '../../types';
import { normalizePlainSubtitleText } from './ass-text';
export interface SubtitleProcessingControllerDeps {
tokenizeSubtitle: (text: string) => Promise<SubtitleData | null>;
emitSubtitle: (payload: SubtitleData) => void;
/**
* Fires when the controller runs out of work: every scheduled line has been
* processed, whether it ended in an emit, a suppressed duplicate, or a
* tokenizer failure. Callers that hold a resource for the duration of
* processing (prefetch pausing) release it here rather than on an emit,
* which is not guaranteed to happen.
*/
onProcessingSettled?: () => void;
logDebug?: (message: string) => void;
now?: () => number;
cacheLimit?: number;
@@ -17,16 +26,37 @@ export interface SubtitleProcessingControllerDeps {
export const DEFAULT_SUBTITLE_TOKENIZATION_CACHE_LIMIT = 2500;
export interface SubtitleProcessingController {
onSubtitleChange: (text: string) => void;
refreshCurrentSubtitle: (textOverride?: string) => void;
/**
* Returns whether processing is now scheduled or already in flight for this
* event. A false return means the controller is idle and will do nothing, so
* onProcessingSettled will not fire; callers that pause work for the duration
* of processing (such as subtitle prefetching) must release it themselves.
*/
onSubtitleChange: (text: string) => boolean;
/** Same contract as onSubtitleChange: whether processing is pending. */
refreshCurrentSubtitle: (textOverride?: string) => boolean;
/**
* Records that this exact text has already been shown plain by someone else
* (autoplay priming paints its first frame before scheduling tokenization),
* so the controller does not repeat that payload on its way to the tokenized
* one.
*/
notePlainSubtitleEmitted: (text: string) => void;
invalidateTokenizationCache: () => void;
preCacheTokenization: (text: string, data: SubtitleData) => void;
consumeCachedSubtitle: (text: string) => SubtitleData | null;
hasCachedSubtitle: (text: string) => boolean;
}
/**
* Prefetched cues and live mpv text are both already decoded from ASS, so the key only
* has to settle whitespace for one authored line to resolve to one entry.
*
* An empty key is not a line: it is whatever normalization reduced to nothing. Callers
* must skip the cache for it rather than let every such input share one entry.
*/
export function normalizeSubtitleCacheKey(text: string): string {
return text.replace(/\r\n/g, '\n').replace(/\\N/g, '\n').replace(/\\n/g, '\n').trim();
return normalizePlainSubtitleText(text);
}
export function createSubtitleProcessingController(
@@ -50,6 +80,9 @@ export function createSubtitleProcessingController(
const getCachedTokenization = (text: string): SubtitleData | null => {
const cacheKey = normalizeSubtitleCacheKey(text);
if (!cacheKey) {
return null;
}
const cached = tokenizationCache.get(cacheKey);
if (!cached) {
return null;
@@ -61,7 +94,11 @@ export function createSubtitleProcessingController(
};
const setCachedTokenization = (text: string, payload: SubtitleData): void => {
tokenizationCache.set(normalizeSubtitleCacheKey(text), payload);
const cacheKey = normalizeSubtitleCacheKey(text);
if (!cacheKey) {
return;
}
tokenizationCache.set(cacheKey, payload);
while (tokenizationCache.size > SUBTITLE_TOKENIZATION_CACHE_LIMIT) {
const firstKey = tokenizationCache.keys().next().value;
if (firstKey !== undefined) {
@@ -164,14 +201,20 @@ export function createSubtitleProcessingController(
(latestText.trim() && cacheGeneration !== lastEmittedGeneration)
) {
processLatest();
return;
}
// Nothing left to do: signal completion even when this run emitted
// nothing (suppressed duplicate, tokenizer failure), or callers waiting
// on the controller would wait forever.
deps.onProcessingSettled?.();
});
};
return {
onSubtitleChange: (text: string) => {
if (text === latestText) {
return;
// A run already in flight for this text will still emit for it.
return processing;
}
latestText = text;
if (
@@ -183,21 +226,28 @@ export function createSubtitleProcessingController(
lastPlainEmittedText = text;
}
processLatest();
return true;
},
refreshCurrentSubtitle: (textOverride?: string) => {
if (typeof textOverride === 'string') {
latestText = textOverride;
}
if (!latestText.trim()) {
return;
// A run in flight will pick this up and emit the empty subtitle, so
// the caller is still waiting on an emit.
return processing;
}
if (
processing ||
(latestText === lastEmittedText && cacheGeneration === lastEmittedGeneration)
) {
return;
if (processing) {
return true;
}
if (latestText === lastEmittedText && cacheGeneration === lastEmittedGeneration) {
return false;
}
processLatest();
return true;
},
notePlainSubtitleEmitted: (text: string) => {
lastPlainEmittedText = text;
},
invalidateTokenizationCache: () => {
tokenizationCache.clear();
@@ -219,7 +269,8 @@ export function createSubtitleProcessingController(
return cached;
},
hasCachedSubtitle: (text: string) => {
return tokenizationCache.has(normalizeSubtitleCacheKey(text));
const cacheKey = normalizeSubtitleCacheKey(text);
return cacheKey.length > 0 && tokenizationCache.has(cacheKey);
},
};
}
+6 -36
View File
@@ -1651,9 +1651,11 @@ test('tokenizeSubtitle clears JLPT level from standalone Yomitan particle token'
assert.equal(result.tokens?.[0]?.jlptLevel, undefined);
});
test('tokenizeSubtitle returns null tokens for empty normalized text', async () => {
test('tokenizeSubtitle returns the normalized text when it comes out empty', async () => {
// Handing back the original would push whatever normalization dropped into app state
// as if it were subtitle text.
const result = await tokenizeSubtitle(' \\n ', makeDeps());
assert.deepEqual(result, { text: ' \\n ', tokens: null });
assert.deepEqual(result, { text: '', tokens: null });
});
test('tokenizeSubtitle normalizes newlines before Yomitan parse request', async () => {
@@ -2934,44 +2936,12 @@ test('tokenizeSubtitle preserves Yomitan compound token when MeCab components ar
return [];
}
if (script.includes('parseText')) {
return [
{
source: 'scanning-parser',
index: 0,
content: [
[
{
text: '取り組んで',
surface: '取り組んで',
reading: 'とりくんで',
headwords: [[{ term: '取り組む' }]],
},
],
[
{
text: 'もらいます',
reading: 'もらいます',
headwords: [[{ term: 'もらう' }]],
},
],
],
},
];
}
return [
{
surface: '取り',
reading: 'とり',
headword: '取る',
headword: '取り組む',
startPos: 0,
endPos: 2,
},
{
surface: '組んで',
reading: 'くんで',
headword: '組む',
startPos: 2,
endPos: 5,
},
{
+57 -8
View File
@@ -27,6 +27,7 @@ import {
} from './tokenizer/yomitan-parser-runtime';
import type { YomitanTermFrequency } from './tokenizer/yomitan-parser-runtime';
import { isKanaChar } from './tokenizer/token-classification';
import { normalizePlainSubtitleText } from './ass-text';
const logger = createLogger('main:tokenizer');
@@ -70,6 +71,7 @@ export interface TokenizerServiceDeps {
getNameMatchImagesEnabled?: () => boolean;
getCharacterNameImage?: (term: string) => CharacterNameImage | null;
getCurrentCharacterDictionaryMediaId?: () => number | null;
getCharacterNameCandidates?: () => { key: string; forms: string[] } | null;
getFrequencyDictionaryEnabled?: () => boolean;
getFrequencyDictionaryMatchMode?: () => FrequencyDictionaryMatchMode;
getFrequencyRank?: FrequencyDictionaryLookup;
@@ -106,6 +108,7 @@ export interface TokenizerDepsRuntimeOptions {
getNameMatchImagesEnabled?: () => boolean;
getCharacterNameImage?: (term: string) => CharacterNameImage | null;
getCurrentCharacterDictionaryMediaId?: () => number | null;
getCharacterNameCandidates?: () => { key: string; forms: string[] } | null;
getFrequencyDictionaryEnabled?: () => boolean;
getFrequencyDictionaryMatchMode?: () => FrequencyDictionaryMatchMode;
getFrequencyRank?: FrequencyDictionaryLookup;
@@ -266,6 +269,7 @@ export function createTokenizerDepsRuntime(
getNameMatchImagesEnabled: options.getNameMatchImagesEnabled,
getCharacterNameImage: options.getCharacterNameImage,
getCurrentCharacterDictionaryMediaId: options.getCurrentCharacterDictionaryMediaId,
getCharacterNameCandidates: options.getCharacterNameCandidates,
getFrequencyDictionaryEnabled: options.getFrequencyDictionaryEnabled,
getFrequencyDictionaryMatchMode: options.getFrequencyDictionaryMatchMode ?? (() => 'headword'),
getFrequencyRank: options.getFrequencyRank,
@@ -716,15 +720,30 @@ function getAnnotationOptions(deps: TokenizerServiceDeps): TokenizerAnnotationOp
};
}
// Per-line stage durations for the pipeline debug log; every field is filled in
// by the stage that awaits the corresponding work.
interface TokenizationStageTimings {
scanMs?: number;
mecabMs?: number;
frequencyMs?: number;
annotateMs?: number;
}
async function parseWithYomitanInternalParser(
text: string,
deps: TokenizerServiceDeps,
options: TokenizerAnnotationOptions,
stageTimings?: TokenizationStageTimings,
): Promise<MergedToken[] | null> {
const scanStartedAtMs = Date.now();
const selectedTokens = await requestYomitanScanTokens(text, deps, logger, {
includeNameMatchMetadata: options.nameMatchEnabled,
currentCharacterDictionaryMediaId: deps.getCurrentCharacterDictionaryMediaId?.() ?? null,
nameCandidates: deps.getCharacterNameCandidates?.() ?? null,
});
if (stageTimings) {
stageTimings.scanMs = Date.now() - scanStartedAtMs;
}
if (!selectedTokens || selectedTokens.length === 0) {
return null;
}
@@ -757,6 +776,7 @@ async function parseWithYomitanInternalParser(
const frequencyRankPromise: Promise<YomitanFrequencyIndex> = options.frequencyEnabled
? (async () => {
const frequencyStartedAtMs = Date.now();
const frequencyMatchMode = options.frequencyMatchMode;
const termReadingList = buildYomitanFrequencyTermReadingList(
normalizedSelectedTokens,
@@ -767,12 +787,17 @@ async function parseWithYomitanInternalParser(
deps,
logger,
);
return buildYomitanFrequencyIndex(yomitanFrequencies);
const frequencyIndex = buildYomitanFrequencyIndex(yomitanFrequencies);
if (stageTimings) {
stageTimings.frequencyMs = Date.now() - frequencyStartedAtMs;
}
return frequencyIndex;
})()
: Promise.resolve({ byPair: new Map(), byTerm: new Map() });
const mecabEnrichmentPromise: Promise<MergedToken[]> = needsMecabPosEnrichment(options)
? (async () => {
const mecabStartedAtMs = Date.now();
try {
const mecabTokens = await deps.tokenizeWithMecab(text);
const enrichTokensWithMecab = deps.enrichTokensWithMecab ?? enrichTokensWithMecabAsync;
@@ -786,6 +811,10 @@ async function parseWithYomitanInternalParser(
`textLength=${text.length}`,
);
return normalizedSelectedTokens;
} finally {
if (stageTimings) {
stageTimings.mecabMs = Date.now() - mecabStartedAtMs;
}
}
})()
: Promise.resolve(normalizedSelectedTokens);
@@ -858,14 +887,14 @@ export async function tokenizeSubtitle(
text: string,
deps: TokenizerServiceDeps,
): Promise<SubtitleData> {
const displayText = text
.replace(/\r\n/g, '\n')
.replace(/\\N/g, '\n')
.replace(/\\n/g, '\n')
.trim();
const displayText = normalizePlainSubtitleText(text);
// ASS decoding already happened upstream (cue parser for files, mpv for live text), so
// all this drops is whitespace -- but a whitespace-only line still normalizes to empty.
// Return the normalized form anyway: handing back the original would put a blank line
// into application state as if it were subtitle text.
if (!displayText) {
return { text, tokens: null };
return { text: displayText, tokens: null };
}
const tokenizeText = displayText
@@ -876,15 +905,35 @@ export async function tokenizeSubtitle(
const annotationOptions = getAnnotationOptions(deps);
annotationOptions.sourceText = tokenizeText;
const yomitanTokens = await parseWithYomitanInternalParser(tokenizeText, deps, annotationOptions);
const stageTimings: TokenizationStageTimings = {};
const startedAtMs = Date.now();
const logStageTimings = (tokenCount: number): void => {
logger.debug(
`Subtitle tokenization stages; textLength=${tokenizeText.length}, tokenCount=${tokenCount}, ` +
`scanMs=${stageTimings.scanMs ?? '-'}, mecabMs=${stageTimings.mecabMs ?? '-'}, ` +
`frequencyMs=${stageTimings.frequencyMs ?? '-'}, annotateMs=${stageTimings.annotateMs ?? '-'}, ` +
`totalMs=${Date.now() - startedAtMs}`,
);
};
const yomitanTokens = await parseWithYomitanInternalParser(
tokenizeText,
deps,
annotationOptions,
stageTimings,
);
if (yomitanTokens && yomitanTokens.length > 0) {
const annotateStartedAtMs = Date.now();
const annotatedTokens = await applyAnnotationStage(yomitanTokens, deps, annotationOptions);
stageTimings.annotateMs = Date.now() - annotateStartedAtMs;
const renderedTokens = applyCharacterNameImages(annotatedTokens, deps, annotationOptions);
logStageTimings(renderedTokens.length);
return {
text: displayText,
tokens: renderedTokens.length > 0 ? renderedTokens : null,
};
}
logStageTimings(0);
return { text: displayText, tokens: null };
}
@@ -0,0 +1,4 @@
// Title prefix of the dictionaries SubMiner generates per media. Lives on its
// own because both the main process and the injected scan runtime match on it,
// and the injected fragments interpolate it into their own source.
export const CHARACTER_DICTIONARY_TITLE_PREFIX = 'SubMiner Character Dictionary';
@@ -366,8 +366,11 @@ export function createReplayMessageStore(messages: GoldenRecordedMessage[]): Rep
};
}
async function runInjectedScriptInVm(script: string, store: ReplayMessageStore): Promise<unknown> {
return await vm.runInNewContext(script, {
// One persistent context per fixture, matching the real parser window: the
// scan runtime installs itself once into globalThis and later per-line call
// scripts reuse it.
function createInjectedScriptVm(store: ReplayMessageStore): (script: string) => Promise<unknown> {
const context = vm.createContext({
chrome: {
runtime: {
lastError: null,
@@ -393,6 +396,7 @@ async function runInjectedScriptInVm(script: string, store: ReplayMessageStore):
Set,
String,
});
return async (script: string) => await vm.runInContext(script, context);
}
export function createReplayTokenizerDeps(fixture: GoldenFixture): TokenizerServiceDeps {
@@ -400,13 +404,14 @@ export function createReplayTokenizerDeps(fixture: GoldenFixture): TokenizerServ
const scriptResults = new Map(
fixture.recording.scripts.map((entry) => [entry.sha256, entry] as const),
);
const runInjectedScriptInVm = createInjectedScriptVm(store);
const parserWindow = {
isDestroyed: () => false,
webContents: {
executeJavaScript: async (script: string) => {
try {
return await runInjectedScriptInVm(script, store);
return await runInjectedScriptInVm(script);
} catch (vmError) {
const recorded = scriptResults.get(hashInjectedScript(script));
if (recorded) {
@@ -8,6 +8,7 @@ import {
isKanaChar,
isKanaOnlyText,
isTokenPos2Excluded,
normalizeKana,
} from './token-classification';
const POS1_EXCLUSIONS = new Set(['助詞']);
@@ -29,6 +30,26 @@ function makeNoun(surface: string): MergedToken {
};
}
test('kana normalization folds halfwidth kana, composing the voiced pairs', () => {
// カ + ゙ is two code points for one character: without composing them, a
// halfwidth word counts as longer than the reading that spells it, which
// disqualifies the reading from known-word matching.
assert.equal(normalizeKana('ガク'), normalizeKana('ガク'));
assert.equal(normalizeKana('パン'), normalizeKana('パン'));
assert.equal(normalizeKana('ミナト'), 'みなと');
assert.ok(isKanaOnlyText('ガク'));
});
test('kana normalization leaves characters other than halfwidth kana alone', () => {
// The composition is scoped to the halfwidth runs: applied to the whole
// string, NFKC would also rewrite these into something the dictionary, the
// known-word list, and the frequency data were never keyed on.
assert.equal(normalizeKana('①ガ'), '①が');
assert.equal(normalizeKana('Aガ'), 'Aが');
assert.equal(normalizeKana('㍑ガ'), '㍑が');
assert.equal(normalizeKana('fiガ'), 'fiが');
});
test('kana classification excludes the katakana-hiragana double hyphen', () => {
assert.equal(isKanaChar(''), false);
assert.equal(isKanaOnlyText(''), false);
@@ -4,8 +4,20 @@ const KATAKANA_TO_HIRAGANA_OFFSET = 0x60;
const KATAKANA_CODEPOINT_START = 0x30a1;
const KATAKANA_CODEPOINT_END = 0x30f6;
// No `u` flag: the range is entirely BMP so it changes nothing here, and
// Bun's unicode-mode matcher mis-handles this class next to certain ligatures.
const HALFWIDTH_KANA_RUN = /[\uff66-\uff9f]+/g;
// NFKC over the halfwidth kana only, never the whole string: it composes the
// voiced pairs (カ + ゙) into single characters so ガク compares equal to ガク
// instead of counting one character longer than the word it spells, but run
// over everything it would also rewrite unrelated text (① → 1, ㍑ → リットル).
function composeHalfwidthKana(text: string): string {
return text.replace(HALFWIDTH_KANA_RUN, (run) => run.normalize('NFKC'));
}
export function normalizeKana(text: string): string {
const raw = text.trim();
const raw = composeHalfwidthKana(text).trim();
if (!raw) {
return '';
}
@@ -0,0 +1,150 @@
// Dictionary classification for the injected scan runtime: which dictionaries
// an entry came from, and whether it is a SubMiner character entry for the
// media being watched. Both walk nested entry data, so both are memoized on the
// entry object by the runtime that hosts them.
import { CHARACTER_DICTIONARY_TITLE_PREFIX } from './character-dictionary-title';
// The prefix is interpolated into generated regex source, so metacharacters in
// it would change what the pattern matches (or fail to compile).
const ESCAPED_TITLE_PREFIX = CHARACTER_DICTIONARY_TITLE_PREFIX.replace(
/[.*+?^${}()|[\]\\]/g,
'\\$&',
);
const TITLE_MEDIA_ID_PATTERN = ESCAPED_TITLE_PREFIX + String.raw`[^\d]*(?:AniList\s*)?(\d+)`;
export const YOMITAN_DICTIONARY_CLASSIFICATION_HELPERS = String.raw`
function normalizeWordClasses(headword) {
if (!Array.isArray(headword?.wordClasses)) { return undefined; }
const classes = headword.wordClasses.filter((wordClass) => typeof wordClass === "string" && wordClass.trim().length > 0);
return classes.length > 0 ? classes : undefined;
}
function appendDictionaryNames(target, value) {
if (!value || typeof value !== 'object') {
return;
}
const candidates = [
value.dictionary,
value.dictionaryName,
value.name,
value.title,
value.dictionaryTitle,
value.dictionaryAlias
];
for (const candidate of candidates) {
if (typeof candidate === 'string' && candidate.trim().length > 0) {
target.push(candidate.trim());
}
}
}
// Memoized on the entry object: termsFind results are cached across
// lines, so the same entries come back for every repeated lookup, and
// each one is classified several times per scan (name pre-pass,
// headword preference, every retry window).
function getDictionaryEntryNames(entry) {
if (!entry || typeof entry !== 'object') { return []; }
const cached = dictionaryEntryNamesCache.get(entry);
if (cached !== undefined) { return cached; }
const names = [];
appendDictionaryNames(names, entry);
for (const definition of entry?.definitions || []) {
appendDictionaryNames(names, definition);
}
for (const frequency of entry?.frequencies || []) {
appendDictionaryNames(names, frequency);
}
for (const pronunciation of entry?.pronunciations || []) {
appendDictionaryNames(names, pronunciation);
}
dictionaryEntryNamesCache.set(entry, names);
return names;
}
// Cached per scan rather than per runtime: the answer depends on
// includeNameMatchMetadata, which is a per-call parameter.
const nameDictionaryEntryCache = new WeakMap();
function isNameDictionaryEntry(entry) {
if (!includeNameMatchMetadata || !entry || typeof entry !== 'object') {
return false;
}
const cached = nameDictionaryEntryCache.get(entry);
if (cached !== undefined) { return cached; }
const isName = getDictionaryEntryNames(entry).some((name) => name.startsWith(${JSON.stringify(CHARACTER_DICTIONARY_TITLE_PREFIX)}));
nameDictionaryEntryCache.set(entry, isName);
return isName;
}
const TITLE_MEDIA_ID_REGEX = new RegExp(${JSON.stringify(TITLE_MEDIA_ID_PATTERN)}, 'i');
function parseSubMinerMediaIdFromString(value) {
const imageMatch = value.match(/\bimg\/m(\d+)-/i);
if (imageMatch) {
const parsed = Number.parseInt(imageMatch[1], 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
const titleMatch = value.match(TITLE_MEDIA_ID_REGEX);
if (titleMatch) {
const parsed = Number.parseInt(titleMatch[1], 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
return null;
}
function parseSubMinerMediaIdCandidate(value) {
if (typeof value === 'number' && Number.isSafeInteger(value) && value > 0) {
return value;
}
if (typeof value === 'string' && /^\d+$/.test(value.trim())) {
const parsed = Number.parseInt(value.trim(), 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
return null;
}
function collectSubMinerMediaIds(value, target) {
if (typeof value === 'string') {
const parsed = parseSubMinerMediaIdFromString(value);
if (parsed !== null) { target.add(parsed); }
return;
}
if (!value || typeof value !== 'object') {
return;
}
if (Array.isArray(value)) {
for (const item of value) { collectSubMinerMediaIds(item, target); }
return;
}
const mediaIdCandidates = [
value.subminerMediaId,
value.subMinerMediaId,
value.characterDictionaryMediaId,
value.data?.subminerMediaId,
value.data?.subMinerMediaId,
value.data?.characterDictionaryMediaId
];
for (const candidate of mediaIdCandidates) {
const parsed = parseSubMinerMediaIdCandidate(candidate);
if (parsed !== null) { target.add(parsed); }
}
for (const child of Object.values(value)) {
collectSubMinerMediaIds(child, target);
}
}
// Walking an entry collects media ids from every nested value, so this
// is the most expensive classification step; memoized on the entry for
// the same reason as the dictionary names above.
function getSubMinerMediaIds(entry) {
if (!entry || typeof entry !== 'object') { return EMPTY_MEDIA_ID_SET; }
const cached = subMinerMediaIdsCache.get(entry);
if (cached !== undefined) { return cached; }
const mediaIds = new Set();
collectSubMinerMediaIds(entry, mediaIds);
subMinerMediaIdsCache.set(entry, mediaIds);
return mediaIds;
}
function isCurrentMediaNameDictionaryEntry(entry) {
if (!isNameDictionaryEntry(entry)) {
return false;
}
if (currentCharacterDictionaryMediaId === null) {
return true;
}
const mediaIds = getSubMinerMediaIds(entry);
return mediaIds.size === 0 || mediaIds.has(currentCharacterDictionaryMediaId);
}
`;
@@ -0,0 +1,135 @@
// Frequency-rank resolution for the injected scan runtime: reads the many
// shapes a Yomitan frequency entry can take and picks the best rank for a
// headword, honouring per-dictionary priority and occurrence-vs-rank mode.
export const YOMITAN_FREQUENCY_HELPERS = String.raw`
function parsePositiveFrequencyNumber(value) {
if (typeof value === 'number' && Number.isFinite(value) && value > 0) {
return Math.max(1, Math.floor(value));
}
if (typeof value === 'string') {
const numericMatch = value.trim().match(/[+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?/)?.[0];
if (!numericMatch) { return null; }
const parsed = Number.parseFloat(numericMatch);
if (!Number.isFinite(parsed) || parsed <= 0) { return null; }
return Math.max(1, Math.floor(parsed));
}
if (Array.isArray(value)) {
for (const item of value) {
const parsed = parsePositiveFrequencyNumber(item);
if (parsed !== null) { return parsed; }
}
}
return null;
}
function parseDisplayFrequencyNumber(value) {
if (typeof value === 'string') {
const leadingDigits = value.trim().match(/^\d+/)?.[0];
if (!leadingDigits) { return null; }
const parsed = Number.parseInt(leadingDigits, 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : null;
}
return parsePositiveFrequencyNumber(value);
}
function getFrequencyDictionaryName(frequency) {
const candidates = [
frequency?.dictionary,
frequency?.dictionaryName,
frequency?.name,
frequency?.title,
frequency?.dictionaryTitle,
frequency?.dictionaryAlias
];
for (const candidate of candidates) {
if (typeof candidate === 'string' && candidate.trim().length > 0) {
return candidate.trim();
}
}
return null;
}
function getBestFrequencyRank(dictionaryEntry, headwordIndex, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
let best = null;
const headwordCount = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords.length : 0;
for (const frequency of dictionaryEntry?.frequencies || []) {
if (!frequency || typeof frequency !== 'object') { continue; }
const frequencyHeadwordIndex = frequency.headwordIndex;
if (typeof frequencyHeadwordIndex === 'number') {
if (frequencyHeadwordIndex !== headwordIndex) { continue; }
} else if (headwordCount > 1) {
continue;
}
const dictionary = getFrequencyDictionaryName(frequency);
if (!dictionary) { continue; }
if (dictionaryFrequencyModeByName[dictionary] === 'occurrence-based') { continue; }
const rank =
parseDisplayFrequencyNumber(frequency.displayValue) ??
parsePositiveFrequencyNumber(frequency.frequency);
if (rank === null) { continue; }
const priorityRaw = dictionaryPriorityByName[dictionary];
const fallbackPriority =
typeof frequency.dictionaryIndex === 'number' && Number.isFinite(frequency.dictionaryIndex)
? Math.max(0, Math.floor(frequency.dictionaryIndex))
: Number.MAX_SAFE_INTEGER;
const priority =
typeof priorityRaw === 'number' && Number.isFinite(priorityRaw)
? Math.max(0, Math.floor(priorityRaw))
: fallbackPriority;
if (best === null || priority < best.priority || (priority === best.priority && rank < best.rank)) {
best = { priority, rank };
}
}
return best?.rank ?? null;
}
function hasExactSource(headword, token, requirePrimary) {
for (const src of headword?.sources || []) {
if (src.originalText !== token) { continue; }
if (requirePrimary && !src.isPrimary) { continue; }
if (src.matchType !== 'exact') { continue; }
return true;
}
return false;
}
function collectExactHeadwordMatches(dictionaryEntries, token, requirePrimary) {
const matches = [];
for (const dictionaryEntry of dictionaryEntries || []) {
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (let headwordIndex = 0; headwordIndex < headwords.length; headwordIndex += 1) {
const headword = headwords[headwordIndex];
if (!hasExactSource(headword, token, requirePrimary)) { continue; }
matches.push({ dictionaryEntry, headword, headwordIndex });
}
}
return matches;
}
function sameHeadword(match, preferredMatch) {
if (!match || !preferredMatch) {
return false;
}
if (match.headword?.term !== preferredMatch.headword?.term) {
return false;
}
const matchReading = typeof match.headword?.reading === 'string' ? match.headword.reading : '';
const preferredReading =
typeof preferredMatch.headword?.reading === 'string' ? preferredMatch.headword.reading : '';
if (!matchReading || !preferredReading) {
return true;
}
return matchReading === preferredReading;
}
function getBestFrequencyRankForMatches(matches, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
let best = null;
for (const match of matches) {
const rank = getBestFrequencyRank(
match.dictionaryEntry,
match.headwordIndex,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
);
if (rank === null) { continue; }
if (best === null || rank < best) {
best = rank;
}
}
return best;
}
`;
@@ -0,0 +1,170 @@
// Furigana distribution for the injected scan runtime: splits a headword and
// its reading into the segments a token carries, including the inflected case
// where the matched source text differs from the dictionary form.
export const YOMITAN_FURIGANA_HELPERS = String.raw`
function createFuriganaSegment(text, reading) { return {text, reading}; }
function getSegmentReadingContribution(segment) {
if (typeof segment.reading === "string" && segment.reading.length > 0) { return segment.reading; }
const segmentText = typeof segment.text === "string" ? segment.text : "";
const isKanaOnly = segmentText.length > 0 && [...segmentText].every((char) => isCodePointKana(char.codePointAt(0)));
return isKanaOnly ? convertHalfwidthKanaToKatakana(segmentText) : "";
}
function getProlongedHiragana(previousCharacter) {
switch (previousCharacter) {
case "あ": case "か": case "が": case "さ": case "ざ": case "た": case "だ": case "な": case "は": case "ば": case "ぱ": case "ま": case "や": case "ら": case "わ": case "ぁ": case "ゃ": case "ゎ": return "あ";
case "い": case "き": case "ぎ": case "し": case "じ": case "ち": case "ぢ": case "に": case "ひ": case "び": case "ぴ": case "み": case "り": case "ぃ": return "い";
case "う": case "く": case "ぐ": case "す": case "ず": case "つ": case "づ": case "ぬ": case "ふ": case "ぶ": case "ぷ": case "む": case "ゆ": case "る": case "ぅ": case "ゅ": return "う";
case "え": case "け": case "げ": case "せ": case "ぜ": case "て": case "で": case "ね": case "へ": case "べ": case "ぺ": case "め": case "れ": case "ぇ": return "え";
case "お": case "こ": case "ご": case "そ": case "ぞ": case "と": case "ど": case "の": case "ほ": case "ぼ": case "ぽ": case "も": case "よ": case "ろ": case "を": case "ぉ": case "ょ": return "う";
default: return null;
}
}
function getFuriganaKanaSegments(text, reading) {
const newSegments = [];
let start = 0;
let state = (reading[0] === text[0]);
for (let i = 1; i < text.length; ++i) {
const newState = (reading[i] === text[i]);
if (state === newState) { continue; }
newSegments.push(createFuriganaSegment(text.substring(start, i), state ? '' : reading.substring(start, i)));
state = newState;
start = i;
}
newSegments.push(createFuriganaSegment(text.substring(start), state ? '' : reading.substring(start)));
return newSegments;
}
function convertKatakanaToHiragana(text, keepProlongedSoundMarks = false) {
let result = '';
const offset = (HIRAGANA_CONVERSION_RANGE[0] - KATAKANA_CONVERSION_RANGE[0]);
for (let char of text) {
const codePoint = char.codePointAt(0);
switch (codePoint) {
case KATAKANA_SMALL_KA_CODE_POINT:
case KATAKANA_SMALL_KE_CODE_POINT:
break;
case KANA_PROLONGED_SOUND_MARK_CODE_POINT:
case HALFWIDTH_KANA_PROLONGED_SOUND_MARK_CODE_POINT:
char = "ー";
if (!keepProlongedSoundMarks && result.length > 0) {
const char2 = getProlongedHiragana(result[result.length - 1]);
if (char2 !== null) { char = char2; }
}
break;
default:
if (isCodePointInRange(codePoint, KATAKANA_CONVERSION_RANGE)) {
char = String.fromCodePoint(codePoint + offset);
break;
}
// Halfwidth katakana folds too, or a name written that way would
// match neither a candidate form nor its own reading.
const halfwidthHiragana = convertHalfwidthKanaCodePointToHiragana(codePoint);
if (halfwidthHiragana !== null) { char = halfwidthHiragana; }
break;
}
result += char;
}
return result;
}
function segmentizeFurigana(reading, readingNormalized, groups, groupsStart) {
const groupCount = groups.length - groupsStart;
if (groupCount <= 0) { return reading.length === 0 ? [] : null; }
const group = groups[groupsStart];
const {isKana, text} = group;
if (isKana) {
if (group.textNormalized !== null && readingNormalized.startsWith(group.textNormalized)) {
const segments = segmentizeFurigana(reading.substring(text.length), readingNormalized.substring(text.length), groups, groupsStart + 1);
if (segments !== null) {
if (reading.startsWith(text)) { segments.unshift(createFuriganaSegment(text, '')); }
else { segments.unshift(...getFuriganaKanaSegments(text, reading)); }
return segments;
}
}
return null;
}
let result = null;
for (let i = reading.length; i >= text.length; --i) {
const segments = segmentizeFurigana(reading.substring(i), readingNormalized.substring(i), groups, groupsStart + 1);
if (segments !== null) {
if (result !== null) { return null; }
segments.unshift(createFuriganaSegment(text, reading.substring(0, i)));
result = segments;
}
if (groupCount === 1) { break; }
}
return result;
}
function distributeFurigana(term, reading) {
if (reading === term) { return [createFuriganaSegment(term, '')]; }
const groups = [];
let groupPre = null;
let isKanaPre = null;
for (const c of term) {
const isKana = isCodePointKana(c.codePointAt(0));
if (isKana === isKanaPre) { groupPre.text += c; }
else {
groupPre = {isKana, text: c, textNormalized: null};
groups.push(groupPre);
isKanaPre = isKana;
}
}
for (const group of groups) {
if (group.isKana) { group.textNormalized = convertKatakanaToHiragana(group.text); }
}
const segments = segmentizeFurigana(reading, convertKatakanaToHiragana(reading), groups, 0);
return segments !== null ? segments : [createFuriganaSegment(term, reading)];
}
function getStemLength(text1, text2) {
const minLength = Math.min(text1.length, text2.length);
if (minLength === 0) { return 0; }
let i = 0;
while (true) {
const char1 = text1.codePointAt(i);
const char2 = text2.codePointAt(i);
if (char1 !== char2) { break; }
const charLength = String.fromCodePoint(char1).length;
i += charLength;
if (i >= minLength) {
if (i > minLength) { i -= charLength; }
break;
}
}
return i;
}
function distributeFuriganaInflected(term, reading, source) {
const termNormalized = convertKatakanaToHiragana(term);
const readingNormalized = convertKatakanaToHiragana(reading);
const sourceNormalized = convertKatakanaToHiragana(source);
let mainText = term;
let stemLength = getStemLength(termNormalized, sourceNormalized);
const readingStemLength = getStemLength(readingNormalized, sourceNormalized);
if (readingStemLength > 0 && readingStemLength >= stemLength) {
mainText = reading;
stemLength = readingStemLength;
reading = source.substring(0, stemLength) + reading.substring(stemLength);
}
const segments = [];
if (stemLength > 0) {
mainText = source.substring(0, stemLength) + mainText.substring(stemLength);
const segments2 = distributeFurigana(mainText, reading);
let consumed = 0;
for (const segment of segments2) {
const start = consumed;
consumed += segment.text.length;
if (consumed < stemLength) { segments.push(segment); }
else if (consumed === stemLength) { segments.push(segment); break; }
else {
if (start < stemLength) { segments.push(createFuriganaSegment(mainText.substring(start, stemLength), '')); }
break;
}
}
}
if (stemLength < source.length) {
const remainder = source.substring(stemLength);
const last = segments[segments.length - 1];
if (last && last.reading.length === 0) { last.text += remainder; }
else { segments.push(createFuriganaSegment(remainder, '')); }
}
return segments;
}
`;
@@ -0,0 +1,45 @@
// Kana classification and normalization for the injected scan runtime: the
// code-point ranges the walk tests every character against, and the folds that
// let halfwidth and katakana spellings compare equal to their dictionary form.
import { HAN_CODE_POINT_RANGES } from '../../text/han-code-points';
export const YOMITAN_KANA_HELPERS = String.raw`
const HIRAGANA_CONVERSION_RANGE = [0x3041, 0x3096];
const KATAKANA_CONVERSION_RANGE = [0x30a1, 0x30f6];
const KANA_PROLONGED_SOUND_MARK_CODE_POINT = 0x30fc;
const KATAKANA_SMALL_KA_CODE_POINT = 0x30f5;
const KATAKANA_SMALL_KE_CODE_POINT = 0x30f6;
const KANA_RANGES = [[0x3040, 0x309f], [0x30a0, 0x30ff], [0xff66, 0xff9f]];
const HALFWIDTH_KATAKANA_RANGE = [0xff66, 0xff9d];
const HALFWIDTH_KANA_PROLONGED_SOUND_MARK_CODE_POINT = 0xff70;
// Folded one code point to one, so every index into a normalized string
// still lines up with the original text — the name-candidate prefilter
// and the furigana stem matching both index back into it. The standalone
// voiced marks (゙ ゚) have no one-character equivalent and stay as they are.
const HALFWIDTH_KATAKANA_TO_HIRAGANA = "をぁぃぅぇぉゃゅょっーあいうえおかきくけこさしすせそたちつてとなにぬねのはひふへほまみむめもやゆよらりるれろわん";
function convertHalfwidthKanaCodePointToHiragana(codePoint) {
if (codePoint < HALFWIDTH_KATAKANA_RANGE[0] || codePoint > HALFWIDTH_KATAKANA_RANGE[1]) { return null; }
return HALFWIDTH_KATAKANA_TO_HIRAGANA[codePoint - HALFWIDTH_KATAKANA_RANGE[0]] || null;
}
// Halfwidth katakana is kana here but not to the rest of the pipeline
// (known-word matching and frequency lookups only fold fullwidth), so a
// reading taken from halfwidth text is written the way the fullwidth
// katakana path already writes it. NFKC rather than the per-code-point
// table: this is the one place where nothing indexes back into the
// result, so a voiced pair (カ + ゙) can compose into the single ガ it
// means instead of leaving a stray combining mark in the reading. Scoped
// to the halfwidth runs, because NFKC over everything else rewrites
// characters that have nothing to do with kana (① → 1, ㍑ → リットル).
function convertHalfwidthKanaToKatakana(text) {
return text.replace(/[ヲ-゚]+/g, (run) => run.normalize("NFKC"));
}
// Han ranges come from the shared table so the scan walk and the character
// dictionary agree on what a kanji is (supplementary planes included).
// Halfwidth katakana counts as Japanese text: a name written that way has
// to reach the greedy pre-pass, which has its own handling for it.
const JAPANESE_RANGES = [[0x3040, 0x30ff], [0xff66, 0xff9f], ...${JSON.stringify(HAN_CODE_POINT_RANGES)}];
function isCodePointInRange(codePoint, range) { return codePoint >= range[0] && codePoint <= range[1]; }
function isCodePointInRanges(codePoint, ranges) { return ranges.some((range) => isCodePointInRange(codePoint, range)); }
function isCodePointKana(codePoint) { return isCodePointInRanges(codePoint, KANA_RANGES); }
function isCodePointJapanese(codePoint) { return isCodePointInRanges(codePoint, JAPANESE_RANGES); }
`;
@@ -0,0 +1,79 @@
// Match selection for the injected scan runtime: picks the headword a position
// tokenizes to, and the longest name or generic match in a window, which is how
// the greedy name pre-pass decides what to reserve.
export const YOMITAN_MATCH_SELECTION_HELPERS = String.raw`
function findLongestNameMatch(dictionaryEntries, textWindow) {
let best = null;
for (const dictionaryEntry of dictionaryEntries || []) {
if (!isCurrentMediaNameDictionaryEntry(dictionaryEntry)) { continue; }
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (let headwordIndex = 0; headwordIndex < headwords.length; headwordIndex += 1) {
const headword = headwords[headwordIndex];
for (const src of headword?.sources || []) {
if (src.matchType !== 'exact' || src.isPrimary !== true) { continue; }
const originalText = typeof src.originalText === 'string' ? src.originalText : '';
if (!originalText || !textWindow.startsWith(originalText)) { continue; }
if (best === null || originalText.length > best.sourceLength) {
best = { dictionaryEntry, headword, headwordIndex, sourceLength: originalText.length };
}
}
}
}
return best;
}
function findLongestGenericMatchLength(dictionaryEntries, textWindow) {
let best = 0;
for (const dictionaryEntry of dictionaryEntries || []) {
if (isNameDictionaryEntry(dictionaryEntry)) { continue; }
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (const headword of headwords) {
for (const src of headword?.sources || []) {
if (src.matchType !== 'exact' || src.isPrimary !== true) { continue; }
const originalText = typeof src.originalText === 'string' ? src.originalText : '';
if (!originalText || !textWindow.startsWith(originalText)) { continue; }
if (originalText.length > best) { best = originalText.length; }
}
}
}
return best;
}
function getPreferredHeadword(dictionaryEntries, token, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
const currentMediaDictionaryEntries =
currentCharacterDictionaryMediaId === null
? (dictionaryEntries || [])
: (dictionaryEntries || []).filter((entry) => {
if (!isNameDictionaryEntry(entry)) { return true; }
return isCurrentMediaNameDictionaryEntry(entry);
});
const exactPrimaryMatches = collectExactHeadwordMatches(currentMediaDictionaryEntries, token, true);
let matchedNameDictionary = false;
if (includeNameMatchMetadata) {
// Every match already comes from currentMediaDictionaryEntries, so
// classifying its own entry is enough.
for (const match of exactPrimaryMatches) {
if (!isCurrentMediaNameDictionaryEntry(match.dictionaryEntry)) { continue; }
matchedNameDictionary = true;
break;
}
}
const preferredMatch = exactPrimaryMatches[0];
if (preferredMatch) {
const exactFrequencyMatches = collectExactHeadwordMatches(currentMediaDictionaryEntries, token, false)
.filter((match) => sameHeadword(match, preferredMatch));
return {
term: preferredMatch.headword.term,
reading: preferredMatch.headword.reading,
wordClasses: normalizeWordClasses(preferredMatch.headword),
isNameMatch:
matchedNameDictionary || isCurrentMediaNameDictionaryEntry(preferredMatch.dictionaryEntry),
frequencyRank: getBestFrequencyRankForMatches(
exactFrequencyMatches.length > 0 ? exactFrequencyMatches : exactPrimaryMatches,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
)
};
}
return null;
}
`;
File diff suppressed because it is too large Load Diff
@@ -3,6 +3,14 @@ import * as fs from 'fs';
import * as http from 'http';
import * as path from 'path';
import { selectYomitanParseTokens } from './parser-selection-stage';
import {
buildYomitanScanCallScript,
buildYomitanScanNameCandidatesScript,
CHARACTER_DICTIONARY_TITLE_PREFIX,
YOMITAN_SCAN_RUNTIME_INSTALL_SCRIPT,
YOMITAN_SCAN_RUNTIME_MISSING_SENTINEL,
type YomitanFrequencyMode,
} from './yomitan-scan-runtime-script';
interface LoggerLike {
error: (message: string, ...args: unknown[]) => void;
@@ -22,8 +30,6 @@ interface YomitanParserRuntimeDeps {
createYomitanExtensionWindow?: (pageName: string) => Promise<BrowserWindow | null>;
}
type YomitanFrequencyMode = 'occurrence-based' | 'rank-based';
export interface YomitanDictionaryInfo {
title: string;
revision?: string | number;
@@ -74,13 +80,19 @@ export interface YomitanAddNoteResult {
}
const DEFAULT_YOMITAN_SCAN_LENGTH = 40;
const CHARACTER_DICTIONARY_TITLE_PREFIX = 'SubMiner Character Dictionary';
const yomitanProfileMetadataByWindow = new WeakMap<BrowserWindow, YomitanProfileMetadata>();
const yomitanProfileDiagnosticsLoggedByWindow = new WeakSet<BrowserWindow>();
const yomitanFrequencyCacheByWindow = new WeakMap<
BrowserWindow,
Map<string, YomitanTermFrequency[]>
>();
// Epoch passed with every scan request; the in-window termsFind cache clears
// itself when the epoch changes (dictionary imports, settings changes).
const yomitanScanCacheEpochByWindow = new WeakMap<BrowserWindow, number>();
function getYomitanScanCacheEpoch(window: BrowserWindow): number {
return yomitanScanCacheEpochByWindow.get(window) ?? 0;
}
function isObject(value: unknown): value is Record<string, unknown> {
return Boolean(value && typeof value === 'object');
@@ -99,6 +111,7 @@ function isScanTokenArray(value: unknown): value is YomitanScanToken[] {
typeof entry.startPos === 'number' &&
typeof entry.endPos === 'number' &&
(entry.isNameMatch === undefined || typeof entry.isNameMatch === 'boolean') &&
(entry.isUnparsedRun === undefined || typeof entry.isUnparsedRun === 'boolean') &&
(entry.frequencyRank === undefined || typeof entry.frequencyRank === 'number') &&
(entry.wordClasses === undefined ||
(Array.isArray(entry.wordClasses) &&
@@ -107,13 +120,9 @@ function isScanTokenArray(value: unknown): value is YomitanScanToken[] {
);
}
function scanTokenSpanKey(token: YomitanScanToken): string {
return `${token.startPos}:${token.endPos}:${token.surface}`;
}
// Maps a parse-selected token to the scanner-token shape carried out of the
// parser runtime. Shared by both selectYomitanParseTokens fallback paths so the
// projected fields stay in sync as the shape changes.
// parser runtime, used by the parseText fallback path when the in-window
// scanner is unavailable.
function toYomitanScanToken(token: {
surface: string;
reading: string;
@@ -132,66 +141,6 @@ function toYomitanScanToken(token: {
};
}
// parseText segmentation is authoritative (it emits filler chunks for text the
// termsFind scanner skips), but only the termsFind scanner carries annotation
// metadata (isNameMatch, frequencyRank, headwordReading, wordClasses). Graft
// scanner tokens onto the parseText segmentation per matching span so one
// unmatched chunk degrades only itself instead of dropping the whole line's
// metadata.
//
// Exception: character-name tokens. The greedy name scan can re-segment text
// around a name (e.g. とヨータ → と + ヨータ instead of とヨー + タ), so
// parseText segmentation cannot be authoritative there. Each name span is
// expanded until it aligns with token boundaries in both segmentations, then
// the parse tokens inside are replaced with the scanner tokens.
function mergeScannerTokensIntoParseTokens(
parseScanTokens: YomitanScanToken[],
scannerTokens: YomitanScanToken[],
): YomitanScanToken[] {
const scannerTokensBySpan = new Map<string, YomitanScanToken>();
for (const token of scannerTokens) {
scannerTokensBySpan.set(scanTokenSpanKey(token), token);
}
const graftedTokens = parseScanTokens.map(
(token) => scannerTokensBySpan.get(scanTokenSpanKey(token)) ?? token,
);
const nameTokens = scannerTokens.filter((token) => token.isNameMatch === true);
if (nameTokens.length === 0) {
return graftedTokens;
}
const regions = nameTokens.map((token) => ({ start: token.startPos, end: token.endPos }));
const allTokens = [...parseScanTokens, ...scannerTokens];
let expanded = true;
while (expanded) {
expanded = false;
for (const region of regions) {
for (const token of allTokens) {
const overlaps = token.startPos < region.end && token.endPos > region.start;
const extendsBeyond = token.startPos < region.start || token.endPos > region.end;
if (overlaps && extendsBeyond) {
region.start = Math.min(region.start, token.startPos);
region.end = Math.max(region.end, token.endPos);
expanded = true;
}
}
}
}
const isInsideNameRegion = (token: YomitanScanToken): boolean =>
regions.some((region) => token.startPos >= region.start && token.endPos <= region.end);
const merged = graftedTokens.filter((token) => !isInsideNameRegion(token));
for (const token of scannerTokens) {
if (isInsideNameRegion(token)) {
merged.push(token);
}
}
merged.sort((a, b) => a.startPos - b.startPos || a.endPos - b.endPos);
return merged;
}
function makeTermReadingCacheKey(term: string, reading: string | null): string {
return `${term}\u0000${reading ?? ''}`;
}
@@ -208,6 +157,7 @@ function getWindowFrequencyCache(window: BrowserWindow): Map<string, YomitanTerm
function clearWindowCaches(window: BrowserWindow): void {
yomitanProfileMetadataByWindow.delete(window);
yomitanFrequencyCacheByWindow.delete(window);
yomitanScanCacheEpochByWindow.set(window, getYomitanScanCacheEpoch(window) + 1);
}
export function clearYomitanParserCachesForWindow(window: BrowserWindow): void {
clearWindowCaches(window);
@@ -704,6 +654,10 @@ async function ensureYomitanParserWindow(
if (readyPromise) {
await readyPromise;
}
// Eagerly install the scan runtime so the first subtitle line does not
// pay the install round trip; failures fall back to the per-request
// install-and-retry path.
await installYomitanScanRuntime(parserWindow).catch(() => {});
return true;
} catch (err) {
@@ -877,668 +831,42 @@ async function serveDictionaryZipOnce<T>(
}
}
const YOMITAN_SCANNING_HELPERS = String.raw`
const HIRAGANA_CONVERSION_RANGE = [0x3041, 0x3096];
const KATAKANA_CONVERSION_RANGE = [0x30a1, 0x30f6];
const KANA_PROLONGED_SOUND_MARK_CODE_POINT = 0x30fc;
const KATAKANA_SMALL_KA_CODE_POINT = 0x30f5;
const KATAKANA_SMALL_KE_CODE_POINT = 0x30f6;
const KANA_RANGES = [[0x3040, 0x309f], [0x30a0, 0x30ff]];
const JAPANESE_RANGES = [[0x3040, 0x30ff], [0x3400, 0x9fff]];
function isCodePointInRange(codePoint, range) { return codePoint >= range[0] && codePoint <= range[1]; }
function isCodePointInRanges(codePoint, ranges) { return ranges.some((range) => isCodePointInRange(codePoint, range)); }
function isCodePointKana(codePoint) { return isCodePointInRanges(codePoint, KANA_RANGES); }
function isCodePointJapanese(codePoint) { return isCodePointInRanges(codePoint, JAPANESE_RANGES); }
function createFuriganaSegment(text, reading) { return {text, reading}; }
function getSegmentReadingContribution(segment) {
if (typeof segment.reading === "string" && segment.reading.length > 0) { return segment.reading; }
const segmentText = typeof segment.text === "string" ? segment.text : "";
const isKanaOnly = segmentText.length > 0 && [...segmentText].every((char) => isCodePointKana(char.codePointAt(0)));
return isKanaOnly ? segmentText : "";
}
function getProlongedHiragana(previousCharacter) {
switch (previousCharacter) {
case "あ": case "か": case "が": case "さ": case "ざ": case "た": case "だ": case "な": case "は": case "ば": case "ぱ": case "ま": case "や": case "ら": case "わ": case "ぁ": case "ゃ": case "ゎ": return "あ";
case "い": case "き": case "ぎ": case "し": case "じ": case "ち": case "ぢ": case "に": case "ひ": case "び": case "ぴ": case "み": case "り": case "ぃ": return "い";
case "う": case "く": case "ぐ": case "す": case "ず": case "つ": case "づ": case "ぬ": case "ふ": case "ぶ": case "ぷ": case "む": case "ゆ": case "る": case "ぅ": case "ゅ": return "う";
case "え": case "け": case "げ": case "せ": case "ぜ": case "て": case "で": case "ね": case "へ": case "べ": case "ぺ": case "め": case "れ": case "ぇ": return "え";
case "お": case "こ": case "ご": case "そ": case "ぞ": case "と": case "ど": case "の": case "ほ": case "ぼ": case "ぽ": case "も": case "よ": case "ろ": case "を": case "ぉ": case "ょ": return "う";
default: return null;
}
}
function getFuriganaKanaSegments(text, reading) {
const newSegments = [];
let start = 0;
let state = (reading[0] === text[0]);
for (let i = 1; i < text.length; ++i) {
const newState = (reading[i] === text[i]);
if (state === newState) { continue; }
newSegments.push(createFuriganaSegment(text.substring(start, i), state ? '' : reading.substring(start, i)));
state = newState;
start = i;
}
newSegments.push(createFuriganaSegment(text.substring(start), state ? '' : reading.substring(start)));
return newSegments;
}
function convertKatakanaToHiragana(text, keepProlongedSoundMarks = false) {
let result = '';
const offset = (HIRAGANA_CONVERSION_RANGE[0] - KATAKANA_CONVERSION_RANGE[0]);
for (let char of text) {
const codePoint = char.codePointAt(0);
switch (codePoint) {
case KATAKANA_SMALL_KA_CODE_POINT:
case KATAKANA_SMALL_KE_CODE_POINT:
break;
case KANA_PROLONGED_SOUND_MARK_CODE_POINT:
if (!keepProlongedSoundMarks && result.length > 0) {
const char2 = getProlongedHiragana(result[result.length - 1]);
if (char2 !== null) { char = char2; }
}
break;
default:
if (isCodePointInRange(codePoint, KATAKANA_CONVERSION_RANGE)) {
char = String.fromCodePoint(codePoint + offset);
}
break;
}
result += char;
}
return result;
}
function segmentizeFurigana(reading, readingNormalized, groups, groupsStart) {
const groupCount = groups.length - groupsStart;
if (groupCount <= 0) { return reading.length === 0 ? [] : null; }
const group = groups[groupsStart];
const {isKana, text} = group;
if (isKana) {
if (group.textNormalized !== null && readingNormalized.startsWith(group.textNormalized)) {
const segments = segmentizeFurigana(reading.substring(text.length), readingNormalized.substring(text.length), groups, groupsStart + 1);
if (segments !== null) {
if (reading.startsWith(text)) { segments.unshift(createFuriganaSegment(text, '')); }
else { segments.unshift(...getFuriganaKanaSegments(text, reading)); }
return segments;
}
}
return null;
}
let result = null;
for (let i = reading.length; i >= text.length; --i) {
const segments = segmentizeFurigana(reading.substring(i), readingNormalized.substring(i), groups, groupsStart + 1);
if (segments !== null) {
if (result !== null) { return null; }
segments.unshift(createFuriganaSegment(text, reading.substring(0, i)));
result = segments;
}
if (groupCount === 1) { break; }
}
return result;
}
function distributeFurigana(term, reading) {
if (reading === term) { return [createFuriganaSegment(term, '')]; }
const groups = [];
let groupPre = null;
let isKanaPre = null;
for (const c of term) {
const isKana = isCodePointKana(c.codePointAt(0));
if (isKana === isKanaPre) { groupPre.text += c; }
else {
groupPre = {isKana, text: c, textNormalized: null};
groups.push(groupPre);
isKanaPre = isKana;
}
}
for (const group of groups) {
if (group.isKana) { group.textNormalized = convertKatakanaToHiragana(group.text); }
}
const segments = segmentizeFurigana(reading, convertKatakanaToHiragana(reading), groups, 0);
return segments !== null ? segments : [createFuriganaSegment(term, reading)];
}
function getStemLength(text1, text2) {
const minLength = Math.min(text1.length, text2.length);
if (minLength === 0) { return 0; }
let i = 0;
while (true) {
const char1 = text1.codePointAt(i);
const char2 = text2.codePointAt(i);
if (char1 !== char2) { break; }
const charLength = String.fromCodePoint(char1).length;
i += charLength;
if (i >= minLength) {
if (i > minLength) { i -= charLength; }
break;
}
}
return i;
}
function distributeFuriganaInflected(term, reading, source) {
const termNormalized = convertKatakanaToHiragana(term);
const readingNormalized = convertKatakanaToHiragana(reading);
const sourceNormalized = convertKatakanaToHiragana(source);
let mainText = term;
let stemLength = getStemLength(termNormalized, sourceNormalized);
const readingStemLength = getStemLength(readingNormalized, sourceNormalized);
if (readingStemLength > 0 && readingStemLength >= stemLength) {
mainText = reading;
stemLength = readingStemLength;
reading = source.substring(0, stemLength) + reading.substring(stemLength);
}
const segments = [];
if (stemLength > 0) {
mainText = source.substring(0, stemLength) + mainText.substring(stemLength);
const segments2 = distributeFurigana(mainText, reading);
let consumed = 0;
for (const segment of segments2) {
const start = consumed;
consumed += segment.text.length;
if (consumed < stemLength) { segments.push(segment); }
else if (consumed === stemLength) { segments.push(segment); break; }
else {
if (start < stemLength) { segments.push(createFuriganaSegment(mainText.substring(start, stemLength), '')); }
break;
}
}
}
if (stemLength < source.length) {
const remainder = source.substring(stemLength);
const last = segments[segments.length - 1];
if (last && last.reading.length === 0) { last.text += remainder; }
else { segments.push(createFuriganaSegment(remainder, '')); }
}
return segments;
}
function parsePositiveFrequencyNumber(value) {
if (typeof value === 'number' && Number.isFinite(value) && value > 0) {
return Math.max(1, Math.floor(value));
}
if (typeof value === 'string') {
const numericMatch = value.trim().match(/[+-]?(\d+(\.\d*)?|\.\d+)([eE][+-]?\d+)?/)?.[0];
if (!numericMatch) { return null; }
const parsed = Number.parseFloat(numericMatch);
if (!Number.isFinite(parsed) || parsed <= 0) { return null; }
return Math.max(1, Math.floor(parsed));
}
if (Array.isArray(value)) {
for (const item of value) {
const parsed = parsePositiveFrequencyNumber(item);
if (parsed !== null) { return parsed; }
}
}
return null;
}
function parseDisplayFrequencyNumber(value) {
if (typeof value === 'string') {
const leadingDigits = value.trim().match(/^\d+/)?.[0];
if (!leadingDigits) { return null; }
const parsed = Number.parseInt(leadingDigits, 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : null;
}
return parsePositiveFrequencyNumber(value);
}
function getFrequencyDictionaryName(frequency) {
const candidates = [
frequency?.dictionary,
frequency?.dictionaryName,
frequency?.name,
frequency?.title,
frequency?.dictionaryTitle,
frequency?.dictionaryAlias
];
for (const candidate of candidates) {
if (typeof candidate === 'string' && candidate.trim().length > 0) {
return candidate.trim();
}
}
return null;
}
function getBestFrequencyRank(dictionaryEntry, headwordIndex, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
let best = null;
const headwordCount = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords.length : 0;
for (const frequency of dictionaryEntry?.frequencies || []) {
if (!frequency || typeof frequency !== 'object') { continue; }
const frequencyHeadwordIndex = frequency.headwordIndex;
if (typeof frequencyHeadwordIndex === 'number') {
if (frequencyHeadwordIndex !== headwordIndex) { continue; }
} else if (headwordCount > 1) {
continue;
}
const dictionary = getFrequencyDictionaryName(frequency);
if (!dictionary) { continue; }
if (dictionaryFrequencyModeByName[dictionary] === 'occurrence-based') { continue; }
const rank =
parseDisplayFrequencyNumber(frequency.displayValue) ??
parsePositiveFrequencyNumber(frequency.frequency);
if (rank === null) { continue; }
const priorityRaw = dictionaryPriorityByName[dictionary];
const fallbackPriority =
typeof frequency.dictionaryIndex === 'number' && Number.isFinite(frequency.dictionaryIndex)
? Math.max(0, Math.floor(frequency.dictionaryIndex))
: Number.MAX_SAFE_INTEGER;
const priority =
typeof priorityRaw === 'number' && Number.isFinite(priorityRaw)
? Math.max(0, Math.floor(priorityRaw))
: fallbackPriority;
if (best === null || priority < best.priority || (priority === best.priority && rank < best.rank)) {
best = { priority, rank };
}
}
return best?.rank ?? null;
}
function hasExactSource(headword, token, requirePrimary) {
for (const src of headword.sources || []) {
if (src.originalText !== token) { continue; }
if (requirePrimary && !src.isPrimary) { continue; }
if (src.matchType !== 'exact') { continue; }
return true;
}
return false;
}
function collectExactHeadwordMatches(dictionaryEntries, token, requirePrimary) {
const matches = [];
for (const dictionaryEntry of dictionaryEntries || []) {
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (let headwordIndex = 0; headwordIndex < headwords.length; headwordIndex += 1) {
const headword = headwords[headwordIndex];
if (!hasExactSource(headword, token, requirePrimary)) { continue; }
matches.push({ dictionaryEntry, headword, headwordIndex });
}
}
return matches;
}
function sameHeadword(match, preferredMatch) {
if (!match || !preferredMatch) {
return false;
}
if (match.headword?.term !== preferredMatch.headword?.term) {
return false;
}
const matchReading = typeof match.headword?.reading === 'string' ? match.headword.reading : '';
const preferredReading =
typeof preferredMatch.headword?.reading === 'string' ? preferredMatch.headword.reading : '';
if (!matchReading || !preferredReading) {
return true;
}
return matchReading === preferredReading;
}
function getBestFrequencyRankForMatches(matches, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
let best = null;
for (const match of matches) {
const rank = getBestFrequencyRank(
match.dictionaryEntry,
match.headwordIndex,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
);
if (rank === null) { continue; }
if (best === null || rank < best) {
best = rank;
}
}
return best;
}
function normalizeWordClasses(headword) {
if (!Array.isArray(headword?.wordClasses)) { return undefined; }
const classes = headword.wordClasses.filter((wordClass) => typeof wordClass === "string" && wordClass.trim().length > 0);
return classes.length > 0 ? classes : undefined;
}
function appendDictionaryNames(target, value) {
if (!value || typeof value !== 'object') {
return;
}
const candidates = [
value.dictionary,
value.dictionaryName,
value.name,
value.title,
value.dictionaryTitle,
value.dictionaryAlias
];
for (const candidate of candidates) {
if (typeof candidate === 'string' && candidate.trim().length > 0) {
target.push(candidate.trim());
}
}
}
function getDictionaryEntryNames(entry) {
const names = [];
appendDictionaryNames(names, entry);
for (const definition of entry?.definitions || []) {
appendDictionaryNames(names, definition);
}
for (const frequency of entry?.frequencies || []) {
appendDictionaryNames(names, frequency);
}
for (const pronunciation of entry?.pronunciations || []) {
appendDictionaryNames(names, pronunciation);
}
return names;
}
function isNameDictionaryEntry(entry) {
if (!includeNameMatchMetadata || !entry || typeof entry !== 'object') {
return false;
}
return getDictionaryEntryNames(entry).some((name) => name.startsWith(${JSON.stringify(CHARACTER_DICTIONARY_TITLE_PREFIX)}));
}
function parseSubMinerMediaIdFromString(value) {
const imageMatch = value.match(/\bimg\/m(\d+)-/i);
if (imageMatch) {
const parsed = Number.parseInt(imageMatch[1], 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
const titleMatch = value.match(/${CHARACTER_DICTIONARY_TITLE_PREFIX}[^\d]*(?:AniList\s*)?(\d+)/i);
if (titleMatch) {
const parsed = Number.parseInt(titleMatch[1], 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
return null;
}
function parseSubMinerMediaIdCandidate(value) {
if (typeof value === 'number' && Number.isSafeInteger(value) && value > 0) {
return value;
}
if (typeof value === 'string' && /^\d+$/.test(value.trim())) {
const parsed = Number.parseInt(value.trim(), 10);
if (Number.isSafeInteger(parsed) && parsed > 0) { return parsed; }
}
return null;
}
function collectSubMinerMediaIds(value, target) {
if (typeof value === 'string') {
const parsed = parseSubMinerMediaIdFromString(value);
if (parsed !== null) { target.add(parsed); }
return;
}
if (!value || typeof value !== 'object') {
return;
}
if (Array.isArray(value)) {
for (const item of value) { collectSubMinerMediaIds(item, target); }
return;
}
const mediaIdCandidates = [
value.subminerMediaId,
value.subMinerMediaId,
value.characterDictionaryMediaId,
value.data?.subminerMediaId,
value.data?.subMinerMediaId,
value.data?.characterDictionaryMediaId
];
for (const candidate of mediaIdCandidates) {
const parsed = parseSubMinerMediaIdCandidate(candidate);
if (parsed !== null) { target.add(parsed); }
}
for (const child of Object.values(value)) {
collectSubMinerMediaIds(child, target);
}
}
function getSubMinerMediaIds(entry) {
const mediaIds = new Set();
collectSubMinerMediaIds(entry, mediaIds);
return mediaIds;
}
function isCurrentMediaNameDictionaryEntry(entry) {
if (!isNameDictionaryEntry(entry)) {
return false;
}
if (currentCharacterDictionaryMediaId === null) {
return true;
}
const mediaIds = getSubMinerMediaIds(entry);
return mediaIds.size === 0 || mediaIds.has(currentCharacterDictionaryMediaId);
}
function findLongestNameMatch(dictionaryEntries, textWindow) {
let best = null;
for (const dictionaryEntry of dictionaryEntries || []) {
if (!isCurrentMediaNameDictionaryEntry(dictionaryEntry)) { continue; }
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (let headwordIndex = 0; headwordIndex < headwords.length; headwordIndex += 1) {
const headword = headwords[headwordIndex];
for (const src of headword?.sources || []) {
if (src.matchType !== 'exact' || src.isPrimary !== true) { continue; }
const originalText = typeof src.originalText === 'string' ? src.originalText : '';
if (!originalText || !textWindow.startsWith(originalText)) { continue; }
if (best === null || originalText.length > best.sourceLength) {
best = { dictionaryEntry, headword, headwordIndex, sourceLength: originalText.length };
}
}
}
}
return best;
}
function findLongestGenericMatchLength(dictionaryEntries, textWindow) {
let best = 0;
for (const dictionaryEntry of dictionaryEntries || []) {
if (isNameDictionaryEntry(dictionaryEntry)) { continue; }
const headwords = Array.isArray(dictionaryEntry?.headwords) ? dictionaryEntry.headwords : [];
for (const headword of headwords) {
for (const src of headword?.sources || []) {
if (src.matchType !== 'exact' || src.isPrimary !== true) { continue; }
const originalText = typeof src.originalText === 'string' ? src.originalText : '';
if (!originalText || !textWindow.startsWith(originalText)) { continue; }
if (originalText.length > best) { best = originalText.length; }
}
}
}
return best;
}
function getPreferredHeadword(dictionaryEntries, token, dictionaryPriorityByName, dictionaryFrequencyModeByName) {
const currentMediaDictionaryEntries =
currentCharacterDictionaryMediaId === null
? (dictionaryEntries || [])
: (dictionaryEntries || []).filter((entry) => {
if (!isNameDictionaryEntry(entry)) { return true; }
return isCurrentMediaNameDictionaryEntry(entry);
});
const exactPrimaryMatches = collectExactHeadwordMatches(currentMediaDictionaryEntries, token, true);
let matchedNameDictionary = false;
if (includeNameMatchMetadata) {
for (const dictionaryEntry of currentMediaDictionaryEntries || []) {
if (!isCurrentMediaNameDictionaryEntry(dictionaryEntry)) { continue; }
for (const match of exactPrimaryMatches) {
if (match.dictionaryEntry !== dictionaryEntry) { continue; }
matchedNameDictionary = true;
break;
}
if (matchedNameDictionary) { break; }
}
}
const preferredMatch = exactPrimaryMatches[0];
if (preferredMatch) {
const exactFrequencyMatches = collectExactHeadwordMatches(currentMediaDictionaryEntries, token, false)
.filter((match) => sameHeadword(match, preferredMatch));
return {
term: preferredMatch.headword.term,
reading: preferredMatch.headword.reading,
wordClasses: normalizeWordClasses(preferredMatch.headword),
isNameMatch:
matchedNameDictionary || isCurrentMediaNameDictionaryEntry(preferredMatch.dictionaryEntry),
frequencyRank: getBestFrequencyRankForMatches(
exactFrequencyMatches.length > 0 ? exactFrequencyMatches : exactPrimaryMatches,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
)
};
}
return null;
}
`;
async function installYomitanScanRuntime(parserWindow: BrowserWindow): Promise<void> {
await parserWindow.webContents.executeJavaScript(YOMITAN_SCAN_RUNTIME_INSTALL_SCRIPT, true);
// A fresh runtime has no candidate list; force the next scan to reinstall it.
yomitanScanNameCandidateKeyByWindow.delete(parserWindow);
}
function buildYomitanScanningScript(
text: string,
profileIndex: number,
scanLength: number,
includeNameMatchMetadata: boolean,
greedyNameScanEnabled: boolean,
currentCharacterDictionaryMediaId: number | null,
dictionaryPriorityByName: Record<string, number>,
dictionaryFrequencyModeByName: Partial<Record<string, YomitanFrequencyMode>>,
): string {
return `
(async () => {
const invoke = (action, params) =>
new Promise((resolve, reject) => {
chrome.runtime.sendMessage({ action, params }, (response) => {
if (chrome.runtime.lastError) {
reject(new Error(chrome.runtime.lastError.message));
// Key of the character-name candidate list currently installed in each parser
// window, so an unchanged list costs nothing per line.
const yomitanScanNameCandidateKeyByWindow = new WeakMap<BrowserWindow, string>();
async function ensureYomitanScanNameCandidates(
parserWindow: BrowserWindow,
nameCandidates: { key: string; forms: string[] } | null,
logger: LoggerLike,
): Promise<void> {
const installedKey = yomitanScanNameCandidateKeyByWindow.get(parserWindow);
const nextKey = nameCandidates?.key ?? '';
if (installedKey === nextKey) {
return;
}
if (!response || typeof response !== "object") {
reject(new Error("Invalid response from Yomitan backend"));
return;
}
if (response.error) {
reject(new Error(response.error.message || "Yomitan backend error"));
return;
}
resolve(response.result);
});
});
${YOMITAN_SCANNING_HELPERS}
const includeNameMatchMetadata = ${includeNameMatchMetadata ? 'true' : 'false'};
const greedyNameScanEnabled = ${greedyNameScanEnabled ? 'true' : 'false'};
const currentCharacterDictionaryMediaId = ${
currentCharacterDictionaryMediaId !== null
? String(currentCharacterDictionaryMediaId)
: 'null'
};
const dictionaryPriorityByName = ${JSON.stringify(dictionaryPriorityByName)};
const dictionaryFrequencyModeByName = ${JSON.stringify(dictionaryFrequencyModeByName)};
const text = ${JSON.stringify(text)};
const details = {matchType: "exact", deinflect: true};
const tokens = [];
const termsFindCache = new Map();
async function termsFindAt(position, windowLength) {
const cacheKey = position + ":" + windowLength;
const cached = termsFindCache.get(cacheKey);
if (cached) { return cached; }
const substring = text.substring(position, position + windowLength);
const result = await invoke("termsFind", { text: substring, details, optionsContext: { index: ${profileIndex} } });
termsFindCache.set(cacheKey, result);
return result;
}
function buildScanToken(position, source, preferredHeadword) {
const reading = typeof preferredHeadword.reading === "string" ? preferredHeadword.reading : "";
const segments = distributeFuriganaInflected(preferredHeadword.term, reading, source);
const tokenPayload = {
surface: segments.map((segment) => segment.text).join("") || source,
reading: segments.map(getSegmentReadingContribution).join(""),
headword: preferredHeadword.term,
headwordReading: reading || undefined,
startPos: position,
endPos: position + source.length,
isNameMatch: includeNameMatchMetadata && preferredHeadword.isNameMatch === true,
frequencyRank:
typeof preferredHeadword.frequencyRank === "number" && Number.isFinite(preferredHeadword.frequencyRank)
? Math.max(1, Math.floor(preferredHeadword.frequencyRank))
: undefined,
};
if (Array.isArray(preferredHeadword.wordClasses) && preferredHeadword.wordClasses.length > 0) {
tokenPayload.wordClasses = preferredHeadword.wordClasses;
}
return tokenPayload;
}
async function findTokenAt(position, windowLength) {
const codePoint = text.codePointAt(position);
const character = String.fromCodePoint(codePoint);
const result = await termsFindAt(position, windowLength);
const dictionaryEntries = Array.isArray(result?.dictionaryEntries) ? result.dictionaryEntries : [];
const originalTextLength = typeof result?.originalTextLength === "number" ? result.originalTextLength : 0;
if (dictionaryEntries.length === 0 || originalTextLength <= 0 || (originalTextLength === character.length && !isCodePointJapanese(codePoint))) {
return { token: null, matchedLength: 0 };
}
const source = text.substring(position, position + originalTextLength);
const preferredHeadword = getPreferredHeadword(
dictionaryEntries,
source,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
try {
await parserWindow.webContents.executeJavaScript(
buildYomitanScanNameCandidatesScript(nameCandidates),
true,
);
if (!preferredHeadword || typeof preferredHeadword.term !== "string") {
return { token: null, matchedLength: originalTextLength };
yomitanScanNameCandidateKeyByWindow.set(parserWindow, nextKey);
} catch (err) {
// The scan falls back to checking every position when the list is absent,
// so a failed install costs speed, never a missed name.
logger.warn?.(
'Failed to install Yomitan character-name scan candidates:',
(err as Error).message,
);
yomitanScanNameCandidateKeyByWindow.delete(parserWindow);
}
return { token: buildScanToken(position, source, preferredHeadword), matchedLength: originalTextLength };
}
// Greedy name pre-pass: character-name matches claim their spans before
// the left-to-right walk, so a longer generic match starting earlier
// (e.g. とヨー → 渡洋) cannot swallow the start of a name (ヨータ).
const nameTokens = [];
if (greedyNameScanEnabled) {
let namePos = 0;
while (namePos < text.length) {
const codePoint = text.codePointAt(namePos);
if (!isCodePointJapanese(codePoint)) {
namePos += String.fromCodePoint(codePoint).length;
continue;
}
const result = await termsFindAt(namePos, ${scanLength});
const dictionaryEntries = Array.isArray(result?.dictionaryEntries) ? result.dictionaryEntries : [];
const textWindow = text.substring(namePos, namePos + ${scanLength});
const nameMatch = findLongestNameMatch(dictionaryEntries, textWindow);
// A name only claims its span when no strictly longer generic word
// starts at the same position (a character named 空 must not split
// 空気). Ties go to the name. Generic matches that start earlier and
// overlap the name are still blocked by the reservation.
if (
!nameMatch ||
findLongestGenericMatchLength(dictionaryEntries, textWindow) > nameMatch.sourceLength
) {
namePos += String.fromCodePoint(codePoint).length;
continue;
}
const source = text.substring(namePos, namePos + nameMatch.sourceLength);
nameTokens.push(buildScanToken(namePos, source, {
term: nameMatch.headword.term,
reading: nameMatch.headword.reading,
wordClasses: normalizeWordClasses(nameMatch.headword),
isNameMatch: true,
frequencyRank: getBestFrequencyRank(
nameMatch.dictionaryEntry,
nameMatch.headwordIndex,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
)
}));
namePos += nameMatch.sourceLength;
}
}
let i = 0;
let nameIndex = 0;
while (i < text.length) {
while (nameIndex < nameTokens.length && nameTokens[nameIndex].startPos < i) { nameIndex += 1; }
const nextNameToken = nameIndex < nameTokens.length ? nameTokens[nameIndex] : null;
if (nextNameToken && nextNameToken.startPos === i) {
tokens.push(nextNameToken);
i = nextNameToken.endPos;
nameIndex += 1;
continue;
}
// Cap the window at the next reserved name span so a generic match
// cannot consume into it.
const windowLength = nextNameToken ? Math.min(${scanLength}, nextNameToken.startPos - i) : ${scanLength};
let attempt = await findTokenAt(i, windowLength);
// Yomitan text normalization can consume characters (whitespace,
// punctuation) beyond the matched term, leaving no headword whose
// source equals the consumed text. Retry with shorter windows so a
// valid prefix term (e.g. a character name before a paren) still
// tokenizes instead of the position being skipped.
let retryLength = Math.min(attempt.matchedLength, windowLength) - 1;
while (!attempt.token && retryLength >= 1) {
const retry = await findTokenAt(i, retryLength);
if (retry.token) {
attempt = retry;
break;
}
retryLength = Math.min(retryLength - 1, retry.matchedLength - 1);
}
if (attempt.token) {
tokens.push(attempt.token);
i += attempt.matchedLength;
continue;
}
i += String.fromCodePoint(text.codePointAt(i)).length;
}
return tokens;
})();
`;
}
export async function requestYomitanParseResults(
@@ -1635,6 +963,20 @@ export async function requestYomitanParseResults(
}
}
// parseText fallback for when the in-window scanner cannot run (script eval
// failure, unexpected payload). The scanner walk is the primary tokenizer and
// emits its own filler runs, so this extra full parse only happens on errors.
async function requestYomitanParseFallbackTokens(
text: string,
deps: YomitanParserRuntimeDeps,
logger: LoggerLike,
): Promise<YomitanScanToken[] | null> {
const parseResults = await requestYomitanParseResults(text, deps, logger);
const selectedTokens = selectYomitanParseTokens(parseResults, () => false, 'headword');
const parseScanTokens = selectedTokens?.map(toYomitanScanToken) ?? null;
return parseScanTokens && parseScanTokens.length > 0 ? parseScanTokens : null;
}
export async function requestYomitanScanTokens(
text: string,
deps: YomitanParserRuntimeDeps,
@@ -1642,6 +984,7 @@ export async function requestYomitanScanTokens(
options?: {
includeNameMatchMetadata?: boolean;
currentCharacterDictionaryMediaId?: number | null;
nameCandidates?: { key: string; forms: string[] } | null;
},
): Promise<YomitanScanToken[] | null> {
const yomitanExt = deps.getYomitanExt();
@@ -1655,10 +998,6 @@ export async function requestYomitanScanTokens(
return null;
}
const parseResults = await requestYomitanParseResults(text, deps, logger);
const selectedParseTokens = selectYomitanParseTokens(parseResults, () => false, 'headword');
const parseScanTokens = selectedParseTokens?.map(toYomitanScanToken) ?? null;
const metadata = await requestYomitanProfileMetadata(parserWindow, logger);
const profileIndex = metadata?.profileIndex ?? 0;
const scanLength = metadata?.scanLength ?? DEFAULT_YOMITAN_SCAN_LENGTH;
@@ -1669,44 +1008,63 @@ export async function requestYomitanScanTokens(
name.startsWith(CHARACTER_DICTIONARY_TITLE_PREFIX),
);
try {
const rawResult = await parserWindow.webContents.executeJavaScript(
buildYomitanScanningScript(
// Candidate name forms let the in-page pre-pass skip positions where no
// character name can start. Installed only when it changes (per media), so
// the per-line call stays a single tiny script.
const nameCandidates = greedyNameScanEnabled ? (options?.nameCandidates ?? null) : null;
await ensureYomitanScanNameCandidates(parserWindow, nameCandidates, logger);
const callScript = buildYomitanScanCallScript({
text,
profileIndex,
scanLength,
includeNameMatchMetadata,
greedyNameScanEnabled,
currentCharacterDictionaryMediaId:
typeof options?.currentCharacterDictionaryMediaId === 'number' &&
Number.isFinite(options.currentCharacterDictionaryMediaId) &&
options.currentCharacterDictionaryMediaId > 0
? Math.floor(options.currentCharacterDictionaryMediaId)
: null,
metadata?.dictionaryPriorityByName ?? {},
metadata?.dictionaryFrequencyModeByName ?? {},
),
true,
);
dictionaryPriorityByName: metadata?.dictionaryPriorityByName ?? {},
dictionaryFrequencyModeByName: metadata?.dictionaryFrequencyModeByName ?? {},
cacheEpoch: getYomitanScanCacheEpoch(parserWindow),
nameCandidateKey: nameCandidates?.key ?? null,
});
try {
let rawResult = await parserWindow.webContents.executeJavaScript(callScript, true);
if (rawResult === YOMITAN_SCAN_RUNTIME_MISSING_SENTINEL) {
// First request for this window, or the page reloaded and dropped the
// installed runtime: install and retry once. The candidate list lives in
// the same page state, so it has to be reinstalled alongside it.
await installYomitanScanRuntime(parserWindow);
await ensureYomitanScanNameCandidates(parserWindow, nameCandidates, logger);
rawResult = await parserWindow.webContents.executeJavaScript(callScript, true);
}
// The scanner reports a line where a position ran out of shrinking-window
// retries: it stopped short of windows an uncapped ladder would have tried,
// so a real term may be sitting in an unparsed run. One parseText for the
// line is the bounded way to get the exhaustive answer back (this is the
// parse the scanner replaced, and it only runs for these rare lines).
if (isObject(rawResult) && rawResult.retryBudgetExhausted === true) {
logger.info?.('Yomitan scanner exhausted its retry budget; parsing the line as a fallback.');
const fallbackTokens = await requestYomitanParseFallbackTokens(text, deps, logger);
if (fallbackTokens) {
return fallbackTokens;
}
rawResult = rawResult.tokens;
}
if (isScanTokenArray(rawResult)) {
if (parseScanTokens && parseScanTokens.length > 0) {
return mergeScannerTokensIntoParseTokens(parseScanTokens, rawResult);
// Filler-only results carry no dictionary match; keep the historical
// contract of returning null so callers fall back to raw text.
return rawResult.some((token) => token.isUnparsedRun !== true) ? rawResult : null;
}
return rawResult;
}
if (Array.isArray(rawResult)) {
const selectedTokens = selectYomitanParseTokens(rawResult, () => false, 'headword');
return selectedTokens?.map(toYomitanScanToken) ?? null;
}
if (parseScanTokens && parseScanTokens.length > 0) {
return parseScanTokens;
}
return null;
logger.error('Yomitan scanner returned an unexpected payload; using parseText fallback.');
return await requestYomitanParseFallbackTokens(text, deps, logger);
} catch (err) {
if (parseScanTokens && parseScanTokens.length > 0) {
return parseScanTokens;
}
logger.error('Yomitan scanner request failed:', (err as Error).message);
return null;
return await requestYomitanParseFallbackTokens(text, deps, logger);
}
}
@@ -0,0 +1,563 @@
// In-page Yomitan scan runtime: the scan walk that gets installed once per
// parser window as globalThis.__subminerYomitanScan, plus the tiny per-line
// call script. Kept separate from the host runtime module so the injected
// script text (which is data, not executed here) does not dominate that file;
// the helper bundle it embeds is composed in yomitan-scanning-helpers-script.ts
// from the yomitan-*-script.ts fragments.
import { YOMITAN_SCANNING_HELPERS } from './yomitan-scanning-helpers-script';
export { CHARACTER_DICTIONARY_TITLE_PREFIX } from './yomitan-scanning-helpers-script';
export type YomitanFrequencyMode = 'occurrence-based' | 'rank-based';
// Bump whenever the install script below changes so already-loaded parser
// windows re-install the new scan runtime instead of running the stale one.
export const YOMITAN_SCAN_RUNTIME_VERSION = 12;
export const YOMITAN_SCAN_RUNTIME_MISSING_SENTINEL = '__subminer-yomitan-scan-runtime-missing__';
export interface YomitanScanRequestParams {
text: string;
profileIndex: number;
scanLength: number;
includeNameMatchMetadata: boolean;
greedyNameScanEnabled: boolean;
currentCharacterDictionaryMediaId: number | null;
dictionaryPriorityByName: Record<string, number>;
dictionaryFrequencyModeByName: Partial<Record<string, YomitanFrequencyMode>>;
cacheEpoch: number;
/**
* Key of the character-name candidate list installed for the current media,
* or null to scan every Japanese position (see the pre-pass prefilter).
*/
nameCandidateKey: string | null;
}
// Installed once per parser window (and re-installed after in-page reloads):
// keeps V8 from re-parsing the helper bundle on every subtitle line, and hosts
// the cross-line termsFind cache. Each subtitle line then only evaluates a tiny
// call into globalThis.__subminerYomitanScan.
export const YOMITAN_SCAN_RUNTIME_INSTALL_SCRIPT = String.raw`
(() => {
if (globalThis.__subminerYomitanScanVersion === ${YOMITAN_SCAN_RUNTIME_VERSION}) {
return true;
}
const invoke = (action, params) =>
new Promise((resolve, reject) => {
chrome.runtime.sendMessage({ action, params }, (response) => {
if (chrome.runtime.lastError) {
reject(new Error(chrome.runtime.lastError.message));
return;
}
if (!response || typeof response !== "object") {
reject(new Error("Invalid response from Yomitan backend"));
return;
}
if (response.error) {
reject(new Error(response.error.message || "Yomitan backend error"));
return;
}
resolve(response.result);
});
});
// Cross-line termsFind LRU keyed by profile + substring: subtitle lines
// repeat particles and inflections constantly, so most lookups hit here.
// Entries hold in-flight promises so concurrent identical lookups dedupe.
const termsFindCache = new Map();
// Two bounds. The key count keeps the map itself small; the accumulated
// dictionary-entry count stands in for retained bytes, because a single
// lookup over a common prefix can hold hundreds of entries with their full
// glossaries and a key-count cap alone would not bound that.
const TERMS_FIND_CACHE_LIMIT = 2000;
const TERMS_FIND_CACHE_DICTIONARY_ENTRY_LIMIT = 20000;
let termsFindCacheDictionaryEntries = 0;
let termsFindCacheEpoch = -1;
function dropCachedTermsFind(cacheKey, entry) {
if (termsFindCache.get(cacheKey) !== entry) { return; }
termsFindCache.delete(cacheKey);
termsFindCacheDictionaryEntries -= entry.dictionaryEntryCount;
}
// Runs on insert and again once a lookup resolves: an entry is only worth
// its estimated weight of 1 until then, so a single oversized response
// would otherwise sit in the cache forever, over the limit and reused.
function evictOverflowingTermsFindEntries() {
while (
termsFindCache.size > TERMS_FIND_CACHE_LIMIT ||
termsFindCacheDictionaryEntries > TERMS_FIND_CACHE_DICTIONARY_ENTRY_LIMIT
) {
const oldest = termsFindCache.entries().next().value;
if (oldest === undefined) { break; }
dropCachedTermsFind(oldest[0], oldest[1]);
}
}
// Classification of a dictionary entry (which dictionaries it came from,
// which media ids it mentions) depends only on the entry object, so it is
// memoized for as long as that object lives. Entries are shared with the
// termsFind cache above, which is what makes this worth keeping: the same
// objects come back for every repeated lookup, on every line.
const dictionaryEntryNamesCache = new WeakMap();
const subMinerMediaIdsCache = new WeakMap();
const EMPTY_MEDIA_ID_SET = new Set();
// Only blind ladder steps are capped (see the retry loop): those are the
// ones that would otherwise degrade into O(scanLength) lookups at a single
// position. Steps the backend guides by reporting a shorter consumed length
// stay uncapped, so a valid prefix term is still found on lines where
// normalization eats a long tail.
const MAX_BLIND_SHRINKING_WINDOW_RETRIES = 4;
// Character-name candidate forms for the current media, installed
// separately from the per-line scan call so the per-line script stays tiny.
// Stored raw here; the normalized lookup index is built inside the scan,
// where the kana-normalization helper is in scope, and reused by key.
let rawNameCandidates = null;
let nameCandidateIndex = null;
globalThis.__subminerYomitanScanSetNameCandidates = (key, forms) => {
if (!key || !Array.isArray(forms) || forms.length === 0) {
rawNameCandidates = null;
nameCandidateIndex = null;
return false;
}
rawNameCandidates = { key, forms };
nameCandidateIndex = null;
return true;
};
globalThis.__subminerYomitanScanVersion = ${YOMITAN_SCAN_RUNTIME_VERSION};
globalThis.__subminerYomitanScan = async (scanParams) => {
const {
text,
profileIndex,
scanLength,
includeNameMatchMetadata,
greedyNameScanEnabled,
currentCharacterDictionaryMediaId,
dictionaryPriorityByName,
dictionaryFrequencyModeByName,
cacheEpoch,
nameCandidateKey
} = scanParams;
if (cacheEpoch !== termsFindCacheEpoch) {
termsFindCache.clear();
termsFindCacheDictionaryEntries = 0;
termsFindCacheEpoch = cacheEpoch;
}
${YOMITAN_SCANNING_HELPERS}
const CAPTION_OPENING_BRACKETS = new Set(["(", "", "[", "", "{", "", "「", "『", "【", "〈", "《", "≪", "", "<"]);
function shouldEmitUnparsedRunAsToken(runText) {
if (!/[\p{L}\p{N}]/u.test(runText)) { return false; }
const firstChar = Array.from(runText.trim())[0];
return firstChar !== undefined && !CAPTION_OPENING_BRACKETS.has(firstChar);
}
function isLookupWorthyCodePoint(codePoint) {
if (isCodePointJapanese(codePoint)) { return true; }
return /[\p{L}\p{N}]/u.test(String.fromCodePoint(codePoint));
}
function isKanaOnlyRunText(runText) {
const chars = Array.from(runText);
return chars.length > 0 && chars.every((char) => isCodePointKana(char.codePointAt(0)));
}
const details = {matchType: "exact", deinflect: true};
const tokens = [];
async function termsFindAt(position, windowLength) {
const substring = text.substring(position, position + windowLength);
const cacheKey = profileIndex + "\u0000" + substring;
const cached = termsFindCache.get(cacheKey);
if (cached !== undefined) {
termsFindCache.delete(cacheKey);
termsFindCache.set(cacheKey, cached);
return await cached.promise;
}
// An in-flight lookup counts as one entry until it resolves; the real
// weight replaces that estimate once the result is known.
const entry = { promise: null, dictionaryEntryCount: 1 };
entry.promise = invoke("termsFind", { text: substring, details, optionsContext: { index: profileIndex } })
.then((result) => {
const resolvedCount =
1 + (Array.isArray(result?.dictionaryEntries) ? result.dictionaryEntries.length : 0);
const isCached = termsFindCache.get(cacheKey) === entry;
if (isCached) {
termsFindCacheDictionaryEntries += resolvedCount - entry.dictionaryEntryCount;
}
entry.dictionaryEntryCount = resolvedCount;
// The real weight can push the cache over its budget, and a single
// response can exceed it on its own, so re-check here.
if (isCached) { evictOverflowingTermsFindEntries(); }
return result;
});
termsFindCache.set(cacheKey, entry);
termsFindCacheDictionaryEntries += entry.dictionaryEntryCount;
evictOverflowingTermsFindEntries();
try {
return await entry.promise;
} catch (error) {
dropCachedTermsFind(cacheKey, entry);
throw error;
}
}
// Text the walk skips accumulates into unparsed runs, mirroring the
// filler chunks the parseText segmentation used to provide: runs stay
// hoverable (flagged isUnparsedRun) unless they are punctuation-only or
// caption-style asides, and kana continuations of a longer headword
// extend the previous token instead.
function flushUnparsedRun(runStart, runEnd) {
if (runStart === null || runEnd <= runStart) { return; }
const runText = text.substring(runStart, runEnd);
const previousToken = tokens[tokens.length - 1];
if (
previousToken &&
previousToken.endPos === runStart &&
isKanaOnlyRunText(runText) &&
typeof previousToken.headword === "string" &&
previousToken.headword.length > previousToken.surface.length &&
previousToken.headword.startsWith(previousToken.surface + runText)
) {
previousToken.surface += runText;
// The run is kana-only, so its reading is itself: append it or the
// reading stops covering the surface, which disables the known-word
// reading fallback (isCompleteReadingForSurface) downstream.
previousToken.reading += runText;
// The run is kana-only, so its reading is itself: append it or the
// reading stops covering the surface, which disables the known-word
// reading fallback (isCompleteReadingForSurface) downstream.
previousToken.endPos = runEnd;
return;
}
if (!shouldEmitUnparsedRunAsToken(runText)) { return; }
tokens.push({
surface: runText,
reading: "",
headword: runText,
startPos: runStart,
endPos: runEnd,
isUnparsedRun: true
});
}
function buildScanToken(position, source, preferredHeadword) {
const reading = typeof preferredHeadword.reading === "string" ? preferredHeadword.reading : "";
const segments = distributeFuriganaInflected(preferredHeadword.term, reading, source);
const tokenPayload = {
surface: segments.map((segment) => segment.text).join("") || source,
reading: segments.map(getSegmentReadingContribution).join(""),
headword: preferredHeadword.term,
headwordReading: reading || undefined,
startPos: position,
endPos: position + source.length,
isNameMatch: includeNameMatchMetadata && preferredHeadword.isNameMatch === true,
frequencyRank:
typeof preferredHeadword.frequencyRank === "number" && Number.isFinite(preferredHeadword.frequencyRank)
? Math.max(1, Math.floor(preferredHeadword.frequencyRank))
: undefined,
};
if (Array.isArray(preferredHeadword.wordClasses) && preferredHeadword.wordClasses.length > 0) {
tokenPayload.wordClasses = preferredHeadword.wordClasses;
}
return tokenPayload;
}
// findTokenAt plus the shrinking-window ladder below it: Yomitan text
// normalization can consume characters (whitespace, punctuation) beyond
// the matched term, leaving no headword whose source equals the consumed
// text. Retry with shorter windows so a valid prefix term (e.g. a
// character name before a paren) still tokenizes instead of the position
// being skipped.
// Every window at or above the consumed length repeats the same result,
// so the next informative window sits just below it. A lookup that
// consumed its whole window reports nothing to aim at, and the step down
// from it is a blind guess: only those are budgeted.
// The window can run past the end of the line, so blindness is judged
// against the text the lookup actually saw.
// Set when a position stopped short of windows an uncapped ladder would
// still have tried; the line then escalates to parseText at the end.
let blindRetryBudgetExhausted = false;
async function resolveTokenAt(position, windowLength) {
let attempt = await findTokenAt(position, windowLength);
const scannedLength = Math.min(windowLength, text.length - position);
let retryLength = Math.min(attempt.matchedLength, scannedLength) - 1;
let stepIsBlind = attempt.matchedLength >= scannedLength;
let blindRetriesRemaining = MAX_BLIND_SHRINKING_WINDOW_RETRIES;
while (!attempt.token && retryLength >= 1) {
if (stepIsBlind) {
if (blindRetriesRemaining <= 0) {
blindRetryBudgetExhausted = true;
break;
}
blindRetriesRemaining -= 1;
}
const retry = await findTokenAt(position, retryLength);
if (retry.token) { return retry; }
const guidedLength = retry.matchedLength - 1;
stepIsBlind = guidedLength >= retryLength - 1;
retryLength = Math.min(retryLength - 1, guidedLength);
}
return attempt;
}
async function findTokenAt(position, windowLength) {
const codePoint = text.codePointAt(position);
const character = String.fromCodePoint(codePoint);
const result = await termsFindAt(position, windowLength);
const dictionaryEntries = Array.isArray(result?.dictionaryEntries) ? result.dictionaryEntries : [];
const originalTextLength = typeof result?.originalTextLength === "number" ? result.originalTextLength : 0;
if (dictionaryEntries.length === 0 || originalTextLength <= 0 || (originalTextLength === character.length && !isCodePointJapanese(codePoint))) {
return { token: null, matchedLength: 0 };
}
const source = text.substring(position, position + originalTextLength);
const preferredHeadword = getPreferredHeadword(
dictionaryEntries,
source,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
);
if (!preferredHeadword || typeof preferredHeadword.term !== "string") {
return { token: null, matchedLength: originalTextLength };
}
return { token: buildScanToken(position, source, preferredHeadword), matchedLength: originalTextLength };
}
// Kana normalization folds halfwidth katakana one code point to one, so an
// unvoiced halfwidth spelling prefix-matches a candidate form like any
// other. What it cannot fold is a voiced pair: カ + ゙ stays two characters
// where the candidate form carries the single が, so the comparison fails
// at that character. That break can sit anywhere inside the name, not
// just at its first character (山ガク starts on a kanji), so the bypass is
// keyed on the region a candidate could cover, not on how it starts.
function isHalfwidthKanaVoicedMarkCodePoint(codePoint) {
return codePoint === 0xff9e || codePoint === 0xff9f;
}
// Build (once per candidate list) a first-character bucket index of the
// normalized name forms, so the pre-pass can reject a position with a
// single map hit instead of a backend round trip.
if (rawNameCandidates && nameCandidateIndex?.key !== rawNameCandidates.key) {
const byFirstChar = new Map();
for (const form of rawNameCandidates.forms) {
const normalized = typeof form === "string" ? convertKatakanaToHiragana(form.trim()) : "";
if (!normalized) { continue; }
const bucket = byFirstChar.get(normalized[0]);
if (bucket) { bucket.push(normalized); } else { byFirstChar.set(normalized[0], [normalized]); }
}
nameCandidateIndex = byFirstChar.size > 0 ? { key: rawNameCandidates.key, byFirstChar } : null;
} else if (!rawNameCandidates) {
nameCandidateIndex = null;
}
// Only meaningful when the installed list matches the media this scan is
// for; otherwise fall back to scanning every position.
const activeNameCandidateIndex =
nameCandidateKey !== null && nameCandidateIndex?.key === nameCandidateKey
? nameCandidateIndex
: null;
const normalizedText = activeNameCandidateIndex ? convertKatakanaToHiragana(text) : "";
// Yomitan collapses emphatic sequences before matching (すっっごーーい →
// すごい), so a stretched name still resolves to its entry. Skipping these
// characters keeps such spellings candidates; the filter only ever grows
// the probe set, so a false positive costs one lookup, never a name.
const EMPHATIC_SKIP_CHARS = new Set(["ぁ", "ぃ", "ぅ", "ぇ", "ぉ", "っ", "ゃ", "ゅ", "ょ", "ー"]);
function matchesCandidateFormAt(form, position) {
let textIndex = position;
for (let formIndex = 0; formIndex < form.length; formIndex += 1) {
while (
textIndex < normalizedText.length &&
normalizedText[textIndex] !== form[formIndex] &&
EMPHATIC_SKIP_CHARS.has(normalizedText[textIndex])
) {
textIndex += 1;
}
if (normalizedText[textIndex] !== form[formIndex]) { return false; }
textIndex += 1;
}
return true;
}
// Where the folding gives up, listed once per line. Matching may skip any
// number of emphatic characters on its way through a form (山ーーーーーーガク),
// so there is no shorter honest bound than the window a name lookup
// covers: scanLength. The list is almost always empty, which is what
// keeps the check below free on ordinary lines.
const halfwidthVoicedMarkPositions = [];
if (activeNameCandidateIndex) {
for (let index = 0; index < text.length; index += 1) {
if (isHalfwidthKanaVoicedMarkCodePoint(text.charCodeAt(index))) {
halfwidthVoicedMarkPositions.push(index);
}
}
}
function hasHalfwidthVoicedMarkInScanWindow(position) {
const end = position + scanLength;
for (const markPosition of halfwidthVoicedMarkPositions) {
if (markPosition >= position && markPosition < end) { return true; }
}
return false;
}
// A name written ガ... folds to か + ゙, so its first character never leads
// to the が bucket the candidate form is filed under. Nothing else can
// find it, so such a position is always worth a probe.
function startsHalfwidthVoicedPair(position, codePoint) {
if (codePoint < 0xff66 || codePoint > 0xff9d) { return false; }
return isHalfwidthKanaVoicedMarkCodePoint(text.charCodeAt(position + 1));
}
function couldNameStartAt(position, codePoint) {
// Nothing starts with a combining voiced mark, whether or not the
// prefilter is active.
if (isHalfwidthKanaVoicedMarkCodePoint(codePoint)) { return false; }
if (!activeNameCandidateIndex) { return true; }
const bucket = activeNameCandidateIndex.byFirstChar.get(normalizedText[position]);
if (!bucket) {
// No candidate begins with this character, and the window search
// below would only ever say yes to positions like this one, so an
// unrelated ガ elsewhere in the line must not drag them in.
return startsHalfwidthVoicedPair(position, codePoint);
}
for (const form of bucket) {
if (matchesCandidateFormAt(form, position)) { return true; }
}
// A candidate does start here but did not match: an unfoldable voiced
// pair anywhere in the window is a reason the comparison could not see
// it (山ガク, 山ーーーーーーガク), so probe rather than drop the name.
return hasHalfwidthVoicedMarkInScanWindow(position);
}
// Greedy name pre-pass: character-name matches claim their spans before
// the left-to-right walk, so a longer generic match starting earlier
// (e.g. とヨー → 渡洋) cannot swallow the start of a name (ヨータ).
const nameTokens = [];
if (greedyNameScanEnabled) {
let namePos = 0;
while (namePos < text.length) {
const codePoint = text.codePointAt(namePos);
if (!isCodePointJapanese(codePoint) || !couldNameStartAt(namePos, codePoint)) {
namePos += String.fromCodePoint(codePoint).length;
continue;
}
const result = await termsFindAt(namePos, scanLength);
const dictionaryEntries = Array.isArray(result?.dictionaryEntries) ? result.dictionaryEntries : [];
const textWindow = text.substring(namePos, namePos + scanLength);
const nameMatch = findLongestNameMatch(dictionaryEntries, textWindow);
// A name only claims its span when no strictly longer generic word
// starts at the same position (a character named 空 must not split
// 空気). Ties go to the name. Generic matches that start earlier and
// overlap the name are still blocked by the reservation.
if (
!nameMatch ||
findLongestGenericMatchLength(dictionaryEntries, textWindow) > nameMatch.sourceLength
) {
namePos += String.fromCodePoint(codePoint).length;
continue;
}
const source = text.substring(namePos, namePos + nameMatch.sourceLength);
nameTokens.push(buildScanToken(namePos, source, {
term: nameMatch.headword.term,
reading: nameMatch.headword.reading,
wordClasses: normalizeWordClasses(nameMatch.headword),
isNameMatch: true,
frequencyRank: getBestFrequencyRank(
nameMatch.dictionaryEntry,
nameMatch.headwordIndex,
dictionaryPriorityByName,
dictionaryFrequencyModeByName
)
}));
namePos += nameMatch.sourceLength;
}
}
// First reserved name span that a match ending at endPos would leave
// half-consumed. Spans the match covers entirely are not returned: those
// lose to the longer word instead of splitting it.
function findSplitNameToken(startIndex, endPos) {
for (let index = startIndex; index < nameTokens.length; index += 1) {
const nameToken = nameTokens[index];
if (nameToken.startPos >= endPos) { return null; }
if (nameToken.endPos > endPos) { return nameToken; }
}
return null;
}
let i = 0;
let nameIndex = 0;
let unparsedRunStart = null;
while (i < text.length) {
while (nameIndex < nameTokens.length && nameTokens[nameIndex].startPos < i) { nameIndex += 1; }
const nextNameToken = nameIndex < nameTokens.length ? nameTokens[nameIndex] : null;
if (nextNameToken && nextNameToken.startPos === i) {
flushUnparsedRun(unparsedRunStart, i);
unparsedRunStart = null;
tokens.push(nextNameToken);
i = nextNameToken.endPos;
nameIndex += 1;
continue;
}
const codePoint = text.codePointAt(i);
// Punctuation and whitespace can never start a token: skip the backend
// round trip entirely. Latin letters and digits stay lookup-worthy
// (terms like Tシャツ start on an ASCII letter).
if (!isLookupWorthyCodePoint(codePoint)) {
if (unparsedRunStart === null) { unparsedRunStart = i; }
i += String.fromCodePoint(codePoint).length;
continue;
}
// A reservation only outranks generic matches that would cut into it.
// Look the position up unrestricted first: a generic word that starts
// earlier and covers the whole name span (写真 over a character named
// 真) is the better reading, so the reservation yields rather than
// splitting the word. Only a match that ends inside a name span gets
// re-run against a window capped at that span.
let attempt = await resolveTokenAt(i, scanLength);
if (attempt.token) {
const splitNameToken = findSplitNameToken(nameIndex, attempt.token.endPos);
if (splitNameToken) {
attempt = await resolveTokenAt(i, splitNameToken.startPos - i);
}
}
if (attempt.token) {
flushUnparsedRun(unparsedRunStart, i);
unparsedRunStart = null;
tokens.push(attempt.token);
i += attempt.matchedLength;
continue;
}
if (unparsedRunStart === null) { unparsedRunStart = i; }
i += String.fromCodePoint(text.codePointAt(i)).length;
}
flushUnparsedRun(unparsedRunStart, text.length);
if (blindRetryBudgetExhausted) {
// A position gave up with shorter windows still worth trying. The walk
// is the only tokenizer now, so stopping there would leave a real term
// as an unparsed run; report it so the host can spend one parseText on
// the line instead of letting the ladder run to O(scanLength) lookups.
return { tokens, retryBudgetExhausted: true };
}
return tokens;
};
return true;
})();
`;
// Installs (or clears) the character-name candidate forms for the current
// media. Runs only when the list changes, not per line. Passing null restores
// the exhaustive every-position pre-pass.
export function buildYomitanScanNameCandidatesScript(
nameCandidates: { key: string; forms: string[] } | null,
): string {
if (!nameCandidates) {
return `
(() => {
if (typeof globalThis.__subminerYomitanScanSetNameCandidates !== "function") {
return false;
}
return globalThis.__subminerYomitanScanSetNameCandidates(null, null);
})();
`;
}
return `
(() => {
if (typeof globalThis.__subminerYomitanScanSetNameCandidates !== "function") {
return false;
}
return globalThis.__subminerYomitanScanSetNameCandidates(
${JSON.stringify(nameCandidates.key)},
${JSON.stringify(nameCandidates.forms)}
);
})();
`;
}
export function buildYomitanScanCallScript(params: YomitanScanRequestParams): string {
return `
(async () => {
if (typeof globalThis.__subminerYomitanScan !== "function") {
return ${JSON.stringify(YOMITAN_SCAN_RUNTIME_MISSING_SENTINEL)};
}
return await globalThis.__subminerYomitanScan(${JSON.stringify(params)});
})();
`;
}
@@ -0,0 +1,304 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { requestYomitanScanTokens } from './yomitan-parser-runtime';
import {
countTermsFindLookups,
createNameScanDeps,
NAME_SCAN_WORDS,
} from './yomitan-scan-test-harness';
// Behaviour of the in-page scan runtime around character names and kana:
// which positions the greedy pre-pass probes, and what the walk makes of
// halfwidth spellings. Driven end to end through requestYomitanScanTokens
// because the runtime only exists inside the parser window.
const NAME_SCAN_LINE = 'ミナトはまだ学校にいない';
test('requestYomitanScanTokens skips name pre-pass lookups where no candidate name can start', async () => {
const exhaustiveLookups: string[] = [];
const exhaustive = await requestYomitanScanTokens(
NAME_SCAN_LINE,
createNameScanDeps(exhaustiveLookups),
{ error: () => undefined },
{ includeNameMatchMetadata: true },
);
const prefilteredLookups: string[] = [];
const prefiltered = await requestYomitanScanTokens(
NAME_SCAN_LINE,
createNameScanDeps(prefilteredLookups),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
// Terms and readings the generated dictionary exposes for this media.
nameCandidates: { key: 'media-1', forms: ['ミナト', 'みなと'] },
},
);
// Same tokenization, including the name match, with fewer round trips.
assert.deepEqual(prefiltered, exhaustive);
assert.equal(prefiltered?.[0]?.surface, 'ミナト');
assert.equal(prefiltered?.[0]?.isNameMatch, true);
assert.ok(
prefilteredLookups.length < exhaustiveLookups.length,
`expected fewer lookups with candidates (${prefilteredLookups.length} vs ${exhaustiveLookups.length})`,
);
// Mid-token positions are exactly what the pre-pass used to probe (a name can
// start mid-token); with candidates they cost nothing, while the main walk's
// own token-start lookups are unaffected.
assert.ok(countTermsFindLookups(exhaustiveLookups, '校に') > 0);
assert.equal(countTermsFindLookups(prefilteredLookups, '校に'), 0);
});
test('requestYomitanScanTokens matches a katakana name from its kana-normalized candidate form', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
NAME_SCAN_LINE,
createNameScanDeps(lookups),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
// Only the hiragana reading is listed; the katakana surface in the line
// must still be found through kana normalization.
nameCandidates: { key: 'media-1', forms: ['みなと'] },
},
);
assert.equal(result?.[0]?.surface, 'ミナト');
assert.equal(result?.[0]?.isNameMatch, true);
});
// Kana normalization folds halfwidth katakana, so a name written that way does
// prefix-match a candidate form — but only if the position counts as Japanese
// in the first place. The generic word here reaches into the name, so only a
// pre-pass reservation can keep the name whole.
const HALFWIDTH_NAME_SCAN_WORDS: Array<[string, string, string, boolean]> = [
['ネコ', 'ネコ', 'ねこ', false],
['まだミ', 'まだミ', 'まだみ', false],
['まだ', 'まだ', 'まだ', false],
['ミナト', 'ミナト', 'みなと', true],
];
test('requestYomitanScanTokens probes halfwidth katakana positions during the name pre-pass', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
'ネコまだミナト',
createNameScanDeps(lookups, HALFWIDTH_NAME_SCAN_WORDS),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
// Fullwidth forms only, as the generated dictionary stores them.
nameCandidates: { key: 'media-1', forms: ['ミナト', 'みなと'] },
},
);
assert.equal(countTermsFindLookups(lookups, 'ミナト'), 1);
// コ is mid-token, so only the pre-pass would ever look it up, and it matches
// no candidate: folding halfwidth made those positions indexable, so they no
// longer cost a round trip apiece.
assert.equal(countTermsFindLookups(lookups, 'コ'), 0);
assert.deepEqual(
result?.map((token) => token.surface),
['ネコ', 'まだ', 'ミナト'],
);
assert.equal(result?.[2]?.isNameMatch, true);
// The reading is written the way the fullwidth katakana path writes it
// (surface spelling, fullwidth): halfwidth kana is not kana to the known-word
// and frequency code downstream, and an empty reading there disables the
// reading fallback entirely.
assert.equal(result?.[2]?.reading, 'ミナト');
assert.equal(result?.[2]?.headwordReading, 'みなと');
});
test('a voiced halfwidth name still bypasses the candidate prefilter', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
'まだガク',
createNameScanDeps(lookups, [
['まだカ', 'まだカ', 'まだか', false],
['まだ', 'まだ', 'まだ', false],
['ガク', 'ガク', 'がく', true],
]),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['ガク', 'がく'] },
},
);
// カ + ゙ folds to か + ゙, which cannot prefix-match が, so the prefilter would
// drop this position; the voiced-mark bypass is what keeps the name.
assert.deepEqual(
result?.map((token) => token.surface),
['まだ', 'ガク'],
);
assert.equal(result?.[1]?.isNameMatch, true);
});
test('an unrelated halfwidth voiced word does not restore the exhaustive pre-pass', async () => {
const baseline: string[] = [];
await requestYomitanScanTokens(
NAME_SCAN_LINE,
createNameScanDeps(baseline),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['ミナト', 'みなと'] },
},
);
const withVoicedTail: string[] = [];
await requestYomitanScanTokens(
`${NAME_SCAN_LINE}ガ`,
createNameScanDeps(withVoicedTail),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['ミナト', 'みなと'] },
},
);
// Mid-token positions are the ones only the pre-pass would ever probe. A ガ
// anywhere in the line used to drag every position within scanLength of it
// back in; now only the voiced pair itself, which the fold cannot index, is
// added to what the line already looked up.
for (const midTokenPrefix of ['ナト', 'だ学', '校に', 'ない']) {
assert.equal(countTermsFindLookups(baseline, midTokenPrefix), 0, midTokenPrefix);
assert.equal(countTermsFindLookups(withVoicedTail, midTokenPrefix), 0, midTokenPrefix);
}
assert.ok(
withVoicedTail.length - baseline.length <= 3,
`expected the ガ tail to add only its own lookups, saw ${JSON.stringify(withVoicedTail)}`,
);
});
test('a mixed-width voiced name survives the candidate prefilter', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
'まだ山ガク',
createNameScanDeps(lookups, [
['まだ山', 'まだ山', 'まだやま', false],
['まだ', 'まだ', 'まだ', false],
['山ガク', '山ガク', 'やまがく', true],
]),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['山ガク', 'やまがく'] },
},
);
// The name starts on a kanji, so the fold only breaks mid-name: 山ガク
// normalizes to 山がく, which still cannot match the candidate 山がく. The
// bypass is keyed on the scan window rather than the first character, so the
// position is still probed and the generic まだ山 cannot swallow the 山.
assert.deepEqual(
result?.map((token) => token.surface),
['まだ', '山ガク'],
);
assert.equal(result?.[1]?.isNameMatch, true);
});
test('a stretched mixed-width voiced name survives the candidate prefilter', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
'まだ山ーーーーーーガク',
createNameScanDeps(lookups, [
['まだ山', 'まだ山', 'まだやま', false],
['まだ', 'まだ', 'まだ', false],
['山ーーーーーーガク', '山ガク', 'やまがく', true],
]),
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['山ガク', 'やまがく'] },
},
);
// Matching skips any number of emphatic characters, so the voiced mark that
// defeats the fold can sit arbitrarily far into the name: the search for it
// has to cover the whole lookup window, not a multiple of the form length.
assert.deepEqual(
result?.map((token) => token.surface),
['まだ', '山ーーーーーーガク'],
);
assert.equal(result?.[1]?.isNameMatch, true);
});
test('halfwidth voiced kana compose into the reading instead of leaving a stray mark', async () => {
const lookups: string[] = [];
const result = await requestYomitanScanTokens(
'ガク パン',
createNameScanDeps(lookups, [
['ガク', 'ガク', 'がく', false],
['パン', 'パン', 'ぱん', false],
]),
{ error: () => undefined },
{ includeNameMatchMetadata: true },
);
// The name pre-pass runs over every position here (no candidate list), but a
// standalone voiced mark can never start a name, so it costs no lookup.
assert.equal(countTermsFindLookups(lookups, '゙'), 0);
assert.equal(countTermsFindLookups(lookups, '゚'), 0);
const readings = (result ?? [])
.filter((token) => token.isUnparsedRun !== true)
.map((token) => [token.surface, token.reading]);
assert.deepEqual(readings, [
['ガク', 'ガク'],
['パン', 'パン'],
]);
});
test('requestYomitanScanTokens falls back to the exhaustive name scan without candidates', async () => {
const withoutLookups: string[] = [];
const withoutCandidates = await requestYomitanScanTokens(
NAME_SCAN_LINE,
createNameScanDeps(withoutLookups),
{ error: () => undefined },
{ includeNameMatchMetadata: true, currentCharacterDictionaryMediaId: 1, nameCandidates: null },
);
assert.equal(withoutCandidates?.[0]?.isNameMatch, true);
// No candidate list means every Japanese position is probed, as before.
assert.ok(countTermsFindLookups(withoutLookups, '校に') > 0);
});
test('requestYomitanScanTokens reinstalls name candidates when the media changes', async () => {
const lookups: string[] = [];
const deps = createNameScanDeps(lookups);
// First media's candidates cannot match this line's name.
const otherMedia = await requestYomitanScanTokens(
NAME_SCAN_LINE,
deps,
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 2,
nameCandidates: { key: 'media-2', forms: ['カズマ'] },
},
);
assert.equal(otherMedia?.[0]?.isNameMatch, undefined);
const correctMedia = await requestYomitanScanTokens(
NAME_SCAN_LINE,
deps,
{ error: () => undefined },
{
includeNameMatchMetadata: true,
currentCharacterDictionaryMediaId: 1,
nameCandidates: { key: 'media-1', forms: ['ミナト'] },
},
);
assert.equal(correctMedia?.[0]?.surface, 'ミナト');
assert.equal(correctMedia?.[0]?.isNameMatch, true);
});
@@ -0,0 +1,166 @@
// Shared harness for the Yomitan parser-runtime and scan-runtime tests: fake
// parser-window deps whose injected scripts run in a vm context, plus the
// backend stubs the scanner tests drive them with. Kept out of the test files
// so the runtime tests and the in-page scanner tests can share one setup.
import * as vm from 'node:vm';
export function createDeps(
executeJavaScript: (script: string) => Promise<unknown>,
options?: {
createYomitanExtensionWindow?: (pageName: string) => Promise<unknown>;
},
) {
const parserWindow = {
isDestroyed: () => false,
webContents: {
executeJavaScript: async (script: string) => await executeJavaScript(script),
},
};
return {
getYomitanExt: () => ({ id: 'ext-id' }) as never,
getYomitanParserWindow: () => parserWindow as never,
setYomitanParserWindow: () => undefined,
getYomitanParserReadyPromise: () => null,
setYomitanParserReadyPromise: () => undefined,
getYomitanParserInitPromise: () => null,
setYomitanParserInitPromise: () => undefined,
createYomitanExtensionWindow: options?.createYomitanExtensionWindow as never,
};
}
function createYomitanScriptSandbox(handler: (action: string, params: unknown) => unknown) {
return {
chrome: {
runtime: {
lastError: null,
sendMessage: (
payload: { action?: string; params?: unknown },
callback: (response: { result?: unknown; error?: { message?: string } }) => void,
) => {
try {
callback({ result: handler(payload.action ?? '', payload.params) });
} catch (error) {
callback({ error: { message: (error as Error).message } });
}
},
},
},
Array,
Error,
JSON,
Map,
Math,
Number,
Object,
Promise,
RegExp,
Set,
String,
};
}
export async function runInjectedYomitanScript(
script: string,
handler: (action: string, params: unknown) => unknown,
): Promise<unknown> {
return await vm.runInNewContext(script, createYomitanScriptSandbox(handler));
}
// Persistent page context shared across executeJavaScript calls, matching the
// real parser window: the scan runtime is installed once via
// globalThis.__subminerYomitanScan and per-line calls reuse it (and its
// cross-line termsFind cache).
function createPersistentYomitanScriptRunner(
handler: (action: string, params: unknown) => unknown,
): (script: string) => Promise<unknown> {
const context = vm.createContext(createYomitanScriptSandbox(handler));
return async (script: string) => await vm.runInContext(script, context);
}
// Deps whose parser window executes every injected script (profile metadata,
// scan runtime install, per-line scan calls, parseText fallback) inside one
// persistent vm context, dispatching backend actions to `handler`.
export function createScanDeps(
handler: (action: string, params: unknown) => unknown,
options?: { onScript?: (script: string) => void },
) {
const runScript = createPersistentYomitanScriptRunner(handler);
return createDeps(async (script) => {
options?.onScript?.(script);
return await runScript(script);
});
}
export function countTermsFindLookups(lookups: string[], prefix: string): number {
return lookups.filter((lookupText) => lookupText.startsWith(prefix)).length;
}
// Backend stub for the greedy name pre-pass: one character name (ミナト) in a
// line of ordinary words, with the SubMiner character dictionary enabled.
export const NAME_SCAN_WORDS: Array<[string, string, string, boolean]> = [
['ミナト', 'ミナト', 'みなと', true],
['は', 'は', 'は', false],
['まだ', 'まだ', 'まだ', false],
['学校', '学校', 'がっこう', false],
['に', 'に', 'に', false],
['いない', 'いる', 'いる', false],
];
export function createNameScanDeps(
lookups: string[],
words: Array<[string, string, string, boolean]> = NAME_SCAN_WORDS,
) {
return createScanDeps((action, params) => {
if (action === 'optionsGetFull') {
return {
profileCurrent: 0,
profiles: [
{
options: {
scanning: { length: 40 },
dictionaries: [
{ name: 'JMdict', enabled: true, id: 0 },
{
name: 'SubMiner Character Dictionary (AniList 1)',
enabled: true,
id: 1,
},
],
},
},
],
};
}
if (action === 'getDictionaryInfo') {
return [];
}
if (action !== 'termsFind') {
throw new Error(`unexpected action: ${action}`);
}
const text = (params as { text?: string } | undefined)?.text ?? '';
lookups.push(text);
for (const [surface, term, reading, isName] of words) {
if (text.startsWith(surface)) {
return {
originalTextLength: surface.length,
dictionaryEntries: [
{
headwords: [
{
term,
reading,
sources: [{ originalText: surface, isPrimary: true, matchType: 'exact' }],
},
],
definitions: [
{ dictionary: isName ? 'SubMiner Character Dictionary (AniList 1)' : 'JMdict' },
],
},
],
};
}
}
return { originalTextLength: 0, dictionaryEntries: [] };
});
}
@@ -0,0 +1,21 @@
// Helper bundle for the in-page Yomitan scan runtime, composed from the
// fragments below. Injected as text into the parser window by
// yomitan-scan-runtime-script.ts, so it is data here, not code this process
// runs. The fragments are concatenated into a single function body and share
// one lexical scope: every function in them is hoisted, but the constants are
// not, so kana stays first — the later fragments read its ranges as they run.
import { YOMITAN_DICTIONARY_CLASSIFICATION_HELPERS } from './yomitan-dictionary-classification-script';
import { YOMITAN_FREQUENCY_HELPERS } from './yomitan-frequency-script';
import { YOMITAN_FURIGANA_HELPERS } from './yomitan-furigana-script';
import { YOMITAN_KANA_HELPERS } from './yomitan-kana-script';
import { YOMITAN_MATCH_SELECTION_HELPERS } from './yomitan-match-selection-script';
export { CHARACTER_DICTIONARY_TITLE_PREFIX } from './character-dictionary-title';
export const YOMITAN_SCANNING_HELPERS = [
YOMITAN_KANA_HELPERS,
YOMITAN_FURIGANA_HELPERS,
YOMITAN_FREQUENCY_HELPERS,
YOMITAN_DICTIONARY_CLASSIFICATION_HELPERS,
YOMITAN_MATCH_SELECTION_HELPERS,
].join('\n');
+54
View File
@@ -0,0 +1,54 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { HAN_CODE_POINT_RANGES, HAN_REGEXP_CLASS_BODY, isHanCodePoint } from './han-code-points';
test('every range boundary is inside the table', () => {
for (const [start, end] of HAN_CODE_POINT_RANGES) {
for (const codePoint of [start, end]) {
assert.ok(isHanCodePoint(codePoint), `expected U+${codePoint.toString(16)} to be Han`);
}
}
// Extension J (Unicode 17) and the Compatibility blocks are the ones a
// BMP-only table used to miss.
assert.ok(isHanCodePoint(0x323b0));
assert.ok(isHanCodePoint(0x33479));
assert.ok(isHanCodePoint(0xf900));
assert.ok(isHanCodePoint(0x2f800));
});
test('no unified ideograph the runtime knows about falls outside the table', () => {
// One direction only: a runtime with older Unicode data simply checks fewer
// code points, where asserting the reverse would fail on Extension J.
const unifiedIdeograph = /\p{Unified_Ideograph}/u;
for (let codePoint = 0x3000; codePoint <= 0x40000; codePoint += 1) {
if (unifiedIdeograph.test(String.fromCodePoint(codePoint))) {
assert.ok(
isHanCodePoint(codePoint),
`expected unified ideograph U+${codePoint.toString(16)} to be in the table`,
);
}
}
});
test('code points just outside the table are rejected', () => {
for (const codePoint of [0x33ff, 0x4dc0, 0xa000, 0x1f000, 0x3347a]) {
assert.equal(
isHanCodePoint(codePoint),
false,
`expected U+${codePoint.toString(16)} not to be Han`,
);
}
});
test('the regexp class body matches the same code points as the predicate', () => {
const classRegExp = new RegExp(`^[${HAN_REGEXP_CLASS_BODY}]$`, 'u');
for (const codePoint of [0x3400, 0x4e00, 0x9fff, 0xf900, 0x20000, 0x323b0, 0x33479]) {
assert.match(String.fromCodePoint(codePoint), classRegExp);
}
for (const codePoint of [0x3040, 0x30ff, 0x33fa, 0x3347a]) {
assert.doesNotMatch(String.fromCodePoint(codePoint), classRegExp);
}
});
+31
View File
@@ -0,0 +1,31 @@
// Single source of truth for "this code point is a Han character", shared by
// the main-process character dictionary and the in-page Yomitan scan runtime.
// The two used to carry separate range lists, and they drifted: a name written
// with a supplementary-plane kanji could enter the generated dictionary while
// the scanner's greedy name pre-pass refused to probe the position.
//
// Ranges rather than \p{Script=Han}: the scan walk tests one code point per
// character of every subtitle line, where an integer compare beats building a
// string for a regex, and the script is injected as text into a page where a
// shared helper cannot be imported.
export const HAN_CODE_POINT_RANGES: ReadonlyArray<readonly [number, number]> = [
[0x3400, 0x4dbf], // Extension A
[0x4e00, 0x9fff], // CJK Unified Ideographs
[0xf900, 0xfaff], // Compatibility Ideographs
[0x20000, 0x2a6df], // Extension B
[0x2a700, 0x2ebef], // Extensions C-F
[0x2ebf0, 0x2ee5f], // Extension I
[0x2f800, 0x2fa1f], // Compatibility Ideographs Supplement
[0x30000, 0x3134f], // Extension G
[0x31350, 0x323af], // Extension H
[0x323b0, 0x33479], // Extension J (Unicode 17)
];
export function isHanCodePoint(codePoint: number): boolean {
return HAN_CODE_POINT_RANGES.some(([start, end]) => codePoint >= start && codePoint <= end);
}
/** The same ranges as a regular expression character class body (needs the `u` flag). */
export const HAN_REGEXP_CLASS_BODY = HAN_CODE_POINT_RANGES.map(
([start, end]) => `\\u{${start.toString(16)}}-\\u{${end.toString(16)}}`,
).join('');
+203
View File
@@ -0,0 +1,203 @@
import assert from 'node:assert/strict';
import fs from 'node:fs';
import path from 'node:path';
import test from 'node:test';
import { parseChangelog, resolveChangelogGroupKey } from './changelog-parse';
const SAMPLE = `# Changelog
## v0.19.2 (2026-08-04)
### Changed
- Subsync: picks both tracks now.
### Fixed
- Overlay: shows the plain line immediately.
<details>
<summary>Internal changes</summary>
### Internal
- Patched \`undici\`.
</details>
## v0.19.1 (2026-08-01)
### Added
- **Word Card Type:**
- Adds a setting.
- Flags clear each other.
## v0.18.0 (2026-07-01)
### Fixed
- Something older.
`;
test('changelog parser reads versions, dates, and sections in file order', () => {
const entries = parseChangelog(SAMPLE);
assert.deepEqual(
entries.map((entry) => `${entry.version}@${entry.date}`),
['0.19.2@2026-08-04', '0.19.1@2026-08-01', '0.18.0@2026-07-01'],
);
assert.deepEqual(
entries[0]?.sections.map((section) => section.heading),
['Changed', 'Fixed', 'Internal'],
);
assert.deepEqual(entries[0]?.sections[1]?.items, [
{ text: 'Overlay: shows the plain line immediately.', children: [] },
]);
});
test('changelog parser flags sections inside the details block as internal', () => {
const entries = parseChangelog(SAMPLE);
const sections = entries[0]?.sections ?? [];
assert.deepEqual(
sections.map((section) => section.internal),
[false, false, true],
);
assert.deepEqual(sections[2]?.items, [{ text: 'Patched `undici`.', children: [] }]);
});
test('changelog parser groups entries by major.minor', () => {
const entries = parseChangelog(SAMPLE);
assert.deepEqual(
entries.map((entry) => entry.groupKey),
['0.19', '0.19', '0.18'],
);
assert.equal(resolveChangelogGroupKey('1.2.3'), '1.2');
});
test('changelog parser keeps bullets that precede any section heading', () => {
const entries = parseChangelog('## v0.1.0 (2025-01-01)\n\n- Initial release.\n');
assert.deepEqual(entries[0]?.sections, [
{
heading: 'Changes',
items: [{ text: 'Initial release.', children: [] }],
internal: false,
},
]);
});
test('changelog parser drops empty sections and tolerates missing dates', () => {
const entries = parseChangelog('## v0.2.0\n\n### Added\n\n### Fixed\n- One fix.\n');
assert.equal(entries[0]?.date, '');
assert.deepEqual(
entries[0]?.sections.map((section) => section.heading),
['Fixed'],
);
});
test('changelog parser keeps indented sub-bullets nested under their lead bullet', () => {
const entries = parseChangelog(SAMPLE);
const added = entries[1]?.sections.find((section) => section.heading === 'Added');
assert.deepEqual(added?.items, [
{
text: '**Word Card Type:**',
children: [
{ text: 'Adds a setting.', children: [] },
{ text: 'Flags clear each other.', children: [] },
],
},
]);
});
test('changelog parser nests three bullet levels and rejoins wrapped lines', () => {
const entries = parseChangelog(
[
'## v0.9.0 (2025-05-05)',
'',
'### Added',
'- Top level',
' - Second level',
' - Third level',
' continued on the next line',
' - Back to second level',
'- Another top level',
'',
].join('\n'),
);
assert.deepEqual(entries[0]?.sections[0]?.items, [
{
text: 'Top level',
children: [
{
text: 'Second level',
children: [{ text: 'Third level continued on the next line', children: [] }],
},
{ text: 'Back to second level', children: [] },
],
},
{ text: 'Another top level', children: [] },
]);
});
test('changelog parser reads prerelease and build metadata version headings', () => {
const entries = parseChangelog(
[
'## v0.16.0 (2026-06-01)',
'',
'### Added',
'- New in 0.16.',
'',
'## v0.15.0-rc.1+build.2 (2026-05-29)',
'',
'### Added',
'- Release candidate note.',
'',
].join('\n'),
);
// The prerelease heading has to become its own entry. Asserting the exact
// version list is what catches the failure mode: a heading the regex misses
// is not skipped, its notes silently fold into the release above it.
assert.deepEqual(
entries.map((entry) => entry.version),
['0.16.0', '0.15.0-rc.1+build.2'],
);
assert.equal(entries[1]?.date, '2026-05-29');
assert.equal(entries[1]?.groupKey, '0.15');
assert.equal(entries[0]?.sections.length, 1);
// The prerelease body has to land on its own entry, not fold into 0.16.0.
assert.deepEqual(entries[1]?.sections, [
{
heading: 'Added',
items: [{ text: 'Release candidate note.', children: [] }],
internal: false,
},
]);
});
test('changelog parser handles the repo CHANGELOG.md', () => {
const markdown = fs.readFileSync(path.join(process.cwd(), 'CHANGELOG.md'), 'utf8');
const entries = parseChangelog(markdown);
assert.ok(entries.length > 3);
for (const entry of entries) {
assert.match(entry.version, /^\d+\.\d+\.\d+/);
assert.ok(entry.sections.length > 0, `expected sections for v${entry.version}`);
for (const section of entry.sections) {
for (const item of section.items) {
assert.ok(item.text.length > 0, `empty bullet in v${entry.version}`);
}
}
}
// Older entries group notes under a bold lead bullet; nesting must survive.
const breaking = entries
.find((entry) => entry.version === '0.15.0')
?.sections.find((section) => section.heading === 'Breaking Changes');
assert.deepEqual(
breaking?.items.map((item) => `${item.text}:${item.children.length}`),
['**Subsync:**:2', '**N+1 Highlighting:**:2'],
);
});
+128
View File
@@ -0,0 +1,128 @@
import type { ChangelogEntry, ChangelogItem, ChangelogSection } from '../../types/changelog';
// Prerelease and build metadata are matched separately: a single `[-+]`-led
// group cannot span `-rc.1+build.2`, and an unmatched heading silently folds
// that release's notes into the previous entry.
const VERSION_HEADING =
/^##\s+v(\d+\.\d+\.\d+(?:-[0-9A-Za-z.-]+)?(?:\+[0-9A-Za-z.-]+)?)\s*(?:\(([^)]*)\))?\s*$/;
const SECTION_HEADING = /^###\s+(.+?)\s*$/;
const BULLET = /^(\s*)[-*]\s+(.*)$/;
/**
* Entries are grouped by `major.minor` so the whole current minor line renders
* expanded, matching how docs-site/changelog.md splits current vs previous.
*/
export function resolveChangelogGroupKey(version: string): string {
const match = version.match(/^(\d+)\.(\d+)/);
if (!match) return version;
return `${match[1]}.${match[2]}`;
}
/**
* Parses the repo CHANGELOG.md into version entries. Bullets keep their inline
* markdown and their nesting: older entries group related notes under a bold
* lead bullet with indented children, and flattening them loses that structure.
*/
export function parseChangelog(markdown: string): ChangelogEntry[] {
const entries: ChangelogEntry[] = [];
let entry: ChangelogEntry | null = null;
let section: ChangelogSection | null = null;
let internal = false;
// Open bullets from outermost to innermost, used to place the next bullet.
let openItems: Array<{ indent: number; item: ChangelogItem }> = [];
function startSection(heading: string): void {
section = { heading, items: [], internal };
openItems = [];
entry?.sections.push(section);
}
function addBullet(indent: number, text: string): void {
if (!section) {
// Bullets before any "###" heading (older entries) land in a generic group.
startSection('Changes');
}
const item: ChangelogItem = { text, children: [] };
while (openItems.length > 0 && (openItems[openItems.length - 1]?.indent ?? 0) >= indent) {
openItems.pop();
}
const parent = openItems[openItems.length - 1];
if (parent) {
parent.item.children.push(item);
} else {
section?.items.push(item);
}
openItems.push({ indent, item });
}
function appendContinuation(text: string): void {
const current = openItems[openItems.length - 1];
if (!current) return;
current.item.text = `${current.item.text} ${text}`;
}
for (const rawLine of markdown.split(/\r?\n/)) {
const line = rawLine.trimEnd();
const trimmed = line.trim();
const versionMatch = trimmed.match(VERSION_HEADING);
if (versionMatch) {
const version = versionMatch[1] ?? '';
entry = {
version,
date: versionMatch[2]?.trim() ?? '',
groupKey: resolveChangelogGroupKey(version),
sections: [],
};
entries.push(entry);
section = null;
internal = false;
openItems = [];
continue;
}
if (!entry) continue;
if (trimmed.startsWith('<details')) {
internal = true;
section = null;
openItems = [];
continue;
}
if (trimmed.startsWith('</details')) {
internal = false;
section = null;
openItems = [];
continue;
}
if (trimmed.startsWith('<summary')) continue;
const sectionMatch = trimmed.match(SECTION_HEADING);
if (sectionMatch) {
startSection(sectionMatch[1] ?? '');
continue;
}
const bulletMatch = line.match(BULLET);
if (bulletMatch) {
addBullet((bulletMatch[1] ?? '').length, bulletMatch[2] ?? '');
continue;
}
// An indented non-bullet line continues the bullet above it, including
// across a blank line: that is CommonMark's continuation paragraph, and
// dropping the open bullets here would silently discard the text.
if (!trimmed) {
continue;
}
if (/^\s/.test(line)) {
appendContinuation(trimmed);
}
}
return entries.map((item) => ({
...item,
sections: item.sections.filter((entrySection) => entrySection.items.length > 0),
}));
}
+68 -1
View File
@@ -1,6 +1,6 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { shouldForceX11ElectronBackend } from './electron-backend';
import { resolveX11ElectronRelaunchArgs, shouldForceX11ElectronBackend } from './electron-backend';
function withPlatform(platform: NodeJS.Platform, run: () => void): void {
const original = Object.getOwnPropertyDescriptor(process, 'platform');
@@ -32,3 +32,70 @@ test('shouldForceX11ElectronBackend is false off Linux', () => {
assert.equal(shouldForceX11ElectronBackend({}), false);
});
});
test('resolveX11ElectronRelaunchArgs adds the raw X11 Ozone argument on unsupported Linux', () => {
assert.deepEqual(
resolveX11ElectronRelaunchArgs(
['--start'],
{
DISPLAY: ':1',
WAYLAND_DISPLAY: 'wayland-0',
XDG_CURRENT_DESKTOP: 'KDE',
},
'linux',
),
['--start', '--ozone-platform=x11'],
);
});
test('resolveX11ElectronRelaunchArgs avoids loops and preserves native Wayland backends', () => {
const kdeWayland = {
DISPLAY: ':1',
WAYLAND_DISPLAY: 'wayland-0',
XDG_CURRENT_DESKTOP: 'KDE',
};
assert.equal(
resolveX11ElectronRelaunchArgs(['--start', '--ozone-platform=x11'], kdeWayland, 'linux'),
null,
);
assert.equal(
resolveX11ElectronRelaunchArgs(
['--start'],
{ ...kdeWayland, HYPRLAND_INSTANCE_SIGNATURE: 'hypr' },
'linux',
),
null,
);
assert.equal(resolveX11ElectronRelaunchArgs(['--start'], kdeWayland, 'darwin'), null);
assert.equal(
resolveX11ElectronRelaunchArgs(
[],
{
...kdeWayland,
SUBMINER_APP_ARGC: '1',
SUBMINER_APP_ARG_0: '--start',
},
'linux',
)?.at(-1),
'--ozone-platform=x11',
);
assert.equal(
resolveX11ElectronRelaunchArgs([], { ...kdeWayland, SUBMINER_X11_BOOTSTRAPPED: '1' }, 'linux'),
null,
);
});
test('resolveX11ElectronRelaunchArgs replaces an explicit unsupported Wayland argument', () => {
assert.deepEqual(
resolveX11ElectronRelaunchArgs(
['--start', '--ozone-platform', 'wayland'],
{
DISPLAY: ':1',
WAYLAND_DISPLAY: 'wayland-0',
XDG_CURRENT_DESKTOP: 'KDE',
},
'linux',
),
['--start', '--ozone-platform=x11'],
);
});
+36 -2
View File
@@ -4,6 +4,9 @@ import { isSupportedWaylandCompositor } from '../../shared/mpv-x11-backend';
const logger = createLogger('core:electron-backend');
export const X11_ELECTRON_BOOTSTRAP_ENV = 'SUBMINER_X11_BOOTSTRAPPED';
const X11_ELECTRON_OZONE_ARG = '--ozone-platform=x11';
function getElectronOzonePlatformHint(env: NodeJS.ProcessEnv = process.env): string | null {
const hint = env.ELECTRON_OZONE_PLATFORM_HINT?.trim().toLowerCase();
if (hint) return hint;
@@ -24,11 +27,42 @@ function getElectronOzonePlatformHint(env: NodeJS.ProcessEnv = process.env): str
* Electron Wayland backend is unsupported); the Hyprland/Sway case is left untouched so
* {@link enforceUnsupportedWaylandMode} can report it.
*/
export function shouldForceX11ElectronBackend(env: NodeJS.ProcessEnv = process.env): boolean {
if (process.platform !== 'linux') return false;
export function shouldForceX11ElectronBackend(
env: NodeJS.ProcessEnv = process.env,
platform: NodeJS.Platform = process.platform,
): boolean {
if (platform !== 'linux') return false;
return !isSupportedWaylandCompositor(env);
}
export function resolveX11ElectronRelaunchArgs(
args: string[],
env: NodeJS.ProcessEnv = process.env,
platform: NodeJS.Platform = process.platform,
): string[] | null {
if (!shouldForceX11ElectronBackend(env, platform)) return null;
if (env[X11_ELECTRON_BOOTSTRAP_ENV] === '1') return null;
const retainedArgs: string[] = [];
let alreadyForced = false;
for (let index = 0; index < args.length; index += 1) {
const arg = args[index];
if (arg === '--ozone-platform') {
const value = args[index + 1];
alreadyForced = value?.trim().toLowerCase() === 'x11';
if (value && !value.startsWith('--')) index += 1;
continue;
}
if (arg?.startsWith('--ozone-platform=')) {
alreadyForced = arg.slice('--ozone-platform='.length).trim().toLowerCase() === 'x11';
continue;
}
if (arg) retainedArgs.push(arg);
}
return alreadyForced ? null : [...retainedArgs, X11_ELECTRON_OZONE_ARG];
}
export function forceX11Backend(args: CliArgs): void {
if (!shouldStartApp(args)) return;
if (!shouldForceX11ElectronBackend()) return;
+17 -1
View File
@@ -52,9 +52,15 @@ function resolveRuntimeDefaultNotificationIconPath(): string | null {
});
}
/**
* Live notifications keyed by `replaceId`. Electron exposes no native "replace this notification"
* flag, so a repeated status closes its predecessor instead of stacking a fresh toast per update.
*/
const notificationsByReplaceId = new Map<string, Electron.Notification>();
export function showDesktopNotification(
title: string,
options: { body?: string; icon?: string },
options: { body?: string; icon?: string; replaceId?: string },
): void {
const notificationOptions: {
title: string;
@@ -98,5 +104,15 @@ export function showDesktopNotification(
}
const notification = new Notification(notificationOptions);
const replaceId = options.replaceId?.trim();
if (replaceId) {
notificationsByReplaceId.get(replaceId)?.close();
notificationsByReplaceId.set(replaceId, notification);
notification.once('close', () => {
if (notificationsByReplaceId.get(replaceId) === notification) {
notificationsByReplaceId.delete(replaceId);
}
});
}
notification.show();
}
+59
View File
@@ -0,0 +1,59 @@
/**
* Loose semver ordering shared by the updater and the changelog UI.
* Returns >0 when `a` is newer, <0 when older, 0 when equal.
*/
export function compareSemverLike(a: string, b: string): number {
const parse = (
value: string,
): {
core: number[];
prerelease: Array<number | string>;
} => {
// Build metadata ("+build.2") is not part of precedence per semver, and
// leaving it attached makes it leak into the prerelease comparison.
const normalized = value.replace(/^v/i, '').split('+', 1)[0] ?? '';
const [coreText = '', ...prereleaseParts] = normalized.split('-');
const core = coreText
.split('.')
.slice(0, 3)
.map((part) => Number.parseInt(part, 10) || 0);
while (core.length < 3) core.push(0);
const prereleaseText = prereleaseParts.join('-');
return {
core,
prerelease: prereleaseText
? prereleaseText.split('.').map((part) => {
const numeric = Number.parseInt(part, 10);
return /^\d+$/.test(part) ? numeric : part;
})
: [],
};
};
const left = parse(a);
const right = parse(b);
for (let i = 0; i < 3; i += 1) {
const diff = (left.core[i] ?? 0) - (right.core[i] ?? 0);
if (diff !== 0) return diff;
}
if (left.prerelease.length === 0 && right.prerelease.length === 0) return 0;
if (left.prerelease.length === 0) return 1;
if (right.prerelease.length === 0) return -1;
const length = Math.max(left.prerelease.length, right.prerelease.length);
for (let i = 0; i < length; i += 1) {
const leftPart = left.prerelease[i];
const rightPart = right.prerelease[i];
if (leftPart === undefined && rightPart === undefined) return 0;
if (leftPart === undefined) return -1;
if (rightPart === undefined) return 1;
if (leftPart === rightPart) continue;
if (typeof leftPart === 'number' && typeof rightPart === 'number') {
return leftPart - rightPart;
}
if (typeof leftPart === 'number') return -1;
if (typeof rightPart === 'number') return 1;
return leftPart > rightPart ? 1 : -1;
}
return 0;
}
@@ -0,0 +1,7 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { normalizeTitleIdentity } from './title-normalization';
test('normalizeTitleIdentity produces a Unicode-aware comparison key', () => {
assert.equal(normalizeTitleIdentity(' BOCCHI・The ROCK!! '), 'bocchi the rock');
});
+8
View File
@@ -0,0 +1,8 @@
export function normalizeTitleIdentity(title: string): string {
return title
.normalize('NFKC')
.toLowerCase()
.replace(/[^\p{L}\p{N}]+/gu, ' ')
.trim()
.replace(/\s+/g, ' ');
}
+16
View File
@@ -25,8 +25,24 @@ import {
applyBackgroundBootstrapCommandLineSwitches,
applyEarlyLinuxCommandLineSwitches,
resolveLinuxPasswordStoreValue,
spawnDetachedApp,
} from './main-entry-runtime';
test('detached app launch policy stays in the startup runtime utilities', () => {
const entrySource = fs.readFileSync(path.join(process.cwd(), 'src/main-entry.ts'), 'utf8');
const runtimeSource = fs.readFileSync(
path.join(process.cwd(), 'src/main-entry-runtime.ts'),
'utf8',
);
assert.equal(typeof spawnDetachedApp, 'function');
assert.doesNotMatch(entrySource, /function spawnDetachedApp/);
assert.match(
runtimeSource,
/child\.once\('error', \(error\) => \{\s*console\.error\([^;]*error\);\s*\}\);\s*child\.unref\(\)/,
);
});
test('background bootstrap exits through Electron so Chromium children shut down', () => {
const exitCodes: number[] = [];
exitBackgroundBootstrap({ exit: (code) => exitCodes.push(code) });
+21
View File
@@ -1,7 +1,9 @@
import fs from 'node:fs';
import os from 'node:os';
import { spawn } from 'node:child_process';
import { CliArgs, hasExplicitCommand, parseArgs, shouldStartApp } from './cli/args';
import { resolveConfigDir } from './config/path-resolution';
import { resolveAppImageMountKeepaliveInvocation } from './main/appimage-mount-keepalive';
const BACKGROUND_ARG = '--background';
const START_ARG = '--start';
@@ -265,6 +267,25 @@ export function exitBackgroundBootstrap(app: BackgroundBootstrapAppLike): void {
app.exit(0);
}
export function spawnDetachedApp(childArgs: string[], env: NodeJS.ProcessEnv): void {
const keepalive = resolveAppImageMountKeepaliveInvocation(env);
const child = keepalive
? spawn(keepalive.command, [...keepalive.args, ...childArgs], {
detached: true,
stdio: 'ignore',
env,
})
: spawn(process.execPath, childArgs, {
detached: true,
stdio: 'ignore',
env,
});
child.once('error', (error) => {
console.error('Failed to spawn detached SubMiner app:', error);
});
child.unref();
}
export function shouldHandleHelpOnlyAtEntry(argv: string[], env: NodeJS.ProcessEnv): boolean {
if (env.ELECTRON_RUN_AS_NODE === '1') return false;
const args = parseCliArgs(argv);
+19 -16
View File
@@ -1,5 +1,4 @@
import os from 'node:os';
import { spawn } from 'node:child_process';
import { app, dialog, shell } from 'electron';
import { printHelp } from './cli/help';
import {
@@ -20,9 +19,9 @@ import {
shouldHandleHelpOnlyAtEntry,
shouldHandleLaunchMpvAtEntry,
shouldHandleStatsDaemonCommandAtEntry,
spawnDetachedApp,
} from './main-entry-runtime';
import { requestSingleInstanceLockEarly } from './main/early-single-instance';
import { resolveAppImageMountKeepaliveInvocation } from './main/appimage-mount-keepalive';
import { readConfiguredWindowsMpvLaunch } from './main-entry-launch-config';
import { isAppControlServerAvailable, sendAppControlCommand } from './shared/app-control-client';
import {
@@ -44,6 +43,10 @@ import {
resolveDefaultLogFilePath,
type LogRotation,
} from './shared/log-files';
import {
resolveX11ElectronRelaunchArgs,
X11_ELECTRON_BOOTSTRAP_ENV,
} from './core/utils/electron-backend';
const DEFAULT_TEXTHOOKER_PORT = 5174;
@@ -296,26 +299,26 @@ async function runEntryProcess(): Promise<void> {
return;
}
if (shouldDetachBackgroundLaunch(process.argv, process.env)) {
const childArgs = hasTransportedStartupArgs(process.env) ? [] : process.argv.slice(1);
const keepalive = resolveAppImageMountKeepaliveInvocation(process.env);
const child = keepalive
? spawn(keepalive.command, [...keepalive.args, ...childArgs], {
detached: true,
stdio: 'ignore',
env: sanitizeBackgroundEnv(process.env),
})
: spawn(process.execPath, childArgs, {
detached: true,
stdio: 'ignore',
env: sanitizeBackgroundEnv(process.env),
});
child.unref();
const x11ChildArgs = resolveX11ElectronRelaunchArgs(childArgs, process.env);
if (shouldDetachBackgroundLaunch(process.argv, process.env)) {
const childEnv = sanitizeBackgroundEnv(process.env);
if (x11ChildArgs) childEnv[X11_ELECTRON_BOOTSTRAP_ENV] = '1';
spawnDetachedApp(x11ChildArgs ?? childArgs, childEnv);
// Let Electron stop bootstrap Chromium children before its AppImage mount is released.
exitBackgroundBootstrap(app);
return;
}
if (x11ChildArgs) {
const childEnv = sanitizeStartupEnv(process.env);
childEnv[X11_ELECTRON_BOOTSTRAP_ENV] = '1';
spawnDetachedApp(x11ChildArgs, childEnv);
exitBackgroundBootstrap(app);
return;
}
startMainProcess();
}
+75 -19
View File
@@ -256,7 +256,6 @@ import {
import {
enforceUnsupportedWaylandMode,
forceX11Backend,
shouldForceX11ElectronBackend,
generateDefaultConfigFile,
resolveConfiguredShortcuts,
resolveKeybindings,
@@ -468,6 +467,8 @@ import { openJimakuModal as openJimakuModalRuntime } from './main/runtime/jimaku
import { openTsukihimeModal as openTsukihimeModalRuntime } from './main/runtime/tsukihime-open';
import { openSubsyncManualModal as openSubsyncManualModalRuntime } from './main/runtime/subsync-open';
import { openSessionHelpModal as openSessionHelpModalRuntime } from './main/runtime/session-help-open';
import { openChangelogModal as openChangelogModalRuntime } from './main/runtime/changelog-open';
import { createChangelogRuntime } from './main/runtime/changelog/changelog-runtime';
import { openCharacterDictionaryManagerModal as openCharacterDictionaryManagerModalRuntime } from './main/runtime/character-dictionary-open';
import { openControllerSelectModal as openControllerSelectModalRuntime } from './main/runtime/controller-select-open';
import { openControllerDebugModal as openControllerDebugModalRuntime } from './main/runtime/controller-debug-open';
@@ -487,6 +488,7 @@ import { createOverlayVisibilityRuntimeService } from './main/overlay-visibility
import { createDiscordPresenceRuntime } from './main/runtime/discord-presence-runtime';
import { createCharacterDictionaryRuntimeService } from './main/character-dictionary-runtime';
import { createCharacterDictionaryImageLookup } from './main/character-dictionary-runtime/image-lookup';
import { createCharacterNameCandidateLookup } from './main/character-dictionary-runtime/name-candidates';
import {
createCharacterDictionaryAutoSyncRuntimeService,
getCharacterDictionaryManagerSnapshot,
@@ -506,6 +508,7 @@ import { createStartupOsdSequencer } from './main/runtime/startup-osd-sequencer'
import {
INSTALL_UPDATE_ACTION_ID,
UPDATE_AVAILABLE_NOTIFICATION_ID,
VIEW_CHANGELOG_ACTION_ID,
} from './main/runtime/update/update-notifications';
import { createOverlayNotificationsRuntime } from './main/runtime/overlay-notifications-runtime';
import {
@@ -593,15 +596,6 @@ if (process.platform === 'linux') {
);
app.commandLine.appendSwitch('password-store', passwordStore);
createLogger('main').debug(`Applied --password-store ${passwordStore}`);
// Pin the overlay to XWayland on unsupported Wayland sessions (everything except
// Hyprland/Sway). `setAlwaysOnTop`/`moveTop` are no-ops under a native Wayland surface,
// so the overlay can only stay above mpv under X11/XWayland. The command-line switch is
// applied at module load (before app init) so it reliably wins over the late env-var
// fallback in forceX11Backend().
if (shouldForceX11ElectronBackend(process.env)) {
app.commandLine.appendSwitch('ozone-platform-hint', 'x11');
createLogger('main').debug('Forced ozone-platform-hint=x11 for XWayland overlay stacking');
}
}
app.setName('SubMiner');
@@ -1816,7 +1810,7 @@ function withCurrentSubtitleTiming(payload: SubtitleData): SubtitleData {
endTime: appState.mpvClient?.currentSubEnd ?? null,
};
}
function emitSubtitlePayload(payload: SubtitleData): void {
function emitSubtitlePayload(payload: SubtitleData, options?: { resumePrefetch?: boolean }): void {
const timedPayload = withCurrentSubtitleTiming(payload);
const currentSubtitleData = appState.currentSubtitleData;
const isAnnotationUpgrade = isSubtitleAnnotationUpgrade(currentSubtitleData, timedPayload);
@@ -1833,7 +1827,13 @@ function emitSubtitlePayload(payload: SubtitleData): void {
}
annotationSubtitleWsService.broadcast(timedPayload, frequencyOptions);
autoplayReadyGate.maybeSignalPluginAutoplayReady(timedPayload, { forceWhilePaused: true });
// resumePrefetch: false marks an emit that is not the end of the work for
// this line; prefetch stays paused until the subtitle processing controller
// settles so it does not compete with the on-screen line for the single
// Yomitan parser window.
if (options?.resumePrefetch !== false) {
subtitlePrefetchService?.resume();
}
}
function getCurrentAutoplaySubtitlePayload(): SubtitleData | null {
const payload = appState.currentSubtitleData;
@@ -1889,7 +1889,17 @@ const buildSubtitleProcessingControllerMainDepsHandler =
createBuildSubtitleProcessingControllerMainDepsHandler({
tokenizeSubtitle: async (text: string) =>
tokenizeSubtitleDeferred ? await tokenizeSubtitleDeferred(text) : { text, tokens: null },
emitSubtitle: (payload) => emitSubtitlePayload(payload),
// Controller emits never release the prefetch pause: the first emit for an
// uncached line is the provisional plain payload, sent before tokenization
// starts, so resuming on it would put prefetch back in contention with the
// on-screen line for the single parser window.
emitSubtitle: (payload) => emitSubtitlePayload(payload, { resumePrefetch: false }),
// The pause is released once the controller has no work left, which covers
// the runs that end without an emit (suppressed duplicate, failed
// tokenization) as well as the ones that deliver a payload.
onProcessingSettled: () => {
subtitlePrefetchService?.resume();
},
logDebug: (message) => {
logger.debug(`[subtitle-processing] ${message}`);
},
@@ -1926,7 +1936,7 @@ const autoplaySubtitlePrimingRuntime = createAutoplaySubtitlePrimingRuntime({
appState.activeParsedSubtitleMediaPath = mediaPath;
},
subtitleProcessingController,
emitSubtitlePayload: (payload) => emitSubtitlePayload(payload),
emitSubtitlePayload: (payload, options) => emitSubtitlePayload(payload, options),
getSubtitlePrefetchService: () => subtitlePrefetchService,
getLastObservedTimePos: () => lastObservedTimePos,
getVisibleOverlayVisible: () => overlayManager.getVisibleOverlayVisible(),
@@ -2532,6 +2542,10 @@ const characterDictionaryAutoSyncRuntime = createCharacterDictionaryAutoSyncRunt
},
{
hasParserWindow: () => Boolean(appState.yomitanParserWindow),
invalidateCharacterDictionaryLookups: () => {
characterDictionaryImageLookup.invalidate();
characterNameCandidateLookup.invalidate();
},
clearParserCaches: () => {
if (appState.yomitanParserWindow) {
clearYomitanParserCachesForWindow(appState.yomitanParserWindow);
@@ -2557,6 +2571,13 @@ const characterDictionaryImageLookup = createCharacterDictionaryImageLookup({
getCurrentMediaId: () => characterDictionaryAutoSyncRuntime.getCurrentMediaId(),
});
// Lets the Yomitan scan runtime skip name lookups at positions where no
// character name can start; absent candidates just mean the exhaustive scan.
const characterNameCandidateLookup = createCharacterNameCandidateLookup({
userDataPath: USER_DATA_PATH,
getCurrentMediaId: () => characterDictionaryAutoSyncRuntime.getCurrentMediaId(),
});
const overlayVisibilityRuntime = createOverlayVisibilityRuntimeService(
createBuildOverlayVisibilityRuntimeMainDepsHandler({
getMainWindow: () => overlayManager.getMainWindow(),
@@ -2828,6 +2849,14 @@ function openSessionHelpOverlay(): void {
);
}
function openChangelogOverlay(): void {
openOverlayHostedModalWithOsd(
openChangelogModalRuntime,
'Changelog overlay unavailable.',
'Failed to open changelog overlay.',
);
}
function openCharacterDictionaryManagerOverlay(): void {
openCharacterDictionaryManagerWithConfigGate({
isCharacterDictionaryEnabled: () => configService.getConfig().subtitleStyle.nameMatchEnabled,
@@ -3971,7 +4000,10 @@ const refreshCurrentSubtitleAfterKnownWordUpdate = (): void => {
}
subtitleProcessingController.invalidateTokenizationCache();
subtitlePrefetchService?.onSeek(lastObservedTimePos);
subtitleProcessingController.refreshCurrentSubtitle(appState.currentSubText);
if (!subtitleProcessingController.refreshCurrentSubtitle(appState.currentSubText)) {
// Idle controller: no settle is coming to release the pause above.
subtitlePrefetchService?.resume();
}
};
let hasAttemptedImmersionTrackerStartup = false;
const ensureImmersionTrackerStarted = (): void => {
@@ -4372,9 +4404,15 @@ const {
emitSubtitlePayload(payload);
},
onSubtitleChange: (text) => {
// Pause only; restarting the prefetch run here would discard in-flight
// tokenization work on every line. Real seeks restart via onTimePosUpdate.
subtitlePrefetchService?.pause();
subtitlePrefetchService?.onSeek(lastObservedTimePos);
subtitleProcessingController.onSubtitleChange(text);
if (!subtitleProcessingController.onSubtitleChange(text)) {
// Repeat of the current text: the controller is idle, so no settle is
// coming to release the pause. Resume now instead of idling prefetch
// for the rest of the cue.
subtitlePrefetchService?.resume();
}
},
refreshDiscordPresence: () => {
discordPresenceRuntime.publishDiscordPresence();
@@ -4618,6 +4656,7 @@ const {
getCharacterNameImage: (term) => characterDictionaryImageLookup.get(term),
getCurrentCharacterDictionaryMediaId: () =>
characterDictionaryAutoSyncRuntime.getCurrentMediaId(),
getCharacterNameCandidates: () => characterNameCandidateLookup.get(),
getFrequencyDictionaryEnabled: () =>
getRuntimeBooleanOption(
'subtitle.annotation.frequency',
@@ -5095,6 +5134,18 @@ flushPendingMpvLogWrites = () => {
void flushMpvLog();
};
const { getChangelogSnapshot } = createChangelogRuntime({
getInstalledVersion: () => app.getVersion(),
getUpdateChannel: () => configService.getConfig().updates.channel,
resourcesPath: process.resourcesPath,
appPath: app.getAppPath(),
dirname: __dirname,
joinPath: (...parts) => path.join(...parts),
fileExists: (candidate) => fs.existsSync(candidate),
readFile: (candidate) => fs.readFileSync(candidate, 'utf8'),
logWarn: (message) => logger.warn(message),
});
const { getUpdateService } = createUpdateServiceRuntime({
userDataPath: USER_DATA_PATH,
getUpdatesConfig: () => configService.getConfig().updates,
@@ -5463,6 +5514,12 @@ const { registerIpcRuntimeHandlers } = composeIpcRuntimeHandlers({
logger.warn('Failed to install update from overlay notification action:', error);
});
}
if (
notificationId === UPDATE_AVAILABLE_NOTIFICATION_ID &&
actionId === VIEW_CHANGELOG_ACTION_ID
) {
openChangelogOverlay();
}
if (actionId === OPEN_ANKI_CARD_ACTION_ID && noteId !== undefined) {
void openAnkiCardFromNotification(noteId).catch((error) => {
logger.warn('Failed to open Anki card from overlay notification action:', error);
@@ -5672,7 +5729,6 @@ const { registerIpcRuntimeHandlers } = composeIpcRuntimeHandlers({
if (result.ok && result.rebuildRequired) {
try {
await characterDictionaryAutoSyncRuntime.runSyncNow();
characterDictionaryImageLookup.invalidate();
} catch (error) {
logger.warn('Failed to rebuild character dictionary after manager override:', error);
}
@@ -5703,7 +5759,6 @@ const { registerIpcRuntimeHandlers } = composeIpcRuntimeHandlers({
if (result.ok && result.rebuildRequired) {
try {
await characterDictionaryAutoSyncRuntime.runSyncNow();
characterDictionaryImageLookup.invalidate();
} catch (error) {
logger.warn('Failed to rebuild character dictionary after manager removal:', error);
}
@@ -5720,7 +5775,6 @@ const { registerIpcRuntimeHandlers } = composeIpcRuntimeHandlers({
if (result.ok && result.rebuildRequired) {
try {
await characterDictionaryAutoSyncRuntime.runSyncNow();
characterDictionaryImageLookup.invalidate();
} catch (error) {
logger.warn('Failed to rebuild character dictionary after manager reorder:', error);
}
@@ -5728,6 +5782,7 @@ const { registerIpcRuntimeHandlers } = composeIpcRuntimeHandlers({
return result;
},
appendClipboardVideoToQueue: () => appendClipboardVideoToQueueHandler(),
getChangelogSnapshot: (options) => getChangelogSnapshot(options),
...playlistBrowserMainDeps,
getImmersionTracker: () => appState.immersionTracker,
},
@@ -6118,6 +6173,7 @@ const { ensureTray: ensureTrayHandler, destroyTray: destroyTrayHandler } =
initializeOverlayRuntime: () => initializeOverlayRuntime(),
isOverlayRuntimeInitialized: () => appState.overlayRuntimeInitialized,
openSessionHelpModal: () => openSessionHelpOverlay(),
openChangelogModal: () => openChangelogOverlay(),
openTexthookerInBrowser: () =>
handleCliCommand(parseArgs(['--texthooker', '--open-browser'])),
showTexthookerPage: () => shouldShowTexthookerTrayEntry(configService.getConfig()),
@@ -1903,7 +1903,7 @@ test('generateForCurrentMedia logs progress while resolving and rebuilding snaps
'[dictionary] current anime guess: The Eminence in Shadow (episode 5)',
'[dictionary] AniList match: The Eminence in Shadow -> AniList 130298',
'[dictionary] snapshot miss for AniList 130298, fetching characters',
'[dictionary] downloaded AniList character page 1 for AniList 130298',
'[dictionary] downloaded AniList character page 1 for AniList 130298 (1 characters)',
'[dictionary] downloading 1 images for AniList 130298',
'[dictionary] stored snapshot for AniList 130298: 16 terms',
'[dictionary] building ZIP for AniList 130298',
+48 -4
View File
@@ -64,6 +64,7 @@ export type {
CharacterDictionarySnapshotProgress,
CharacterDictionarySnapshotProgressCallbacks,
CharacterDictionarySnapshotResult,
CharacterDictionarySnapshotStageProgress,
MergedCharacterDictionaryBuildResult,
} from './character-dictionary-runtime/types';
@@ -363,19 +364,28 @@ export function createCharacterDictionaryRuntimeService(deps: CharacterDictionar
deps.logInfo?.(`[dictionary] snapshot stale for AniList ${mediaId}: ${refreshReason}`);
}
const progressMediaTitle = mediaTitleHint || `AniList ${mediaId}`;
progress?.onGenerating?.({
mediaId,
mediaTitle: mediaTitleHint || `AniList ${mediaId}`,
mediaTitle: progressMediaTitle,
});
deps.logInfo?.(`[dictionary] snapshot miss for AniList ${mediaId}, fetching characters`);
const { mediaTitle: fetchedMediaTitle, characters } = await fetchCharactersForMedia(
mediaId,
beforeRequest,
(page) => {
(page, charactersSoFar) => {
deps.logInfo?.(
`[dictionary] downloaded AniList character page ${page} for AniList ${mediaId}`,
`[dictionary] downloaded AniList character page ${page} for AniList ${mediaId} (${charactersSoFar} characters)`,
);
progress?.onGenerateProgress?.({
mediaId,
mediaTitle: progressMediaTitle,
stage: 'characters',
completed: charactersSoFar,
total: null,
page,
});
},
);
if (characters.length === 0) {
@@ -403,12 +413,26 @@ export function createCharacterDictionaryRuntimeService(deps: CharacterDictionar
);
}
let hasAttemptedImageDownload = false;
let attemptedImageCount = 0;
for (const entry of allImageUrls) {
if (hasAttemptedImageDownload) {
await sleepMs(CHARACTER_IMAGE_DOWNLOAD_DELAY_MS);
}
hasAttemptedImageDownload = true;
const image = await downloadCharacterImage(entry.url, entry.id);
attemptedImageCount += 1;
progress?.onGenerateProgress?.({
mediaId,
mediaTitle: progressMediaTitle,
stage: 'images',
completed: attemptedImageCount,
total: allImageUrls.length,
});
if (attemptedImageCount % 100 === 0) {
deps.logInfo?.(
`[dictionary] downloaded ${attemptedImageCount}/${allImageUrls.length} images for AniList ${mediaId}`,
);
}
if (!image) continue;
if (entry.kind === 'character') {
imagesByCharacterId.set(entry.id, {
@@ -425,11 +449,31 @@ export function createCharacterDictionaryRuntimeService(deps: CharacterDictionar
const nameSplitTokenizerAvailable = isNameSplitTokenizerAvailable();
const resolvedNameSplits = nameSplitTokenizerAvailable
? await resolveJapaneseNameSplits(characters, deps.tokenizeJapaneseName!, deps.logWarn)
? await resolveJapaneseNameSplits(
characters,
deps.tokenizeJapaneseName!,
deps.logWarn,
(completed, total) => {
progress?.onGenerateProgress?.({
mediaId,
mediaTitle: progressMediaTitle,
stage: 'names',
completed,
total,
});
},
)
: undefined;
const nameSplitSource =
resolvedNameSplits && resolvedNameSplits.size > 0 ? 'mecab' : 'heuristic';
progress?.onGenerateProgress?.({
mediaId,
mediaTitle: progressMediaTitle,
stage: 'saving',
completed: 0,
total: null,
});
const snapshot = buildSnapshotFromCharacters(
mediaId,
fetchedMediaTitle || mediaTitleHint || `AniList ${mediaId}`,
@@ -1,7 +1,7 @@
export const ANILIST_GRAPHQL_URL = 'https://graphql.anilist.co';
export const ANILIST_REQUEST_DELAY_MS = 2000;
export const CHARACTER_IMAGE_DOWNLOAD_DELAY_MS = 250;
export const CHARACTER_DICTIONARY_FORMAT_VERSION = 19;
export const CHARACTER_DICTIONARY_FORMAT_VERSION = 20;
export const CHARACTER_DICTIONARY_MERGED_TITLE = 'SubMiner Character Dictionary';
export const HONORIFIC_SUFFIXES = [
@@ -278,7 +278,7 @@ export async function fetchAniListMediaCandidateById(
export async function fetchCharactersForMedia(
mediaId: number,
beforeRequest?: () => Promise<void>,
onPageFetched?: (page: number) => void,
onPageFetched?: (page: number, charactersSoFar: number) => void,
): Promise<{
mediaTitle: string;
characters: CharacterRecord[];
@@ -345,7 +345,6 @@ export async function fetchCharactersForMedia(
},
beforeRequest,
);
onPageFetched?.(page);
const media = data.Media;
if (!media) {
@@ -415,6 +414,8 @@ export async function fetchCharactersForMedia(
});
}
onPageFetched?.(page, characters.length);
const hasNextPage = Boolean(media.characters?.pageInfo?.hasNextPage);
if (!hasNextPage) {
break;
@@ -0,0 +1,163 @@
import assert from 'node:assert/strict';
import * as fs from 'fs';
import * as os from 'os';
import * as path from 'path';
import test from 'node:test';
import { CHARACTER_DICTIONARY_FORMAT_VERSION } from './constants';
import { createCharacterNameCandidateLookup } from './name-candidates';
function writeSnapshot(outputDir: string, mediaId: number, entries: Array<[string, string]>): void {
const snapshotsDir = path.join(outputDir, 'snapshots');
fs.mkdirSync(snapshotsDir, { recursive: true });
fs.writeFileSync(
path.join(snapshotsDir, `anilist-${mediaId}.json`),
JSON.stringify({
formatVersion: CHARACTER_DICTIONARY_FORMAT_VERSION,
mediaId,
mediaTitle: `title-${mediaId}`,
entryCount: entries.length,
updatedAt: 1,
termEntries: entries.map(([term, reading]) => [
term,
reading,
'name main',
'',
100,
[],
0,
'',
]),
images: [],
}),
);
}
function withTempDir<T>(run: (dir: string) => T): T {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-name-candidates-'));
try {
return run(dir);
} finally {
fs.rmSync(dir, { recursive: true, force: true });
}
}
test('collects terms and readings for the current media', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [
['ミナト', 'みなと'],
['湊', 'みなと'],
]);
writeSnapshot(dir, 2, [['カズマ', 'かずま']]);
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => 1,
});
const candidates = lookup.get();
assert.ok(candidates);
assert.deepEqual([...candidates.forms].sort(), ['みなと', 'ミナト', '湊'].sort());
// Deduplicated: both entries share the みなと reading.
assert.equal(candidates.forms.length, 3);
});
});
test('returns null without a media scope so the scanner stays exhaustive', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [['ミナト', 'みなと']]);
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => null,
});
assert.equal(lookup.get(), null);
});
});
test('returns null for a media with no cached snapshot', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [['ミナト', 'みなと']]);
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => 999,
});
assert.equal(lookup.get(), null);
});
});
test('key changes when the snapshot content changes', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [['ミナト', 'みなと']]);
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => 1,
});
const first = lookup.get();
writeSnapshot(dir, 1, [
['ミナト', 'みなと'],
['アクア', 'あくあ'],
]);
lookup.invalidate();
const second = lookup.get();
assert.ok(first && second);
assert.notEqual(first.key, second.key);
assert.equal(second.forms.length, 4);
});
});
// The lookup runs once per subtitle line, so it must not stat the snapshot
// directory every call. Asserted behaviorally: an unannounced on-disk change is
// invisible until the recheck interval elapses, which can only be true if the
// filesystem is not consulted per lookup.
test('does not re-read the snapshot directory on every lookup', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [['ミナト', 'みなと']]);
let nowMs = 1_000_000;
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => 1,
now: () => nowMs,
});
assert.equal(lookup.get()?.forms.length, 2);
writeSnapshot(dir, 1, [
['ミナト', 'みなと'],
['アクア', 'あくあ'],
]);
nowMs += 1000;
assert.equal(lookup.get()?.forms.length, 2, 'expected the cached list within the interval');
nowMs += 10_000;
assert.equal(lookup.get()?.forms.length, 4, 'expected a refresh past the interval');
});
});
test('invalidate picks up a snapshot change immediately', () => {
withTempDir((dir) => {
writeSnapshot(dir, 1, [['ミナト', 'みなと']]);
let nowMs = 1_000_000;
const lookup = createCharacterNameCandidateLookup({
outputDir: dir,
getCurrentMediaId: () => 1,
now: () => nowMs,
});
assert.equal(lookup.get()?.forms.length, 2);
writeSnapshot(dir, 1, [
['ミナト', 'みなと'],
['アクア', 'あくあ'],
]);
nowMs += 1;
lookup.invalidate();
assert.equal(lookup.get()?.forms.length, 4);
});
});
@@ -0,0 +1,159 @@
import * as fs from 'fs';
import * as path from 'path';
import { readCachedSnapshots } from './cache';
import type { CharacterDictionarySnapshot } from './types';
// Candidate name forms for the greedy name pre-pass in the Yomitan scan
// runtime. The scanner otherwise has to ask the backend at every Japanese
// position, because a character name can start mid-token; knowing which forms
// exist lets it look up only where a name can actually begin.
//
// A form is any string Yomitan could match a character entry by: the term and
// its reading. Both come from the dictionary SubMiner generated, so the pair is
// the complete matchable set for an entry. Callers treat a missing list as
// "scan every position", so a stale or absent snapshot costs speed, never a
// missed name.
function getSnapshotsDir(outputDir: string): string {
return path.join(outputDir, 'snapshots');
}
function collectSnapshotNameForms(snapshot: CharacterDictionarySnapshot): string[] {
const forms = new Set<string>();
for (const entry of snapshot.termEntries) {
const term = typeof entry[0] === 'string' ? entry[0].trim() : '';
if (term) {
forms.add(term);
}
const reading = typeof entry[1] === 'string' ? entry[1].trim() : '';
if (reading) {
forms.add(reading);
}
}
return [...forms];
}
// The signature grows with the size of the dictionary library, and it rides
// along in every per-line scan call, so it is folded into a fixed-width digest
// first. Collisions only matter against the immediately previous signature (the
// runtime compares keys for equality), and FNV-1a over the file list is far
// beyond what that needs.
function digestSnapshotDirectorySignature(signature: string): string {
let hash = 0x811c9dc5;
for (let index = 0; index < signature.length; index += 1) {
hash ^= signature.charCodeAt(index);
hash = Math.imul(hash, 0x01000193);
}
return (hash >>> 0).toString(36);
}
function getSnapshotDirectorySignature(outputDir: string): string {
let entries: fs.Dirent[] = [];
try {
entries = fs.readdirSync(getSnapshotsDir(outputDir), { withFileTypes: true });
} catch {
return '';
}
const parts: string[] = [];
for (const entry of entries) {
if (!entry.isFile() || !/^anilist-\d+\.json$/.test(entry.name)) {
continue;
}
try {
const stat = fs.statSync(path.join(getSnapshotsDir(outputDir), entry.name));
parts.push(`${entry.name}:${stat.mtimeMs}:${stat.size}`);
} catch {
// Ignore files that disappear during a refresh; the next lookup rebuilds.
}
}
return parts.sort().join('|');
}
export interface CharacterNameCandidateSet {
/** Identifies this exact form list, so the scan runtime can cache it. */
key: string;
forms: string[];
}
// This lookup is consulted once per subtitle line, so it must not stat the
// snapshot directory every time. Dictionary writes are rare and always call
// invalidate(), which forces the next lookup to re-read; the interval only
// bounds staleness from changes made behind our back.
const SNAPSHOT_SIGNATURE_RECHECK_INTERVAL_MS = 5000;
export function createCharacterNameCandidateLookup(deps: {
userDataPath?: string;
outputDir?: string;
getCurrentMediaId?: () => number | null | undefined;
now?: () => number;
}): {
get: (mediaId?: number | null) => CharacterNameCandidateSet | null;
invalidate: () => void;
} {
const outputDir =
deps.outputDir ??
(deps.userDataPath ? path.join(deps.userDataPath, 'character-dictionaries') : '');
const now = deps.now ?? (() => Date.now());
let signature: string | null = null;
let lastSignatureCheckAtMs = 0;
let formsByMediaId = new Map<number, string[]>();
function refreshIfNeeded(): void {
if (!outputDir) {
formsByMediaId = new Map<number, string[]>();
signature = '';
return;
}
const nowMs = now();
if (
signature !== null &&
nowMs - lastSignatureCheckAtMs < SNAPSHOT_SIGNATURE_RECHECK_INTERVAL_MS
) {
return;
}
lastSignatureCheckAtMs = nowMs;
const nextSignature = getSnapshotDirectorySignature(outputDir);
if (nextSignature === signature) {
return;
}
signature = nextSignature;
formsByMediaId = new Map<number, string[]>();
for (const snapshot of readCachedSnapshots(outputDir)) {
const forms = collectSnapshotNameForms(snapshot);
if (forms.length > 0) {
formsByMediaId.set(snapshot.mediaId, forms);
}
}
}
return {
get(mediaId?: number | null): CharacterNameCandidateSet | null {
refreshIfNeeded();
const rawMediaId = mediaId ?? deps.getCurrentMediaId?.() ?? null;
const normalizedMediaId =
typeof rawMediaId === 'number' && Number.isFinite(rawMediaId) && rawMediaId > 0
? Math.floor(rawMediaId)
: null;
// Without a media scope the pre-pass would need every character of every
// cached title, which is both slow to match and pointless: report no
// candidates so the scanner keeps its exhaustive behavior.
if (normalizedMediaId === null) {
return null;
}
const forms = formsByMediaId.get(normalizedMediaId);
if (!forms || forms.length === 0) {
return null;
}
return {
key: `${digestSnapshotDirectorySignature(signature ?? '')}:${normalizedMediaId}`,
forms,
};
},
invalidate(): void {
signature = null;
lastSignatureCheckAtMs = 0;
},
};
}
@@ -1,3 +1,4 @@
import { isHanCodePoint } from '../../core/text/han-code-points';
import { HONORIFIC_SUFFIXES } from './constants';
import type { JapaneseNameParts, NameReadings, ResolvedNameSplits } from './types';
@@ -26,10 +27,12 @@ export function buildReading(term: string): string {
return katakanaToHiragana(compact);
}
// Code points, not code units: a supplementary-plane kanji (𠮷, U+20BB7) is a
// surrogate pair, and reading only the high surrogate would classify a real
// single-character name as non-kanji and drop it.
export function containsKanji(value: string): boolean {
for (const char of value) {
const code = char.charCodeAt(0);
if ((code >= 0x4e00 && code <= 0x9fff) || (code >= 0x3400 && code <= 0x4dbf)) {
if (isHanCodePoint(char.codePointAt(0) ?? 0)) {
return true;
}
}
@@ -86,8 +86,10 @@ export async function resolveJapaneseNameSplits(
characters: CharacterRecord[],
tokenize: NameSplitTokenizer,
logWarn?: (message: string) => void,
onCharacterResolved?: (completed: number, total: number) => void,
): Promise<Map<string, ResolvedNameSplit>> {
const splits = new Map<string, ResolvedNameSplit>();
let resolvedCharacters = 0;
for (const character of characters) {
const familyHintReading = buildReadingFromHint(character.lastNameHint?.trim() || '');
const givenHintReading = buildReadingFromHint(character.firstNameHint?.trim() || '');
@@ -113,6 +115,8 @@ export async function resolveJapaneseNameSplits(
splits.set(name, { family, given });
}
}
resolvedCharacters += 1;
onCharacterResolved?.(resolvedCharacters, characters.length);
}
return splits;
}
@@ -36,3 +36,134 @@ test('buildNameTerms adds surname honorifics from Japanese localized aliases', (
assert.ok(terms.includes('馬渕さん'));
assert.ok(!terms.includes('송치'));
});
test('buildNameTerms drops the disambiguator letter of a mob character name', () => {
const terms = buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'Joshi A',
nativeName: '女子A',
}),
);
// ア would match every あ〜 in the subtitles; the letter is a disambiguator
// (Girl A / Girl B), not a name.
assert.ok(!terms.includes('ア'));
assert.ok(!terms.includes('アさん'));
assert.ok(terms.includes('女子A'));
assert.ok(terms.includes('ジョシア'));
});
test('buildNameTerms keeps a character whose whole name is one kana', () => {
const terms = buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'A',
nativeName: 'あ',
}),
);
// The mob-label rule only judges parts a name was split into; a name the
// source gives us whole is the character's actual name.
assert.ok(terms.includes('あ'));
assert.ok(terms.includes('あさん'));
// Romanized forms are never lookup targets (the subtitles are Japanese), and
// the single-kana alias "A" transliterates to is dropped as a collision.
assert.ok(!terms.includes('A'));
assert.ok(!terms.includes('ア'));
});
test('buildNameTerms keeps a one-character name written in another script', () => {
const terms = buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'Byeol',
nativeName: '별',
alternativeNames: ['Я'],
}),
);
assert.ok(terms.includes('별'));
assert.ok(terms.includes('별さん'));
assert.ok(terms.includes('Я'));
});
test('buildNameTerms yields nothing for a character whose only name is a bare letter', () => {
// Documented policy rather than an oversight: a romanized name is never a
// term on its own (the subtitles are Japanese), and the single kana a bare
// letter transliterates to would match every あ〜 in the line.
assert.deepEqual(
buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'A',
nativeName: '',
}),
),
[],
);
});
test('buildNameTerms keeps one-character split parts that are not mob labels', () => {
const hangul = buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'Byeol Kim',
nativeName: '별 김',
}),
);
assert.ok(hangul.includes('별'));
assert.ok(hangul.includes('김'));
const middleDot = buildNameTerms(
characterRecord({
firstNameHint: '',
lastNameHint: '',
fullName: 'A Be',
nativeName: 'ア・ベ',
}),
);
assert.ok(middleDot.includes('ア'));
assert.ok(middleDot.includes('ベ'));
});
test('buildNameTerms keeps a single-kanji name part', () => {
// The name is an alias, not the native name, so the parts come from the
// space split rather than from the native-name split.
const terms = buildNameTerms(
characterRecord({
firstNameHint: 'Sora',
lastNameHint: 'Yamada',
fullName: 'Sora Yamada',
nativeName: '',
alternativeNames: ['山田 空'],
}),
);
assert.ok(terms.includes('山田'));
assert.ok(terms.includes('空'));
});
test('buildNameTerms keeps a single supplementary-plane kanji name part', () => {
// 𠮷 (U+20BB7) is a surrogate pair: a code-unit kanji check reads only the
// high surrogate and drops the part as if it were a mob disambiguator.
const terms = buildNameTerms(
characterRecord({
firstNameHint: 'Tsukasa',
lastNameHint: 'Yoshi',
fullName: 'Tsukasa Yoshi',
nativeName: '',
alternativeNames: ['𠮷 司'],
}),
);
assert.ok(terms.includes('𠮷'));
assert.ok(terms.includes('司'));
});
@@ -1,3 +1,4 @@
import { HAN_REGEXP_CLASS_BODY } from '../../core/text/han-code-points';
import { HONORIFIC_SUFFIXES } from './constants';
import {
addRomanizedKanaAliases,
@@ -42,11 +43,29 @@ export function expandRawNameVariants(rawName: string): string[] {
return [...variants];
}
// The label AniList appends to unnamed mob characters: one letter or digit,
// halfwidth or fullwidth (女子A / "Joshi A" / 女子1). Nothing else qualifies —
// a one-character part in any script is a real name part (별 김, ア・ベ, 山田 空).
const SINGLE_LABEL_CHARACTER = /^[0-9A-Za-z\uff10-\uff19\uff21-\uff3a\uff41-\uff5a]$/u;
// Judged on split parts only: a name the source gives us whole in a script the
// subtitles can contain is kept whatever it looks like, because a character
// really can be called あ or 별. (A romanized name is a separate matter: it is
// never a term on its own, only a source of kana aliases. See below.)
function isUsableNameSplitPart(part: string): boolean {
return !SINGLE_LABEL_CHARACTER.test(part);
}
// Kana, Han (shared ranges), and the marks that only ever appear inside a
// Japanese name: iteration marks and the small ka/ke used in place names.
const JAPANESE_NAME_CHARACTERS = new RegExp(
`^[\\u3040-\\u30ff${HAN_REGEXP_CLASS_BODY}\u3005\u3006\u30f5\u30f6\u30fc]+$`,
'u',
);
export function isJapaneseNameSplitCandidate(name: string): boolean {
const compact = name.replace(/[\s\u3000・・·•]/g, '');
return (
containsKanji(compact) && /^[\u3040-\u30ff\u3400-\u4dbf\u4e00-\u9fff々〆ヵヶー]+$/.test(compact)
);
return containsKanji(compact) && JAPANESE_NAME_CHARACTERS.test(compact);
}
function addJapaneseNameParts(
@@ -97,8 +116,11 @@ export function buildNameTerms(
const split = name.split(/[\s\u3000]+/).filter((part) => part.trim().length > 0);
if (split.length === 2) {
target.add(split[0]!);
target.add(split[1]!);
for (const part of split) {
if (isUsableNameSplitPart(part)) {
target.add(part);
}
}
}
const splitByMiddleDot = name
@@ -107,9 +129,11 @@ export function buildNameTerms(
.filter((part) => part.length > 0);
if (splitByMiddleDot.length >= 2) {
for (const part of splitByMiddleDot) {
if (isUsableNameSplitPart(part)) {
target.add(part);
}
}
}
if (target === base) {
addJapaneseNameParts(character, name, base, resolvedSplits);
@@ -117,7 +141,15 @@ export function buildNameTerms(
}
}
// Romanized names never become terms themselves — the subtitles are Japanese,
// so "Joshi A" would never appear in one — they only contribute the kana a
// Japanese writer would spell them with.
for (const alias of addRomanizedKanaAliases(romanizedBase)) {
// Except when the whole name is one letter: it transliterates to a single
// kana (A → ア) that matches every あ〜 in the subtitles. A character whose
// only recorded name is a bare letter therefore yields no terms at all,
// which is the intended outcome: those are unnamed mob characters.
if ([...alias].length === 1) continue;
base.add(alias);
}
@@ -122,9 +122,24 @@ export type CharacterDictionarySnapshotProgress = {
mediaTitle: string;
};
export type CharacterDictionarySnapshotStage = 'characters' | 'images' | 'names' | 'saving';
/**
* Fine-grained generation progress. `total` is null while the work size is still unknown (AniList
* paginates characters, so the character count only settles on the last page).
*/
export type CharacterDictionarySnapshotStageProgress = CharacterDictionarySnapshotProgress & {
stage: CharacterDictionarySnapshotStage;
completed: number;
total: number | null;
/** AniList page currently being downloaded; only set during the `characters` stage. */
page?: number;
};
export type CharacterDictionarySnapshotProgressCallbacks = {
onChecking?: (progress: CharacterDictionarySnapshotProgress) => void;
onGenerating?: (progress: CharacterDictionarySnapshotProgress) => void;
onGenerateProgress?: (progress: CharacterDictionarySnapshotStageProgress) => void;
};
export type MergedCharacterDictionaryBuildResult = {
@@ -3,7 +3,7 @@ import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import test from 'node:test';
import { buildDictionaryZip } from './zip';
import { buildDictionaryZip, readDictionaryZipRevision } from './zip';
import type { CharacterDictionaryTermEntry } from './types';
function makeTempDir(): string {
@@ -105,3 +105,63 @@ test('buildDictionaryZip writes a valid stored zip without fs.writeFileSync', ()
cleanupDir(tempDir);
}
});
test('readDictionaryZipRevision reads the built revision and rejects foreign archives', () => {
const dir = makeTempDir();
try {
const zipPath = path.join(dir, 'merged.zip');
buildDictionaryZip(
zipPath,
'SubMiner Character Dictionary',
'Character names',
'rev-42',
[
{
term: 'ルフィ',
reading: 'ルフィ',
role: 'main',
glossary: [],
} as unknown as CharacterDictionaryTermEntry,
],
[],
);
assert.equal(readDictionaryZipRevision(zipPath), 'rev-42');
assert.equal(readDictionaryZipRevision(path.join(dir, 'missing.zip')), null);
const archive = fs.readFileSync(zipPath);
const truncatedPath = path.join(dir, 'truncated.zip');
fs.writeFileSync(truncatedPath, archive.subarray(0, 40));
assert.equal(readDictionaryZipRevision(truncatedPath), null);
// An archive cut short after index.json still holds a readable revision, but importing it
// would hand Yomitan a half-written file: the missing end-of-central-directory record has to
// reject it. One byte off the end is enough to make the record incomplete.
for (const missingBytes of [1, 22, archive.length - 200]) {
const cutPath = path.join(dir, `cut-${missingBytes}.zip`);
fs.writeFileSync(cutPath, archive.subarray(0, archive.length - missingBytes));
assert.equal(readDictionaryZipRevision(cutPath), null, `cut of ${missingBytes} bytes`);
}
// Same size, corrupt directory: a record overwritten in place has to be rejected too.
const centralStart = archive.readUInt32LE(archive.length - 22 + 16);
const brokenSignaturePath = path.join(dir, 'broken-signature.zip');
const brokenSignature = Buffer.from(archive);
brokenSignature.writeUInt32LE(0xdeadbeef, centralStart);
fs.writeFileSync(brokenSignaturePath, brokenSignature);
assert.equal(readDictionaryZipRevision(brokenSignaturePath), null);
const brokenLengthPath = path.join(dir, 'broken-length.zip');
const brokenLength = Buffer.from(archive);
// Name length that runs the walk past the end of the directory.
brokenLength.writeUInt16LE(0xffff, centralStart + 28);
fs.writeFileSync(brokenLengthPath, brokenLength);
assert.equal(readDictionaryZipRevision(brokenLengthPath), null);
const foreignPath = path.join(dir, 'foreign.zip');
fs.writeFileSync(foreignPath, Buffer.from('not a zip at all', 'utf8'));
assert.equal(readDictionaryZipRevision(foreignPath), null);
} finally {
cleanupDir(dir);
}
});
+18 -1
View File
@@ -1,5 +1,5 @@
import * as path from 'path';
import { writeStoredZip } from '../../shared/stored-zip';
import { readStoredZipFirstFile, writeStoredZip } from '../../shared/stored-zip';
import { ensureDir } from './fs-utils';
import type { CharacterDictionarySnapshotImage, CharacterDictionaryTermEntry } from './types';
@@ -31,6 +31,23 @@ function createTagBank(): Array<[string, string, number, string, number]> {
];
}
/**
* Revision recorded inside a built dictionary ZIP, or null when the archive is missing, truncated,
* or not one of ours. `index.json` is always the first entry written by {@link buildDictionaryZip}.
*/
export function readDictionaryZipRevision(zipPath: string): string | null {
const firstFile = readStoredZipFirstFile(zipPath);
if (!firstFile || firstFile.name !== 'index.json') {
return null;
}
try {
const index = JSON.parse(firstFile.data.toString('utf8')) as { revision?: unknown };
return typeof index.revision === 'string' && index.revision.length > 0 ? index.revision : null;
} catch {
return null;
}
}
export function buildDictionaryZip(
outputPath: string,
dictionaryTitle: string,
+2
View File
@@ -109,6 +109,7 @@ export interface MainIpcRuntimeServiceDepsParams {
removeCharacterDictionaryManagedEntry?: IpcDepsRuntimeOptions['removeCharacterDictionaryManagedEntry'];
moveCharacterDictionaryManagedEntry?: IpcDepsRuntimeOptions['moveCharacterDictionaryManagedEntry'];
appendClipboardVideoToQueue: IpcDepsRuntimeOptions['appendClipboardVideoToQueue'];
getChangelogSnapshot?: IpcDepsRuntimeOptions['getChangelogSnapshot'];
getPlaylistBrowserSnapshot: IpcDepsRuntimeOptions['getPlaylistBrowserSnapshot'];
appendPlaylistBrowserFile: IpcDepsRuntimeOptions['appendPlaylistBrowserFile'];
playPlaylistBrowserIndex: IpcDepsRuntimeOptions['playPlaylistBrowserIndex'];
@@ -302,6 +303,7 @@ export function createMainIpcRuntimeServiceDeps(
removeCharacterDictionaryManagedEntry: params.removeCharacterDictionaryManagedEntry,
moveCharacterDictionaryManagedEntry: params.moveCharacterDictionaryManagedEntry,
appendClipboardVideoToQueue: params.appendClipboardVideoToQueue,
getChangelogSnapshot: params.getChangelogSnapshot,
getPlaylistBrowserSnapshot: params.getPlaylistBrowserSnapshot,
appendPlaylistBrowserFile: params.appendPlaylistBrowserFile,
playPlaylistBrowserIndex: params.playPlaylistBrowserIndex,
+34 -11
View File
@@ -223,7 +223,7 @@ test('update overlay notification action triggers install flow', () => {
assert.match(runtimeSource, /fallbackClient\.openNoteInBrowser\(noteId\)/);
});
test('subtitle change re-prioritizes prefetch around live playback before tokenizing current line', () => {
test('subtitle change pauses prefetch without restarting its run before tokenizing current line', () => {
const source = readMainSource();
const actionBlock = source.match(
/onSubtitleChange:\s*\(text\)\s*=>\s*\{(?<body>[\s\S]*?)\n \},\n refreshDiscordPresence:/,
@@ -231,15 +231,19 @@ test('subtitle change re-prioritizes prefetch around live playback before tokeni
assert.ok(actionBlock);
assert.match(actionBlock, /subtitlePrefetchService\?\.pause\(\);/);
assert.match(actionBlock, /subtitlePrefetchService\?\.onSeek\(lastObservedTimePos\);/);
assert.match(actionBlock, /subtitleProcessingController\.onSubtitleChange\(text\);/);
// Restarting the run per line (onSeek) discards in-flight prefetch work;
// only real seeks restart via onTimePosUpdate.
assert.doesNotMatch(actionBlock, /subtitlePrefetchService\?\.onSeek\(/);
assert.match(actionBlock, /subtitleProcessingController\.onSubtitleChange\(text\)/);
assert.ok(
actionBlock.indexOf('subtitlePrefetchService?.pause();') <
actionBlock.indexOf('subtitlePrefetchService?.onSeek(lastObservedTimePos);'),
actionBlock.indexOf('subtitleProcessingController.onSubtitleChange(text)'),
);
assert.ok(
actionBlock.indexOf('subtitlePrefetchService?.onSeek(lastObservedTimePos);') <
actionBlock.indexOf('subtitleProcessingController.onSubtitleChange(text);'),
// A repeated subtitle emits nothing, so the pause has to be released here or
// prefetching idles until the next distinct line.
assert.match(
actionBlock,
/if \(!subtitleProcessingController\.onSubtitleChange\(text\)\) \{[\s\S]*?subtitlePrefetchService\?\.resume\(\);/,
);
});
@@ -489,16 +493,35 @@ test('known-word updates invalidate prefetched tokenizations before refreshing c
assert.match(actionBlock, /subtitlePrefetchService\?\.onSeek\(lastObservedTimePos\);/);
assert.match(
actionBlock,
/subtitleProcessingController\.refreshCurrentSubtitle\(appState\.currentSubText\);/,
/if \(!subtitleProcessingController\.refreshCurrentSubtitle\(appState\.currentSubText\)\) \{[\s\S]*?subtitlePrefetchService\?\.resume\(\);/,
);
assert.ok(
actionBlock.indexOf('subtitleProcessingController.invalidateTokenizationCache();') <
actionBlock.indexOf(
'subtitleProcessingController.refreshCurrentSubtitle(appState.currentSubText);',
'subtitleProcessingController.refreshCurrentSubtitle(appState.currentSubText)',
),
);
});
test('subtitle processing controller resumes prefetch on settle, not on its emits', () => {
const source = readMainSource();
const depsBlock = source.match(
/createBuildSubtitleProcessingControllerMainDepsHandler\(\{(?<body>[\s\S]*?)\n \}\);/,
)?.groups?.body;
assert.ok(depsBlock);
// A controller emit can be the provisional plain payload sent before the
// scan runs, so it must not release the prefetch pause.
assert.match(
depsBlock,
/emitSubtitle: \(payload\) => emitSubtitlePayload\(payload, \{ resumePrefetch: false \}\),/,
);
assert.match(
depsBlock,
/onProcessingSettled: \(\) => \{\s+subtitlePrefetchService\?\.resume\(\);/,
);
});
test('manual visible overlay changes notify mpv plugin visibility state', () => {
const source = readMainSource();
const setBlock = source.match(
@@ -593,7 +616,7 @@ test('YouTube media cache lifecycle routes through configured status notificatio
test('subtitle broadcasts share one frequency options snapshot per emitted payload', () => {
const source = readMainSource();
const emitBlock = source.match(
/function emitSubtitlePayload\(payload: SubtitleData\): void \{(?<body>[\s\S]*?)\n\}/,
/function emitSubtitlePayload\([\s\S]*?\): void \{(?<body>[\s\S]*?)\n\}/,
)?.groups?.body;
const frequencyOptionsSnapshot = emitBlock?.match(
/const frequencyDictionary = configService\.getConfig\(\)\.subtitleStyle\.frequencyDictionary;(?<body>[\s\S]*?)\n \};/,
@@ -616,7 +639,7 @@ test('subtitle broadcasts share one frequency options snapshot per emitted paylo
test('annotation upgrades skip the duplicate basic websocket event', () => {
const source = readMainSource();
const emitBlock = source.match(
/function emitSubtitlePayload\(payload: SubtitleData\): void \{(?<body>[\s\S]*?)\n\}/,
/function emitSubtitlePayload\([\s\S]*?\): void \{(?<body>[\s\S]*?)\n\}/,
)?.groups?.body;
assert.ok(emitBlock);

Some files were not shown because too many files have changed in this diff Show More