Compare commits

..

9 Commits

Author SHA1 Message Date
sudacode 4a7c9b04f1 fix(stats): enforce cleanup mode exclusivity and JSON requests
- Reject conflicting explicit cleanup modes
- Require application/json for duplicate-line maintenance requests
2026-08-12 00:49:32 -07:00
sudacode 9fd2dfd7a0 fix(cli): preserve stats cleanup lookback validation
- Parse the full inline lookback-days value before validation
2026-08-11 22:31:56 -07:00
sudacode 046a87a826 fix(stats): harden duplicate-line cleanup and tracking
- Reject malformed requests and sub-day cleanup windows
- Reset deduplication across media and subtitle-track changes
- Handle karaoke bursts ending with a long hold frame
2026-08-11 22:20:50 -07:00
sudacode fa4c734e30 docs(stats): document duplicate-line cleanup flag rules and forwarded app flags 2026-08-11 18:41:29 -07:00
sudacode b64e1264dc test(stats): click backdrop, not disabled button, in cleanup close test
- Assert the Close button is disabled while apply is in flight
- Click the always-enabled backdrop instead, since that's the path that actually reaches the close guard
2026-08-11 00:16:46 -07:00
sudacode c2561f8c4c fix(stats): fix duplicate-line cleanup reload tracking
- Track reload-owed state separately from the displayed result so a later scan or window change can't erase it before the reload runs
- Refuse to close mid-apply so the pending reload isn't dropped
- Add tests covering these edge cases
2026-08-11 00:04:18 -07:00
sudacode 039aed79c3 fix(stats): fix duplicate-line cleanup edge cases
- Drain the full write queue before scanning, not just one batch, so pending burst rows aren't missed
- Floor lookback-days before the positivity check so a sub-day value no longer collapses to a zero-day window
- Let the parsed cue list override the streaming dedup heuristic wherever it covers a line
- Reject combining --lifetime with --duplicate-lines and non-positive --lookback-days values
- Widen the burst frame bound to catch heavier typesetting while still sparing longer-spaced runs
- Keep the cleanup modal open until the user closes it so they can read the result before it unmounts
2026-08-10 23:50:48 -07:00
sudacode 684ab9eaff fix(stats): stop counting duplicate typeset subtitle lines
- Collapse animation-burst subtitle lines (karaoke OPs, animated signs) at ingest time using the same dedup rules the subtitle sidebar already applies, so repeated frames no longer flood "Top Repeated Words"
- Add retroactive cleanup for stats already affected: a "Duplicates" scanner/cleaner in the Vocabulary tab and `subminer stats cleanup --duplicate-lines` (`--dry-run`, `--lookback-days`) on the CLI
- Only subtitle lines and the vocabulary counts they feed are touched; watch time and lines-seen totals are left as recorded
2026-08-10 23:18:43 -07:00
sudacode 7b0fbdf254 fix(subtitles): collapse duplicate ASS events and decode text once (#186) 2026-08-10 22:21:44 -07:00
75 changed files with 3756 additions and 3820 deletions
+5
View File
@@ -0,0 +1,5 @@
type: fixed
area: stats
- Typeset subtitles no longer flood the stats. Karaoke openings and animated signs are authored as one subtitle event per animation frame, and immersion tracking counted every frame, which was enough to put an OP lyric at the top of "Top Repeated Words" for good. Lines are now collapsed on the way in using the same rules the subtitle sidebar already applies: matching parsed timings record exactly the cues the sidebar shows, while shifted, changing, or unparsed sources use a strict fallback where identical, contiguous, sub-0.1s lines stop counting after a few frames. Ordinary repeated dialogue and rewatches are unaffected.
- Added a cleanup for stats already affected. The Vocabulary tab has a **Duplicates** button that scans a chosen window (7 days through all time), shows the bursts it found and the word and kanji counts they added, and collapses each run to one line once confirmed. `subminer stats cleanup --duplicate-lines` does the same from the terminal, with `--dry-run` and `--lookback-days <n>`. Only subtitle lines and the vocabulary counts they feed are touched; watch time and lines-seen totals are left as recorded.
-6
View File
@@ -1,6 +0,0 @@
type: added
area: stats
- Library: duplicate cards for the same show can now be combined. Press "Select" above the library grid, tick the cards, and use "Merge Selected"; the dialog picks which entry to keep and moves every episode onto it. Sessions, mined cards, and watch time are preserved, the emptied entries disappear, and remembered title aliases keep future episodes on the merged card.
- Library: episodes can be reassigned to another library entry from the "→" button on an episode row, which is the fix when one file lands under a stray title (e.g. an episode name parsed as the series). Emptying an entry this way removes it and returns to the grid.
- Library: exact AniList title matches with compatible seasons fold duplicate cards automatically. Fuzzy same-AniList matches appear as dismissible "Possible duplicate" reviews instead of changing the library without confirmation; conflicting explicit seasons are left alone.
+5
View File
@@ -0,0 +1,5 @@
type: fixed
area: subtitles
- Heavily typeset ASS scripts (karaoke OP/ED, sign work) no longer fill the subtitle sidebar with garbage. Vector drawing runs (`\p1``\p0`) are no longer shown as subtitle text (e.g. `m 20 0 b 10 0 0 10 0 20 …`), and duplicate events in a parsed subtitle file collapse into one cue: identical text over an identical span (layered "shadow" copies), and per-frame animation bursts. An ASS burst has to prove itself with authoring evidence — a temporal tag (`\t`, `\move`, karaoke timing), an animated `Effect` column, or override values that change from event to event — plus one shared style and actor, so three rapid `えっ` reactions from three characters, or a sign repeated with the same static `\clip`, stay separate. SRT and VTT carry no such metadata, so there the run has to be at least five contiguous events all shorter than 0.1s, which is where ASS-to-SRT conversion leaves karaoke frames.
- Subtitle text is now decoded from ASS exactly once, where it enters the app, and matches how mpv renders the same line (including `\N`, `\n`, `\h`, unclosed `{`, and the fact that `\{` is not an escape). Renderer, timing tracker, tokenizer and the tokenization cache take that decoded text as-is instead of each re-deriving it, so one authored line can no longer produce two different cache keys, and a cue that normalizes to nothing is no longer stored as subtitle text or cached under an empty key.
+27 -8
View File
@@ -34,7 +34,7 @@ The same immersion data powers the stats dashboard.
- In-app overlay: focus the visible overlay, then press the key from `stats.toggleKey` (default: `` ` `` / `Backquote`).
- Launcher command: run `subminer stats` to start the local stats server on demand (it also opens the dashboard in your browser when `stats.autoOpenBrowser` is enabled; the default is `false`).
- Background server: run `subminer stats -b` to start or reuse a dedicated background stats daemon without keeping the launcher attached, and `subminer stats -s` to stop that daemon.
- Maintenance commands: run `subminer stats cleanup` or `subminer stats cleanup -v` to backfill/repair vocabulary metadata (`headword`, `reading`, POS) and purge stale or excluded rows from `imm_words` on demand; `subminer stats cleanup -l` repairs lifetime summary tables. `subminer stats rebuild` and `subminer stats backfill` rebuild or backfill rollup data.
- Maintenance commands: run `subminer stats cleanup` or `subminer stats cleanup -v` to backfill/repair vocabulary metadata (`headword`, `reading`, POS) and purge stale or excluded rows from `imm_words` on demand; `subminer stats cleanup -l` repairs lifetime summary tables; `subminer stats cleanup --duplicate-lines` collapses repeated lines left behind by typeset subtitles (see [Repeated Line Cleanup](#repeated-line-cleanup)). `subminer stats rebuild` and `subminer stats backfill` rebuild or backfill rollup data.
- Browser page: open `http://127.0.0.1:6969` directly if the local stats server is already running.
### Dashboard Tabs
@@ -57,13 +57,6 @@ Jellyfin stream URLs are normalized to stable item links before stats titles are
When YouTube channel metadata is available, the Library tab groups videos by creator/channel and treats each tracked video as an episode-like entry inside that channel section.
A library entry is identified by its parsed title plus any detected season, so the same show can end up on several cards when releases disagree about the title or omit the season tag. Two fixes are available:
- **Merge duplicates.** Hit **Select** above the grid, tick the cards that are the same show, and choose **Merge Selected**. Pick which entry to keep in the dialog; every episode moves onto it and the other cards are removed. Nothing is deleted, so sessions, mined cards and watch time all carry over. SubMiner remembers the merged title variants, so future episodes parsed with one of those names join the kept entry instead of recreating a duplicate card.
- **Move a single episode.** Hover an episode row in a title's episode list and use the **→** button to reassign it to another library entry. If that was the entry's last episode, the now-empty card is removed and you are returned to the grid.
Once cover art resolves a series to an AniList entry, cards with compatible seasons are folded together automatically only when the searched title exactly matches an AniList title or synonym. A fuzzy result that points at an AniList entry already used by another card appears as a **Possible duplicate** review above the Library grid instead. Choose **Review merge** to compare the cards and pick which one to keep, or **Not duplicates** to dismiss that suggestion permanently. Entries with conflicting explicit season numbers are left alone rather than merged or suggested.
Open a title and use **Delete Entry** in its header to remove a mistakenly tracked show outright. This deletes every episode of that title along with their sessions, subtitle lines, rollups and cover art, drops the words and kanji that were only seen there, and removes the card from the Library grid. Individual episodes and sessions can still be deleted on their own from the episode list and session rows. Entry deletion is refused while that title is the one currently playing.
![Stats Library](/screenshots/stats-library.png)
@@ -132,6 +125,32 @@ Secondary subtitle text (typically English translations) is stored alongside pri
The Vocabulary tab toolbar includes an **Exclusions** button for hiding words from all vocabulary views. Excluded words are stored in the immersion database, with older browser localStorage exclusions imported on first load after upgrade. They can be managed (restored or cleared) from the exclusion modal. Exclusions affect stat cards, charts, the frequency rank table, and the word list.
### Repeated Line Cleanup
Karaoke openings and animated signs are authored as one subtitle event per animation frame, all carrying the same text. Playback reports every one of those frames, so a single OP lyric could be recorded hundreds of times and dominate "Top Repeated Words".
Recording now collapses those runs as they happen, matching what the subtitle sidebar shows:
- When the active subtitle source has been parsed, its cue list has already had duplicate events and animation bursts merged. A line landing inside a surviving cue but after that cue's start is a frame the sidebar merged away, and is not recorded.
- When no parsed cue covers the live timing, including while a subtitle source is changing or shifted, the strict metadata-free rule applies: a run of identical, contiguous lines each shorter than 0.1s stops being recorded after a few frames. Ordinary repeated dialogue, and lines held for a normal beat, always record.
For stats recorded before this, the Vocabulary tab toolbar has a **Duplicates** button:
- Pick how far back to look (7 days, 30 days, 90 days, 1 year, or all time). A narrower window does less work and keeps older history untouched.
- **Scan** reports the bursts found, the lines they added, and the word and kanji counts they inflated, without writing anything.
- **Clean Up** applies exactly what the scan reported: each run collapses to its first line (extended to cover the run), and the removed lines' word and kanji occurrences are subtracted from the vocabulary aggregates.
The same thing runs from the terminal:
```bash
subminer stats cleanup --duplicate-lines --dry-run --lookback-days 30
subminer stats cleanup --duplicate-lines --lookback-days 30
```
`--duplicate-lines` (short: `-d`) picks the cleanup mode, so it cannot be combined with `--vocab` or `--lifetime`, and `--dry-run` and `--lookback-days <days>` only apply to it. Omitting `--lookback-days` scans all history; the value must be at least one day.
Runs never cross a session boundary, so rewatching an episode keeps both watches. Session telemetry (watch time, lines seen, tokens seen) and the rollups derived from it are left as recorded: they are cumulative samples taken during playback, and cannot be recomputed for sessions whose raw rows have since been pruned.
## Retention Defaults
By default, SubMiner keeps all retention tables and raw data (`0` means keep all) while continuing daily/monthly rollup maintenance:
+1
View File
@@ -151,6 +151,7 @@ subminer stats -b # start background stats daemon
| `subminer stats` | Start the stats server (opens the dashboard when `stats.autoOpenBrowser` is on) |
| `subminer stats -b` / `-s` | Start/reuse or stop the background stats daemon |
| `subminer stats cleanup` | Backfill vocabulary metadata and prune stale rows (`-v` vocab, `-l` lifetime summaries) |
| `subminer stats cleanup -d` | Collapse repeated lines from typeset subs (`--dry-run`, `--lookback-days <n>`) |
| `subminer stats rebuild` / `backfill` | Rebuild or backfill rollup data |
| `subminer doctor` | Dependency + config + socket diagnostics (`--refresh-known-words` refreshes the known-word cache) |
| `subminer settings` | Open the SubMiner settings window |
+5 -1
View File
@@ -95,6 +95,8 @@ subminer texthooker # Texthooker-only mode (-o also opens the brow
subminer stats -b # Start/reuse the background stats daemon
subminer stats -s # Stop the background stats daemon
subminer stats cleanup # Backfill vocabulary metadata, prune stale rows
subminer stats cleanup -d --dry-run # Preview cleanup of repeated typeset subtitle lines
subminer stats cleanup -d --lookback-days 30 # Clean only lines recorded in the last 30 days
subminer stats rebuild # Rebuild rollup data
subminer doctor --refresh-known-words # Refresh the known-word cache
subminer logs -e # Export a sanitized log ZIP and print its path
@@ -107,6 +109,8 @@ subminer app --stop # Stop the background app
subminer --version # Print the launcher's version
```
`stats cleanup` runs one mode per invocation: `-v`/`--vocab` (the default), `-l`/`--lifetime`, or `-d`/`--duplicate-lines`; explicitly selected modes cannot be combined. `--dry-run` and `--lookback-days <days>` apply to `--duplicate-lines` only and are rejected without it; `--lookback-days` must be at least one day, and leaving it off scans all history.
Jellyfin, cross-machine sync, and character-dictionary commands have their own sections: [Jellyfin](/jellyfin-integration), [Sync Between Machines](/launcher-script#sync-between-machines), and [Character Dictionary](/character-dictionary).
</details>
@@ -137,7 +141,7 @@ SubMiner.AppImage --start --log-level debug # Verbose logging without dev mode
SubMiner.AppImage --help # Show all options
```
The remaining flags are internal or scripting-only surfaces: the `--jellyfin-*` family (login, library listing, item playback, cast announce), `--sync-cli` (the app's headless sync entrypoint that `subminer sync` proxies to), `--dictionary-candidates` / `--dictionary-select`, and `--playback-feedback <text>`. Run `SubMiner.AppImage --help` for the complete list. The previous `--open-animetosho` flag is still accepted as a deprecated alias for `--open-tsukihime`.
The remaining flags are internal or scripting-only surfaces: the `--jellyfin-*` family (login, library listing, item playback, cast announce), `--sync-cli` (the app's headless sync entrypoint that `subminer sync` proxies to), the `--stats-cleanup-*` family that `subminer stats cleanup` forwards (`--stats-cleanup-vocab`, `--stats-cleanup-lifetime`, `--stats-cleanup-duplicate-lines`, and its `--stats-cleanup-dry-run` / `--stats-cleanup-lookback-days <days>` modifiers), `--dictionary-candidates` / `--dictionary-select`, and `--playback-feedback <text>`. Run `SubMiner.AppImage --help` for the complete list. The previous `--open-animetosho` flag is still accepted as a deprecated alias for `--open-tsukihime`.
</details>
@@ -64,18 +64,23 @@ External subtitle files only (SRT, VTT, ASS). Embedded subtitle tracks are out o
A cue parser extracts both timing and text content from subtitle files for prefetching.
**Parsed cue structure:**
```typescript
interface SubtitleCue {
startTime: number; // seconds
endTime: number; // seconds
text: string; // raw subtitle text
text: string; // plain text, decoded from the source format
}
```
**Supported formats:**
- SRT/VTT: Regex-based parsing of timing lines + text content between timing blocks.
- ASS: Parse `[Events]` section, extract `Dialogue:` lines, split on the first 9 commas only (ASS v4+ has 10 fields; the last field is Text which can itself contain commas). Strip ASS override tags (`{\...}`) from the text before storing.
ASS text fields contain inline override tags like `{\b1}`, `{\an8}`, `{\fad(200,300)}`. The cue parser strips these during extraction so the tokenizer receives clean text.
- ASS: Parse `[Events]` section, extract `Dialogue:` lines, read the field order from the `Format:` row, and take everything after the Text field index as the text (Text can itself contain commas).
**ASS decoding.** The parser is where ASS text is decoded, once, via `assToPlainText()` in `src/core/services/ass-text.ts`. That decoder mirrors mpv's `ass_to_plaintext` so a cue read from a file reads identically to the same line arriving live on `sub-text`: `{...}` override blocks are markup, `\pN … \p0` vector drawing runs are dropped rather than shown as text, `\N`/`\n`/`\h` are the only escapes (`\{`, `\}` and `\\` are not), and an unclosed `{` is rendered verbatim. Every layer downstream — renderer, timing tracker, tokenizer, tokenization cache keys — receives plain text and uses `normalizePlainSubtitleText()` for whitespace only, so nothing decodes the same string twice and one authored line always maps to one cache key.
**Duplicate collapsing.** Typeset scripts emit one `Dialogue:` event per animation frame, plus layered copies of the same line. The parser collapses identical text over an identical span unconditionally, and collapses contiguous same-text runs of at least three events when the run looks like an animation. For ASS that means shared style and actor plus authoring evidence: a temporal tag (`\t`, `\move`, `\k`/`\kf`/`\ko`/`\K`, or anything wrapped in `\t(...)`), an animated `Effect` column (`Karaoke`, `Banner`, `Scroll`), or override values that change across the run. Static tags shared by every event (`\pos`, an identical `\clip`) are not evidence. SRT/VTT carry no such metadata, so there collapsing needs at least five contiguous events all under 0.1s — the frame timing left behind by ASS-to-SRT conversion. The parser keeps this authoring metadata (style, actor, layer, `Effect`, parsed override commands, source order) private; `parseSubtitleCues()` returns only `SubtitleCue`.
#### Prefetch Service Lifecycle
@@ -153,6 +158,7 @@ tokens (already have frequencyRank values from parser-level applyFrequencyRanks)
### Dependency Analysis
All annotations either depend on MeCab POS data or benefit from running after it:
- **Known word marking:** Needs base tokens (surface/headword). No POS dependency, but no reason to run separately.
- **Frequency filtering:** Uses `pos1Exclusions` and `pos2Exclusions` to clear frequency ranks on excluded tokens (particles, noise). Depends on MeCab POS data.
- **JLPT marking:** Uses `shouldIgnoreJlptForMecabPos1` to filter. Depends on MeCab POS data.
@@ -169,18 +175,14 @@ function annotateTokens(tokens, deps, options): MergedToken[] {
// Single pass: known word + frequency filtering + JLPT computed together
const annotated = tokens.map((token) => {
const isKnown = nPlusOneEnabled
? token.isKnown || computeIsKnown(token, deps)
: false;
const isKnown = nPlusOneEnabled ? token.isKnown || computeIsKnown(token, deps) : false;
// Filter frequency rank using POS exclusions (rank values already set at parser level)
const frequencyRank = frequencyEnabled
? filterFrequencyRank(token, pos1Exclusions, pos2Exclusions)
: undefined;
const jlptLevel = jlptEnabled
? computeJlptLevel(token, deps.getJlptLevel)
: undefined;
const jlptLevel = jlptEnabled ? computeJlptLevel(token, deps.getJlptLevel) : undefined;
return { ...token, isKnown, frequencyRank, jlptLevel };
});
@@ -221,6 +223,7 @@ Replace `document.createElement('span')` calls in the renderer with `templateSpa
### Current Behavior
In `renderWithTokens` (`subtitle-render.ts`), each render cycle:
1. Clears DOM with `innerHTML = ''`
2. Creates a `DocumentFragment`
3. Calls `document.createElement('span')` for each token (~10-15 per subtitle)
@@ -257,7 +260,7 @@ Full recycling (collecting old nodes, clearing attributes, reusing them) require
## Combined Impact Summary
| Scenario | Before | After | Improvement |
|----------|--------|-------|-------------|
| --------------------------------- | ---------- | ---------- | ----------- |
| Normal playback (prefetch-warmed) | ~200-320ms | ~30-50ms | ~80-85% |
| Cache hit (repeated subtitle) | ~72ms | ~55-65ms | ~10-20% |
| Cache miss (immediate seek) | ~200-320ms | ~150-260ms | ~20-25% |
@@ -267,16 +270,19 @@ Full recycling (collecting old nodes, clearing attributes, reusing them) require
## Files Summary
### New Files
- `src/core/services/subtitle-prefetch.ts`
- `src/core/services/subtitle-cue-parser.ts`
### Modified Files
- `src/core/services/subtitle-processing-controller.ts` (expose `preCacheTokenization`)
- `src/core/services/tokenizer/annotation-stage.ts` (batched single-pass)
- `src/renderer/subtitle-render.ts` (template cloneNode)
- `src/main.ts` (wire up prefetch service)
### Test Files
- New tests for subtitle cue parser (SRT, VTT, ASS formats)
- New tests for subtitle prefetch service (priority window, seek, pause/resume)
- Updated tests for annotation stage (same behavior, new implementation)
+1 -1
View File
@@ -24,7 +24,7 @@ Read when: you need to find the owner module for a behavior or test surface
- Subtitle/token pipeline: `src/core/services/subtitle-*.ts`, `src/core/services/tokenizer*`, `src/core/services/tokenizer/`, `src/subsync/`
- Anki workflow: `src/anki-integration/`, `src/core/services/anki-jimaku*.ts`
- Immersion tracking: `src/core/services/immersion-tracker/`
Includes stats storage/query schema such as `imm_videos`, `imm_media_art`, and `imm_youtube_videos` for per-video and YouTube-specific library metadata. Library-entry identity aliases and merge recommendations are persisted alongside this schema; the stats HTTP and SPA layers only expose and present those domain decisions.
Includes stats storage/query schema such as `imm_videos`, `imm_media_art`, and `imm_youtube_videos` for per-video and YouTube-specific library metadata.
- AniList tracking + character dictionary: `src/core/services/anilist/`, `src/main/runtime/composers/anilist-*`, `src/main/character-dictionary-runtime.ts`, `src/main/character-dictionary-runtime/`
- Jellyfin integration: `src/core/services/jellyfin*.ts`, `src/main/runtime/composers/jellyfin-*`
- Window trackers: `src/window-trackers/`
+9
View File
@@ -157,6 +157,15 @@ export async function runStatsCommand(
if (args.statsCleanupLifetime) {
forwarded.push('--stats-cleanup-lifetime');
}
if (args.statsCleanupDuplicateLines) {
forwarded.push('--stats-cleanup-duplicate-lines');
}
if (args.statsCleanupDryRun) {
forwarded.push('--stats-cleanup-dry-run');
}
if (args.statsCleanupLookbackDays) {
forwarded.push('--stats-cleanup-lookback-days', String(args.statsCleanupLookbackDays));
}
if (shouldForwardLogLevel(args.logLevel)) {
forwarded.push('--log-level', args.logLevel);
}
+12
View File
@@ -134,6 +134,9 @@ test('applyInvocationsToArgs maps config and jellyfin invocation state', () => {
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
statsCleanupLookbackDays: null,
statsLogLevel: null,
syncTriggered: false,
syncCliTokens: [],
@@ -185,6 +188,9 @@ test('applyInvocationsToArgs maps settings invocation to settings window', () =>
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
statsCleanupLookbackDays: null,
statsLogLevel: null,
syncTriggered: false,
syncCliTokens: [],
@@ -229,6 +235,9 @@ test('applyInvocationsToArgs fails when config invocation has no action', () =>
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
statsCleanupLookbackDays: null,
statsLogLevel: null,
syncTriggered: false,
syncCliTokens: [],
@@ -271,6 +280,9 @@ test('applyInvocationsToArgs maps texthooker browser-open request', () => {
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
statsCleanupLookbackDays: null,
statsLogLevel: null,
syncTriggered: false,
syncCliTokens: [],
+7
View File
@@ -162,6 +162,8 @@ export function createDefaultArgs(
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
doctor: false,
doctorRefreshKnownWords: false,
logsExport: false,
@@ -258,6 +260,11 @@ export function applyInvocationsToArgs(parsed: Args, invocations: CliInvocations
if (invocations.statsCleanup) parsed.statsCleanup = true;
if (invocations.statsCleanupVocab) parsed.statsCleanupVocab = true;
if (invocations.statsCleanupLifetime) parsed.statsCleanupLifetime = true;
if (invocations.statsCleanupDuplicateLines) parsed.statsCleanupDuplicateLines = true;
if (invocations.statsCleanupDryRun) parsed.statsCleanupDryRun = true;
if (invocations.statsCleanupLookbackDays !== null) {
parsed.statsCleanupLookbackDays = invocations.statsCleanupLookbackDays;
}
if (invocations.dictionaryTarget) {
parsed.dictionaryTarget = parseDictionaryTarget(invocations.dictionaryTarget);
} else if (
+47 -3
View File
@@ -37,6 +37,9 @@ export interface CliInvocations {
statsCleanup: boolean;
statsCleanupVocab: boolean;
statsCleanupLifetime: boolean;
statsCleanupDuplicateLines: boolean;
statsCleanupDryRun: boolean;
statsCleanupLookbackDays: number | null;
statsLogLevel: string | null;
syncTriggered: boolean;
syncCliTokens: string[];
@@ -53,6 +56,16 @@ export interface CliInvocations {
texthookerOpenBrowser: boolean;
}
/** `--lookback-days` narrows the duplicate-line cleanup; fractions are floored. */
function parseStatsLookbackDays(value: unknown): number | null {
if (typeof value !== 'string' && typeof value !== 'number') return null;
const days = Number(value);
if (!Number.isFinite(days) || days < 1) {
throw new Error('Stats --lookback-days must be at least one day.');
}
return Math.floor(days);
}
function applyRootOptions(program: Command): void {
program
.option(
@@ -169,6 +182,9 @@ export function parseCliPrograms(
let statsCleanup = false;
let statsCleanupVocab = false;
let statsCleanupLifetime = false;
let statsCleanupDuplicateLines = false;
let statsCleanupDryRun = false;
let statsCleanupLookbackDays: number | null = null;
let statsLogLevel: string | null = null;
let syncTriggered = false;
let syncCliTokens: string[] = [];
@@ -269,6 +285,9 @@ export function parseCliPrograms(
.option('-s, --stop', 'Stop the background stats server')
.option('-v, --vocab', 'Clean vocabulary rows in the stats database')
.option('-l, --lifetime', 'Rebuild lifetime summary rows from retained data')
.option('-d, --duplicate-lines', 'Collapse repeated subtitle lines from typeset animations')
.option('--dry-run', 'Report what a cleanup would remove without changing anything')
.option('--lookback-days <days>', 'Only clean lines recorded in the last N days')
.option('--log-level <level>', 'Log level')
.action((action: string | undefined, options: Record<string, unknown>) => {
statsTriggered = true;
@@ -289,13 +308,35 @@ export function parseCliPrograms(
if (normalizedAction && (statsBackground || statsStop)) {
throw new Error('Stats background and stop flags cannot be combined with stats actions.');
}
if (normalizedAction !== 'cleanup' && (options.vocab === true || options.lifetime === true)) {
throw new Error('Stats --vocab and --lifetime flags require the cleanup action.');
if (
normalizedAction !== 'cleanup' &&
(options.vocab === true || options.lifetime === true || options.duplicateLines === true)
) {
throw new Error(
'Stats --vocab, --lifetime and --duplicate-lines flags require the cleanup action.',
);
}
if (
options.duplicateLines !== true &&
(options.dryRun === true || options.lookbackDays !== undefined)
) {
throw new Error('Stats --dry-run and --lookback-days require --duplicate-lines.');
}
if (normalizedAction === 'cleanup') {
statsCleanup = true;
statsCleanupLifetime = options.lifetime === true;
statsCleanupVocab = statsCleanupLifetime ? false : options.vocab !== false;
statsCleanupDuplicateLines = options.duplicateLines === true;
const explicitModeCount = [options.vocab, options.lifetime, options.duplicateLines].filter(
(value) => value === true,
).length;
if (explicitModeCount > 1) {
throw new Error('Stats cleanup runs one mode at a time.');
}
// Vocabulary cleanup stays the default so `stats cleanup` keeps its old meaning.
statsCleanupVocab =
statsCleanupLifetime || statsCleanupDuplicateLines ? false : options.vocab !== false;
statsCleanupDryRun = options.dryRun === true;
statsCleanupLookbackDays = parseStatsLookbackDays(options.lookbackDays);
} else if (normalizedAction === 'rebuild' || normalizedAction === 'backfill') {
statsCleanup = true;
statsCleanupLifetime = true;
@@ -483,6 +524,9 @@ export function parseCliPrograms(
statsCleanup,
statsCleanupVocab,
statsCleanupLifetime,
statsCleanupDuplicateLines,
statsCleanupDryRun,
statsCleanupLookbackDays,
statsLogLevel,
syncTriggered,
syncCliTokens,
+73 -1
View File
@@ -232,6 +232,75 @@ test('parseArgs maps lifetime stats cleanup flag', () => {
assert.equal(parsed.statsCleanupLifetime, true);
});
test('parseArgs maps duplicate-line stats cleanup flags', () => {
const parsed = parseArgs(
['stats', 'cleanup', '--duplicate-lines', '--dry-run', '--lookback-days', '30'],
'subminer',
{},
);
assert.equal(parsed.statsCleanup, true);
assert.equal(parsed.statsCleanupVocab, false);
assert.equal(parsed.statsCleanupDuplicateLines, true);
assert.equal(parsed.statsCleanupDryRun, true);
assert.equal(parsed.statsCleanupLookbackDays, 30);
const fractional = parseArgs(
['stats', 'cleanup', '--duplicate-lines', '--lookback-days', '1.5'],
'subminer',
{},
);
assert.equal(fractional.statsCleanupLookbackDays, 1);
});
test('parseArgs rejects duplicate-line flags without the duplicate-lines mode', () => {
const error = withProcessExitIntercept(() => {
parseArgs(['stats', 'cleanup', '--dry-run'], 'subminer', {});
});
assert.equal(error.code, 1);
assert.match(error.stderr, /--dry-run and --lookback-days require --duplicate-lines/);
});
test('parseArgs rejects an empty lookback value outside duplicate-line cleanup', () => {
const error = withProcessExitIntercept(() => {
parseArgs(['stats', '--lookback-days', ''], 'subminer', {});
});
assert.equal(error.code, 1);
assert.match(error.stderr, /--dry-run and --lookback-days require --duplicate-lines/);
});
test('parseArgs rejects combining explicit cleanup modes', () => {
for (const modes of [
['--lifetime', '--duplicate-lines'],
['--vocab', '--duplicate-lines'],
['--vocab', '--lifetime'],
]) {
const error = withProcessExitIntercept(() => {
parseArgs(['stats', 'cleanup', ...modes], 'subminer', {});
});
assert.equal(error.code, 1);
assert.match(error.stderr, /Stats cleanup runs one mode at a time/);
}
});
test('parseArgs rejects unusable lookback windows', () => {
for (const value of ['0', '0.5', '-5', 'soon']) {
const error = withProcessExitIntercept(() => {
parseArgs(
['stats', 'cleanup', '--duplicate-lines', '--lookback-days', value],
'subminer',
{},
);
});
assert.equal(error.code, 1);
assert.match(error.stderr, /--lookback-days must be at least one day/);
}
});
test('parseArgs rejects cleanup-only stats flags without cleanup action', () => {
const error = withProcessExitIntercept(() => {
parseArgs(['stats', '--vocab'], 'subminer', {});
@@ -239,7 +308,10 @@ test('parseArgs rejects cleanup-only stats flags without cleanup action', () =>
assert.equal(error.code, 1);
assert.match(error.message, /exit:1/);
assert.match(error.stderr, /Stats --vocab and --lifetime flags require the cleanup action/);
assert.match(
error.stderr,
/Stats --vocab, --lifetime and --duplicate-lines flags require the cleanup action/,
);
});
test('parseArgs maps stats rebuild action to cleanup lifetime mode', () => {
+3
View File
@@ -142,6 +142,9 @@ export interface Args {
statsCleanup?: boolean;
statsCleanupVocab?: boolean;
statsCleanupLifetime?: boolean;
statsCleanupDuplicateLines?: boolean;
statsCleanupDryRun?: boolean;
statsCleanupLookbackDays?: number;
dictionaryTarget?: string;
doctor: boolean;
doctorRefreshKnownWords: boolean;
+24
View File
@@ -399,6 +399,30 @@ test('hasExplicitCommand and shouldStartApp preserve command intent', () => {
assert.equal(statsLifetimeRebuild.statsCleanupLifetime, true);
assert.equal(statsLifetimeRebuild.statsCleanupVocab, false);
assert.throws(
() =>
parseArgs([
'--stats',
'--stats-cleanup',
'--stats-cleanup-duplicate-lines',
'--stats-cleanup-lookback-days',
'0.5',
]),
/at least one day/,
);
assert.equal(
parseArgs([
'--stats',
'--stats-cleanup',
'--stats-cleanup-duplicate-lines',
'--stats-cleanup-lookback-days',
'1.5',
]).statsCleanupLookbackDays,
1,
);
assert.equal(parseArgs(['--stats-cleanup-lookback-days=30']).statsCleanupLookbackDays, 30);
assert.throws(() => parseArgs(['--stats-cleanup-lookback-days=30=oops']), /at least one day/);
const jellyfinLibraries = parseArgs(['--jellyfin-libraries']);
assert.equal(jellyfinLibraries.jellyfinLibraries, true);
assert.equal(hasExplicitCommand(jellyfinLibraries), true);
+22 -1
View File
@@ -64,6 +64,9 @@ export interface CliArgs {
statsCleanup?: boolean;
statsCleanupVocab?: boolean;
statsCleanupLifetime?: boolean;
statsCleanupDuplicateLines?: boolean;
statsCleanupDryRun?: boolean;
statsCleanupLookbackDays?: number;
statsResponsePath?: string;
jellyfin: boolean;
jellyfinLogin: boolean;
@@ -109,6 +112,14 @@ export interface CliArgs {
export type CliCommandSource = 'initial' | 'second-instance';
function parseStatsCleanupLookbackDays(value: string | undefined): number {
const days = Number(value);
if (!Number.isFinite(days) || days < 1) {
throw new Error('Stats --lookback-days must be at least one day.');
}
return Math.floor(days);
}
export function parseArgs(argv: string[]): CliArgs {
const args: CliArgs = {
background: false,
@@ -167,6 +178,8 @@ export function parseArgs(argv: string[]): CliArgs {
statsCleanup: false,
statsCleanupVocab: false,
statsCleanupLifetime: false,
statsCleanupDuplicateLines: false,
statsCleanupDryRun: false,
jellyfin: false,
jellyfinLogin: false,
jellyfinLogout: false,
@@ -368,7 +381,15 @@ export function parseArgs(argv: string[]): CliArgs {
} else if (arg === '--stats-cleanup') args.statsCleanup = true;
else if (arg === '--stats-cleanup-vocab') args.statsCleanupVocab = true;
else if (arg === '--stats-cleanup-lifetime') args.statsCleanupLifetime = true;
else if (arg.startsWith('--stats-response-path=')) {
else if (arg === '--stats-cleanup-duplicate-lines') args.statsCleanupDuplicateLines = true;
else if (arg === '--stats-cleanup-dry-run') args.statsCleanupDryRun = true;
else if (arg.startsWith('--stats-cleanup-lookback-days=')) {
args.statsCleanupLookbackDays = parseStatsCleanupLookbackDays(
arg.slice('--stats-cleanup-lookback-days='.length),
);
} else if (arg === '--stats-cleanup-lookback-days') {
args.statsCleanupLookbackDays = parseStatsCleanupLookbackDays(readValue(argv[i + 1]));
} else if (arg.startsWith('--stats-response-path=')) {
const value = arg.split('=', 2)[1];
if (value) args.statsResponsePath = value;
} else if (arg === '--stats-response-path') {
@@ -1,256 +0,0 @@
import test from 'node:test';
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import type { DatabaseSync } from '../immersion-tracker/sqlite';
type ImmersionTrackerService = import('../immersion-tracker-service').ImmersionTrackerService;
type ImmersionTrackerServiceCtor =
typeof import('../immersion-tracker-service').ImmersionTrackerService;
let trackerCtor: ImmersionTrackerServiceCtor | null = null;
async function loadTrackerCtor(): Promise<ImmersionTrackerServiceCtor> {
if (trackerCtor) return trackerCtor;
const mod = await import('../immersion-tracker-service');
trackerCtor = mod.ImmersionTrackerService;
return trackerCtor;
}
function makeDbPath(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-write-queue-test-'));
return path.join(dir, 'immersion.sqlite');
}
function cleanupDbPath(dbPath: string): void {
const dir = path.dirname(dbPath);
if (!fs.existsSync(dir)) return;
fs.rmSync(dir, { recursive: true, force: true });
}
interface TrackerInternals {
db: DatabaseSync;
queue: unknown[];
recordWrite: (write: Record<string, unknown>) => void;
mergeAnime: (targetAnimeId: number, sourceAnimeIds: number[]) => Promise<unknown>;
moveVideoToAnime: (videoId: number, targetAnimeId: number) => Promise<unknown>;
rebuildLifetimeSummaries: () => Promise<unknown>;
flushNow: () => void;
}
test('mergeAnime fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.mergeAnime(1, [2]), /queue did not drain/i);
assert.deepEqual(
internals.db
.prepare('SELECT anime_id AS animeId FROM imm_anime ORDER BY anime_id')
.all()
.map((row) => (row as { animeId: number }).animeId),
[1, 2],
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('moveVideoToAnime fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.moveVideoToAnime(2, 1), /queue did not drain/i);
assert.equal(
(
internals.db
.prepare('SELECT anime_id AS animeId FROM imm_videos WHERE video_id = 2')
.get() as {
animeId: number;
}
).animeId,
2,
);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('rebuildLifetimeSummaries fails closed when queued writes cannot drain', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 1);
internals.flushNow = () => {};
await assert.rejects(internals.rebuildLifetimeSummaries(), /queue did not drain/i);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
function seedTwoEntries(db: DatabaseSync): void {
db.exec(`
INSERT INTO imm_anime (anime_id, normalized_title_key, canonical_title, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'show', 'Show', 1000, 1000), (2, 'show season 1', 'Show Season 1', 1000, 1000);
INSERT INTO imm_videos (video_id, video_key, canonical_title, anime_id, source_type, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'local:/tmp/a.mkv', 'A', 1, 1, 0, 1440000, 1000, 1000),
(2, 'local:/tmp/b.mkv', 'B', 2, 1, 0, 1440000, 1000, 1000);
INSERT INTO imm_sessions (session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, active_watched_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'drain-session', 2, '1000', '2000', 2, 1000, 1000, 2000);
`);
}
function queueSubtitleLines(tracker: TrackerInternals, count: number): void {
for (let index = 0; index < count; index += 1) {
tracker.recordWrite({
kind: 'subtitleLine',
sessionId: 1,
videoId: 2,
lineIndex: index,
segmentStartMs: index * 1000,
segmentEndMs: index * 1000 + 900,
text: `line ${index}`,
wordOccurrences: [],
kanjiOccurrences: [],
firstSeen: 1000,
lastSeen: 2000,
});
}
}
/**
* Queued last so it sits past the first batch. Lifetime `total_lines_seen`
* reads this counter, not a COUNT over imm_subtitle_lines, so the rebuilt
* summary only reflects the session once the queue is drained all the way.
*/
function queueTelemetry(tracker: TrackerInternals, linesSeen: number): void {
tracker.recordWrite({
kind: 'telemetry',
sessionId: 1,
sampleMs: 3000,
lastMediaMs: 3000,
totalWatchedMs: 4000,
activeWatchedMs: 3500,
linesSeen,
tokensSeen: linesSeen * 5,
cardsMined: 2,
lookupCount: 0,
lookupHits: 0,
yomitanLookupCount: 0,
pauseCount: 0,
pauseMs: 0,
seekForwardCount: 0,
seekBackwardCount: 0,
mediaBufferEvents: 0,
});
}
function lifetimeForAnime(
db: DatabaseSync,
animeId: number,
): { linesSeen: number; activeMs: number; cards: number } | null {
const row = db
.prepare(
`SELECT total_lines_seen AS linesSeen, total_active_ms AS activeMs, total_cards AS cards
FROM imm_lifetime_anime WHERE anime_id = ?`,
)
.get(animeId) as { linesSeen: number; activeMs: number; cards: number } | undefined;
return row
? { linesSeen: Number(row.linesSeen), activeMs: Number(row.activeMs), cards: Number(row.cards) }
: null;
}
function countLinesForAnime(db: DatabaseSync, animeId: number): number {
const row = db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id = ?')
.get(animeId) as { total: number };
return Number(row.total);
}
/**
* Both entry points rebuild the lifetime summaries, which recompute from the
* database. A single flushNow() only writes one batch off the front of the
* queue, so anything past `batchSize` would still be unwritten when the rebuild
* reads.
*/
test('mergeAnime drains a queue larger than one batch before rebuilding summaries', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 8);
queueTelemetry(internals, 8);
assert.ok(internals.queue.length > 2, 'expected more queued writes than one batch');
await internals.mergeAnime(1, [2]);
assert.equal(internals.queue.length, 0);
assert.equal(countLinesForAnime(internals.db, 1), 8);
// The surviving entry's summary was rebuilt from the fully drained queue.
const lifetime = lifetimeForAnime(internals.db, 1);
assert.equal(lifetime?.linesSeen, 8);
assert.equal(lifetime?.activeMs, 3500);
assert.equal(lifetime?.cards, 2);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
test('moveVideoToAnime drains a queue larger than one batch before rebuilding summaries', async () => {
const dbPath = makeDbPath();
let tracker: ImmersionTrackerService | null = null;
try {
const Ctor = await loadTrackerCtor();
tracker = new Ctor({ dbPath, policy: { batchSize: 2 } });
const internals = tracker as unknown as TrackerInternals;
seedTwoEntries(internals.db);
queueSubtitleLines(internals, 8);
queueTelemetry(internals, 8);
await internals.moveVideoToAnime(2, 1);
assert.equal(internals.queue.length, 0);
assert.equal(countLinesForAnime(internals.db, 1), 8);
const lifetime = lifetimeForAnime(internals.db, 1);
assert.equal(lifetime?.linesSeen, 8);
assert.equal(lifetime?.activeMs, 3500);
assert.equal(lifetime?.cards, 2);
} finally {
tracker?.destroy();
cleanupDbPath(dbPath);
}
});
+174 -191
View File
@@ -1032,6 +1032,180 @@ describe('stats server API routes', () => {
]);
});
it('POST /api/stats/maintenance/duplicate-lines forwards the window and dry-run flag', async () => {
let seenOptions: unknown = null;
const summary = {
dryRun: true,
lookbackDays: 30,
scannedLines: 900,
burstGroups: 2,
removedLines: 180,
removedWordOccurrences: 540,
removedKanjiOccurrences: 120,
samples: [],
};
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async (options: unknown) => {
seenOptions = options;
return summary;
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ dryRun: true, lookbackDays: 30 }),
});
assert.equal(res.status, 200);
assert.deepEqual(await res.json(), summary);
assert.deepEqual(seenOptions, { dryRun: true, lookbackDays: 30 });
});
it('POST /api/stats/maintenance/duplicate-lines rejects cross-origin simple requests', async () => {
let cleanupCalls = 0;
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async () => {
cleanupCalls += 1;
throw new Error('cleanup must not run');
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: {
'Content-Type': 'text/plain',
Origin: 'https://attacker.example',
},
body: JSON.stringify({ dryRun: false, lookbackDays: null }),
});
assert.equal(res.status, 415);
assert.equal(cleanupCalls, 0);
});
it('POST /api/stats/maintenance/duplicate-lines rejects a window shorter than a day', async () => {
let cleanupCalls = 0;
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async () => {
cleanupCalls += 1;
return {
dryRun: true,
lookbackDays: null,
scannedLines: 0,
burstGroups: 0,
removedLines: 0,
removedWordOccurrences: 0,
removedKanjiOccurrences: 0,
samples: [],
};
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ dryRun: true, lookbackDays: 0.5 }),
});
assert.equal(res.status, 400);
assert.equal(cleanupCalls, 0);
});
it('POST /api/stats/maintenance/duplicate-lines floors a fractional multi-day window', async () => {
let seenOptions: unknown = null;
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async (options: unknown) => {
seenOptions = options;
return {
dryRun: true,
lookbackDays: 1,
scannedLines: 0,
burstGroups: 0,
removedLines: 0,
removedWordOccurrences: 0,
removedKanjiOccurrences: 0,
samples: [],
};
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ dryRun: true, lookbackDays: 1.5 }),
});
assert.equal(res.status, 200);
assert.deepEqual(seenOptions, { dryRun: true, lookbackDays: 1 });
});
it('POST /api/stats/maintenance/duplicate-lines accepts an explicit empty object for all history', async () => {
let seenOptions: unknown = null;
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async (options: unknown) => {
seenOptions = options;
return {
dryRun: false,
lookbackDays: null,
scannedLines: 0,
burstGroups: 0,
removedLines: 0,
removedWordOccurrences: 0,
removedKanjiOccurrences: 0,
samples: [],
};
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: '{}',
});
assert.equal(res.status, 200);
assert.deepEqual(seenOptions, { dryRun: false, lookbackDays: null });
});
for (const malformed of [
{ name: 'a missing body', body: undefined },
{ name: 'malformed JSON', body: '{' },
{ name: 'JSON null', body: 'null' },
{ name: 'a JSON array', body: '[]' },
]) {
it(`POST /api/stats/maintenance/duplicate-lines rejects ${malformed.name}`, async () => {
let cleanupCalls = 0;
const app = createStatsApp(
createMockTracker({
cleanupDuplicateSubtitleLines: async () => {
cleanupCalls += 1;
throw new Error('cleanup must not run');
},
}),
);
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: malformed.body,
});
assert.equal(res.status, 400);
assert.equal(cleanupCalls, 0);
});
}
it('PUT /api/stats/excluded-words rejects malformed rows', async () => {
const app = createStatsApp(createMockTracker());
@@ -1053,55 +1227,6 @@ describe('stats server API routes', () => {
assert.equal(body[0].canonicalTitle, 'Little Witch Academia');
});
it('GET /api/stats/anime/merge-recommendations returns pending duplicate pairs', async () => {
const app = createStatsApp(
createMockTracker({
getAnimeMergeRecommendations: async () => [{ recommendationId: 4, animeIds: [1, 2] }],
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations');
assert.equal(res.status, 200);
assert.deepEqual(await res.json(), {
recommendations: [{ recommendationId: 4, animeIds: [1, 2] }],
});
});
it('DELETE /api/stats/anime/merge-recommendations/:id dismisses a pending pair', async () => {
let dismissedId: number | null = null;
const app = createStatsApp(
createMockTracker({
dismissAnimeMergeRecommendation: async (recommendationId: number) => {
dismissedId = recommendationId;
return true;
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations/4', {
method: 'DELETE',
});
assert.equal(res.status, 200);
assert.equal(dismissedId, 4);
assert.deepEqual(await res.json(), { ok: true });
});
it('DELETE /api/stats/anime/merge-recommendations/:id reports missing recommendations', async () => {
const app = createStatsApp(
createMockTracker({
dismissAnimeMergeRecommendation: async () => false,
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/merge-recommendations/99', {
method: 'DELETE',
});
assert.equal(res.status, 404);
});
it('GET /api/stats/anime/:animeId returns anime detail with episodes', async () => {
const app = createStatsApp(createMockTracker());
const res = await app.request('/api/stats/anime/1');
@@ -3073,148 +3198,6 @@ Aligned English subtitle
assert.equal(deleteCalls, 0);
});
it('POST /api/stats/anime/:animeId/merge folds the given entries into the target', async () => {
let merged: { targetAnimeId: number; sourceAnimeIds: number[] } | null = null;
const app = createStatsApp(
createMockTracker({
mergeAnime: async (targetAnimeId: number, sourceAnimeIds: number[]) => {
merged = { targetAnimeId, sourceAnimeIds };
return {
survivingAnimeId: targetAnimeId,
mergedAnimeIds: sourceAnimeIds,
movedVideos: 3,
};
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
// The target repeated in the sources must not delete the entry we keep.
body: '{"sourceAnimeIds":[8,9,8,7]}',
});
assert.equal(res.status, 200);
assert.deepEqual(merged, { targetAnimeId: 7, sourceAnimeIds: [8, 9] });
assert.deepEqual(await res.json(), {
ok: true,
animeId: 7,
mergedAnimeIds: [8, 9],
movedVideos: 3,
});
});
it('POST /api/stats/anime/:animeId/merge rejects an empty or malformed source list', async () => {
let mergeCalls = 0;
const app = createStatsApp(
createMockTracker({
mergeAnime: async () => {
mergeCalls += 1;
return { survivingAnimeId: 7, mergedAnimeIds: [], movedVideos: 0 };
},
} as Partial<ImmersionTrackerService>),
);
for (const body of [
'{"sourceAnimeIds":[]}',
'{"sourceAnimeIds":[7]}',
'{"sourceAnimeIds":0}',
]) {
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body,
});
assert.equal(res.status, 400);
}
assert.equal(mergeCalls, 0);
});
it('PATCH /api/stats/media/:videoId/anime moves the episode to another entry', async () => {
let moved: { videoId: number; animeId: number } | null = null;
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async (videoId: number, animeId: number) => {
moved = { videoId, animeId };
return { targetAnimeId: animeId, previousAnimeId: 4, removedPreviousAnime: true };
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":7}',
});
assert.equal(res.status, 200);
assert.deepEqual(moved, { videoId: 12, animeId: 7 });
assert.deepEqual(await res.json(), {
ok: true,
animeId: 7,
previousAnimeId: 4,
removedPreviousAnime: true,
});
});
it('POST /api/stats/anime/:animeId/merge reports a merge that folded nothing as 404', async () => {
const app = createStatsApp(
createMockTracker({
mergeAnime: async (targetAnimeId: number) => ({
survivingAnimeId: targetAnimeId,
mergedAnimeIds: [],
movedVideos: 0,
}),
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/anime/7/merge', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: '{"sourceAnimeIds":[8]}',
});
assert.equal(res.status, 404);
});
it('PATCH /api/stats/media/:videoId/anime reports an unknown target as 404', async () => {
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async () => {
throw new Error('Unknown episode or target library entry');
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":99}',
});
assert.equal(res.status, 404);
});
it('PATCH /api/stats/media/:videoId/anime does not disguise storage failures as 404', async () => {
const app = createStatsApp(
createMockTracker({
moveVideoToAnime: async () => {
throw new Error('database is locked');
},
} as Partial<ImmersionTrackerService>),
);
const res = await app.request('/api/stats/media/12/anime', {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: '{"animeId":7}',
});
assert.notEqual(res.status, 404);
assert.equal(res.status >= 500, true);
});
it('POST /api/stats/anki/browse returns 400 for missing noteId', async () => {
const app = createStatsApp(createMockTracker());
const res = await app.request('/api/stats/anki/browse', { method: 'POST' });
@@ -327,7 +327,6 @@ export function createCoverArtFetcher(
titleEnglish: selected.title?.english ?? null,
titleNative: selected.title?.native ?? null,
episodesTotal: selected.episodes ?? null,
exactTitleMatch: resolution?.exactTitleMatch ?? false,
});
logger.info(
@@ -156,49 +156,6 @@ test('season 1 resolves to the anchor without relation lookups', async () => {
assert.deepEqual(relationLookups, []);
});
test('season 2 preserves the anchor exact-title evidence through a sequel resolution', async () => {
const { execute } = createExecutor(OREGAIRU_SEARCH, OREGAIRU_RELATIONS);
const result = await resolveAnilistSeasonMedia(
{ title: 'My Teen Romantic Comedy SNAFU', season: 2, episode: 1 },
{ execute },
);
assert.equal(result?.id, 20698);
assert.equal(result?.via, 'sequel-chain');
assert.equal(result?.exactTitleMatch, true);
});
test('reports an exact normalized synonym match as strong evidence', async () => {
const { execute } = createExecutor([
{
id: 1,
episodes: 12,
format: 'TV',
title: { english: 'Hitori Gotoh Story' },
synonyms: ['BOCCHI THE ROCK'],
},
]);
const result = await resolveAnilistSeasonMedia({ title: 'Bocchi the Rock!' }, { execute });
assert.equal(result?.exactTitleMatch, true);
});
test('reports a fuzzy-only search result as weak evidence', async () => {
const { execute } = createExecutor([
{
id: 1,
episodes: 12,
format: 'TV',
title: { english: 'Actual Show' },
},
]);
const result = await resolveAnilistSeasonMedia({ title: 'Unrelated Release' }, { execute });
assert.equal(result?.exactTitleMatch, false);
});
test('strips a season marker already present in the parsed title', async () => {
const { execute, searches } = createExecutor(OREGAIRU_SEARCH, OREGAIRU_RELATIONS);
const result = await resolveAnilistSeasonMedia(
+12 -16
View File
@@ -9,8 +9,6 @@
* reports `seasonResolved: false` so callers can refuse to act instead of guessing.
*/
import { normalizeTitleIdentity } from '../../utils/title-normalization';
export interface AnilistSeasonMediaTitle {
romaji?: string | null;
english?: string | null;
@@ -44,8 +42,6 @@ export interface AnilistSeasonResolution {
seasonResolved: boolean;
requestedSeason: number | null;
via: AnilistSeasonResolutionVia;
/** Exact normalized match against an AniList title or synonym. */
exactTitleMatch: boolean;
}
export interface ResolveAnilistSeasonMediaInput {
@@ -119,6 +115,10 @@ const SEASONAL_FORMAT_PRIORITY = ['TV', 'TV_SHORT', 'ONA'];
const MAX_SEQUEL_HOPS = 12;
function normalizeTitle(value: string): string {
return value.trim().toLowerCase().replace(/\s+/g, ' ');
}
/**
* Drops season markers a release name carries but AniList titles never do,
* so "Some Show Season 3" and "Some Show S3" both search as "Some Show".
@@ -136,7 +136,7 @@ function mediaTitles(media: AnilistSeasonMedia): string[] {
const synonyms = Array.isArray(media.synonyms) ? media.synonyms : [];
return [media.title?.english, media.title?.romaji, media.title?.native, ...synonyms]
.filter((value): value is string => typeof value === 'string' && value.trim().length > 0)
.map((value) => normalizeTitleIdentity(value));
.map((value) => normalizeTitle(value));
}
function displayTitle(media: AnilistSeasonMedia, fallback: string): string {
@@ -176,7 +176,6 @@ function toResolution(
season: number | null,
via: AnilistSeasonResolutionVia,
seasonResolved: boolean,
exactTitleMatch: boolean,
): AnilistSeasonResolution {
return {
id: media.id,
@@ -186,7 +185,6 @@ function toResolution(
seasonResolved,
requestedSeason: season,
via,
exactTitleMatch,
};
}
@@ -211,10 +209,9 @@ export function pickAnchorMedia(
: media;
const pool = episodeFiltered.length > 0 ? episodeFiltered : media;
const targets = [
normalizeTitleIdentity(title),
normalizeTitleIdentity(stripSeasonSuffix(title)),
].filter((value, index, all) => value.length > 0 && all.indexOf(value) === index);
const targets = [normalizeTitle(title), normalizeTitle(stripSeasonSuffix(title))].filter(
(value, index, all) => value.length > 0 && all.indexOf(value) === index,
);
const scored = pool.map((entry, index) => {
const candidateTitles = mediaTitles(entry);
@@ -370,10 +367,9 @@ export async function resolveAnilistSeasonMedia(
episode: season === null || season <= 1 ? input.episode : null,
});
if (!anchor) return null;
const exactTitleMatch = mediaTitles(anchor).includes(normalizeTitleIdentity(searchTitle));
if (season === null || season <= 1) {
return toResolution(anchor, searchTitle, season, 'anchor', true, exactTitleMatch);
return toResolution(anchor, searchTitle, season, 'anchor', true);
}
let chainError: unknown = null;
@@ -387,7 +383,7 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] season ${season} of "${searchTitle}" resolved via sequel chain: ${displayTitle(viaChain, searchTitle)} (${viaChain.id})`,
);
return toResolution(viaChain, searchTitle, season, 'sequel-chain', true, exactTitleMatch);
return toResolution(viaChain, searchTitle, season, 'sequel-chain', true);
}
const viaAirOrder = pickByAirOrder(anchor, season, media);
@@ -395,7 +391,7 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] season ${season} of "${searchTitle}" resolved via air order: ${displayTitle(viaAirOrder, searchTitle)} (${viaAirOrder.id})`,
);
return toResolution(viaAirOrder, searchTitle, season, 'air-order', true, exactTitleMatch);
return toResolution(viaAirOrder, searchTitle, season, 'air-order', true);
}
// The chain failed for transport reasons rather than because the season is absent;
@@ -407,5 +403,5 @@ export async function resolveAnilistSeasonMedia(
deps.logInfo?.(
`[anilist] could not resolve season ${season} of "${searchTitle}"; falling back to ${displayTitle(anchor, searchTitle)} (${anchor.id})`,
);
return toResolution(anchor, searchTitle, season, 'anchor', false, exactTitleMatch);
return toResolution(anchor, searchTitle, season, 'anchor', false);
}
+195
View File
@@ -0,0 +1,195 @@
import { test } from 'node:test';
import assert from 'node:assert/strict';
import {
assOverrideSignature,
assToPlainText,
collectAssOverrideCommands,
extractAssOverrideBlocks,
hasAssTemporalOverride,
isAnimatedAssEffectKind,
isAssTemporalCommand,
normalizePlainSubtitleText,
parseAssEffectField,
} from './ass-text';
test('assToPlainText drops vector drawing runs', () => {
assert.equal(
assToPlainText(
'{\\an5\\pos(730,1042)\\p1\\blur1}m 20 0 b 10 0 0 10 0 20 b 0 31 10 40 20 40 {\\p0}',
),
'',
);
});
test('assToPlainText keeps text around drawing runs on the same event', () => {
assert.equal(
assToPlainText('{\\p1}m 0 0 l 10 10{\\p0}本文{\\p1}m 5 5 l 6 6{\\p0}続き'),
'本文続き',
);
});
test('assToPlainText leaves \\pos alone when no drawing mode is active', () => {
assert.equal(assToPlainText('{\\pos(960,1068)\\bord3}位置指定'), '位置指定');
});
test('assToPlainText does not read \\pos as a drawing tag', () => {
assert.equal(assToPlainText('{\\p1\\pos(1,2)}m 0 0 l 5 5'), '');
});
test('assToPlainText resolves line-break and space escapes', () => {
assert.equal(assToPlainText('一行目\\N二行目'), '一行目\n二行目');
assert.equal(assToPlainText('一行目\\n二行目'), '一行目\n二行目');
assert.equal(assToPlainText('一行目\\N二行目', ' '), '一行目 二行目');
assert.equal(assToPlainText('間\\h隔'), '間 隔');
});
test('assToPlainText matches mpv on brace and backslash sequences', () => {
// mpv has no `\{` / `\}` / `\\` escapes: the backslashes are literal text and the
// braces still open and close an override block.
assert.equal(assToPlainText('\\{注\\}'), '\\');
assert.equal(assToPlainText('\\\\N'), '\\\n');
});
test('assToPlainText renders an unclosed override block verbatim', () => {
// mpv shows the stray brace; guessing where the block ended can eat a whole line.
assert.equal(assToPlainText('本文{\\pos(1,2)'), '本文{\\pos(1,2)');
});
test('assToPlainText is idempotent', () => {
const samples = [
'{\\an5\\p1}m 0 0 l 5 5{\\p0}本文',
'\\{注\\}',
'\\\\N',
'本文{\\pos(1,2)',
'一行目\\N二行目\\h終わり',
];
for (const sample of samples) {
const once = assToPlainText(sample);
assert.equal(assToPlainText(once), once, sample);
}
});
test('assToPlainText normalizes CRLF before converting', () => {
assert.equal(assToPlainText('一行目\r\n二行目'), '一行目\n二行目');
});
test('normalizePlainSubtitleText settles whitespace without decoding ASS', () => {
// A brace reaching this layer is literal text mpv chose to show, not markup.
assert.equal(normalizePlainSubtitleText('本文{\\pos(1,2)'), '本文{\\pos(1,2)');
assert.equal(normalizePlainSubtitleText('一行目\\N二行目'), '一行目\n二行目');
assert.equal(
normalizePlainSubtitleText('一行目\\N二行目', { collapseLineBreaks: true }),
'一行目 二行目',
);
assert.equal(normalizePlainSubtitleText(' 余白 ', { trim: false }), ' 余白 ');
});
test('normalizePlainSubtitleText is idempotent', () => {
for (const sample of ['一行目\\N二行目', '間\\h隔', '本文{\\pos(1,2)', ' 余白 ']) {
const once = normalizePlainSubtitleText(sample);
assert.equal(normalizePlainSubtitleText(once), once, sample);
}
});
test('extractAssOverrideBlocks returns block contents', () => {
assert.deepEqual(extractAssOverrideBlocks('{\\an8}上{\\fad(200,200)}下'), [
'\\an8',
'\\fad(200,200)',
]);
assert.deepEqual(extractAssOverrideBlocks('括弧なし'), []);
});
test('collectAssOverrideCommands captures names and arguments from blocks only', () => {
const commands = collectAssOverrideCommands('{\\pos(1,2)\\1c&HFFFFFF&\\kf30}歌詞');
assert.deepEqual(commands, [
{ name: 'pos', args: '1,2', animated: false },
{ name: '1c', args: '&HFFFFFF&', animated: false },
{ name: 'kf', args: '30', animated: false },
]);
// A `\pos(...)` sitting in visible text is not typesetting markup.
assert.deepEqual(collectAssOverrideCommands('\\pos(730,1042) と書いてある'), []);
});
test('collectAssOverrideCommands marks tags animated by a wrapping \\t', () => {
const commands = collectAssOverrideCommands('{\\clip(0,0,10,10)\\t(0,500,\\frz30)}文字');
assert.deepEqual(
commands.map((command) => [command.name, command.animated]),
[
['clip', false],
['t', false],
['frz', true],
],
);
assert.equal(hasAssTemporalOverride(commands), true);
});
test('collectAssOverrideCommands stops descending into deeply nested \\t tags', () => {
// Nested far past the recursion cap. Uncapped, this recurses once per level, and a
// pathological line (real files reach one or two levels) overflows the stack.
const nesting = 32;
const block = `{${'\\t(0,500,'.repeat(nesting)}\\frz30${')'.repeat(nesting)}}文字`;
const commands = collectAssOverrideCommands(block);
// The outer `\t` plus one per allowed recursion level, and nothing from below the cap.
assert.equal(commands.length, 9);
assert.deepEqual(new Set(commands.map((command) => command.name)), new Set(['t']));
assert.equal(hasAssTemporalOverride(commands), true);
});
test('hasAssTemporalOverride ignores static placement and shape tags', () => {
assert.equal(
hasAssTemporalOverride(collectAssOverrideCommands('{\\pos(1,2)\\clip(m 1 1)\\blur2}文字')),
false,
);
assert.equal(hasAssTemporalOverride(collectAssOverrideCommands('{\\move(1,2,3,4)}文字')), true);
});
test('isAssTemporalCommand covers only intrinsically animated tags', () => {
for (const command of ['t', 'move', 'k', 'kf', 'ko', 'K']) {
assert.equal(isAssTemporalCommand(command), true, command);
}
for (const command of ['clip', 'iclip', 'frz', 'fscx', 'blur', 'be', 'pos', 'fad']) {
assert.equal(isAssTemporalCommand(command), false, command);
}
});
test('assOverrideSignature distinguishes events by their override values', () => {
const first = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 1 1)}歌詞'));
const second = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 2 2)}歌詞'));
const repeat = assOverrideSignature(collectAssOverrideCommands('{\\clip(m 1 1)}別の行'));
assert.notEqual(first, second);
assert.equal(first, repeat);
});
test('parseAssEffectField classifies the event-level Effect column', () => {
assert.equal(parseAssEffectField(''), 'none');
assert.equal(parseAssEffectField(' '), 'none');
assert.equal(parseAssEffectField('Banner;20;1;0'), 'banner');
assert.equal(parseAssEffectField('Scroll up;0;0;30;10'), 'scroll');
assert.equal(parseAssEffectField('Scroll down;0;0;30;10'), 'scroll');
assert.equal(parseAssEffectField('Karaoke'), 'karaoke');
assert.equal(parseAssEffectField('fx-template'), 'other');
});
test('parseAssEffectField matches stock effect names exactly', () => {
// Custom effect names that merely start with a stock name are not stock effects.
assert.equal(parseAssEffectField('scrolling-credit'), 'other');
assert.equal(parseAssEffectField('bannerfx;1'), 'other');
assert.equal(parseAssEffectField('karaoke-template'), 'other');
assert.equal(parseAssEffectField('Scroll'), 'other');
});
test('isAnimatedAssEffectKind covers the stock animated effects only', () => {
assert.equal(isAnimatedAssEffectKind('karaoke'), true);
assert.equal(isAnimatedAssEffectKind('banner'), true);
assert.equal(isAnimatedAssEffectKind('scroll'), true);
// Typesetting groups put static template names in this column too.
assert.equal(isAnimatedAssEffectKind('other'), false);
assert.equal(isAnimatedAssEffectKind('none'), false);
});
+280
View File
@@ -0,0 +1,280 @@
/*
* ASS/SSA text handling, split into two deliberately distinct contracts:
*
* assToPlainText() raw ASS event text -> plain text. Ingestion only.
* normalizePlainSubtitleText() already-decoded text -> display/lookup form.
*
* Subtitle text is decoded from ASS exactly once, at the point it enters the app: the
* file cue parser does it for sidecar/embedded scripts, and mpv does it for live text
* (`sub-text` is already run through mpv's own `ass_to_plaintext`). Everything
* downstream -- renderer, timing tracker, tokenizer, tokenization cache keys -- gets
* plain text and only normalizes whitespace, so no layer decodes the same string twice.
*
* assToPlainText mirrors mpv's `ass_to_plaintext` rather than inventing its own rules,
* so a cue parsed from a file reads the same as the same line arriving live:
* - `{...}` override blocks are markup
* - `\pN ... \p0` runs are vector paths, not dialogue
* - `\N`, `\n` and `\h` are the only escapes; `\{`, `\}` and `\\` are NOT escapes,
* so `\{注\}` decodes to a lone backslash exactly as mpv renders it
* - an unclosed `{` is rendered verbatim instead of swallowing the rest of the line
* Because the decoder never emits an escape or a closed brace, running it twice is a
* no-op -- but downstream code should still use normalizePlainSubtitleText.
*/
/** What `\N` and `\n` become. */
export type AssLineBreak = '\n' | ' ';
// `\p<n>` with n > 0 switches libass into vector-drawing mode: everything until the
// next `\p0` is a path (`m 20 0 b 10 0 ...`), not dialogue. The negative lookahead keeps
// `\pos(...)` from being read as a drawing tag.
const ASS_DRAWING_SCALE_PATTERN = /\\p(?![a-zA-Z])(\d*)/g;
function readDrawingScale(block: string): number | null {
ASS_DRAWING_SCALE_PATTERN.lastIndex = 0;
let scale: number | null = null;
let match: RegExpExecArray | null;
// Drawing mode is whatever the last `\p` tag in this block set it to.
while ((match = ASS_DRAWING_SCALE_PATTERN.exec(block)) !== null) {
scale = match[1] ? Number(match[1]) : 0;
}
return scale;
}
/** Resolve `\N`, `\n` and `\h`. The only text-level escapes libass recognises. */
function resolveWhitespaceEscapes(text: string, lineBreak: AssLineBreak): string {
return text.replace(/\\([Nnh])/g, (_match, escaped: string) =>
escaped === 'h' ? ' ' : lineBreak,
);
}
/** Strip `{...}` override blocks and the drawing runs they enable. */
function stripAssMarkup(raw: string): string {
let out = '';
let cursor = 0;
let drawing = false;
while (cursor < raw.length) {
if (raw[cursor] !== '{') {
if (!drawing) {
out += raw[cursor];
}
cursor += 1;
continue;
}
const close = raw.indexOf('}', cursor + 1);
if (close === -1) {
// mpv shows an unclosed `{` and everything after it. Guessing where the block was
// meant to end can eat a whole line of dialogue.
if (!drawing) {
out += raw.slice(cursor);
}
break;
}
const scale = readDrawingScale(raw.slice(cursor, close + 1));
if (scale !== null) {
drawing = scale > 0;
}
cursor = close + 1;
}
return out;
}
/**
* Decode a raw ASS/SSA event text field. Call this once, where the text enters the app;
* downstream layers take the result as plain text.
*/
export function assToPlainText(text: string, lineBreak: AssLineBreak = '\n'): string {
if (!text) return '';
return resolveWhitespaceEscapes(stripAssMarkup(text.replace(/\r\n/g, '\n')), lineBreak);
}
export interface NormalizePlainSubtitleTextOptions {
/** Fold every line break into a single space. */
collapseLineBreaks?: boolean;
trim?: boolean;
}
/**
* Whitespace normalization for text that has already been decoded -- by mpv for live
* subtitles, by the cue parser for files. Override blocks and drawing runs are none of
* this function's business; a `{` that reaches here is literal text mpv chose to show.
*
* `\N`/`\n`/`\h` are still folded, because subtitle sources outside the ASS path (asbplayer
* and other websocket clients) forward them raw and the display layer has to cope.
*/
export function normalizePlainSubtitleText(
text: string,
options: NormalizePlainSubtitleTextOptions = {},
): string {
if (!text) return '';
const { collapseLineBreaks = false, trim = true } = options;
let normalized = resolveWhitespaceEscapes(
text.replace(/\r\n/g, '\n'),
collapseLineBreaks ? ' ' : '\n',
);
if (collapseLineBreaks) {
normalized = normalized.replace(/\n/g, ' ').replace(/\s+/g, ' ');
}
return trim ? normalized.trim() : normalized;
}
/** The contents of each `{...}` block, without the braces. */
export function extractAssOverrideBlocks(text: string): string[] {
const blocks: string[] = [];
let cursor = 0;
while (cursor < text.length) {
const open = text.indexOf('{', cursor);
if (open === -1) {
break;
}
const close = text.indexOf('}', open + 1);
if (close === -1) {
break;
}
blocks.push(text.slice(open + 1, close));
cursor = close + 1;
}
return blocks;
}
export interface AssOverrideCommand {
/** Tag name without the backslash, e.g. `pos`, `kf`, `1c`. */
name: string;
/** Everything the tag was given, e.g. `960,1068` for `\pos(960,1068)`. */
args: string;
/** Nested inside a `\t(...)` argument, so its value is animated over the event. */
animated: boolean;
}
const ASS_OVERRIDE_NAME_PATTERN = /[1-4]?[a-zA-Z]+/y;
function readCommandArgs(block: string, start: number): { args: string; next: number } {
if (block[start] === '(') {
let depth = 0;
for (let i = start; i < block.length; i += 1) {
if (block[i] === '(') depth += 1;
else if (block[i] === ')') {
depth -= 1;
if (depth === 0) {
return { args: block.slice(start + 1, i), next: i + 1 };
}
}
}
return { args: block.slice(start + 1), next: block.length };
}
const nextTag = block.indexOf('\\', start);
const end = nextTag === -1 ? block.length : nextTag;
return { args: block.slice(start, end), next: end };
}
// `\t(...)` can wrap another `\t(...)`, and nothing in the format stops an author (or a
// malformed file) from nesting them thousands deep. Real typesetting never goes past one
// or two levels, so stop recursing well before the call stack is at risk.
const MAX_ANIMATION_NESTING_DEPTH = 8;
function parseOverrideBlock(
block: string,
animated: boolean,
into: AssOverrideCommand[],
depth = 0,
): void {
let cursor = 0;
while (cursor < block.length) {
if (block[cursor] !== '\\') {
cursor += 1;
continue;
}
ASS_OVERRIDE_NAME_PATTERN.lastIndex = cursor + 1;
const nameMatch = ASS_OVERRIDE_NAME_PATTERN.exec(block);
if (!nameMatch) {
cursor += 1;
continue;
}
const name = nameMatch[0];
const { args, next } = readCommandArgs(block, cursor + 1 + name.length);
into.push({ name, args: args.trim(), animated });
// `\t(0,500,\frz30)` animates whatever it wraps, so record the inner tags too.
if (name === 't' && args.includes('\\') && depth < MAX_ANIMATION_NESTING_DEPTH) {
parseOverrideBlock(args, true, into, depth + 1);
}
cursor = next;
}
}
/**
* Override commands with their arguments, in source order. Only `{...}` blocks are
* inspected, so a `\pos(...)` sitting in visible text is never mistaken for markup.
*/
export function collectAssOverrideCommands(text: string): AssOverrideCommand[] {
const commands: AssOverrideCommand[] = [];
for (const block of extractAssOverrideBlocks(text)) {
parseOverrideBlock(block, false, commands);
}
return commands;
}
// Tags that are animated by definition: `\t` interpolates, `\move` travels, and the
// karaoke tags advance a highlight across the event's own duration. Everything else --
// `\pos`, `\clip`, `\frz`, `\blur`, `\fad` -- is a static value for the event, so its
// presence says nothing about whether neighbouring events form one animation.
const ASS_TEMPORAL_COMMANDS = new Set(['t', 'move', 'k', 'kf', 'ko', 'K']);
export function isAssTemporalCommand(name: string): boolean {
return ASS_TEMPORAL_COMMANDS.has(name);
}
/** True when the event animates on its own, or animates a static tag through `\t(...)`. */
export function hasAssTemporalOverride(commands: readonly AssOverrideCommand[]): boolean {
return commands.some((command) => command.animated || isAssTemporalCommand(command.name));
}
/**
* Canonical form of an event's override values, for comparing consecutive events. Two
* events with the same signature were typeset identically, so neither is a frame of an
* animation the other belongs to.
*/
export function assOverrideSignature(commands: readonly AssOverrideCommand[]): string {
return commands.map((command) => `${command.name}(${command.args})`).join('|');
}
export type AssEffectKind = 'none' | 'banner' | 'scroll' | 'karaoke' | 'other';
// The stock effects, matched exactly. Typesetting groups put their own template names in
// this column -- `scrolling-credit` is a static sign, not libass's `Scroll up` -- so a
// prefix match would hand out animation evidence to arbitrary custom effects.
const STOCK_ASS_EFFECTS = new Map<string, AssEffectKind>([
['banner', 'banner'],
['scroll up', 'scroll'],
['scroll down', 'scroll'],
['karaoke', 'karaoke'],
]);
/**
* The event-level `Effect` column. The stock values (`Banner;...`, `Scroll up;...`,
* `Scroll down;...`, `Karaoke`) all animate; anything else is a custom name and lands in
* `other`.
*/
export function parseAssEffectField(raw: string): AssEffectKind {
const value = raw.trim().toLowerCase();
if (!value) return 'none';
const name = value.split(';', 1)[0]!.trim();
return STOCK_ASS_EFFECTS.get(name) ?? 'other';
}
const ANIMATED_ASS_EFFECT_KINDS = new Set<AssEffectKind>(['banner', 'scroll', 'karaoke']);
export function isAnimatedAssEffectKind(kind: AssEffectKind): boolean {
return ANIMATED_ASS_EFFECT_KINDS.has(kind);
}
+38 -82
View File
@@ -91,20 +91,16 @@ import {
markVideoWatched,
upsertCoverArt,
} from './immersion-tracker/query-maintenance';
import {
cleanupDuplicateSubtitleLines,
type DuplicateSubtitleLineCleanupOptions,
type DuplicateSubtitleLineCleanupSummary,
} from './immersion-tracker/duplicate-line-cleanup';
import { repairJellyfinStreamVideoLinks } from './immersion-tracker/jellyfin-link-repair';
import {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
repairLegacySeasonlessAnimeRows,
resolveAnimeAnilistConflict,
type AnimeMergeRecommendation,
} from './immersion-tracker/anime-season-repair';
import {
mergeAnimeRecords,
moveVideoToAnime as moveVideoToAnimeQuery,
type AnimeMergeSummary,
type VideoMoveSummary,
} from './immersion-tracker/anime-merge';
import {
buildVideoKey,
deriveCanonicalTitle,
@@ -604,8 +600,21 @@ export class ImmersionTrackerService {
});
}
/**
* Collapse animation bursts that earlier versions recorded frame by frame. The whole
* queue is drained first so a burst still waiting to be written is scanned as stored
* rows rather than surviving the cleanup and landing a moment after it.
*/
async cleanupDuplicateSubtitleLines(
options: DuplicateSubtitleLineCleanupOptions = {},
): Promise<DuplicateSubtitleLineCleanupSummary> {
this.drainQueue();
return cleanupDuplicateSubtitleLines(this.db, options);
}
async rebuildLifetimeSummaries(): Promise<LifetimeRebuildSummary> {
this.requireWriteQueueDrained('rebuilding lifetime summaries');
this.flushTelemetry(true);
this.flushNow();
return rebuildLifetimeSummaryTables(this.db);
}
@@ -672,14 +681,6 @@ export class ImmersionTrackerService {
return getAnimeLibrary(this.db);
}
async getAnimeMergeRecommendations(): Promise<AnimeMergeRecommendation[]> {
return getAnimeMergeRecommendations(this.db);
}
async dismissAnimeMergeRecommendation(recommendationId: number): Promise<boolean> {
return dismissAnimeMergeRecommendation(this.db, recommendationId);
}
async getAnimeDetail(animeId: number): Promise<AnimeDetailRow | null> {
this.relinkYoutubeAnimeLibrary();
return getAnimeDetail(this.db, animeId);
@@ -772,63 +773,6 @@ export class ImmersionTrackerService {
deleteAnimeQuery(this.db, animeId);
}
/**
* Fold duplicate library entries into one. Sources that hold the currently
* playing episode are fine: the videos move, nothing is deleted out from
* under the active session.
*/
async mergeAnime(targetAnimeId: number, sourceAnimeIds: number[]): Promise<AnimeMergeSummary> {
const pendingVideoId = this.sessionState?.videoId;
if (pendingVideoId !== undefined) {
await this.pendingAnimeMetadataUpdates.get(pendingVideoId);
}
// This rebuilds the lifetime summaries, which recompute from the database:
// queued writes have to land first or the active session is dropped from
// the merged totals.
this.requireWriteQueueDrained('merging library entries');
return mergeAnimeRecords(this.db, targetAnimeId, sourceAnimeIds);
}
async moveVideoToAnime(videoId: number, targetAnimeId: number): Promise<VideoMoveSummary> {
await this.pendingAnimeMetadataUpdates.get(videoId);
this.requireWriteQueueDrained('moving an episode');
return moveVideoToAnimeQuery(this.db, videoId, targetAnimeId);
}
/**
* Persist every queued write before a caller recomputes summaries from the
* database.
*
* A single `flushNow()` is not enough: forced telemetry is appended to the
* back of the queue while `flushNow()` writes at most `batchSize` entries off
* the front, so a busy session leaves the newest sample unwritten. Stops as
* soon as a pass makes no progress — a rolled-back batch is pushed back onto
* the queue, and looping on that would spin forever.
*
* Returns false when the queue could not be emptied. Summary-rebuilding
* callers fail closed in that case.
*/
private drainWriteQueue(context: string): boolean {
this.flushTelemetry(true);
while (this.queue.length > 0) {
const pending = this.queue.length;
this.flushNow();
if (this.queue.length >= pending) {
this.logger.warn(
`Immersion tracker queue did not drain before ${context}; summaries may lag by ${this.queue.length} writes`,
);
return false;
}
}
return true;
}
private requireWriteQueueDrained(context: string): void {
if (!this.drainWriteQueue(context)) {
throw new Error(`Immersion tracker queue did not drain before ${context}`);
}
}
async reassignAnimeAnilist(
animeId: number,
info: {
@@ -841,13 +785,7 @@ export class ImmersionTrackerService {
coverUrl?: string | null;
},
): Promise<void> {
// The user is acting on this entry, so it is the one that survives when
// another row already claims the same AniList id.
const repair = resolveAnimeAnilistConflict(this.db, animeId, info.anilistId, {
survivor: 'target',
matchConfidence: 'manual',
});
if (repair.anilistAssignmentBlocked) return;
const repair = resolveAnimeAnilistConflict(this.db, animeId, info.anilistId);
this.db
.prepare(
`
@@ -1878,6 +1816,24 @@ export class ImmersionTrackerService {
}
}
/**
* Write out everything queued, not just the next batch.
*
* `flushNow` writes at most `batchSize` entries and does nothing at all while the write
* lock is held, so a maintenance pass that runs straight after it can still be reading
* a database that is missing rows. Each pass has to shrink the queue to continue: a
* failed flush puts its batch back, and looping on that would never finish.
*/
private drainQueue(): void {
while (this.queue.length > 0) {
const pendingBefore = this.queue.length;
this.flushNow();
if (this.queue.length >= pendingBefore) {
return;
}
}
}
private flushSingle(write: QueuedWrite): void {
executeQueuedWrite(write, this.preparedStatements);
}
@@ -1,700 +0,0 @@
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import test from 'node:test';
import { Database } from '../sqlite.js';
import type { DatabaseSync } from '../sqlite.js';
import { applyPragmas, ensureSchema, getOrCreateAnimeRecord } from '../storage.js';
import { mergeAnimeRecords, moveVideoToAnime } from '../anime-merge.js';
import {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
resolveAnimeAnilistConflict,
} from '../anime-season-repair.js';
import { updateAnimeAnilistInfo } from '../query-maintenance.js';
const BASE_MS = 1_700_000_000_000;
function makeDbPath(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-anime-merge-test-'));
return path.join(dir, 'immersion.sqlite');
}
function cleanupDbPath(dbPath: string): void {
const dir = path.dirname(dbPath);
if (!fs.existsSync(dir)) return;
fs.rmSync(dir, { recursive: true, force: true });
}
function withDb(work: (db: DatabaseSync) => void): void {
const dbPath = makeDbPath();
const db = new Database(dbPath);
try {
applyPragmas(db);
ensureSchema(db);
work(db);
} finally {
db.close();
cleanupDbPath(dbPath);
}
}
interface AnimeSeed {
animeId: number;
key: string;
title: string;
anilistId?: number | null;
titleRomaji?: string | null;
}
function insertAnime(db: DatabaseSync, seed: AnimeSeed): void {
db.prepare(
`INSERT INTO imm_anime(anime_id, normalized_title_key, canonical_title, anilist_id, title_romaji, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, ?, ?, ?)`,
).run(
seed.animeId,
seed.key,
seed.title,
seed.anilistId ?? null,
seed.titleRomaji ?? null,
BASE_MS,
BASE_MS,
);
}
interface EpisodeSeed {
videoId: number;
animeId: number;
season?: number | null;
episode?: number;
activeMs?: number;
cards?: number;
}
/** One episode with one ended session, so lifetime rebuilds have something to sum. */
function insertEpisode(db: DatabaseSync, seed: EpisodeSeed): void {
const activeMs = seed.activeMs ?? 1000;
const cards = seed.cards ?? 1;
db.prepare(
`INSERT INTO imm_videos(video_id, video_key, anime_id, canonical_title, source_type, parsed_title, parsed_season, parsed_episode, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, 1, 'Show', ?, ?, 1, 1440000, ?, ?)`,
).run(
seed.videoId,
`local:/tmp/show-${seed.videoId}.mkv`,
seed.animeId,
`Show ${seed.videoId}`,
seed.season ?? null,
seed.episode ?? seed.videoId,
BASE_MS,
BASE_MS,
);
db.prepare(
`INSERT INTO imm_sessions(session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, active_watched_ms, cards_mined, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?, ?, 2, ?, ?, ?, ?)`,
).run(
seed.videoId,
`session-${seed.videoId}`,
seed.videoId,
String(BASE_MS),
String(BASE_MS + activeMs),
activeMs,
cards,
BASE_MS,
BASE_MS,
);
db.prepare(
`INSERT INTO imm_subtitle_lines(session_id, video_id, anime_id, line_index, text, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, 1, ?, ?, ?)`,
).run(seed.videoId, seed.videoId, seed.animeId, `line ${seed.videoId}`, BASE_MS, BASE_MS);
}
function animeIds(db: DatabaseSync): number[] {
return (
db.prepare('SELECT anime_id AS id FROM imm_anime ORDER BY anime_id').all() as Array<{
id: number;
}>
).map((row) => row.id);
}
function videoAnimeId(db: DatabaseSync, videoId: number): number | null {
return (
db.prepare('SELECT anime_id AS id FROM imm_videos WHERE video_id = ?').get(videoId) as {
id: number | null;
}
).id;
}
function lineAnimeIds(db: DatabaseSync, animeId: number): number {
return Number(
(
db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id = ?')
.get(animeId) as { total: number }
).total,
);
}
test('mergeAnimeRecords folds episodes, lines and lifetime totals into the target', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1', anilistId: 555 });
insertEpisode(db, { videoId: 1, animeId: 1, activeMs: 1000, cards: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1, activeMs: 2000, cards: 3 });
const summary = mergeAnimeRecords(db, 1, [2]);
assert.equal(summary.survivingAnimeId, 1);
assert.deepEqual(summary.mergedAnimeIds, [2]);
assert.equal(summary.movedVideos, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.equal(lineAnimeIds(db, 1), 2);
const lifetime = db
.prepare(
'SELECT total_active_ms AS activeMs, total_cards AS cards, episodes_started AS episodes FROM imm_lifetime_anime WHERE anime_id = 1',
)
.get() as { activeMs: number; cards: number; episodes: number };
assert.equal(lifetime.activeMs, 3000);
assert.equal(lifetime.cards, 4);
assert.equal(lifetime.episodes, 2);
});
});
test('mergeAnimeRecords repoints subtitle lines recorded before the anime link landed', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
// Lines are written with the video's anime_id at the time, which is NULL
// until the async title parse assigns one.
db.prepare(
`INSERT INTO imm_subtitle_lines(session_id, video_id, anime_id, line_index, text, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (2, 2, NULL, 2, 'unlinked line', ?, ?)`,
).run(BASE_MS, BASE_MS);
mergeAnimeRecords(db, 1, [2]);
assert.equal(lineAnimeIds(db, 1), 3);
const orphaned = Number(
(
db
.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines WHERE anime_id IS NULL')
.get() as { total: number }
).total,
);
assert.equal(orphaned, 0);
});
});
test('mergeAnimeRecords inherits metadata the target is missing without clobbering its own', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', titleRomaji: 'Shou' });
insertAnime(db, {
animeId: 2,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 555,
titleRomaji: 'Show Romaji',
});
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
mergeAnimeRecords(db, 1, [2]);
const row = db
.prepare(
'SELECT canonical_title AS title, anilist_id AS anilistId, title_romaji AS romaji FROM imm_anime WHERE anime_id = 1',
)
.get() as { title: string; anilistId: number | null; romaji: string | null };
assert.equal(row.title, 'Show');
// anilist_id is UNIQUE, so inheriting it proves the source row was gone first.
assert.equal(row.anilistId, 555);
assert.equal(row.romaji, 'Shou');
});
});
test('mergeAnimeRecords preserves source title identities as aliases of the survivor', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
db.prepare(
`INSERT INTO imm_anime_title_aliases(normalized_title_key, anime_id, CREATED_DATE, LAST_UPDATE_DATE)
VALUES ('show s01', 2, ?, ?)`,
).run(BASE_MS, BASE_MS);
mergeAnimeRecords(db, 1, [2]);
const fromSourceTitle = getOrCreateAnimeRecord(db, {
parsedTitle: 'Show Season 1',
canonicalTitle: 'Show Season 1',
seasonScope: 1,
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
const fromTransferredAlias = getOrCreateAnimeRecord(db, {
parsedTitle: 'Show S01',
canonicalTitle: 'Show S01',
anilistId: null,
titleRomaji: null,
titleEnglish: null,
titleNative: null,
metadataJson: null,
});
assert.equal(fromSourceTitle, 1);
assert.equal(fromTransferredAlias, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(
(
db.prepare('SELECT canonical_title AS title FROM imm_anime WHERE anime_id = 1').get() as {
title: string;
}
).title,
'Show',
);
});
});
test('mergeAnimeRecords ignores unknown targets and self-merges', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
assert.deepEqual(mergeAnimeRecords(db, 99, [1]).mergedAnimeIds, []);
assert.deepEqual(mergeAnimeRecords(db, 1, [1]).mergedAnimeIds, []);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 1), 1);
});
});
test('moveVideoToAnime moves one episode and prunes the emptied entry', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'stray', title: 'Stray Episode Title', anilistId: 777 });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, activeMs: 5000, cards: 2 });
const summary = moveVideoToAnime(db, 2, 1);
assert.equal(summary.targetAnimeId, 1);
assert.equal(summary.previousAnimeId, 2);
assert.equal(summary.removedPreviousAnime, true);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.equal(lineAnimeIds(db, 1), 2);
const lifetime = db
.prepare('SELECT total_active_ms AS activeMs FROM imm_lifetime_anime WHERE anime_id = 1')
.get() as { activeMs: number };
assert.equal(lifetime.activeMs, 6000);
// The stray entry's AniList link is dropped, not inherited: a move makes no
// claim that the two entries are the same show.
const target = db
.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 1')
.get() as { anilistId: number | null };
assert.equal(target.anilistId, null);
});
});
test('moveVideoToAnime is a no-op when the episode is already in the target entry', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
const summary = moveVideoToAnime(db, 1, 1);
assert.equal(summary.targetAnimeId, 1);
assert.equal(summary.previousAnimeId, 1);
assert.equal(summary.removedPreviousAnime, false);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 1), 1);
});
});
test('moveVideoToAnime keeps the source entry when other episodes remain', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertAnime(db, { animeId: 2, key: 'other', title: 'Other' });
insertEpisode(db, { videoId: 1, animeId: 2 });
insertEpisode(db, { videoId: 2, animeId: 2 });
const summary = moveVideoToAnime(db, 2, 1);
assert.equal(summary.removedPreviousAnime, false);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 1), 2);
assert.equal(videoAnimeId(db, 2), 1);
});
});
test('moveVideoToAnime rejects unknown episodes and targets', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show' });
insertEpisode(db, { videoId: 1, animeId: 1 });
assert.throws(() => moveVideoToAnime(db, 99, 1));
assert.throws(() => moveVideoToAnime(db, 1, 99));
assert.equal(videoAnimeId(db, 1), 1);
});
});
test('resolveAnimeAnilistConflict folds a seasonless duplicate into the entry that owns the id', () => {
withDb((db) => {
// Same show, split because one release tagged S01 and the other did not.
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.survivingAnimeId, 1);
assert.equal(summary.movedVideos, 1);
assert.equal(summary.deletedAnimeRows, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
});
});
test('resolveAnimeAnilistConflict recommends a weak title collision instead of merging it', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.repaired, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('automatic AniList update leaves a weak collision unassigned for user review', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Actual Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
exactTitleMatch: false,
});
const target = db
.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 2')
.get() as {
anilistId: number | null;
};
assert.equal(target.anilistId, null);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('dismissed weak collision stays dismissed when automatic resolution repeats', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(dismissAnimeMergeRecommendation(db, 1), true);
resolveAnimeAnilistConflict(db, 2, 163132);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('dismissed recommendation prevents a later exact automatic merge of the pair', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'weak' });
assert.equal(dismissAnimeMergeRecommendation(db, 1), true);
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'exact' });
assert.equal(summary.repaired, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('manual merge clears recommendations involving the absorbed entry', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
mergeAnimeRecords(db, 1, [2]);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('resolveAnimeAnilistConflict keeps the target entry when the user drove the change', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { survivor: 'target' });
assert.equal(summary.survivingAnimeId, 2);
assert.deepEqual(animeIds(db), [2]);
assert.equal(videoAnimeId(db, 1), 2);
const row = db.prepare('SELECT anilist_id AS id FROM imm_anime WHERE anime_id = 2').get() as {
id: number | null;
};
assert.equal(row.id, 163132);
});
});
test('resolveAnimeAnilistConflict falls back to season redistribution for multi-season rows', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 1, season: 2 });
insertEpisode(db, { videoId: 3, animeId: 2, season: 1 });
resolveAnimeAnilistConflict(db, 2, 163132);
// The mixed row is split by season instead of being poured onto one card.
const titles = (
db.prepare('SELECT canonical_title AS title FROM imm_anime ORDER BY title').all() as Array<{
title: string;
}>
).map((row) => row.title);
assert.deepEqual(titles, ['Show Season 1', 'Show Season 2']);
assert.equal(videoAnimeId(db, 1), 2);
assert.equal(videoAnimeId(db, 3), 2);
assert.notEqual(videoAnimeId(db, 2), 2);
});
});
test('resolveAnimeAnilistConflict leaves explicit incompatible seasons and assignments unchanged', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { matchConfidence: 'exact' });
assert.equal(summary.repaired, 0);
assert.equal(summary.movedVideos, 0);
assert.equal(summary.deletedAnimeRows, 0);
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(videoAnimeId(db, 2), 2);
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: 163132 },
{ animeId: 2, anilistId: null },
]);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('manual AniList resolution reassigns across explicit seasons without merging them', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132, { survivor: 'target' });
assert.equal(summary.anilistAssignmentBlocked, false);
assert.deepEqual(animeIds(db), [1, 2]);
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: null },
{ animeId: 2, anilistId: 163132 },
]);
});
});
test('automatic AniList update does not transfer an assignment across explicit seasons', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'show season 1',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show',
});
insertAnime(db, { animeId: 2, key: 'show season 2', title: 'Show Season 2' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
exactTitleMatch: true,
});
const assignments = db
.prepare(
'SELECT anime_id AS animeId, anilist_id AS anilistId FROM imm_anime ORDER BY anime_id',
)
.all() as Array<{ animeId: number; anilistId: number | null }>;
assert.deepEqual(assignments, [
{ animeId: 1, anilistId: 163132 },
{ animeId: 2, anilistId: null },
]);
assert.equal(videoAnimeId(db, 1), 1);
assert.equal(videoAnimeId(db, 2), 2);
});
});
test('automatic AniList update with unknown match confidence validates stored titles', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'actual show',
title: 'Actual Show',
anilistId: 163132,
titleRomaji: 'Actual Show',
});
insertAnime(db, { animeId: 2, key: 'unrelated release', title: 'Unrelated Release' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
updateAnimeAnilistInfo(db, 2, {
anilistId: 163132,
titleRomaji: 'Actual Show',
titleEnglish: null,
titleNative: null,
episodesTotal: 12,
});
assert.deepEqual(animeIds(db), [1, 2]);
assert.equal(videoAnimeId(db, 2), 2);
assert.deepEqual(getAnimeMergeRecommendations(db), [{ recommendationId: 1, animeIds: [1, 2] }]);
});
});
test('stored AniList titles ignore season suffixes when validating an automatic merge', () => {
withDb((db) => {
insertAnime(db, {
animeId: 1,
key: 'legacy show',
title: 'Show Season 1',
anilistId: 163132,
titleRomaji: 'Show Season 1',
});
insertAnime(db, { animeId: 2, key: 'show season 1', title: 'Show Season 1' });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 1 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(summary.deletedAnimeRows, 1);
assert.deepEqual(animeIds(db), [1]);
assert.equal(videoAnimeId(db, 2), 1);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
test('resolveAnimeAnilistConflict leaves an entry that already links elsewhere alone', () => {
withDb((db) => {
insertAnime(db, { animeId: 1, key: 'show', title: 'Show', anilistId: 163132 });
insertAnime(db, { animeId: 2, key: 'show s2', title: 'Show Season 2', anilistId: 999 });
insertEpisode(db, { videoId: 1, animeId: 1, season: 1 });
insertEpisode(db, { videoId: 2, animeId: 2, season: 2 });
const summary = resolveAnimeAnilistConflict(db, 2, 163132);
assert.equal(videoAnimeId(db, 2), 2);
assert.ok(animeIds(db).includes(2));
assert.equal(
(
db.prepare('SELECT anilist_id AS anilistId FROM imm_anime WHERE anime_id = 2').get() as {
anilistId: number;
}
).anilistId,
999,
);
assert.equal(summary.repaired, 0);
assert.equal(summary.movedVideos, 0);
assert.deepEqual(getAnimeMergeRecommendations(db), []);
});
});
@@ -0,0 +1,349 @@
import assert from 'node:assert/strict';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import test from 'node:test';
import { Database } from '../sqlite.js';
import type { DatabaseSync } from '../sqlite.js';
import { ensureSchema } from '../storage.js';
import { cleanupDuplicateSubtitleLines } from '../duplicate-line-cleanup.js';
const DAY_MS = 86_400_000;
const BASE_MS = 1_700_000_000_000;
const WORD_ID = 1;
interface SeedLine {
session: number;
text: string;
startMs: number;
endMs: number;
/** Recording wall-clock, i.e. what the lookback window filters on. */
createdMs?: number;
}
function makeDbPath(): string {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-duplicate-line-test-'));
return path.join(dir, 'immersion.sqlite');
}
function cleanupDbPath(dbPath: string): void {
const dir = path.dirname(dbPath);
if (!fs.existsSync(dir)) return;
fs.rmSync(dir, { recursive: true, force: true });
}
/** One episode, two sessions of it, and one word occurrence per seeded line. */
function seed(db: DatabaseSync, lines: SeedLine[]): void {
db.exec(`
INSERT INTO imm_anime(anime_id, normalized_title_key, canonical_title, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'show', 'Show', ${BASE_MS}, ${BASE_MS});
INSERT INTO imm_videos(video_id, video_key, anime_id, canonical_title, source_type, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 'v1', 1, 'Ep 1', 1, 1, 1440000, ${BASE_MS}, ${BASE_MS});
INSERT INTO imm_sessions(session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (1, 's1', 1, '${BASE_MS}', '${BASE_MS + 1000}', 2, ${BASE_MS}, ${BASE_MS}),
(2, 's2', 1, '${BASE_MS + DAY_MS}', '${BASE_MS + DAY_MS + 1000}', 2, ${BASE_MS}, ${BASE_MS});
INSERT INTO imm_words(id, headword, word, reading, part_of_speech, pos1, first_seen, last_seen, frequency)
VALUES (${WORD_ID}, '飛び上がる', '飛び上がる', '', 'verb', '動詞', ${Math.floor(BASE_MS / 1000)}, ${Math.floor(BASE_MS / 1000)}, 0);
`);
const insertLine = db.prepare(
`INSERT INTO imm_subtitle_lines(
line_id, session_id, video_id, anime_id, line_index,
segment_start_ms, segment_end_ms, text, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, 1, 1, ?, ?, ?, ?, ?, ?)`,
);
const insertOccurrence = db.prepare(
`INSERT INTO imm_word_line_occurrences(line_id, word_id, occurrence_count, seen_ms)
VALUES (?, ?, 1, ?)`,
);
lines.forEach((line, index) => {
const lineId = index + 1;
const lineIndex = index + 1;
const createdMs = line.createdMs ?? BASE_MS;
insertLine.run(
lineId,
line.session,
lineIndex,
line.startMs,
line.endMs,
line.text,
createdMs,
createdMs,
);
insertOccurrence.run(lineId, WORD_ID, createdMs);
});
db.exec(`
UPDATE imm_words SET frequency = (
SELECT COALESCE(SUM(o.occurrence_count), 0)
FROM imm_word_line_occurrences o WHERE o.word_id = imm_words.id
)
`);
}
function createDb(lines: SeedLine[]): { db: DatabaseSync; dbPath: string } {
const dbPath = makeDbPath();
const db = new Database(dbPath);
ensureSchema(db);
seed(db, lines);
return { db, dbPath };
}
/** A typeset line mpv reported once per animation frame. */
function karaokeFrames(
session: number,
text: string,
startMs: number,
frames: number,
frameMs: number,
): SeedLine[] {
return Array.from({ length: frames }, (_, index) => ({
session,
text,
startMs: startMs + index * frameMs,
endMs: startMs + (index + 1) * frameMs,
}));
}
function countLines(db: DatabaseSync): number {
return (db.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines').get() as { total: number })
.total;
}
function wordFrequency(db: DatabaseSync): number {
const row = db.prepare('SELECT frequency FROM imm_words WHERE id = ?').get(WORD_ID) as {
frequency: number;
} | null;
return row?.frequency ?? 0;
}
test('a karaoke burst collapses to one line and gives back its word counts', () => {
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 40, 40),
{ session: 1, text: 'おはよう', startMs: 20_000, endMs: 22_000 },
]);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 1);
assert.equal(summary.removedLines, 39);
assert.equal(summary.removedWordOccurrences, 39);
assert.equal(countLines(db), 2);
assert.equal(wordFrequency(db), 2);
// The surviving line covers the whole run, the way the parsed cue would.
const kept = db
.prepare(
'SELECT segment_start_ms AS startMs, segment_end_ms AS endMs FROM imm_subtitle_lines WHERE line_id = 1',
)
.get() as { startMs: number; endMs: number };
assert.equal(kept.startMs, 10_000);
assert.equal(kept.endMs, 10_000 + 40 * 40);
assert.equal(summary.samples.length, 1);
assert.equal(summary.samples[0]!.text, '飛び上がる');
assert.equal(summary.samples[0]!.frames, 40);
assert.equal(summary.samples[0]!.videoTitle, 'Ep 1');
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('ordinary repeated dialogue survives', () => {
// Six contiguous `飛び上がる`, each held for a normal beat rather than a frame.
const lines = Array.from({ length: 6 }, (_, index) => ({
session: 1,
text: '飛び上がる',
startMs: 5_000 + index * 800,
endMs: 5_000 + (index + 1) * 800,
}));
const { db, dbPath } = createDb(lines);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 0);
assert.equal(summary.removedLines, 0);
assert.equal(countLines(db), 6);
assert.equal(wordFrequency(db), 6);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a long run of quarter-second frames is still a burst', () => {
// Between the timing-only bound (0.1s) and the animation-frame bound (0.3s): heavier
// typesetting lands here, and the run length is what makes it conclusive.
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 10_000, 6, 250));
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 1);
assert.equal(summary.removedLines, 5);
assert.equal(countLines(db), 1);
assert.equal(wordFrequency(db), 1);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a qualifying short-frame burst may end with one long hold frame', () => {
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 8, 40),
{ session: 1, text: '飛び上がる', startMs: 10_320, endMs: 12_320 },
]);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 1);
assert.equal(summary.removedLines, 8);
assert.equal(countLines(db), 1);
assert.equal(wordFrequency(db), 1);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a long event before the final frame prevents burst cleanup', () => {
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 5, 40),
{ session: 1, text: '飛び上がる', startMs: 10_200, endMs: 12_200 },
{ session: 1, text: '飛び上がる', startMs: 12_200, endMs: 12_240 },
]);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 0);
assert.equal(countLines(db), 7);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a run of frames longer than the animation bound survives', () => {
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 10_000, 6, 400));
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 0);
assert.equal(countLines(db), 6);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a short run below the threshold survives', () => {
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 1_000, 4, 40));
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 0);
assert.equal(countLines(db), 4);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('the same line in a rewatch session is never merged into the first watch', () => {
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40),
...karaokeFrames(2, '飛び上がる', 10_000, 6, 40),
]);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 2);
assert.equal(summary.removedLines, 10);
// One surviving line per session, not one across both.
assert.equal(countLines(db), 2);
assert.equal(wordFrequency(db), 2);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a gap between runs splits them', () => {
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40),
...karaokeFrames(1, '飛び上がる', 60_000, 6, 40),
]);
try {
const summary = cleanupDuplicateSubtitleLines(db);
assert.equal(summary.burstGroups, 2);
assert.equal(countLines(db), 2);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('a dry run reports what an apply would do and writes nothing', () => {
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 10_000, 40, 40));
try {
const preview = cleanupDuplicateSubtitleLines(db, { dryRun: true });
assert.equal(preview.dryRun, true);
assert.equal(preview.removedLines, 39);
assert.equal(countLines(db), 40);
assert.equal(wordFrequency(db), 40);
const applied = cleanupDuplicateSubtitleLines(db);
assert.equal(applied.removedLines, preview.removedLines);
assert.equal(applied.removedWordOccurrences, preview.removedWordOccurrences);
assert.equal(countLines(db), 1);
} finally {
db.close();
cleanupDbPath(dbPath);
}
});
test('the lookback window leaves older bursts alone', () => {
const recentMs = BASE_MS;
const oldMs = BASE_MS - 40 * DAY_MS;
const { db, dbPath } = createDb([
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40).map((line) => ({
...line,
createdMs: oldMs,
})),
...karaokeFrames(2, '飛び上がる', 10_000, 6, 40).map((line) => ({
...line,
createdMs: recentMs,
})),
]);
globalThis.__subminerTestNowMs = BASE_MS;
try {
const summary = cleanupDuplicateSubtitleLines(db, { lookbackDays: 30 });
assert.equal(summary.lookbackDays, 30);
assert.equal(summary.scannedLines, 6);
assert.equal(summary.burstGroups, 1);
assert.equal(summary.removedLines, 5);
// Six untouched old frames plus the one surviving recent line.
assert.equal(countLines(db), 7);
assert.equal(wordFrequency(db), 7);
} finally {
globalThis.__subminerTestNowMs = undefined;
db.close();
cleanupDbPath(dbPath);
}
});
@@ -1,169 +0,0 @@
import type { DatabaseSync } from './sqlite';
import { animeSeasonsAreMergeCompatible, getParsedSeasonsForAnime } from './anime-merge';
import { toDbTimestamp } from './query-shared';
import { normalizeAnimeIdentityKey } from './storage';
import { nowMs } from './time';
export interface AnimeMergeRecommendation {
recommendationId: number;
animeIds: [number, number];
}
export interface AnimeConflictRecommendationOptions {
survivor?: 'target' | 'existing';
/** Automatic matches must be exact; manual assignment is authoritative. */
matchConfidence?: 'exact' | 'weak' | 'manual';
}
interface AnimeTitleRow {
canonical_title: string;
title_romaji: string | null;
title_english: string | null;
title_native: string | null;
}
function getAnimeTitles(db: DatabaseSync, animeId: number): AnimeTitleRow | null {
return db
.prepare(
`SELECT canonical_title, title_romaji, title_english, title_native
FROM imm_anime
WHERE anime_id = ?`,
)
.get(animeId) as AnimeTitleRow | null;
}
function getParsedTitles(db: DatabaseSync, animeId: number): Array<string | null> {
return (
db.prepare('SELECT parsed_title FROM imm_videos WHERE anime_id = ?').all(animeId) as Array<{
parsed_title: string | null;
}>
).map((row) => row.parsed_title);
}
function stripSeasonIdentitySuffix(title: string): string {
return title
.replace(/\bseason\s*\d{1,2}\b/gi, ' ')
.replace(/\b\d{1,2}(?:st|nd|rd|th)\s+season\b/gi, ' ')
.replace(/\bs\d{1,2}\b/gi, ' ');
}
export function hasExactStoredTitleMatch(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
): boolean {
const target = getAnimeTitles(db, targetAnimeId);
const conflict = getAnimeTitles(db, conflictAnimeId);
if (!target || !conflict) return false;
const targetKeys = [target.canonical_title, ...getParsedTitles(db, targetAnimeId)]
.filter((title): title is string => Boolean(title?.trim()))
.map((title) => normalizeAnimeIdentityKey(stripSeasonIdentitySuffix(title)))
.filter(Boolean);
const anilistTitleKeys = [
conflict.title_romaji,
conflict.title_english,
conflict.title_native,
conflict.canonical_title,
]
.filter((title): title is string => Boolean(title?.trim()))
.map((title) => normalizeAnimeIdentityKey(stripSeasonIdentitySuffix(title)))
.filter(Boolean);
return targetKeys.some((key) => anilistTitleKeys.includes(key));
}
export function shouldRecommendAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
options: AnimeConflictRecommendationOptions,
): boolean {
if (options.survivor === 'target' || options.matchConfidence === 'manual') return false;
if (
!animeSeasonsAreMergeCompatible(
getParsedSeasonsForAnime(db, targetAnimeId),
getParsedSeasonsForAnime(db, conflictAnimeId),
)
) {
return false;
}
return (
options.matchConfidence === 'weak' ||
(options.matchConfidence === undefined &&
!hasExactStoredTitleMatch(db, targetAnimeId, conflictAnimeId))
);
}
export function recordAnimeMergeRecommendation(
db: DatabaseSync,
firstCandidateAnimeId: number,
secondCandidateAnimeId: number,
anilistId: number,
): void {
const firstAnimeId = Math.min(firstCandidateAnimeId, secondCandidateAnimeId);
const secondAnimeId = Math.max(firstCandidateAnimeId, secondCandidateAnimeId);
const timestamp = toDbTimestamp(nowMs());
db.prepare(
`INSERT INTO imm_anime_merge_recommendations(
first_anime_id, second_anime_id, anilist_id, status, CREATED_DATE, LAST_UPDATE_DATE
) VALUES (?, ?, ?, 'pending', ?, ?)
ON CONFLICT(first_anime_id, second_anime_id, anilist_id) DO UPDATE SET
LAST_UPDATE_DATE = excluded.LAST_UPDATE_DATE`,
).run(firstAnimeId, secondAnimeId, anilistId, timestamp, timestamp);
}
export function hasDismissedAnimeMergeRecommendation(
db: DatabaseSync,
firstCandidateAnimeId: number,
secondCandidateAnimeId: number,
): boolean {
const firstAnimeId = Math.min(firstCandidateAnimeId, secondCandidateAnimeId);
const secondAnimeId = Math.max(firstCandidateAnimeId, secondCandidateAnimeId);
return Boolean(
db
.prepare(
`SELECT 1
FROM imm_anime_merge_recommendations
WHERE first_anime_id = ?
AND second_anime_id = ?
AND status = 'dismissed'
LIMIT 1`,
)
.get(firstAnimeId, secondAnimeId),
);
}
export function getAnimeMergeRecommendations(db: DatabaseSync): AnimeMergeRecommendation[] {
return (
db
.prepare(
`SELECT recommendation_id AS recommendationId,
first_anime_id AS firstAnimeId,
second_anime_id AS secondAnimeId
FROM imm_anime_merge_recommendations
WHERE status = 'pending'
ORDER BY recommendation_id ASC`,
)
.all() as Array<{
recommendationId: number;
firstAnimeId: number;
secondAnimeId: number;
}>
).map((row) => ({
recommendationId: row.recommendationId,
animeIds: [row.firstAnimeId, row.secondAnimeId],
}));
}
export function dismissAnimeMergeRecommendation(
db: DatabaseSync,
recommendationId: number,
): boolean {
const result = db
.prepare(
`UPDATE imm_anime_merge_recommendations
SET status = 'dismissed', LAST_UPDATE_DATE = ?
WHERE recommendation_id = ? AND status = 'pending'`,
)
.run(toDbTimestamp(nowMs()), recommendationId) as { changes: number };
return result.changes > 0;
}
@@ -1,292 +0,0 @@
import type { DatabaseSync } from './sqlite';
import { rebuildLifetimeSummariesInTransaction } from './lifetime';
import { toDbTimestamp } from './query-shared';
import { nowMs } from './time';
/** Thrown when a move names an episode or destination entry that is not there. */
export const UNKNOWN_MOVE_TARGET_MESSAGE = 'Unknown episode or target library entry';
export interface AnimeMergeSummary {
/** Library entry that owns every moved episode once the merge finishes. */
survivingAnimeId: number;
/** Entries that were folded into the survivor and deleted. */
mergedAnimeIds: number[];
movedVideos: number;
}
export interface VideoMoveSummary {
targetAnimeId: number;
/** Previous owner, or null when the episode had no library entry yet. */
previousAnimeId: number | null;
/** True when the previous owner was left empty and pruned. */
removedPreviousAnime: boolean;
}
interface AnimeMetadataRow {
normalized_title_key: string;
anilist_id: number | null;
title_romaji: string | null;
title_english: string | null;
title_native: string | null;
episodes_total: number | null;
description: string | null;
}
function emptyMergeSummary(survivingAnimeId: number): AnimeMergeSummary {
return { survivingAnimeId, mergedAnimeIds: [], movedVideos: 0 };
}
function runInTransaction<T>(db: DatabaseSync, work: () => T): T {
db.exec('BEGIN IMMEDIATE');
try {
const result = work();
db.exec('COMMIT');
return result;
} catch (error) {
db.exec('ROLLBACK');
throw error;
}
}
function readAnimeMetadata(db: DatabaseSync, animeId: number): AnimeMetadataRow | null {
return (db
.prepare(
`
SELECT normalized_title_key, anilist_id, title_romaji, title_english, title_native, episodes_total, description
FROM imm_anime
WHERE anime_id = ?
`,
)
.get(animeId) ?? null) as AnimeMetadataRow | null;
}
function animeExists(db: DatabaseSync, animeId: number): boolean {
return Boolean(db.prepare('SELECT 1 FROM imm_anime WHERE anime_id = ?').get(animeId));
}
function hasAnimeReferences(db: DatabaseSync, animeId: number): boolean {
const row = db
.prepare(
`
SELECT 1 AS found
WHERE EXISTS (SELECT 1 FROM imm_videos WHERE anime_id = ?)
OR EXISTS (SELECT 1 FROM imm_subtitle_lines WHERE anime_id = ?)
`,
)
.get(animeId, animeId) as { found: number } | null;
return Boolean(row);
}
/**
* Distinct explicit seasons behind a library entry. Videos with no parsed
* season are ignored, so an entry built from `Show - 03.mkv` style filenames
* reports an empty set rather than a bogus season.
*/
export function getParsedSeasonsForAnime(db: DatabaseSync, animeId: number): Set<number> {
const rows = db
.prepare(
`
SELECT DISTINCT parsed_season AS season
FROM imm_videos
WHERE anime_id = ?
AND parsed_season IS NOT NULL
AND parsed_season > 0
`,
)
.all(animeId) as Array<{ season: number }>;
return new Set(rows.map((row) => row.season));
}
/**
* Two entries are safe to fold together when neither spans more than one
* explicit season and they do not disagree about which season that is. A
* seasonless entry is compatible with anything single-season: those are the
* `Show - 03.mkv` vs `Show.S01E03.mkv` splits that produce duplicate cards.
*/
export function animeSeasonsAreMergeCompatible(a: Set<number>, b: Set<number>): boolean {
if (a.size > 1 || b.size > 1) return false;
if (a.size === 0 || b.size === 0) return true;
return [...a][0] === [...b][0];
}
/**
* Fill in whatever the target is missing from a source row that is on its way
* out. Must run after the source row is deleted: imm_anime.anilist_id is
* UNIQUE, so the two rows cannot hold the same id at once.
*/
function absorbAnimeMetadata(
db: DatabaseSync,
targetAnimeId: number,
source: AnimeMetadataRow | null,
updatedAt: string,
): void {
if (!source) return;
db.prepare(
`
UPDATE imm_anime
SET
anilist_id = COALESCE(anilist_id, ?),
title_romaji = COALESCE(title_romaji, ?),
title_english = COALESCE(title_english, ?),
title_native = COALESCE(title_native, ?),
episodes_total = COALESCE(episodes_total, ?),
description = COALESCE(description, ?),
LAST_UPDATE_DATE = ?
WHERE anime_id = ?
`,
).run(
source.anilist_id,
source.title_romaji,
source.title_english,
source.title_native,
source.episodes_total,
source.description,
updatedAt,
targetAnimeId,
);
}
/**
* Fold `sourceAnimeIds` into `targetAnimeId`: every episode and subtitle line
* is repointed, metadata the target is missing is inherited from the sources,
* and the emptied source rows are deleted.
*
* Assumes the caller already holds a write transaction and rebuilds the
* lifetime summaries afterwards; use {@link mergeAnimeRecords} otherwise.
*/
export function mergeAnimeRecordsInTransaction(
db: DatabaseSync,
targetAnimeId: number,
sourceAnimeIds: number[],
): AnimeMergeSummary {
const summary = emptyMergeSummary(targetAnimeId);
if (!animeExists(db, targetAnimeId)) {
return summary;
}
const updatedAt = toDbTimestamp(nowMs());
const sourceVideosStmt = db.prepare(
'SELECT video_id AS videoId FROM imm_videos WHERE anime_id = ?',
);
const moveVideosStmt = db.prepare(
'UPDATE imm_videos SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE anime_id = ?',
);
// Repointed per video rather than by anime_id: lines recorded before the
// async title parse assigns the link are stored with a NULL anime_id, and
// matching on the source id would strand them unattributed.
const moveLinesStmt = db.prepare(
'UPDATE imm_subtitle_lines SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE video_id = ?',
);
const dropLifetimeStmt = db.prepare('DELETE FROM imm_lifetime_anime WHERE anime_id = ?');
const sourceAliasesStmt = db.prepare(
'SELECT normalized_title_key AS normalizedTitleKey FROM imm_anime_title_aliases WHERE anime_id = ?',
);
const upsertAliasStmt = db.prepare(
`INSERT INTO imm_anime_title_aliases(normalized_title_key, anime_id, CREATED_DATE, LAST_UPDATE_DATE)
VALUES (?, ?, ?, ?)
ON CONFLICT(normalized_title_key) DO UPDATE SET
anime_id = excluded.anime_id,
LAST_UPDATE_DATE = excluded.LAST_UPDATE_DATE`,
);
const dropSourceAliasesStmt = db.prepare(
'DELETE FROM imm_anime_title_aliases WHERE anime_id = ?',
);
const dropAnimeStmt = db.prepare('DELETE FROM imm_anime WHERE anime_id = ?');
for (const sourceAnimeId of new Set(sourceAnimeIds)) {
if (sourceAnimeId === targetAnimeId || !animeExists(db, sourceAnimeId)) {
continue;
}
const sourceMetadata = readAnimeMetadata(db, sourceAnimeId);
const sourceAliases = sourceAliasesStmt.all(sourceAnimeId) as Array<{
normalizedTitleKey: string;
}>;
const sourceVideoIds = (sourceVideosStmt.all(sourceAnimeId) as Array<{ videoId: number }>).map(
(row) => row.videoId,
);
const moved = moveVideosStmt.run(targetAnimeId, updatedAt, sourceAnimeId) as {
changes: number;
};
for (const videoId of sourceVideoIds) {
moveLinesStmt.run(targetAnimeId, updatedAt, videoId);
}
dropSourceAliasesStmt.run(sourceAnimeId);
for (const alias of [
...(sourceMetadata ? [sourceMetadata.normalized_title_key] : []),
...sourceAliases.map((row) => row.normalizedTitleKey),
]) {
upsertAliasStmt.run(alias, targetAnimeId, updatedAt, updatedAt);
}
dropLifetimeStmt.run(sourceAnimeId);
dropAnimeStmt.run(sourceAnimeId);
absorbAnimeMetadata(db, targetAnimeId, sourceMetadata, updatedAt);
summary.mergedAnimeIds.push(sourceAnimeId);
summary.movedVideos += moved.changes;
}
return summary;
}
export function mergeAnimeRecords(
db: DatabaseSync,
targetAnimeId: number,
sourceAnimeIds: number[],
): AnimeMergeSummary {
return runInTransaction(db, () => {
const summary = mergeAnimeRecordsInTransaction(db, targetAnimeId, sourceAnimeIds);
if (summary.mergedAnimeIds.length > 0) {
rebuildLifetimeSummariesInTransaction(db);
}
return summary;
});
}
/**
* Move a single episode to another library entry, pruning the previous owner
* when it is left with nothing.
*/
export function moveVideoToAnime(
db: DatabaseSync,
videoId: number,
targetAnimeId: number,
): VideoMoveSummary {
return runInTransaction(db, () => {
const videoRow = db
.prepare('SELECT anime_id AS animeId FROM imm_videos WHERE video_id = ?')
.get(videoId) as { animeId: number | null } | null;
if (!videoRow || !animeExists(db, targetAnimeId)) {
throw new Error(UNKNOWN_MOVE_TARGET_MESSAGE);
}
const previousAnimeId = videoRow.animeId;
if (previousAnimeId === targetAnimeId) {
return { targetAnimeId, previousAnimeId, removedPreviousAnime: false };
}
const updatedAt = toDbTimestamp(nowMs());
db.prepare('UPDATE imm_videos SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE video_id = ?').run(
targetAnimeId,
updatedAt,
videoId,
);
db.prepare(
'UPDATE imm_subtitle_lines SET anime_id = ?, LAST_UPDATE_DATE = ? WHERE video_id = ?',
).run(targetAnimeId, updatedAt, videoId);
let removedPreviousAnime = false;
if (previousAnimeId !== null && !hasAnimeReferences(db, previousAnimeId)) {
// The emptied entry's metadata is deliberately dropped rather than
// absorbed. A move says "this episode belongs elsewhere", not "these are
// the same show", and the entry being emptied is usually a mis-parse
// whose AniList link would be wrong for the target.
db.prepare('DELETE FROM imm_lifetime_anime WHERE anime_id = ?').run(previousAnimeId);
db.prepare('DELETE FROM imm_anime WHERE anime_id = ?').run(previousAnimeId);
removedPreviousAnime = true;
}
rebuildLifetimeSummariesInTransaction(db);
return { targetAnimeId, previousAnimeId, removedPreviousAnime };
});
}
@@ -1,16 +1,4 @@
import type { DatabaseSync } from './sqlite';
import {
animeSeasonsAreMergeCompatible,
getParsedSeasonsForAnime,
mergeAnimeRecordsInTransaction,
} from './anime-merge';
import {
hasExactStoredTitleMatch,
hasDismissedAnimeMergeRecommendation,
recordAnimeMergeRecommendation,
shouldRecommendAnilistConflict,
type AnimeConflictRecommendationOptions,
} from './anime-merge-recommendations';
import { getOrCreateAnimeRecord } from './storage';
import { toDbTimestamp } from './query-shared';
import { nowMs } from './time';
@@ -20,33 +8,8 @@ export interface AnimeSeasonRepairSummary {
repaired: number;
movedVideos: number;
deletedAnimeRows: number;
/**
* Entry that owns the videos afterwards when two rows were folded together,
* so callers can keep pointing at a row that still exists.
*/
survivingAnimeId: number | null;
/** True when an ambiguous AniList collision was saved for user review. */
mergeRecommended: boolean;
/** True when automatic metadata must not assign the colliding AniList id. */
anilistAssignmentBlocked: boolean;
}
export interface AnimeAnilistConflictOptions extends AnimeConflictRecommendationOptions {
/**
* Which row keeps its identity when two entries claim the same AniList id.
* `existing` (the default) keeps the row that already held the id, so
* automatic cover-art resolution does not rename a card under the user;
* `target` keeps the row the user is acting on.
*/
survivor?: 'target' | 'existing';
}
export {
dismissAnimeMergeRecommendation,
getAnimeMergeRecommendations,
type AnimeMergeRecommendation,
} from './anime-merge-recommendations';
interface AnimeRow {
anime_id: number;
anilist_id: number | null;
@@ -75,9 +38,6 @@ function emptySummary(scanned = 0): AnimeSeasonRepairSummary {
repaired: 0,
movedVideos: 0,
deletedAnimeRows: 0,
survivingAnimeId: null,
mergeRecommended: false,
anilistAssignmentBlocked: false,
};
}
@@ -89,9 +49,6 @@ function mergeSummary(
target.repaired += source.repaired;
target.movedVideos += source.movedVideos;
target.deletedAnimeRows += source.deletedAnimeRows;
target.survivingAnimeId = source.survivingAnimeId ?? target.survivingAnimeId;
target.mergeRecommended ||= source.mergeRecommended;
target.anilistAssignmentBlocked ||= source.anilistAssignmentBlocked;
return target;
}
@@ -344,18 +301,10 @@ export function repairLegacySeasonlessAnimeRows(db: DatabaseSync): AnimeSeasonRe
});
}
/**
* Two library entries cannot both hold the same AniList id
* (`imm_anime.anilist_id` is UNIQUE). Fold an automatic collision only when
* exact title evidence and compatible parsed seasons make it safe. Persist a
* review recommendation for compatible weak matches. Fall back to legacy
* season redistribution when the conflicting row spans several seasons.
*/
export function resolveAnimeAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
anilistId: number,
options: AnimeAnilistConflictOptions = {},
): AnimeSeasonRepairSummary {
const conflict = db
.prepare(
@@ -372,96 +321,10 @@ export function resolveAnimeAnilistConflict(
return emptySummary();
}
return runInTransaction(db, () => {
const targetRow = getAnimeRow(db, targetAnimeId);
if (
options.survivor !== 'target' &&
targetRow?.anilist_id != null &&
targetRow.anilist_id !== anilistId
) {
// An automatic lookup disagreeing with an existing explicit link is a
// mis-resolution, not evidence that either row should move or merge.
return emptySummary(1);
}
const isManual = options.survivor === 'target' || options.matchConfidence === 'manual';
if (!isManual && hasDismissedAnimeMergeRecommendation(db, targetAnimeId, conflict.animeId)) {
const summary = emptySummary(1);
summary.anilistAssignmentBlocked = true;
return summary;
}
const targetSeasons = getParsedSeasonsForAnime(db, targetAnimeId);
const conflictSeasons = getParsedSeasonsForAnime(db, conflict.animeId);
if (
!isManual &&
targetSeasons.size === 1 &&
conflictSeasons.size === 1 &&
[...targetSeasons][0] !== [...conflictSeasons][0]
) {
const summary = emptySummary(1);
summary.anilistAssignmentBlocked = true;
return summary;
}
if (canMergeAnilistConflict(db, targetAnimeId, conflict.animeId, anilistId, options)) {
const survivingAnimeId = options.survivor === 'target' ? targetAnimeId : conflict.animeId;
const absorbedAnimeId = survivingAnimeId === targetAnimeId ? conflict.animeId : targetAnimeId;
const merge = mergeAnimeRecordsInTransaction(db, survivingAnimeId, [absorbedAnimeId]);
const summary = emptySummary(1);
summary.movedVideos = merge.movedVideos;
summary.deletedAnimeRows = merge.mergedAnimeIds.length;
if (merge.mergedAnimeIds.length > 0) {
summary.repaired = 1;
// Only reported once a row really absorbed the other, so callers never
// follow this to an anime id that was never written.
summary.survivingAnimeId = survivingAnimeId;
}
// Lifetime summaries are rebuilt by the caller off this summary, the same
// as the redistribution path below.
return summary;
}
if (shouldRecommendAnilistConflict(db, targetAnimeId, conflict.animeId, options)) {
recordAnimeMergeRecommendation(db, targetAnimeId, conflict.animeId, anilistId);
const summary = emptySummary(1);
summary.mergeRecommended = true;
return summary;
}
return redistributeAnimeRowByParsedSeasonsInTransaction(db, conflict.animeId, {
return runInTransaction(db, () =>
redistributeAnimeRowByParsedSeasonsInTransaction(db, conflict.animeId, {
transferAnilistToAnimeId: targetAnimeId,
overwriteTargetAnilist: true,
});
});
}
function canMergeAnilistConflict(
db: DatabaseSync,
targetAnimeId: number,
conflictAnimeId: number,
anilistId: number,
options: AnimeAnilistConflictOptions,
): boolean {
const targetRow = getAnimeRow(db, targetAnimeId);
if (!targetRow) {
// Nothing to merge with a row that no longer exists (a stale id from the
// caller); fall through to the redistribution path.
return false;
}
if (options.survivor !== 'target') {
// The target is the row about to disappear here, so an existing link of its
// own means this is a mis-resolution rather than a duplicate: leave it be.
if (targetRow.anilist_id != null && targetRow.anilist_id !== anilistId) {
return false;
}
}
if (
options.matchConfidence === 'weak' ||
(options.matchConfidence === undefined &&
!hasExactStoredTitleMatch(db, targetAnimeId, conflictAnimeId))
) {
return false;
}
return animeSeasonsAreMergeCompatible(
getParsedSeasonsForAnime(db, targetAnimeId),
getParsedSeasonsForAnime(db, conflictAnimeId),
}),
);
}
@@ -0,0 +1,386 @@
/*
* Retroactive removal of animation-burst subtitle lines from the stats database.
*
* Before the live ingest gate existed, a karaoke OP recorded one line -- and one count
* for every word in it -- per animation frame, which is enough to put an OP lyric at the
* top of "Top Repeated Words" for good. This module finds those runs in what is already
* stored and takes them back down to one line.
*
* Only timing is available here: the stored text has been stripped of ASS markup, so the
* authoring evidence the file-level parser uses (`\t`, `\move`, karaoke timing, a
* changing override signature) is long gone. What is left is a run of identical,
* contiguous, short-lived lines inside a single session.
*
* The run has to be as long as the timing-only rule in `subtitle-cue-dedup` demands, but
* its short frames may be as long as the animation-frame bound rather than the much
* tighter timing-only one. A qualifying run may end with one longer hold, which is a
* common karaoke shape. Five or more repeats of the same text, each ending where the next
* begins, is already conclusive on its own -- no dialogue does that -- and the tighter
* bound would walk straight past the heavier typesetting that motivated this, where
* frames sit nearer a quarter of a second. Both bounds are options, so a cautious run can
* ask for more, and a dry run always reports before anything is removed.
*
* Scope: subtitle lines, their word/kanji occurrences, and the `imm_words`/`imm_kanji`
* aggregates those occurrences feed. Session telemetry (`lines_seen`, `tokens_seen`) and
* the rollups derived from it are left alone; they are cumulative samples taken at record
* time, and for sessions whose raw rows have since been pruned they cannot be recomputed.
*/
import type { DatabaseSync } from './sqlite';
import {
ANIMATION_FRAME_MAX_SECONDS,
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
MIN_TIMING_ONLY_FRAMES,
} from '../subtitle-burst-constants';
import {
applyLexicalRemovals,
makePlaceholders,
planLexicalRemovalsForLines,
toDbTimestamp,
} from './query-shared';
import { nowMs } from './time';
const MS_PER_DAY = 86_400_000;
/** SQLite caps bound parameters per statement; stay well under it. */
const LINE_ID_BATCH_SIZE = 400;
const DEFAULT_SAMPLE_LIMIT = 20;
export interface DuplicateSubtitleLineCleanupOptions {
/** Only consider lines recorded within this many days. Null or omitted = all history. */
lookbackDays?: number | null;
/** Measure without writing. */
dryRun?: boolean;
/** Identical contiguous lines needed before a run counts as an animation. */
minRunLength?: number;
/** Longest a single event may last and still look like an animation frame. */
maxFrameSeconds?: number;
/** How many of the largest runs to describe in the summary. */
sampleLimit?: number;
}
export interface DuplicateSubtitleLineBurst {
sessionId: number;
videoId: number;
text: string;
/** Kept line, extended to cover the whole run. */
keptLineId: number;
removedLineIds: number[];
startMs: number;
endMs: number;
}
export interface DuplicateSubtitleLineSample {
videoId: number;
videoTitle: string | null;
text: string;
frames: number;
removedLines: number;
startMs: number;
endMs: number;
}
export interface DuplicateSubtitleLineCleanupSummary {
dryRun: boolean;
lookbackDays: number | null;
scannedLines: number;
burstGroups: number;
removedLines: number;
removedWordOccurrences: number;
removedKanjiOccurrences: number;
samples: DuplicateSubtitleLineSample[];
}
export interface StoredSubtitleLineRow {
lineId: number;
sessionId: number;
videoId: number;
text: string;
startMs: number;
endMs: number;
}
interface ResolvedBounds {
lookbackDays: number | null;
minRunLength: number;
maxFrameMs: number;
gapToleranceMs: number;
sampleLimit: number;
}
function resolveBounds(options: DuplicateSubtitleLineCleanupOptions): ResolvedBounds {
const lookbackDays =
typeof options.lookbackDays === 'number' && Number.isFinite(options.lookbackDays)
? Math.max(1, Math.floor(options.lookbackDays))
: null;
const minRunLength =
typeof options.minRunLength === 'number' && Number.isFinite(options.minRunLength)
? Math.max(2, Math.floor(options.minRunLength))
: MIN_TIMING_ONLY_FRAMES;
const maxFrameSeconds =
typeof options.maxFrameSeconds === 'number' && options.maxFrameSeconds > 0
? options.maxFrameSeconds
: ANIMATION_FRAME_MAX_SECONDS;
const sampleLimit =
typeof options.sampleLimit === 'number' && options.sampleLimit >= 0
? Math.floor(options.sampleLimit)
: DEFAULT_SAMPLE_LIMIT;
return {
lookbackDays,
minRunLength,
maxFrameMs: Math.round(maxFrameSeconds * 1000),
gapToleranceMs: Math.round(DUPLICATE_CUE_GAP_TOLERANCE_SECONDS * 1000),
sampleLimit,
};
}
/**
* `CREATED_DATE` holds epoch milliseconds on rows this app wrote, but older and synced
* rows can carry seconds, so normalize before comparing against the cutoff.
*/
const CREATED_MS_SQL = `
CASE
WHEN sl.CREATED_DATE < 10000000000 THEN sl.CREATED_DATE * 1000
ELSE sl.CREATED_DATE
END`;
function readCandidateLines(db: DatabaseSync, bounds: ResolvedBounds): StoredSubtitleLineRow[] {
const scope =
bounds.lookbackDays === null
? ''
: `AND sl.CREATED_DATE IS NOT NULL AND ${CREATED_MS_SQL} >= ?`;
const params = bounds.lookbackDays === null ? [] : [nowMs() - bounds.lookbackDays * MS_PER_DAY];
return db
.prepare(
`SELECT
sl.line_id AS lineId,
sl.session_id AS sessionId,
sl.video_id AS videoId,
sl.text AS text,
sl.segment_start_ms AS startMs,
sl.segment_end_ms AS endMs
FROM imm_subtitle_lines sl
WHERE sl.segment_start_ms IS NOT NULL
AND sl.segment_end_ms IS NOT NULL
${scope}
ORDER BY sl.session_id, sl.video_id, sl.segment_start_ms, sl.line_id`,
)
.all(...params) as StoredSubtitleLineRow[];
}
function isBurst(run: StoredSubtitleLineRow[], bounds: ResolvedBounds): boolean {
if (run.length < bounds.minRunLength) {
return false;
}
const isShortFrame = (row: StoredSubtitleLineRow): boolean =>
row.endMs - row.startMs <= bounds.maxFrameMs;
if (run.every(isShortFrame)) {
return true;
}
// Karaoke commonly finishes its short animation frames with one long hold. Only the
// final event may exceed the frame bound, and the strict short-frame threshold must
// already have been met before it.
return (
run.length - 1 >= bounds.minRunLength &&
run.slice(0, -1).every(isShortFrame) &&
!isShortFrame(run[run.length - 1]!)
);
}
function toBurst(run: StoredSubtitleLineRow[]): DuplicateSubtitleLineBurst {
const [first] = run;
return {
sessionId: first!.sessionId,
videoId: first!.videoId,
text: first!.text,
keptLineId: first!.lineId,
removedLineIds: run.slice(1).map((row) => row.lineId),
startMs: first!.startMs,
endMs: run.reduce((latest, row) => Math.max(latest, row.endMs), first!.endMs),
};
}
/**
* Group stored lines into animation runs.
*
* Runs never cross a session, which is what keeps a rewatch intact: the same episode
* watched twice stores the same line twice, and those two belong to different sessions.
*/
export function findDuplicateSubtitleLineBursts(
rows: readonly StoredSubtitleLineRow[],
options: DuplicateSubtitleLineCleanupOptions = {},
): DuplicateSubtitleLineBurst[] {
const bounds = resolveBounds(options);
const bursts: DuplicateSubtitleLineBurst[] = [];
let run: StoredSubtitleLineRow[] = [];
let chainEndMs = 0;
const closeRun = (): void => {
if (run.length > 1 && isBurst(run, bounds)) {
bursts.push(toBurst(run));
}
run = [];
};
for (const row of rows) {
const previous = run[run.length - 1];
const continuesRun =
previous !== undefined &&
previous.sessionId === row.sessionId &&
previous.videoId === row.videoId &&
previous.text === row.text &&
row.startMs <= chainEndMs + bounds.gapToleranceMs;
if (continuesRun) {
run.push(row);
chainEndMs = Math.max(chainEndMs, row.endMs);
continue;
}
closeRun();
run = [row];
chainEndMs = row.endMs;
}
closeRun();
return bursts;
}
function chunk<T>(values: T[], size: number): T[][] {
const chunks: T[][] = [];
for (let i = 0; i < values.length; i += size) {
chunks.push(values.slice(i, i + size));
}
return chunks;
}
function buildSamples(
db: DatabaseSync,
bursts: DuplicateSubtitleLineBurst[],
sampleLimit: number,
): DuplicateSubtitleLineSample[] {
if (sampleLimit === 0 || bursts.length === 0) {
return [];
}
const largest = [...bursts]
.sort((a, b) => b.removedLineIds.length - a.removedLineIds.length)
.slice(0, sampleLimit);
const videoIds = [...new Set(largest.map((burst) => burst.videoId))];
const titles = new Map<number, string>();
for (const batch of chunk(videoIds, LINE_ID_BATCH_SIZE)) {
const rows = db
.prepare(
`SELECT video_id AS videoId, canonical_title AS title
FROM imm_videos
WHERE video_id IN (${makePlaceholders(batch)})`,
)
.all(...batch) as Array<{ videoId: number; title: string | null }>;
for (const row of rows) {
if (row.title) titles.set(row.videoId, row.title);
}
}
return largest.map((burst) => ({
videoId: burst.videoId,
videoTitle: titles.get(burst.videoId) ?? null,
text: burst.text,
frames: burst.removedLineIds.length + 1,
removedLines: burst.removedLineIds.length,
startMs: burst.startMs,
endMs: burst.endMs,
}));
}
function sumRemovedOccurrences(
db: DatabaseSync,
table: 'imm_word_line_occurrences' | 'imm_kanji_line_occurrences',
lineIds: number[],
): number {
let total = 0;
for (const batch of chunk(lineIds, LINE_ID_BATCH_SIZE)) {
const row = db
.prepare(
`SELECT COALESCE(SUM(occurrence_count), 0) AS total
FROM ${table}
WHERE line_id IN (${makePlaceholders(batch)})`,
)
.get(...batch) as { total: number } | null;
total += row?.total ?? 0;
}
return total;
}
function applyBursts(db: DatabaseSync, bursts: DuplicateSubtitleLineBurst[]): void {
const removedLineIds = bursts.flatMap((burst) => burst.removedLineIds);
const currentMs = toDbTimestamp(nowMs());
db.exec('BEGIN IMMEDIATE');
try {
for (const batch of chunk(removedLineIds, LINE_ID_BATCH_SIZE)) {
const placeholders = makePlaceholders(batch);
// Measured before the delete, applied after it: `applyLexicalRemovals` checks the
// surviving occurrences to decide whether a zeroed count really means the word is
// gone, so the rows it inspects have to be the post-delete ones.
const plan = planLexicalRemovalsForLines(db, batch);
db.prepare(`DELETE FROM imm_word_line_occurrences WHERE line_id IN (${placeholders})`).run(
...batch,
);
db.prepare(`DELETE FROM imm_kanji_line_occurrences WHERE line_id IN (${placeholders})`).run(
...batch,
);
db.prepare(`DELETE FROM imm_subtitle_lines WHERE line_id IN (${placeholders})`).run(...batch);
applyLexicalRemovals(db, plan);
}
const extendStmt = db.prepare(
`UPDATE imm_subtitle_lines
SET segment_end_ms = ?, LAST_UPDATE_DATE = ?
WHERE line_id = ? AND (segment_end_ms IS NULL OR segment_end_ms < ?)`,
);
for (const burst of bursts) {
extendStmt.run(burst.endMs, currentMs, burst.keptLineId, burst.endMs);
}
db.exec('COMMIT');
} catch (error) {
db.exec('ROLLBACK');
throw error;
}
}
/**
* Collapse stored animation bursts down to one line each.
*
* A dry run measures exactly what an apply would remove, using the same scan, so the
* numbers shown in a confirmation prompt are the numbers that will happen.
*/
export function cleanupDuplicateSubtitleLines(
db: DatabaseSync,
options: DuplicateSubtitleLineCleanupOptions = {},
): DuplicateSubtitleLineCleanupSummary {
const bounds = resolveBounds(options);
const dryRun = options.dryRun === true;
const rows = readCandidateLines(db, bounds);
const bursts = findDuplicateSubtitleLineBursts(rows, options);
const removedLineIds = bursts.flatMap((burst) => burst.removedLineIds);
const summary: DuplicateSubtitleLineCleanupSummary = {
dryRun,
lookbackDays: bounds.lookbackDays,
scannedLines: rows.length,
burstGroups: bursts.length,
removedLines: removedLineIds.length,
removedWordOccurrences: sumRemovedOccurrences(db, 'imm_word_line_occurrences', removedLineIds),
removedKanjiOccurrences: sumRemovedOccurrences(
db,
'imm_kanji_line_occurrences',
removedLineIds,
),
samples: buildSamples(db, bursts, bounds.sampleLimit),
};
if (dryRun || removedLineIds.length === 0) {
return summary;
}
applyBursts(db, bursts);
return summary;
}
@@ -418,7 +418,6 @@ export function updateAnimeAnilistInfo(
titleEnglish: string | null;
titleNative: string | null;
episodesTotal: number | null;
exactTitleMatch?: boolean;
},
): void {
const row = db.prepare('SELECT anime_id FROM imm_videos WHERE video_id = ?').get(videoId) as {
@@ -426,11 +425,7 @@ export function updateAnimeAnilistInfo(
} | null;
if (!row?.anime_id) return;
const repair = resolveAnimeAnilistConflict(db, row.anime_id, info.anilistId, {
matchConfidence:
info.exactTitleMatch === true ? 'exact' : info.exactTitleMatch === false ? 'weak' : undefined,
});
if (repair.mergeRecommended || repair.anilistAssignmentBlocked) return;
const repair = resolveAnimeAnilistConflict(db, row.anime_id, info.anilistId);
const targetRow = db
.prepare('SELECT anime_id FROM imm_videos WHERE video_id = ?')
.get(videoId) as {
@@ -268,6 +268,19 @@ export function planLexicalRemovalsForSessions(
return planLexicalRemovals(db, `sl.session_id IN (${makePlaceholders(sessionIds)})`, sessionIds);
}
/**
* Measure what deleting these individual subtitle lines removes from the vocabulary
* tables. Used by the duplicate-line cleanup, which drops animation frames out of the
* middle of sessions that otherwise stay intact.
*/
export function planLexicalRemovalsForLines(
db: DatabaseSync,
lineIds: number[],
): LexicalRemovalPlan {
if (lineIds.length === 0) return EMPTY_LEXICAL_REMOVAL_PLAN;
return planLexicalRemovals(db, `sl.line_id IN (${makePlaceholders(lineIds)})`, lineIds);
}
/** Measure what deleting these videos removes from the vocabulary tables. */
export function planLexicalRemovalsForVideos(
db: DatabaseSync,
+10 -42
View File
@@ -1,6 +1,5 @@
import { createHash } from 'node:crypto';
import { parseMediaInfo } from '../../../jimaku/utils';
import { normalizeTitleIdentity } from '../../utils/title-normalization';
import type { DatabaseSync } from './sqlite';
import { nowMs } from './time';
import { SCHEMA_VERSION } from './types';
@@ -320,7 +319,14 @@ export function applyPragmas(db: DatabaseSync): void {
db.exec(`PRAGMA journal_size_limit = ${WAL_JOURNAL_SIZE_LIMIT_BYTES}`);
}
export const normalizeAnimeIdentityKey = normalizeTitleIdentity;
export function normalizeAnimeIdentityKey(title: string): string {
return title
.normalize('NFKC')
.toLowerCase()
.replace(/[^\p{L}\p{N}]+/gu, ' ')
.trim()
.replace(/\s+/g, ' ');
}
function normalizeSeasonScope(value: number | null | undefined): number | null {
if (typeof value !== 'number' || !Number.isSafeInteger(value) || value <= 0) {
@@ -524,36 +530,6 @@ function ensureStatsExcludedWordsTable(db: DatabaseSync): void {
`);
}
function ensureAnimeMergeTables(db: DatabaseSync): void {
db.exec(`
CREATE TABLE IF NOT EXISTS imm_anime_title_aliases(
normalized_title_key TEXT PRIMARY KEY,
anime_id INTEGER NOT NULL,
CREATED_DATE TEXT,
LAST_UPDATE_DATE TEXT,
FOREIGN KEY(anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE
);
CREATE INDEX IF NOT EXISTS idx_anime_title_aliases_anime_id
ON imm_anime_title_aliases(anime_id);
CREATE TABLE IF NOT EXISTS imm_anime_merge_recommendations(
recommendation_id INTEGER PRIMARY KEY AUTOINCREMENT,
first_anime_id INTEGER NOT NULL,
second_anime_id INTEGER NOT NULL,
anilist_id INTEGER NOT NULL,
status TEXT NOT NULL DEFAULT 'pending' CHECK(status IN ('pending', 'dismissed')),
CREATED_DATE TEXT,
LAST_UPDATE_DATE TEXT,
CHECK(first_anime_id < second_anime_id),
UNIQUE(first_anime_id, second_anime_id, anilist_id),
FOREIGN KEY(first_anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE,
FOREIGN KEY(second_anime_id) REFERENCES imm_anime(anime_id) ON DELETE CASCADE
);
CREATE INDEX IF NOT EXISTS idx_anime_merge_recommendations_status
ON imm_anime_merge_recommendations(status, recommendation_id);
`);
}
export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput): number {
const seasonScope = normalizeSeasonScope(input.seasonScope);
const identityTitle = buildSeasonScopedAnimeTitle(input.parsedTitle, seasonScope);
@@ -574,14 +550,8 @@ export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput
const byNormalizedTitle = db
.prepare('SELECT anime_id FROM imm_anime WHERE normalized_title_key = ?')
.get(normalizedTitleKey) as { anime_id: number } | null;
const byTitleAlias = db
.prepare('SELECT anime_id FROM imm_anime_title_aliases WHERE normalized_title_key = ?')
.get(normalizedTitleKey) as { anime_id: number } | null;
const existing = byAnilistId ?? byNormalizedTitle ?? byTitleAlias;
const existing = byAnilistId ?? byNormalizedTitle;
if (existing?.anime_id) {
// An alias remembers an intentionally merged-away spelling. Reusing it
// must not rename the survivor back to that discarded display title.
const canonicalTitleUpdate = byAnilistId || byNormalizedTitle ? canonicalTitle : null;
db.prepare(
`
UPDATE imm_anime
@@ -596,7 +566,7 @@ export function getOrCreateAnimeRecord(db: DatabaseSync, input: AnimeRecordInput
WHERE anime_id = ?
`,
).run(
canonicalTitleUpdate,
canonicalTitle,
input.anilistId,
input.titleRomaji,
input.titleEnglish,
@@ -781,7 +751,6 @@ export function ensureSchema(db: DatabaseSync): void {
if (currentVersion?.schema_version === SCHEMA_VERSION) {
ensureLifetimeSummaryTables(db);
ensureStatsExcludedWordsTable(db);
ensureAnimeMergeTables(db);
return;
}
@@ -830,7 +799,6 @@ export function ensureSchema(db: DatabaseSync): void {
FOREIGN KEY(anime_id) REFERENCES imm_anime(anime_id) ON DELETE SET NULL
);
`);
ensureAnimeMergeTables(db);
db.exec(`
CREATE TABLE IF NOT EXISTS imm_sessions(
session_id INTEGER PRIMARY KEY AUTOINCREMENT,
+1 -1
View File
@@ -1,4 +1,4 @@
export const SCHEMA_VERSION = 20;
export const SCHEMA_VERSION = 19;
export const DEFAULT_QUEUE_CAP = 1_000;
export const DEFAULT_BATCH_SIZE = 25;
export const DEFAULT_FLUSH_INTERVAL_MS = 500;
@@ -1,14 +1,13 @@
import type { Hono } from 'hono';
import { statsJson } from '../../../types/stats-http-contract.js';
import { UNKNOWN_MOVE_TARGET_MESSAGE } from '../immersion-tracker/anime-merge.js';
import type { ImmersionTrackerService } from '../immersion-tracker-service.js';
import {
buildSentenceSearchOptions,
enrichSessionsWithKnownWordMetrics,
parseBooleanQuery,
parseDuplicateLineCleanupBody,
parseExcludedWordsBody,
parseIntQuery,
parsePositiveIdList,
} from './route-support.js';
export function registerStatsLibraryRoutes(
@@ -42,6 +41,19 @@ export function registerStatsLibraryRoutes(
return c.json(statsJson('setExcludedWords', { ok: true }));
});
// Collapse animation bursts older versions recorded frame by frame. `dryRun` measures
// the same scan without writing, so the confirmation the user sees is the real cost.
app.post('/api/stats/maintenance/duplicate-lines', async (c) => {
const contentType = c.req.header('content-type')?.split(';', 1)[0]?.trim().toLowerCase();
if (contentType !== 'application/json') return c.body(null, 415);
const body = await c.req.json().catch(() => null);
const options = parseDuplicateLineCleanupBody(body);
if (!options) return c.body(null, 400);
const { dryRun, lookbackDays } = options;
const result = await tracker.cleanupDuplicateSubtitleLines({ dryRun, lookbackDays });
return c.json(statsJson('duplicateLineCleanup', result));
});
app.get('/api/stats/vocabulary/occurrences', async (c) => {
const headword = (c.req.query('headword') ?? '').trim();
const word = (c.req.query('word') ?? '').trim();
@@ -132,19 +144,6 @@ export function registerStatsLibraryRoutes(
return c.json(statsJson('animeLibrary', rows));
});
app.get('/api/stats/anime/merge-recommendations', async (c) => {
const recommendations = await tracker.getAnimeMergeRecommendations();
return c.json(statsJson('animeMergeRecommendations', { recommendations }));
});
app.delete('/api/stats/anime/merge-recommendations/:recommendationId', async (c) => {
const recommendationId = parseIntQuery(c.req.param('recommendationId'), 0);
if (recommendationId <= 0) return c.body(null, 400);
const dismissed = await tracker.dismissAnimeMergeRecommendation(recommendationId);
if (!dismissed) return c.body(null, 404);
return c.json(statsJson('dismissAnimeMergeRecommendation', { ok: true }));
});
app.get('/api/stats/anime/:animeId', async (c) => {
const animeId = parseIntQuery(c.req.param('animeId'), 0);
if (animeId <= 0) return c.body(null, 400);
@@ -212,50 +211,4 @@ export function registerStatsLibraryRoutes(
await tracker.deleteAnime(animeId);
return c.json(statsJson('deleteAnime', { ok: true }));
});
app.post('/api/stats/anime/:animeId/merge', async (c) => {
const animeId = parseIntQuery(c.req.param('animeId'), 0);
if (animeId <= 0) return c.body(null, 400);
const body = await c.req.json().catch(() => null);
const sourceAnimeIds = parsePositiveIdList(body?.sourceAnimeIds).filter((id) => id !== animeId);
if (sourceAnimeIds.length === 0) return c.body(null, 400);
const summary = await tracker.mergeAnime(animeId, sourceAnimeIds);
// Nothing folded means the target or every source was already gone, so the
// caller should not be told the merge succeeded.
if (summary.mergedAnimeIds.length === 0) return c.body(null, 404);
return c.json(
statsJson('mergeAnime', {
ok: true,
animeId: summary.survivingAnimeId,
mergedAnimeIds: summary.mergedAnimeIds,
movedVideos: summary.movedVideos,
}),
);
});
app.patch('/api/stats/media/:videoId/anime', async (c) => {
const videoId = parseIntQuery(c.req.param('videoId'), 0);
if (videoId <= 0) return c.body(null, 400);
const body = await c.req.json().catch(() => null);
const animeId = Number.isSafeInteger(body?.animeId) ? (body.animeId as number) : 0;
if (animeId <= 0) return c.body(null, 400);
try {
const summary = await tracker.moveVideoToAnime(videoId, animeId);
return c.json(
statsJson('moveVideoToAnime', {
ok: true,
animeId: summary.targetAnimeId,
previousAnimeId: summary.previousAnimeId,
removedPreviousAnime: summary.removedPreviousAnime,
}),
);
} catch (error) {
// Only a missing episode or entry is a 404; storage failures must not be
// reported to the caller as "not found".
if (error instanceof Error && error.message === UNKNOWN_MOVE_TARGET_MESSAGE) {
return c.body(null, 404);
}
throw error;
}
});
}
+29 -12
View File
@@ -88,6 +88,35 @@ export function parseExcludedWordsBody(body: unknown): StatsExcludedWord[] | nul
return words;
}
/**
* Read a duplicate-line cleanup request. An explicit object with no lookback scans all
* history. Invalid bodies and invalid windows are rejected instead of broadening scope.
*/
export function parseDuplicateLineCleanupBody(body: unknown): {
dryRun: boolean;
lookbackDays: number | null;
} | null {
if (!body || typeof body !== 'object' || Array.isArray(body)) {
return null;
}
const source = body as Record<string, unknown>;
if (source.dryRun !== undefined && typeof source.dryRun !== 'boolean') {
return null;
}
const rawLookback = source.lookbackDays;
if (
rawLookback !== undefined &&
rawLookback !== null &&
(typeof rawLookback !== 'number' || !Number.isFinite(rawLookback) || rawLookback < 1)
) {
return null;
}
return {
dryRun: source.dryRun === true,
lookbackDays: typeof rawLookback === 'number' ? Math.floor(rawLookback) : null,
};
}
export function loadKnownWordsSet(cachePath: string | undefined): Set<string> | null {
if (!cachePath || !existsSync(cachePath)) return null;
try {
@@ -170,18 +199,6 @@ export async function enrichSessionsWithKnownWordMetrics<
);
}
/** Deduplicated positive integer ids from an untrusted JSON body field. */
export function parsePositiveIdList(raw: unknown): number[] {
if (!Array.isArray(raw)) return [];
const ids = new Set<number>();
for (const value of raw) {
if (Number.isSafeInteger(value) && (value as number) > 0) {
ids.add(value as number);
}
}
return [...ids];
}
export function parseBooleanQuery(raw: string | undefined, fallback: boolean): boolean {
if (raw === undefined) return fallback;
const normalized = raw.trim().toLowerCase();
@@ -0,0 +1,40 @@
/*
* Thresholds that decide when a run of repeated subtitle events is one animation.
*
* Three consumers have to agree on these numbers or the same karaoke line is one cue in
* the sidebar and two hundred in the stats: the file-level cue dedup
* (`subtitle-cue-dedup`), the live gate that decides what immersion stats record
* (`subtitle-line-dedup-gate`), and the retroactive database cleanup
* (`immersion-tracker/duplicate-line-cleanup`).
*/
/**
* Back-to-back frames of the same animation are authored flush against each other; a
* tiny tolerance absorbs the centisecond rounding of the ASS timestamp format.
*/
export const DUPLICATE_CUE_GAP_TOLERANCE_SECONDS = 0.05;
/**
* A burst is a *sequence*. Two adjacent events are two events, not an animation --
* characters do repeat each other, and a repeated line can legitimately be short.
*/
export const MIN_BURST_EVENTS = 3;
/**
* Real dialogue holds on screen for about a second, so a run with a couple of much
* shorter events among them looks like frames. Used only alongside authoring evidence.
*/
export const ANIMATION_FRAME_MAX_SECONDS = 0.3;
/** A karaoke run usually ends on a long "hold" frame, so not every event is short. */
export const MIN_TAGGED_BURST_FRAMES = 2;
/**
* SRT and VTT carry no authoring metadata at all, so timing is the only signal available
* -- which makes it the easiest one to get wrong. ASS->SRT conversion leaves frames at
* ~0.04s, well under any real utterance, and a burst leaves many of them behind. Both
* bounds are deliberately far stricter than the ASS path: a run of ordinary short lines
* (`えっ` traded between characters) must not clear them.
*/
export const TIMING_ONLY_FRAME_MAX_SECONDS = 0.1;
export const MIN_TIMING_ONLY_FRAMES = 5;
+180
View File
@@ -0,0 +1,180 @@
/*
* Duplicate/animation-burst collapsing for parsed subtitle cues.
*
* Split out of the cue parser so the parsing rules and the "is this run one animation?"
* heuristics can be read -- and tested -- on their own. The parser owns the cue shape;
* this module only decides which cues survive.
*/
import { hasAssTemporalOverride, isAnimatedAssEffectKind } from './ass-text';
import {
ANIMATION_FRAME_MAX_SECONDS,
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
MIN_BURST_EVENTS,
MIN_TAGGED_BURST_FRAMES,
MIN_TIMING_ONLY_FRAMES,
TIMING_ONLY_FRAME_MAX_SECONDS,
} from './subtitle-burst-constants';
import type {
AnnotatedSubtitleCue,
SubtitleCue,
SubtitleSourceFormat,
} from './subtitle-cue-parser';
function cueKey(cue: SubtitleCue): string {
return `${cue.startTime}|${cue.endTime}|${cue.text}`;
}
/**
* Identical text over an identical span is redundant however it was authored -- most
* often a layered ASS event stacking a shadow copy under the visible one.
*/
function collapseExactDuplicates(cues: AnnotatedSubtitleCue[]): AnnotatedSubtitleCue[] {
const seen = new Set<string>();
return cues.filter((cue) => {
const key = cueKey(cue);
if (seen.has(key)) {
return false;
}
seen.add(key);
return true;
});
}
function countFramesShorterThan(run: AnnotatedSubtitleCue[], maxSeconds: number): number {
return run.filter((cue) => cue.endTime - cue.startTime < maxSeconds).length;
}
/**
* Evidence that a run of ASS events is one animation rather than several authored lines.
* A static tag says nothing on its own -- three events sharing one `\clip(...)` are three
* signs -- so the tag has to be temporal by nature (`\t`, `\move`, karaoke timing, or
* anything wrapped in `\t(...)`), an animated `Effect` column, or a value that actually
* changes from event to event, which is how per-frame typesetting is authored.
*/
export function hasAssAnimationEvidence(run: AnnotatedSubtitleCue[]): boolean {
if (run.every((cue) => hasAssTemporalOverride(cue.overrides))) {
return true;
}
if (run.every((cue) => isAnimatedAssEffectKind(cue.effectKind))) {
return true;
}
const [first] = run;
const everyEventTypeset = run.every((cue) => cue.overrides.length > 0);
const signatureChanges = run.some((cue) => cue.overrideSignature !== first!.overrideSignature);
return everyEventTypeset && signatureChanges;
}
export function isAnimationBurst(
run: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): boolean {
if (run.length < MIN_BURST_EVENTS) {
return false;
}
if (format === 'srt') {
return (
run.length >= MIN_TIMING_ONLY_FRAMES &&
countFramesShorterThan(run, TIMING_ONLY_FRAME_MAX_SECONDS) === run.length
);
}
if (countFramesShorterThan(run, ANIMATION_FRAME_MAX_SECONDS) < MIN_TAGGED_BURST_FRAMES) {
return false;
}
// One animation belongs to one styled, one named source line. Two characters trading
// the same short word are two styles or two actors, and never merge.
const [first] = run;
if (run.some((cue) => cue.style !== first!.style || cue.name !== first!.name)) {
return false;
}
return hasAssAnimationEvidence(run);
}
/**
* Karaoke and sign typesetting emits one Dialogue event per animation frame, all carrying
* the same visible text over a contiguous span. Collapse each such run into a single cue.
*
* Only runs that look like animation collapse. Two ordinary lines that happen to repeat
* -- several characters each saying `おはよう` in turn, a positioned sign redrawn with a
* different fade -- stay separate, because merging them would destroy real mineable lines.
*/
function collapseAnimationBursts(
cues: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): AnnotatedSubtitleCue[] {
const indicesByText = new Map<string, number[]>();
cues.forEach((cue, index) => {
const bucket = indicesByText.get(cue.text);
if (bucket) {
bucket.push(index);
} else {
indicesByText.set(cue.text, [index]);
}
});
const dropped = new Set<number>();
const extendedEnd = new Map<number, number>();
for (const indices of indicesByText.values()) {
if (indices.length < MIN_BURST_EVENTS) {
continue;
}
let runStart = 0;
while (runStart < indices.length) {
let runEnd = runStart;
let chainEnd = cues[indices[runStart]!]!.endTime;
while (runEnd + 1 < indices.length) {
const next = cues[indices[runEnd + 1]!]!;
if (next.startTime > chainEnd + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS) {
break;
}
chainEnd = Math.max(chainEnd, next.endTime);
runEnd += 1;
}
const run = indices.slice(runStart, runEnd + 1).map((index) => cues[index]!);
if (isAnimationBurst(run, format)) {
for (let i = runStart + 1; i <= runEnd; i += 1) {
dropped.add(indices[i]!);
}
extendedEnd.set(indices[runStart]!, chainEnd);
}
runStart = runEnd + 1;
}
}
if (dropped.size === 0) {
return cues;
}
const merged: AnnotatedSubtitleCue[] = [];
cues.forEach((cue, index) => {
if (dropped.has(index)) {
return;
}
const end = extendedEnd.get(index);
merged.push(end !== undefined && end > cue.endTime ? { ...cue, endTime: end } : cue);
});
return merged;
}
/**
* Collapse redundant cues. Input must already be sorted by non-decreasing `startTime`,
* ties broken by `endTime` then source `order` -- burst detection chains events by
* comparing each one against the running end of the events before it, so an unsorted
* list breaks runs apart and leaves the frames behind.
*/
export function mergeDuplicateCues(
cues: AnnotatedSubtitleCue[],
format: SubtitleSourceFormat,
): AnnotatedSubtitleCue[] {
return collapseAnimationBursts(collapseExactDuplicates(cues), format);
}
+393 -3
View File
@@ -91,6 +91,17 @@ test('parseSrtCues skips malformed timing lines gracefully', () => {
assert.equal(cues[0]!.text, '有効');
});
test('parseSubtitleCues strips complete brace blocks from SRT and VTT text', () => {
const content = ['1', '00:00:01,000 --> 00:00:02,000', '彼は{謎}と言った', ''].join('\n');
for (const filename of ['test.srt', 'test.vtt']) {
const cues = parseSubtitleCues(content, filename);
assert.equal(cues.length, 1, filename);
assert.equal(cues[0]!.text, '彼はと言った', filename);
}
});
test('parseAssCues parses basic ASS dialogue lines', () => {
const content = [
'[Script Info]',
@@ -137,7 +148,9 @@ test('parseAssCues handles text containing commas', () => {
assert.equal(cues[0]!.text, 'はい、そうです、ね');
});
test('parseAssCues handles \\N line breaks', () => {
test('parseAssCues decodes \\N line breaks into real newlines', () => {
// ASS is decoded once, here at ingestion, so cue text matches what mpv hands over for
// the same line played live.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
@@ -146,7 +159,7 @@ test('parseAssCues handles \\N line breaks', () => {
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '一行目\\N二行目');
assert.equal(cues[0]!.text, '一行目\n二行目');
});
test('parseAssCues strips HTML-like markup while preserving ASS line breaks', () => {
@@ -158,7 +171,46 @@ test('parseAssCues strips HTML-like markup while preserving ASS line breaks', ()
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '一行目\\N二行目');
assert.equal(cues[0]!.text, '一行目\n二行目');
});
test('parseAssCues drops vector drawing runs enabled by \\p', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 1,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\an5\\pos(730,1042)\\p1\\blur1}m 20 0 b 10 0 0 10 0 20 b 0 31 10 40 20 40 {\\p0}',
'Dialogue: 0,0:00:05.00,0:00:08.00,Default,,0,0,0,,これは字幕',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues.length, 1);
assert.equal(cues[0]!.text, 'これは字幕');
});
test('parseAssCues keeps text that follows a \\p0 reset on the same line', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\p1}m 0 0 l 10 10{\\p0}本文{\\p1}m 5 5 l 6 6{\\p0}続き',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues.length, 1);
assert.equal(cues[0]!.text, '本文続き');
});
test('parseAssCues leaves \\pos untouched when no drawing mode is active', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,{\\pos(960,1068)\\bord3}位置指定',
].join('\n');
const cues = parseAssCues(content);
assert.equal(cues[0]!.text, '位置指定');
});
test('parseAssCues returns empty for content without Events section', () => {
@@ -258,6 +310,344 @@ test('parseSubtitleCues returns cues sorted by start time', () => {
assert.equal(cues[1]!.text, '二番目');
});
test('parseSubtitleCues collapses per-frame karaoke duplicates into one cue', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}過ぎ去ってしまう瞬間を',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}過ぎ去ってしまう瞬間を',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,,{\\clip(m 3 3)}過ぎ去ってしまう瞬間を',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[0]!.endTime, 3.55);
assert.equal(cues[0]!.text, '過ぎ去ってしまう瞬間を');
});
test('parseSubtitleCues keeps back-to-back plain dialogue repeats separate', () => {
// Several characters greeting in turn: distinct utterances that happen to abut.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:04:05.67,0:04:06.82,Dial_JP,,0,0,0,,おはよう',
'Dialogue: 0,0:04:06.82,0:04:07.56,Dial_JP,,0,0,0,,おはよう',
'Dialogue: 0,0:04:07.56,0:04:08.78,Dial_JP,,0,0,0,,おはよう',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
assert.equal(cues[0]!.endTime, 246.82);
assert.equal(cues[2]!.startTime, 247.56);
});
test('parseSubtitleCues collapses exact duplicate cues even without effect tags', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:04.00,Default,,0,0,0,,重なった行',
'Dialogue: 1,0:00:01.00,0:00:04.00,Default,,0,0,0,,重なった行',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
});
test('parseSubtitleCues collapses tag-less animation frames in converted SRT', () => {
// ASS -> SRT conversion drops override tags, so only the ~0.04s frame timing remains.
const lines = ['1', '00:00:07,870 --> 00:00:07,910', 'Kaguya Wants to be Confessed to', ''];
for (let i = 1; i < 8; i++) {
const start = 7910 + (i - 1) * 40;
const end = start + 40;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(end)}`, 'Kaguya Wants to be Confessed to', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 7.87);
});
test('parseSubtitleCues keeps identical lines that recur far apart', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:02.00,Default,,0,0,0,,なんで',
'Dialogue: 0,0:05:00.00,0:05:01.00,Default,,0,0,0,,なんで',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[1]!.startTime, 300.0);
});
test('parseSubtitleCues keeps two positioned signs that repeat the same text', () => {
// Both carry override tags, but `\pos` and `\fad` are static placement, not animation.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:01:00.00,0:01:03.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(200,200)}第一話',
'Dialogue: 0,0:01:03.00,0:01:06.00,Sign,,0,0,0,,{\\pos(960,900)\\fad(200,200)}第一話',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[1]!.startTime, 63.0);
});
test('parseSubtitleCues keeps a run of ordinary positioned lines separate', () => {
// Three events is a sequence, but none of them runs at animation-frame speed.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:01:00.00,0:01:02.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
'Dialogue: 0,0:01:02.00,0:01:04.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
'Dialogue: 0,0:01:04.00,0:01:06.00,Sign,,0,0,0,,{\\pos(960,120)\\fad(100,100)}止まれ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps a short repeated SRT pair without burst evidence', () => {
const content = [
'1',
'00:00:01,000 --> 00:00:01,200',
'えっ',
'',
'2',
'00:00:01,200 --> 00:00:01,400',
'えっ',
'',
].join('\n');
const cues = parseSubtitleCues(content, 'test.srt');
assert.equal(cues.length, 2);
});
test('parseSubtitleCues collapses a burst marked only by the Effect column', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,Karaoke,歌詞',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,Karaoke,歌詞',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,Karaoke,歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.endTime, 3.55);
});
test('parseSubtitleCues keeps a second karaoke burst that starts after a gap', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}リフレイン',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}リフレイン',
'Dialogue: 0,0:00:01.09,0:00:03.00,OP_JP,,0,0,0,,{\\clip(m 3 3)}リフレイン',
'Dialogue: 0,0:00:20.00,0:00:20.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}リフレイン',
'Dialogue: 0,0:00:20.05,0:00:20.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}リフレイン',
'Dialogue: 0,0:00:20.09,0:00:22.00,OP_JP,,0,0,0,,{\\clip(m 3 3)}リフレイン',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.equal(cues[0]!.endTime, 3.0);
assert.equal(cues[1]!.startTime, 20.0);
assert.equal(cues[1]!.endTime, 22.0);
});
test('parseSubtitleCues does not merge a burst into unrelated dialogue between frames', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,OP_JP,,0,0,0,,{\\clip(m 1 1)}歌詞',
'Dialogue: 0,0:00:01.02,0:00:03.00,Dial_JP,,0,0,0,,別のセリフ',
'Dialogue: 0,0:00:01.05,0:00:01.09,OP_JP,,0,0,0,,{\\clip(m 2 2)}歌詞',
'Dialogue: 0,0:00:01.09,0:00:03.55,OP_JP,,0,0,0,,{\\clip(m 3 3)}歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 2);
assert.deepEqual(
cues.map((cue) => cue.text),
['歌詞', '別のセリフ'],
);
assert.equal(cues[0]!.endTime, 3.55);
});
test('parseSubtitleCues keeps rapid ASS lines from different actors separate', () => {
// Three 200ms `えっ` reactions traded between characters. Fast, adjacent and identical,
// but authored as three lines: different styles and different actors.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_A,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_B,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_C,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues reads the speaker column when it is spelled Actor', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Actor, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_JP,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_JP,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_JP,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues does not treat a custom Effect name as animation', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,scrolling-credit,制作',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,scrolling-credit,制作',
'Dialogue: 0,0:00:01.40,0:00:01.60,Sign,,0,0,0,scrolling-credit,制作',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps rapid ASS lines that share a style but not an actor', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Dial_JP,アリス,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.40,Dial_JP,ボブ,0,0,0,,えっ',
'Dialogue: 0,0:00:01.40,0:00:01.60,Dial_JP,キャロル,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues keeps untagged rapid ASS repeats separate', () => {
// No overrides at all: timing-only evidence is an SRT/VTT fallback and must not apply
// to ASS, where the absence of typesetting is itself evidence of plain dialogue.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.05,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.05,0:00:01.10,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.10,0:00:01.15,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.15,0:00:01.20,Dial_JP,,0,0,0,,えっ',
'Dialogue: 0,0:00:01.20,0:00:01.25,Dial_JP,,0,0,0,,えっ',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 5);
});
test('parseSubtitleCues keeps repeated signs sharing one static clip', () => {
// `\clip` is a static shape for the event. Three events with the identical clip were
// typeset the same way, so none of them is a frame of the others.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
'Dialogue: 0,0:00:01.40,0:00:01.60,Sign,,0,0,0,,{\\clip(0,0,100,100)}注意',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 3);
});
test('parseSubtitleCues collapses a sign animated through \\t', () => {
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
'Dialogue: 0,0:00:01.20,0:00:01.40,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
'Dialogue: 0,0:00:01.40,0:00:03.00,Sign,,0,0,0,,{\\pos(10,10)\\t(0,200,\\frz30)}回る',
].join('\n');
const cues = parseSubtitleCues(content, 'test.ass');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.endTime, 3.0);
});
test('parseSubtitleCues keeps a short repeated SRT run above the frame threshold', () => {
// Five contiguous 200ms cues: a sequence, but nowhere near animation-frame speed.
const lines: string[] = [];
for (let i = 0; i < 5; i++) {
const start = 1000 + i * 200;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(start + 200)}`, 'えっ', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 5);
});
test('parseSubtitleCues keeps a short SRT frame run below the minimum length', () => {
// Four 40ms frames: frame-speed, but too few to tell an animation from an artefact.
const lines: string[] = [];
for (let i = 0; i < 4; i++) {
const start = 7870 + i * 40;
const at = (ms: number) =>
`00:00:0${Math.floor(ms / 1000)},${String(ms % 1000).padStart(3, '0')}`;
lines.push(String(i + 1), `${at(start)} --> ${at(start + 40)}`, 'タイトル', '');
}
const cues = parseSubtitleCues(lines.join('\n'), 'test.srt');
assert.equal(cues.length, 4);
});
test('parseSubtitleCues applies ASS burst rules to ASS content behind an .srt filename', () => {
// The extension lies, so the SRT parser finds nothing and the content-sniffing fallback
// takes over -- which has to carry the `ass` source format with it, or the far stricter
// timing-only thresholds would let this karaoke burst through as three cues.
const content = [
'[Events]',
'Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text',
'Dialogue: 0,0:00:01.00,0:00:01.20,Karaoke,,0,0,0,,{\\k20}歌詞',
'Dialogue: 0,0:00:01.20,0:00:01.40,Karaoke,,0,0,0,,{\\k20}歌詞',
'Dialogue: 0,0:00:01.40,0:00:03.00,Karaoke,,0,0,0,,{\\k20}歌詞',
].join('\n');
const cues = parseSubtitleCues(content, 'test.srt');
assert.equal(cues.length, 1);
assert.equal(cues[0]!.startTime, 1.0);
assert.equal(cues[0]!.endTime, 3.0);
assert.equal(cues[0]!.text, '歌詞');
});
test('parseSubtitleCues detects subtitle formats from remote URLs', () => {
const assContent = [
'[Events]',
+160 -34
View File
@@ -1,9 +1,46 @@
import {
assOverrideSignature,
assToPlainText,
collectAssOverrideCommands,
parseAssEffectField,
type AssEffectKind,
type AssOverrideCommand,
} from './ass-text';
import { mergeDuplicateCues } from './subtitle-cue-dedup';
export interface SubtitleCue {
startTime: number;
endTime: number;
text: string;
}
/**
* Everything the parser knows about a source event, shared only with the dedup engine.
* Deduplication needs the authoring context -- which style the line belongs to, which
* override commands it carries, whether the `Effect` column was set -- to tell a karaoke
* burst apart from two characters saying the same word in turn. None of it is meaningful
* outside the parser, so the public API stays `{startTime, endTime, text}`.
*/
export interface AnnotatedSubtitleCue extends SubtitleCue {
/** Text exactly as authored, override blocks and all. */
rawText: string;
style: string;
layer: number;
/** ASS `Name`/`Actor` column. */
name: string;
/** ASS `Effect` column, verbatim. */
effect: string;
effectKind: AssEffectKind;
/** Override commands found in `{...}` blocks, with their arguments. */
overrides: readonly AssOverrideCommand[];
/** Canonical form of `overrides`, for spotting values that change across a run. */
overrideSignature: string;
/** Position in the source file, so sorting by time stays deterministic across layers. */
order: number;
}
export type SubtitleSourceFormat = 'ass' | 'srt';
const HTML_SUBTITLE_TAG_PATTERN = /<\/?[A-Za-z][^>\n]*>/g;
const SRT_TIMING_PATTERN =
@@ -23,12 +60,21 @@ function parseTimestamp(
);
}
/**
* The single ASS decode for the file path: cues leave the parser as plain text with real
* line breaks, matching what mpv hands over for the same line played live. No layer
* downstream decodes ASS again.
*/
function sanitizeSubtitleCueText(text: string): string {
return text.replace(ASS_OVERRIDE_TAG_PATTERN, '').replace(HTML_SUBTITLE_TAG_PATTERN, '').trim();
return assToPlainText(text, '\n').replace(HTML_SUBTITLE_TAG_PATTERN, '').trim();
}
export function parseSrtCues(content: string): SubtitleCue[] {
const cues: SubtitleCue[] = [];
function toPublicCues(cues: AnnotatedSubtitleCue[]): SubtitleCue[] {
return cues.map(({ startTime, endTime, text }) => ({ startTime, endTime, text }));
}
function parseAnnotatedSrtCues(content: string): AnnotatedSubtitleCue[] {
const cues: AnnotatedSubtitleCue[] = [];
const lines = content.split(/\r?\n/);
let i = 0;
@@ -60,20 +106,39 @@ export function parseSrtCues(content: string): SubtitleCue[] {
i += 1;
}
const text = sanitizeSubtitleCueText(textLines.join('\n'));
const rawText = textLines.join('\n');
const text = sanitizeSubtitleCueText(rawText);
if (text) {
cues.push({ startTime, endTime, text });
cues.push({
startTime,
endTime,
text,
rawText,
style: '',
layer: 0,
name: '',
effect: '',
effectKind: 'none',
// SRT and VTT carry no authoring metadata, and the dedup engine never reads
// overrides for those formats -- collecting them would be parsing for nobody.
overrides: [],
overrideSignature: '',
order: cues.length,
});
}
}
return cues;
}
const ASS_OVERRIDE_TAG_PATTERN = /\{[^}]*\}/g;
export function parseSrtCues(content: string): SubtitleCue[] {
return toPublicCues(parseAnnotatedSrtCues(content));
}
const ASS_TIMING_PATTERN = /^(\d+):(\d{2}):(\d{2})\.(\d{1,2})$/;
const ASS_FORMAT_PREFIX = 'Format:';
const ASS_DIALOGUE_PREFIX = 'Dialogue:';
const ASS_NAME_FIELD_ALIASES = ['name', 'actor'];
function parseAssTimestamp(raw: string): number | null {
const match = ASS_TIMING_PATTERN.exec(raw.trim());
@@ -87,13 +152,43 @@ function parseAssTimestamp(raw: string): number | null {
return hours * 3600 + minutes * 60 + seconds + centiseconds / 100;
}
export function parseAssCues(content: string): SubtitleCue[] {
const cues: SubtitleCue[] = [];
function readField(fields: string[], index: number): string {
return index >= 0 && index < fields.length ? fields[index]!.trim() : '';
}
function findFieldIndex(formatFields: string[], aliases: string[]): number {
for (const alias of aliases) {
const index = formatFields.indexOf(alias);
if (index >= 0) {
return index;
}
}
return -1;
}
function parseAnnotatedAssCues(content: string): AnnotatedSubtitleCue[] {
const cues: AnnotatedSubtitleCue[] = [];
const lines = content.split(/\r?\n/);
let inEventsSection = false;
let startFieldIndex = -1;
let endFieldIndex = -1;
let textFieldIndex = -1;
const fieldIndex = {
start: -1,
end: -1,
text: -1,
style: -1,
layer: -1,
name: -1,
effect: -1,
};
const resetFieldIndex = () => {
fieldIndex.start = -1;
fieldIndex.end = -1;
fieldIndex.text = -1;
fieldIndex.style = -1;
fieldIndex.layer = -1;
fieldIndex.name = -1;
fieldIndex.effect = -1;
};
for (const line of lines) {
const trimmed = line.trim();
@@ -101,9 +196,7 @@ export function parseAssCues(content: string): SubtitleCue[] {
if (trimmed.startsWith('[') && trimmed.endsWith(']')) {
inEventsSection = trimmed.toLowerCase() === '[events]';
if (!inEventsSection) {
startFieldIndex = -1;
endFieldIndex = -1;
textFieldIndex = -1;
resetFieldIndex();
}
continue;
}
@@ -117,9 +210,15 @@ export function parseAssCues(content: string): SubtitleCue[] {
.slice(ASS_FORMAT_PREFIX.length)
.split(',')
.map((field) => field.trim().toLowerCase());
startFieldIndex = formatFields.indexOf('start');
endFieldIndex = formatFields.indexOf('end');
textFieldIndex = formatFields.indexOf('text');
fieldIndex.start = formatFields.indexOf('start');
fieldIndex.end = formatFields.indexOf('end');
fieldIndex.text = formatFields.indexOf('text');
fieldIndex.style = formatFields.indexOf('style');
fieldIndex.layer = formatFields.indexOf('layer');
// Aegisub writes the speaker column as `Actor`; the v4+ spec calls it `Name`.
// Missing it costs the burst check its speaker guard, so both spellings count.
fieldIndex.name = findFieldIndex(formatFields, ASS_NAME_FIELD_ALIASES);
fieldIndex.effect = formatFields.indexOf('effect');
continue;
}
@@ -127,34 +226,57 @@ export function parseAssCues(content: string): SubtitleCue[] {
continue;
}
if (startFieldIndex < 0 || endFieldIndex < 0 || textFieldIndex < 0) {
if (fieldIndex.start < 0 || fieldIndex.end < 0 || fieldIndex.text < 0) {
continue;
}
const fields = trimmed.slice(ASS_DIALOGUE_PREFIX.length).split(',');
if (
startFieldIndex >= fields.length ||
endFieldIndex >= fields.length ||
textFieldIndex >= fields.length
fieldIndex.start >= fields.length ||
fieldIndex.end >= fields.length ||
fieldIndex.text >= fields.length
) {
continue;
}
const startTime = parseAssTimestamp(fields[startFieldIndex]!);
const endTime = parseAssTimestamp(fields[endFieldIndex]!);
const startTime = parseAssTimestamp(fields[fieldIndex.start]!);
const endTime = parseAssTimestamp(fields[fieldIndex.end]!);
if (startTime === null || endTime === null) {
continue;
}
const text = sanitizeSubtitleCueText(fields.slice(textFieldIndex).join(','));
if (text) {
cues.push({ startTime, endTime, text });
const rawText = fields.slice(fieldIndex.text).join(',');
const text = sanitizeSubtitleCueText(rawText);
if (!text) {
continue;
}
const effect = readField(fields, fieldIndex.effect);
const layer = Number(readField(fields, fieldIndex.layer));
const overrides = collectAssOverrideCommands(rawText);
cues.push({
startTime,
endTime,
text,
rawText,
style: readField(fields, fieldIndex.style),
layer: Number.isFinite(layer) ? layer : 0,
name: readField(fields, fieldIndex.name),
effect,
effectKind: parseAssEffectField(effect),
overrides,
overrideSignature: assOverrideSignature(overrides),
order: cues.length,
});
}
return cues;
}
export function parseAssCues(content: string): SubtitleCue[] {
return toPublicCues(parseAnnotatedAssCues(content));
}
function detectSubtitleFormat(source: string): 'srt' | 'vtt' | 'ass' | 'ssa' | null {
const [normalizedSource = source] =
(() => {
@@ -173,27 +295,31 @@ function detectSubtitleFormat(source: string): 'srt' | 'vtt' | 'ass' | 'ssa' | n
export function parseSubtitleCues(content: string, filename: string): SubtitleCue[] {
const format = detectSubtitleFormat(filename);
let cues: SubtitleCue[];
let cues: AnnotatedSubtitleCue[];
let sourceFormat: SubtitleSourceFormat = 'srt';
switch (format) {
case 'srt':
case 'vtt':
cues = parseSrtCues(content);
cues = parseAnnotatedSrtCues(content);
break;
case 'ass':
case 'ssa':
cues = parseAssCues(content);
cues = parseAnnotatedAssCues(content);
sourceFormat = 'ass';
break;
default:
cues = [];
}
if (cues.length === 0) {
const assCues = parseAssCues(content);
const srtCues = parseSrtCues(content);
cues = assCues.length >= srtCues.length ? assCues : srtCues;
const assCues = parseAnnotatedAssCues(content);
const srtCues = parseAnnotatedSrtCues(content);
const preferAss = assCues.length >= srtCues.length;
cues = preferAss ? assCues : srtCues;
sourceFormat = preferAss && assCues.length > 0 ? 'ass' : 'srt';
}
cues.sort((a, b) => a.startTime - b.startTime);
return cues;
cues.sort((a, b) => a.startTime - b.startTime || a.endTime - b.endTime || a.order - b.order);
return toPublicCues(mergeDuplicateCues(cues, sourceFormat));
}
@@ -0,0 +1,169 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { createSubtitleLineDedupGate } from './subtitle-line-dedup-gate';
import type { SubtitleCue } from '../../types';
function karaokeFrames(text: string, start: number, frames: number, frameSeconds: number) {
return Array.from({ length: frames }, (_, index) => ({
text,
startSec: start + index * frameSeconds,
endSec: start + (index + 1) * frameSeconds,
}));
}
test('parsed cues drop the frames the sidebar already collapsed', () => {
// What `mergeDuplicateCues` leaves behind for a karaoke run: one cue over the run.
const cues: SubtitleCue[] = [
{ startTime: 10, endTime: 14, text: '飛び上がる' },
{ startTime: 14, endTime: 16, text: 'もしも' },
];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
const recorded = karaokeFrames('飛び上がる', 10, 40, 0.04).filter((sample) =>
gate.shouldRecord(sample),
);
assert.equal(recorded.length, 1);
assert.equal(recorded[0]!.startSec, 10);
assert.equal(gate.shouldRecord({ text: 'もしも', startSec: 14, endSec: 16 }), true);
});
test('parsed cues keep separate lines that merely repeat', () => {
const cues: SubtitleCue[] = [
{ startTime: 3, endTime: 3.4, text: 'えっ' },
{ startTime: 3.4, endTime: 3.9, text: 'えっ' },
{ startTime: 3.9, endTime: 4.5, text: 'えっ' },
];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
const recorded = cues.filter((cue) =>
gate.shouldRecord({ text: cue.text, startSec: cue.startTime, endSec: cue.endTime }),
);
assert.equal(recorded.length, 3);
});
test('parsed cues outrank the streaming heuristic for short repeated cues', () => {
// Long enough to trip the timing-only rule, but the parser saw these with full
// lookahead and kept them, so every one of them is a line the sidebar shows.
const cues: SubtitleCue[] = Array.from({ length: 8 }, (_, index) => ({
startTime: 3 + index * 0.08,
endTime: 3 + (index + 1) * 0.08,
text: 'えっ',
}));
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
const recorded = cues.filter((cue) =>
gate.shouldRecord({ text: cue.text, startSec: cue.startTime, endSec: cue.endTime }),
);
assert.equal(recorded.length, 8);
});
test('parsed cues preserve legitimately separate cues only 40ms apart', () => {
const cues: SubtitleCue[] = Array.from({ length: 8 }, (_, index) => ({
startTime: 3 + index * 0.04,
endTime: 3 + (index + 1) * 0.04,
text: 'えっ',
}));
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
const recorded = cues.filter((cue) =>
gate.shouldRecord({ text: cue.text, startSec: cue.startTime, endSec: cue.endTime }),
);
assert.equal(recorded.length, 8);
});
test('a line whose timing does not match any cue still records', () => {
// A shifted track, an embedded sub nobody parsed: no match, no drop.
const cues: SubtitleCue[] = [{ startTime: 10, endTime: 14, text: '飛び上がる' }];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 42, endSec: 44 }), true);
});
test('shifted parsed text falls back to streaming burst detection', () => {
const cues: SubtitleCue[] = [{ startTime: 10, endTime: 14, text: '飛び上がる' }];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
const recorded = karaokeFrames('飛び上がる', 42, 40, 0.04).filter((sample) =>
gate.shouldRecord(sample),
);
assert.equal(recorded.length, 4);
});
test('replacing the parsed cue source forgets a streaming run', () => {
let cues: SubtitleCue[] = [];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
karaokeFrames('飛び上がる', 42, 20, 0.04).forEach((sample) => gate.shouldRecord(sample));
cues = [];
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 42.8, endSec: 42.84 }), true);
});
test('without parsed cues a long run of identical short frames stops recording', () => {
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
const recorded = karaokeFrames('ひとしずく', 0, 200, 0.04).filter((sample) =>
gate.shouldRecord(sample),
);
assert.equal(recorded.length, 4);
});
test('without parsed cues ordinary repeated dialogue keeps recording', () => {
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
// Six contiguous `えっ`, each held for a normal beat rather than an animation frame.
const recorded = karaokeFrames('えっ', 0, 6, 0.6).filter((sample) => gate.shouldRecord(sample));
assert.equal(recorded.length, 6);
});
test('the same event offered twice does not advance the run', () => {
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
// mpv fires the timing handler once for `sub-start` and once for `sub-end`.
for (let i = 0; i < 8; i += 1) {
assert.equal(gate.shouldRecord({ text: '待って', startSec: 5, endSec: 5.05 }), true);
}
});
test('a gap between frames starts a new run', () => {
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
const first = karaokeFrames('もし', 0, 6, 0.04).filter((sample) => gate.shouldRecord(sample));
const second = karaokeFrames('もし', 30, 6, 0.04).filter((sample) => gate.shouldRecord(sample));
assert.equal(first.length, 4);
assert.equal(second.length, 4);
});
test('reset forgets the streaming run', () => {
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
karaokeFrames('もし', 0, 20, 0.04).forEach((sample) => gate.shouldRecord(sample));
gate.reset();
assert.equal(gate.shouldRecord({ text: 'もし', startSec: 0.8, endSec: 0.84 }), true);
});
test('reset ignores stale parsed cues until the source publishes a new cue list', () => {
let cues: SubtitleCue[] = [{ startTime: 10, endTime: 14, text: '飛び上がる' }];
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 10, endSec: 10.04 }), true);
gate.reset();
const recordedWithStaleCues = karaokeFrames('飛び上がる', 10.04, 8, 0.04).filter((sample) =>
gate.shouldRecord(sample),
);
assert.equal(recordedWithStaleCues.length, 4);
cues = [{ startTime: 20, endTime: 24, text: '飛び上がる' }];
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 20, endSec: 20.04 }), true);
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 20.04, endSec: 20.08 }), false);
});
@@ -0,0 +1,208 @@
/*
* Decides which live mpv subtitle lines reach the immersion stats.
*
* The sidebar reads a parsed subtitle file, so it can collapse an animation burst with
* full lookahead (`subtitle-cue-dedup`). Stats are fed from mpv's `sub-start`/`sub-end`
* properties instead -- one event per animation frame, each with its own start time --
* so without a gate a karaoke OP counts its lyrics once per frame and buries every real
* word in the vocabulary charts.
*
* Two layers, in order:
*
* 1. When the active source has been parsed, its cue list has *already* been collapsed.
* A live line that lands inside a surviving cue of the same text, but after that
* cue's start, is a frame the sidebar merged away, so stats drop it too. This is the
* layer that keeps the two views consistent by construction.
* 2. Otherwise (embedded track nobody parsed, a source whose timings mpv has shifted)
* fall back to timing alone. No authoring metadata is available live -- mpv delivers
* `sub-text-ass` after `sub-start`/`sub-end`, so any ASS text read here belongs to the
* previous event -- which puts this layer in the same position as the SRT path in
* `subtitle-cue-dedup`, and it uses that path's deliberately strict bounds.
*/
import { normalizePlainSubtitleText } from './ass-text';
import {
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
MIN_TIMING_ONLY_FRAMES,
TIMING_ONLY_FRAME_MAX_SECONDS,
} from './subtitle-burst-constants';
import type { SubtitleCue } from './subtitle-cue-parser';
export interface SubtitleLineSample {
text: string;
startSec: number;
endSec: number;
}
export interface SubtitleLineDedupGateDeps {
/** Cues for the active source, already collapsed by the parser. */
getParsedCues: () => readonly SubtitleCue[] | null | undefined;
}
export interface SubtitleLineDedupGate {
/** False when this line is an animation frame of a line already recorded. */
shouldRecord: (sample: SubtitleLineSample) => boolean;
/** Forget run state and ignore the current cue list until its source is replaced. */
reset: () => void;
}
interface CueSpan {
startTime: number;
endTime: number;
}
interface StreamingRunState {
text: string;
startMs: number;
chainEndSec: number;
/** Contiguous identical short frames seen so far, including the recorded first one. */
frames: number;
}
/** Exact cue identity, separate from the looser tolerance used to chain adjacent frames. */
const CUE_START_IDENTITY_TOLERANCE_SECONDS = 0.005;
function normalizeLineText(text: string): string {
return normalizePlainSubtitleText(text, { collapseLineBreaks: true });
}
function buildSpansByText(cues: readonly SubtitleCue[]): Map<string, CueSpan[]> {
const spansByText = new Map<string, CueSpan[]>();
for (const cue of cues) {
const key = normalizeLineText(cue.text);
if (!key) continue;
const span = { startTime: cue.startTime, endTime: cue.endTime };
const existing = spansByText.get(key);
if (existing) {
existing.push(span);
} else {
spansByText.set(key, [span]);
}
}
return spansByText;
}
/**
* A frame the parser merged away: the same text, starting inside a surviving cue but
* after it began.
*
* Starting a cue always wins over falling inside one. The first frame of a collapsed run
* starts *at* the merged cue, and a line the parser deliberately kept separate -- three
* characters trading `えっ` back to back -- begins exactly where the one before it ends.
*/
function isMergedAwayFrame(spans: readonly CueSpan[], startSec: number): boolean | null {
const coveringSpans = spans.filter(
(span) =>
startSec >= span.startTime - CUE_START_IDENTITY_TOLERANCE_SECONDS &&
startSec <= span.endTime + CUE_START_IDENTITY_TOLERANCE_SECONDS,
);
if (coveringSpans.length === 0) {
return null;
}
const startsOwnCue = spans.some(
(span) => Math.abs(startSec - span.startTime) <= CUE_START_IDENTITY_TOLERANCE_SECONDS,
);
if (startsOwnCue) {
return false;
}
return coveringSpans.some(
(span) =>
startSec > span.startTime + CUE_START_IDENTITY_TOLERANCE_SECONDS &&
startSec <= span.endTime + CUE_START_IDENTITY_TOLERANCE_SECONDS,
);
}
export function createSubtitleLineDedupGate(
deps: SubtitleLineDedupGateDeps,
): SubtitleLineDedupGate {
let indexedCues: readonly SubtitleCue[] | null | undefined;
let ignoredCuesAfterReset: readonly SubtitleCue[] | null | undefined;
let spansByText: Map<string, CueSpan[]> = new Map();
let run: StreamingRunState | null = null;
const lookupSpans = (text: string): CueSpan[] | null => {
const cues = deps.getParsedCues() ?? null;
if (ignoredCuesAfterReset !== undefined) {
if (cues === ignoredCuesAfterReset) {
return null;
}
ignoredCuesAfterReset = undefined;
}
if (cues !== indexedCues) {
indexedCues = cues;
spansByText = cues?.length ? buildSpansByText(cues) : new Map();
run = null;
}
return spansByText.get(text) ?? null;
};
/**
* Timing-only burst detection over a stream. Without lookahead the run can only be
* recognised from the inside, so the first frames of a burst are recorded and the rest
* dropped -- an OP costs a handful of counted lines instead of several hundred.
*/
const advanceStreamingRun = (text: string, sample: SubtitleLineSample): boolean => {
const startMs = Math.round(sample.startSec * 1000);
// mpv reports `sub-start` and `sub-end` separately, so one event can be offered
// twice. The same start is the same frame, never the next one in a run.
if (run && run.text === text && run.startMs === startMs) {
run.chainEndSec = Math.max(run.chainEndSec, sample.endSec);
return run.frames < MIN_TIMING_ONLY_FRAMES;
}
const isShortFrame = sample.endSec - sample.startSec < TIMING_ONLY_FRAME_MAX_SECONDS;
// Frames are authored flush against each other, but typesetters do overlap them, so
// the chain only requires forward progress that stays inside the running end.
const continuesRun =
run !== null &&
run.text === text &&
isShortFrame &&
startMs > run.startMs &&
sample.startSec <= run.chainEndSec + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS;
if (continuesRun && run) {
run.startMs = startMs;
run.chainEndSec = Math.max(run.chainEndSec, sample.endSec);
run.frames += 1;
} else {
run = {
text,
startMs,
chainEndSec: sample.endSec,
frames: isShortFrame ? 1 : 0,
};
}
return run.frames < MIN_TIMING_ONLY_FRAMES;
};
return {
shouldRecord: (sample) => {
const text = normalizeLineText(sample.text);
if (!text) {
return true;
}
// The parsed cue list has the final say wherever it covers this line. Falling
// through to the streaming heuristic would let it drop cues the parser looked at
// with full lookahead and deliberately kept apart, which is the disagreement
// between sidebar and stats this gate exists to prevent.
const spans = lookupSpans(text);
if (spans) {
const mergedAway = isMergedAwayFrame(spans, sample.startSec);
if (mergedAway !== null) {
run = null;
return !mergedAway;
}
}
return advanceStreamingRun(text, sample);
},
reset: () => {
run = null;
ignoredCuesAfterReset = deps.getParsedCues() ?? null;
indexedCues = undefined;
spansByText = new Map();
},
};
}
@@ -115,6 +115,21 @@ test('subtitle processing does not emit plain payload for cached lines', async (
assert.deepEqual(emitted, [{ text: '字幕', tokens: [] }]);
});
test('text that normalizes to nothing is never cached', () => {
const controller = createSubtitleProcessingController({
tokenizeSubtitle: async (text) => ({ text, tokens: [] }),
emitSubtitle: () => {},
});
// Two different inputs both reduce to an empty key; sharing one entry would serve the
// first one's tokens for the second.
controller.preCacheTokenization(' ', { text: ' ', tokens: [] });
assert.equal(controller.hasCachedSubtitle(' '), false);
assert.equal(controller.hasCachedSubtitle('\\n'), false);
assert.equal(controller.consumeCachedSubtitle('\\n'), null);
});
test('subtitle processing shows plain line while tokenization is still pending', async () => {
const emitted: SubtitleData[] = [];
let resolveTokenization: ((value: SubtitleData | null) => void) | undefined;
@@ -1,4 +1,5 @@
import type { SubtitleData } from '../../types';
import { normalizePlainSubtitleText } from './ass-text';
export interface SubtitleProcessingControllerDeps {
tokenizeSubtitle: (text: string) => Promise<SubtitleData | null>;
@@ -47,8 +48,15 @@ export interface SubtitleProcessingController {
hasCachedSubtitle: (text: string) => boolean;
}
/**
* Prefetched cues and live mpv text are both already decoded from ASS, so the key only
* has to settle whitespace for one authored line to resolve to one entry.
*
* An empty key is not a line: it is whatever normalization reduced to nothing. Callers
* must skip the cache for it rather than let every such input share one entry.
*/
export function normalizeSubtitleCacheKey(text: string): string {
return text.replace(/\r\n/g, '\n').replace(/\\N/g, '\n').replace(/\\n/g, '\n').trim();
return normalizePlainSubtitleText(text);
}
export function createSubtitleProcessingController(
@@ -72,6 +80,9 @@ export function createSubtitleProcessingController(
const getCachedTokenization = (text: string): SubtitleData | null => {
const cacheKey = normalizeSubtitleCacheKey(text);
if (!cacheKey) {
return null;
}
const cached = tokenizationCache.get(cacheKey);
if (!cached) {
return null;
@@ -83,7 +94,11 @@ export function createSubtitleProcessingController(
};
const setCachedTokenization = (text: string, payload: SubtitleData): void => {
tokenizationCache.set(normalizeSubtitleCacheKey(text), payload);
const cacheKey = normalizeSubtitleCacheKey(text);
if (!cacheKey) {
return;
}
tokenizationCache.set(cacheKey, payload);
while (tokenizationCache.size > SUBTITLE_TOKENIZATION_CACHE_LIMIT) {
const firstKey = tokenizationCache.keys().next().value;
if (firstKey !== undefined) {
@@ -254,7 +269,8 @@ export function createSubtitleProcessingController(
return cached;
},
hasCachedSubtitle: (text: string) => {
return tokenizationCache.has(normalizeSubtitleCacheKey(text));
const cacheKey = normalizeSubtitleCacheKey(text);
return cacheKey.length > 0 && tokenizationCache.has(cacheKey);
},
};
}
+4 -2
View File
@@ -1651,9 +1651,11 @@ test('tokenizeSubtitle clears JLPT level from standalone Yomitan particle token'
assert.equal(result.tokens?.[0]?.jlptLevel, undefined);
});
test('tokenizeSubtitle returns null tokens for empty normalized text', async () => {
test('tokenizeSubtitle returns the normalized text when it comes out empty', async () => {
// Handing back the original would push whatever normalization dropped into app state
// as if it were subtitle text.
const result = await tokenizeSubtitle(' \\n ', makeDeps());
assert.deepEqual(result, { text: ' \\n ', tokens: null });
assert.deepEqual(result, { text: '', tokens: null });
});
test('tokenizeSubtitle normalizes newlines before Yomitan parse request', async () => {
+7 -6
View File
@@ -27,6 +27,7 @@ import {
} from './tokenizer/yomitan-parser-runtime';
import type { YomitanTermFrequency } from './tokenizer/yomitan-parser-runtime';
import { isKanaChar } from './tokenizer/token-classification';
import { normalizePlainSubtitleText } from './ass-text';
const logger = createLogger('main:tokenizer');
@@ -886,14 +887,14 @@ export async function tokenizeSubtitle(
text: string,
deps: TokenizerServiceDeps,
): Promise<SubtitleData> {
const displayText = text
.replace(/\r\n/g, '\n')
.replace(/\\N/g, '\n')
.replace(/\\n/g, '\n')
.trim();
const displayText = normalizePlainSubtitleText(text);
// ASS decoding already happened upstream (cue parser for files, mpv for live text), so
// all this drops is whitespace -- but a whitespace-only line still normalizes to empty.
// Return the normalized form anyway: handing back the original would put a blank line
// into application state as if it were subtitle text.
if (!displayText) {
return { text, tokens: null };
return { text: displayText, tokens: null };
}
const tokenizeText = displayText
@@ -1,7 +0,0 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { normalizeTitleIdentity } from './title-normalization';
test('normalizeTitleIdentity produces a Unicode-aware comparison key', () => {
assert.equal(normalizeTitleIdentity(' BOCCHI・The ROCK!! '), 'bocchi the rock');
});
-8
View File
@@ -1,8 +0,0 @@
export function normalizeTitleIdentity(title: string): string {
return title
.normalize('NFKC')
.toLowerCase()
.replace(/[^\p{L}\p{N}]+/gu, ' ')
.trim()
.replace(/\s+/g, ' ');
}
@@ -267,3 +267,123 @@ test('flushPlaybackPositionOnMediaPathClear ignores disconnected mpv time-pos re
assert.deepEqual(recorded, [42]);
});
test('media and subtitle-track transitions reset live subtitle-line deduplication', () => {
const recordedStarts: number[] = [];
const handlers = createBuildBindMpvMainEventHandlersMainDepsHandler({
appState: {
initialArgs: null,
overlayRuntimeInitialized: true,
mpvClient: null,
immersionTracker: {
recordSubtitleLine: (_text: string, start: number) => recordedStarts.push(start),
},
subtitleTimingTracker: null,
activeParsedSubtitleCues: null,
currentMediaPath: '/video-a.mkv',
currentSubText: '',
currentSubAssText: '',
playbackPaused: null,
previousSecondarySubVisibility: false,
},
getQuitOnDisconnectArmed: () => false,
scheduleQuitCheck: () => {},
quitApp: () => {},
reportJellyfinRemoteStopped: () => {},
syncOverlayMpvSubtitleSuppression: () => {},
maybeRunAnilistPostWatchUpdate: async () => {},
logSubtitleTimingError: () => {},
broadcastToOverlayWindows: () => {},
onSubtitleChange: () => {},
ensureImmersionTrackerInitialized: () => {},
updateCurrentMediaPath: () => {},
restoreMpvSubVisibility: () => {},
resetSubtitleSidebarEmbeddedLayout: () => {},
getCurrentAnilistMediaKey: () => null,
resetAnilistMediaTracking: () => {},
maybeProbeAnilistDuration: () => {},
ensureAnilistMediaGuess: () => {},
syncImmersionMediaState: () => {},
updateCurrentMediaTitle: () => {},
resetAnilistMediaGuessState: () => {},
reportJellyfinRemoteProgress: () => {},
updateSubtitleRenderMetrics: () => {},
refreshDiscordPresence: () => {},
})();
for (let index = 0; index < 8; index += 1) {
handlers.recordImmersionSubtitleLine('待って', index * 0.04, (index + 1) * 0.04);
}
assert.equal(recordedStarts.length, 4);
handlers.updateCurrentMediaPath('/video-b.mkv');
handlers.recordImmersionSubtitleLine('待って', 0.32, 0.36);
assert.equal(recordedStarts.length, 5);
for (let index = 9; index < 16; index += 1) {
handlers.recordImmersionSubtitleLine('待って', index * 0.04, (index + 1) * 0.04);
}
assert.equal(recordedStarts.length, 8);
assert.equal(typeof handlers.onSubtitleTrackChange, 'function');
handlers.onSubtitleTrackChange?.(2);
handlers.recordImmersionSubtitleLine('待って', 0.64, 0.68);
assert.equal(recordedStarts.length, 9);
});
test('subtitle-track transitions ignore stale parsed cues until replacement cues arrive', () => {
const recordedStarts: number[] = [];
const appState = {
initialArgs: null,
overlayRuntimeInitialized: true,
mpvClient: null,
immersionTracker: {
recordSubtitleLine: (_text: string, start: number) => recordedStarts.push(start),
},
subtitleTimingTracker: null,
activeParsedSubtitleCues: [{ startTime: 10, endTime: 14, text: '飛び上がる' }],
currentMediaPath: '/video-a.mkv',
currentSubText: '',
currentSubAssText: '',
playbackPaused: null,
previousSecondarySubVisibility: false,
};
const handlers = createBuildBindMpvMainEventHandlersMainDepsHandler({
appState,
getQuitOnDisconnectArmed: () => false,
scheduleQuitCheck: () => {},
quitApp: () => {},
reportJellyfinRemoteStopped: () => {},
syncOverlayMpvSubtitleSuppression: () => {},
maybeRunAnilistPostWatchUpdate: async () => {},
logSubtitleTimingError: () => {},
broadcastToOverlayWindows: () => {},
onSubtitleChange: () => {},
ensureImmersionTrackerInitialized: () => {},
updateCurrentMediaPath: () => {},
restoreMpvSubVisibility: () => {},
resetSubtitleSidebarEmbeddedLayout: () => {},
getCurrentAnilistMediaKey: () => null,
resetAnilistMediaTracking: () => {},
maybeProbeAnilistDuration: () => {},
ensureAnilistMediaGuess: () => {},
syncImmersionMediaState: () => {},
updateCurrentMediaTitle: () => {},
resetAnilistMediaGuessState: () => {},
reportJellyfinRemoteProgress: () => {},
updateSubtitleRenderMetrics: () => {},
refreshDiscordPresence: () => {},
})();
handlers.recordImmersionSubtitleLine('飛び上がる', 10, 10.04);
handlers.onSubtitleTrackChange?.(2);
for (let index = 1; index <= 8; index += 1) {
handlers.recordImmersionSubtitleLine('飛び上がる', 10 + index * 0.04, 10 + (index + 1) * 0.04);
}
assert.equal(recordedStarts.length, 5);
appState.activeParsedSubtitleCues = [{ startTime: 20, endTime: 24, text: '飛び上がる' }];
handlers.recordImmersionSubtitleLine('飛び上がる', 20, 20.04);
handlers.recordImmersionSubtitleLine('飛び上がる', 20.04, 20.08);
assert.deepEqual(recordedStarts.slice(-1), [20]);
});
+19 -5
View File
@@ -1,4 +1,5 @@
import type { MergedToken, SubtitleData } from '../../types';
import { createSubtitleLineDedupGate } from '../../core/services/subtitle-line-dedup-gate';
import type { MergedToken, SubtitleCue, SubtitleData } from '../../types';
type AnilistPostWatchRunOptions = {
watchedSeconds?: number;
@@ -34,6 +35,7 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
subtitleTimingTracker: {
recordSubtitle?: (text: string, start: number, end: number, secondaryText?: string) => void;
} | null;
activeParsedSubtitleCues?: SubtitleCue[] | null;
currentMediaPath?: string | null;
currentSubText: string;
currentSubAssText: string;
@@ -86,6 +88,11 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
deps.ensureImmersionTrackerInitialized();
deps.appState.immersionTracker?.recordPlaybackPosition?.(normalizedTimeSec);
};
// mpv reports every animation frame of a typeset line as its own subtitle event, so
// stats have to collapse bursts the same way the parsed cue list already does.
const immersionLineDedupGate = createSubtitleLineDedupGate({
getParsedCues: () => deps.appState.activeParsedSubtitleCues,
});
const hasInitialPlaybackQuitOnDisconnectArg = (): boolean =>
Boolean(
deps.appState.initialArgs?.managedPlayback ||
@@ -110,6 +117,9 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
if (!tracker?.recordSubtitleLine) {
return;
}
if (!immersionLineDedupGate.shouldRecord({ text, startSec: start, endSec: end })) {
return;
}
const secondaryText = deps.appState.mpvClient?.currentSecondarySubText || null;
const cachedTokens =
deps.appState.currentSubtitleData?.text === text
@@ -159,9 +169,10 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
logSubtitleProcessingDebug: deps.logSubtitleProcessingDebug
? (message: string) => deps.logSubtitleProcessingDebug!(message)
: undefined,
onSubtitleTrackChange: deps.onSubtitleTrackChange
? (sid: number | null) => deps.onSubtitleTrackChange!(sid)
: undefined,
onSubtitleTrackChange: (sid: number | null) => {
immersionLineDedupGate.reset();
deps.onSubtitleTrackChange?.(sid);
},
onSubtitleTrackListChange: deps.onSubtitleTrackListChange
? (trackList: unknown[] | null) => deps.onSubtitleTrackListChange!(trackList)
: undefined,
@@ -173,7 +184,10 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
deps.broadcastToOverlayWindows('subtitle-ass:set', text),
broadcastSecondarySubtitle: (text: string) =>
deps.broadcastToOverlayWindows('secondary-subtitle:set', text),
updateCurrentMediaPath: (path: string) => deps.updateCurrentMediaPath(path),
updateCurrentMediaPath: (path: string) => {
immersionLineDedupGate.reset();
deps.updateCurrentMediaPath(path);
},
restoreMpvSubVisibility: () => deps.restoreMpvSubVisibility(),
resetSubtitleSidebarEmbeddedLayout: () => deps.resetSubtitleSidebarEmbeddedLayout?.(),
getCurrentAnilistMediaKey: () => deps.getCurrentAnilistMediaKey(),
@@ -200,6 +200,59 @@ test('stats cli command fails when immersion tracking is disabled', async () =>
]);
});
test('stats cli command runs a duplicate-line cleanup preview without touching the dashboard', async () => {
const { handler, calls, responses } = makeHandler({
getImmersionTracker: () => ({
cleanupDuplicateSubtitleLines: async (options: {
dryRun?: boolean;
lookbackDays?: number | null;
}) => ({
dryRun: options.dryRun === true,
lookbackDays: options.lookbackDays ?? null,
scannedLines: 900,
burstGroups: 2,
removedLines: 180,
removedWordOccurrences: 540,
removedKanjiOccurrences: 120,
samples: [
{
videoId: 7,
videoTitle: 'Ep 1',
text: '飛び上がる',
frames: 90,
removedLines: 89,
startMs: 1000,
endMs: 5000,
},
],
}),
}),
});
await handler(
{
statsResponsePath: '/tmp/subminer-stats-response.json',
statsCleanup: true,
statsCleanupDuplicateLines: true,
statsCleanupDryRun: true,
statsCleanupLookbackDays: 30,
},
'initial',
);
assert.deepEqual(calls, [
'ensureImmersionTrackerStarted',
'info:Stats duplicate-line cleanup preview (last 30d): scanned=900 bursts=2 removedLines=180 removedWordCounts=540 removedKanjiCounts=120',
'info: Ep 1: "飛び上がる" x90',
]);
assert.deepEqual(responses, [
{
responsePath: '/tmp/subminer-stats-response.json',
payload: { ok: true },
},
]);
});
test('stats cli command runs vocab cleanup instead of opening dashboard when cleanup mode is requested', async () => {
const { handler, calls, responses } = makeHandler({
getImmersionTracker: () => ({
+30
View File
@@ -1,6 +1,7 @@
import fs from 'node:fs';
import path from 'node:path';
import type { CliArgs, CliCommandSource } from '../../cli/args';
import type { DuplicateSubtitleLineCleanupSummary } from '../../core/services/immersion-tracker/duplicate-line-cleanup';
import type {
LifetimeRebuildSummary,
VocabularyCleanupSummary,
@@ -50,6 +51,10 @@ export function createRunStatsCliCommandHandler(deps: {
ensureVocabularyCleanupTokenizerReady?: () => Promise<void> | void;
getImmersionTracker: () => {
cleanupVocabularyStats?: () => Promise<VocabularyCleanupSummary>;
cleanupDuplicateSubtitleLines?: (options: {
dryRun?: boolean;
lookbackDays?: number | null;
}) => Promise<DuplicateSubtitleLineCleanupSummary>;
rebuildLifetimeSummaries?: () => Promise<LifetimeRebuildSummary>;
} | null;
ensureStatsServerStarted: () => string;
@@ -83,6 +88,9 @@ export function createRunStatsCliCommandHandler(deps: {
| 'statsCleanup'
| 'statsCleanupVocab'
| 'statsCleanupLifetime'
| 'statsCleanupDuplicateLines'
| 'statsCleanupDryRun'
| 'statsCleanupLookbackDays'
>,
source: CliCommandSource,
): Promise<void> => {
@@ -126,6 +134,7 @@ export function createRunStatsCliCommandHandler(deps: {
const cleanupModes = [
args.statsCleanupVocab ? 'vocab' : null,
args.statsCleanupLifetime ? 'lifetime' : null,
args.statsCleanupDuplicateLines ? 'duplicate-lines' : null,
].filter(Boolean);
if (cleanupModes.length !== 1) {
throw new Error('Choose exactly one stats cleanup mode.');
@@ -142,6 +151,27 @@ export function createRunStatsCliCommandHandler(deps: {
writeResponseSafe(args.statsResponsePath, { ok: true });
return;
}
if (args.statsCleanupDuplicateLines && tracker.cleanupDuplicateSubtitleLines) {
const result = await tracker.cleanupDuplicateSubtitleLines({
dryRun: args.statsCleanupDryRun === true,
lookbackDays: args.statsCleanupLookbackDays ?? null,
});
const window =
result.lookbackDays === null ? 'all history' : `last ${result.lookbackDays}d`;
deps.logInfo(
`Stats duplicate-line cleanup ${result.dryRun ? 'preview' : 'complete'} (${window}): ` +
`scanned=${result.scannedLines} bursts=${result.burstGroups} ` +
`removedLines=${result.removedLines} removedWordCounts=${result.removedWordOccurrences} ` +
`removedKanjiCounts=${result.removedKanjiOccurrences}`,
);
for (const sample of result.samples.slice(0, 5)) {
deps.logInfo(
` ${sample.videoTitle ?? `video ${sample.videoId}`}: "${sample.text}" x${sample.frames}`,
);
}
writeResponseSafe(args.statsResponsePath, { ok: true });
return;
}
if (!args.statsCleanupLifetime || !tracker.rebuildLifetimeSummaries) {
throw new Error('Stats cleanup mode is not available.');
}
+18
View File
@@ -1004,6 +1004,24 @@ test('normalizeSubtitle collapses explicit line breaks when collapseLineBreaks i
);
});
test('normalizeSubtitle leaves already-decoded text alone', () => {
// Primary subtitle text is decoded from ASS once, upstream: by mpv for live lines and
// by the cue parser for prefetched ones. A brace that survives that is literal text.
assert.equal(normalizeSubtitle('本文{\\pos(1,2)'), '本文{\\pos(1,2)');
assert.equal(normalizeSubtitle(' 余白 ', false), ' 余白 ');
});
test('prepareSecondarySubtitleLines drops ASS vector drawing runs', () => {
assert.deepEqual(
prepareSecondarySubtitleLines(
'{\\an5\\pos(730,1042)\\p1\\blur1}m 20 0 b 10 0 0 10 0 20 b 0 31 10 40 20 40 {\\p0}',
),
[],
);
assert.deepEqual(prepareSecondarySubtitleLines('{\\p1}m 0 0 l 10 10{\\p0}本文'), ['本文']);
assert.deepEqual(prepareSecondarySubtitleLines('{\\pos(960,1068)\\bord3}位置指定'), ['位置指定']);
});
test('shouldRenderTokenizedSubtitle enables token rendering when tokens exist', () => {
assert.equal(shouldRenderTokenizedSubtitle(5), true);
assert.equal(shouldRenderTokenizedSubtitle(0), false);
+8 -11
View File
@@ -5,6 +5,7 @@ import type {
SubtitleData,
SubtitleRendererStyleConfig,
} from '../types';
import { assToPlainText, normalizePlainSubtitleText } from '../core/services/ass-text.js';
import type { RendererContext } from './context';
import { PRIMARY_SUB_VISIBLE_ON_YOMITAN_POPUP_CLASS } from './yomitan-popup.js';
@@ -42,17 +43,10 @@ function isWhitespaceOnly(value: string): boolean {
return value.trim().length === 0;
}
// Text reaching the overlay has already been decoded from ASS -- by mpv for live lines,
// by the cue parser for prefetched ones -- so this only settles line breaks.
export function normalizeSubtitle(text: string, trim = true, collapseLineBreaks = false): string {
if (!text) return '';
let normalized = text.replace(/\\N/g, '\n').replace(/\\n/g, '\n');
normalized = normalized.replace(/\{[^}]*\}/g, '');
if (collapseLineBreaks) {
normalized = normalized.replace(/\n/g, ' ');
normalized = normalized.replace(/\s+/g, ' ');
}
return trim ? normalized.trim() : normalized;
return normalizePlainSubtitleText(text, { trim, collapseLineBreaks });
}
const HEX_COLOR_PATTERN = /^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{4}|[0-9a-fA-F]{6}|[0-9a-fA-F]{8})$/;
@@ -672,7 +666,10 @@ function isKaraokeLikeLineSet(lines: string[]): boolean {
}
export function prepareSecondarySubtitleLines(text: string): string[] {
const normalized = normalizeSubtitle(text, true, false);
// The one display-side ASS decode: secondary text also reaches the overlay from
// websocket clients that forward their source line untouched, so unlike the primary
// path it cannot assume mpv already decoded it.
const normalized = assToPlainText(text).trim();
if (!normalized) return [];
+6 -13
View File
@@ -16,6 +16,8 @@
* along with this program. If not, see <https://www.gnu.org/licenses/>.
*/
import { normalizePlainSubtitleText } from './core/services/ass-text';
interface TimingEntry {
startTime: number;
endTime: number;
@@ -191,23 +193,14 @@ export class SubtitleTimingTracker {
return costs[shorter.length] || 0;
}
// Both sides take text mpv has already decoded from ASS; only whitespace differs
// between the lookup key (single line) and the display form (line breaks kept).
private normalizeText(text: string): string {
return text
.replace(/\\N/g, ' ')
.replace(/\\n/g, ' ')
.replace(/\n/g, ' ')
.replace(/{[^}]*}/g, '')
.replace(/\s+/g, ' ')
.trim();
return normalizePlainSubtitleText(text, { collapseLineBreaks: true });
}
private prepareDisplayText(text: string): string {
// Convert ASS/SSA newlines to real newlines, strip tags
return text
.replace(/\\N/g, '\n')
.replace(/\\n/g, '\n')
.replace(/{[^}]*}/g, '')
.trim();
return normalizePlainSubtitleText(text);
}
private startCleanup(): void {
+12 -41
View File
@@ -18,6 +18,7 @@ import type {
SessionTimelinePoint,
StatsAnkiNoteInfo,
StatsCoverImagesData,
StatsDuplicateLineCleanupResult,
StatsExcludedWord,
StreakCalendarDay,
TrendsDashboardData,
@@ -31,6 +32,13 @@ export type StatsTrendRange = '7d' | '30d' | '90d' | '365d' | 'all';
export type StatsTrendGroupBy = 'day' | 'month';
export type StatsMineMode = 'word' | 'sentence' | 'audio';
/** Body of `POST /api/stats/maintenance/duplicate-lines`. */
export interface StatsDuplicateLineCleanupRequest {
dryRun?: boolean;
/** Null means every recorded line, whatever its age. */
lookbackDays?: number | null;
}
export interface StatsSessionKnownWordsTimelinePoint {
linesSeen: number;
knownWordsSeen: number;
@@ -100,39 +108,6 @@ export interface StatsAnkiNotesInfoRequest {
noteIds: number[];
}
export interface StatsMergeAnimeRequest {
sourceAnimeIds: number[];
}
export interface StatsMoveVideoRequest {
animeId: number;
}
export interface StatsAnimeMergeRecommendation {
recommendationId: number;
animeIds: [number, number];
}
export interface StatsAnimeMergeRecommendationsResponse {
recommendations: StatsAnimeMergeRecommendation[];
}
export interface StatsMergeAnimeResponse {
ok: true;
/** Library entry that owns every merged episode afterwards. */
animeId: number;
mergedAnimeIds: number[];
movedVideos: number;
}
export interface StatsMoveVideoResponse {
ok: true;
animeId: number;
previousAnimeId: number | null;
/** True when the previous entry was emptied by the move and removed. */
removedPreviousAnime: boolean;
}
export interface StatsOkResponse {
ok: true;
}
@@ -158,6 +133,7 @@ export interface StatsJsonResponseMap {
vocabulary: VocabularyEntry[];
excludedWords: StatsExcludedWord[];
setExcludedWords: StatsOkResponse;
duplicateLineCleanup: StatsDuplicateLineCleanupResult;
wordOccurrences: VocabularyOccurrenceEntry[];
sentenceSearch: SentenceSearchResult[];
kanji: KanjiEntry[];
@@ -167,7 +143,6 @@ export interface StatsJsonResponseMap {
mediaLibrary: MediaLibraryItem[];
mediaDetail: MediaDetailData;
animeLibrary: AnimeLibraryItem[];
animeMergeRecommendations: StatsAnimeMergeRecommendationsResponse;
animeDetail: AnimeDetailData;
animeWords: AnimeWord[];
animeRollups: DailyRollup[];
@@ -176,9 +151,6 @@ export interface StatsJsonResponseMap {
deleteSession: StatsOkResponse;
deleteVideo: StatsOkResponse;
deleteAnime: StatsOkResponse;
mergeAnime: StatsMergeAnimeResponse;
moveVideoToAnime: StatsMoveVideoResponse;
dismissAnimeMergeRecommendation: StatsOkResponse;
anilistSearch: StatsAnilistSearchResult[];
knownWords: string[];
knownWordsSummary: StatsKnownWordsSummary;
@@ -215,6 +187,9 @@ export interface StatsHttpClient {
getVocabulary: (limit?: number) => Promise<VocabularyEntry[]>;
getExcludedWords: () => Promise<StatsExcludedWord[]>;
setExcludedWords: (words: StatsExcludedWord[]) => Promise<void>;
cleanupDuplicateLines: (
options?: StatsDuplicateLineCleanupRequest,
) => Promise<StatsDuplicateLineCleanupResult>;
getWordOccurrences: (
headword: string,
word: string,
@@ -236,7 +211,6 @@ export interface StatsHttpClient {
getMediaLibrary: () => Promise<MediaLibraryItem[]>;
getMediaDetail: (videoId: number) => Promise<MediaDetailData>;
getAnimeLibrary: () => Promise<AnimeLibraryItem[]>;
getAnimeMergeRecommendations: () => Promise<StatsAnimeMergeRecommendationsResponse>;
getAnimeDetail: (animeId: number) => Promise<AnimeDetailData>;
getAnimeWords: (animeId: number, limit?: number) => Promise<AnimeWord[]>;
getAnimeRollups: (animeId: number, limit?: number) => Promise<DailyRollup[]>;
@@ -259,9 +233,6 @@ export interface StatsHttpClient {
deleteSessions: (sessionIds: number[]) => Promise<void>;
deleteVideo: (videoId: number) => Promise<void>;
deleteAnime: (animeId: number) => Promise<void>;
mergeAnime: (targetAnimeId: number, sourceAnimeIds: number[]) => Promise<StatsMergeAnimeResponse>;
moveVideoToAnime: (videoId: number, animeId: number) => Promise<StatsMoveVideoResponse>;
dismissAnimeMergeRecommendation: (recommendationId: number) => Promise<void>;
getKnownWords: () => Promise<string[]>;
getKnownWordsSummary: () => Promise<StatsKnownWordsSummary>;
getAnimeKnownWordsSummary: (animeId: number) => Promise<StatsKnownWordsSummary>;
+23
View File
@@ -82,6 +82,29 @@ export interface StatsExcludedWord {
reading: string;
}
/** One animation burst the duplicate-line cleanup found in the stats database. */
export interface StatsDuplicateLineSample {
videoId: number;
videoTitle: string | null;
text: string;
/** Events recorded for this run, including the one that is kept. */
frames: number;
removedLines: number;
startMs: number;
endMs: number;
}
export interface StatsDuplicateLineCleanupResult {
dryRun: boolean;
lookbackDays: number | null;
scannedLines: number;
burstGroups: number;
removedLines: number;
removedWordOccurrences: number;
removedKanjiOccurrences: number;
samples: StatsDuplicateLineSample[];
}
export interface StatsCoverImage {
contentType: string;
dataUrl: string;
+3 -26
View File
@@ -5,45 +5,22 @@ import type { AnimeLibraryItem } from '../../types/stats';
interface AnimeCardProps {
anime: AnimeLibraryItem;
onClick: () => void;
/** While selecting, clicking the card toggles it instead of opening it. */
selectable?: boolean;
selected?: boolean;
}
export function AnimeCard({
anime,
onClick,
selectable = false,
selected = false,
}: AnimeCardProps) {
export function AnimeCard({ anime, onClick }: AnimeCardProps) {
return (
<button
type="button"
onClick={onClick}
aria-pressed={selectable ? selected : undefined}
className={`group bg-ctp-surface0 border rounded-lg overflow-hidden hover:shadow-lg hover:shadow-ctp-blue/10 transition-all duration-200 hover:-translate-y-1 text-left w-full ${
selected ? 'border-ctp-blue' : 'border-ctp-surface1 hover:border-ctp-blue/50'
}`}
className="group bg-ctp-surface0 border border-ctp-surface1 rounded-lg overflow-hidden hover:border-ctp-blue/50 hover:shadow-lg hover:shadow-ctp-blue/10 transition-all duration-200 hover:-translate-y-1 text-left w-full"
>
<div className="overflow-hidden relative">
<div className="overflow-hidden">
<AnimeCoverImage
animeId={anime.animeId}
title={anime.canonicalTitle}
coverRetryToken={anime.anilistId ?? 0}
className="w-full aspect-[3/4] rounded-t-lg transition-transform duration-200 group-hover:scale-105"
/>
{selectable && (
<span
aria-hidden="true"
className={`absolute top-2 left-2 w-5 h-5 rounded border flex items-center justify-center text-xs ${
selected
? 'bg-ctp-blue border-ctp-blue text-ctp-base'
: 'bg-ctp-crust/70 border-ctp-surface2 text-transparent'
}`}
>
{'✓'}
</span>
)}
</div>
<div className="p-3">
<div className="text-sm font-medium text-ctp-text truncate">{anime.canonicalTitle}</div>
@@ -25,8 +25,6 @@ interface AnimeDetailViewProps {
* keeps showing the previous title's art.
*/
onAnilistRelinked?: () => void;
/** Called after an episode is reassigned to another entry. */
onEpisodeMoved?: () => void;
}
type Range = 14 | 30 | 90;
@@ -152,7 +150,6 @@ export function AnimeDetailView({
onOpenEpisodeDetail,
onAnimeDeleted,
onAnilistRelinked,
onEpisodeMoved,
}: AnimeDetailViewProps) {
const { data, loading, error, reload } = useAnimeDetail(animeId);
const [showAnilistSelector, setShowAnilistSelector] = useState(false);
@@ -226,13 +223,6 @@ export function AnimeDetailView({
<AnimeOverviewStats detail={detail} knownWordsSummary={knownWordsSummary} />
<EpisodeList
episodes={episodes}
animeId={animeId}
onEpisodeMoved={(removedPreviousAnime) => {
onEpisodeMoved?.();
// The last episode taking the entry with it leaves nothing to show.
if (removedPreviousAnime) onBack();
else reload();
}}
onOpenDetail={onOpenEpisodeDetail ? (videoId) => onOpenEpisodeDetail(videoId) : undefined}
/>
<AnimeWatchChart animeId={animeId} />
@@ -1,232 +0,0 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { Window } from 'happy-dom';
import { act, useState } from 'react';
import { createRoot } from 'react-dom/client';
import { apiClient } from '../../lib/api-client';
import type { AnimeLibraryItem } from '../../types/stats';
import { AnimeMergeDialog } from './AnimeMergeDialog';
import { LibraryEntryPicker } from './LibraryEntryPicker';
interface TestWindow extends Window {
IS_REACT_ACT_ENVIRONMENT?: boolean;
}
function installDom(): () => void {
const previousWindow = globalThis.window;
const previousDocument = globalThis.document;
const previousHTMLElement = globalThis.HTMLElement;
const previousIsReactActEnvironment = (
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT;
const window = new Window() as TestWindow;
Object.defineProperty(globalThis, 'window', { value: window, configurable: true });
Object.defineProperty(globalThis, 'document', { value: window.document, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: window.HTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
return () => {
Object.defineProperty(globalThis, 'window', { value: previousWindow, configurable: true });
Object.defineProperty(globalThis, 'document', { value: previousDocument, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: previousHTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = previousIsReactActEnvironment;
};
}
function libraryItem(animeId: number, title: string): AnimeLibraryItem {
return {
animeId,
canonicalTitle: title,
anilistId: null,
totalSessions: 1,
totalActiveMs: 1000,
totalCards: 0,
totalTokensSeen: 0,
episodeCount: 1,
episodesTotal: null,
lastWatchedMs: 1,
};
}
test('AnimeMergeDialog focuses its close control, closes on Escape, and restores focus', async () => {
const uninstallDom = installDom();
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
function Harness() {
const [open, setOpen] = useState(false);
return (
<>
<button type="button" onClick={() => setOpen(true)}>
Review merge
</button>
{open ? (
<AnimeMergeDialog
entries={[libraryItem(1, 'Show'), libraryItem(2, 'Show Season 1')]}
onClose={() => setOpen(false)}
onMerged={() => undefined}
/>
) : null}
</>
);
}
await act(async () => root.render(<Harness />));
const trigger = container.querySelector('button') as HTMLButtonElement;
trigger.focus();
await act(async () => trigger.click());
assert.equal(document.activeElement?.getAttribute('aria-label'), 'Close');
await act(async () => {
document.dispatchEvent(new window.KeyboardEvent('keydown', { key: 'Escape' }));
});
assert.equal(container.querySelector('[role="dialog"]'), null);
assert.equal(document.activeElement, trigger);
await act(async () => root.unmount());
} finally {
uninstallDom();
}
});
test('AnimeMergeDialog keeps keyboard focus inside the modal', async () => {
const uninstallDom = installDom();
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(
<AnimeMergeDialog
entries={[libraryItem(1, 'Show'), libraryItem(2, 'Show Season 1')]}
onClose={() => undefined}
onMerged={() => undefined}
/>,
);
});
const dialog = container.querySelector('[role="dialog"]') as HTMLElement;
const focusable = [...dialog.querySelectorAll('button:not([disabled])')] as HTMLButtonElement[];
const first = focusable[0];
const last = focusable.at(-1);
assert.ok(first);
assert.ok(last);
last.focus();
await act(async () => {
document.dispatchEvent(new window.KeyboardEvent('keydown', { key: 'Tab' }));
});
assert.equal(document.activeElement, first);
first.focus();
await act(async () => {
document.dispatchEvent(new window.KeyboardEvent('keydown', { key: 'Tab', shiftKey: true }));
});
assert.equal(document.activeElement, last);
await act(async () => root.unmount());
} finally {
uninstallDom();
}
});
test('LibraryEntryPicker focuses search, closes on Escape, and restores focus', async () => {
const uninstallDom = installDom();
const original = apiClient.getAnimeLibrary;
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show'),
]) as typeof apiClient.getAnimeLibrary;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
function Harness() {
const [open, setOpen] = useState(false);
return (
<>
<button type="button" onClick={() => setOpen(true)}>
Move
</button>
{open ? (
<LibraryEntryPicker
heading="Move episode"
onSelect={() => undefined}
onClose={() => setOpen(false)}
/>
) : null}
</>
);
}
await act(async () => root.render(<Harness />));
const trigger = container.querySelector('button') as HTMLButtonElement;
trigger.focus();
await act(async () => trigger.click());
assert.equal(document.activeElement?.getAttribute('placeholder'), 'Search library...');
await act(async () => {
document.dispatchEvent(new window.KeyboardEvent('keydown', { key: 'Escape' }));
});
assert.equal(container.querySelector('[role="dialog"]'), null);
assert.equal(document.activeElement, trigger);
await act(async () => root.unmount());
} finally {
apiClient.getAnimeLibrary = original;
uninstallDom();
}
});
test('LibraryEntryPicker cannot be dismissed while a move is in flight', async () => {
const uninstallDom = installDom();
const original = apiClient.getAnimeLibrary;
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show'),
]) as typeof apiClient.getAnimeLibrary;
let closeCalls = 0;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(
<LibraryEntryPicker
heading="Move episode"
busyAnimeId={1}
onSelect={() => undefined}
onClose={() => {
closeCalls += 1;
}}
/>,
);
});
const closeButton = container.querySelector('button[aria-label="Close"]') as HTMLButtonElement;
assert.equal(closeButton.disabled, true);
await act(async () => {
closeButton.click();
container.firstElementChild?.dispatchEvent(new window.MouseEvent('click', { bubbles: true }));
document.dispatchEvent(new window.KeyboardEvent('keydown', { key: 'Escape' }));
});
assert.equal(closeCalls, 0);
await act(async () => root.unmount());
} finally {
apiClient.getAnimeLibrary = original;
uninstallDom();
}
});
@@ -1,164 +0,0 @@
import { useId, useRef, useState } from 'react';
import { apiClient } from '../../lib/api-client';
import { formatDuration, formatNumber } from '../../lib/formatters';
import { useModalFocus } from '../../hooks/useModalFocus';
import { AnimeCoverImage } from './AnimeCoverImage';
import type { AnimeLibraryItem } from '../../types/stats';
interface AnimeMergeDialogProps {
entries: AnimeLibraryItem[];
onClose: () => void;
onMerged: (survivingAnimeId: number) => void;
}
/** Biggest entry first: the one most likely to carry the right title and art. */
function pickDefaultKeeper(entries: AnimeLibraryItem[]): number {
const best = [...entries].sort(
(a, b) => b.episodeCount - a.episodeCount || b.totalActiveMs - a.totalActiveMs,
)[0];
return best?.animeId ?? 0;
}
export function AnimeMergeDialog({ entries, onClose, onMerged }: AnimeMergeDialogProps) {
const headingId = useId();
const dialogRef = useRef<HTMLDivElement>(null);
const closeButtonRef = useRef<HTMLButtonElement>(null);
const [keeperId, setKeeperId] = useState(() => pickDefaultKeeper(entries));
const [merging, setMerging] = useState(false);
const [error, setError] = useState<string | null>(null);
const totalEpisodes = entries.reduce((sum, entry) => sum + entry.episodeCount, 0);
const totalCards = entries.reduce((sum, entry) => sum + entry.totalCards, 0);
const totalActiveMs = entries.reduce((sum, entry) => sum + entry.totalActiveMs, 0);
useModalFocus({
dialogRef,
initialFocusRef: closeButtonRef,
dismissDisabled: merging,
onDismiss: onClose,
});
const handleMerge = async () => {
const sourceAnimeIds = entries
.map((entry) => entry.animeId)
.filter((animeId) => animeId !== keeperId);
if (sourceAnimeIds.length === 0) return;
setMerging(true);
setError(null);
try {
const result = await apiClient.mergeAnime(keeperId, sourceAnimeIds);
onMerged(result.animeId);
} catch (err) {
setError(err instanceof Error ? err.message : 'Failed to merge these entries.');
setMerging(false);
}
};
// Dismissing mid-request would leave the caller unaware of a merge that is
// still going to land, so the backdrop and close button are inert until it
// resolves.
const handleDismiss = () => {
if (!merging) onClose();
};
return (
<div
className="fixed inset-0 z-50 flex items-start justify-center pt-[10vh]"
onClick={handleDismiss}
>
<div className="absolute inset-0 bg-ctp-crust/70 backdrop-blur-[2px]" />
<div
ref={dialogRef}
role="dialog"
aria-modal="true"
aria-labelledby={headingId}
className="relative bg-ctp-base border border-ctp-surface1 rounded-xl shadow-2xl w-full max-w-lg max-h-[70vh] flex flex-col animate-fade-in"
onClick={(e) => e.stopPropagation()}
>
<div className="p-4 border-b border-ctp-surface1">
<div className="flex items-center justify-between">
<h3 id={headingId} className="text-sm font-semibold text-ctp-text">
Merge {entries.length} Library Entries
</h3>
<button
ref={closeButtonRef}
type="button"
onClick={handleDismiss}
disabled={merging}
aria-label="Close"
className="text-ctp-overlay2 hover:text-ctp-text text-lg leading-none disabled:opacity-50"
>
{'✕'}
</button>
</div>
<p className="text-xs text-ctp-overlay2 mt-2">
Pick the entry to keep. Every episode moves onto it and the others are removed; no
sessions or mined cards are deleted.
</p>
</div>
<div className="flex-1 overflow-y-auto p-2">
{entries.map((entry) => (
<button
key={entry.animeId}
type="button"
disabled={merging}
aria-pressed={keeperId === entry.animeId}
onClick={() => setKeeperId(entry.animeId)}
className={`w-full flex items-center gap-3 p-2.5 rounded-lg transition-colors text-left disabled:opacity-50 ${
keeperId === entry.animeId ? 'bg-ctp-surface1' : 'hover:bg-ctp-surface0'
}`}
>
<span
aria-hidden="true"
className={`w-4 h-4 rounded-full border shrink-0 ${
keeperId === entry.animeId
? 'border-ctp-blue bg-ctp-blue'
: 'border-ctp-surface2 bg-transparent'
}`}
/>
<AnimeCoverImage
animeId={entry.animeId}
title={entry.canonicalTitle}
coverRetryToken={entry.anilistId ?? 0}
className="w-10 h-14 rounded shrink-0"
/>
<div className="min-w-0 flex-1">
<div className="text-sm text-ctp-text truncate">{entry.canonicalTitle}</div>
<div className="text-xs text-ctp-overlay2 mt-0.5">
{entry.episodeCount} episode{entry.episodeCount !== 1 ? 's' : ''} ·{' '}
{formatDuration(entry.totalActiveMs)} · {formatNumber(entry.totalCards)} cards
</div>
</div>
{keeperId === entry.animeId ? (
<span className="text-xs text-ctp-blue shrink-0">Keep</span>
) : null}
</button>
))}
</div>
<div className="p-4 border-t border-ctp-surface1 space-y-2">
{error ? (
<div role="alert" className="text-xs text-ctp-red">
{error}
</div>
) : null}
<div className="flex items-center justify-between gap-3">
<div className="text-xs text-ctp-overlay2">
Result: {totalEpisodes} episode{totalEpisodes !== 1 ? 's' : ''} ·{' '}
{formatDuration(totalActiveMs)} · {formatNumber(totalCards)} cards
</div>
<button
type="button"
disabled={merging}
onClick={() => void handleMerge()}
className="px-3 py-1.5 rounded-lg bg-ctp-blue/15 border border-ctp-blue/40 text-xs text-ctp-blue hover:bg-ctp-blue/25 transition-colors disabled:opacity-50"
>
{merging ? 'Merging…' : 'Merge Entries'}
</button>
</div>
</div>
</div>
</div>
);
}
@@ -1,416 +0,0 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { Window } from 'happy-dom';
import { act } from 'react';
import { createRoot } from 'react-dom/client';
import { apiClient } from '../../lib/api-client';
import type { AnimeLibraryItem, StatsMergeAnimeResponse } from '../../types/stats';
import { AnimeTab } from './AnimeTab';
interface TestWindow extends Window {
IS_REACT_ACT_ENVIRONMENT?: boolean;
}
function installDom(): () => void {
const previousWindow = globalThis.window;
const previousDocument = globalThis.document;
const previousHTMLElement = globalThis.HTMLElement;
const previousIsReactActEnvironment = (
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT;
const window = new Window() as TestWindow;
Object.defineProperty(globalThis, 'window', { value: window, configurable: true });
Object.defineProperty(globalThis, 'document', { value: window.document, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: window.HTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
return () => {
Object.defineProperty(globalThis, 'window', { value: previousWindow, configurable: true });
Object.defineProperty(globalThis, 'document', { value: previousDocument, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: previousHTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = previousIsReactActEnvironment;
};
}
function libraryItem(animeId: number, title: string, episodeCount: number): AnimeLibraryItem {
return {
animeId,
canonicalTitle: title,
anilistId: null,
totalSessions: 1,
totalActiveMs: 1000,
totalCards: 1,
totalTokensSeen: 0,
episodeCount,
episodesTotal: null,
lastWatchedMs: animeId,
};
}
function findButton(container: Element, label: string): HTMLElement {
const match = [...container.querySelectorAll('button')].find((button) =>
(button.textContent ?? '').includes(label),
);
assert.ok(match, `expected a "${label}" button`);
return match as unknown as HTMLElement;
}
/** Library cards only expose aria-pressed while selection mode is on. */
function cardButtons(container: Element): HTMLButtonElement[] {
return [...container.querySelectorAll('button[aria-pressed]')] as unknown as HTMLButtonElement[];
}
function mergeButton(container: Element): HTMLButtonElement {
const match = [...container.querySelectorAll('button')].find(
(button) => (button.textContent ?? '').trim() === 'Merge Selected',
);
assert.ok(match, 'expected a "Merge Selected" button');
return match as unknown as HTMLButtonElement;
}
test('AnimeTab merges the selected duplicate entries into the chosen keeper', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
mergeAnime: apiClient.mergeAnime,
};
// Two cards for one show, the split this feature exists to undo.
let entries = [libraryItem(1, 'Show', 2), libraryItem(2, 'Show Season 1', 1)];
let libraryFetches = 0;
let mergeCall: { targetAnimeId: number; sourceAnimeIds: number[] } | null = null;
apiClient.getAnimeLibrary = (async () => {
libraryFetches += 1;
return entries;
}) as typeof apiClient.getAnimeLibrary;
apiClient.mergeAnime = (async (targetAnimeId: number, sourceAnimeIds: number[]) => {
mergeCall = { targetAnimeId, sourceAnimeIds };
entries = [libraryItem(1, 'Show', 3)];
return {
ok: true,
animeId: targetAnimeId,
mergedAnimeIds: sourceAnimeIds,
movedVideos: 1,
} satisfies StatsMergeAnimeResponse;
}) as typeof apiClient.mergeAnime;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(<AnimeTab />);
});
assert.equal(libraryFetches, 1);
await act(async () => {
findButton(container, 'Select').click();
});
// Nothing to merge until at least two entries are picked.
assert.equal(mergeButton(container).disabled, true);
// Sorted by last watched, so the season-tagged duplicate comes first.
const cards = cardButtons(container);
assert.equal(cards.length, 2);
assert.match(cards[0]?.textContent ?? '', /Show Season 1/);
await act(async () => {
cards[0]?.click();
});
assert.equal(mergeButton(container).disabled, true);
await act(async () => {
cardButtons(container)[1]?.click();
});
assert.equal(mergeButton(container).disabled, false);
await act(async () => {
mergeButton(container).click();
});
// The dialog defaults to the entry with the most episodes.
assert.match(container.textContent ?? '', /Merge 2 Library Entries/);
await act(async () => {
findButton(container, 'Merge Entries').click();
});
assert.deepEqual(mergeCall, { targetAnimeId: 1, sourceAnimeIds: [2] });
assert.equal(libraryFetches, 2);
// Selection mode closes and the grid is back to a single card.
assert.doesNotMatch(container.textContent ?? '', /Merge 2 Library Entries/);
assert.doesNotMatch(container.textContent ?? '', /Show Season 1/);
await act(async () => {
root.unmount();
});
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab keeps a suggested duplicate visible until it is reviewed and merged', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
dismissAnimeMergeRecommendation: apiClient.dismissAnimeMergeRecommendation,
mergeAnime: apiClient.mergeAnime,
};
let entries = [libraryItem(1, 'Show', 2), libraryItem(2, 'Show Season 1', 1)];
let recommendations = [{ recommendationId: 41, animeIds: [1, 2] }];
let mergeCall: { targetAnimeId: number; sourceAnimeIds: number[] } | null = null;
apiClient.getAnimeLibrary = (async () => entries) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => ({
recommendations,
})) as typeof apiClient.getAnimeMergeRecommendations;
apiClient.dismissAnimeMergeRecommendation = (async () =>
undefined) as typeof apiClient.dismissAnimeMergeRecommendation;
apiClient.mergeAnime = (async (targetAnimeId: number, sourceAnimeIds: number[]) => {
mergeCall = { targetAnimeId, sourceAnimeIds };
entries = [libraryItem(targetAnimeId, 'Show', 3)];
recommendations = [];
return {
ok: true,
animeId: targetAnimeId,
mergedAnimeIds: sourceAnimeIds,
movedVideos: 1,
} satisfies StatsMergeAnimeResponse;
}) as typeof apiClient.mergeAnime;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(<AnimeTab />);
});
assert.match(container.textContent ?? '', /Possible duplicate/);
assert.match(container.textContent ?? '', /Show/);
assert.match(container.textContent ?? '', /Show Season 1/);
await act(async () => {
findButton(container, 'Review merge').click();
});
assert.match(container.textContent ?? '', /Merge 2 Library Entries/);
const keeper = [...container.querySelectorAll('button[aria-pressed]')].find((button) =>
(button.textContent ?? '').includes('Show Season 1'),
) as HTMLButtonElement | undefined;
assert.ok(keeper);
await act(async () => {
keeper.click();
});
await act(async () => {
findButton(container, 'Merge Entries').click();
});
assert.deepEqual(mergeCall, { targetAnimeId: 2, sourceAnimeIds: [1] });
assert.doesNotMatch(container.textContent ?? '', /Possible duplicate/);
assert.doesNotMatch(container.textContent ?? '', /Show Season 1/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab dismisses a false-positive duplicate recommendation', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
dismissAnimeMergeRecommendation: apiClient.dismissAnimeMergeRecommendation,
};
let dismissedId: number | null = null;
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show', 2),
libraryItem(2, 'Different Show', 1),
]) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => ({
recommendations: [{ recommendationId: 73, animeIds: [1, 2] }],
})) as typeof apiClient.getAnimeMergeRecommendations;
apiClient.dismissAnimeMergeRecommendation = (async (recommendationId: number) => {
dismissedId = recommendationId;
}) as typeof apiClient.dismissAnimeMergeRecommendation;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => root.render(<AnimeTab />));
await act(async () => {
findButton(container, 'Not duplicates').click();
});
assert.equal(dismissedId, 73);
assert.doesNotMatch(container.textContent ?? '', /Possible duplicate/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab refreshes the library and recommendations when the window regains focus', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
};
let entries = [libraryItem(1, 'Show', 2), libraryItem(2, 'Show Season 1', 1)];
let libraryFetches = 0;
let recommendationFetches = 0;
apiClient.getAnimeLibrary = (async () => {
libraryFetches += 1;
return entries;
}) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => {
recommendationFetches += 1;
return { recommendations: [] };
}) as typeof apiClient.getAnimeMergeRecommendations;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => root.render(<AnimeTab />));
entries = [libraryItem(1, 'Show', 3)];
await act(async () => {
window.dispatchEvent(new window.Event('focus'));
});
assert.equal(libraryFetches, 2);
assert.equal(recommendationFetches, 2);
assert.doesNotMatch(container.textContent ?? '', /Show Season 1/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab keeps a recommendation visible through a transient refresh failure', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
};
let failRecommendations = false;
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show', 2),
libraryItem(2, 'Show Season 1', 1),
]) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => {
if (failRecommendations) throw new Error('temporary failure');
return { recommendations: [{ recommendationId: 41, animeIds: [1, 2] }] };
}) as typeof apiClient.getAnimeMergeRecommendations;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => root.render(<AnimeTab />));
assert.match(container.textContent ?? '', /Possible duplicate/);
failRecommendations = true;
await act(async () => window.dispatchEvent(new window.Event('focus')));
assert.match(container.textContent ?? '', /Possible duplicate/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab retains a recommendation and reports a failed dismissal', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
dismissAnimeMergeRecommendation: apiClient.dismissAnimeMergeRecommendation,
};
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show', 2),
libraryItem(2, 'Show Season 1', 1),
]) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => ({
recommendations: [{ recommendationId: 41, animeIds: [1, 2] }],
})) as typeof apiClient.getAnimeMergeRecommendations;
apiClient.dismissAnimeMergeRecommendation = (async () => {
throw new Error('offline');
}) as typeof apiClient.dismissAnimeMergeRecommendation;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => root.render(<AnimeTab />));
await act(async () => findButton(container, 'Not duplicates').click());
assert.match(container.textContent ?? '', /Possible duplicate/);
assert.match(container.textContent ?? '', /Could not dismiss this suggestion/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
test('AnimeTab keeps an open recommendation review stable during background refresh', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
getAnimeMergeRecommendations: apiClient.getAnimeMergeRecommendations,
};
let recommendations = [{ recommendationId: 41, animeIds: [1, 2] as [number, number] }];
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Show', 2),
libraryItem(2, 'Show Season 1', 1),
]) as typeof apiClient.getAnimeLibrary;
apiClient.getAnimeMergeRecommendations = (async () => ({
recommendations,
})) as typeof apiClient.getAnimeMergeRecommendations;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => root.render(<AnimeTab />));
await act(async () => findButton(container, 'Review merge').click());
assert.match(container.textContent ?? '', /Merge 2 Library Entries/);
recommendations = [];
await act(async () => window.dispatchEvent(new window.Event('focus')));
assert.match(container.textContent ?? '', /Merge 2 Library Entries/);
assert.match(container.textContent ?? '', /Show Season 1/);
await act(async () => root.unmount());
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
+2 -118
View File
@@ -9,8 +9,6 @@ import {
} from '../../lib/library-card-size';
import { AnimeCard } from './AnimeCard';
import { AnimeDetailView } from './AnimeDetailView';
import { AnimeMergeDialog } from './AnimeMergeDialog';
import { DuplicateReviewStrip } from './DuplicateReviewStrip';
type SortKey = 'lastWatched' | 'watchTime' | 'cards' | 'episodes';
@@ -55,17 +53,7 @@ export function AnimeTab({
onNavigateToWord,
onOpenEpisodeDetail,
}: AnimeTabProps) {
const {
anime,
loading,
error,
reload,
recommendations,
dismissRecommendation,
dismissingRecommendationId,
recommendationActionError,
clearRecommendation,
} = useAnimeLibrary();
const { anime, loading, error, reload } = useAnimeLibrary();
const [search, setSearch] = useState('');
const [sortKey, setSortKey] = useState<SortKey>('lastWatched');
const [cardSize, setCardSize] = useState<LibraryCardSize>(() =>
@@ -74,23 +62,6 @@ export function AnimeTab({
),
);
const [selectedAnimeId, setSelectedAnimeId] = useState<number | null>(null);
const [selectionMode, setSelectionMode] = useState(false);
const [checkedAnimeIds, setCheckedAnimeIds] = useState<number[]>([]);
const [showMergeDialog, setShowMergeDialog] = useState(false);
const [reviewRecommendationId, setReviewRecommendationId] = useState<number | null>(null);
const [reviewAnimeIds, setReviewAnimeIds] = useState<[number, number] | null>(null);
function toggleChecked(animeId: number): void {
setCheckedAnimeIds((ids) =>
ids.includes(animeId) ? ids.filter((id) => id !== animeId) : [...ids, animeId],
);
}
function exitSelectionMode(): void {
setSelectionMode(false);
setCheckedAnimeIds([]);
setShowMergeDialog(false);
}
function handleCardSizeChange(size: LibraryCardSize): void {
setCardSize(size);
@@ -115,22 +86,6 @@ export function AnimeTab({
}, [anime, search, sortKey]);
const totalMs = anime.reduce((sum, a) => sum + a.totalActiveMs, 0);
const checkedEntries = checkedAnimeIds
.map((animeId) => anime.find((entry) => entry.animeId === animeId))
.filter((entry): entry is (typeof anime)[number] => entry !== undefined);
const hydratedRecommendations = recommendations
.map((recommendation) => ({
...recommendation,
entries: recommendation.animeIds
.map((animeId) => anime.find((entry) => entry.animeId === animeId))
.filter((entry): entry is (typeof anime)[number] => entry !== undefined),
}))
.filter((recommendation) => recommendation.entries.length >= 2);
const activeRecommendation = hydratedRecommendations[0] ?? null;
const reviewEntries = (reviewAnimeIds ?? [])
.map((animeId) => anime.find((entry) => entry.animeId === animeId))
.filter((entry): entry is (typeof anime)[number] => entry !== undefined);
const mergeEntries = reviewRecommendationId !== null ? reviewEntries : checkedEntries;
if (selectedAnimeId !== null) {
return (
@@ -145,7 +100,6 @@ export function AnimeTab({
}
onAnimeDeleted={reload}
onAnilistRelinked={reload}
onEpisodeMoved={reload}
/>
);
}
@@ -189,57 +143,11 @@ export function AnimeTab({
</button>
))}
</div>
<button
type="button"
onClick={() => (selectionMode ? exitSelectionMode() : setSelectionMode(true))}
title="Select several entries to merge them into one"
className={`px-2 py-2 rounded-lg border text-xs shrink-0 transition-colors ${
selectionMode
? 'bg-ctp-blue/15 border-ctp-blue/40 text-ctp-blue'
: 'bg-ctp-surface0 border-ctp-surface1 text-ctp-overlay2 hover:text-ctp-subtext0'
}`}
>
{selectionMode ? 'Cancel' : 'Select'}
</button>
<div className="text-xs text-ctp-overlay2 shrink-0">
{filtered.length} titles · {formatDuration(totalMs)}
</div>
</div>
{activeRecommendation ? (
<DuplicateReviewStrip
entries={activeRecommendation.entries}
current={1}
total={hydratedRecommendations.length}
dismissing={dismissingRecommendationId === activeRecommendation.recommendationId}
error={recommendationActionError}
onReview={() => {
setReviewRecommendationId(activeRecommendation.recommendationId);
setReviewAnimeIds(activeRecommendation.animeIds);
setShowMergeDialog(true);
}}
onDismiss={() => void dismissRecommendation(activeRecommendation.recommendationId)}
/>
) : null}
{selectionMode && (
<div className="flex items-center justify-between gap-3 bg-ctp-surface0 border border-ctp-surface1 rounded-lg px-3 py-2">
<div className="text-xs text-ctp-overlay2">
{checkedEntries.length === 0
? 'Pick the duplicate entries to combine'
: `${checkedEntries.length} selected`}
</div>
<button
type="button"
disabled={checkedEntries.length < 2}
onClick={() => setShowMergeDialog(true)}
className="px-3 py-1.5 rounded-lg bg-ctp-blue/15 border border-ctp-blue/40 text-xs text-ctp-blue hover:bg-ctp-blue/25 transition-colors disabled:opacity-40 disabled:cursor-not-allowed"
>
Merge Selected
</button>
</div>
)}
{filtered.length === 0 ? (
<div className="text-sm text-ctp-overlay2 p-4">No titles found</div>
) : (
@@ -248,35 +156,11 @@ export function AnimeTab({
<AnimeCard
key={item.animeId}
anime={item}
selectable={selectionMode}
selected={checkedAnimeIds.includes(item.animeId)}
onClick={() =>
selectionMode ? toggleChecked(item.animeId) : setSelectedAnimeId(item.animeId)
}
onClick={() => setSelectedAnimeId(item.animeId)}
/>
))}
</div>
)}
{showMergeDialog && mergeEntries.length >= 2 && (
<AnimeMergeDialog
entries={mergeEntries}
onClose={() => {
setShowMergeDialog(false);
setReviewRecommendationId(null);
setReviewAnimeIds(null);
}}
onMerged={() => {
if (reviewRecommendationId !== null) {
clearRecommendation(reviewRecommendationId);
}
exitSelectionMode();
setReviewRecommendationId(null);
setReviewAnimeIds(null);
reload();
}}
/>
)}
</div>
);
}
@@ -1,66 +0,0 @@
import type { AnimeLibraryItem } from '../../types/stats';
interface DuplicateReviewStripProps {
entries: AnimeLibraryItem[];
current: number;
total: number;
dismissing: boolean;
error?: string | null;
onReview: () => void;
onDismiss: () => void;
}
export function DuplicateReviewStrip({
entries,
current,
total,
dismissing,
error = null,
onReview,
onDismiss,
}: DuplicateReviewStripProps) {
return (
<aside
aria-label="Possible duplicate library entries"
className="relative overflow-hidden rounded-lg border border-ctp-yellow/25 bg-ctp-yellow/[0.06] px-3 py-2.5"
>
<div className="absolute inset-y-0 left-0 w-0.5 bg-ctp-yellow/70" aria-hidden="true" />
<div className="flex items-center gap-3">
<div className="min-w-0 flex-1">
<div className="flex items-center gap-2">
<span className="text-xs font-medium text-ctp-yellow">Possible duplicate</span>
{total > 1 ? (
<span className="text-[10px] tabular-nums text-ctp-overlay1">
{current} of {total}
</span>
) : null}
</div>
<p className="mt-0.5 truncate text-xs text-ctp-subtext0">
{entries.map((entry) => entry.canonicalTitle).join(' · ')}
</p>
</div>
<button
type="button"
disabled={dismissing}
onClick={onDismiss}
className="shrink-0 rounded-md px-2.5 py-1.5 text-xs text-ctp-overlay2 transition-colors hover:bg-ctp-surface0 hover:text-ctp-text disabled:opacity-50"
>
{dismissing ? 'Dismissing…' : 'Not duplicates'}
</button>
<button
type="button"
disabled={dismissing}
onClick={onReview}
className="shrink-0 rounded-md border border-ctp-yellow/35 bg-ctp-yellow/10 px-2.5 py-1.5 text-xs font-medium text-ctp-yellow transition-colors hover:bg-ctp-yellow/20 disabled:opacity-50"
>
Review merge
</button>
</div>
{error ? (
<p role="alert" className="mt-1.5 text-xs text-ctp-red">
{error}
</p>
) : null}
</aside>
);
}
+1 -62
View File
@@ -4,39 +4,21 @@ import { apiClient } from '../../lib/api-client';
import { confirmEpisodeDelete } from '../../lib/delete-confirm';
import { buildLookupRateDisplay } from '../../lib/yomitan-lookup';
import { EpisodeDetail } from './EpisodeDetail';
import { LibraryEntryPicker } from './LibraryEntryPicker';
import type { AnimeEpisode } from '../../types/stats';
/**
* Row actions that only appear on hover. Keyboard focus and pointers with no
* hover (touch) reveal them too, otherwise those users cannot reach the button
* at all.
*/
const HOVER_REVEALED =
'opacity-0 group-hover:opacity-100 focus-visible:opacity-100 [@media(hover:none)]:opacity-100';
interface EpisodeListProps {
episodes: AnimeEpisode[];
/** Entry these episodes currently belong to; excluded from the move picker. */
animeId?: number;
onEpisodeDeleted?: () => void;
/** Fires after an episode is reassigned, so the caller can refetch. */
onEpisodeMoved?: (removedPreviousAnime: boolean) => void;
onOpenDetail?: (videoId: number) => void;
}
export function EpisodeList({
episodes: initialEpisodes,
animeId,
onEpisodeDeleted,
onEpisodeMoved,
onOpenDetail,
}: EpisodeListProps) {
const [expandedVideoId, setExpandedVideoId] = useState<number | null>(null);
const [episodes, setEpisodes] = useState(initialEpisodes);
const [movingEpisode, setMovingEpisode] = useState<AnimeEpisode | null>(null);
const [moveTargetId, setMoveTargetId] = useState<number | null>(null);
const [moveError, setMoveError] = useState<string | null>(null);
if (episodes.length === 0) return null;
@@ -69,22 +51,6 @@ export function EpisodeList({
onEpisodeDeleted?.();
};
const handleMoveEpisode = async (videoId: number, targetAnimeId: number) => {
setMoveTargetId(targetAnimeId);
setMoveError(null);
try {
const result = await apiClient.moveVideoToAnime(videoId, targetAnimeId);
setEpisodes((prev) => prev.filter((ep) => ep.videoId !== videoId));
if (expandedVideoId === videoId) setExpandedVideoId(null);
setMovingEpisode(null);
onEpisodeMoved?.(result.removedPreviousAnime);
} catch (err) {
setMoveError(err instanceof Error ? err.message : 'Failed to move this episode.');
} finally {
setMoveTargetId(null);
}
};
const watchedCount = episodes.filter((ep) => ep.watched).length;
return (
@@ -198,28 +164,14 @@ export function EpisodeList({
>
{'\u2713'}
</button>
<button
type="button"
onClick={(e) => {
e.stopPropagation();
setMoveError(null);
setMovingEpisode(ep);
}}
className={`w-5 h-5 rounded border border-ctp-surface2 text-transparent hover:border-ctp-blue/50 hover:text-ctp-blue focus-visible:text-ctp-blue hover:bg-ctp-blue/10 transition-colors text-xs flex items-center justify-center ${HOVER_REVEALED}`}
title="Move to another library entry"
aria-label="Move to another library entry"
>
{'\u2192'}
</button>
<button
type="button"
onClick={(e) => {
e.stopPropagation();
void handleDeleteEpisode(ep.videoId, ep.canonicalTitle);
}}
className={`w-5 h-5 rounded border border-ctp-surface2 text-transparent hover:border-ctp-red/50 hover:text-ctp-red focus-visible:text-ctp-red hover:bg-ctp-red/10 transition-colors text-xs flex items-center justify-center ${HOVER_REVEALED}`}
className="w-5 h-5 rounded border border-ctp-surface2 text-transparent hover:border-ctp-red/50 hover:text-ctp-red hover:bg-ctp-red/10 transition-colors opacity-0 group-hover:opacity-100 text-xs flex items-center justify-center"
title="Delete episode"
aria-label="Delete episode"
>
{'\u2715'}
</button>
@@ -239,19 +191,6 @@ export function EpisodeList({
</tbody>
</table>
</div>
{movingEpisode && (
<LibraryEntryPicker
heading={`Move "${movingEpisode.canonicalTitle}" To`}
excludeAnimeIds={animeId != null ? [animeId] : []}
busyAnimeId={moveTargetId}
error={moveError}
onSelect={(entry) => void handleMoveEpisode(movingEpisode.videoId, entry.animeId)}
onClose={() => {
setMovingEpisode(null);
setMoveError(null);
}}
/>
)}
</div>
);
}
@@ -1,159 +0,0 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { Window } from 'happy-dom';
import { act } from 'react';
import { createRoot } from 'react-dom/client';
import { apiClient } from '../../lib/api-client';
import type { AnimeEpisode, AnimeLibraryItem, StatsMoveVideoResponse } from '../../types/stats';
import { EpisodeList } from './EpisodeList';
interface TestWindow extends Window {
IS_REACT_ACT_ENVIRONMENT?: boolean;
}
function installDom(): () => void {
const previousWindow = globalThis.window;
const previousDocument = globalThis.document;
const previousHTMLElement = globalThis.HTMLElement;
const previousIsReactActEnvironment = (
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT;
const window = new Window() as TestWindow;
Object.defineProperty(globalThis, 'window', { value: window, configurable: true });
Object.defineProperty(globalThis, 'document', { value: window.document, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: window.HTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
return () => {
Object.defineProperty(globalThis, 'window', { value: previousWindow, configurable: true });
Object.defineProperty(globalThis, 'document', { value: previousDocument, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: previousHTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = previousIsReactActEnvironment;
};
}
function episode(videoId: number, title: string): AnimeEpisode {
return {
videoId,
episode: videoId,
season: null,
durationMs: 1_440_000,
endedMediaMs: null,
watched: 0,
canonicalTitle: title,
totalSessions: 1,
totalActiveMs: 1000,
totalCards: 0,
totalTokensSeen: 0,
totalYomitanLookupCount: 0,
lastWatchedMs: 1,
};
}
function libraryItem(animeId: number, title: string): AnimeLibraryItem {
return {
animeId,
canonicalTitle: title,
anilistId: null,
totalSessions: 1,
totalActiveMs: 1000,
totalCards: 0,
totalTokensSeen: 0,
episodeCount: 1,
episodesTotal: null,
lastWatchedMs: 1,
};
}
function findButtonByTitle(container: Element, title: string): HTMLElement {
const match = [...container.querySelectorAll('button')].find(
(button) => button.getAttribute('title') === title,
);
assert.ok(match, `expected a button titled "${title}"`);
return match as unknown as HTMLElement;
}
function findButtonByText(container: Element, text: string): HTMLElement {
const match = [...container.querySelectorAll('button')].find((button) =>
(button.textContent ?? '').includes(text),
);
assert.ok(match, `expected a "${text}" button`);
return match as unknown as HTMLElement;
}
test('EpisodeList moves an episode to the library entry picked in the dialog', async () => {
const uninstallDom = installDom();
const original = {
getAnimeLibrary: apiClient.getAnimeLibrary,
moveVideoToAnime: apiClient.moveVideoToAnime,
};
let moveCall: { videoId: number; animeId: number } | null = null;
let movedResult: boolean | null = null;
apiClient.getAnimeLibrary = (async () => [
libraryItem(1, 'Current Entry'),
libraryItem(2, 'Real Series'),
]) as typeof apiClient.getAnimeLibrary;
apiClient.moveVideoToAnime = (async (videoId: number, animeId: number) => {
moveCall = { videoId, animeId };
return {
ok: true,
animeId,
previousAnimeId: 1,
removedPreviousAnime: true,
} satisfies StatsMoveVideoResponse;
}) as typeof apiClient.moveVideoToAnime;
try {
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(
<EpisodeList
episodes={[episode(5, 'Stray Episode')]}
animeId={1}
onEpisodeMoved={(removedPreviousAnime) => {
movedResult = removedPreviousAnime;
}}
/>,
);
});
await act(async () => {
findButtonByTitle(container, 'Move to another library entry').click();
});
assert.match(container.textContent ?? '', /Move "Stray Episode" To/);
// The entry the episode already belongs to is not offered as a target.
assert.doesNotMatch(container.textContent ?? '', /Current Entry/);
await act(async () => {
findButtonByText(container, 'Real Series').click();
});
assert.deepEqual(moveCall, { videoId: 5, animeId: 2 });
assert.equal(movedResult, true);
// The row leaves this entry's list and the picker closes.
assert.doesNotMatch(container.textContent ?? '', /Stray Episode/);
await act(async () => {
root.unmount();
});
} finally {
Object.assign(apiClient, original);
uninstallDom();
}
});
@@ -1,168 +0,0 @@
import { useEffect, useId, useMemo, useRef, useState } from 'react';
import { apiClient } from '../../lib/api-client';
import { formatDuration } from '../../lib/formatters';
import { useModalFocus } from '../../hooks/useModalFocus';
import { AnimeCoverImage } from './AnimeCoverImage';
import type { AnimeLibraryItem } from '../../types/stats';
interface LibraryEntryPickerProps {
heading: string;
/** Entries that cannot be picked, typically the one being moved away from. */
excludeAnimeIds?: number[];
initialQuery?: string;
busyAnimeId?: number | null;
error?: string | null;
onSelect: (entry: AnimeLibraryItem) => void;
onClose: () => void;
}
export function LibraryEntryPicker({
heading,
excludeAnimeIds = [],
initialQuery = '',
busyAnimeId = null,
error = null,
onSelect,
onClose,
}: LibraryEntryPickerProps) {
const [entries, setEntries] = useState<AnimeLibraryItem[] | null>(null);
const [loadFailed, setLoadFailed] = useState(false);
const [query, setQuery] = useState(initialQuery);
const inputRef = useRef<HTMLInputElement>(null);
const dialogRef = useRef<HTMLDivElement>(null);
const headingId = useId();
const searchId = useId();
const busy = busyAnimeId !== null;
useEffect(() => {
let cancelled = false;
apiClient
.getAnimeLibrary()
.then((data) => {
if (!cancelled) setEntries(data);
})
.catch(() => {
// Distinct from an empty library: telling the user "no other titles"
// when the request failed hides a retryable error.
if (cancelled) return;
setEntries([]);
setLoadFailed(true);
});
return () => {
cancelled = true;
};
}, []);
useModalFocus({
dialogRef,
initialFocusRef: inputRef,
dismissDisabled: busy,
onDismiss: onClose,
});
const handleDismiss = () => {
if (!busy) onClose();
};
const excluded = useMemo(() => new Set(excludeAnimeIds), [excludeAnimeIds]);
const visible = useMemo(() => {
const term = query.trim().toLowerCase();
return (entries ?? [])
.filter((entry) => !excluded.has(entry.animeId))
.filter((entry) => !term || entry.canonicalTitle.toLowerCase().includes(term))
.sort((a, b) => b.lastWatchedMs - a.lastWatchedMs);
}, [entries, excluded, query]);
return (
<div
className="fixed inset-0 z-50 flex items-start justify-center pt-[10vh]"
onClick={handleDismiss}
>
<div className="absolute inset-0 bg-ctp-crust/70 backdrop-blur-[2px]" />
<div
ref={dialogRef}
role="dialog"
aria-modal="true"
aria-labelledby={headingId}
className="relative bg-ctp-base border border-ctp-surface1 rounded-xl shadow-2xl w-full max-w-lg max-h-[70vh] flex flex-col animate-fade-in"
onClick={(e) => e.stopPropagation()}
>
<div className="p-4 border-b border-ctp-surface1">
<div className="flex items-center justify-between mb-3">
<h3 id={headingId} className="text-sm font-semibold text-ctp-text">
{heading}
</h3>
<button
type="button"
onClick={handleDismiss}
disabled={busy}
aria-label="Close"
className="text-ctp-overlay2 hover:text-ctp-text text-lg leading-none disabled:opacity-50"
>
{'✕'}
</button>
</div>
<label htmlFor={searchId} className="sr-only">
Search library
</label>
<input
ref={inputRef}
id={searchId}
type="text"
value={query}
onChange={(e) => setQuery(e.target.value)}
placeholder="Search library..."
className="w-full bg-ctp-surface0 border border-ctp-surface1 rounded-lg px-3 py-2 text-sm text-ctp-text placeholder:text-ctp-overlay2 focus:outline-none focus:border-ctp-blue"
/>
{error ? (
<div role="alert" className="text-xs text-ctp-red mt-2">
{error}
</div>
) : null}
</div>
<div className="flex-1 overflow-y-auto p-2">
{entries === null && <div className="text-xs text-ctp-overlay2 p-3">Loading...</div>}
{loadFailed && (
<div role="alert" className="text-xs text-ctp-red p-3">
Could not load the library. Close this dialog and try again.
</div>
)}
{!loadFailed && entries !== null && visible.length === 0 && (
<div className="text-xs text-ctp-overlay2 p-3">
{query.trim() ? 'No matches' : 'No other titles'}
</div>
)}
{visible.map((entry) => (
<button
key={entry.animeId}
type="button"
disabled={busy}
onClick={() => onSelect(entry)}
className="w-full flex items-center gap-3 p-2.5 rounded-lg hover:bg-ctp-surface0 transition-colors text-left disabled:opacity-50"
>
<AnimeCoverImage
animeId={entry.animeId}
title={entry.canonicalTitle}
coverRetryToken={entry.anilistId ?? 0}
className="w-10 h-14 rounded shrink-0"
/>
<div className="min-w-0 flex-1">
<div className="text-sm text-ctp-text truncate">{entry.canonicalTitle}</div>
<div className="text-xs text-ctp-overlay2 mt-0.5">
{entry.episodeCount} episode{entry.episodeCount !== 1 ? 's' : ''} ·{' '}
{formatDuration(entry.totalActiveMs)}
</div>
</div>
{busyAnimeId === entry.animeId ? (
<span className="text-xs text-ctp-blue shrink-0">Moving...</span>
) : (
<span className="text-xs text-ctp-overlay2 shrink-0">Select</span>
)}
</button>
))}
</div>
</div>
</div>
);
}
@@ -0,0 +1,254 @@
import assert from 'node:assert/strict';
import test from 'node:test';
import { Window } from 'happy-dom';
import { act } from 'react';
import { createRoot } from 'react-dom/client';
import { apiClient } from '../../lib/api-client';
import type { StatsDuplicateLineCleanupResult } from '../../types/stats';
import { DuplicateLineCleanup } from './DuplicateLineCleanup';
interface TestWindow extends Window {
IS_REACT_ACT_ENVIRONMENT?: boolean;
}
function installDom(): () => void {
const previousWindow = globalThis.window;
const previousDocument = globalThis.document;
const previousHTMLElement = globalThis.HTMLElement;
const previousISReactActEnvironment = (
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT;
const window = new Window() as TestWindow;
Object.defineProperty(globalThis, 'window', { value: window, configurable: true });
Object.defineProperty(globalThis, 'document', { value: window.document, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: window.HTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = true;
return () => {
Object.defineProperty(globalThis, 'window', { value: previousWindow, configurable: true });
Object.defineProperty(globalThis, 'document', { value: previousDocument, configurable: true });
Object.defineProperty(globalThis, 'HTMLElement', {
value: previousHTMLElement,
configurable: true,
});
(
globalThis as typeof globalThis & { IS_REACT_ACT_ENVIRONMENT?: boolean }
).IS_REACT_ACT_ENVIRONMENT = previousISReactActEnvironment;
};
}
function findButton(container: Element, label: string): HTMLButtonElement {
const match = [...container.querySelectorAll('button')].find(
(button) => (button.textContent ?? '').trim() === label,
);
assert.ok(match, `expected a "${label}" button`);
return match as unknown as HTMLButtonElement;
}
/** The backdrop stays clickable during an apply, so it reaches the guard in `close`. */
function findBackdrop(container: Element): HTMLButtonElement {
const match = container.querySelector('button[aria-label="Close duplicate line cleanup"]');
assert.ok(match, 'expected the backdrop close button');
return match as unknown as HTMLButtonElement;
}
function deferred<T>(): { promise: Promise<T>; resolve: (value: T) => void } {
let resolve!: (value: T) => void;
const promise = new Promise<T>((done) => {
resolve = done;
});
return { promise, resolve };
}
function summary(
overrides: Partial<StatsDuplicateLineCleanupResult> = {},
): StatsDuplicateLineCleanupResult {
return {
dryRun: false,
lookbackDays: 30,
scannedLines: 900,
burstGroups: 2,
removedLines: 180,
removedWordOccurrences: 540,
removedKanjiOccurrences: 120,
samples: [],
...overrides,
};
}
interface Harness {
container: Element;
cleanedCalls: () => number;
closedCalls: () => number;
teardown: () => void;
}
async function mount(cleanup: (typeof apiClient)['cleanupDuplicateLines']): Promise<Harness> {
const uninstallDom = installDom();
const originalCleanup = apiClient.cleanupDuplicateLines;
apiClient.cleanupDuplicateLines = cleanup;
let cleaned = 0;
let closed = 0;
const container = document.createElement('div');
document.body.append(container);
const root = createRoot(container);
await act(async () => {
root.render(
<DuplicateLineCleanup
onClose={() => {
closed += 1;
}}
onCleaned={() => {
cleaned += 1;
}}
/>,
);
});
return {
container,
cleanedCalls: () => cleaned,
closedCalls: () => closed,
teardown: () => {
apiClient.cleanupDuplicateLines = originalCleanup;
uninstallDom();
},
};
}
test('a reload is still owed after a later scan replaces the applied result', async () => {
const harness = await mount(async ({ dryRun } = {}) => summary({ dryRun: dryRun === true }));
try {
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Clean Up').click();
});
assert.equal(harness.cleanedCalls(), 0, 'reload must wait for the result to be read');
// The follow-up scan clears the applied summary, but the rows are already gone.
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Close').click();
});
assert.equal(harness.cleanedCalls(), 1);
assert.equal(harness.closedCalls(), 1);
} finally {
harness.teardown();
}
});
test('a reload is still owed after the lookback window changes', async () => {
const harness = await mount(async ({ dryRun } = {}) => summary({ dryRun: dryRun === true }));
try {
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Clean Up').click();
});
await act(async () => {
findButton(harness.container, '7 days').click();
});
await act(async () => {
findButton(harness.container, 'Close').click();
});
assert.equal(harness.cleanedCalls(), 1);
} finally {
harness.teardown();
}
});
test('closing is refused while an apply is in flight', async () => {
const pending = deferred<StatsDuplicateLineCleanupResult>();
const harness = await mount(async ({ dryRun } = {}) =>
dryRun === true ? summary({ dryRun: true }) : pending.promise,
);
try {
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Clean Up').click();
});
assert.equal(findButton(harness.container, 'Close').disabled, true);
// The backdrop is never disabled, so this is the path that has to be refused.
await act(async () => {
findBackdrop(harness.container).click();
});
assert.equal(harness.closedCalls(), 0, 'the modal must stay open mid-apply');
assert.equal(harness.cleanedCalls(), 0);
await act(async () => {
pending.resolve(summary({ removedLines: 12 }));
await pending.promise;
});
await act(async () => {
findButton(harness.container, 'Close').click();
});
assert.equal(harness.closedCalls(), 1);
assert.equal(harness.cleanedCalls(), 1);
} finally {
harness.teardown();
}
});
test('a scan on its own owes no reload', async () => {
const harness = await mount(async ({ dryRun } = {}) => summary({ dryRun: dryRun === true }));
try {
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Close').click();
});
assert.equal(harness.cleanedCalls(), 0);
assert.equal(harness.closedCalls(), 1);
} finally {
harness.teardown();
}
});
test('an apply that removes nothing owes no reload', async () => {
// The scan saw work to do, but by the time it ran another cleanup had taken it.
const harness = await mount(async ({ dryRun } = {}) =>
dryRun === true ? summary({ dryRun: true }) : summary({ burstGroups: 0, removedLines: 0 }),
);
try {
await act(async () => {
findButton(harness.container, 'Scan').click();
});
await act(async () => {
findButton(harness.container, 'Clean Up').click();
});
await act(async () => {
findButton(harness.container, 'Close').click();
});
assert.equal(harness.cleanedCalls(), 0);
assert.equal(harness.closedCalls(), 1);
} finally {
harness.teardown();
}
});
@@ -0,0 +1,202 @@
import { useCallback, useState } from 'react';
import { getStatsClient } from '../../hooks/useStatsApi';
import { formatNumber } from '../../lib/formatters';
import type { StatsDuplicateLineCleanupResult } from '../../types/stats';
interface DuplicateLineCleanupProps {
onClose: () => void;
/** Called after rows are actually removed, so the charts can reload. */
onCleaned: () => void;
}
const LOOKBACK_CHOICES: Array<{ label: string; days: number | null }> = [
{ label: '7 days', days: 7 },
{ label: '30 days', days: 30 },
{ label: '90 days', days: 90 },
{ label: '1 year', days: 365 },
{ label: 'All time', days: null },
];
function formatTimecode(ms: number): string {
const totalSeconds = Math.max(0, Math.floor(ms / 1000));
const minutes = Math.floor(totalSeconds / 60);
const seconds = totalSeconds % 60;
return `${minutes}:${String(seconds).padStart(2, '0')}`;
}
export function DuplicateLineCleanup({ onClose, onCleaned }: DuplicateLineCleanupProps) {
const [lookbackDays, setLookbackDays] = useState<number | null>(30);
const [preview, setPreview] = useState<StatsDuplicateLineCleanupResult | null>(null);
const [applied, setApplied] = useState<StatsDuplicateLineCleanupResult | null>(null);
const [busy, setBusy] = useState<'scan' | 'apply' | null>(null);
const [error, setError] = useState<string | null>(null);
// Survives everything the displayed result does not: another scan, a different window.
// Rows are gone from the moment an apply succeeds, so the reload is owed until it runs.
const [needsReload, setNeedsReload] = useState(false);
const run = useCallback(
async (dryRun: boolean) => {
setBusy(dryRun ? 'scan' : 'apply');
setError(null);
try {
const result = await getStatsClient().cleanupDuplicateLines({ dryRun, lookbackDays });
if (dryRun) {
setPreview(result);
setApplied(null);
} else {
setApplied(result);
setPreview(null);
if (result.removedLines > 0) {
setNeedsReload(true);
}
}
} catch (cause) {
setError(cause instanceof Error ? cause.message : String(cause));
} finally {
setBusy(null);
}
},
[lookbackDays],
);
// Reloading the vocabulary tables unmounts this modal along with the rest of the tab,
// so it waits for the user to close: they get to read what was removed first. Closing
// is refused mid-apply, which would drop the reload on the floor along with the report.
const close = useCallback(() => {
if (busy === 'apply') {
return;
}
if (needsReload) {
onCleaned();
}
onClose();
}, [busy, needsReload, onCleaned, onClose]);
const result = applied ?? preview;
const nothingToDo = preview !== null && preview.removedLines === 0;
return (
<div className="fixed inset-0 z-50">
<button
type="button"
aria-label="Close duplicate line cleanup"
className="absolute inset-0 bg-ctp-crust/70 backdrop-blur-[2px]"
onClick={close}
/>
<div className="absolute inset-x-0 top-1/2 mx-auto max-w-xl -translate-y-1/2 rounded-xl border border-ctp-surface1 bg-ctp-mantle shadow-2xl">
<div className="flex items-center justify-between border-b border-ctp-surface1 px-5 py-4">
<h2 className="text-sm font-semibold text-ctp-text">Duplicate Lines</h2>
<button
type="button"
disabled={busy === 'apply'}
className="rounded-md border border-ctp-surface2 px-3 py-1.5 text-xs font-medium text-ctp-subtext0 transition hover:border-ctp-blue hover:text-ctp-blue disabled:opacity-50"
onClick={close}
>
Close
</button>
</div>
<div className="space-y-4 px-5 py-4">
<p className="text-xs leading-relaxed text-ctp-subtext0">
Typeset subtitles karaoke openings, animated signs are authored as one event per
animation frame, and older versions counted every frame as its own line. This finds
those runs and collapses each one back to a single line, giving back the word and kanji
counts they inflated. Ordinary repeated dialogue is left alone.
</p>
<div>
<div className="mb-2 text-xs font-medium text-ctp-subtext1">Look back over</div>
<div className="flex flex-wrap gap-2">
{LOOKBACK_CHOICES.map((choice) => (
<button
key={choice.label}
type="button"
disabled={busy !== null}
onClick={() => {
setLookbackDays(choice.days);
setPreview(null);
setApplied(null);
}}
className={`rounded-lg border px-3 py-1.5 text-xs transition disabled:opacity-50 ${
lookbackDays === choice.days
? 'border-ctp-blue/50 bg-ctp-surface2 text-ctp-text'
: 'border-ctp-surface1 bg-ctp-surface0 text-ctp-overlay2 hover:text-ctp-subtext0'
}`}
>
{choice.label}
</button>
))}
</div>
</div>
{error && (
<div className="rounded-lg border border-ctp-red/30 bg-ctp-red/10 px-3 py-2 text-xs text-ctp-red">
{error}
</div>
)}
{result && (
<div className="rounded-lg bg-ctp-surface0 px-4 py-3">
<div className="text-sm text-ctp-text">
{applied
? `Removed ${formatNumber(applied.removedLines)} repeated lines`
: nothingToDo
? 'No animation bursts found in this window'
: `Found ${formatNumber(preview!.burstGroups)} bursts covering ${formatNumber(preview!.removedLines)} extra lines`}
</div>
<div className="mt-1 text-xs text-ctp-overlay2">
{formatNumber(result.scannedLines)} lines scanned ·{' '}
{formatNumber(result.removedWordOccurrences)} word counts ·{' '}
{formatNumber(result.removedKanjiOccurrences)} kanji counts
{applied ? ' removed' : ' would be removed'}
</div>
{result.samples.length > 0 && (
<div className="mt-3 max-h-52 space-y-1.5 overflow-y-auto">
{result.samples.map((sample) => (
<div
key={`${sample.videoId}:${sample.startMs}:${sample.text}`}
className="flex items-center justify-between gap-3 rounded-md bg-ctp-mantle px-3 py-1.5"
>
<div className="min-w-0">
<div className="truncate text-xs text-ctp-text">{sample.text}</div>
<div className="truncate text-[11px] text-ctp-overlay1">
{sample.videoTitle ?? `Video ${sample.videoId}`} ·{' '}
{formatTimecode(sample.startMs)}
</div>
</div>
<span className="shrink-0 text-xs text-ctp-peach">×{sample.frames}</span>
</div>
))}
</div>
)}
</div>
)}
<div className="flex items-center justify-end gap-2">
<button
type="button"
disabled={busy !== null}
onClick={() => void run(true)}
className="rounded-md border border-ctp-surface2 px-3 py-1.5 text-xs font-medium text-ctp-subtext0 transition hover:border-ctp-blue hover:text-ctp-blue disabled:opacity-50"
>
{busy === 'scan' ? 'Scanning…' : 'Scan'}
</button>
<button
type="button"
disabled={busy !== null || preview === null || nothingToDo}
onClick={() => void run(false)}
className="rounded-md border border-ctp-red/30 px-3 py-1.5 text-xs font-medium text-ctp-red transition hover:bg-ctp-red/10 disabled:opacity-40"
>
{busy === 'apply' ? 'Cleaning…' : 'Clean Up'}
</button>
</div>
<p className="text-[11px] text-ctp-overlay1">
Scan first: cleanup removes rows and cannot be undone. Session watch time and lines-seen
totals are left untouched.
</p>
</div>
</div>
</div>
);
}
@@ -5,6 +5,7 @@ import { WordList } from './WordList';
import { KanjiBreakdown } from './KanjiBreakdown';
import { KanjiDetailPanel } from './KanjiDetailPanel';
import { ExclusionManager } from './ExclusionManager';
import { DuplicateLineCleanup } from './DuplicateLineCleanup';
import { formatNumber } from '../../lib/formatters';
import { TrendChart } from '../trends/TrendChart';
import { FrequencyRankTable } from './FrequencyRankTable';
@@ -34,10 +35,11 @@ export function VocabularyTab({
onRemoveExclusion,
onClearExclusions,
}: VocabularyTabProps) {
const { words, kanji, knownWords, loading, error } = useVocabulary();
const { words, kanji, knownWords, loading, error, reload } = useVocabulary();
const [selectedKanjiId, setSelectedKanjiId] = useState<number | null>(null);
const [hideNames, setHideNames] = useState(false);
const [showExclusionManager, setShowExclusionManager] = useState(false);
const [showDuplicateLineCleanup, setShowDuplicateLineCleanup] = useState(false);
const hasNames = useMemo(() => words.some(isProperNoun), [words]);
const filteredWords = useMemo(() => {
@@ -129,6 +131,13 @@ export function VocabularyTab({
Hide Names
</button>
)}
<button
type="button"
onClick={() => setShowDuplicateLineCleanup(true)}
className="shrink-0 rounded-lg border border-ctp-surface1 bg-ctp-surface0 px-3 py-2 text-xs text-ctp-overlay2 transition-colors hover:text-ctp-subtext0"
>
Duplicates
</button>
<button
type="button"
onClick={() => setShowExclusionManager(true)}
@@ -193,6 +202,13 @@ export function VocabularyTab({
onClose={() => setShowExclusionManager(false)}
/>
)}
{showDuplicateLineCleanup && (
<DuplicateLineCleanup
onClose={() => setShowDuplicateLineCleanup(false)}
onCleaned={reload}
/>
)}
</div>
);
}
+4 -71
View File
@@ -1,16 +1,11 @@
import { useCallback, useState, useEffect } from 'react';
import { getStatsClient } from './useStatsApi';
import type { AnimeLibraryItem, StatsAnimeMergeRecommendation } from '../types/stats';
const BACKGROUND_REFRESH_MS = 30_000;
import type { AnimeLibraryItem } from '../types/stats';
export function useAnimeLibrary() {
const [anime, setAnime] = useState<AnimeLibraryItem[]>([]);
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const [recommendations, setRecommendations] = useState<StatsAnimeMergeRecommendation[]>([]);
const [dismissingRecommendationId, setDismissingRecommendationId] = useState<number | null>(null);
const [recommendationActionError, setRecommendationActionError] = useState<string | null>(null);
const [reloadToken, setReloadToken] = useState(0);
const reload = useCallback(() => {
@@ -19,14 +14,10 @@ export function useAnimeLibrary() {
useEffect(() => {
let cancelled = false;
const client = getStatsClient();
client
getStatsClient()
.getAnimeLibrary()
.then((data) => {
if (!cancelled) {
setAnime(data);
setError(null);
}
if (!cancelled) setAnime(data);
})
.catch((err: Error) => {
if (!cancelled) setError(err.message);
@@ -34,68 +25,10 @@ export function useAnimeLibrary() {
.finally(() => {
if (!cancelled) setLoading(false);
});
// Recommendation support is deliberately non-blocking. An older backend
// should still be able to display its library even when this endpoint is
// unavailable.
client
.getAnimeMergeRecommendations()
.then((data) => {
if (!cancelled) {
setRecommendations(data.recommendations);
setRecommendationActionError(null);
}
})
.catch(() => {
// Preserve the last confirmed set. A transient polling failure should
// not make a pending review silently disappear.
});
return () => {
cancelled = true;
};
}, [reloadToken]);
useEffect(() => {
const refreshOnFocus = () => reload();
const interval = window.setInterval(reload, BACKGROUND_REFRESH_MS);
window.addEventListener('focus', refreshOnFocus);
return () => {
window.clearInterval(interval);
window.removeEventListener('focus', refreshOnFocus);
};
}, [reload]);
const dismissRecommendation = useCallback(async (recommendationId: number) => {
setDismissingRecommendationId(recommendationId);
setRecommendationActionError(null);
try {
await getStatsClient().dismissAnimeMergeRecommendation(recommendationId);
setRecommendations((current) =>
current.filter((item) => item.recommendationId !== recommendationId),
);
} catch {
setRecommendationActionError('Could not dismiss this suggestion. Try again.');
} finally {
setDismissingRecommendationId(null);
}
}, []);
const clearRecommendation = useCallback((recommendationId: number) => {
setRecommendations((current) =>
current.filter((item) => item.recommendationId !== recommendationId),
);
setRecommendationActionError(null);
}, []);
return {
anime,
loading,
error,
reload,
recommendations,
dismissRecommendation,
dismissingRecommendationId,
recommendationActionError,
clearRecommendation,
};
return { anime, loading, error, reload };
}
-73
View File
@@ -1,73 +0,0 @@
import { useEffect, useRef, type RefObject } from 'react';
const FOCUSABLE_SELECTOR = [
'button:not([disabled])',
'input:not([disabled])',
'select:not([disabled])',
'textarea:not([disabled])',
'a[href]',
'[tabindex]:not([tabindex="-1"])',
].join(',');
interface UseModalFocusOptions {
dialogRef: RefObject<HTMLElement | null>;
initialFocusRef: RefObject<HTMLElement | null>;
dismissDisabled?: boolean;
onDismiss: () => void;
}
export function useModalFocus({
dialogRef,
initialFocusRef,
dismissDisabled = false,
onDismiss,
}: UseModalFocusOptions): void {
const dismissDisabledRef = useRef(dismissDisabled);
const onDismissRef = useRef(onDismiss);
dismissDisabledRef.current = dismissDisabled;
onDismissRef.current = onDismiss;
useEffect(() => {
const previouslyFocused =
document.activeElement instanceof HTMLElement ? document.activeElement : null;
initialFocusRef.current?.focus();
const handleKeyDown = (event: KeyboardEvent) => {
if (event.key === 'Escape') {
if (!dismissDisabledRef.current) {
event.preventDefault();
onDismissRef.current();
}
return;
}
if (event.key !== 'Tab') return;
const dialog = dialogRef.current;
if (!dialog) return;
const focusable = [...dialog.querySelectorAll<HTMLElement>(FOCUSABLE_SELECTOR)];
const first = focusable[0];
const last = focusable.at(-1);
if (!first || !last) return;
if (
event.shiftKey &&
(document.activeElement === first || !dialog.contains(document.activeElement))
) {
event.preventDefault();
last.focus();
} else if (
!event.shiftKey &&
(document.activeElement === last || !dialog.contains(document.activeElement))
) {
event.preventDefault();
first.focus();
}
};
document.addEventListener('keydown', handleKeyDown);
return () => {
document.removeEventListener('keydown', handleKeyDown);
previouslyFocused?.focus();
};
}, [dialogRef, initialFocusRef]);
}
+6 -3
View File
@@ -1,4 +1,4 @@
import { useState, useEffect } from 'react';
import { useState, useEffect, useCallback } from 'react';
import { getStatsClient } from './useStatsApi';
import type { VocabularyEntry, KanjiEntry } from '../types/stats';
@@ -8,6 +8,9 @@ export function useVocabulary() {
const [knownWords, setKnownWords] = useState<Set<string>>(new Set());
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
// Bumped by `reload` after maintenance rewrites the vocabulary tables.
const [reloadToken, setReloadToken] = useState(0);
const reload = useCallback(() => setReloadToken((token) => token + 1), []);
useEffect(() => {
let cancelled = false;
@@ -46,7 +49,7 @@ export function useVocabulary() {
return () => {
cancelled = true;
};
}, []);
}, [reloadToken]);
return { words, kanji, knownWords, loading, error };
return { words, kanji, knownWords, loading, error, reload };
}
-39
View File
@@ -44,45 +44,6 @@ test('getAnimeCoverUrl appends retry tokens for late cover refreshes', () => {
);
});
test('getAnimeMergeRecommendations loads pending duplicate pairs', async () => {
const originalFetch = globalThis.fetch;
let seenUrl = '';
globalThis.fetch = (async (input: RequestInfo | URL) => {
seenUrl = String(input);
return new Response(
JSON.stringify({ recommendations: [{ recommendationId: 4, animeIds: [7, 8] }] }),
{ status: 200, headers: { 'Content-Type': 'application/json' } },
);
}) as typeof globalThis.fetch;
try {
const result = await apiClient.getAnimeMergeRecommendations();
assert.equal(seenUrl, `${BASE_URL}/api/stats/anime/merge-recommendations`);
assert.deepEqual(result.recommendations, [{ recommendationId: 4, animeIds: [7, 8] }]);
} finally {
globalThis.fetch = originalFetch;
}
});
test('dismissAnimeMergeRecommendation sends DELETE for the selected suggestion', async () => {
const originalFetch = globalThis.fetch;
let seenUrl = '';
let seenMethod = '';
globalThis.fetch = (async (input: RequestInfo | URL, init?: RequestInit) => {
seenUrl = String(input);
seenMethod = init?.method ?? 'GET';
return new Response(JSON.stringify({ ok: true }), { status: 200 });
}) as typeof globalThis.fetch;
try {
await apiClient.dismissAnimeMergeRecommendation(4);
assert.equal(seenUrl, `${BASE_URL}/api/stats/anime/merge-recommendations/4`);
assert.equal(seenMethod, 'DELETE');
} finally {
globalThis.fetch = originalFetch;
}
});
test('getCoverImages batches anime and media cover requests', async () => {
const originalFetch = globalThis.fetch;
let seenUrl = '';
+15 -30
View File
@@ -6,13 +6,11 @@ import type {
StatsAnkiNotesInfoRequest,
StatsCoverImagesRequest,
StatsDeleteSessionsRequest,
StatsDuplicateLineCleanupRequest,
StatsDuplicateLineCleanupResult,
StatsExcludedWordsRequest,
StatsHttpClient,
StatsJsonResponseMap,
StatsMergeAnimeRequest,
StatsMergeAnimeResponse,
StatsMoveVideoRequest,
StatsMoveVideoResponse,
StatsTrendGroupBy,
StatsTrendRange,
StatsVideoWatchedRequest,
@@ -106,6 +104,19 @@ export const apiClient = {
body: JSON.stringify({ words } satisfies StatsExcludedWordsRequest),
});
},
cleanupDuplicateLines: async (
options: StatsDuplicateLineCleanupRequest = {},
): Promise<StatsDuplicateLineCleanupResult> => {
const res = await fetchResponse('/api/stats/maintenance/duplicate-lines', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
dryRun: options.dryRun === true,
lookbackDays: options.lookbackDays ?? null,
} satisfies StatsDuplicateLineCleanupRequest),
});
return res.json() as Promise<StatsDuplicateLineCleanupResult>;
},
getWordOccurrences: (headword: string, word: string, reading: string, limit = 50, offset = 0) =>
fetchJson(
'wordOccurrences',
@@ -129,8 +140,6 @@ export const apiClient = {
getMediaLibrary: () => fetchJson('mediaLibrary', '/api/stats/media'),
getMediaDetail: (videoId: number) => fetchJson('mediaDetail', `/api/stats/media/${videoId}`),
getAnimeLibrary: () => fetchJson('animeLibrary', '/api/stats/anime'),
getAnimeMergeRecommendations: () =>
fetchJson('animeMergeRecommendations', '/api/stats/anime/merge-recommendations'),
getAnimeDetail: (animeId: number) => fetchJson('animeDetail', `/api/stats/anime/${animeId}`),
getAnimeWords: (animeId: number, limit = 50) =>
fetchJson('animeWords', `/api/stats/anime/${animeId}/words?limit=${limit}`),
@@ -200,30 +209,6 @@ export const apiClient = {
fetchResponse(`/api/stats/anime/${animeId}`, { method: 'DELETE' }),
);
},
mergeAnime: async (
targetAnimeId: number,
sourceAnimeIds: number[],
): Promise<StatsMergeAnimeResponse> => {
const res = await fetchResponse(`/api/stats/anime/${targetAnimeId}/merge`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ sourceAnimeIds } satisfies StatsMergeAnimeRequest),
});
return res.json() as Promise<StatsMergeAnimeResponse>;
},
moveVideoToAnime: async (videoId: number, animeId: number): Promise<StatsMoveVideoResponse> => {
const res = await fetchResponse(`/api/stats/media/${videoId}/anime`, {
method: 'PATCH',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ animeId } satisfies StatsMoveVideoRequest),
});
return res.json() as Promise<StatsMoveVideoResponse>;
},
dismissAnimeMergeRecommendation: async (recommendationId: number): Promise<void> => {
await fetchResponse(`/api/stats/anime/merge-recommendations/${recommendationId}`, {
method: 'DELETE',
});
},
getKnownWords: () => fetchJson('knownWords', '/api/stats/known-words'),
getKnownWordsSummary: () => fetchJson('knownWordsSummary', '/api/stats/known-words-summary'),
getAnimeKnownWordsSummary: (animeId: number) =>