mirror of
https://github.com/ksyasuda/SubMiner.git
synced 2026-08-13 01:55:50 -07:00
fix(stats): stop counting duplicate typeset subtitle lines
- Collapse animation-burst subtitle lines (karaoke OPs, animated signs) at ingest time using the same dedup rules the subtitle sidebar already applies, so repeated frames no longer flood "Top Repeated Words" - Add retroactive cleanup for stats already affected: a "Duplicates" scanner/cleaner in the Vocabulary tab and `subminer stats cleanup --duplicate-lines` (`--dry-run`, `--lookback-days`) on the CLI - Only subtitle lines and the vocabulary counts they feed are touched; watch time and lines-seen totals are left as recorded
This commit is contained in:
@@ -0,0 +1,5 @@
|
|||||||
|
type: fixed
|
||||||
|
area: stats
|
||||||
|
|
||||||
|
- Typeset subtitles no longer flood the stats. Karaoke openings and animated signs are authored as one subtitle event per animation frame, and immersion tracking counted every frame, which was enough to put an OP lyric at the top of "Top Repeated Words" for good. Lines are now collapsed on the way in using the same rules the subtitle sidebar already applies: when the active subtitle file has been parsed, stats record exactly the cues the sidebar shows, and for sources with no parsed cue list a run of identical, contiguous, sub-0.1s lines stops counting after a few frames. Ordinary repeated dialogue and rewatches are unaffected.
|
||||||
|
- Added a cleanup for stats already affected. The Vocabulary tab has a **Duplicates** button that scans a chosen window (7 days through all time), shows the bursts it found and the word and kanji counts they added, and collapses each run to one line once confirmed. `subminer stats cleanup --duplicate-lines` does the same from the terminal, with `--dry-run` and `--lookback-days <n>`. Only subtitle lines and the vocabulary counts they feed are touched; watch time and lines-seen totals are left as recorded.
|
||||||
@@ -34,7 +34,7 @@ The same immersion data powers the stats dashboard.
|
|||||||
- In-app overlay: focus the visible overlay, then press the key from `stats.toggleKey` (default: `` ` `` / `Backquote`).
|
- In-app overlay: focus the visible overlay, then press the key from `stats.toggleKey` (default: `` ` `` / `Backquote`).
|
||||||
- Launcher command: run `subminer stats` to start the local stats server on demand (it also opens the dashboard in your browser when `stats.autoOpenBrowser` is enabled; the default is `false`).
|
- Launcher command: run `subminer stats` to start the local stats server on demand (it also opens the dashboard in your browser when `stats.autoOpenBrowser` is enabled; the default is `false`).
|
||||||
- Background server: run `subminer stats -b` to start or reuse a dedicated background stats daemon without keeping the launcher attached, and `subminer stats -s` to stop that daemon.
|
- Background server: run `subminer stats -b` to start or reuse a dedicated background stats daemon without keeping the launcher attached, and `subminer stats -s` to stop that daemon.
|
||||||
- Maintenance commands: run `subminer stats cleanup` or `subminer stats cleanup -v` to backfill/repair vocabulary metadata (`headword`, `reading`, POS) and purge stale or excluded rows from `imm_words` on demand; `subminer stats cleanup -l` repairs lifetime summary tables. `subminer stats rebuild` and `subminer stats backfill` rebuild or backfill rollup data.
|
- Maintenance commands: run `subminer stats cleanup` or `subminer stats cleanup -v` to backfill/repair vocabulary metadata (`headword`, `reading`, POS) and purge stale or excluded rows from `imm_words` on demand; `subminer stats cleanup -l` repairs lifetime summary tables; `subminer stats cleanup --duplicate-lines` collapses repeated lines left behind by typeset subtitles (see [Repeated Line Cleanup](#repeated-line-cleanup)). `subminer stats rebuild` and `subminer stats backfill` rebuild or backfill rollup data.
|
||||||
- Browser page: open `http://127.0.0.1:6969` directly if the local stats server is already running.
|
- Browser page: open `http://127.0.0.1:6969` directly if the local stats server is already running.
|
||||||
|
|
||||||
### Dashboard Tabs
|
### Dashboard Tabs
|
||||||
@@ -125,6 +125,30 @@ Secondary subtitle text (typically English translations) is stored alongside pri
|
|||||||
|
|
||||||
The Vocabulary tab toolbar includes an **Exclusions** button for hiding words from all vocabulary views. Excluded words are stored in the immersion database, with older browser localStorage exclusions imported on first load after upgrade. They can be managed (restored or cleared) from the exclusion modal. Exclusions affect stat cards, charts, the frequency rank table, and the word list.
|
The Vocabulary tab toolbar includes an **Exclusions** button for hiding words from all vocabulary views. Excluded words are stored in the immersion database, with older browser localStorage exclusions imported on first load after upgrade. They can be managed (restored or cleared) from the exclusion modal. Exclusions affect stat cards, charts, the frequency rank table, and the word list.
|
||||||
|
|
||||||
|
### Repeated Line Cleanup
|
||||||
|
|
||||||
|
Karaoke openings and animated signs are authored as one subtitle event per animation frame, all carrying the same text. Playback reports every one of those frames, so a single OP lyric could be recorded hundreds of times and dominate "Top Repeated Words".
|
||||||
|
|
||||||
|
Recording now collapses those runs as they happen, matching what the subtitle sidebar shows:
|
||||||
|
|
||||||
|
- When the active subtitle source has been parsed, its cue list has already had duplicate events and animation bursts merged. A line landing inside a surviving cue but after that cue's start is a frame the sidebar merged away, and is not recorded.
|
||||||
|
- Otherwise only timing is available, so the strict metadata-free rule applies: a run of identical, contiguous lines each shorter than 0.1s stops being recorded after a few frames. Ordinary repeated dialogue, and lines held for a normal beat, always record.
|
||||||
|
|
||||||
|
For stats recorded before this, the Vocabulary tab toolbar has a **Duplicates** button:
|
||||||
|
|
||||||
|
- Pick how far back to look (7 days, 30 days, 90 days, 1 year, or all time). A narrower window does less work and keeps older history untouched.
|
||||||
|
- **Scan** reports the bursts found, the lines they added, and the word and kanji counts they inflated, without writing anything.
|
||||||
|
- **Clean Up** applies exactly what the scan reported: each run collapses to its first line (extended to cover the run), and the removed lines' word and kanji occurrences are subtracted from the vocabulary aggregates.
|
||||||
|
|
||||||
|
The same thing runs from the terminal:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
subminer stats cleanup --duplicate-lines --dry-run --lookback-days 30
|
||||||
|
subminer stats cleanup --duplicate-lines --lookback-days 30
|
||||||
|
```
|
||||||
|
|
||||||
|
Runs never cross a session boundary, so rewatching an episode keeps both watches. Session telemetry (watch time, lines seen, tokens seen) and the rollups derived from it are left as recorded: they are cumulative samples taken during playback, and cannot be recomputed for sessions whose raw rows have since been pruned.
|
||||||
|
|
||||||
## Retention Defaults
|
## Retention Defaults
|
||||||
|
|
||||||
By default, SubMiner keeps all retention tables and raw data (`0` means keep all) while continuing daily/monthly rollup maintenance:
|
By default, SubMiner keeps all retention tables and raw data (`0` means keep all) while continuing daily/monthly rollup maintenance:
|
||||||
|
|||||||
@@ -151,6 +151,7 @@ subminer stats -b # start background stats daemon
|
|||||||
| `subminer stats` | Start the stats server (opens the dashboard when `stats.autoOpenBrowser` is on) |
|
| `subminer stats` | Start the stats server (opens the dashboard when `stats.autoOpenBrowser` is on) |
|
||||||
| `subminer stats -b` / `-s` | Start/reuse or stop the background stats daemon |
|
| `subminer stats -b` / `-s` | Start/reuse or stop the background stats daemon |
|
||||||
| `subminer stats cleanup` | Backfill vocabulary metadata and prune stale rows (`-v` vocab, `-l` lifetime summaries) |
|
| `subminer stats cleanup` | Backfill vocabulary metadata and prune stale rows (`-v` vocab, `-l` lifetime summaries) |
|
||||||
|
| `subminer stats cleanup -d` | Collapse repeated lines from typeset subs (`--dry-run`, `--lookback-days <n>`) |
|
||||||
| `subminer stats rebuild` / `backfill` | Rebuild or backfill rollup data |
|
| `subminer stats rebuild` / `backfill` | Rebuild or backfill rollup data |
|
||||||
| `subminer doctor` | Dependency + config + socket diagnostics (`--refresh-known-words` refreshes the known-word cache) |
|
| `subminer doctor` | Dependency + config + socket diagnostics (`--refresh-known-words` refreshes the known-word cache) |
|
||||||
| `subminer settings` | Open the SubMiner settings window |
|
| `subminer settings` | Open the SubMiner settings window |
|
||||||
|
|||||||
@@ -95,6 +95,7 @@ subminer texthooker # Texthooker-only mode (-o also opens the brow
|
|||||||
subminer stats -b # Start/reuse the background stats daemon
|
subminer stats -b # Start/reuse the background stats daemon
|
||||||
subminer stats -s # Stop the background stats daemon
|
subminer stats -s # Stop the background stats daemon
|
||||||
subminer stats cleanup # Backfill vocabulary metadata, prune stale rows
|
subminer stats cleanup # Backfill vocabulary metadata, prune stale rows
|
||||||
|
subminer stats cleanup -d --dry-run # Preview cleanup of repeated typeset subtitle lines
|
||||||
subminer stats rebuild # Rebuild rollup data
|
subminer stats rebuild # Rebuild rollup data
|
||||||
subminer doctor --refresh-known-words # Refresh the known-word cache
|
subminer doctor --refresh-known-words # Refresh the known-word cache
|
||||||
subminer logs -e # Export a sanitized log ZIP and print its path
|
subminer logs -e # Export a sanitized log ZIP and print its path
|
||||||
|
|||||||
@@ -157,6 +157,15 @@ export async function runStatsCommand(
|
|||||||
if (args.statsCleanupLifetime) {
|
if (args.statsCleanupLifetime) {
|
||||||
forwarded.push('--stats-cleanup-lifetime');
|
forwarded.push('--stats-cleanup-lifetime');
|
||||||
}
|
}
|
||||||
|
if (args.statsCleanupDuplicateLines) {
|
||||||
|
forwarded.push('--stats-cleanup-duplicate-lines');
|
||||||
|
}
|
||||||
|
if (args.statsCleanupDryRun) {
|
||||||
|
forwarded.push('--stats-cleanup-dry-run');
|
||||||
|
}
|
||||||
|
if (args.statsCleanupLookbackDays) {
|
||||||
|
forwarded.push('--stats-cleanup-lookback-days', String(args.statsCleanupLookbackDays));
|
||||||
|
}
|
||||||
if (shouldForwardLogLevel(args.logLevel)) {
|
if (shouldForwardLogLevel(args.logLevel)) {
|
||||||
forwarded.push('--log-level', args.logLevel);
|
forwarded.push('--log-level', args.logLevel);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -134,6 +134,9 @@ test('applyInvocationsToArgs maps config and jellyfin invocation state', () => {
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
|
statsCleanupLookbackDays: null,
|
||||||
statsLogLevel: null,
|
statsLogLevel: null,
|
||||||
syncTriggered: false,
|
syncTriggered: false,
|
||||||
syncCliTokens: [],
|
syncCliTokens: [],
|
||||||
@@ -185,6 +188,9 @@ test('applyInvocationsToArgs maps settings invocation to settings window', () =>
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
|
statsCleanupLookbackDays: null,
|
||||||
statsLogLevel: null,
|
statsLogLevel: null,
|
||||||
syncTriggered: false,
|
syncTriggered: false,
|
||||||
syncCliTokens: [],
|
syncCliTokens: [],
|
||||||
@@ -229,6 +235,9 @@ test('applyInvocationsToArgs fails when config invocation has no action', () =>
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
|
statsCleanupLookbackDays: null,
|
||||||
statsLogLevel: null,
|
statsLogLevel: null,
|
||||||
syncTriggered: false,
|
syncTriggered: false,
|
||||||
syncCliTokens: [],
|
syncCliTokens: [],
|
||||||
@@ -271,6 +280,9 @@ test('applyInvocationsToArgs maps texthooker browser-open request', () => {
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
|
statsCleanupLookbackDays: null,
|
||||||
statsLogLevel: null,
|
statsLogLevel: null,
|
||||||
syncTriggered: false,
|
syncTriggered: false,
|
||||||
syncCliTokens: [],
|
syncCliTokens: [],
|
||||||
|
|||||||
@@ -162,6 +162,8 @@ export function createDefaultArgs(
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
doctor: false,
|
doctor: false,
|
||||||
doctorRefreshKnownWords: false,
|
doctorRefreshKnownWords: false,
|
||||||
logsExport: false,
|
logsExport: false,
|
||||||
@@ -258,6 +260,11 @@ export function applyInvocationsToArgs(parsed: Args, invocations: CliInvocations
|
|||||||
if (invocations.statsCleanup) parsed.statsCleanup = true;
|
if (invocations.statsCleanup) parsed.statsCleanup = true;
|
||||||
if (invocations.statsCleanupVocab) parsed.statsCleanupVocab = true;
|
if (invocations.statsCleanupVocab) parsed.statsCleanupVocab = true;
|
||||||
if (invocations.statsCleanupLifetime) parsed.statsCleanupLifetime = true;
|
if (invocations.statsCleanupLifetime) parsed.statsCleanupLifetime = true;
|
||||||
|
if (invocations.statsCleanupDuplicateLines) parsed.statsCleanupDuplicateLines = true;
|
||||||
|
if (invocations.statsCleanupDryRun) parsed.statsCleanupDryRun = true;
|
||||||
|
if (invocations.statsCleanupLookbackDays !== null) {
|
||||||
|
parsed.statsCleanupLookbackDays = invocations.statsCleanupLookbackDays;
|
||||||
|
}
|
||||||
if (invocations.dictionaryTarget) {
|
if (invocations.dictionaryTarget) {
|
||||||
parsed.dictionaryTarget = parseDictionaryTarget(invocations.dictionaryTarget);
|
parsed.dictionaryTarget = parseDictionaryTarget(invocations.dictionaryTarget);
|
||||||
} else if (
|
} else if (
|
||||||
|
|||||||
@@ -37,6 +37,9 @@ export interface CliInvocations {
|
|||||||
statsCleanup: boolean;
|
statsCleanup: boolean;
|
||||||
statsCleanupVocab: boolean;
|
statsCleanupVocab: boolean;
|
||||||
statsCleanupLifetime: boolean;
|
statsCleanupLifetime: boolean;
|
||||||
|
statsCleanupDuplicateLines: boolean;
|
||||||
|
statsCleanupDryRun: boolean;
|
||||||
|
statsCleanupLookbackDays: number | null;
|
||||||
statsLogLevel: string | null;
|
statsLogLevel: string | null;
|
||||||
syncTriggered: boolean;
|
syncTriggered: boolean;
|
||||||
syncCliTokens: string[];
|
syncCliTokens: string[];
|
||||||
@@ -53,6 +56,16 @@ export interface CliInvocations {
|
|||||||
texthookerOpenBrowser: boolean;
|
texthookerOpenBrowser: boolean;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** `--lookback-days` narrows the duplicate-line cleanup; anything unusable means no limit. */
|
||||||
|
function parseStatsLookbackDays(value: unknown): number | null {
|
||||||
|
if (typeof value !== 'string' && typeof value !== 'number') return null;
|
||||||
|
const days = Number(value);
|
||||||
|
if (!Number.isFinite(days) || days <= 0) {
|
||||||
|
throw new Error('Stats --lookback-days must be a positive number of days.');
|
||||||
|
}
|
||||||
|
return Math.floor(days);
|
||||||
|
}
|
||||||
|
|
||||||
function applyRootOptions(program: Command): void {
|
function applyRootOptions(program: Command): void {
|
||||||
program
|
program
|
||||||
.option(
|
.option(
|
||||||
@@ -169,6 +182,9 @@ export function parseCliPrograms(
|
|||||||
let statsCleanup = false;
|
let statsCleanup = false;
|
||||||
let statsCleanupVocab = false;
|
let statsCleanupVocab = false;
|
||||||
let statsCleanupLifetime = false;
|
let statsCleanupLifetime = false;
|
||||||
|
let statsCleanupDuplicateLines = false;
|
||||||
|
let statsCleanupDryRun = false;
|
||||||
|
let statsCleanupLookbackDays: number | null = null;
|
||||||
let statsLogLevel: string | null = null;
|
let statsLogLevel: string | null = null;
|
||||||
let syncTriggered = false;
|
let syncTriggered = false;
|
||||||
let syncCliTokens: string[] = [];
|
let syncCliTokens: string[] = [];
|
||||||
@@ -269,6 +285,9 @@ export function parseCliPrograms(
|
|||||||
.option('-s, --stop', 'Stop the background stats server')
|
.option('-s, --stop', 'Stop the background stats server')
|
||||||
.option('-v, --vocab', 'Clean vocabulary rows in the stats database')
|
.option('-v, --vocab', 'Clean vocabulary rows in the stats database')
|
||||||
.option('-l, --lifetime', 'Rebuild lifetime summary rows from retained data')
|
.option('-l, --lifetime', 'Rebuild lifetime summary rows from retained data')
|
||||||
|
.option('-d, --duplicate-lines', 'Collapse repeated subtitle lines from typeset animations')
|
||||||
|
.option('--dry-run', 'Report what a cleanup would remove without changing anything')
|
||||||
|
.option('--lookback-days <days>', 'Only clean lines recorded in the last N days')
|
||||||
.option('--log-level <level>', 'Log level')
|
.option('--log-level <level>', 'Log level')
|
||||||
.action((action: string | undefined, options: Record<string, unknown>) => {
|
.action((action: string | undefined, options: Record<string, unknown>) => {
|
||||||
statsTriggered = true;
|
statsTriggered = true;
|
||||||
@@ -289,13 +308,29 @@ export function parseCliPrograms(
|
|||||||
if (normalizedAction && (statsBackground || statsStop)) {
|
if (normalizedAction && (statsBackground || statsStop)) {
|
||||||
throw new Error('Stats background and stop flags cannot be combined with stats actions.');
|
throw new Error('Stats background and stop flags cannot be combined with stats actions.');
|
||||||
}
|
}
|
||||||
if (normalizedAction !== 'cleanup' && (options.vocab === true || options.lifetime === true)) {
|
if (
|
||||||
throw new Error('Stats --vocab and --lifetime flags require the cleanup action.');
|
normalizedAction !== 'cleanup' &&
|
||||||
|
(options.vocab === true || options.lifetime === true || options.duplicateLines === true)
|
||||||
|
) {
|
||||||
|
throw new Error(
|
||||||
|
'Stats --vocab, --lifetime and --duplicate-lines flags require the cleanup action.',
|
||||||
|
);
|
||||||
|
}
|
||||||
|
if (options.duplicateLines !== true && (options.dryRun === true || options.lookbackDays)) {
|
||||||
|
throw new Error('Stats --dry-run and --lookback-days require --duplicate-lines.');
|
||||||
}
|
}
|
||||||
if (normalizedAction === 'cleanup') {
|
if (normalizedAction === 'cleanup') {
|
||||||
statsCleanup = true;
|
statsCleanup = true;
|
||||||
statsCleanupLifetime = options.lifetime === true;
|
statsCleanupLifetime = options.lifetime === true;
|
||||||
statsCleanupVocab = statsCleanupLifetime ? false : options.vocab !== false;
|
statsCleanupDuplicateLines = options.duplicateLines === true;
|
||||||
|
if (statsCleanupLifetime && statsCleanupDuplicateLines) {
|
||||||
|
throw new Error('Stats cleanup runs one mode at a time.');
|
||||||
|
}
|
||||||
|
// Vocabulary cleanup stays the default so `stats cleanup` keeps its old meaning.
|
||||||
|
statsCleanupVocab =
|
||||||
|
statsCleanupLifetime || statsCleanupDuplicateLines ? false : options.vocab !== false;
|
||||||
|
statsCleanupDryRun = options.dryRun === true;
|
||||||
|
statsCleanupLookbackDays = parseStatsLookbackDays(options.lookbackDays);
|
||||||
} else if (normalizedAction === 'rebuild' || normalizedAction === 'backfill') {
|
} else if (normalizedAction === 'rebuild' || normalizedAction === 'backfill') {
|
||||||
statsCleanup = true;
|
statsCleanup = true;
|
||||||
statsCleanupLifetime = true;
|
statsCleanupLifetime = true;
|
||||||
@@ -483,6 +518,9 @@ export function parseCliPrograms(
|
|||||||
statsCleanup,
|
statsCleanup,
|
||||||
statsCleanupVocab,
|
statsCleanupVocab,
|
||||||
statsCleanupLifetime,
|
statsCleanupLifetime,
|
||||||
|
statsCleanupDuplicateLines,
|
||||||
|
statsCleanupDryRun,
|
||||||
|
statsCleanupLookbackDays,
|
||||||
statsLogLevel,
|
statsLogLevel,
|
||||||
syncTriggered,
|
syncTriggered,
|
||||||
syncCliTokens,
|
syncCliTokens,
|
||||||
|
|||||||
@@ -232,6 +232,29 @@ test('parseArgs maps lifetime stats cleanup flag', () => {
|
|||||||
assert.equal(parsed.statsCleanupLifetime, true);
|
assert.equal(parsed.statsCleanupLifetime, true);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test('parseArgs maps duplicate-line stats cleanup flags', () => {
|
||||||
|
const parsed = parseArgs(
|
||||||
|
['stats', 'cleanup', '--duplicate-lines', '--dry-run', '--lookback-days', '30'],
|
||||||
|
'subminer',
|
||||||
|
{},
|
||||||
|
);
|
||||||
|
|
||||||
|
assert.equal(parsed.statsCleanup, true);
|
||||||
|
assert.equal(parsed.statsCleanupVocab, false);
|
||||||
|
assert.equal(parsed.statsCleanupDuplicateLines, true);
|
||||||
|
assert.equal(parsed.statsCleanupDryRun, true);
|
||||||
|
assert.equal(parsed.statsCleanupLookbackDays, 30);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('parseArgs rejects duplicate-line flags without the duplicate-lines mode', () => {
|
||||||
|
const error = withProcessExitIntercept(() => {
|
||||||
|
parseArgs(['stats', 'cleanup', '--dry-run'], 'subminer', {});
|
||||||
|
});
|
||||||
|
|
||||||
|
assert.equal(error.code, 1);
|
||||||
|
assert.match(error.stderr, /--dry-run and --lookback-days require --duplicate-lines/);
|
||||||
|
});
|
||||||
|
|
||||||
test('parseArgs rejects cleanup-only stats flags without cleanup action', () => {
|
test('parseArgs rejects cleanup-only stats flags without cleanup action', () => {
|
||||||
const error = withProcessExitIntercept(() => {
|
const error = withProcessExitIntercept(() => {
|
||||||
parseArgs(['stats', '--vocab'], 'subminer', {});
|
parseArgs(['stats', '--vocab'], 'subminer', {});
|
||||||
@@ -239,7 +262,10 @@ test('parseArgs rejects cleanup-only stats flags without cleanup action', () =>
|
|||||||
|
|
||||||
assert.equal(error.code, 1);
|
assert.equal(error.code, 1);
|
||||||
assert.match(error.message, /exit:1/);
|
assert.match(error.message, /exit:1/);
|
||||||
assert.match(error.stderr, /Stats --vocab and --lifetime flags require the cleanup action/);
|
assert.match(
|
||||||
|
error.stderr,
|
||||||
|
/Stats --vocab, --lifetime and --duplicate-lines flags require the cleanup action/,
|
||||||
|
);
|
||||||
});
|
});
|
||||||
|
|
||||||
test('parseArgs maps stats rebuild action to cleanup lifetime mode', () => {
|
test('parseArgs maps stats rebuild action to cleanup lifetime mode', () => {
|
||||||
|
|||||||
@@ -142,6 +142,9 @@ export interface Args {
|
|||||||
statsCleanup?: boolean;
|
statsCleanup?: boolean;
|
||||||
statsCleanupVocab?: boolean;
|
statsCleanupVocab?: boolean;
|
||||||
statsCleanupLifetime?: boolean;
|
statsCleanupLifetime?: boolean;
|
||||||
|
statsCleanupDuplicateLines?: boolean;
|
||||||
|
statsCleanupDryRun?: boolean;
|
||||||
|
statsCleanupLookbackDays?: number;
|
||||||
dictionaryTarget?: string;
|
dictionaryTarget?: string;
|
||||||
doctor: boolean;
|
doctor: boolean;
|
||||||
doctorRefreshKnownWords: boolean;
|
doctorRefreshKnownWords: boolean;
|
||||||
|
|||||||
+14
-1
@@ -64,6 +64,9 @@ export interface CliArgs {
|
|||||||
statsCleanup?: boolean;
|
statsCleanup?: boolean;
|
||||||
statsCleanupVocab?: boolean;
|
statsCleanupVocab?: boolean;
|
||||||
statsCleanupLifetime?: boolean;
|
statsCleanupLifetime?: boolean;
|
||||||
|
statsCleanupDuplicateLines?: boolean;
|
||||||
|
statsCleanupDryRun?: boolean;
|
||||||
|
statsCleanupLookbackDays?: number;
|
||||||
statsResponsePath?: string;
|
statsResponsePath?: string;
|
||||||
jellyfin: boolean;
|
jellyfin: boolean;
|
||||||
jellyfinLogin: boolean;
|
jellyfinLogin: boolean;
|
||||||
@@ -167,6 +170,8 @@ export function parseArgs(argv: string[]): CliArgs {
|
|||||||
statsCleanup: false,
|
statsCleanup: false,
|
||||||
statsCleanupVocab: false,
|
statsCleanupVocab: false,
|
||||||
statsCleanupLifetime: false,
|
statsCleanupLifetime: false,
|
||||||
|
statsCleanupDuplicateLines: false,
|
||||||
|
statsCleanupDryRun: false,
|
||||||
jellyfin: false,
|
jellyfin: false,
|
||||||
jellyfinLogin: false,
|
jellyfinLogin: false,
|
||||||
jellyfinLogout: false,
|
jellyfinLogout: false,
|
||||||
@@ -368,7 +373,15 @@ export function parseArgs(argv: string[]): CliArgs {
|
|||||||
} else if (arg === '--stats-cleanup') args.statsCleanup = true;
|
} else if (arg === '--stats-cleanup') args.statsCleanup = true;
|
||||||
else if (arg === '--stats-cleanup-vocab') args.statsCleanupVocab = true;
|
else if (arg === '--stats-cleanup-vocab') args.statsCleanupVocab = true;
|
||||||
else if (arg === '--stats-cleanup-lifetime') args.statsCleanupLifetime = true;
|
else if (arg === '--stats-cleanup-lifetime') args.statsCleanupLifetime = true;
|
||||||
else if (arg.startsWith('--stats-response-path=')) {
|
else if (arg === '--stats-cleanup-duplicate-lines') args.statsCleanupDuplicateLines = true;
|
||||||
|
else if (arg === '--stats-cleanup-dry-run') args.statsCleanupDryRun = true;
|
||||||
|
else if (arg.startsWith('--stats-cleanup-lookback-days=')) {
|
||||||
|
const value = Number(arg.split('=', 2)[1]);
|
||||||
|
if (Number.isFinite(value) && value > 0) args.statsCleanupLookbackDays = Math.floor(value);
|
||||||
|
} else if (arg === '--stats-cleanup-lookback-days') {
|
||||||
|
const value = Number(readValue(argv[i + 1]));
|
||||||
|
if (Number.isFinite(value) && value > 0) args.statsCleanupLookbackDays = Math.floor(value);
|
||||||
|
} else if (arg.startsWith('--stats-response-path=')) {
|
||||||
const value = arg.split('=', 2)[1];
|
const value = arg.split('=', 2)[1];
|
||||||
if (value) args.statsResponsePath = value;
|
if (value) args.statsResponsePath = value;
|
||||||
} else if (arg === '--stats-response-path') {
|
} else if (arg === '--stats-response-path') {
|
||||||
|
|||||||
@@ -1032,6 +1032,64 @@ describe('stats server API routes', () => {
|
|||||||
]);
|
]);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
it('POST /api/stats/maintenance/duplicate-lines forwards the window and dry-run flag', async () => {
|
||||||
|
let seenOptions: unknown = null;
|
||||||
|
const summary = {
|
||||||
|
dryRun: true,
|
||||||
|
lookbackDays: 30,
|
||||||
|
scannedLines: 900,
|
||||||
|
burstGroups: 2,
|
||||||
|
removedLines: 180,
|
||||||
|
removedWordOccurrences: 540,
|
||||||
|
removedKanjiOccurrences: 120,
|
||||||
|
samples: [],
|
||||||
|
};
|
||||||
|
const app = createStatsApp(
|
||||||
|
createMockTracker({
|
||||||
|
cleanupDuplicateSubtitleLines: async (options: unknown) => {
|
||||||
|
seenOptions = options;
|
||||||
|
return summary;
|
||||||
|
},
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
|
||||||
|
const res = await app.request('/api/stats/maintenance/duplicate-lines', {
|
||||||
|
method: 'POST',
|
||||||
|
headers: { 'Content-Type': 'application/json' },
|
||||||
|
body: JSON.stringify({ dryRun: true, lookbackDays: 30 }),
|
||||||
|
});
|
||||||
|
|
||||||
|
assert.equal(res.status, 200);
|
||||||
|
assert.deepEqual(await res.json(), summary);
|
||||||
|
assert.deepEqual(seenOptions, { dryRun: true, lookbackDays: 30 });
|
||||||
|
});
|
||||||
|
|
||||||
|
it('POST /api/stats/maintenance/duplicate-lines treats a missing body as an apply over all history', async () => {
|
||||||
|
let seenOptions: unknown = null;
|
||||||
|
const app = createStatsApp(
|
||||||
|
createMockTracker({
|
||||||
|
cleanupDuplicateSubtitleLines: async (options: unknown) => {
|
||||||
|
seenOptions = options;
|
||||||
|
return {
|
||||||
|
dryRun: false,
|
||||||
|
lookbackDays: null,
|
||||||
|
scannedLines: 0,
|
||||||
|
burstGroups: 0,
|
||||||
|
removedLines: 0,
|
||||||
|
removedWordOccurrences: 0,
|
||||||
|
removedKanjiOccurrences: 0,
|
||||||
|
samples: [],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
}),
|
||||||
|
);
|
||||||
|
|
||||||
|
const res = await app.request('/api/stats/maintenance/duplicate-lines', { method: 'POST' });
|
||||||
|
|
||||||
|
assert.equal(res.status, 200);
|
||||||
|
assert.deepEqual(seenOptions, { dryRun: false, lookbackDays: null });
|
||||||
|
});
|
||||||
|
|
||||||
it('PUT /api/stats/excluded-words rejects malformed rows', async () => {
|
it('PUT /api/stats/excluded-words rejects malformed rows', async () => {
|
||||||
const app = createStatsApp(createMockTracker());
|
const app = createStatsApp(createMockTracker());
|
||||||
|
|
||||||
|
|||||||
@@ -91,6 +91,11 @@ import {
|
|||||||
markVideoWatched,
|
markVideoWatched,
|
||||||
upsertCoverArt,
|
upsertCoverArt,
|
||||||
} from './immersion-tracker/query-maintenance';
|
} from './immersion-tracker/query-maintenance';
|
||||||
|
import {
|
||||||
|
cleanupDuplicateSubtitleLines,
|
||||||
|
type DuplicateSubtitleLineCleanupOptions,
|
||||||
|
type DuplicateSubtitleLineCleanupSummary,
|
||||||
|
} from './immersion-tracker/duplicate-line-cleanup';
|
||||||
import { repairJellyfinStreamVideoLinks } from './immersion-tracker/jellyfin-link-repair';
|
import { repairJellyfinStreamVideoLinks } from './immersion-tracker/jellyfin-link-repair';
|
||||||
import {
|
import {
|
||||||
repairLegacySeasonlessAnimeRows,
|
repairLegacySeasonlessAnimeRows,
|
||||||
@@ -595,6 +600,18 @@ export class ImmersionTrackerService {
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Collapse animation bursts that earlier versions recorded frame by frame. Pending
|
||||||
|
* writes are flushed first so a burst that is still queued is scanned as stored rows
|
||||||
|
* rather than surviving the cleanup and reappearing seconds later.
|
||||||
|
*/
|
||||||
|
async cleanupDuplicateSubtitleLines(
|
||||||
|
options: DuplicateSubtitleLineCleanupOptions = {},
|
||||||
|
): Promise<DuplicateSubtitleLineCleanupSummary> {
|
||||||
|
this.flushNow();
|
||||||
|
return cleanupDuplicateSubtitleLines(this.db, options);
|
||||||
|
}
|
||||||
|
|
||||||
async rebuildLifetimeSummaries(): Promise<LifetimeRebuildSummary> {
|
async rebuildLifetimeSummaries(): Promise<LifetimeRebuildSummary> {
|
||||||
this.flushTelemetry(true);
|
this.flushTelemetry(true);
|
||||||
this.flushNow();
|
this.flushNow();
|
||||||
|
|||||||
@@ -0,0 +1,279 @@
|
|||||||
|
import assert from 'node:assert/strict';
|
||||||
|
import fs from 'node:fs';
|
||||||
|
import os from 'node:os';
|
||||||
|
import path from 'node:path';
|
||||||
|
import test from 'node:test';
|
||||||
|
import { Database } from '../sqlite.js';
|
||||||
|
import type { DatabaseSync } from '../sqlite.js';
|
||||||
|
import { ensureSchema } from '../storage.js';
|
||||||
|
import { cleanupDuplicateSubtitleLines } from '../duplicate-line-cleanup.js';
|
||||||
|
|
||||||
|
const DAY_MS = 86_400_000;
|
||||||
|
const BASE_MS = 1_700_000_000_000;
|
||||||
|
const WORD_ID = 1;
|
||||||
|
|
||||||
|
interface SeedLine {
|
||||||
|
session: number;
|
||||||
|
text: string;
|
||||||
|
startMs: number;
|
||||||
|
endMs: number;
|
||||||
|
/** Recording wall-clock, i.e. what the lookback window filters on. */
|
||||||
|
createdMs?: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
function makeDbPath(): string {
|
||||||
|
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'subminer-duplicate-line-test-'));
|
||||||
|
return path.join(dir, 'immersion.sqlite');
|
||||||
|
}
|
||||||
|
|
||||||
|
function cleanupDbPath(dbPath: string): void {
|
||||||
|
const dir = path.dirname(dbPath);
|
||||||
|
if (!fs.existsSync(dir)) return;
|
||||||
|
fs.rmSync(dir, { recursive: true, force: true });
|
||||||
|
}
|
||||||
|
|
||||||
|
/** One episode, two sessions of it, and one word occurrence per seeded line. */
|
||||||
|
function seed(db: DatabaseSync, lines: SeedLine[]): void {
|
||||||
|
db.exec(`
|
||||||
|
INSERT INTO imm_anime(anime_id, normalized_title_key, canonical_title, CREATED_DATE, LAST_UPDATE_DATE)
|
||||||
|
VALUES (1, 'show', 'Show', ${BASE_MS}, ${BASE_MS});
|
||||||
|
INSERT INTO imm_videos(video_id, video_key, anime_id, canonical_title, source_type, watched, duration_ms, CREATED_DATE, LAST_UPDATE_DATE)
|
||||||
|
VALUES (1, 'v1', 1, 'Ep 1', 1, 1, 1440000, ${BASE_MS}, ${BASE_MS});
|
||||||
|
INSERT INTO imm_sessions(session_id, session_uuid, video_id, started_at_ms, ended_at_ms, status, CREATED_DATE, LAST_UPDATE_DATE)
|
||||||
|
VALUES (1, 's1', 1, '${BASE_MS}', '${BASE_MS + 1000}', 2, ${BASE_MS}, ${BASE_MS}),
|
||||||
|
(2, 's2', 1, '${BASE_MS + DAY_MS}', '${BASE_MS + DAY_MS + 1000}', 2, ${BASE_MS}, ${BASE_MS});
|
||||||
|
INSERT INTO imm_words(id, headword, word, reading, part_of_speech, pos1, first_seen, last_seen, frequency)
|
||||||
|
VALUES (${WORD_ID}, '飛び上がる', '飛び上がる', '', 'verb', '動詞', ${Math.floor(BASE_MS / 1000)}, ${Math.floor(BASE_MS / 1000)}, 0);
|
||||||
|
`);
|
||||||
|
|
||||||
|
const insertLine = db.prepare(
|
||||||
|
`INSERT INTO imm_subtitle_lines(
|
||||||
|
line_id, session_id, video_id, anime_id, line_index,
|
||||||
|
segment_start_ms, segment_end_ms, text, CREATED_DATE, LAST_UPDATE_DATE)
|
||||||
|
VALUES (?, ?, 1, 1, ?, ?, ?, ?, ?, ?)`,
|
||||||
|
);
|
||||||
|
const insertOccurrence = db.prepare(
|
||||||
|
`INSERT INTO imm_word_line_occurrences(line_id, word_id, occurrence_count, seen_ms)
|
||||||
|
VALUES (?, ?, 1, ?)`,
|
||||||
|
);
|
||||||
|
|
||||||
|
lines.forEach((line, index) => {
|
||||||
|
const lineId = index + 1;
|
||||||
|
const createdMs = line.createdMs ?? BASE_MS;
|
||||||
|
insertLine.run(
|
||||||
|
lineId,
|
||||||
|
line.session,
|
||||||
|
lineId,
|
||||||
|
line.startMs,
|
||||||
|
line.endMs,
|
||||||
|
line.text,
|
||||||
|
createdMs,
|
||||||
|
createdMs,
|
||||||
|
);
|
||||||
|
insertOccurrence.run(lineId, WORD_ID, createdMs);
|
||||||
|
});
|
||||||
|
|
||||||
|
db.exec(`
|
||||||
|
UPDATE imm_words SET frequency = (
|
||||||
|
SELECT COALESCE(SUM(o.occurrence_count), 0)
|
||||||
|
FROM imm_word_line_occurrences o WHERE o.word_id = imm_words.id
|
||||||
|
)
|
||||||
|
`);
|
||||||
|
}
|
||||||
|
|
||||||
|
function createDb(lines: SeedLine[]): { db: DatabaseSync; dbPath: string } {
|
||||||
|
const dbPath = makeDbPath();
|
||||||
|
const db = new Database(dbPath);
|
||||||
|
ensureSchema(db);
|
||||||
|
seed(db, lines);
|
||||||
|
return { db, dbPath };
|
||||||
|
}
|
||||||
|
|
||||||
|
/** A typeset line mpv reported once per animation frame. */
|
||||||
|
function karaokeFrames(
|
||||||
|
session: number,
|
||||||
|
text: string,
|
||||||
|
startMs: number,
|
||||||
|
frames: number,
|
||||||
|
frameMs: number,
|
||||||
|
): SeedLine[] {
|
||||||
|
return Array.from({ length: frames }, (_, index) => ({
|
||||||
|
session,
|
||||||
|
text,
|
||||||
|
startMs: startMs + index * frameMs,
|
||||||
|
endMs: startMs + (index + 1) * frameMs,
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
|
function countLines(db: DatabaseSync): number {
|
||||||
|
return (db.prepare('SELECT COUNT(*) AS total FROM imm_subtitle_lines').get() as { total: number })
|
||||||
|
.total;
|
||||||
|
}
|
||||||
|
|
||||||
|
function wordFrequency(db: DatabaseSync): number {
|
||||||
|
const row = db.prepare('SELECT frequency FROM imm_words WHERE id = ?').get(WORD_ID) as {
|
||||||
|
frequency: number;
|
||||||
|
} | null;
|
||||||
|
return row?.frequency ?? 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
test('a karaoke burst collapses to one line and gives back its word counts', () => {
|
||||||
|
const { db, dbPath } = createDb([
|
||||||
|
...karaokeFrames(1, '飛び上がる', 10_000, 40, 40),
|
||||||
|
{ session: 1, text: 'おはよう', startMs: 20_000, endMs: 22_000 },
|
||||||
|
]);
|
||||||
|
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db);
|
||||||
|
|
||||||
|
assert.equal(summary.burstGroups, 1);
|
||||||
|
assert.equal(summary.removedLines, 39);
|
||||||
|
assert.equal(summary.removedWordOccurrences, 39);
|
||||||
|
assert.equal(countLines(db), 2);
|
||||||
|
assert.equal(wordFrequency(db), 2);
|
||||||
|
|
||||||
|
// The surviving line covers the whole run, the way the parsed cue would.
|
||||||
|
const kept = db
|
||||||
|
.prepare(
|
||||||
|
'SELECT segment_start_ms AS startMs, segment_end_ms AS endMs FROM imm_subtitle_lines WHERE line_id = 1',
|
||||||
|
)
|
||||||
|
.get() as { startMs: number; endMs: number };
|
||||||
|
assert.equal(kept.startMs, 10_000);
|
||||||
|
assert.equal(kept.endMs, 10_000 + 40 * 40);
|
||||||
|
|
||||||
|
assert.equal(summary.samples.length, 1);
|
||||||
|
assert.equal(summary.samples[0]!.text, '飛び上がる');
|
||||||
|
assert.equal(summary.samples[0]!.frames, 40);
|
||||||
|
assert.equal(summary.samples[0]!.videoTitle, 'Ep 1');
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('ordinary repeated dialogue survives', () => {
|
||||||
|
// Six contiguous `飛び上がる`, each held for a normal beat rather than a frame.
|
||||||
|
const lines = Array.from({ length: 6 }, (_, index) => ({
|
||||||
|
session: 1,
|
||||||
|
text: '飛び上がる',
|
||||||
|
startMs: 5_000 + index * 800,
|
||||||
|
endMs: 5_000 + (index + 1) * 800,
|
||||||
|
}));
|
||||||
|
const { db, dbPath } = createDb(lines);
|
||||||
|
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db);
|
||||||
|
|
||||||
|
assert.equal(summary.burstGroups, 0);
|
||||||
|
assert.equal(summary.removedLines, 0);
|
||||||
|
assert.equal(countLines(db), 6);
|
||||||
|
assert.equal(wordFrequency(db), 6);
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('a short run below the threshold survives', () => {
|
||||||
|
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 1_000, 4, 40));
|
||||||
|
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db);
|
||||||
|
|
||||||
|
assert.equal(summary.burstGroups, 0);
|
||||||
|
assert.equal(countLines(db), 4);
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('the same line in a rewatch session is never merged into the first watch', () => {
|
||||||
|
const { db, dbPath } = createDb([
|
||||||
|
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40),
|
||||||
|
...karaokeFrames(2, '飛び上がる', 10_000, 6, 40),
|
||||||
|
]);
|
||||||
|
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db);
|
||||||
|
|
||||||
|
assert.equal(summary.burstGroups, 2);
|
||||||
|
assert.equal(summary.removedLines, 10);
|
||||||
|
// One surviving line per session, not one across both.
|
||||||
|
assert.equal(countLines(db), 2);
|
||||||
|
assert.equal(wordFrequency(db), 2);
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('a gap between runs splits them', () => {
|
||||||
|
const { db, dbPath } = createDb([
|
||||||
|
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40),
|
||||||
|
...karaokeFrames(1, '飛び上がる', 60_000, 6, 40),
|
||||||
|
]);
|
||||||
|
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db);
|
||||||
|
|
||||||
|
assert.equal(summary.burstGroups, 2);
|
||||||
|
assert.equal(countLines(db), 2);
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('a dry run reports what an apply would do and writes nothing', () => {
|
||||||
|
const { db, dbPath } = createDb(karaokeFrames(1, '飛び上がる', 10_000, 40, 40));
|
||||||
|
|
||||||
|
try {
|
||||||
|
const preview = cleanupDuplicateSubtitleLines(db, { dryRun: true });
|
||||||
|
|
||||||
|
assert.equal(preview.dryRun, true);
|
||||||
|
assert.equal(preview.removedLines, 39);
|
||||||
|
assert.equal(countLines(db), 40);
|
||||||
|
assert.equal(wordFrequency(db), 40);
|
||||||
|
|
||||||
|
const applied = cleanupDuplicateSubtitleLines(db);
|
||||||
|
assert.equal(applied.removedLines, preview.removedLines);
|
||||||
|
assert.equal(applied.removedWordOccurrences, preview.removedWordOccurrences);
|
||||||
|
assert.equal(countLines(db), 1);
|
||||||
|
} finally {
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('the lookback window leaves older bursts alone', () => {
|
||||||
|
const recentMs = BASE_MS;
|
||||||
|
const oldMs = BASE_MS - 40 * DAY_MS;
|
||||||
|
const { db, dbPath } = createDb([
|
||||||
|
...karaokeFrames(1, '飛び上がる', 10_000, 6, 40).map((line) => ({
|
||||||
|
...line,
|
||||||
|
createdMs: oldMs,
|
||||||
|
})),
|
||||||
|
...karaokeFrames(2, '飛び上がる', 10_000, 6, 40).map((line) => ({
|
||||||
|
...line,
|
||||||
|
createdMs: recentMs,
|
||||||
|
})),
|
||||||
|
]);
|
||||||
|
|
||||||
|
globalThis.__subminerTestNowMs = BASE_MS;
|
||||||
|
try {
|
||||||
|
const summary = cleanupDuplicateSubtitleLines(db, { lookbackDays: 30 });
|
||||||
|
|
||||||
|
assert.equal(summary.lookbackDays, 30);
|
||||||
|
assert.equal(summary.scannedLines, 6);
|
||||||
|
assert.equal(summary.burstGroups, 1);
|
||||||
|
assert.equal(summary.removedLines, 5);
|
||||||
|
// Six untouched old frames plus the one surviving recent line.
|
||||||
|
assert.equal(countLines(db), 7);
|
||||||
|
assert.equal(wordFrequency(db), 7);
|
||||||
|
} finally {
|
||||||
|
globalThis.__subminerTestNowMs = undefined;
|
||||||
|
db.close();
|
||||||
|
cleanupDbPath(dbPath);
|
||||||
|
}
|
||||||
|
});
|
||||||
@@ -0,0 +1,366 @@
|
|||||||
|
/*
|
||||||
|
* Retroactive removal of animation-burst subtitle lines from the stats database.
|
||||||
|
*
|
||||||
|
* Before the live ingest gate existed, a karaoke OP recorded one line -- and one count
|
||||||
|
* for every word in it -- per animation frame, which is enough to put an OP lyric at the
|
||||||
|
* top of "Top Repeated Words" for good. This module finds those runs in what is already
|
||||||
|
* stored and takes them back down to one line.
|
||||||
|
*
|
||||||
|
* Only timing is available here: the stored text has been stripped of ASS markup, so the
|
||||||
|
* authoring evidence the file-level parser uses (`\t`, `\move`, karaoke timing, a
|
||||||
|
* changing override signature) is long gone. The rule is therefore the strict,
|
||||||
|
* metadata-free one -- a long run of identical, contiguous, short-lived lines inside a
|
||||||
|
* single session -- and every bound is configurable so a cautious run can ask for more.
|
||||||
|
*
|
||||||
|
* Scope: subtitle lines, their word/kanji occurrences, and the `imm_words`/`imm_kanji`
|
||||||
|
* aggregates those occurrences feed. Session telemetry (`lines_seen`, `tokens_seen`) and
|
||||||
|
* the rollups derived from it are left alone; they are cumulative samples taken at record
|
||||||
|
* time, and for sessions whose raw rows have since been pruned they cannot be recomputed.
|
||||||
|
*/
|
||||||
|
|
||||||
|
import type { DatabaseSync } from './sqlite';
|
||||||
|
import {
|
||||||
|
ANIMATION_FRAME_MAX_SECONDS,
|
||||||
|
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
|
||||||
|
MIN_TIMING_ONLY_FRAMES,
|
||||||
|
} from '../subtitle-burst-constants';
|
||||||
|
import {
|
||||||
|
applyLexicalRemovals,
|
||||||
|
makePlaceholders,
|
||||||
|
planLexicalRemovalsForLines,
|
||||||
|
toDbTimestamp,
|
||||||
|
} from './query-shared';
|
||||||
|
import { nowMs } from './time';
|
||||||
|
|
||||||
|
const MS_PER_DAY = 86_400_000;
|
||||||
|
/** SQLite caps bound parameters per statement; stay well under it. */
|
||||||
|
const LINE_ID_BATCH_SIZE = 400;
|
||||||
|
const DEFAULT_SAMPLE_LIMIT = 20;
|
||||||
|
|
||||||
|
export interface DuplicateSubtitleLineCleanupOptions {
|
||||||
|
/** Only consider lines recorded within this many days. Null or omitted = all history. */
|
||||||
|
lookbackDays?: number | null;
|
||||||
|
/** Measure without writing. */
|
||||||
|
dryRun?: boolean;
|
||||||
|
/** Identical contiguous lines needed before a run counts as an animation. */
|
||||||
|
minRunLength?: number;
|
||||||
|
/** Longest a single event may last and still look like an animation frame. */
|
||||||
|
maxFrameSeconds?: number;
|
||||||
|
/** How many of the largest runs to describe in the summary. */
|
||||||
|
sampleLimit?: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface DuplicateSubtitleLineBurst {
|
||||||
|
sessionId: number;
|
||||||
|
videoId: number;
|
||||||
|
text: string;
|
||||||
|
/** Kept line, extended to cover the whole run. */
|
||||||
|
keptLineId: number;
|
||||||
|
removedLineIds: number[];
|
||||||
|
startMs: number;
|
||||||
|
endMs: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface DuplicateSubtitleLineSample {
|
||||||
|
videoId: number;
|
||||||
|
videoTitle: string | null;
|
||||||
|
text: string;
|
||||||
|
frames: number;
|
||||||
|
removedLines: number;
|
||||||
|
startMs: number;
|
||||||
|
endMs: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface DuplicateSubtitleLineCleanupSummary {
|
||||||
|
dryRun: boolean;
|
||||||
|
lookbackDays: number | null;
|
||||||
|
scannedLines: number;
|
||||||
|
burstGroups: number;
|
||||||
|
removedLines: number;
|
||||||
|
removedWordOccurrences: number;
|
||||||
|
removedKanjiOccurrences: number;
|
||||||
|
samples: DuplicateSubtitleLineSample[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StoredSubtitleLineRow {
|
||||||
|
lineId: number;
|
||||||
|
sessionId: number;
|
||||||
|
videoId: number;
|
||||||
|
text: string;
|
||||||
|
startMs: number;
|
||||||
|
endMs: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface ResolvedBounds {
|
||||||
|
lookbackDays: number | null;
|
||||||
|
minRunLength: number;
|
||||||
|
maxFrameMs: number;
|
||||||
|
gapToleranceMs: number;
|
||||||
|
sampleLimit: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
function resolveBounds(options: DuplicateSubtitleLineCleanupOptions): ResolvedBounds {
|
||||||
|
const lookbackDays =
|
||||||
|
typeof options.lookbackDays === 'number' && Number.isFinite(options.lookbackDays)
|
||||||
|
? Math.max(1, Math.floor(options.lookbackDays))
|
||||||
|
: null;
|
||||||
|
const minRunLength =
|
||||||
|
typeof options.minRunLength === 'number' && Number.isFinite(options.minRunLength)
|
||||||
|
? Math.max(2, Math.floor(options.minRunLength))
|
||||||
|
: MIN_TIMING_ONLY_FRAMES;
|
||||||
|
const maxFrameSeconds =
|
||||||
|
typeof options.maxFrameSeconds === 'number' && options.maxFrameSeconds > 0
|
||||||
|
? options.maxFrameSeconds
|
||||||
|
: ANIMATION_FRAME_MAX_SECONDS;
|
||||||
|
const sampleLimit =
|
||||||
|
typeof options.sampleLimit === 'number' && options.sampleLimit >= 0
|
||||||
|
? Math.floor(options.sampleLimit)
|
||||||
|
: DEFAULT_SAMPLE_LIMIT;
|
||||||
|
return {
|
||||||
|
lookbackDays,
|
||||||
|
minRunLength,
|
||||||
|
maxFrameMs: Math.round(maxFrameSeconds * 1000),
|
||||||
|
gapToleranceMs: Math.round(DUPLICATE_CUE_GAP_TOLERANCE_SECONDS * 1000),
|
||||||
|
sampleLimit,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* `CREATED_DATE` holds epoch milliseconds on rows this app wrote, but older and synced
|
||||||
|
* rows can carry seconds, so normalize before comparing against the cutoff.
|
||||||
|
*/
|
||||||
|
const CREATED_MS_SQL = `
|
||||||
|
CASE
|
||||||
|
WHEN sl.CREATED_DATE < 10000000000 THEN sl.CREATED_DATE * 1000
|
||||||
|
ELSE sl.CREATED_DATE
|
||||||
|
END`;
|
||||||
|
|
||||||
|
function readCandidateLines(db: DatabaseSync, bounds: ResolvedBounds): StoredSubtitleLineRow[] {
|
||||||
|
const scope =
|
||||||
|
bounds.lookbackDays === null
|
||||||
|
? ''
|
||||||
|
: `AND sl.CREATED_DATE IS NOT NULL AND ${CREATED_MS_SQL} >= ?`;
|
||||||
|
const params = bounds.lookbackDays === null ? [] : [nowMs() - bounds.lookbackDays * MS_PER_DAY];
|
||||||
|
return db
|
||||||
|
.prepare(
|
||||||
|
`SELECT
|
||||||
|
sl.line_id AS lineId,
|
||||||
|
sl.session_id AS sessionId,
|
||||||
|
sl.video_id AS videoId,
|
||||||
|
sl.text AS text,
|
||||||
|
sl.segment_start_ms AS startMs,
|
||||||
|
sl.segment_end_ms AS endMs
|
||||||
|
FROM imm_subtitle_lines sl
|
||||||
|
WHERE sl.segment_start_ms IS NOT NULL
|
||||||
|
AND sl.segment_end_ms IS NOT NULL
|
||||||
|
${scope}
|
||||||
|
ORDER BY sl.session_id, sl.video_id, sl.segment_start_ms, sl.line_id`,
|
||||||
|
)
|
||||||
|
.all(...params) as StoredSubtitleLineRow[];
|
||||||
|
}
|
||||||
|
|
||||||
|
function isBurst(run: StoredSubtitleLineRow[], bounds: ResolvedBounds): boolean {
|
||||||
|
if (run.length < bounds.minRunLength) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return run.every((row) => row.endMs - row.startMs <= bounds.maxFrameMs);
|
||||||
|
}
|
||||||
|
|
||||||
|
function toBurst(run: StoredSubtitleLineRow[]): DuplicateSubtitleLineBurst {
|
||||||
|
const [first] = run;
|
||||||
|
return {
|
||||||
|
sessionId: first!.sessionId,
|
||||||
|
videoId: first!.videoId,
|
||||||
|
text: first!.text,
|
||||||
|
keptLineId: first!.lineId,
|
||||||
|
removedLineIds: run.slice(1).map((row) => row.lineId),
|
||||||
|
startMs: first!.startMs,
|
||||||
|
endMs: run.reduce((latest, row) => Math.max(latest, row.endMs), first!.endMs),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Group stored lines into animation runs.
|
||||||
|
*
|
||||||
|
* Runs never cross a session, which is what keeps a rewatch intact: the same episode
|
||||||
|
* watched twice stores the same line twice, and those two belong to different sessions.
|
||||||
|
*/
|
||||||
|
export function findDuplicateSubtitleLineBursts(
|
||||||
|
rows: readonly StoredSubtitleLineRow[],
|
||||||
|
options: DuplicateSubtitleLineCleanupOptions = {},
|
||||||
|
): DuplicateSubtitleLineBurst[] {
|
||||||
|
const bounds = resolveBounds(options);
|
||||||
|
const bursts: DuplicateSubtitleLineBurst[] = [];
|
||||||
|
let run: StoredSubtitleLineRow[] = [];
|
||||||
|
let chainEndMs = 0;
|
||||||
|
|
||||||
|
const closeRun = (): void => {
|
||||||
|
if (run.length > 1 && isBurst(run, bounds)) {
|
||||||
|
bursts.push(toBurst(run));
|
||||||
|
}
|
||||||
|
run = [];
|
||||||
|
};
|
||||||
|
|
||||||
|
for (const row of rows) {
|
||||||
|
const previous = run[run.length - 1];
|
||||||
|
const continuesRun =
|
||||||
|
previous !== undefined &&
|
||||||
|
previous.sessionId === row.sessionId &&
|
||||||
|
previous.videoId === row.videoId &&
|
||||||
|
previous.text === row.text &&
|
||||||
|
row.startMs <= chainEndMs + bounds.gapToleranceMs;
|
||||||
|
|
||||||
|
if (continuesRun) {
|
||||||
|
run.push(row);
|
||||||
|
chainEndMs = Math.max(chainEndMs, row.endMs);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
closeRun();
|
||||||
|
run = [row];
|
||||||
|
chainEndMs = row.endMs;
|
||||||
|
}
|
||||||
|
closeRun();
|
||||||
|
|
||||||
|
return bursts;
|
||||||
|
}
|
||||||
|
|
||||||
|
function chunk<T>(values: T[], size: number): T[][] {
|
||||||
|
const chunks: T[][] = [];
|
||||||
|
for (let i = 0; i < values.length; i += size) {
|
||||||
|
chunks.push(values.slice(i, i + size));
|
||||||
|
}
|
||||||
|
return chunks;
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildSamples(
|
||||||
|
db: DatabaseSync,
|
||||||
|
bursts: DuplicateSubtitleLineBurst[],
|
||||||
|
sampleLimit: number,
|
||||||
|
): DuplicateSubtitleLineSample[] {
|
||||||
|
if (sampleLimit === 0 || bursts.length === 0) {
|
||||||
|
return [];
|
||||||
|
}
|
||||||
|
const largest = [...bursts]
|
||||||
|
.sort((a, b) => b.removedLineIds.length - a.removedLineIds.length)
|
||||||
|
.slice(0, sampleLimit);
|
||||||
|
const videoIds = [...new Set(largest.map((burst) => burst.videoId))];
|
||||||
|
const titles = new Map<number, string>();
|
||||||
|
for (const batch of chunk(videoIds, LINE_ID_BATCH_SIZE)) {
|
||||||
|
const rows = db
|
||||||
|
.prepare(
|
||||||
|
`SELECT video_id AS videoId, canonical_title AS title
|
||||||
|
FROM imm_videos
|
||||||
|
WHERE video_id IN (${makePlaceholders(batch)})`,
|
||||||
|
)
|
||||||
|
.all(...batch) as Array<{ videoId: number; title: string | null }>;
|
||||||
|
for (const row of rows) {
|
||||||
|
if (row.title) titles.set(row.videoId, row.title);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return largest.map((burst) => ({
|
||||||
|
videoId: burst.videoId,
|
||||||
|
videoTitle: titles.get(burst.videoId) ?? null,
|
||||||
|
text: burst.text,
|
||||||
|
frames: burst.removedLineIds.length + 1,
|
||||||
|
removedLines: burst.removedLineIds.length,
|
||||||
|
startMs: burst.startMs,
|
||||||
|
endMs: burst.endMs,
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
|
function sumRemovedOccurrences(
|
||||||
|
db: DatabaseSync,
|
||||||
|
table: 'imm_word_line_occurrences' | 'imm_kanji_line_occurrences',
|
||||||
|
lineIds: number[],
|
||||||
|
): number {
|
||||||
|
let total = 0;
|
||||||
|
for (const batch of chunk(lineIds, LINE_ID_BATCH_SIZE)) {
|
||||||
|
const row = db
|
||||||
|
.prepare(
|
||||||
|
`SELECT COALESCE(SUM(occurrence_count), 0) AS total
|
||||||
|
FROM ${table}
|
||||||
|
WHERE line_id IN (${makePlaceholders(batch)})`,
|
||||||
|
)
|
||||||
|
.get(...batch) as { total: number } | null;
|
||||||
|
total += row?.total ?? 0;
|
||||||
|
}
|
||||||
|
return total;
|
||||||
|
}
|
||||||
|
|
||||||
|
function applyBursts(db: DatabaseSync, bursts: DuplicateSubtitleLineBurst[]): void {
|
||||||
|
const removedLineIds = bursts.flatMap((burst) => burst.removedLineIds);
|
||||||
|
const currentMs = toDbTimestamp(nowMs());
|
||||||
|
|
||||||
|
db.exec('BEGIN IMMEDIATE');
|
||||||
|
try {
|
||||||
|
for (const batch of chunk(removedLineIds, LINE_ID_BATCH_SIZE)) {
|
||||||
|
const placeholders = makePlaceholders(batch);
|
||||||
|
// Measured before the delete, applied after it: `applyLexicalRemovals` checks the
|
||||||
|
// surviving occurrences to decide whether a zeroed count really means the word is
|
||||||
|
// gone, so the rows it inspects have to be the post-delete ones.
|
||||||
|
const plan = planLexicalRemovalsForLines(db, batch);
|
||||||
|
db.prepare(`DELETE FROM imm_word_line_occurrences WHERE line_id IN (${placeholders})`).run(
|
||||||
|
...batch,
|
||||||
|
);
|
||||||
|
db.prepare(`DELETE FROM imm_kanji_line_occurrences WHERE line_id IN (${placeholders})`).run(
|
||||||
|
...batch,
|
||||||
|
);
|
||||||
|
db.prepare(`DELETE FROM imm_subtitle_lines WHERE line_id IN (${placeholders})`).run(...batch);
|
||||||
|
applyLexicalRemovals(db, plan);
|
||||||
|
}
|
||||||
|
|
||||||
|
const extendStmt = db.prepare(
|
||||||
|
`UPDATE imm_subtitle_lines
|
||||||
|
SET segment_end_ms = ?, LAST_UPDATE_DATE = ?
|
||||||
|
WHERE line_id = ? AND (segment_end_ms IS NULL OR segment_end_ms < ?)`,
|
||||||
|
);
|
||||||
|
for (const burst of bursts) {
|
||||||
|
extendStmt.run(burst.endMs, currentMs, burst.keptLineId, burst.endMs);
|
||||||
|
}
|
||||||
|
db.exec('COMMIT');
|
||||||
|
} catch (error) {
|
||||||
|
db.exec('ROLLBACK');
|
||||||
|
throw error;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Collapse stored animation bursts down to one line each.
|
||||||
|
*
|
||||||
|
* A dry run measures exactly what an apply would remove, using the same scan, so the
|
||||||
|
* numbers shown in a confirmation prompt are the numbers that will happen.
|
||||||
|
*/
|
||||||
|
export function cleanupDuplicateSubtitleLines(
|
||||||
|
db: DatabaseSync,
|
||||||
|
options: DuplicateSubtitleLineCleanupOptions = {},
|
||||||
|
): DuplicateSubtitleLineCleanupSummary {
|
||||||
|
const bounds = resolveBounds(options);
|
||||||
|
const dryRun = options.dryRun === true;
|
||||||
|
const rows = readCandidateLines(db, bounds);
|
||||||
|
const bursts = findDuplicateSubtitleLineBursts(rows, options);
|
||||||
|
const removedLineIds = bursts.flatMap((burst) => burst.removedLineIds);
|
||||||
|
|
||||||
|
const summary: DuplicateSubtitleLineCleanupSummary = {
|
||||||
|
dryRun,
|
||||||
|
lookbackDays: bounds.lookbackDays,
|
||||||
|
scannedLines: rows.length,
|
||||||
|
burstGroups: bursts.length,
|
||||||
|
removedLines: removedLineIds.length,
|
||||||
|
removedWordOccurrences: sumRemovedOccurrences(db, 'imm_word_line_occurrences', removedLineIds),
|
||||||
|
removedKanjiOccurrences: sumRemovedOccurrences(
|
||||||
|
db,
|
||||||
|
'imm_kanji_line_occurrences',
|
||||||
|
removedLineIds,
|
||||||
|
),
|
||||||
|
samples: buildSamples(db, bursts, bounds.sampleLimit),
|
||||||
|
};
|
||||||
|
|
||||||
|
if (dryRun || removedLineIds.length === 0) {
|
||||||
|
return summary;
|
||||||
|
}
|
||||||
|
|
||||||
|
applyBursts(db, bursts);
|
||||||
|
return summary;
|
||||||
|
}
|
||||||
@@ -268,6 +268,19 @@ export function planLexicalRemovalsForSessions(
|
|||||||
return planLexicalRemovals(db, `sl.session_id IN (${makePlaceholders(sessionIds)})`, sessionIds);
|
return planLexicalRemovals(db, `sl.session_id IN (${makePlaceholders(sessionIds)})`, sessionIds);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Measure what deleting these individual subtitle lines removes from the vocabulary
|
||||||
|
* tables. Used by the duplicate-line cleanup, which drops animation frames out of the
|
||||||
|
* middle of sessions that otherwise stay intact.
|
||||||
|
*/
|
||||||
|
export function planLexicalRemovalsForLines(
|
||||||
|
db: DatabaseSync,
|
||||||
|
lineIds: number[],
|
||||||
|
): LexicalRemovalPlan {
|
||||||
|
if (lineIds.length === 0) return EMPTY_LEXICAL_REMOVAL_PLAN;
|
||||||
|
return planLexicalRemovals(db, `sl.line_id IN (${makePlaceholders(lineIds)})`, lineIds);
|
||||||
|
}
|
||||||
|
|
||||||
/** Measure what deleting these videos removes from the vocabulary tables. */
|
/** Measure what deleting these videos removes from the vocabulary tables. */
|
||||||
export function planLexicalRemovalsForVideos(
|
export function planLexicalRemovalsForVideos(
|
||||||
db: DatabaseSync,
|
db: DatabaseSync,
|
||||||
|
|||||||
@@ -5,6 +5,7 @@ import {
|
|||||||
buildSentenceSearchOptions,
|
buildSentenceSearchOptions,
|
||||||
enrichSessionsWithKnownWordMetrics,
|
enrichSessionsWithKnownWordMetrics,
|
||||||
parseBooleanQuery,
|
parseBooleanQuery,
|
||||||
|
parseDuplicateLineCleanupBody,
|
||||||
parseExcludedWordsBody,
|
parseExcludedWordsBody,
|
||||||
parseIntQuery,
|
parseIntQuery,
|
||||||
} from './route-support.js';
|
} from './route-support.js';
|
||||||
@@ -40,6 +41,15 @@ export function registerStatsLibraryRoutes(
|
|||||||
return c.json(statsJson('setExcludedWords', { ok: true }));
|
return c.json(statsJson('setExcludedWords', { ok: true }));
|
||||||
});
|
});
|
||||||
|
|
||||||
|
// Collapse animation bursts older versions recorded frame by frame. `dryRun` measures
|
||||||
|
// the same scan without writing, so the confirmation the user sees is the real cost.
|
||||||
|
app.post('/api/stats/maintenance/duplicate-lines', async (c) => {
|
||||||
|
const body = await c.req.json().catch(() => null);
|
||||||
|
const { dryRun, lookbackDays } = parseDuplicateLineCleanupBody(body);
|
||||||
|
const result = await tracker.cleanupDuplicateSubtitleLines({ dryRun, lookbackDays });
|
||||||
|
return c.json(statsJson('duplicateLineCleanup', result));
|
||||||
|
});
|
||||||
|
|
||||||
app.get('/api/stats/vocabulary/occurrences', async (c) => {
|
app.get('/api/stats/vocabulary/occurrences', async (c) => {
|
||||||
const headword = (c.req.query('headword') ?? '').trim();
|
const headword = (c.req.query('headword') ?? '').trim();
|
||||||
const word = (c.req.query('word') ?? '').trim();
|
const word = (c.req.query('word') ?? '').trim();
|
||||||
|
|||||||
@@ -88,6 +88,23 @@ export function parseExcludedWordsBody(body: unknown): StatsExcludedWord[] | nul
|
|||||||
return words;
|
return words;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Read a duplicate-line cleanup request. An absent or unusable `lookbackDays` scans all
|
||||||
|
* history, which is what the CLI does; only a positive number narrows the window.
|
||||||
|
*/
|
||||||
|
export function parseDuplicateLineCleanupBody(body: unknown): {
|
||||||
|
dryRun: boolean;
|
||||||
|
lookbackDays: number | null;
|
||||||
|
} {
|
||||||
|
const source = body && typeof body === 'object' ? (body as Record<string, unknown>) : {};
|
||||||
|
const rawLookback = source.lookbackDays;
|
||||||
|
const lookbackDays =
|
||||||
|
typeof rawLookback === 'number' && Number.isFinite(rawLookback) && rawLookback > 0
|
||||||
|
? Math.floor(rawLookback)
|
||||||
|
: null;
|
||||||
|
return { dryRun: source.dryRun === true, lookbackDays };
|
||||||
|
}
|
||||||
|
|
||||||
export function loadKnownWordsSet(cachePath: string | undefined): Set<string> | null {
|
export function loadKnownWordsSet(cachePath: string | undefined): Set<string> | null {
|
||||||
if (!cachePath || !existsSync(cachePath)) return null;
|
if (!cachePath || !existsSync(cachePath)) return null;
|
||||||
try {
|
try {
|
||||||
|
|||||||
@@ -0,0 +1,40 @@
|
|||||||
|
/*
|
||||||
|
* Thresholds that decide when a run of repeated subtitle events is one animation.
|
||||||
|
*
|
||||||
|
* Three consumers have to agree on these numbers or the same karaoke line is one cue in
|
||||||
|
* the sidebar and two hundred in the stats: the file-level cue dedup
|
||||||
|
* (`subtitle-cue-dedup`), the live gate that decides what immersion stats record
|
||||||
|
* (`subtitle-line-dedup-gate`), and the retroactive database cleanup
|
||||||
|
* (`immersion-tracker/duplicate-line-cleanup`).
|
||||||
|
*/
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Back-to-back frames of the same animation are authored flush against each other; a
|
||||||
|
* tiny tolerance absorbs the centisecond rounding of the ASS timestamp format.
|
||||||
|
*/
|
||||||
|
export const DUPLICATE_CUE_GAP_TOLERANCE_SECONDS = 0.05;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A burst is a *sequence*. Two adjacent events are two events, not an animation --
|
||||||
|
* characters do repeat each other, and a repeated line can legitimately be short.
|
||||||
|
*/
|
||||||
|
export const MIN_BURST_EVENTS = 3;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Real dialogue holds on screen for about a second, so a run with a couple of much
|
||||||
|
* shorter events among them looks like frames. Used only alongside authoring evidence.
|
||||||
|
*/
|
||||||
|
export const ANIMATION_FRAME_MAX_SECONDS = 0.3;
|
||||||
|
|
||||||
|
/** A karaoke run usually ends on a long "hold" frame, so not every event is short. */
|
||||||
|
export const MIN_TAGGED_BURST_FRAMES = 2;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* SRT and VTT carry no authoring metadata at all, so timing is the only signal available
|
||||||
|
* -- which makes it the easiest one to get wrong. ASS->SRT conversion leaves frames at
|
||||||
|
* ~0.04s, well under any real utterance, and a burst leaves many of them behind. Both
|
||||||
|
* bounds are deliberately far stricter than the ASS path: a run of ordinary short lines
|
||||||
|
* (`えっ` traded between characters) must not clear them.
|
||||||
|
*/
|
||||||
|
export const TIMING_ONLY_FRAME_MAX_SECONDS = 0.1;
|
||||||
|
export const MIN_TIMING_ONLY_FRAMES = 5;
|
||||||
@@ -7,31 +7,20 @@
|
|||||||
*/
|
*/
|
||||||
|
|
||||||
import { hasAssTemporalOverride, isAnimatedAssEffectKind } from './ass-text';
|
import { hasAssTemporalOverride, isAnimatedAssEffectKind } from './ass-text';
|
||||||
|
import {
|
||||||
|
ANIMATION_FRAME_MAX_SECONDS,
|
||||||
|
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
|
||||||
|
MIN_BURST_EVENTS,
|
||||||
|
MIN_TAGGED_BURST_FRAMES,
|
||||||
|
MIN_TIMING_ONLY_FRAMES,
|
||||||
|
TIMING_ONLY_FRAME_MAX_SECONDS,
|
||||||
|
} from './subtitle-burst-constants';
|
||||||
import type {
|
import type {
|
||||||
AnnotatedSubtitleCue,
|
AnnotatedSubtitleCue,
|
||||||
SubtitleCue,
|
SubtitleCue,
|
||||||
SubtitleSourceFormat,
|
SubtitleSourceFormat,
|
||||||
} from './subtitle-cue-parser';
|
} from './subtitle-cue-parser';
|
||||||
|
|
||||||
// Back-to-back frames of the same animation are authored flush against each other; a
|
|
||||||
// tiny tolerance absorbs the centisecond rounding of the ASS timestamp format.
|
|
||||||
const DUPLICATE_CUE_GAP_TOLERANCE_SECONDS = 0.05;
|
|
||||||
// A burst is a *sequence*. Two adjacent events are two events, not an animation --
|
|
||||||
// characters do repeat each other, and a repeated line can legitimately be short.
|
|
||||||
const MIN_BURST_EVENTS = 3;
|
|
||||||
// Real dialogue holds on screen for about a second, so a run with a couple of much
|
|
||||||
// shorter events among them looks like frames. Used only alongside authoring evidence.
|
|
||||||
const ANIMATION_FRAME_MAX_SECONDS = 0.3;
|
|
||||||
// A karaoke run usually ends on a long "hold" frame, so not every event is short.
|
|
||||||
const MIN_TAGGED_BURST_FRAMES = 2;
|
|
||||||
// SRT and VTT carry no authoring metadata at all, so timing is the only signal available
|
|
||||||
// -- which makes it the easiest one to get wrong. ASS->SRT conversion leaves frames at
|
|
||||||
// ~0.04s, well under any real utterance, and a burst leaves many of them behind. Both
|
|
||||||
// bounds are deliberately far stricter than the ASS path: a run of ordinary short lines
|
|
||||||
// (`えっ` traded between characters) must not clear them.
|
|
||||||
const TIMING_ONLY_FRAME_MAX_SECONDS = 0.1;
|
|
||||||
const MIN_TIMING_ONLY_FRAMES = 5;
|
|
||||||
|
|
||||||
function cueKey(cue: SubtitleCue): string {
|
function cueKey(cue: SubtitleCue): string {
|
||||||
return `${cue.startTime}|${cue.endTime}|${cue.text}`;
|
return `${cue.startTime}|${cue.endTime}|${cue.text}`;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,99 @@
|
|||||||
|
import assert from 'node:assert/strict';
|
||||||
|
import test from 'node:test';
|
||||||
|
import { createSubtitleLineDedupGate } from './subtitle-line-dedup-gate';
|
||||||
|
import type { SubtitleCue } from '../../types';
|
||||||
|
|
||||||
|
function karaokeFrames(text: string, start: number, frames: number, frameSeconds: number) {
|
||||||
|
return Array.from({ length: frames }, (_, index) => ({
|
||||||
|
text,
|
||||||
|
startSec: start + index * frameSeconds,
|
||||||
|
endSec: start + (index + 1) * frameSeconds,
|
||||||
|
}));
|
||||||
|
}
|
||||||
|
|
||||||
|
test('parsed cues drop the frames the sidebar already collapsed', () => {
|
||||||
|
// What `mergeDuplicateCues` leaves behind for a karaoke run: one cue over the run.
|
||||||
|
const cues: SubtitleCue[] = [
|
||||||
|
{ startTime: 10, endTime: 14, text: '飛び上がる' },
|
||||||
|
{ startTime: 14, endTime: 16, text: 'もしも' },
|
||||||
|
];
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
|
||||||
|
|
||||||
|
const recorded = karaokeFrames('飛び上がる', 10, 40, 0.1).filter((sample) =>
|
||||||
|
gate.shouldRecord(sample),
|
||||||
|
);
|
||||||
|
|
||||||
|
assert.equal(recorded.length, 1);
|
||||||
|
assert.equal(recorded[0]!.startSec, 10);
|
||||||
|
assert.equal(gate.shouldRecord({ text: 'もしも', startSec: 14, endSec: 16 }), true);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('parsed cues keep separate lines that merely repeat', () => {
|
||||||
|
const cues: SubtitleCue[] = [
|
||||||
|
{ startTime: 3, endTime: 3.4, text: 'えっ' },
|
||||||
|
{ startTime: 3.4, endTime: 3.9, text: 'えっ' },
|
||||||
|
{ startTime: 3.9, endTime: 4.5, text: 'えっ' },
|
||||||
|
];
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
|
||||||
|
|
||||||
|
const recorded = cues.filter((cue) =>
|
||||||
|
gate.shouldRecord({ text: cue.text, startSec: cue.startTime, endSec: cue.endTime }),
|
||||||
|
);
|
||||||
|
|
||||||
|
assert.equal(recorded.length, 3);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('a line whose timing does not match any cue still records', () => {
|
||||||
|
// A shifted track, an embedded sub nobody parsed: no match, no drop.
|
||||||
|
const cues: SubtitleCue[] = [{ startTime: 10, endTime: 14, text: '飛び上がる' }];
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => cues });
|
||||||
|
|
||||||
|
assert.equal(gate.shouldRecord({ text: '飛び上がる', startSec: 42, endSec: 44 }), true);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('without parsed cues a long run of identical short frames stops recording', () => {
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
|
||||||
|
|
||||||
|
const recorded = karaokeFrames('ひとしずく', 0, 200, 0.04).filter((sample) =>
|
||||||
|
gate.shouldRecord(sample),
|
||||||
|
);
|
||||||
|
|
||||||
|
assert.equal(recorded.length, 4);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('without parsed cues ordinary repeated dialogue keeps recording', () => {
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
|
||||||
|
|
||||||
|
// Six contiguous `えっ`, each held for a normal beat rather than an animation frame.
|
||||||
|
const recorded = karaokeFrames('えっ', 0, 6, 0.6).filter((sample) => gate.shouldRecord(sample));
|
||||||
|
|
||||||
|
assert.equal(recorded.length, 6);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('the same event offered twice does not advance the run', () => {
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
|
||||||
|
|
||||||
|
// mpv fires the timing handler once for `sub-start` and once for `sub-end`.
|
||||||
|
for (let i = 0; i < 8; i += 1) {
|
||||||
|
assert.equal(gate.shouldRecord({ text: '待って', startSec: 5, endSec: 5.05 }), true);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
test('a gap between frames starts a new run', () => {
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
|
||||||
|
|
||||||
|
const first = karaokeFrames('もし', 0, 6, 0.04).filter((sample) => gate.shouldRecord(sample));
|
||||||
|
const second = karaokeFrames('もし', 30, 6, 0.04).filter((sample) => gate.shouldRecord(sample));
|
||||||
|
|
||||||
|
assert.equal(first.length, 4);
|
||||||
|
assert.equal(second.length, 4);
|
||||||
|
});
|
||||||
|
|
||||||
|
test('reset forgets the streaming run', () => {
|
||||||
|
const gate = createSubtitleLineDedupGate({ getParsedCues: () => null });
|
||||||
|
|
||||||
|
karaokeFrames('もし', 0, 20, 0.04).forEach((sample) => gate.shouldRecord(sample));
|
||||||
|
gate.reset();
|
||||||
|
|
||||||
|
assert.equal(gate.shouldRecord({ text: 'もし', startSec: 0.8, endSec: 0.84 }), true);
|
||||||
|
});
|
||||||
@@ -0,0 +1,184 @@
|
|||||||
|
/*
|
||||||
|
* Decides which live mpv subtitle lines reach the immersion stats.
|
||||||
|
*
|
||||||
|
* The sidebar reads a parsed subtitle file, so it can collapse an animation burst with
|
||||||
|
* full lookahead (`subtitle-cue-dedup`). Stats are fed from mpv's `sub-start`/`sub-end`
|
||||||
|
* properties instead -- one event per animation frame, each with its own start time --
|
||||||
|
* so without a gate a karaoke OP counts its lyrics once per frame and buries every real
|
||||||
|
* word in the vocabulary charts.
|
||||||
|
*
|
||||||
|
* Two layers, in order:
|
||||||
|
*
|
||||||
|
* 1. When the active source has been parsed, its cue list has *already* been collapsed.
|
||||||
|
* A live line that lands inside a surviving cue of the same text, but after that
|
||||||
|
* cue's start, is a frame the sidebar merged away, so stats drop it too. This is the
|
||||||
|
* layer that keeps the two views consistent by construction.
|
||||||
|
* 2. Otherwise (embedded track nobody parsed, a source whose timings mpv has shifted)
|
||||||
|
* fall back to timing alone. No authoring metadata is available live -- mpv delivers
|
||||||
|
* `sub-text-ass` after `sub-start`/`sub-end`, so any ASS text read here belongs to the
|
||||||
|
* previous event -- which puts this layer in the same position as the SRT path in
|
||||||
|
* `subtitle-cue-dedup`, and it uses that path's deliberately strict bounds.
|
||||||
|
*/
|
||||||
|
|
||||||
|
import { normalizePlainSubtitleText } from './ass-text';
|
||||||
|
import {
|
||||||
|
DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
|
||||||
|
MIN_TIMING_ONLY_FRAMES,
|
||||||
|
TIMING_ONLY_FRAME_MAX_SECONDS,
|
||||||
|
} from './subtitle-burst-constants';
|
||||||
|
import type { SubtitleCue } from './subtitle-cue-parser';
|
||||||
|
|
||||||
|
export interface SubtitleLineSample {
|
||||||
|
text: string;
|
||||||
|
startSec: number;
|
||||||
|
endSec: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface SubtitleLineDedupGateDeps {
|
||||||
|
/** Cues for the active source, already collapsed by the parser. */
|
||||||
|
getParsedCues: () => readonly SubtitleCue[] | null | undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface SubtitleLineDedupGate {
|
||||||
|
/** False when this line is an animation frame of a line already recorded. */
|
||||||
|
shouldRecord: (sample: SubtitleLineSample) => boolean;
|
||||||
|
/** Forget the streaming run state, e.g. when playback moves to another file. */
|
||||||
|
reset: () => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface CueSpan {
|
||||||
|
startTime: number;
|
||||||
|
endTime: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface StreamingRunState {
|
||||||
|
text: string;
|
||||||
|
startMs: number;
|
||||||
|
chainEndSec: number;
|
||||||
|
/** Contiguous identical short frames seen so far, including the recorded first one. */
|
||||||
|
frames: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
function normalizeLineText(text: string): string {
|
||||||
|
return normalizePlainSubtitleText(text, { collapseLineBreaks: true });
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildSpansByText(cues: readonly SubtitleCue[]): Map<string, CueSpan[]> {
|
||||||
|
const spansByText = new Map<string, CueSpan[]>();
|
||||||
|
for (const cue of cues) {
|
||||||
|
const key = normalizeLineText(cue.text);
|
||||||
|
if (!key) continue;
|
||||||
|
const span = { startTime: cue.startTime, endTime: cue.endTime };
|
||||||
|
const existing = spansByText.get(key);
|
||||||
|
if (existing) {
|
||||||
|
existing.push(span);
|
||||||
|
} else {
|
||||||
|
spansByText.set(key, [span]);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return spansByText;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* A frame the parser merged away: the same text, starting inside a surviving cue but
|
||||||
|
* after it began.
|
||||||
|
*
|
||||||
|
* Starting a cue always wins over falling inside one. The first frame of a collapsed run
|
||||||
|
* starts *at* the merged cue, and a line the parser deliberately kept separate -- three
|
||||||
|
* characters trading `えっ` back to back -- begins exactly where the one before it ends.
|
||||||
|
*/
|
||||||
|
function isMergedAwayFrame(spans: readonly CueSpan[], startSec: number): boolean {
|
||||||
|
const startsOwnCue = spans.some(
|
||||||
|
(span) => Math.abs(startSec - span.startTime) <= DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
|
||||||
|
);
|
||||||
|
if (startsOwnCue) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return spans.some(
|
||||||
|
(span) =>
|
||||||
|
startSec > span.startTime + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS &&
|
||||||
|
startSec <= span.endTime + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function createSubtitleLineDedupGate(
|
||||||
|
deps: SubtitleLineDedupGateDeps,
|
||||||
|
): SubtitleLineDedupGate {
|
||||||
|
let indexedCues: readonly SubtitleCue[] | null = null;
|
||||||
|
let spansByText: Map<string, CueSpan[]> = new Map();
|
||||||
|
let run: StreamingRunState | null = null;
|
||||||
|
|
||||||
|
const lookupSpans = (text: string): CueSpan[] | null => {
|
||||||
|
const cues = deps.getParsedCues();
|
||||||
|
if (!cues?.length) {
|
||||||
|
indexedCues = null;
|
||||||
|
spansByText = new Map();
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
if (cues !== indexedCues) {
|
||||||
|
indexedCues = cues;
|
||||||
|
spansByText = buildSpansByText(cues);
|
||||||
|
}
|
||||||
|
return spansByText.get(text) ?? null;
|
||||||
|
};
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Timing-only burst detection over a stream. Without lookahead the run can only be
|
||||||
|
* recognised from the inside, so the first frames of a burst are recorded and the rest
|
||||||
|
* dropped -- an OP costs a handful of counted lines instead of several hundred.
|
||||||
|
*/
|
||||||
|
const advanceStreamingRun = (text: string, sample: SubtitleLineSample): boolean => {
|
||||||
|
const startMs = Math.round(sample.startSec * 1000);
|
||||||
|
// mpv reports `sub-start` and `sub-end` separately, so one event can be offered
|
||||||
|
// twice. The same start is the same frame, never the next one in a run.
|
||||||
|
if (run && run.text === text && run.startMs === startMs) {
|
||||||
|
run.chainEndSec = Math.max(run.chainEndSec, sample.endSec);
|
||||||
|
return run.frames < MIN_TIMING_ONLY_FRAMES;
|
||||||
|
}
|
||||||
|
|
||||||
|
const isShortFrame = sample.endSec - sample.startSec < TIMING_ONLY_FRAME_MAX_SECONDS;
|
||||||
|
// Frames are authored flush against each other, but typesetters do overlap them, so
|
||||||
|
// the chain only requires forward progress that stays inside the running end.
|
||||||
|
const continuesRun =
|
||||||
|
run !== null &&
|
||||||
|
run.text === text &&
|
||||||
|
isShortFrame &&
|
||||||
|
startMs > run.startMs &&
|
||||||
|
sample.startSec <= run.chainEndSec + DUPLICATE_CUE_GAP_TOLERANCE_SECONDS;
|
||||||
|
|
||||||
|
if (continuesRun && run) {
|
||||||
|
run.startMs = startMs;
|
||||||
|
run.chainEndSec = Math.max(run.chainEndSec, sample.endSec);
|
||||||
|
run.frames += 1;
|
||||||
|
} else {
|
||||||
|
run = {
|
||||||
|
text,
|
||||||
|
startMs,
|
||||||
|
chainEndSec: sample.endSec,
|
||||||
|
frames: isShortFrame ? 1 : 0,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
return run.frames < MIN_TIMING_ONLY_FRAMES;
|
||||||
|
};
|
||||||
|
|
||||||
|
return {
|
||||||
|
shouldRecord: (sample) => {
|
||||||
|
const text = normalizeLineText(sample.text);
|
||||||
|
if (!text) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const spans = lookupSpans(text);
|
||||||
|
if (spans && isMergedAwayFrame(spans, sample.startSec)) {
|
||||||
|
run = null;
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
return advanceStreamingRun(text, sample);
|
||||||
|
},
|
||||||
|
reset: () => {
|
||||||
|
run = null;
|
||||||
|
},
|
||||||
|
};
|
||||||
|
}
|
||||||
@@ -1,4 +1,5 @@
|
|||||||
import type { MergedToken, SubtitleData } from '../../types';
|
import { createSubtitleLineDedupGate } from '../../core/services/subtitle-line-dedup-gate';
|
||||||
|
import type { MergedToken, SubtitleCue, SubtitleData } from '../../types';
|
||||||
|
|
||||||
type AnilistPostWatchRunOptions = {
|
type AnilistPostWatchRunOptions = {
|
||||||
watchedSeconds?: number;
|
watchedSeconds?: number;
|
||||||
@@ -34,6 +35,7 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
|
|||||||
subtitleTimingTracker: {
|
subtitleTimingTracker: {
|
||||||
recordSubtitle?: (text: string, start: number, end: number, secondaryText?: string) => void;
|
recordSubtitle?: (text: string, start: number, end: number, secondaryText?: string) => void;
|
||||||
} | null;
|
} | null;
|
||||||
|
activeParsedSubtitleCues?: SubtitleCue[] | null;
|
||||||
currentMediaPath?: string | null;
|
currentMediaPath?: string | null;
|
||||||
currentSubText: string;
|
currentSubText: string;
|
||||||
currentSubAssText: string;
|
currentSubAssText: string;
|
||||||
@@ -86,6 +88,11 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
|
|||||||
deps.ensureImmersionTrackerInitialized();
|
deps.ensureImmersionTrackerInitialized();
|
||||||
deps.appState.immersionTracker?.recordPlaybackPosition?.(normalizedTimeSec);
|
deps.appState.immersionTracker?.recordPlaybackPosition?.(normalizedTimeSec);
|
||||||
};
|
};
|
||||||
|
// mpv reports every animation frame of a typeset line as its own subtitle event, so
|
||||||
|
// stats have to collapse bursts the same way the parsed cue list already does.
|
||||||
|
const immersionLineDedupGate = createSubtitleLineDedupGate({
|
||||||
|
getParsedCues: () => deps.appState.activeParsedSubtitleCues,
|
||||||
|
});
|
||||||
const hasInitialPlaybackQuitOnDisconnectArg = (): boolean =>
|
const hasInitialPlaybackQuitOnDisconnectArg = (): boolean =>
|
||||||
Boolean(
|
Boolean(
|
||||||
deps.appState.initialArgs?.managedPlayback ||
|
deps.appState.initialArgs?.managedPlayback ||
|
||||||
@@ -110,6 +117,9 @@ export function createBuildBindMpvMainEventHandlersMainDepsHandler(deps: {
|
|||||||
if (!tracker?.recordSubtitleLine) {
|
if (!tracker?.recordSubtitleLine) {
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
if (!immersionLineDedupGate.shouldRecord({ text, startSec: start, endSec: end })) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
const secondaryText = deps.appState.mpvClient?.currentSecondarySubText || null;
|
const secondaryText = deps.appState.mpvClient?.currentSecondarySubText || null;
|
||||||
const cachedTokens =
|
const cachedTokens =
|
||||||
deps.appState.currentSubtitleData?.text === text
|
deps.appState.currentSubtitleData?.text === text
|
||||||
|
|||||||
@@ -200,6 +200,59 @@ test('stats cli command fails when immersion tracking is disabled', async () =>
|
|||||||
]);
|
]);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test('stats cli command runs a duplicate-line cleanup preview without touching the dashboard', async () => {
|
||||||
|
const { handler, calls, responses } = makeHandler({
|
||||||
|
getImmersionTracker: () => ({
|
||||||
|
cleanupDuplicateSubtitleLines: async (options: {
|
||||||
|
dryRun?: boolean;
|
||||||
|
lookbackDays?: number | null;
|
||||||
|
}) => ({
|
||||||
|
dryRun: options.dryRun === true,
|
||||||
|
lookbackDays: options.lookbackDays ?? null,
|
||||||
|
scannedLines: 900,
|
||||||
|
burstGroups: 2,
|
||||||
|
removedLines: 180,
|
||||||
|
removedWordOccurrences: 540,
|
||||||
|
removedKanjiOccurrences: 120,
|
||||||
|
samples: [
|
||||||
|
{
|
||||||
|
videoId: 7,
|
||||||
|
videoTitle: 'Ep 1',
|
||||||
|
text: '飛び上がる',
|
||||||
|
frames: 90,
|
||||||
|
removedLines: 89,
|
||||||
|
startMs: 1000,
|
||||||
|
endMs: 5000,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
}),
|
||||||
|
}),
|
||||||
|
});
|
||||||
|
|
||||||
|
await handler(
|
||||||
|
{
|
||||||
|
statsResponsePath: '/tmp/subminer-stats-response.json',
|
||||||
|
statsCleanup: true,
|
||||||
|
statsCleanupDuplicateLines: true,
|
||||||
|
statsCleanupDryRun: true,
|
||||||
|
statsCleanupLookbackDays: 30,
|
||||||
|
},
|
||||||
|
'initial',
|
||||||
|
);
|
||||||
|
|
||||||
|
assert.deepEqual(calls, [
|
||||||
|
'ensureImmersionTrackerStarted',
|
||||||
|
'info:Stats duplicate-line cleanup preview (last 30d): scanned=900 bursts=2 removedLines=180 removedWordCounts=540 removedKanjiCounts=120',
|
||||||
|
'info: Ep 1: "飛び上がる" x90',
|
||||||
|
]);
|
||||||
|
assert.deepEqual(responses, [
|
||||||
|
{
|
||||||
|
responsePath: '/tmp/subminer-stats-response.json',
|
||||||
|
payload: { ok: true },
|
||||||
|
},
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
test('stats cli command runs vocab cleanup instead of opening dashboard when cleanup mode is requested', async () => {
|
test('stats cli command runs vocab cleanup instead of opening dashboard when cleanup mode is requested', async () => {
|
||||||
const { handler, calls, responses } = makeHandler({
|
const { handler, calls, responses } = makeHandler({
|
||||||
getImmersionTracker: () => ({
|
getImmersionTracker: () => ({
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
import fs from 'node:fs';
|
import fs from 'node:fs';
|
||||||
import path from 'node:path';
|
import path from 'node:path';
|
||||||
import type { CliArgs, CliCommandSource } from '../../cli/args';
|
import type { CliArgs, CliCommandSource } from '../../cli/args';
|
||||||
|
import type { DuplicateSubtitleLineCleanupSummary } from '../../core/services/immersion-tracker/duplicate-line-cleanup';
|
||||||
import type {
|
import type {
|
||||||
LifetimeRebuildSummary,
|
LifetimeRebuildSummary,
|
||||||
VocabularyCleanupSummary,
|
VocabularyCleanupSummary,
|
||||||
@@ -50,6 +51,10 @@ export function createRunStatsCliCommandHandler(deps: {
|
|||||||
ensureVocabularyCleanupTokenizerReady?: () => Promise<void> | void;
|
ensureVocabularyCleanupTokenizerReady?: () => Promise<void> | void;
|
||||||
getImmersionTracker: () => {
|
getImmersionTracker: () => {
|
||||||
cleanupVocabularyStats?: () => Promise<VocabularyCleanupSummary>;
|
cleanupVocabularyStats?: () => Promise<VocabularyCleanupSummary>;
|
||||||
|
cleanupDuplicateSubtitleLines?: (options: {
|
||||||
|
dryRun?: boolean;
|
||||||
|
lookbackDays?: number | null;
|
||||||
|
}) => Promise<DuplicateSubtitleLineCleanupSummary>;
|
||||||
rebuildLifetimeSummaries?: () => Promise<LifetimeRebuildSummary>;
|
rebuildLifetimeSummaries?: () => Promise<LifetimeRebuildSummary>;
|
||||||
} | null;
|
} | null;
|
||||||
ensureStatsServerStarted: () => string;
|
ensureStatsServerStarted: () => string;
|
||||||
@@ -83,6 +88,9 @@ export function createRunStatsCliCommandHandler(deps: {
|
|||||||
| 'statsCleanup'
|
| 'statsCleanup'
|
||||||
| 'statsCleanupVocab'
|
| 'statsCleanupVocab'
|
||||||
| 'statsCleanupLifetime'
|
| 'statsCleanupLifetime'
|
||||||
|
| 'statsCleanupDuplicateLines'
|
||||||
|
| 'statsCleanupDryRun'
|
||||||
|
| 'statsCleanupLookbackDays'
|
||||||
>,
|
>,
|
||||||
source: CliCommandSource,
|
source: CliCommandSource,
|
||||||
): Promise<void> => {
|
): Promise<void> => {
|
||||||
@@ -126,6 +134,7 @@ export function createRunStatsCliCommandHandler(deps: {
|
|||||||
const cleanupModes = [
|
const cleanupModes = [
|
||||||
args.statsCleanupVocab ? 'vocab' : null,
|
args.statsCleanupVocab ? 'vocab' : null,
|
||||||
args.statsCleanupLifetime ? 'lifetime' : null,
|
args.statsCleanupLifetime ? 'lifetime' : null,
|
||||||
|
args.statsCleanupDuplicateLines ? 'duplicate-lines' : null,
|
||||||
].filter(Boolean);
|
].filter(Boolean);
|
||||||
if (cleanupModes.length !== 1) {
|
if (cleanupModes.length !== 1) {
|
||||||
throw new Error('Choose exactly one stats cleanup mode.');
|
throw new Error('Choose exactly one stats cleanup mode.');
|
||||||
@@ -142,6 +151,27 @@ export function createRunStatsCliCommandHandler(deps: {
|
|||||||
writeResponseSafe(args.statsResponsePath, { ok: true });
|
writeResponseSafe(args.statsResponsePath, { ok: true });
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
if (args.statsCleanupDuplicateLines && tracker.cleanupDuplicateSubtitleLines) {
|
||||||
|
const result = await tracker.cleanupDuplicateSubtitleLines({
|
||||||
|
dryRun: args.statsCleanupDryRun === true,
|
||||||
|
lookbackDays: args.statsCleanupLookbackDays ?? null,
|
||||||
|
});
|
||||||
|
const window =
|
||||||
|
result.lookbackDays === null ? 'all history' : `last ${result.lookbackDays}d`;
|
||||||
|
deps.logInfo(
|
||||||
|
`Stats duplicate-line cleanup ${result.dryRun ? 'preview' : 'complete'} (${window}): ` +
|
||||||
|
`scanned=${result.scannedLines} bursts=${result.burstGroups} ` +
|
||||||
|
`removedLines=${result.removedLines} removedWordCounts=${result.removedWordOccurrences} ` +
|
||||||
|
`removedKanjiCounts=${result.removedKanjiOccurrences}`,
|
||||||
|
);
|
||||||
|
for (const sample of result.samples.slice(0, 5)) {
|
||||||
|
deps.logInfo(
|
||||||
|
` ${sample.videoTitle ?? `video ${sample.videoId}`}: "${sample.text}" x${sample.frames}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
writeResponseSafe(args.statsResponsePath, { ok: true });
|
||||||
|
return;
|
||||||
|
}
|
||||||
if (!args.statsCleanupLifetime || !tracker.rebuildLifetimeSummaries) {
|
if (!args.statsCleanupLifetime || !tracker.rebuildLifetimeSummaries) {
|
||||||
throw new Error('Stats cleanup mode is not available.');
|
throw new Error('Stats cleanup mode is not available.');
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -18,6 +18,7 @@ import type {
|
|||||||
SessionTimelinePoint,
|
SessionTimelinePoint,
|
||||||
StatsAnkiNoteInfo,
|
StatsAnkiNoteInfo,
|
||||||
StatsCoverImagesData,
|
StatsCoverImagesData,
|
||||||
|
StatsDuplicateLineCleanupResult,
|
||||||
StatsExcludedWord,
|
StatsExcludedWord,
|
||||||
StreakCalendarDay,
|
StreakCalendarDay,
|
||||||
TrendsDashboardData,
|
TrendsDashboardData,
|
||||||
@@ -31,6 +32,13 @@ export type StatsTrendRange = '7d' | '30d' | '90d' | '365d' | 'all';
|
|||||||
export type StatsTrendGroupBy = 'day' | 'month';
|
export type StatsTrendGroupBy = 'day' | 'month';
|
||||||
export type StatsMineMode = 'word' | 'sentence' | 'audio';
|
export type StatsMineMode = 'word' | 'sentence' | 'audio';
|
||||||
|
|
||||||
|
/** Body of `POST /api/stats/maintenance/duplicate-lines`. */
|
||||||
|
export interface StatsDuplicateLineCleanupRequest {
|
||||||
|
dryRun?: boolean;
|
||||||
|
/** Null means every recorded line, whatever its age. */
|
||||||
|
lookbackDays?: number | null;
|
||||||
|
}
|
||||||
|
|
||||||
export interface StatsSessionKnownWordsTimelinePoint {
|
export interface StatsSessionKnownWordsTimelinePoint {
|
||||||
linesSeen: number;
|
linesSeen: number;
|
||||||
knownWordsSeen: number;
|
knownWordsSeen: number;
|
||||||
@@ -125,6 +133,7 @@ export interface StatsJsonResponseMap {
|
|||||||
vocabulary: VocabularyEntry[];
|
vocabulary: VocabularyEntry[];
|
||||||
excludedWords: StatsExcludedWord[];
|
excludedWords: StatsExcludedWord[];
|
||||||
setExcludedWords: StatsOkResponse;
|
setExcludedWords: StatsOkResponse;
|
||||||
|
duplicateLineCleanup: StatsDuplicateLineCleanupResult;
|
||||||
wordOccurrences: VocabularyOccurrenceEntry[];
|
wordOccurrences: VocabularyOccurrenceEntry[];
|
||||||
sentenceSearch: SentenceSearchResult[];
|
sentenceSearch: SentenceSearchResult[];
|
||||||
kanji: KanjiEntry[];
|
kanji: KanjiEntry[];
|
||||||
@@ -178,6 +187,9 @@ export interface StatsHttpClient {
|
|||||||
getVocabulary: (limit?: number) => Promise<VocabularyEntry[]>;
|
getVocabulary: (limit?: number) => Promise<VocabularyEntry[]>;
|
||||||
getExcludedWords: () => Promise<StatsExcludedWord[]>;
|
getExcludedWords: () => Promise<StatsExcludedWord[]>;
|
||||||
setExcludedWords: (words: StatsExcludedWord[]) => Promise<void>;
|
setExcludedWords: (words: StatsExcludedWord[]) => Promise<void>;
|
||||||
|
cleanupDuplicateLines: (
|
||||||
|
options?: StatsDuplicateLineCleanupRequest,
|
||||||
|
) => Promise<StatsDuplicateLineCleanupResult>;
|
||||||
getWordOccurrences: (
|
getWordOccurrences: (
|
||||||
headword: string,
|
headword: string,
|
||||||
word: string,
|
word: string,
|
||||||
|
|||||||
@@ -82,6 +82,29 @@ export interface StatsExcludedWord {
|
|||||||
reading: string;
|
reading: string;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** One animation burst the duplicate-line cleanup found in the stats database. */
|
||||||
|
export interface StatsDuplicateLineSample {
|
||||||
|
videoId: number;
|
||||||
|
videoTitle: string | null;
|
||||||
|
text: string;
|
||||||
|
/** Events recorded for this run, including the one that is kept. */
|
||||||
|
frames: number;
|
||||||
|
removedLines: number;
|
||||||
|
startMs: number;
|
||||||
|
endMs: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface StatsDuplicateLineCleanupResult {
|
||||||
|
dryRun: boolean;
|
||||||
|
lookbackDays: number | null;
|
||||||
|
scannedLines: number;
|
||||||
|
burstGroups: number;
|
||||||
|
removedLines: number;
|
||||||
|
removedWordOccurrences: number;
|
||||||
|
removedKanjiOccurrences: number;
|
||||||
|
samples: StatsDuplicateLineSample[];
|
||||||
|
}
|
||||||
|
|
||||||
export interface StatsCoverImage {
|
export interface StatsCoverImage {
|
||||||
contentType: string;
|
contentType: string;
|
||||||
dataUrl: string;
|
dataUrl: string;
|
||||||
|
|||||||
@@ -0,0 +1,183 @@
|
|||||||
|
import { useCallback, useState } from 'react';
|
||||||
|
import { getStatsClient } from '../../hooks/useStatsApi';
|
||||||
|
import { formatNumber } from '../../lib/formatters';
|
||||||
|
import type { StatsDuplicateLineCleanupResult } from '../../types/stats';
|
||||||
|
|
||||||
|
interface DuplicateLineCleanupProps {
|
||||||
|
onClose: () => void;
|
||||||
|
/** Called after rows are actually removed, so the charts can reload. */
|
||||||
|
onCleaned: () => void;
|
||||||
|
}
|
||||||
|
|
||||||
|
const LOOKBACK_CHOICES: Array<{ label: string; days: number | null }> = [
|
||||||
|
{ label: '7 days', days: 7 },
|
||||||
|
{ label: '30 days', days: 30 },
|
||||||
|
{ label: '90 days', days: 90 },
|
||||||
|
{ label: '1 year', days: 365 },
|
||||||
|
{ label: 'All time', days: null },
|
||||||
|
];
|
||||||
|
|
||||||
|
function formatTimecode(ms: number): string {
|
||||||
|
const totalSeconds = Math.max(0, Math.floor(ms / 1000));
|
||||||
|
const minutes = Math.floor(totalSeconds / 60);
|
||||||
|
const seconds = totalSeconds % 60;
|
||||||
|
return `${minutes}:${String(seconds).padStart(2, '0')}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function DuplicateLineCleanup({ onClose, onCleaned }: DuplicateLineCleanupProps) {
|
||||||
|
const [lookbackDays, setLookbackDays] = useState<number | null>(30);
|
||||||
|
const [preview, setPreview] = useState<StatsDuplicateLineCleanupResult | null>(null);
|
||||||
|
const [applied, setApplied] = useState<StatsDuplicateLineCleanupResult | null>(null);
|
||||||
|
const [busy, setBusy] = useState<'scan' | 'apply' | null>(null);
|
||||||
|
const [error, setError] = useState<string | null>(null);
|
||||||
|
|
||||||
|
const run = useCallback(
|
||||||
|
async (dryRun: boolean) => {
|
||||||
|
setBusy(dryRun ? 'scan' : 'apply');
|
||||||
|
setError(null);
|
||||||
|
try {
|
||||||
|
const result = await getStatsClient().cleanupDuplicateLines({ dryRun, lookbackDays });
|
||||||
|
if (dryRun) {
|
||||||
|
setPreview(result);
|
||||||
|
setApplied(null);
|
||||||
|
} else {
|
||||||
|
setApplied(result);
|
||||||
|
setPreview(null);
|
||||||
|
onCleaned();
|
||||||
|
}
|
||||||
|
} catch (cause) {
|
||||||
|
setError(cause instanceof Error ? cause.message : String(cause));
|
||||||
|
} finally {
|
||||||
|
setBusy(null);
|
||||||
|
}
|
||||||
|
},
|
||||||
|
[lookbackDays, onCleaned],
|
||||||
|
);
|
||||||
|
|
||||||
|
const result = applied ?? preview;
|
||||||
|
const nothingToDo = preview !== null && preview.removedLines === 0;
|
||||||
|
|
||||||
|
return (
|
||||||
|
<div className="fixed inset-0 z-50">
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
aria-label="Close duplicate line cleanup"
|
||||||
|
className="absolute inset-0 bg-ctp-crust/70 backdrop-blur-[2px]"
|
||||||
|
onClick={onClose}
|
||||||
|
/>
|
||||||
|
<div className="absolute inset-x-0 top-1/2 mx-auto max-w-xl -translate-y-1/2 rounded-xl border border-ctp-surface1 bg-ctp-mantle shadow-2xl">
|
||||||
|
<div className="flex items-center justify-between border-b border-ctp-surface1 px-5 py-4">
|
||||||
|
<h2 className="text-sm font-semibold text-ctp-text">Duplicate Lines</h2>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
className="rounded-md border border-ctp-surface2 px-3 py-1.5 text-xs font-medium text-ctp-subtext0 transition hover:border-ctp-blue hover:text-ctp-blue"
|
||||||
|
onClick={onClose}
|
||||||
|
>
|
||||||
|
Close
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div className="space-y-4 px-5 py-4">
|
||||||
|
<p className="text-xs leading-relaxed text-ctp-subtext0">
|
||||||
|
Typeset subtitles — karaoke openings, animated signs — are authored as one event per
|
||||||
|
animation frame, and older versions counted every frame as its own line. This finds
|
||||||
|
those runs and collapses each one back to a single line, giving back the word and kanji
|
||||||
|
counts they inflated. Ordinary repeated dialogue is left alone.
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<div>
|
||||||
|
<div className="mb-2 text-xs font-medium text-ctp-subtext1">Look back over</div>
|
||||||
|
<div className="flex flex-wrap gap-2">
|
||||||
|
{LOOKBACK_CHOICES.map((choice) => (
|
||||||
|
<button
|
||||||
|
key={choice.label}
|
||||||
|
type="button"
|
||||||
|
disabled={busy !== null}
|
||||||
|
onClick={() => {
|
||||||
|
setLookbackDays(choice.days);
|
||||||
|
setPreview(null);
|
||||||
|
setApplied(null);
|
||||||
|
}}
|
||||||
|
className={`rounded-lg border px-3 py-1.5 text-xs transition disabled:opacity-50 ${
|
||||||
|
lookbackDays === choice.days
|
||||||
|
? 'border-ctp-blue/50 bg-ctp-surface2 text-ctp-text'
|
||||||
|
: 'border-ctp-surface1 bg-ctp-surface0 text-ctp-overlay2 hover:text-ctp-subtext0'
|
||||||
|
}`}
|
||||||
|
>
|
||||||
|
{choice.label}
|
||||||
|
</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{error && (
|
||||||
|
<div className="rounded-lg border border-ctp-red/30 bg-ctp-red/10 px-3 py-2 text-xs text-ctp-red">
|
||||||
|
{error}
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
{result && (
|
||||||
|
<div className="rounded-lg bg-ctp-surface0 px-4 py-3">
|
||||||
|
<div className="text-sm text-ctp-text">
|
||||||
|
{applied
|
||||||
|
? `Removed ${formatNumber(applied.removedLines)} repeated lines`
|
||||||
|
: nothingToDo
|
||||||
|
? 'No animation bursts found in this window'
|
||||||
|
: `Found ${formatNumber(preview!.burstGroups)} bursts covering ${formatNumber(preview!.removedLines)} extra lines`}
|
||||||
|
</div>
|
||||||
|
<div className="mt-1 text-xs text-ctp-overlay2">
|
||||||
|
{formatNumber(result.scannedLines)} lines scanned ·{' '}
|
||||||
|
{formatNumber(result.removedWordOccurrences)} word counts ·{' '}
|
||||||
|
{formatNumber(result.removedKanjiOccurrences)} kanji counts
|
||||||
|
{applied ? ' removed' : ' would be removed'}
|
||||||
|
</div>
|
||||||
|
|
||||||
|
{result.samples.length > 0 && (
|
||||||
|
<div className="mt-3 max-h-52 space-y-1.5 overflow-y-auto">
|
||||||
|
{result.samples.map((sample) => (
|
||||||
|
<div
|
||||||
|
key={`${sample.videoId}:${sample.startMs}:${sample.text}`}
|
||||||
|
className="flex items-center justify-between gap-3 rounded-md bg-ctp-mantle px-3 py-1.5"
|
||||||
|
>
|
||||||
|
<div className="min-w-0">
|
||||||
|
<div className="truncate text-xs text-ctp-text">{sample.text}</div>
|
||||||
|
<div className="truncate text-[11px] text-ctp-overlay1">
|
||||||
|
{sample.videoTitle ?? `Video ${sample.videoId}`} ·{' '}
|
||||||
|
{formatTimecode(sample.startMs)}
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<span className="shrink-0 text-xs text-ctp-peach">×{sample.frames}</span>
|
||||||
|
</div>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
|
<div className="flex items-center justify-end gap-2">
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
disabled={busy !== null}
|
||||||
|
onClick={() => void run(true)}
|
||||||
|
className="rounded-md border border-ctp-surface2 px-3 py-1.5 text-xs font-medium text-ctp-subtext0 transition hover:border-ctp-blue hover:text-ctp-blue disabled:opacity-50"
|
||||||
|
>
|
||||||
|
{busy === 'scan' ? 'Scanning…' : 'Scan'}
|
||||||
|
</button>
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
disabled={busy !== null || preview === null || nothingToDo}
|
||||||
|
onClick={() => void run(false)}
|
||||||
|
className="rounded-md border border-ctp-red/30 px-3 py-1.5 text-xs font-medium text-ctp-red transition hover:bg-ctp-red/10 disabled:opacity-40"
|
||||||
|
>
|
||||||
|
{busy === 'apply' ? 'Cleaning…' : 'Clean Up'}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
<p className="text-[11px] text-ctp-overlay1">
|
||||||
|
Scan first: cleanup removes rows and cannot be undone. Session watch time and lines-seen
|
||||||
|
totals are left untouched.
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -5,6 +5,7 @@ import { WordList } from './WordList';
|
|||||||
import { KanjiBreakdown } from './KanjiBreakdown';
|
import { KanjiBreakdown } from './KanjiBreakdown';
|
||||||
import { KanjiDetailPanel } from './KanjiDetailPanel';
|
import { KanjiDetailPanel } from './KanjiDetailPanel';
|
||||||
import { ExclusionManager } from './ExclusionManager';
|
import { ExclusionManager } from './ExclusionManager';
|
||||||
|
import { DuplicateLineCleanup } from './DuplicateLineCleanup';
|
||||||
import { formatNumber } from '../../lib/formatters';
|
import { formatNumber } from '../../lib/formatters';
|
||||||
import { TrendChart } from '../trends/TrendChart';
|
import { TrendChart } from '../trends/TrendChart';
|
||||||
import { FrequencyRankTable } from './FrequencyRankTable';
|
import { FrequencyRankTable } from './FrequencyRankTable';
|
||||||
@@ -34,10 +35,11 @@ export function VocabularyTab({
|
|||||||
onRemoveExclusion,
|
onRemoveExclusion,
|
||||||
onClearExclusions,
|
onClearExclusions,
|
||||||
}: VocabularyTabProps) {
|
}: VocabularyTabProps) {
|
||||||
const { words, kanji, knownWords, loading, error } = useVocabulary();
|
const { words, kanji, knownWords, loading, error, reload } = useVocabulary();
|
||||||
const [selectedKanjiId, setSelectedKanjiId] = useState<number | null>(null);
|
const [selectedKanjiId, setSelectedKanjiId] = useState<number | null>(null);
|
||||||
const [hideNames, setHideNames] = useState(false);
|
const [hideNames, setHideNames] = useState(false);
|
||||||
const [showExclusionManager, setShowExclusionManager] = useState(false);
|
const [showExclusionManager, setShowExclusionManager] = useState(false);
|
||||||
|
const [showDuplicateLineCleanup, setShowDuplicateLineCleanup] = useState(false);
|
||||||
|
|
||||||
const hasNames = useMemo(() => words.some(isProperNoun), [words]);
|
const hasNames = useMemo(() => words.some(isProperNoun), [words]);
|
||||||
const filteredWords = useMemo(() => {
|
const filteredWords = useMemo(() => {
|
||||||
@@ -129,6 +131,13 @@ export function VocabularyTab({
|
|||||||
Hide Names
|
Hide Names
|
||||||
</button>
|
</button>
|
||||||
)}
|
)}
|
||||||
|
<button
|
||||||
|
type="button"
|
||||||
|
onClick={() => setShowDuplicateLineCleanup(true)}
|
||||||
|
className="shrink-0 rounded-lg border border-ctp-surface1 bg-ctp-surface0 px-3 py-2 text-xs text-ctp-overlay2 transition-colors hover:text-ctp-subtext0"
|
||||||
|
>
|
||||||
|
Duplicates
|
||||||
|
</button>
|
||||||
<button
|
<button
|
||||||
type="button"
|
type="button"
|
||||||
onClick={() => setShowExclusionManager(true)}
|
onClick={() => setShowExclusionManager(true)}
|
||||||
@@ -193,6 +202,13 @@ export function VocabularyTab({
|
|||||||
onClose={() => setShowExclusionManager(false)}
|
onClose={() => setShowExclusionManager(false)}
|
||||||
/>
|
/>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
|
{showDuplicateLineCleanup && (
|
||||||
|
<DuplicateLineCleanup
|
||||||
|
onClose={() => setShowDuplicateLineCleanup(false)}
|
||||||
|
onCleaned={reload}
|
||||||
|
/>
|
||||||
|
)}
|
||||||
</div>
|
</div>
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
import { useState, useEffect } from 'react';
|
import { useState, useEffect, useCallback } from 'react';
|
||||||
import { getStatsClient } from './useStatsApi';
|
import { getStatsClient } from './useStatsApi';
|
||||||
import type { VocabularyEntry, KanjiEntry } from '../types/stats';
|
import type { VocabularyEntry, KanjiEntry } from '../types/stats';
|
||||||
|
|
||||||
@@ -8,6 +8,9 @@ export function useVocabulary() {
|
|||||||
const [knownWords, setKnownWords] = useState<Set<string>>(new Set());
|
const [knownWords, setKnownWords] = useState<Set<string>>(new Set());
|
||||||
const [loading, setLoading] = useState(true);
|
const [loading, setLoading] = useState(true);
|
||||||
const [error, setError] = useState<string | null>(null);
|
const [error, setError] = useState<string | null>(null);
|
||||||
|
// Bumped by `reload` after maintenance rewrites the vocabulary tables.
|
||||||
|
const [reloadToken, setReloadToken] = useState(0);
|
||||||
|
const reload = useCallback(() => setReloadToken((token) => token + 1), []);
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
let cancelled = false;
|
let cancelled = false;
|
||||||
@@ -46,7 +49,7 @@ export function useVocabulary() {
|
|||||||
return () => {
|
return () => {
|
||||||
cancelled = true;
|
cancelled = true;
|
||||||
};
|
};
|
||||||
}, []);
|
}, [reloadToken]);
|
||||||
|
|
||||||
return { words, kanji, knownWords, loading, error };
|
return { words, kanji, knownWords, loading, error, reload };
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -6,6 +6,8 @@ import type {
|
|||||||
StatsAnkiNotesInfoRequest,
|
StatsAnkiNotesInfoRequest,
|
||||||
StatsCoverImagesRequest,
|
StatsCoverImagesRequest,
|
||||||
StatsDeleteSessionsRequest,
|
StatsDeleteSessionsRequest,
|
||||||
|
StatsDuplicateLineCleanupRequest,
|
||||||
|
StatsDuplicateLineCleanupResult,
|
||||||
StatsExcludedWordsRequest,
|
StatsExcludedWordsRequest,
|
||||||
StatsHttpClient,
|
StatsHttpClient,
|
||||||
StatsJsonResponseMap,
|
StatsJsonResponseMap,
|
||||||
@@ -102,6 +104,19 @@ export const apiClient = {
|
|||||||
body: JSON.stringify({ words } satisfies StatsExcludedWordsRequest),
|
body: JSON.stringify({ words } satisfies StatsExcludedWordsRequest),
|
||||||
});
|
});
|
||||||
},
|
},
|
||||||
|
cleanupDuplicateLines: async (
|
||||||
|
options: StatsDuplicateLineCleanupRequest = {},
|
||||||
|
): Promise<StatsDuplicateLineCleanupResult> => {
|
||||||
|
const res = await fetchResponse('/api/stats/maintenance/duplicate-lines', {
|
||||||
|
method: 'POST',
|
||||||
|
headers: { 'Content-Type': 'application/json' },
|
||||||
|
body: JSON.stringify({
|
||||||
|
dryRun: options.dryRun === true,
|
||||||
|
lookbackDays: options.lookbackDays ?? null,
|
||||||
|
} satisfies StatsDuplicateLineCleanupRequest),
|
||||||
|
});
|
||||||
|
return res.json() as Promise<StatsDuplicateLineCleanupResult>;
|
||||||
|
},
|
||||||
getWordOccurrences: (headword: string, word: string, reading: string, limit = 50, offset = 0) =>
|
getWordOccurrences: (headword: string, word: string, reading: string, limit = 50, offset = 0) =>
|
||||||
fetchJson(
|
fetchJson(
|
||||||
'wordOccurrences',
|
'wordOccurrences',
|
||||||
|
|||||||
Reference in New Issue
Block a user