YouTube's sentence-level ASR emits long caption rows with a placeholder
d="3000", while the caption actually displays until the next window
event. Trusting `d` made long lines vanish mid-speech and leave a blank
gap until the next cue.
Rolling auto-caption documents (rows with a="1") now end each cue at the
next event timestamp instead of t + d, matching YouTube's own display
timing. Manual and non-rolling TimedText keep duration-based timing so
real silence gaps are preserved.