Skip to main content
When a source publishes captions, Refmatter derives a transcript during acquisition and keeps it as a durable derivative next to the poster frame. Manual captions are preferred; automatic captions are used when they are all the source has.

In the reference

source is captions for a human-authored track or automatic_captions for the platform’s speech recognition. null means the source had no usable track; the reference is still complete.

Full text and timings

The transcript object can also be downloaded as JSON through GET /v1/media/{mediaId}/download-url with the mediaId from the summary.

Language choice

English is preferred when several tracks exist; otherwise the first manual track wins, then automatic captions. Per-request language selection is planned.

Retention

Transcripts are Tier A: they stay as long as the reference exists, including after the media bytes expire. A source with no captions gets no transcript; speech recognition on the acquired audio is planned as a fallback.