CueSift取幕
Media terminology guide

Captions vs Subtitles vs Transcripts: The Practical Difference

Captions synchronize same-language speech and meaningful sounds with video. Subtitles commonly translate dialogue for another language. A transcript presents the spoken content as a readable document and may omit timing. In Chinese, the word 字幕 often covers both captions and subtitles, so context matters.

Published: August 9, 2026Last verified: August 9, 2026

Short answer

Use captions when the text must represent what a viewer would otherwise hear, including speech, speaker changes, and important sounds. Use subtitles when translating dialogue for another language. Use a transcript when the main goal is reading, searching, quoting, note-taking, or repurposing the spoken content.

These are content roles, not file extensions. The same SRT or VTT container can carry captions or subtitles, while a transcript can be plain TXT or a timestamped document.

Try CueSift with a public link

The difference in one table

OutputTypical contentTimingBest useCommon formats
CaptionsSame-language speech plus meaningful audio informationSynchronizedAccessibility and watching without soundSRT / VTT / SCC
SubtitlesUsually translated dialogue and relevant on-screen textSynchronizedLocalization and foreign-language viewingSRT / VTT / TTML
TranscriptSpoken content arranged for readingOptionalNotes, research, search, and content reuseTXT / DOCX / HTML

Choose by the audience's need

A viewer cannot hear the audio

Provide synchronized captions with dialogue, meaningful sound cues, and speaker identification. Do not assume translated dialogue alone covers accessibility needs.

A viewer speaks another language

Provide translated subtitles that preserve meaning, names, and relevant on-screen text. Review cultural context instead of translating each line in isolation.

A reader needs notes or quotes

Provide a transcript organized for reading. Add timestamps when source checking matters, and verify names and numbers before quoting.

An editor needs reusable timed text

Keep timestamps in SRT or VTT, edit cue boundaries, and export a separate TXT transcript when a reading version is also needed.

One scene, three different outputs

This original example shows why the labels are not interchangeable. The source scene contains two speakers, a door sound, and one Chinese translation target.

Closed captions
00:01.000 --> 00:04.000
MAYA: Did you save the draft?

00:04.100 --> 00:05.000
[door closes]

00:05.100 --> 00:07.500
LEO: Yes, I sent it this morning.
Chinese subtitles
00:01.000 --> 00:04.000
你保存草稿了吗?

00:05.100 --> 00:07.500
保存了,我今天早上发出去了。
Transcript

Maya asks whether Leo saved the draft. After a door closes, Leo says he saved it and sent it that morning.

Original editorial example. It is a terminology demonstration, not a transcript copied from a real video.

What CueSift exports

CueSift can turn an eligible public video into editable timed segments. Export SRT or VTT when those segments need to stay synchronized with video. Export TXT when the main goal is reading, notes, or content reuse.

The file format does not guarantee content quality. Review speaker names, sound cues, translations, punctuation, and timestamps according to the audience you are serving.

Frequently asked questions

Are captions and subtitles the same thing?

They are both timed text tracks, but their content goals differ. Captions usually represent same-language speech plus relevant sounds for accessibility, while subtitles commonly translate dialogue for viewers of another language. Everyday usage can blur the terms.

Does a transcript include timestamps?

Not necessarily. YouTube's glossary defines a transcript as unformatted and untimed verbatim text. Many tools also offer timestamped transcripts, so check the actual output instead of relying only on the label.

Which output is best for accessibility?

Synchronized captions are central for viewers who cannot hear the audio, and they should include meaningful non-speech sounds and speaker identification when needed. Some users also need a descriptive transcript, accessible player controls, or audio description.

Can subtitles be converted into a transcript?

Yes. Remove cue numbers and timecodes, join or regroup the text into readable paragraphs, preserve speaker and sound information that still matters, and proofread the result for repeated or missing fragments.