Captions vs Subtitles: What's the Difference and When to Use Each
Captions vs subtitles: what separates them, the open vs closed distinction that changes your workflow, and how to pick the right one for every platform.
Updated August 6, 2026
Most people use "captions" and "subtitles" as if they were the same word wearing two hats. They are not. They were invented for different viewers, they carry different information, and on some platforms picking the wrong one is the difference between a video that lands and a video that scrolls past on mute. The good news: once you see what each one is actually for, the choice takes about five seconds.
What subtitles assume
Subtitles were born in foreign-language cinema. They assume the viewer can hear the audio perfectly well but does not speak the language. So subtitles carry one thing: the spoken words, transcribed or translated. No sound effects, no "[door creaks]," no speaker labels. If a doorbell rings mid-scene, subtitles stay silent about it: they carry what people say, not the sounds around them.
Reach for subtitles when the barrier is language, not hearing: a French film for an English audience, your English tutorial shown to viewers in Brazil, a product demo you want to sell in three markets.
What captions assume: you cannot hear the audio at all
Captions assume the opposite. The viewer cannot hear, either because they are deaf or hard of hearing, or because the sound is off. So captions carry everything the audio was doing: the dialogue, plus non-speech sound written out ("[tense music]," "[crowd cheering]"), plus who is speaking when it is not obvious.
That extra layer is why captions, not subtitles, are the accessibility standard, and why broadcast, education, and most corporate work require them. It is also why they are the right call for social feeds, where the vast majority of people watch with the sound off. On mute, your video is a silent film, and captions are the only reason it still makes sense.
One regional wrinkle worth knowing: in the United States the word is usually "closed captions," while much of the UK and Europe says "subtitles" for the same accessibility track. If a client hands you a spec, read what they describe, not just the label they used.
Burned into the picture, or a separate file?
Captions and subtitles each come in two forms, and the form decides how the file behaves everywhere it travels.
Open means burned into the picture. The text is part of the video pixels, always on, impossible to switch off. It travels with the file no matter where it is re-uploaded or downloaded, and you control exactly how it looks.
Closed means the words ship separately: a file that travels next to the video, an SRT, which is just the lines and their timings in plain text, or a track the platform stores and lets viewers switch on. The viewer turns it on or off, can often pick a language, and the platform decides some of the styling.
Burned-in wins on social for a mechanical reason rather than a stylistic one. TikTok, Reels, and Shorts autoplay on mute, and burned-in text is part of the picture, so it shows wherever the picture shows. There is no toggle for a viewer to find and no setting for an app to respect or ignore. A separate track that the feed never reads helps no one.
A separate file wins on control and reach. On YouTube or Vimeo, an SRT lets viewers switch captions on, change language, and read them at their own size, and it is what accessibility tools go looking for. It is also plain text, so you can open it and fix a word by hand.
Usually the answer is both, and taking both is not a compromise. Burned in for the muted feed, the file for the platform and for anyone who needs to toggle. Two viewing contexts, two files, one video.
Which one your video needs
Where the video is going decides it:
- Short-form social (TikTok, Reels, Shorts): burned-in captions, styled and placed high so the app's buttons do not bury them. Most viewers are on mute.
- YouTube long-form: a separate SRT so viewers can toggle and translate, and optionally a burned-in version for the first few silent-autoplay seconds.
- Selling in more than one language: subtitles as selectable closed tracks, one per language, so each viewer picks their own.
- Accessibility or compliance (broadcast, courses, corporate): closed captions with sound descriptions and speaker labels, not just a transcript of the words.
Read that last line carefully if a spec is going to be checked against your file. The speaker labels it asks for are not something captioning here produces. What listens to your audio measures words and where they land, never who is speaking, so the lines come back in order with nobody's name on any of them. Naming them is your edit. On a two-person interview it is quick, since questions and answers alternate and an SRT is plain text, so the names go straight in. On six people around a table it is a job of its own, and worth pricing in before you promise a compliance date.
What makes captions good, not just present
Getting captions onto a video is easy. Getting them right is a craft, and viewers feel the difference even when they cannot name it:
- Timing is the whole game. Each line should land on the word being spoken, not half a beat early, not a beat late. Everything else is a distant second.
- Short cues read, long ones stall. A cue is the few words that stand on screen at any one moment. One every one to three seconds, six or seven words at most, broken where the phrase breaks rather than mid-clause.
- Style for the phone, not the desktop. A plain, unfussy typeface, a soft shadow or a translucent background so the words survive bright footage, and on vertical video a placement lifted clear of the app's own text along the bottom.
- Share the bottom of the frame on purpose. Name bars and title strips live down there too, so say where each one sits before either gets made rather than discovering the collision on playback.
- Color is a tool, not decoration. One accent color on the word that matters, at most. No rainbow, no bouncing letters. Legibility first.
Getting both from one video
You do not have to run this as two separate jobs. Add subtitles to the video in one request: upload it to Chat Octopus, describe the look you want, and it transcribes the speech with word-level timing and produces the captions for you. Say "bold white captions with a soft shadow, placed high for vertical" and that is what you get, delivered as a captioned MP4 with the text burned in and a standalone SRT whose timings match it line for line. The open version for your social feed and the closed file for the platform, from one request.
SRT comes back by default because nearly everything takes it. Ask for VTT or ASS when the place you are posting to wants one of those two instead, and that is the file you get.
Accuracy has a shape, and knowing the shape saves you a full replay. Ordinary speech in a quiet room comes back clean. What slips is the vocabulary nobody outside your company uses: brand names, surnames, the product you launched last week. The step that listens is handed the audio on its own, with none of the context you typed around it, so an unfamiliar name arrives as the closest ordinary word that fits the sound, correctly timed and confidently wrong. A strong accent or a room with traffic in it widens the same crack. Play the result once with your list of names beside you. That check costs one playthrough of your own video, and it is the whole of the quality control.
The rest stays a conversation. Ask for a language you do not have yet ("now make Spanish captions from the English ones") and both versions live in the same thread. Spot a typo? Name the line and what it should say, "fix 'their' to 'there' at 2:14", and the correction lands on that line alone. The video itself gets built again around it, which is machine time rather than yours: you are not hunting along a timeline for the moment or re-cutting anything you had already finished. If you have not worked this way before, here is how a Chat Octopus conversation flows, and if your source is an audio recording rather than a video, it can transcribe that first and caption from there.
Captions or subtitles, open or closed, is not a trivia question. It is a decision about who is watching and how. Once you know that, the file practically picks itself.