Guide

Captions vs Subtitles: What's the Difference and When to Use Each

Captions vs subtitles: what separates them, whether the text should be burned into the picture or shipped as a file viewers switch on, and how to pick the right one for every platform.

Updated August 6, 2026

Most people use "captions" and "subtitles" as if they were the same word wearing two hats. They are not. They were invented for different viewers, they carry different information, and on some platforms picking the wrong one is the difference between a video that lands and a video that scrolls past on mute. Once you see what each one is actually for, the choice takes about five seconds.

What subtitles assume

Subtitles were born in foreign-language cinema. They assume the viewer can hear the audio perfectly well but does not speak the language. So subtitles carry only the spoken words, transcribed or translated, and nothing more. No sound effects, no "[door creaks]," no speaker labels. If a doorbell rings mid-scene, subtitles stay silent about it, because they carry what people say, not the sounds around them.

Reach for subtitles when the barrier is language, not hearing: a French film for an English audience, your English tutorial shown to viewers in Brazil, a product demo you want to sell in three markets.

What captions assume: you cannot hear the audio at all

Captions assume the opposite. The viewer cannot hear, either because they are deaf or hard of hearing, or because the sound is off. So captions carry everything the audio was doing: the dialogue, plus non-speech sound written out ("[tense music]," "[crowd cheering]"), plus who is speaking when it is not obvious.

That extra layer is why captions, not subtitles, are the accessibility standard, and why broadcast, education, and most corporate work require them. It is also why they are the right call for social feeds, where most people watch with the sound off. On mute, your video is a silent film, and captions are the only reason it still makes sense.

In the United States, this is usually called "captions" on a spec sheet, sometimes with a "CC" next to it. In the UK and across much of Europe, the word for the same full package is just "subtitles," with no separate term for the deaf-or-hard-of-hearing version. If a client hands you a spec, read what they describe, not just the label they used.

Burned into the picture, or a separate file?

Captions and subtitles each come in two forms, and the form decides how the file behaves everywhere it travels.

Open means burned into the picture. The text is part of the video pixels, always on, impossible to switch off. It travels with the file no matter where it is re-uploaded or downloaded, and you control exactly how it looks.

Closed means the words ship separately, as a file that travels next to the video. That file is usually an SRT, just the lines and their timings written out in plain text, though a platform can also store the words themselves as a track. Either way, the viewer turns it on or off, can often pick a language, and the platform decides some of the styling.

Burned-in wins on social for a mechanical reason rather than a stylistic one. TikTok, Reels, and Shorts autoplay on mute, and burned-in text is part of the picture, so it shows wherever the picture shows. There is no toggle for a viewer to find and no setting for an app to respect or ignore. A separate track that the feed never reads helps no one.

A separate file wins on control and reach. On YouTube or Vimeo, an SRT lets viewers switch captions on, change language, and read them at their own size, and it is what accessibility tools go looking for. It is also plain text, so you can open it and fix a word by hand.

Usually the answer is both, and taking both is not a compromise. Burned in for the muted feed, the file for the platform and for anyone who needs to toggle.

Which one your video needs

Where the video is going decides it:

  • Short-form social (TikTok, Reels, Shorts): burned-in captions, styled and placed high so the app's buttons do not bury them. Most viewers are on mute.
  • YouTube long-form: a separate SRT so viewers can toggle and translate, plus a burned-in version if you want something on screen during the first few seconds a preview or embed plays muted, before anyone taps to unmute.
  • Selling in more than one language: subtitles as selectable closed tracks, one per language, so each viewer picks their own.
  • Accessibility or compliance (broadcast, courses, corporate): closed captions with sound descriptions and speaker labels, not just a transcript of the words.

On request, captions can also carry the sound descriptions a compliance spec asks for, tags like [tense music], because captioning hears the music and the room as well as the words. The speaker labels in that same line are a different matter, worth reading carefully if a spec is going to be checked against your file: captioning turns speech into words, not into names, so the lines come back in order with nobody identified. You can ask for names anyway. What you get back is a guess, pulled from context in the words rather than recognized by voice, so you are the one confirming it is right. Typing the labels in yourself works just as well. On a two-person interview that check is quick, since questions and answers alternate and an SRT is plain text, so the names go straight in. On six people around a table it is a job of its own, worth pricing in before you promise a compliance date, and it gets easier if each voice was recorded on its own track to begin with, so the merge already has the right name on each line.

What makes captions good, not just present

Getting captions onto a video is easy. Getting them right is a craft, and viewers feel the difference even when they cannot name it.

Timing is the whole game. Each line should land on the word being spoken, not half a beat early, not a beat late, and everything else is a distant second. Short cues read. Long ones stall. A cue is the few words that stand on screen at any one moment. One every one to three seconds, six or seven words at most, broken where the phrase breaks rather than mid-clause. Style for the phone, not the desktop: a plain, unfussy typeface, a soft shadow or a translucent background so the words survive bright footage, and on vertical video a placement lifted clear of the app's own text along the bottom. The bottom of the frame gets shared on purpose too. A name bar (a person's name and title sliding in the first time they speak) and a title strip (a heading naming what is happening on screen) both want that same real estate, so say where each one sits before either gets made rather than discovering the collision on playback. Color is a tool, not decoration: one accent color on the word that matters, at most. No rainbow, no bouncing letters. Legibility first.

Getting both from one video

You do not have to run this as two separate jobs. Add subtitles to the video in one request: upload it to Chat Octopus, describe the look you want, and it transcribes the speech with word-level timing, meaning every word carries its own start and finish, and produces the captions for you. Say "bold white captions with a soft shadow, placed high for vertical" in the request and that is what you get: a captioned MP4 with the text burned in, plus a standalone SRT whose timings come from the same pass, even though the two files group the words differently. The open version for your social feed and the closed file for the platform, from one request.

SRT comes back by default because nearly every platform reads it. Two other formats exist for when the destination is picky: VTT, built for players sitting inside a web page, and ASS, built for stylized subtitles with the positioning and color baked into the file itself. Ask for either one and that is the file you get.

Ordinary speech in a quiet room comes back clean. What slips is the vocabulary nobody outside your company uses: brand names, surnames, the product you launched last week. Captioning hears the audio on its own, with none of the context you typed around it, so an unfamiliar name arrives as the closest ordinary word that fits the sound, correctly timed and confidently wrong. A strong accent, or a room with traffic in it, makes the same kind of miss more likely. Put the names' spellings in your first message and they arrive as corrections rather than mistakes you have to go hunting for. Play the result once anyway with your list of names beside you: that check costs one playthrough of your own video, and it is the whole of the quality control.

The rest stays a conversation. Ask for a language you do not have yet ("now make Spanish captions from the English ones") and both versions live in the same thread. Spot a typo? Name the line and what it should say, "fix 'their' to 'there' at 2:14", and the correction lands on that line alone. The video itself gets built again around it, which is machine time rather than yours: you are not hunting along a timeline for the moment or re-cutting anything you had already finished. If you have not worked this way before, here is how a Chat Octopus conversation flows, and if your source is an audio recording rather than a video, it can transcribe that first and caption from there.

Captions or subtitles, open or closed, is not a trivia question. It is a decision about who is watching and how. Get it wrong on the first try and fixing it is still not a redo: say what is off, on which line, and the rebuild is machine time, not yours.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.