All creator guides

Captions · Creator guide

Video Captions People Can Actually Read

Readable captions preserve the spoken meaning, arrive with the voice and give viewers enough time to read. Check accuracy, phrase grouping, timing, contrast and placement before adding animation.

By EditorOP · · Sources checked on this date

Use this guide in your next edit.

Ask your preferred assistant for a summary and a creator checklist. These links include the public article title and URL.

If the prompt does not appear, copy it into your chat.

Imagine a microphone tutorial that says, “Don’t put the microphone behind the camera.” One transcription error turns the caption into “Do put the microphone behind the camera.” The font can be beautiful, the timing can be perfect, and the advice is still reversed.

That invented example captures the real editing problem: captions carry your message. Their quality depends on several small decisions that a preset cannot make for every sentence. For Reels, TikToks, and Shorts, review those basics before choosing an animation. A viewer should be able to follow the explanation while still seeing what you are demonstrating.

What should video captions include?

Captions represent the audio that a viewer needs to understand the video. That includes dialogue, speaker identification when necessary, and meaningful sounds. W3C’s explanation of its prerecorded-caption criterion makes this scope explicit. A few highlighted keywords do not provide the same information as a complete caption track. Read W3C’s caption guidance.

For a solo explanation, the speaker may be obvious. For a remote interview, an off-camera reply might need a name. For a repair tutorial, “[motor starts]” could communicate whether the fix worked. Label sounds that explain the action, mood, or response; avoid filling the screen with descriptions of every incidental noise.

Constructed examples: preserve the information carried by audio
Audio momentIncomplete captionMore informative caption
An off-camera producer answers a questionYes, we’re recording.[Producer] Yes, we’re recording.
A timer sounds while the creator looks upNo text[timer beeps]
A speaker rejects a suggested settingUse automatic exposure.Don’t use automatic exposure.

The final row is an accuracy correction. The first two supply information that the image alone may not explain.

Start by proofreading names, numbers, and negations

Automatic transcription is a useful first pass. Give words that change the meaning a deliberate second pass: “can” versus “can’t,” a decimal point, a product name, a price, or a measurement.

YouTube warns that automatic captions can misrepresent speech because of factors including accents, dialects, mispronunciations, and background noise. Its guidance tells creators to review and correct them, including on Shorts. YouTube’s automatic-caption guidance.

Read along with the final audio at normal speed. Do this after removing retakes and changing playback speed, so the text describes the version you will publish. Keep the creator’s phrasing and tone. If the speech says “kind of,” “probably,” or “in my experience,” removing that qualification can make the claim stronger than the speaker intended.

If a sentence remains difficult to caption faithfully, improve the spoken edit or rerecord it. Silently substituting a simpler claim in the captions creates two versions of your advice.

How fast should captions appear?

Measure the amount of text against its actual display time. Two useful measures are:

  • Characters per second (CPS): characters in the caption ÷ seconds visible.
  • Words per minute (WPM): words in the caption ÷ seconds visible × 60.

For the calculations below, spaces and punctuation count as characters; line breaks do not. “Your microphone is too far away.” contains six words and 32 characters.

Original calculation: the same sentence at three display durations
Time visibleCharacters per secondWords per minute
1 second32 CPS360 WPM
2 seconds16 CPS180 WPM
2.5 seconds12.8 CPS144 WPM

These are arithmetic examples, not measured audience results. They explain why the same line can feel different when your editor shortens its duration. Also check individual captions: a comfortable average across the whole reel can hide one dense, rapidly disappearing instruction.

Netflix’s U.S. English timed-text guide sets limits of up to 20 CPS for adult programs and 17 CPS for children’s programs. These are delivery requirements for Netflix content, useful as reference points rather than a proven optimum for social video. Netflix’s English timed-text guide.

Use the numbers to flag captions for review. When a cue is dense, ask whether the viewer also needs to inspect a chart, identify a small button, or watch a hand movement. Give that moment room in the edit. If more time would push captions into unrelated speech, change the cut or narration instead of leaving the old words on screen.

Break lines where the sentence makes sense

Group words into readable phrases. A line break should help someone recognize the idea quickly.

Netflix’s English guide uses a maximum of two lines and recommends keeping related grammatical units together, such as a first and last name or a noun and its adjective. Its 42-character line limit is a delivery rule, not a reason to squeeze 42 characters across a narrow phone frame. Netflix’s line-treatment guidance.

Here is the same sentence with two different breaks:

Constructed line-break example: identical words, different grouping
Awkward breakClearer break
Move the light to
your left, then turn it down.
Move the light to your left,
then turn it down.

Both contain the same words. The second keeps the first instruction together and starts the next instruction on its own line.

For a vertical video, begin with one or two short lines that fit comfortably at the intended text size. Review the phrase as a whole. A fixed rule such as “three words per caption” will split some sentences cleanly and break others at the worst possible place.

Illustration: one instruction, two caption treatments. These constructed layouts show the whole instruction at once and the first of three phrase cues. They are not product screenshots or a performance comparison.

Whole instruction at once

Lighting tutorial

Move the light to your left, turn it down, and keep the background darker than your face.

Several actions compete in one text block.

One phrase cue at a time

Lighting tutorial

Move the light
to your left.

First cue shown. Follow it with “Turn it down,” then “Keep the background darker than your face.”

Word highlighting can sit inside a stable phrase. If your style replaces every word individually, compare it with the phrase version on a phone, especially for instructions containing names, conditions, or several steps. Keep every phrase synchronized with the words actually spoken.

Choose contrast and placement on the moving image

Watch the entire shot when choosing a caption color. A white word that looks clear over a dark shirt may disappear when the creator lifts a white product into the frame.

WCAG’s minimum-contrast criterion uses 4.5:1 for ordinary text and 3:1 for qualifying large text, with exceptions. Those ratios provide a useful design reference, but a video editor’s font-size number alone does not establish how large the text appears to a viewer. W3C’s contrast explanation.

A solid background panel or sufficiently substantial outline can help keep the text distinct as the image changes. Check the highlighted state too. Readability can disappear exactly when a key word changes to your brand color.

Choose placement around the video’s purpose. Keep a face visible during an explanation, a hand visible during a demonstration, and labels visible during a screen recording. W3C’s caption definition also says captions should not obstruct relevant information. W3C’s caption definitions.

Preview the upload with the destination app’s controls visible. Account names, descriptions, navigation, and reaction buttons share the frame. This practical check is more reliable than treating one fixed pixel margin as safe across every interface.

Decide how the viewer will receive the captions

Open captions are part of the video image and cannot be switched off. Closed captions are a separate track that supported players can turn on or off. W3C describes both approaches and notes that some players let viewers customize caption presentation. W3C’s open and closed caption guidance.

Keep a corrected text-and-timing version when your workflow permits it. It is easier to update, translate, or reuse than words that exist only inside an exported MP4. YouTube supports caption-file uploads, while TikTok’s documentation describes creator captions you can edit. YouTube caption uploads and TikTok’s accessibility guidance explain those options.

If you publish a styled export alongside a platform caption track, check playback with that track enabled. Make sure two versions of the same sentence do not cover the demonstration or create an unreadable stack of text.

Run a final caption review before publishing

First, listen and read together. Correct the transcript, speaker changes, and timing. Then watch without sound and ask whether the essential audio information is still available. Finally, inspect the video on a phone with the app interface visible.

Use this checklist for the final export:

  • Names, numbers, units, and negations match the audio.
  • Captions describe the final cut, including speed changes.
  • Necessary speaker labels and meaningful sounds are included.
  • Phrase breaks preserve the sentence’s meaning.
  • Dense captions have enough time in context.
  • Text stays distinct over every shot and highlight state.
  • Faces, actions, labels, and platform controls remain visible.
  • Any separate caption track has been checked alongside the styled video.

The downloadable worksheet includes the microphone sentence at all three illustrative durations, plus blank rows for your own video. Record each flagged cue, its start and end time, the reading-rate calculation, and the correction you plan to make. Its example durations are not recommended targets.

Readable captions give the viewer another way to follow your work. The studies above do not measure Reel completion rates or establish a universal retention increase. Apply their lessons to accuracy and viewing comfort, then evaluate the response using your own video’s audience and context.

Use this on your next video

Download the editable worksheet. Its example rows are illustrations; replace them with observations from your own footage.

Caption review worksheet (CSV)

Sources and how this guide was made

Prepared with AI assistance and checked against the linked primary sources. Editing recommendations and constructed examples are EditorOP’s editorial guidance. They are not a controlled performance study or a promise of more views.

To reference this guide: EditorOP. Video Captions People Can Actually Read. 4 October 2026. Canonical article.