
Subtitles And Silent Viewing
Most social video is watched without sound. Burned-in captions are not an accessibility feature, they are the difference between a video that communicates and one that is scrolled past.
A video that requires audio to make sense is a video that fails for the majority of the people who see it.
Feeds autoplay muted. People watch in offices, on public transport, in bed next to somebody asleep. The sound is an opt-in, and the opt-in happens after the viewer has already decided whether to stay.
Which means the first three seconds have to work silently or the rest of the video does not happen.
Burned-in beats platform captions
Platform auto-captions are better than nothing and worse than the alternative, for three reasons.
They can be turned off, and on some surfaces they default to off.
They are styled by the platform, which means they sit wherever the platform puts them, often behind the interface elements.
They get the words wrong, particularly on product names, trade terms and Australian place names. A caption that says the wrong product name on screen is worse than no caption.
Burned-in captions are part of the file. They always appear, always in the right place, always with the right spelling.
What good captions look like
Short lines, one to three words at a time for fast-cut content, a full short phrase for talking head. Long lines force reading instead of watching.
Positioned clear of the interface. The bottom of the frame is covered by captions, usernames and buttons depending on the platform. Keeping text in the middle band is the only safe answer across surfaces.
High contrast, with a stroke or a plate so the text stays readable over footage that changes brightness.
Synced tightly. Captions that lag behind the audio are more distracting than no captions, and captions that run ahead give away the line before it is said.
Correct. Proper nouns, trademarks and technical terms checked by somebody who knows them. This is the single most common failure in otherwise well-produced content.
The first frame has to carry the claim
Separate from captions, and more important.
A hook that only exists in the audio is invisible. The opening claim goes on screen as text, large, in the first frame, so somebody who sees the video for half a second with no sound still receives it.
That text is not a caption. It is the headline, and it is written as deliberately as an ad headline because that is what it is.
Sound still matters
None of this is an argument for silent video.
The viewers who do turn the sound on are the ones who watch longest and convert best. Audio quality, music choice and the clarity of the voice are what hold them, and a video that was treated as a silent asset will lose them.
The standard is a video that communicates fully without sound and rewards the viewer who adds it.
What this changes upstream
Three things in production.
Framing leaves room for text, which means composing with a safe band rather than filling the frame and adding text over faces later.
Scripts are written to be read as well as heard, which usually means shorter sentences and fewer dependent clauses.
The transcript is produced as part of the edit, because it is the caption source, the content source for written pieces, and the thing that makes a video searchable.
The accessibility point is not separate
Captions make content work for viewers who are deaf or hard of hearing, and they make content work for everyone in a noisy or quiet environment.
Those are the same requirement. Treating it as a compliance item produces captions nobody checks. Treating it as the primary delivery mechanism produces content that works, and the compliance follows.
Written by David Eid. Published .
Read next.
Contact the Ignis Team
Send through your details and we will audit your business before we reply.




