Making your social video accessible

WCAG puts a hard number on flashing content and on text contrast. Platform docs tell you where auto-captions fall short. Here is the whole checklist, with the thresholds attached.

8 minute read

Accessibility and reach are the same piece of work

The moral argument for accessible video holds on its own, but it is not the one that changes behaviour, so start with the other one: the same work buys you reach. Verizon Media and Publicis Media surveyed 5,616 US adults aged 18 to 54 in April 2019 and found that 69% watch video with the sound off in public, a quarter do it at home, and 80% say they are more likely to watch a video all the way through when captions are available. Set that alongside the World Health Organization's March 2026 fact sheet, which puts 430 million people — over 5% of the world — as needing rehabilitation for disabling hearing loss, and projects 2.5 billion living with some degree of hearing loss by 2050. One decision serves both audiences.

More likely to finish a video with captions80%Watch with sound off in public69%Watch with sound off at home25%
Self-reported viewing behaviour from Verizon Media and Publicis Media's online survey of 5,616 US adults aged 18 to 54, fielded in April 2019. The three bars answer three different questions from the same survey rather than forming a scale, and the data is US-only and several years old — read it as evidence that sound-off viewing is normal, not as a current global figure.

Burned-in or closed captions is a per-platform decision

For short-form the answer is settled and covered elsewhere on this blog: burn the text into the pixels, because feeds autoplay muted and a separate track may never render. What gets missed is that burned-in and closed captions are not rivals everywhere. On a long-form YouTube or Vimeo upload, a caption file is machine-readable — it feeds search, the transcript panel and translation — and a viewer who needs it can resize or restyle it, which text baked into the frame can never do. So split the rule by format rather than by principle. Vertical clip going to TikTok, Reels or Shorts: burn it in. Full-length upload: supply a corrected caption file and leave the control with the viewer. On the long-form piece there is no reason not to do both.

Auto-captions are a first draft, and the platforms admit it

Every major platform generates captions automatically now, and every one of them hedges in its own documentation. YouTube's help pages say machine captions can misrepresent speech through mispronunciation, accents, dialects or background noise, and tell creators to review the output and fix anything transcribed wrongly. TikTok makes its auto-captions editable on the upload screen for the same reason. The words that break are predictably the load-bearing ones: names, brands, product terms, figures. There is a second gap automation does not close at all. A genuine caption track also carries speaker labels and non-speech sound, which is the actual line between captions and subtitles. Speech recognition transcribes talking. It does not tell a deaf viewer which of two people is speaking, or that something just crashed off-screen.

Contrast: 4.5:1 is the number, and video makes it move

WCAG gives you something concrete to aim at. Success criterion 1.4.3, at conformance level AA, asks for a contrast ratio of at least 4.5:1 between text and its background, relaxing to 3:1 for large text — which the standard defines as 18 point, or 14 point bold. Caption type on a phone usually counts as large, so 3:1 is the floor and 4.5:1 is the comfortable target. The complication video adds is that your background is moving. White captions that clear 4.5:1 against a dark jacket fail two seconds later when the speaker steps in front of a window, so checking a single frame proves very little. The reliable fix is to stop depending on the footage at all: sit the text on a solid or semi-opaque plate, or give it a heavy stroke, so the ratio becomes a property of your design.

Flashing: there is a specific published threshold

Flashing is the one item here where getting it wrong can physically harm someone, and it happens to carry the least ambiguous rule. WCAG success criterion 2.3.1, at level A, requires that nothing flashes more than three times in any one-second period, unless the flash stays below the general and red flash thresholds or covers less than roughly a quarter of any 10-degree region of the screen. A general flash, in the standard's terms, is a pair of opposing changes in relative luminance of 10% or more where the darker state sits below 0.80. The Epilepsy Society puts photosensitivity at around 3% of people with epilepsy and names 16 to 25 flashes a second as the highest-risk band. Short-form breaches are rarely deliberate: single-frame white flash cuts stacked on the beat, glitch transitions on every cut, montages running faster than three shots a second, pulsing backgrounds, looping stripe overlays. TikTok warns creators about photosensitive effects, but a warning is a fallback, not a design brief.

Alt text belongs to the stills, not the video

Alt text does not apply to a video file. It applies to everything still that travels alongside it: thumbnails, carousel frames, quote cards, the screenshot you drop into a LinkedIn post. Instagram takes alt text under Advanced Settings before you publish and lets you edit it afterwards from the post menu; its automatic version runs on object recognition and produces something generic, so overwriting it is fifteen seconds well spent. LinkedIn puts an Alt button directly under the image on desktop and mobile. Write what the image is doing rather than cataloguing what is in it — a description of the point the chart makes beats a description of the chart — and skip openers like photo of, because the screen reader has already said it is an image.

Audio description, and the version you can actually do

Audio description is the item most creators skip, usually for good reason. WCAG 1.2.5, at level AA, asks for audio description on prerecorded video: a narration track covering the actions, scene changes and on-screen text that a blind viewer would otherwise miss. Producing one per clip is not realistic solo, and the standard hands you the way out — if the important visual information is already conveyed in the main audio, no separate description is required. That converts a production task into a scripting habit. Read the number on screen out loud. Name the thing you are holding up. Describe the before and the after instead of only cutting between them. For talking-head content, done consistently, that meets the intent without a second audio track.

Make it a pass, not a project

None of this survives a rushed export, which is the real reason accessible clips stay rare. The failures are always in the last mile: a name the transcriber mangled, a caption line sitting across someone's mouth, text that looked fine on a monitor and disappeared on a phone. Put the check before the render rather than after it — keep the transcript and the caption styling editable right up to the burn, then watch each clip once at phone size with the sound off. A clipping tool that produces the transcript, keeps it correctable and lets you restyle before burning in, FrameOS included, turns that into a two-minute pass instead of a re-edit. The accessibility check and the quality check are the same pass, which is the only reason it gets done.

FAQ

Do captions actually increase views?

The most-cited evidence is a Verizon Media and Publicis Media survey of 5,616 US adults from April 2019, in which 80% said they are more likely to watch a video in full when captions are available and 69% reported watching with sound off in public. It is self-reported and now dated, but it lines up with how muted, autoplaying feeds behave.

What contrast ratio should on-screen text use?

WCAG 1.4.3 at level AA asks for 4.5:1 against the background, or 3:1 for large text — 18 point, or 14 point bold. Most caption type on a phone qualifies as large. Because video backgrounds move, put the text on a plate or give it a heavy stroke rather than checking one frame and hoping.

How much flashing is safe in a video?

WCAG 2.3.1 sets the line at no more than three flashes in any one second, unless the flash stays under the general and red flash thresholds or covers less than about a quarter of any 10-degree region of the screen. The Epilepsy Society identifies 16 to 25 flashes per second as the highest-risk range.

Can you add alt text to a video?

No — alt text is for still images. On social video the accessible equivalents are captions for the speech, a written description or transcript in the post, and audio description for visual information the soundtrack does not carry. Do add alt text to the thumbnail, carousel frames and any image posted alongside the clip.

Sources

Related pages