# Word-level captions

Canonical URL: https://getvidiyo.app/glossary/word-level-captions
Last updated: 2026-10-08

> Definition: Word-level captions are captions timed to each individual spoken word rather than to whole sentences, so each word can appear, highlight or animate at the exact moment it is said.

Word-level captions follow the speaker's mouth instead of the sentence. Rather than a whole line appearing at once, each word arrives, lights up or pops as it's spoken. This is the caption style behind most Reels, TikToks and Shorts: a few words on screen at a time, with the current word highlighted.

## How do word-level captions work?

Word-level captions need a timestamp for every word, not just every sentence. A speech recognition model produces the transcript, and an alignment step works out each word's start and end time from the audio. The captions are then grouped into short chunks of a few words, and each word's timing drives its highlight or animation. The quality of the alignment decides whether the highlight feels locked to the voice or slightly late.

## What are common word-level caption styles?

The common word-level caption styles are:

| Style | What happens on each word |
|---|---|
| Word highlight | The spoken word changes colour within its group |
| Karaoke pill | A coloured pill slides under the spoken word |
| Pop or slam | Each word pops or slams in as it's said |
| Wipe or typewriter | Words wipe on, or letters type on with a caret |
| Bold stack | Big uppercase words, the spoken one in the accent colour |

## Word-level captions vs sentence captions: what's the difference?

Sentence captions show a full line for its whole duration, like traditional subtitles, while word-level captions reveal or emphasise words one at a time. Sentence captions are calmer and better for long-form and accessibility. Word-level captions add energy and guide attention, which is why short-form creators use them, but they're harder to read in long stretches.

## How do you add word-level captions in Vidiyo?

In Vidiyo, captions are word-anchored by default. Vidiyo transcribes your audio on your Mac with Whisper, using word timing that matches the model, so each caption word is tied to the moment it's spoken and stays in sync when you cut. Pick a style in the Captions panel's Presets tab, or ask in chat ("word-by-word captions, highlight in yellow"). Styles include subtitle, minimal, word highlight, word pop, marker, editorial, karaoke pill, slam, wipe, bold stack, typewriter and rise, where the last six animate per word. Words whose timing the aligner couldn't place confidently are marked as estimated, and you can fix any word or timing in the Transcript tab.

## Frequently asked questions

### How are word-level timestamps created?

A speech recognition model transcribes the audio, then an alignment step estimates when each word starts and ends. Whisper-based tools commonly do this with a timing method applied to the model's attention.

### Are word-by-word captions better for engagement?

They're the dominant style in short-form video because they guide the eye and make silent viewing easy to follow. There's no universal proof they always perform better, so test them against simpler styles for your audience.

### Why are some words in my captions slightly off?

Fast speech, overlapping voices, music and mumbled words make alignment harder. Good tools mark uncertain words so you can check and nudge them by hand.

## Related terms

- [Burned-in captions](https://getvidiyo.app/glossary/burned-in-captions)
- [Whisper transcription](https://getvidiyo.app/glossary/whisper-transcription)
- [Kinetic typography](https://getvidiyo.app/glossary/kinetic-typography)
- [SRT file](https://getvidiyo.app/glossary/srt-file)

## Where Vidiyo does this

- [Auto captions that are transcribed on your Mac](https://getvidiyo.app/features/auto-captions)
- [How to make word-by-word captions](https://getvidiyo.app/blog/word-by-word-captions)
