# How to make word-by-word captions

Canonical URL: https://getvidiyo.app/blog/word-by-word-captions
Last updated: 2026-10-08
Author: Ajay Pawriya (https://getvidiyo.app/about)
Published: 2026-10-08
Category: Captions

> TL;DR: Word-by-word captions show or highlight each word as it's spoken, so you need a timestamp for every word, not just every sentence. Get word timing from a speech model, group two to four words per caption for short-form, then animate the active word. Vidiyo does this locally and keeps the timing through every cut.

Word-by-word captions are the bouncing, highlighting captions you see on most Reels and Shorts: a short phrase on screen, with the word being spoken lit up, popped or slammed in. They look like a styling trick, but the hard part is timing. You need to know when every single word starts and ends, and you need that timing to survive the edit.

## What are word-by-word captions?

Word-by-word captions reveal or highlight each word at the moment it's spoken. Ordinary subtitles only need a start and end time per line; word-by-word captions need a start and end time per word. That extra timing is what lets a highlight land exactly on "free" instead of drifting across the whole sentence.

There are two common forms:

- **One phrase, active word highlighted.** Three or four words stay on screen and the spoken one changes colour, gets a marker behind it, or has a pill slide under it.
- **Words arriving one by one.** Each word pops, rises, wipes or slams in as it's said, building the phrase on screen.

Both rely on the same data. The glossary entry on [word-level captions](https://getvidiyo.app/glossary/word-level-captions) covers the term itself.

## Where do word-level timestamps come from?

They come from a speech model that aligns each word to the audio. You can't get them by hand at any sensible speed, and a plain transcript doesn't have them. Whisper-based tools produce them: whisper.cpp, for example, has an experimental word-level mode (`-ml 1`) that writes one word per segment.

Vidiyo runs Whisper on your Mac and then aligns each word using Whisper's timing anchors and measured silences, so a word doesn't stretch across the pause after it. When the aligner isn't confident about a word, it keeps an estimated time and flags it, rather than guessing silently. You fix those few by hand in the Transcript tab.

## How many words should be on screen at once?

Two to four words is the sweet spot for short-form, because it reads in one glance and keeps the highlight close to the speech. Fewer words feel punchier but can flicker; more words start to feel like subtitles. Match the group size to the energy of the video.

| Video type | Words per group | Why |
|---|---|---|
| Hype, punchlines, hooks | 1 to 3 | Each word lands as its own beat |
| Talking-head Reels and Shorts | 2 to 4 | Readable in one glance, highlight stays close |
| Explainers and product demos | 3 to 5 | Viewers follow ideas, not single words |
| Long-form, calm content | 5 to 8 | Closer to subtitles, less visual noise |

Groups should also break at punctuation. Vidiyo ends a group early at a full stop, question mark or exclamation mark, so a sentence never shares a caption with the start of the next one. If you're timing by hand, do the same.

Reading speed still matters at small group sizes. The [DCMP Captioning Key](https://dcmp.org/learn/captioningkey/601) suggests a presentation rate of up to 160 words per minute for upper-level material. A fast talker can go well past that, which is one more reason to cut pauses and filler before you caption.

## Which word-by-word styles are there?

Most styles fall into three families: highlight, arrive and stack. Pick one per video and stay with it; mixing styles inside one video reads as noise. These are the animated presets in Vidiyo, with what each one suits.

| Style | What the spoken word does | Suits |
|---|---|---|
| Word highlight | Changes to the highlight colour | Talking heads, the safe default |
| Marker | Sits on a marker-coloured block | Bold, editorial creators |
| Pill karaoke | A pill slides under it | Friendly creator-style videos |
| Word pop | Pops in larger | Upbeat shorts |
| Slam | Heavy uppercase words land one by one | Hype and punchlines, 1 to 3 words |
| Wipe | Wipes on | Clean, modern brands |
| Bold stack | Huge uppercase, spoken word highlighted | High energy, 2 to 3 words |
| Typewriter | Letters type on behind a caret | Tech, code, explanations |
| Rise | Rises in | Calm, premium tone |

For exact sizes, colours and positions for each style, see [caption styles for Reels and Shorts](https://getvidiyo.app/blog/caption-styles-for-reels-and-shorts).

## How do I make word-by-word captions in Vidiyo?

Generate captions, pick an animated preset, set the words per group, and fix any flagged words. Every caption word is tied to the word you spoke, so you can keep cutting after you caption and nothing drifts.

1. **Generate.** In the Captions tab, click **Generate captions**. Transcription runs on your Mac. English works out of the box; other languages use the language pack.
2. **Pick a style.** In **Presets**, choose Word highlight, Pill karaoke, Slam or another animated style. The preview shows the real per-word animation.
3. **Set the group size.** Set **Words / group** in the caption inspector's Typography settings, or ask for "three words per caption". It runs from 1 to 12 words; the default is 4.
4. **Fix words.** Correct names and misheard words in the **Transcript** tab. Corrections attach to the word, so they survive later cuts and restyles.
5. **Place them.** Ask the assistant to "keep the captions off my face and the Reels buttons". It measures the face in the frame and checks the platform overlays, then reports anything still covered.

Or describe the whole thing in one message, if you've connected Claude or ChatGPT: "Add pill karaoke captions, three words at a time, highlight in our brand yellow, outlined so they read over the bright window." The assistant watches the frames to judge whether the captions are legible over the footage, not just whether they fit.

The part you'd notice a week later: cut a sentence from the middle of a captioned video and the captions after it simply move with the speech. A burned-in or SRT-based workflow would make you redo every caption after the cut.

## How do I make word-by-word captions without Vidiyo?

Get word timestamps from Whisper, then build the captions in a traditional editor or burn them in as one-word cues. It works, but it's slow, and every later cut means re-timing everything after it.

1. **Get word timing.** Run whisper.cpp with `-ml 1` and `-osrt`. You get an SRT with one word per cue, each with its own start and end time.
2. **Choose a delivery.**
   - **One word at a time:** burn that SRT in with FFmpeg's subtitles filter. You get single words appearing in sequence, with no highlight inside a phrase.
   - **Phrase with highlight:** in a traditional editor, add one text layer per group and keyframe the colour of each word at its timestamp. Expect it to take a while per minute of video.
3. **Edit last.** Because the timing is baked in, finish the cut before you caption. Any change afterwards shifts everything after it.

The honest summary: word-by-word captions are easy when the editor knows which word each caption belongs to, and tedious when it only knows times.

## What makes word-by-word captions look amateur?

A few habits give it away, and all are easy to avoid:

- **Every style at once.** One preset per video. Let the content change, not the font.
- **Groups that wrap to three lines.** In vertical video the third line usually drops into the app's caption area. Use fewer words per group.
- **Highlight colour that fights the brand.** Use your accent colour, or the default if you don't have one.
- **No outline over busy footage.** A thin dark outline keeps white text readable over a window or a light shirt.
- **Typos in the hook.** The first caption gets the most eyes. Check names and numbers there first.

If you're starting from raw footage, caption after cutting silences and filler, not before. Tighter speech gives tighter captions, and the highlight has less dead air to sit through.

## Honest limits

- Word timing is aligned automatically; a word the aligner isn't sure of is flagged for you to check.
- Vidiyo exports 1080p at 30 fps, and timelines max out at 5 minutes.
- Captions are burned into the exported MP4. Export SRT for a separate file.

## Frequently asked questions

### What are word-by-word captions?

Captions where each word appears, lights up or animates at the moment it's spoken. They need word-level timestamps, unlike ordinary subtitles, which only need a start and end time per line.

### How many words should a word-by-word caption show?

For Reels and Shorts, two to four words per group is a good default. Use one to three for very punchy styles like slam, and five to eight words for calmer long-form subtitles.

### Can I make word-by-word captions with an SRT file?

Only roughly. SRT has no per-word timing inside a line, so the usual trick is one word per cue. That shows words one at a time but can't highlight a word within a visible phrase.

### Do Vidiyo's word-by-word captions stay synced after editing?

Yes. Each caption word is anchored to its transcript word, so trims, splits and deleted pauses re-time the captions automatically, and words you cut drop out.

## Related

- [Auto captions that are transcribed on your Mac](https://getvidiyo.app/features/auto-captions)
- [Word-level captions](https://getvidiyo.app/glossary/word-level-captions)
- [Kinetic typography](https://getvidiyo.app/glossary/kinetic-typography)
- [The best caption styles for Reels and Shorts](https://getvidiyo.app/blog/caption-styles-for-reels-and-shorts)
- [How to add captions to a video on Mac for free](https://getvidiyo.app/blog/add-captions-to-video-mac-free)
- [How do I add captions in Vidiyo?](https://getvidiyo.app/docs/captions)

## Sources

- [whisper.cpp README: word-level timestamp](https://github.com/ggml-org/whisper.cpp) (accessed 2026-10-08)
- [whisper.cpp command-line options](https://github.com/ggml-org/whisper.cpp/blob/master/examples/cli/README.md) (accessed 2026-10-08)
- [DCMP Captioning Key: Presentation rate](https://dcmp.org/learn/captioningkey/601) (accessed 2026-10-08)

Download Vidiyo for Mac (free during the beta): https://getvidiyo.app/download
