# Transcript-based vs frame-aware AI video editing

Canonical URL: https://getvidiyo.app/blog/transcript-based-vs-frame-aware-ai-video-editing
Last updated: 2026-10-08
Author: Ajay Pawriya (https://getvidiyo.app/about)
Published: 2026-10-08
Category: Editing with AI

> TL;DR: Transcript-based AI editing decides cuts from the words alone, which is fast and precise for removing fillers, pauses and retakes. Frame-aware AI editing also looks at rendered frames, so it can catch what text can't show: a caption over your face, a head cropped by a zoom, a black frame. Most talking-head edits need both.

Image: Illustration of a filmstrip of a skateboarder with one frame highlighted under a playhead, captioned AI that sees every frame

Most AI video editors start from the transcript, and for good reason: words are cheap to process and precise to cut on. The question is what happens after the cut. A transcript can tell an AI that you said "um". It can't tell it that the caption it just added is sitting across your mouth.

## What is transcript-based AI video editing?

Transcript-based AI editing decides every cut from the words: the AI reads a word-timed transcript, chooses which words, sentences or takes to remove, and the editor cuts the matching video. It's fast, cheap and precise for speech cleanup, because each word has a start and end time to cut against.

This is the approach behind [transcript-based editing](https://getvidiyo.app/glossary/transcript-based-editing) in general, and it works well for the core of a talking-head edit: fillers, false starts, repeated takes, tangents and long pauses between words. What it can't do is see. Every visual decision (framing, captions, graphics, colour) is made blind, or not made at all.

## What is frame-aware AI video editing?

Frame-aware AI editing adds sight: the AI looks at still frames of the video, ideally the composed result with captions, text and graphics drawn in, and uses them to check and correct its edits. It can confirm that a punch-in kept your face in frame, or that a title doesn't cover a caption.

Language models read images, not video files. Claude, for example, accepts JPEG, PNG, GIF and WebP images ([Anthropic docs](https://platform.claude.com/docs/en/build-with-claude/vision)). So a frame-aware editor has to render the frames it wants checked and send them as pictures. Which frames it picks, and whether they include your captions and graphics, decides how much it actually catches.

## What can each one catch?

Transcript-based AI catches problems in what was said; frame-aware AI catches problems in what's on screen. The overlap is small, which is why a talking-head edit benefits from both. This table lists common problems and which approach can detect each one.

| Problem | Transcript only | Frame-aware |
|---|---|---|
| Filler words, false starts, repeated takes | Yes | Yes, via the transcript |
| Long pauses between words | Yes, from word timings | Yes |
| Caption covering your face | No | Yes |
| Face cropped after a punch-in | No | Yes |
| Title overlapping a caption or on-screen text | No | Yes |
| Black, blank or frozen frames | No | Yes |
| Captions under the Reels or Shorts interface | No | Yes, if it knows the overlay |
| Jump cut with identical framing | Partly (it knows where cuts are) | Yes |
| Colour too dark or washed out | No | Yes |
| Music too loud under the voice | No | No, needs audio measurement |

The last row matters: neither approach can hear. Audio needs its own measurements, such as integrated loudness and the level of voice against music.

## Why don't all AI editors look at frames?

Because images cost far more than words. Anthropic's docs give the formula: an image costs its width divided by 28 times its height divided by 28 in visual tokens. A 960 by 540 frame is about 700 tokens, and a minute of 30 fps video has 1,800 frames. Sending everything would mean over a million tokens a minute.

So no practical editor sends every frame. Frame-aware tools sample: a few frames around each cut, the opening, the ending, moments where text appears. The skill is in choosing frames where edits tend to go wrong, and in pairing them with measurements that don't need a model to look at all, like where the face is in each shot.

## How does Vidiyo combine the two?

Vidiyo's assistant cuts from the transcript and checks with frames and measurements. It reads a word-timed transcript to choose cuts, then looks at up to six rendered frames at a time, or a labelled contact sheet of the risky moments, with faces, caption boxes, on-screen text and platform overlays outlined.

The measured part runs after every edit without the AI having to ask. Vidiyo checks for captions over faces, cropped faces, titles over captions, frames with no picture, dead air, a slow opening and raw [jump cuts](https://getvidiyo.app/glossary/jump-cut), plus loudness against a -14 LUFS target. The assistant has to fix errors before it calls the edit done. Caption placement uses the same face measurements, so [captions](https://getvidiyo.app/features/auto-captions) land clear of your face by default.

This is the practical meaning of "the AI watches the frames": not that it views every frame like a person, but that it checks the composed result it made, at the moments most likely to be wrong.

## How do you get frame checks without a frame-aware editor?

You do them yourself, at the moments a transcript-based edit tends to break. If your AI tool only reads text, a short visual pass before export catches most of what it misses. Scrub to each of these and look at the frame for two seconds:

1. The very first frame. No blink, black or half transition.
2. Just after every cut. Is the framing identical on both sides? Then it reads as a jump.
3. Every caption change. Is any caption over your mouth or under the app interface on a phone?
4. Every zoom or punch-in. Is your forehead or chin cut off?
5. Every title or graphic entrance. Does it overlap captions?
6. The last frame. Does it end cleanly on the payoff?

You can also export stills at those times and ask any AI that reads images to review them. It's slower than a built-in check, but it closes most of the gap.

## Which one should you use?

Use transcript-based editing when the job is speech cleanup and you'll handle visuals yourself. Use a frame-aware editor when you want the AI to own the whole edit, including captions, zooms and graphics, because those are exactly the parts a transcript can't verify. For a talking head that's going on social, that's usually the case.

Some examples make the split clearer:

- **A podcast clip you'll finish in another editor.** Transcript-based is enough. You only need the words cleaned up.
- **A vertical short with captions and punch-ins.** Frame-aware. Captions over your face and cropped zooms are the most common mistakes, and only frames reveal them.
- **A product demo with screen recordings.** Frame-aware. The AI needs to see whether a zoom or caption hides the button you're talking about.
- **A long interview you're cutting down.** Start transcript-based to find the story, then check the frames around every cut.

If you're comparing tools on this point, our [2026 landscape of AI video editors](https://getvidiyo.app/blog/ai-video-editors-2026-what-they-actually-see) lists what each one says its AI looks at.

## Honest limits

- Frame-aware checks sample frames; they don't watch every frame of the video.
- Neither approach can hear the mix. Vidiyo balances audio from loudness measurements.

## Frequently asked questions

### What is transcript-based video editing?

Transcript-based editing means editing a video by editing its text: delete a word or sentence in the transcript and the matching video is cut. AI editors that work this way decide what to cut by reading the transcript only.

### What does frame-aware mean in AI video editing?

A frame-aware AI editor looks at still frames of the video, ideally the composed result with captions and graphics, not just the words. That lets it check framing, caption placement and visual problems that a transcript can't describe.

### Why doesn't every AI editor look at every frame?

Images are expensive for language models. By Anthropic's published formula, one 960 by 540 frame costs about 700 visual tokens, and a minute of 30 fps video has 1,800 frames. Frame-aware editors sample the frames that matter instead.

### Is Vidiyo transcript-based or frame-aware?

Both. Vidiyo's assistant cuts from a word-timed transcript and also looks at rendered frames and contact sheets, backed by measured face, caption and loudness checks after every edit.

## Related

- [Transcript-based editing](https://getvidiyo.app/glossary/transcript-based-editing)
- [Edit video by chatting](https://getvidiyo.app/features/chat-video-editing)
- [Auto captions that are transcribed on your Mac](https://getvidiyo.app/features/auto-captions)
- [What is agentic video editing?](https://getvidiyo.app/blog/what-is-agentic-video-editing)
- [AI video editors in 2026: what they actually see](https://getvidiyo.app/blog/ai-video-editors-2026-what-they-actually-see)

## Sources

- [Vision (Claude API docs): resolution and token cost](https://platform.claude.com/docs/en/build-with-claude/vision) (accessed 2026-10-08)

Download Vidiyo for Mac (free during the beta): https://getvidiyo.app/download
