# Cinematic captions for video (/embedded-captions)

/embedded-captions is a Claude Code skill in the Video & Motion section. It is maintained by heygen-com and published under the Apache-2.0 license. Adds captions to a talking-head video without touching the footage, from quiet verbatim text to words placed behind the speaker.

- Web version: https://skills.sgomez.dev/en/s/embedded-captions
- Section: [Video & Motion](https://skills.sgomez.dev/en/video.md)
- Author: heygen-com
- License: Apache-2.0
- Source: https://github.com/heygen-com/hyperframes/tree/5969c6894ba42f9cc55941503be51af2e96e2b3a/skills/embedded-captions
- Updated 28 Sept 2026

## Use it when

- You have a single-speaker video and want captions without editing the footage
- You want quiet, verbatim, readable subtitles
- You want a bold visual effect, with words behind the speaker, for a key moment

## Not for

- Multi-shot footage: split it into single shots before applying it
- Editing or trimming the video: it only adds captions and leaves the picture intact

## What you get

A final MP4 with the captions composited over your original footage, which is delivered unmodified, after frame previews and timing and overlap checks.

## How to ask for it

- `Add captions to this talking-head video`
- `Put cinematic captions behind me in this video`
- `I want flashy TikTok-style effect captions on this clip`

## Install

macOS · Linux:

```
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bash
```

Windows:

```
irm https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.ps1 | iex
```


## Author's description

Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.

## Questions about this skill

### Does it change my original video?

No. The footage is delivered untouched and the captions are the only thing added.

### What styles does it offer?

A catalogue of 35 identities. By default it uses a quiet verbatim caption along the lower part of the frame and saves words placed behind the speaker for the peak moment.

### Do I need external services?

The workflow, including transcription and subject matting, runs locally. It needs Node, FFmpeg and a few dependencies installed in the project.


- [How we review this](https://skills.sgomez.dev/en/methodology.md)
