Cinematic captions for video
/embedded-captions
/embedded-captions is a Claude Code skill in the Video & Motion section. It is maintained by heygen-com and published under the Apache-2.0 license. Adds captions to a talking-head video without touching the footage, from quiet verbatim text to words placed behind the speaker.
Author's description
Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.
Use it when
- You have a single-speaker video and want captions without editing the footage
- You want quiet, verbatim, readable subtitles
- You want a bold visual effect, with words behind the speaker, for a key moment
Not for
- Multi-shot footage: split it into single shots before applying it
- Editing or trimming the video: it only adds captions and leaves the picture intact
What you get
A final MP4 with the captions composited over your original footage, which is delivered unmodified, after frame previews and timing and overlap checks.
How to ask for it
Add captions to this talking-head videoPut cinematic captions behind me in this videoI want flashy TikTok-style effect captions on this clip
Install
curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bashAfter installing with the script, type /embedded-captions. Using Cursor, Windsurf or Codex? Platform guides
Questions about this skill
- Does it change my original video?
- No. The footage is delivered untouched and the captions are the only thing added.
- What styles does it offer?
- A catalogue of 35 identities. By default it uses a quiet verbatim caption along the lower part of the frame and saves words placed behind the speaker for the peak moment.
- Do I need external services?
- The workflow, including transcription and subject matting, runs locally. It needs Node, FFmpeg and a few dependencies installed in the project.
Demo coming soon