Skip to content
By heygen-com

Cinematic captions for video

/embedded-captions

/embedded-captions is a Claude Code skill in the Video & Motion section. It is maintained by heygen-com and published under the Apache-2.0 license. Adds captions to a talking-head video without touching the footage, from quiet verbatim text to words placed behind the speaker.

Author's description

Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it.

Use it when

  • You have a single-speaker video and want captions without editing the footage
  • You want quiet, verbatim, readable subtitles
  • You want a bold visual effect, with words behind the speaker, for a key moment

Not for

  • Multi-shot footage: split it into single shots before applying it
  • Editing or trimming the video: it only adds captions and leaves the picture intact

What you get

A final MP4 with the captions composited over your original footage, which is delivered unmodified, after frame previews and timing and overlap checks.

How to ask for it

  • Add captions to this talking-head video
  • Put cinematic captions behind me in this video
  • I want flashy TikTok-style effect captions on this clip

Install

curl -fsSL https://raw.githubusercontent.com/sgomez-dev/claude-skills/main/install.sh | bash

After installing with the script, type /embedded-captions. Using Cursor, Windsurf or Codex? Platform guides

Questions about this skill

Does it change my original video?
No. The footage is delivered untouched and the captions are the only thing added.
What styles does it offer?
A catalogue of 35 identities. By default it uses a quiet verbatim caption along the lower part of the frame and saves words placed behind the speaker for the peak moment.
Do I need external services?
The workflow, including transcription and subject matting, runs locally. It needs Node, FFmpeg and a few dependencies installed in the project.

Demo coming soon

↑↓ move · Enter open · Esc close