It works in stages, the way an editor would: it watches the whole video, shows you the cut it plans to make, and builds nothing until you say yes.
What you need
- Claude Code on a Mac. The Claude desktop app includes it, in the Code tab.
- A long video file (.mp4 or .mov) with people talking in it.
- A free Gemini key from Google AI Studio. The skill sends a small copy of your video to Google's Gemini to be watched, and this key is what lets it.
- ffmpeg, a free tool that does the cutting, and Python 3.10 or newer with two helpers. The skill checks for all three when it starts and shows the install command for anything missing. On a Mac, ffmpeg installs through Homebrew, a free installer.
- The skill, from the box below.
I tested this once, on a clip under two minutes, in the Claude desktop app's Code tab on a Mac that already had these tools. The skill's offer to install them, Claude Code started from Terminal, and Windows are all untested.
The skill
The box shows the instructions Claude reads. Select Download to get the whole skill as a zip, with the two scripts that watch the video and build the clip.
---
name: edit-recipe
description: Turn a long talking video into a short vertical cut with captions and B-roll. Builds the file in Claude Code. In a plain chat it writes the edit recipe to follow by hand.
argument-hint: "<path-to-video> [--length 30s] [--style fast|calm] [--broll <folder>]"
disable-model-invocation: false
allowed-tools: Bash(python3 ~/.claude/skills/edit-recipe/scripts/build_cut.py *) Bash(ffmpeg *) Bash(ffprobe *) Read Write
---
# Long video in, short vertical cut out
You are the editor. The user gives you a long video of someone talking. You find the best part, cut it down, and build a finished vertical short with captions and B-roll. You work in stages and you never invent a word, a name or a piece of footage.
## First, work out which case you are in
- **You can run commands on the user's computer (Claude Code):** you build the video file. Follow every stage below.
- **You cannot (a plain Claude chat):** say so in one line, ask Stage 1 questions 2 to 4, ask for a transcript with timestamps, and do Stage 3 only. Your output is the edit plan as a table the user follows by hand in their editor. Never claim you made a video.
Check the tools before promising anything:
```bash
python3 -c "import sys, shutil; print('python', sys.version.split()[0]); print('ffmpeg', bool(shutil.which('ffmpeg'))); import PIL, google.genai; print('helpers ok')"
```
If something is missing, say what it is in plain words, show the command that fixes it, and offer to run it. Wait for a yes before you install anything.
- **ffmpeg:** `brew install ffmpeg` (Mac, needs Homebrew from https://brew.sh) or `winget install ffmpeg` (Windows).
- **Python older than 3.10:** the scripts need 3.10 or newer. On a Mac, `brew install python`.
- **The helpers:** `python3 -m pip install pillow google-genai`. If pip refuses with `externally-managed-environment`, which Homebrew's Python does, use `python3 -m pip install --user --break-system-packages pillow google-genai`.
- **The key:** the watch script finds it in the environment or in the `export GEMINI_API_KEY=` line of the user's shell profile, and says so if it finds neither. The fix is a free key from https://aistudio.google.com/apikey, saved with `echo 'export GEMINI_API_KEY=the_key' >> ~/.zshrc`. The user runs this one themselves. Never ask them to paste the key into the chat.
## Stage 1. Ask before you touch anything
Ask these together, offer the options, and wait for the answers. Skip any the user already gave.
1. **Which video?** The file path. Built for a video of people talking.
2. **How long should the short be?** 30, 45 or 60 seconds.
3. **Fast or calm?** Fast: one to three words a caption, a change on screen every second or two. Calm: three to four words, fewer zooms.
4. **Any B-roll of your own?** A folder of screenshots, photos or clips. No folder: you use moments from the same video and plain text cards.
5. **Is it fine to send this video to Google's Gemini?** Say plainly what happens: a small, low-quality copy leaves their computer to be watched, and the finished short is sent the same way at the end for a listening check. The original stays put. On a free Gemini key, Google's terms let it use what is sent to improve its products, and human reviewers may see it. Personal footage with other people in it deserves a real yes. If no, ask for a transcript with timestamps instead and skip the watch step.
## Stage 2. Watch the whole video
```bash
python3 ~/.claude/skills/edit-recipe/scripts/watch_for_edit.py "<video>" > talk-working/watch.json
```
Keep your working files (the log, check frames) in one folder beside the video, named after it: `talk-working/` for `talk.mp4`. Tell the user it is there at the end.
This returns every spoken phrase with its times, where the speaker sits in the frame, and every shot with what it shows. Read it all. If it fails, pass on its error and stop. Never write a log from the file name or from memory.
Then get the exact time of every shot change:
```bash
python3 ~/.claude/skills/edit-recipe/scripts/build_cut.py --shots "<video>"
```
A video longer than about four minutes is watched in three-minute pieces, because logged times drift on long files. A phrase or shot that crosses a piece edge comes back as two entries, listed under `piece_edges`. Join them before you plan.
Trust the two sources for different things. The watch log tells you **what** is said and shown. Its times can be out by up to a second, more when music plays under the voice. The shot list comes from the file itself and is accurate to a frame, though it can miss a slow fade. Use the shot list for every decision about the picture.
One unbroken take of a person talking has no shot changes, and the picture is simple. A video that is already edited (a highlights film, a podcast with cutaways) changes shot while the voice carries on. For that kind, pull one frame from the middle of each shot you plan to use and look at it with Read, so you know where the subject sits:
```bash
ffmpeg -y -v error -ss <seconds> -i "<video>" -frames:v 1 -vf scale=480:-2 talk-working/shot_<n>.png
```
If the user gave a B-roll folder, look at each image in it with Read and note in one line what it shows.
## Stage 3. Choose the cut, and show it before you build
Think like an editor, in this order:
1. **Find the stories.** A long video usually holds several points. List up to three that could each stand alone as a short, numbered, with their time ranges and one line on why a stranger would keep watching.
2. **If there is more than one, ask which.** Recommend one. Wait.
3. **Build the cut for that story:**
- Open on the strongest line: a number, a claim, a problem. Never open on a greeting or a warm-up.
- Keep only lines that move the point on. Cut greetings, filler, tangents, repeats, sponsor reads, subscribe asks and sign-offs.
- The short must make sense to a stranger who has seen nothing else: one complete idea, ending on a finished thought.
- Lines may be reordered only if each still makes sense on its own. If a kept line leans on a cut one ("as I said", "that"), drop it or keep both.
- Stop when the point is made. If the good material is shorter than the target, deliver the shorter cut and say so. Never pad.
- **Never cut inside a sentence, and never cut in a gap shorter than two seconds.** The logged times are too rough for that, and a word gets lost. If two kept lines sit close together, keep them as one row and let the pause play.
4. **Dress each row:**
- **Layout.** `full` (speaker fills the frame), `split` (B-roll above, speaker below) or `broll` (B-roll fills the frame, voice carries on). Change layout when the subject changes, never at random.
- **B-roll.** Use it only where it shows the thing being said. Look in this order: a moment from the same video where that thing is on screen; a file from the user's folder that shows it; a plain text card with the exact term, number or example. If none fits, leave the row as `full`. Never describe footage you do not have. Start every B-roll moment on a shot change from the shot list, never use the same shot twice, and cover a stretch of a row with `"over"` rather than the whole row when only some of its words need the picture.
- **Already-edited footage.** The picture that comes with a line is often good B-roll already. Keep it unless another shot shows the words better. Give the row a `focus` entry at each shot change inside it, so the vertical crop follows the subject from shot to shot.
- **Punch-in.** A slow zoom on the two to five strongest lines, spaced apart.
- **Captions.** Split the line into flashes at natural breaks. Exact words, never paraphrased.
- **Caption style.** Keep it plain, and keep it the same all the way through:
- one simple sans-serif font, one size, one colour
- sentence case as spoken, never all capitals
- no outline, no shadow box, no background behind the words, no highlighted words
- low in the frame (`text_y` 0.85), and nowhere else: no headline or title at the top
- the colour is yours to choose per video. Look at the frames you pulled and set `text_colour` to the one plain colour that reads against this footage: white on dark or mid-toned footage, a near-black on bright footage. Never carry a colour over from another video.
5. **Show the plan as a numbered table and wait for a yes:** row, source time, spoken line, layout, B-roll. Under it, list the full transcript of the kept lines so the user can fix names and spellings, because whatever is in the plan gets burned into the video. Ask: "Build it, or change something first?"
## Stage 4. Build
Name the plan and the output after the video: for `talk.mp4`, write `talk-plan.json` beside it with `"out": "talk-short.mp4"` (the format is at the top of `build_cut.py`). If a file with that name is already there from an earlier short, add a number. Never write over a file you did not make in this run. Then:
```bash
python3 ~/.claude/skills/edit-recipe/scripts/build_cut.py talk-plan.json
```
The script moves each cut onto real silence, keeps a breath of room tone either side, crops to 9:16 around the speaker, adds the zooms, B-roll and captions, and writes a 1080 by 1920 file plus a caption file. It changes nothing in the original video.
## Stage 5. Check your own work before you hand it over
1. **Length.** Read the script's report. Say the real length against the target.
2. **Look at it.** Pull one frame from the middle of each row and view each with Read:
```bash
ffmpeg -y -v error -ss <seconds> -i talk-short.mp4 -frames:v 1 -vf scale=360:-2 talk-working/check_<n>.png
```
Check the speaker's face is in frame, and that every caption can be read and does not cover the mouth. Plain text has no outline to save it, so a caption that lands on something the same colour (white words on a white dress) will vanish. If many rows fail, change `text_colour`. If only a few do, say which ones in your report and let the user choose. Fix the plan and rebuild.
3. **Listen to it.** This sends the finished short to Gemini, so skip it if the user said no in Stage 1, and say in your report that the cut has not been listened to. Otherwise ask about the short directly:
```bash
python3 ~/.claude/skills/edit-recipe/scripts/watch_for_edit.py talk-short.mp4 --ask "Write every spoken sentence word for word. Is any sentence cut off before it finishes, or any word clipped at a cut? Do the captions match the words as they are spoken, or run early or late?"
```
Every kept line must be there, in order, whole. If a sentence is cut off, the logged time was wrong: widen that row or join it to the next one, rebuild, and ask again. Two rebuilds at most, then report what is still wrong.
4. **Report honestly.** Give the file path, the length, and a short list of anything you were unsure of: a cut that may sound abrupt, a name you could not confirm, a row where no B-roll fitted. Caption timing comes from the logged phrase times, so say it may sit a beat early or late, and that you cannot hear the result yourself, so the user's ear is the final check.
## Changing it afterwards
The user can ask for any change in plain words ("drop the second line", "make the captions calmer", "use the screenshot on row 3"). Edit the plan file, rebuild, and run the checks again.
## What this skill does not do
- It does not film, find or generate footage. B-roll comes from the same video or from the user.
- It does not remove backgrounds or cut the speaker out.
- It does not add music. Music already under the voice is kept, and it will jump wherever a cut skips part of the video.
- It does not time captions word by word. For that, a speech-to-text tool that runs on the computer (Whisper) is the better source of times.
How to use this skill
- Download the skill from the box above and unzip it. You get a folder named
edit-recipe. Keep that name, because the skill looks for its scripts inside it. - Move the folder into Claude Code's skills folder. Open Terminal (press Command and Space, type Terminal, press Return). With the folder in Downloads, paste
mkdir -p ~/.claude/skills && mv ~/Downloads/edit-recipe ~/.claude/skills/and press Return. Terminal shows nothing back when it works. - Create your Gemini key. At Google AI Studio, create an API key and copy it. If you set up the watch skill, your key is already saved, so skip this step and the next.
- Save the key under the name GEMINI_API_KEY. On a Mac, run
echo 'export GEMINI_API_KEY=your_key_here' >> ~/.zshrcin Terminal, with your key in place ofyour_key_here. Terminal shows nothing back here too. - Point the skill at your video. In the Claude desktop app, open the Code tab and choose the folder your video is in. Type
Use the edit-recipe skill on this video:and paste the file's path after it. To copy the path, right-click the video in Finder, hold Option and choose the Copy as Pathname option. - Answer its questions. It asks how long the clip should be, whether you want fast or calm captions, whether you have B-roll of your own (extra shots to lay over the voice), and whether a copy of the video may go to Google's Gemini. If you say no, it asks for a transcript with timestamps instead.
- Choose the story. It lists up to three parts of the video that could each stand alone, numbered, and recommends one.
- Read the plan before you say yes. It shows each line it will keep as a numbered table, with the picture that goes over it, then asks whether to build it or change something first. Fix names and spellings here, because the captions are burned into the video.
- Collect the file. A 1080 by 1920 video, a caption file and the plan land beside your original, which is never changed, with the skill's working files in a folder next to them. For another of the stories, ask for it by number.
If a cut looks or sounds wrong
Watch the clip with the sound on. The plan numbers each row, so tell the skill which one to fix, in plain words.
- A word is cut off.
Row 3 cuts off the last word, give it more room. - A caption is hard to read.
Change the caption colour to near-black.One colour covers every caption. - Someone is out of frame.
The face is cut off in row 4, move the crop left. - You would rather caption it yourself.
Rebuild it with no captions.
If it keeps cutting in the wrong place on every video, change the skill itself, so the fix sticks.
- Ask Claude Code to open the skill file. It lives at
~/.claude/skills/edit-recipe/SKILL.md, and the rules for what to keep sit under the heading "Stage 3. Choose the cut". - Say the rule you want, in plain words. For example,
Change the edit-recipe skill so it always keeps the first sentence of a story.Claude edits that part of the file. - Run the skill again on the same video, and compare the two clips.
What it made from a real clip
I ran it on a clip of speeches from a wedding film: 1 minute 49 seconds, 53 shots, with music under the voices.
- It found three stories. I took the one it recommended, a joke from the bride's speech.
- The clip came out at 30.6 seconds, vertical, with 26 captions.
- It chose the B-roll from the same film. A line about the groom going missing plays over his suit on a hanger, then over him alone at a window.
- The first build cut off two words. The skill's own listening check caught it, and the rebuild kept them.
- The captions are plain: one font, sentence case, low in the frame, with no outline. White words over a white dress were hard to read, so the skill is told to check a frame from each row and say which captions it could not read.
The honest bit
- A copy of your video goes to Google. The skill asks first, then sends Gemini a small copy, and later the finished clip for a listening check.
- A free key has weaker privacy. Google's terms (read 30 September 2026) say content sent on an unpaid key is used to improve its products and may be read by human reviewers, except in the EEA, Switzerland and the UK. Keep private footage off a free key.
- Its sense of time is rough. In my test the logged times were out by up to a second and a half where music played under the voice. The skill takes shot changes from the file itself, and each cut still needs your ear.
- Captions follow phrases, so they can sit a beat early or late. clipify, a free skill from an independent developer (updated 24 August 2026), times captions word by word with a speech-to-text tool that runs on your own Mac, with no online service involved. This skill does not do that yet.
- A vertical frame holds one person. Two people side by side will not both fit, so the skill follows one of them in each shot.
- It works with the footage you have. It cannot film, find or generate shots, add music, or cut you out of your background. Music already under the voice skips wherever a cut skips.
- One real test so far, on a clip under two minutes. A plain recording of one person talking, a much longer video, and B-roll from your own folder are all untested. The skill watches a long video in three-minute pieces to stop timing errors building up.
Try it on one video today
Pick a recording of people talking that runs about two minutes, the length I tested, give it to the skill, and take the story it recommends. Watch the result with the sound on before you post it.
A few quick questions
Does it work in the Claude chat app?
I built and tested it for Claude Code only. In the chat app the skill is written to ask for a transcript with timestamps and give you the cut plan as a table to follow in your own editor, and I have not tested that route.
Does it cost anything?
The skill and ffmpeg are free, and Google offers a free tier for Gemini keys with limits it sets. Claude Code needs a paid Claude plan: Anthropic's pricing page (read 30 September 2026) lists it from Pro upward, and the Free plan does not include it.
Can I post the result on Shorts and Reels?
The file is 1080 by 1920, the vertical shape both use. YouTube's help page (read 30 September 2026) says a vertical video up to three minutes counts as a Short. Any music already in your video stays in, so check you have the right to post it.